nixos/portcullis: Tune NICs, IOMMU and EEE
Router-sized 4096-entry rings on every port with GRO kept across forwarding, the IOMMU in passthrough mode, and EEE pinned off on the I226-V ports as one trigger for their link-drop erratum. The ring and GRO settings are .link files, so they land on a device add event rather than at switch time. Also document the tuning deliberately not done -- coalescing and PCIe ASPM -- and what measuring the NICs' ESP offload found, since the esp4_offload modules are software batching and easy to mistake for it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -405,6 +405,41 @@ Two things follow for configuration:
|
|||||||
AES-GCM. GCM resolves to a single fused accelerated implementation, while the authenticate-only
|
AES-GCM. GCM resolves to a single fused accelerated implementation, while the authenticate-only
|
||||||
path falls back to a generic `authenc(hmac(sha256),ecb(cipher_null))` composition.
|
path falls back to a generic `authenc(hmac(sha256),ecb(cipher_null))` composition.
|
||||||
|
|
||||||
|
#### NIC crypto offload
|
||||||
|
|
||||||
|
No mesh uses it, and the `esp4_offload` / `esp6_offload` modules above are unrelated to it — those
|
||||||
|
are software GSO/GRO batching. Hardware ESP offload is a separate XFRM feature that Libreswan only
|
||||||
|
requests for connections setting `nic-offload=yes`, which the
|
||||||
|
[`l2mesh` module](../nixos/modules/l2mesh.nix) does not.
|
||||||
|
|
||||||
|
[`portcullis`](sites/colony/portcullis.md)'s 82599ES ports advertise `esp-hw-offload` and the
|
||||||
|
offload does work, but not for anything the meshes could use. Measured on the box by installing
|
||||||
|
SAs directly with `ip xfrm` and watching `ixgbe`'s `tx_ipsec` counter:
|
||||||
|
|
||||||
|
| SA | Result |
|
||||||
|
|---|---|
|
||||||
|
| `et10g-0`, transport, AES-GCM-128 | offload active — `mode crypto` against the physical port |
|
||||||
|
| `et10g-0`, transport, AES-GCM-256 | rejected: *"IPsec hw offload only supports keys up to 128 bits with a 32 bit salt"* |
|
||||||
|
| `et10g-0`, **tunnel** mode | rejected: *"Unsupported mode for ipsec offload"* |
|
||||||
|
| `lan-hi` (a VLAN on `et10g-0`) | **accepted with the offload silently dropped** — software crypto |
|
||||||
|
| `et2g5-0` (I226-V) | accepted, offload silently dropped — `igc` has none |
|
||||||
|
|
||||||
|
Two traps are worth knowing. Binding an SA to a device that cannot offload is **not** an error:
|
||||||
|
`xfrm_dev_state_add` returns success having cleared the device, so the SA looks fine and quietly
|
||||||
|
runs in software. And a VLAN interface never offloads — it carries no `xfrmdev_ops`, and
|
||||||
|
`esp-hw-offload` reads `off [fixed]` on it even when its parent supports the feature.
|
||||||
|
|
||||||
|
The second trap is the decisive one here. Even with the offload genuinely active against
|
||||||
|
`et10g-0`, driving traffic through the SA left `tx_ipsec` at zero, because the packets egress
|
||||||
|
`lan-hi` and the kernel only offloads when the SA's device matches the egress device. Every address
|
||||||
|
`portcullis` holds is on a VLAN, so an SA would have to be bound to an untagged physical port to
|
||||||
|
see the hardware at all.
|
||||||
|
|
||||||
|
So adopting it would mean dropping to a 128-bit key to suit one NIC family, keeping the underlay
|
||||||
|
off VLANs, and forgoing `udpEncapsulation` — `xfrm_dev_offload_ok` refuses any SA carrying
|
||||||
|
`encap`. That is not a trade worth making for a mesh that has to run across boxes with no offload
|
||||||
|
at all.
|
||||||
|
|
||||||
#### pcrypt
|
#### pcrypt
|
||||||
|
|
||||||
`pcrypt` parallelises an SA's crypto across cores via padata and does lift the per-SA ceiling. It
|
`pcrypt` parallelises an SA's crypto across cores via padata and does lift the per-SA ceiling. It
|
||||||
|
|||||||
@@ -18,6 +18,11 @@ routing currently done by the [`estuary`](estuary.md) VM.
|
|||||||
| Network | Four Intel I226-V 2.5 GbE ports (`et2g5-0`…`et2g5-3`) and one dual-port Intel 82599ES 10 GbE SFP+ card (`et10g-0`, `et10g-1`) |
|
| Network | Four Intel I226-V 2.5 GbE ports (`et2g5-0`…`et2g5-3`) and one dual-port Intel 82599ES 10 GbE SFP+ card (`et10g-0`, `et10g-1`) |
|
||||||
| Management | JetKVM (HDMI/USB KVM with virtual media) |
|
| Management | JetKVM (HDMI/USB KVM with virtual media) |
|
||||||
|
|
||||||
|
The PCIe layout constrains what the cards can reach. The 82599ES sits behind a gen2 x4 link giving
|
||||||
|
16 Gb/s for **both** its ports together, so one port runs at line rate but the pair is
|
||||||
|
oversubscribed. The NVMe is on a x1 root port, capped near 7.9 Gb/s regardless of the drive. Each
|
||||||
|
I226-V has its own x1 link and is not constrained.
|
||||||
|
|
||||||
## Role
|
## Role
|
||||||
|
|
||||||
Not yet in service. The eventual job is to be the physical edge for the colony site at Nikhef,
|
Not yet in service. The eventual job is to be the physical edge for the colony site at Nikhef,
|
||||||
@@ -57,6 +62,46 @@ home — `fergal` goes to Nikhef with it.
|
|||||||
|
|
||||||
The other SFP+ port, `et10g-1`, is unused.
|
The other SFP+ port, `et10g-1`, is unused.
|
||||||
|
|
||||||
|
### Interface tuning
|
||||||
|
|
||||||
|
Every port takes router-sized 4096-entry rings rather than the driver defaults, matching the other
|
||||||
|
routers here, and enables `GenericReceiveOffloadUDPForwarding` so GRO batching survives forwarding
|
||||||
|
once the box carries UDP-encapsulated traffic. Both are `.link` settings, so they apply on the next
|
||||||
|
device event rather than at switch time — a reboot is the reliable way to land a change to them.
|
||||||
|
|
||||||
|
Interrupt coalescing is deliberately left alone. `igc` reports `rx-usecs` 3 and `ixgbe` reports 1,
|
||||||
|
which are the drivers' markers for dynamic ITR rather than literal microseconds; writing a
|
||||||
|
plausible-looking value there replaces adaptive moderation with a fixed one.
|
||||||
|
|
||||||
|
### I226-V erratum
|
||||||
|
|
||||||
|
The I226-V link-drop erratum is driven by PCIe ASPM, Energy Efficient Ethernet and stale NIC
|
||||||
|
firmware. ASPM, the dominant cause, is off across the whole box for the reason in
|
||||||
|
[Power](#power) below. EEE is held off by a udev rule invoking `ethtool`, as `systemd.link` has no
|
||||||
|
knob for it. `igc` already leaves EEE off on these ports, so the rule pins a driver default rather
|
||||||
|
than correcting one, and keeps it from drifting on a kernel bump. Firmware is the remaining item:
|
||||||
|
the ports report NVM `2.13` (EEPROM version word `0x2013`, which `igc` prints as the `2013` in
|
||||||
|
`ethtool -i`), behind the `2.29`/`2.32` images that circulate. Intel does not publish the I226-V
|
||||||
|
NVM image, so updating means third-party firmware and is best attempted while the box is at home
|
||||||
|
and the JetKVM is attached.
|
||||||
|
|
||||||
|
## Power
|
||||||
|
|
||||||
|
The SoC side is already at its floor and needs no tuning: the package draws around 0.75 W idle with
|
||||||
|
cores in C10 essentially all the time, under `intel_pstate` on the `powersave` governor.
|
||||||
|
|
||||||
|
Platform idle is capped instead, and deliberately left that way. The ACPI FADT declares that the
|
||||||
|
system does not support PCIe ASPM, so the OS defers to firmware, every root port advertises ASPM as
|
||||||
|
unsupported and every endpoint sits with it disabled. Deep package C-states need every PCIe link in
|
||||||
|
L1, so the package never leaves C3. `pcie_aspm=force` is the usual answer and is **not** used here:
|
||||||
|
the 82599ES advertises only L0s with an unlimited exit latency, so no amount of forcing reaches the
|
||||||
|
deep states while that card is fitted, and the only links it would actually change are the four
|
||||||
|
I226-V ones — the exact configuration behind the erratum above. Recovering that power is a firmware
|
||||||
|
question for the mini PC, and only worthwhile once the 82599ES is gone.
|
||||||
|
|
||||||
|
`iommu=pt` puts host devices in passthrough so the forwarding path does not pay DMA translation,
|
||||||
|
while leaving the IOMMU available.
|
||||||
|
|
||||||
## Storage
|
## Storage
|
||||||
|
|
||||||
A single NVMe SSD, following the usual tmpfs-root layout: a 2 GiB ESP at `/boot`, then one LVM PV
|
A single NVMe SSD, following the usual tmpfs-root layout: a 2 GiB ESP at `/boot`, then one LVM PV
|
||||||
|
|||||||
@@ -32,6 +32,14 @@ in
|
|||||||
inherit (lib) mkMerge;
|
inherit (lib) mkMerge;
|
||||||
inherit (lib.my) mkVLAN networkdAssignment;
|
inherit (lib.my) mkVLAN networkdAssignment;
|
||||||
inherit (lib.my.c) networkd;
|
inherit (lib.my.c) networkd;
|
||||||
|
|
||||||
|
# Router-sized rings rather than the driver defaults, and GRO kept across
|
||||||
|
# forwarding so UDP-encapsulated traffic stays batched.
|
||||||
|
nicTuning = {
|
||||||
|
RxBufferSize = 4096;
|
||||||
|
TxBufferSize = 4096;
|
||||||
|
GenericReceiveOffloadUDPForwarding = true;
|
||||||
|
};
|
||||||
in
|
in
|
||||||
{
|
{
|
||||||
hardware = {
|
hardware = {
|
||||||
@@ -43,7 +51,9 @@ in
|
|||||||
|
|
||||||
boot = {
|
boot = {
|
||||||
kernelModules = [ "kvm-intel" ];
|
kernelModules = [ "kvm-intel" ];
|
||||||
kernelParams = [ "intel_iommu=on" ];
|
# Passthrough mode keeps the IOMMU available without paying DMA translation
|
||||||
|
# on the forwarding path.
|
||||||
|
kernelParams = [ "intel_iommu=on" "iommu=pt" ];
|
||||||
initrd = {
|
initrd = {
|
||||||
availableKernelModules = [ "xhci_pci" "nvme" "usb_storage" "usbhid" "sd_mod" "sr_mod" ];
|
availableKernelModules = [ "xhci_pci" "nvme" "usb_storage" "usbhid" "sd_mod" "sr_mod" ];
|
||||||
kernelModules = [ "dm-snapshot" ];
|
kernelModules = [ "dm-snapshot" ];
|
||||||
@@ -69,6 +79,13 @@ in
|
|||||||
|
|
||||||
networking = { inherit domain; };
|
networking = { inherit domain; };
|
||||||
|
|
||||||
|
# The I226-V link-drop erratum is driven by EEE as well as ASPM. The driver already
|
||||||
|
# leaves EEE off, so this pins a default rather than changing one; systemd.link has
|
||||||
|
# no knob for it.
|
||||||
|
services.udev.extraRules = ''
|
||||||
|
ACTION=="add", SUBSYSTEM=="net", DRIVERS=="igc", RUN+="${pkgs.ethtool}/bin/ethtool --set-eee $name eee off"
|
||||||
|
'';
|
||||||
|
|
||||||
environment.systemPackages = with pkgs; [
|
environment.systemPackages = with pkgs; [
|
||||||
pciutils
|
pciutils
|
||||||
usbutils
|
usbutils
|
||||||
@@ -87,28 +104,28 @@ in
|
|||||||
links = {
|
links = {
|
||||||
"10-et2g5-0" = {
|
"10-et2g5-0" = {
|
||||||
matchConfig.PermanentMACAddress = "00:d0:b4:05:ed:48";
|
matchConfig.PermanentMACAddress = "00:d0:b4:05:ed:48";
|
||||||
linkConfig.Name = "et2g5-0";
|
linkConfig = nicTuning // { Name = "et2g5-0"; };
|
||||||
};
|
};
|
||||||
"10-et2g5-1" = {
|
"10-et2g5-1" = {
|
||||||
matchConfig.PermanentMACAddress = "00:d0:b4:05:ed:49";
|
matchConfig.PermanentMACAddress = "00:d0:b4:05:ed:49";
|
||||||
linkConfig.Name = "et2g5-1";
|
linkConfig = nicTuning // { Name = "et2g5-1"; };
|
||||||
};
|
};
|
||||||
"10-et2g5-2" = {
|
"10-et2g5-2" = {
|
||||||
matchConfig.PermanentMACAddress = "00:d0:b4:05:ed:4a";
|
matchConfig.PermanentMACAddress = "00:d0:b4:05:ed:4a";
|
||||||
linkConfig.Name = "et2g5-2";
|
linkConfig = nicTuning // { Name = "et2g5-2"; };
|
||||||
};
|
};
|
||||||
"10-et2g5-3" = {
|
"10-et2g5-3" = {
|
||||||
matchConfig.PermanentMACAddress = "00:d0:b4:05:ed:4b";
|
matchConfig.PermanentMACAddress = "00:d0:b4:05:ed:4b";
|
||||||
linkConfig.Name = "et2g5-3";
|
linkConfig = nicTuning // { Name = "et2g5-3"; };
|
||||||
};
|
};
|
||||||
|
|
||||||
"11-et10g-0" = {
|
"11-et10g-0" = {
|
||||||
matchConfig.PermanentMACAddress = "60:be:b4:2e:9b:a2";
|
matchConfig.PermanentMACAddress = "60:be:b4:2e:9b:a2";
|
||||||
linkConfig.Name = "et10g-0";
|
linkConfig = nicTuning // { Name = "et10g-0"; };
|
||||||
};
|
};
|
||||||
"11-et10g-1" = {
|
"11-et10g-1" = {
|
||||||
matchConfig.PermanentMACAddress = "60:be:b4:2e:9b:a3";
|
matchConfig.PermanentMACAddress = "60:be:b4:2e:9b:a3";
|
||||||
linkConfig.Name = "et10g-1";
|
linkConfig = nicTuning // { Name = "et10g-1"; };
|
||||||
};
|
};
|
||||||
};
|
};
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user