@rpycroft Thank you for the writeup. The elimination you did up front — cables,
link speed, compression, RAID vs AHCI, all three storage nodes, BIOS, iPXE
files — saved me a lot of guessing, and the commit you linked is the right one.

You’re right that it’s the r8169 switch, but not quite in the way it looks.
The in-kernel driver isn’t worse than the vendor one here. It exposed something
we’d had broken for nine years without knowing.

Two kernel options that are on by default upstream have been off in the FOS
kernel config since 2016:

# CONFIG_PCIEASPM is not set # CONFIG_PCI_MMCONFIG is not set

ASPM is PCIe link power management — it lets the link drop into a low power
state when idle and wake when there’s traffic. Waking up is the expensive part.

Nobody noticed for nine years because the Realtek vendor drivers we used to
ship handled ASPM themselves. Our r8168 build had dynamic ASPM compiled in,
which effectively switched ASPM off any time packets were moving. It didn’t
matter that the kernel couldn’t manage it, because the driver never asked the
kernel to.

r8169 does it the normal way: it asks the PCI core to disable L1 once, when the
card is probed. That’s where this bites. With CONFIG_PCIEASPM off, the function
it calls isn’t a function at all — it’s a stub that returns “success” without
doing anything. So r8169 is told the OS disabled L1, marks ASPM as being under
OS control, and then goes on to enable ASPM and L1.2 on the card — on a
kernel with no ASPM code in it that disabled nothing. Not one line in the logs
to say so.

That “falling back to CSI” line you included is the second option. Without
CONFIG_PCI_MMCONFIG the kernel only sees the first 256 bytes of each device’s
config space, and the registers controlling the L1 sub-states live past that.
So even a kernel that wanted to fix this couldn’t reach them. Including that
line is what let me tie the two together, so thank you for pasting it.

That also explains the two things in your report that looked strangest:

Why UEFI and not legacy — ASPM is programmed by the firmware, and Dell’s UEFI
path turns L1 on where the legacy/CSM path leaves it off. Our kernel couldn’t
change it either way, so whatever the firmware picked stuck. Legacy was never
actually faster; you were getting a machine where ASPM was already off.

Why deploys and not captures — a deploy is your client receiving. The link
goes quiet between bursts from the server, drops into L1.2, and pays the wake
cost over and over. A capture is the client sending, which keeps the link
busy so it never gets the chance to sleep. That’s your 15 GB/min upload sitting
next to a 1.2 GB/min download on the same cable.

Worth saying before anyone suggests it: adding pcie_aspm=off to the kernel
arguments does nothing on a FOS kernel. The code registering that argument is
inside the same #ifdef, so it’s compiled out too. It would look like you tried
the standard fix and it didn’t help.

I’ve built you an experimental x64 kernel with both options turned back on:

https://github.com/FOGProject/fos/releases/tag/EXP_20260805-123232 sha256 6fadc7204889bc76fe943520319044940c24654365a523d5df7adf8327a88d25

It’s the kernel only — your init is untouched and doesn’t need changing. Back
up your current /var/www/html/fog/service/ipxe/bzImage, drop this one in its
place, and deploy to one of the 3070s. Putting the old file back is the whole
rollback if it doesn’t help.

If you want to confirm the diagnosis yourself first, it takes about a minute
and doesn’t need my kernel at all. Boot a 3070 to a shell in debug mode:

lspci -nn | grep -i ethernet # note the Realtek address, e.g. 02:00.0 lspci -t # note the root port above it, e.g. 00:1c.5 setpci -s 02:00.0 CAP_EXP+10.w setpci -s 00:1c.5 CAP_EXP+10.w

The bottom two bits of each value are the ASPM setting — 0 is off, 2 is L1,
3 is L0s+L1. Run it once booted UEFI and once booted legacy. If I have this
right, UEFI shows a 2 or a 3 and legacy shows a 0.

And to watch it fix itself, clear both (root port first) and rerun the deploy
without rebooting:

setpci -s 00:1c.5 CAP_EXP+10.w=0000 setpci -s 02:00.0 CAP_EXP+10.w=0000

That’s not a fix you can keep — it’s gone on the next boot — but if the speed
jumps back to 6-10 GB/min, that confirms it before you swap any files.

One loose end: could you post the raw output of ethtool -k <interface>? You
mentioned tx checksumming being off and I want to be sure I’m not waving away a
second, separate problem. It shouldn’t be able to cause a download-only
slowdown, but I’d rather look than assume.

Thanks for the report and for offering to test — I’m taking you up on it.