Categories

  • 13k Topics
    115k Posts
    Tom ElliottT

    @rpycroft Thank you for the writeup. The elimination you did up front — cables,
    link speed, compression, RAID vs AHCI, all three storage nodes, BIOS, iPXE
    files — saved me a lot of guessing, and the commit you linked is the right one.

    You’re right that it’s the r8169 switch, but not quite in the way it looks.
    The in-kernel driver isn’t worse than the vendor one here. It exposed something
    we’d had broken for nine years without knowing.

    Two kernel options that are on by default upstream have been off in the FOS
    kernel config since 2016:

    # CONFIG_PCIEASPM is not set # CONFIG_PCI_MMCONFIG is not set

    ASPM is PCIe link power management — it lets the link drop into a low power
    state when idle and wake when there’s traffic. Waking up is the expensive part.

    Nobody noticed for nine years because the Realtek vendor drivers we used to
    ship handled ASPM themselves. Our r8168 build had dynamic ASPM compiled in,
    which effectively switched ASPM off any time packets were moving. It didn’t
    matter that the kernel couldn’t manage it, because the driver never asked the
    kernel to.

    r8169 does it the normal way: it asks the PCI core to disable L1 once, when the
    card is probed. That’s where this bites. With CONFIG_PCIEASPM off, the function
    it calls isn’t a function at all — it’s a stub that returns “success” without
    doing anything. So r8169 is told the OS disabled L1, marks ASPM as being under
    OS control, and then goes on to enable ASPM and L1.2 on the card — on a
    kernel with no ASPM code in it that disabled nothing. Not one line in the logs
    to say so.

    That “falling back to CSI” line you included is the second option. Without
    CONFIG_PCI_MMCONFIG the kernel only sees the first 256 bytes of each device’s
    config space, and the registers controlling the L1 sub-states live past that.
    So even a kernel that wanted to fix this couldn’t reach them. Including that
    line is what let me tie the two together, so thank you for pasting it.

    That also explains the two things in your report that looked strangest:

    Why UEFI and not legacy — ASPM is programmed by the firmware, and Dell’s UEFI
    path turns L1 on where the legacy/CSM path leaves it off. Our kernel couldn’t
    change it either way, so whatever the firmware picked stuck. Legacy was never
    actually faster; you were getting a machine where ASPM was already off.

    Why deploys and not captures — a deploy is your client receiving. The link
    goes quiet between bursts from the server, drops into L1.2, and pays the wake
    cost over and over. A capture is the client sending, which keeps the link
    busy so it never gets the chance to sleep. That’s your 15 GB/min upload sitting
    next to a 1.2 GB/min download on the same cable.

    Worth saying before anyone suggests it: adding pcie_aspm=off to the kernel
    arguments does nothing on a FOS kernel. The code registering that argument is
    inside the same #ifdef, so it’s compiled out too. It would look like you tried
    the standard fix and it didn’t help.

    I’ve built you an experimental x64 kernel with both options turned back on:

    https://github.com/FOGProject/fos/releases/tag/EXP_20260805-123232 sha256 6fadc7204889bc76fe943520319044940c24654365a523d5df7adf8327a88d25

    It’s the kernel only — your init is untouched and doesn’t need changing. Back
    up your current /var/www/html/fog/service/ipxe/bzImage, drop this one in its
    place, and deploy to one of the 3070s. Putting the old file back is the whole
    rollback if it doesn’t help.

    If you want to confirm the diagnosis yourself first, it takes about a minute
    and doesn’t need my kernel at all. Boot a 3070 to a shell in debug mode:

    lspci -nn | grep -i ethernet # note the Realtek address, e.g. 02:00.0 lspci -t # note the root port above it, e.g. 00:1c.5 setpci -s 02:00.0 CAP_EXP+10.w setpci -s 00:1c.5 CAP_EXP+10.w

    The bottom two bits of each value are the ASPM setting — 0 is off, 2 is L1,
    3 is L0s+L1. Run it once booted UEFI and once booted legacy. If I have this
    right, UEFI shows a 2 or a 3 and legacy shows a 0.

    And to watch it fix itself, clear both (root port first) and rerun the deploy
    without rebooting:

    setpci -s 00:1c.5 CAP_EXP+10.w=0000 setpci -s 02:00.0 CAP_EXP+10.w=0000

    That’s not a fix you can keep — it’s gone on the next boot — but if the speed
    jumps back to 6-10 GB/min, that confirms it before you swap any files.

    One loose end: could you post the raw output of ethtool -k <interface>? You
    mentioned tx checksumming being off and I want to be sure I’m not waving away a
    second, separate problem. It shouldn’t be able to cause a download-only
    slowdown, but I’d rather look than assume.

    Thanks for the report and for offering to test — I’m taking you up on it.

  • Get the latest news on what's happening.
    184 Topics
    825 Posts
    A

    @Tom-Elliott I really appreciate that you are putting effort into providing more frequent releases, which makes it easier for everyone to deploy new security fixes in time. Keep up the good work!

  • View tutorials or talk about FOG in general.
    2k Topics
    19k Posts
    8

    @Cpasjuste

    Thank you, I appreciate that. And yes, I did spend a lot of time on it.

    I agree that it bypasses the official installer in order to make use of Docker environment variables; however, at this stage, I don’t believe that FOG is going to change so dramatically that it’ll break it often.

    I will add that if this could get adopted as an official docker image, we could work with the FOG team to make sure nothing gets to pushed to the FOG code that would break the container until we can adjust the container as well.

  • Report bugs, request features, or get the latest progress.
    2k Topics
    21k Posts
    L

    Hi everyone, / @george1421

    I’ve been trying to capture the SD card of a Raspberry Pi 4 and I’ve hit a wall. I am booting via a USB stick using the EDK2 UEFI firmware, and I have explicitly set the System Table Selection to <Devicetree> in the UEFI settings.

    I managed to get past the iPXE phase without issues, and the capture task successfully loads the FOS kernel and init. However, the boot process stops right after with the following error (see attached image):
    No network interfaces found, your kernel is most probably missing the correct driver!

    Current Setup & Debug info:

    FOG version: 1.5.10.2149

    Hardware: Raspberry Pi 4 (trying to capture the built-in SD card)

    Boot method: USB stick with UEFI firmware

    Kernel version: 6.18.38

    Init version: 20260726

    (No PCI network hardware detected - obviously, as it’s the SoC’s built-in NIC)

    I think it seems the current experimental ARM kernel is missing the built-in Pi 4 network driver (bcmgenet). Unfortunately, using an external USB-to-LAN adapter is not an option for me, as the UEFI firmware refuses to boot from it in my environment.

    My question is:
    Given that I am booting from a USB stick and making it exactly to this point in the capture task, is there any chance or workaround to successfully capture the Pi 4’s SD card? Is there a custom/specific kernel available that includes the built-in Pi network driver, or any other method I can use to bypass this missing driver issue?

    Any help or guidance would be greatly appreciated!
    20260803_152648.jpg

108

Online

12.7k

Users

17.6k

Topics

156.8k

Posts