• Recent
    • Unsolved
    • Tags
    • Popular
    • Users
    • Groups
    • Search
    • Register
    • Login

    Windows 11 image captured from VM hangs indefinitely at Dell logo on Latitude 3400 (multiple units, multiple NVMe brands) — works fine on HP

    Scheduled Pinned Locked Moved General
    14 Posts 2 Posters 760 Views
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • Tom ElliottT
      Tom Elliott @servicedesk.pianezza
      last edited by

      @servicedesk-pianezza Those are clean tests and they kill my theory. Correcting the record, and then I think your
      own last test moves this a long way.

      The stale-metadata idea is dead. A zeroed disk with a fresh deploy still hangs, and
      mdadm --examine found no superblock either side of the wipe. Drop it.

      I also replayed our GPT restore path locally against a Windows 11 resizable image, onto a
      disk deliberately smaller than the captured one — the same order FOS uses: dd of d1.mbr,
      sgdisk -z, sgdisk -gl, then the filldisk table through sfdisk. The result verifies clean:
      sgdisk -v reports no problems, the protective MBR is a single 0xEE entry spanning the whole
      device, and first/last usable sectors match the target. So a malformed partition table is not
      what we are looking at either.

      Now the part I think you undersold. You reached the FOG iPXE menu and chose “Boot from hard
      disk”, with the deployed drive fitted, and then it hung. That means the firmware finished
      POST, brought up the NIC, ran iPXE and drew a menu — all with that drive present. The
      firmware is not the thing that hangs.
      It hands off to bootmgfw.efi and the hang is after
      that point.

      Which reframes the symptom. “Stuck at the Dell logo” is not the firmware stalling. It is the
      Windows boot chain hanging before anything repaints the screen, so the OEM logo simply stays
      up. F12 being dead is expected there — the firmware gave up the keyboard at handoff. It also
      explains the missing Automatic Repair: Windows’ boot-failure counter is incremented by the
      boot manager, and a hang never reaches the code that does it.

      So the question is now why this image hangs in early Windows boot on a Whiskey Lake Latitude
      and not on a 12th-gen HP. Two things to do, in this order.

      Deploy the same image to the same Dell as Single Disk (Not Resizable). This splits the
      problem in half for the cost of one deploy. Both 2020 reports on this hardware said
      non-resizable worked where resizable did not — see banana123 in topic 14147, “Using Multiple
      Partition Image - Single Disk (Not Resizable) DOES work fine”. If non-resizable boots, the
      fault is in our resize path, it is ours, and I will want d1.minimum.partitions,
      d1.fixed_size_partitions and a full debug-deploy transcript. If it hangs too, the resize
      path is exonerated and it is the image.

      Make Windows tell you where it stops. Boot a Windows installer USB on the hung machine,
      Shift-F10 for a command prompt, find the ESP letter with diskpart, then:

      bcdedit /store X:\EFI\Microsoft\Boot\BCD /set {default} sos on
      bcdedit /store X:\EFI\Microsoft\Boot\BCD /set {globalsettings} bootmenupolicy legacy
      

      sos replaces the logo with the list of boot drivers as they load, so the screen names the
      last thing it got to instead of showing you a logo. bootmenupolicy legacy gives you the F8
      menu, and Safe Mode is itself a useful result.

      Last question, because you have not said it anywhere in the thread: was the image captured
      after sysprep /generalize /oobe /shutdown, or from a VM that had simply been shut down? An
      image captured without generalize carries the source machine’s driver and device state, and
      booting on one chipset but not another is the usual way that shows up.

      Please help us build the FOG community with everyone involved. It's not just about coding - way more we need people to test things, update documentation and most importantly work on uniting the community of people enjoying and working on FOG! Get in contact with me (chat bubble in the top right corner) if you want to join in.

      Web GUI issue? Please check apache error (debian/ubuntu: /var/log/apache2/error.log, centos/fedora/rhel: /var/log/httpd/error_log) and php-fpm log (/var/log/php*-fpm.log)

      Please support FOG if you like it: https://wiki.fogproject.org/wiki/index.php/Support_FOG

      1 Reply Last reply Reply Quote 0
      • S
        servicedesk.pianezza
        last edited by

        Hi all, @Tom-Elliott
        Update, and I want to push back on the VBS theory — I don’t think that’s actually it, despite what looked like supporting evidence.

        New, cleaner data point: with the deployed drive installed, the machine doesn’t just hang at the Dell logo during normal boot — it also hangs during “Preparing to enter BIOS Setup” when pressing F2. That’s firmware enumerating storage before Setup even opens, with no OS/bootloader/VBS involved at that stage. I confirmed this persists even after a full CMOS battery removal (physically pulled the coin cell, left it disconnected several minutes, reseated) — so NVRAM/CMOS state is ruled out too.

        The clean isolation test: same physical NVMe drive, same machine —

        • Drive wiped (sgdisk -Z + dd zero on first/last 10MB) with nothing deployed → boots fine, F2 reaches Setup normally.
        • Drive with the Windows 11 image deployed (via FOG, resizable or not) → both normal boot AND F2/Setup entry hang.
        • Drive with a Windows 10 image deployed via the identical FOG/Hyper-V capture pipeline → boots fine, every time, F2 works.

        So it’s specific to whatever ends up on the drive after a Windows 11 deploy — and since F2/Setup-entry is affected too, this is firmware choking while scanning the disk itself, not a Windows boot-chain issue. I’d chased VBS/HVCI/TPM for a while (disabling the master switch reduced but did not eliminate the intermittent hang — went from failing most of the time to failing ~1 in 6-10 boots, which in hindsight might just be noise/coincidence rather than a real effect from that change).

        Given F2 itself is affected, I think this points back toward something written to the disk outside the partitions FOG restores — back to your original stale-metadata theory, just not RAID/mdadm specifically. Could this be something IMSM/Intel RST related that mdadm --examine doesn’t recognize, or a vendor-specific NVMe log/telemetry region that only Windows 11’s install process (vs. 10’s) writes to during setup? Genuinely unsure what else to check here — happy to run any other read-only diagnostics (nvme-cli, smartctl, anything) against the raw device to compare a “good” (Win10 or wiped) drive against a “bad” (Win11) one, sector-region by sector-region if needed.

        TThanks

        Tom ElliottT 1 Reply Last reply Reply Quote 0
        • Tom ElliottT
          Tom Elliott @servicedesk.pianezza
          last edited by

          @servicedesk-pianezza You are right and I was wrong. F2 hanging settles it. “Preparing to enter BIOS Setup” runs
          before any bootloader, so the firmware is the thing that stalls, and my reading of the iPXE
          test was bad. Ignore that whole post.

          Your Windows 10 control is the best evidence in this thread. Same capture pipeline, same
          FOG, same drive, same machine, and it boots. Together with “resizable and non-resizable both
          hang”, that clears our resize path and it clears FOG’s table writing as a general fault: we
          are laying down the captured disk faithfully, and this particular captured disk upsets this
          particular firmware.

          So the question is narrow now: which bytes? Two things, and the first one you can do from
          your desk.

          Diff the two tables. You have a good drive and a bad drive from the same pipeline. In a
          debug task on each, capture:

          sfdisk -d /dev/nvme0n1
          sgdisk -v /dev/nvme0n1
          gdisk -l /dev/nvme0n1
          

          Post both sets. Partition count, order, types, attributes and the end-of-disk figures are
          the things the firmware reads at enumeration, and a diff of Win10-good against Win11-bad
          names the difference without any guessing.

          Then bisect the disk, in this order. Start from a deployed drive that reproduces the
          hang, and after each step power off, power on, press F2, and note whether Setup opens.

          Destroy the partition table only, leaving every byte of partition data where it lies:

          sgdisk -Z /dev/nvme0n1
          

          If Setup now opens, the firmware is choking on the partition table — layout, types or
          attributes — and not on anything inside the partitions. If it still hangs, the data is the
          problem, so redeploy and zero the start of the ESP, which is the only partition the firmware
          reads:

          dd if=/dev/zero of=/dev/nvme0n1p1 bs=1M count=1
          

          If Setup opens after that, it is the ESP filesystem, and we can bisect it file by file from
          there.

          Three questions while you are at it. Which Windows 11 build is the image, and which build
          was the Windows 10 one? Was the Hyper-V VM Generation 2 with Secure Boot and a vTPM
          attached? And what BIOS version are the 3400s on — Dell’s last for that model is in the 1.3x
          range, and firmware is the component under suspicion now.

          Also worth saying plainly: if this turns out to be the Latitude 3400 firmware choking on a
          partition layout it does not like, there may be nothing for FOG to fix beyond documenting
          it. Two earlier reporters on this hardware ended up rebuilding the image from a clean install
          rather than finding a cause. I would rather find it, and your Win10 control is the first
          thing anyone has produced that makes that realistic.

          Please help us build the FOG community with everyone involved. It's not just about coding - way more we need people to test things, update documentation and most importantly work on uniting the community of people enjoying and working on FOG! Get in contact with me (chat bubble in the top right corner) if you want to join in.

          Web GUI issue? Please check apache error (debian/ubuntu: /var/log/apache2/error.log, centos/fedora/rhel: /var/log/httpd/error_log) and php-fpm log (/var/log/php*-fpm.log)

          Please support FOG if you like it: https://wiki.fogproject.org/wiki/index.php/Support_FOG

          1 Reply Last reply Reply Quote 0
          • S
            servicedesk.pianezza
            last edited by

            Hi @Tom-Elliott
            Here’s the table comparison you asked for:

            Win10 (23H2), SSSTC drive:
            label: gpt, protective MBR
            p1: start=2048 size=204800 EF00 EFI
            p2: start=206848 size=32768 0C01 MSR
            p3: start=239616 size=498573824 (237.7 GiB) 0700 Basic Data
            p4: start=498813440 size=1304576 (637.0 MiB) 2700 WinRE
            Free: 2157 sectors (1.1 MiB), no problems found

            Win11 (25H2), Toshiba drive:
            label: gpt, protective MBR
            p1: start=2048 size=204800 EF00 EFI
            p2: start=206848 size=32768 0C01 MSR
            p3: start=239616 size=498344448 (237.6 GiB) 0700 Basic Data
            p4: start=498584064 size=1533952 (749.0 MiB) 2700 WinRE
            Free: 2157 sectors (1.1 MiB), no problems found

            Type GUIDs, attributes, partition order, and start sectors for p1/p2/p3 are identical between both. The only difference is where the p3/p4 boundary falls — Win11’s WinRE partition is ~112MB larger, which just shifts that one boundary. No extra/hidden partitions, no reordering, no odd attributes on either side, no MBR anomaly (both show “protective” cleanly). I don’t see anything in the table itself that looks like a red flag.

            Given how clean this diff is, I’m not expecting the bisection’s first step (sgdisk -Z, table-only) to change anything, but running it now anyway as instructed. Will report whether F2/Setup opens after that, and if not, move to zeroing the ESP start next.

            Thanks

            1 Reply Last reply Reply Quote 0
            • S
              servicedesk.pianezza
              last edited by

              Hi @Tom-Elliott
              Update on the bisection: re-deploying and re-testing multiple times shows this is NOT deterministic even with what should be an identical table each time. Most attempts hang at F2, but occasionally one boots through — same image, same deploy process, same resulting table.

              So sgdisk -Z (empty table) → 100% reliable, boots/F2 every time.
              Populated table (Win11 deploy) → hangs most of the time, but not always — same exact deploy repeated shows different outcomes.

              Given that, I don’t think this is “a wrong byte value” in the table (that would be reproducible 100% of the time either way). This looks more like the Latitude 3400 firmware having a genuine timing/race issue enumerating a populated multi-partition GPT NVMe at POST — mostly failing, occasionally succeeding, regardless of the specific bytes. That would also explain why the two 2020 forum reports on this same hardware family “fixed” it by wiping on another machine: not because specific stale bytes were the cause, but because an empty/simple table happens to avoid triggering whatever race condition the firmware has with a fully populated one.

              If that’s right, this isn’t something byte-level bisection will resolve — it’s a firmware reliability issue with this NVMe controller + BIOS 1.39.0 combination when handling multi-partition GPT disks, independent of FOG, the OS, or the capture pipeline. I don’t have a way to test that hypothesis further without either a Dell hardware diagnostic tool or a firmware engineer’s tools.

              At this point, given BIOS 1.39.0 is the latest available for this model, I don’t think there’s a software fix on my end. Documenting this for anyone else who hits it on Latitude 3400 with an NVMe-heavy image: expect intermittent F2/boot hangs with populated GPT tables on this platform, and budget for multiple retry attempts as a practical workaround, since a full wipe-and-redeploy doesn’t reliably avoid it either (confirmed just now — it just probabilistically improves odds by luck, not a guaranteed fix).

              Thanks for pushing the bisection methodology — even though it landed on “flaky firmware” rather than a fixable root cause, it at least closes the loop definitively and rules out FOG/the image as the actual fault.

              1 Reply Last reply Reply Quote 0
              • S
                servicedesk.pianezza
                last edited by

                Hi @Tom-Elliott
                Finally have a clean, confirmed answer — and I apologize for the back-and-forth on disk labeling in my last few posts, that was entirely on me getting physical drives and PowerShell disk numbers crossed. This is now verified cleanly across two different physical NVMe drives (SSSTC and Toshiba), each tested in both states:

                FOG-deployed Windows 11 (either drive) → Protective MBR (LBA0) has:

                • Real x86 bootstrap code present (the classic “Invalid partition table” / “Missing operating system” stub)
                • Partition entry: End CHS = FF FF FF (sentinel), Size field = exact sector count (e.g. AF 32 CF 1D)
                  → Hangs at F2/Setup and normal boot, on both drives, reproducibly.

                Same drive, wiped and repartitioned natively via Windows diskpart (clean, convert gpt, create partition efi/msr/primary), then had real Windows boot files copied onto the ESP via robocopy (so it’s not just an empty table — it has genuine bootmgfw.efi, CIPolicies, language resources, etc.) → Protective MBR has:

                • Boot code region entirely zeroed
                • Partition entry: End CHS = FE 7F 99 (specific computed value), Size field = FF FF FF FF (sentinel)
                  → Boots/reaches F2 Setup reliably, every time, tested repeatedly on both drives.

                So the variable isn’t the drive brand, isn’t the GPT header (identical format on both), isn’t the partition table entries (identical GUIDs/types/order) — it’s specifically how the Protective MBR at LBA0 is written. sgdisk/gdisk (what FOG’s restore uses) writes a real bootstrap stub (likely copied forward from the source Windows install during Partclone’s capture, since GPT disks normally don’t need real MBR boot code) combined with an exact size field. diskpart writes it the opposite way: zeroed boot code, sentinel size, specific CHS.

                Given the 3400’s firmware hangs specifically at “Preparing to enter BIOS Setup” (confirmed earlier — this is firmware disk enumeration, before any OS involvement) when it encounters the sgdisk-style Protective MBR, my best guess is a legacy/CSM-compatibility path in this BIOS reads that boot code region and/or the exact-size field and gets stuck — possibly trying to validate or execute the bootstrap stub even though the system is UEFI-only, or choking on an exact size value in a field it expects to be a sentinel.

                Question for you: does FOG’s restore process (or the underlying sgdisk/gdisk call) preserve/write real x86 boot code into the Protective MBR, and is there a flag to zero that region and/or force the sentinel size convention instead? If sgdisk has an option for this (or if it’s something Partclone carries over from the source capture rather than something sgdisk actively writes), that would be great to know — happy to test a patched version if you can point me to where in the FOG scripts this happens.

                This feels like a genuinely actionable, reproducible finding now — thanks for sticking with the bisection methodology, it’s what got us here.

                Thanks

                Tom ElliottT 1 Reply Last reply Reply Quote 0
                • Tom ElliottT
                  Tom Elliott @servicedesk.pianezza
                  last edited by

                  @servicedesk-pianezza Yes, FOG writes that boot code, and I can show you where. Your finding holds up against our
                  source and against a reproduction here.

                  Where it comes from. On capture, saveGRUB() in funcs.sh copies the first 1 MiB of the
                  source disk into d1.mbr with dd — LBA0 included, boot code and all. On deploy,
                  clearPartitionTables() runs sgdisk -Z, which does clear the MBR, and then restoreGRUB()
                  writes d1.mbr straight back over it. The next two commands are sgdisk -z, which destroys
                  the GPT structures only and leaves the MBR bytes alone, and sgdisk -gl, which rewrites only
                  the protective partition entry. So the captured Windows bootstrap survives the whole
                  sequence, and gdisk supplies the entry in its own convention: EndCHS ff ff ff and the exact
                  sector count.

                  I replayed that sequence here on a loop-backed disk with one of our own Windows 11 resizable
                  images. The result matches what you found byte for byte in shape:

                  LBA0  : 33 c0 8e d0 bc 00 7c 8e ...      (Windows MBR stub)
                  0x1BE : 00 00 02 00 ee ff ff ff 01 00 00 00 af 32 cf 1d
                  

                  Why it is not simply a bug. On a BIOS-booted GPT Linux disk that same region is GRUB’s
                  boot.img, and d1.grub.mbr exists precisely so we keep it. We cannot blanket-zero LBA0 on
                  every GPT restore. On a Windows GPT image it is dead weight — Windows only boots GPT through
                  UEFI — so a targeted change is available to us. But we should change the right bytes.

                  Which byte is it? Your diskpart disk differs from ours in three places at once, so we do
                  not yet know which one the firmware chokes on. Three one-liners settle it. Start from a fresh
                  deploy that hangs, run one of them, power off, power on, press F2. Redeploy between tests so
                  each one is measured on its own:

                  # 1 - the bootstrap, nothing else
                  dd if=/dev/zero of=/dev/nvme0n1 bs=1 count=446 conv=notrunc
                  
                  # 2 - the size field -> sentinel
                  printf '\xff\xff\xff\xff' | dd of=/dev/nvme0n1 bs=1 seek=458 conv=notrunc
                  
                  # 3 - EndCHS -> the diskpart value
                  printf '\xfe\xff\xff' | dd of=/dev/nvme0n1 bs=1 seek=451 conv=notrunc
                  

                  Each writes inside LBA0 only and leaves the GPT untouched. Test 1 is the one I expect to
                  matter, on the theory you already stated — a CSM path in that firmware reading or validating
                  a bootstrap it should be ignoring.

                  Name the byte and the fix is small: for a GPT image whose OS is Windows, clear that field
                  during the restore and leave Linux images alone. That also gives the two 2020 reports on this
                  hardware an explanation, which is worth having on its own.

                  Please help us build the FOG community with everyone involved. It's not just about coding - way more we need people to test things, update documentation and most importantly work on uniting the community of people enjoying and working on FOG! Get in contact with me (chat bubble in the top right corner) if you want to join in.

                  Web GUI issue? Please check apache error (debian/ubuntu: /var/log/apache2/error.log, centos/fedora/rhel: /var/log/httpd/error_log) and php-fpm log (/var/log/php*-fpm.log)

                  Please support FOG if you like it: https://wiki.fogproject.org/wiki/index.php/Support_FOG

                  1 Reply Last reply Reply Quote 0
                  • S
                    servicedesk.pianezza
                    last edited by

                    This post is deleted!
                    1 Reply Last reply Reply Quote 0
                    • S
                      servicedesk.pianezza
                      last edited by

                      Hi @Tom-Elliott,
                      follow-up on this thread: I never found the root cause of the intermittent firmware hang on the Latitude 3400 (16 hypotheses ruled out, including the Protective MBR fields we discussed, thanks again for the help with that), but deploying these machines with diskpart + DISM from a WinPE booted via wimboot (through a small FOG plugin) has worked reliably so far, while every other model stays on FOS/Partclone. The attached report covers the investigation, the workaround and a few side findings (WinPE deploys leaving the task queued, Secure Boot/MOK, a vsftpd home-directory pitfall). If anyone has seen the same hang or has ideas about the cause, I’d like to hear it.Report_Latitude3400_FOG_BootIssue_EN.md.txt

                      Thanks for your help

                      Tom ElliottT 1 Reply Last reply Reply Quote 0
                      • Tom ElliottT
                        Tom Elliott @servicedesk.pianezza
                        last edited by

                        @servicedesk-pianezza Follow-up on the findings from your report. Four changes are merged. Three came from your
                        report directly. The fourth came out of triaging it.

                        A failed snapin download now answers an HTTP status. Snapins::stream() set no status
                        on its error paths, so the client received 200 with the error text as the file body, hashed
                        that text, and reported the same wrong SHA-512 on every retry. A storage-node or FTP failure
                        therefore read as a permanently corrupt snapin. The download paths now set the status before
                        any body is written: 404 when the snapin does not exist, 503 when the storage node is
                        unreachable, when the file cannot be read, or when no storage group or node is assigned. No
                        client change is needed. Both clients already treat a non-2xx as a failed transfer, log the
                        reason, and check in without computing a hash. (GH-1817)

                        The installer now ensures the service account’s home directory on every run. This is the
                        root cause of your FTP failure. configureUsers() created /home/fogproject and set its
                        ownership only in the branch that runs when the account does not yet exist. On every later
                        install the account was found, the step printed “Skipped”, and the home directory was never
                        examined. So a home directory that was removed could not be restored by re-installing, which
                        is exactly what you observed. vsftpd answers “cannot change directory” with no home
                        directory, and that breaks every snapin transfer. A new _ensureSvcUserHome() runs for both
                        branches, before anything writes into that directory. A path that exists as a file now fails
                        with that cause named, instead of being retried on every install. (GH-1820)

                        wipefs is now built into FOS. This is larger than the diagnostic you were missing.
                        restoreLVM() in the FOS libraries has always called wipefs -a to clear stale filesystem
                        and RAID signatures before it recreates a physical volume. The binary was never in the image,
                        and busybox has no wipefs applet, so that call has always been “command not found” with its
                        output and exit status discarded. pvcreate -ff forced past the leftover signatures, so
                        nothing appeared broken. A build check now refuses a configuration that calls wipefs
                        without building it. (FOS GH-188)

                        An unchecked “remove file data” box deleted the file anyway. We found this while
                        triaging your report, and it is the one with no undo. On the image and snapin edit pages the
                        checkbox handler returned early when the box came off, leaving the previous state in place.
                        Ticking the box, changing your mind, and deleting still removed the files from the storage
                        node. The server also accepted an explicit andFile=0 as a yes on the single-item path,
                        while the bulk path required the value. Both are corrected. (GH-1819)

                        Three further items from your report are already corrected in 1.6, so they need no change:

                        • Task timestamps are stored in UTC, and the timezone setting became a display setting. Your
                          two-hour offset is the old behavior.
                        • The installer no longer rewrites its vhost file wholesale. It splices its own block between
                          markers and preserves everything outside them. Note the limit: an Alias or RewriteCond
                          placed inside FOG’s markers is still replaced. Put yours outside them and it survives.
                        • The certificate tree is split into zones. The snapin CA path points at the root certificate,
                          and the web CA is a separate file. The crossed symlink you found belongs to the older
                          layout.

                        All four changes are on the development branch and on the release-candidate branch, so they
                        reach you in the next release candidate.

                        On the Latitude 3400 hang itself, nothing has changed. It is not a FOG defect, your winpe3400
                        plugin is the right answer for those machines, and the thread is the best record of it that
                        exists. The one hypothesis still untested is the drive’s TRIM state: Partclone writes only used
                        blocks and never trims, while diskpart format and a full wipe both do. If you ever have a
                        hanging drive to spare, Optimize-Volume -DriveLetter C -ReTrim in the HP, then back in the
                        Dell, then F2, would settle it.

                        Thank you again. Your write-up found three real defects, and one of them could have destroyed
                        an image.

                        Please help us build the FOG community with everyone involved. It's not just about coding - way more we need people to test things, update documentation and most importantly work on uniting the community of people enjoying and working on FOG! Get in contact with me (chat bubble in the top right corner) if you want to join in.

                        Web GUI issue? Please check apache error (debian/ubuntu: /var/log/apache2/error.log, centos/fedora/rhel: /var/log/httpd/error_log) and php-fpm log (/var/log/php*-fpm.log)

                        Please support FOG if you like it: https://wiki.fogproject.org/wiki/index.php/Support_FOG

                        1 Reply Last reply Reply Quote 0
                        • 1 / 1
                        • First post
                          Last post

                        53

                        Online

                        12.8k

                        Users

                        17.7k

                        Topics

                        157.2k

                        Posts
                        Copyright © 2012-2026 FOG Project