• Recent
    • Unsolved
    • Tags
    • Popular
    • Users
    • Groups
    • Search
    • Register
    • Login

    Windows 11 image captured from VM hangs indefinitely at Dell logo on Latitude 3400 (multiple units, multiple NVMe brands) — works fine on HP

    Scheduled Pinned Locked Moved General
    15 Posts 2 Posters 780 Views
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • S
      servicedesk.pianezza
      last edited by

      Hi all, @Tom-Elliott
      Update, and I want to push back on the VBS theory — I don’t think that’s actually it, despite what looked like supporting evidence.

      New, cleaner data point: with the deployed drive installed, the machine doesn’t just hang at the Dell logo during normal boot — it also hangs during “Preparing to enter BIOS Setup” when pressing F2. That’s firmware enumerating storage before Setup even opens, with no OS/bootloader/VBS involved at that stage. I confirmed this persists even after a full CMOS battery removal (physically pulled the coin cell, left it disconnected several minutes, reseated) — so NVRAM/CMOS state is ruled out too.

      The clean isolation test: same physical NVMe drive, same machine —

      • Drive wiped (sgdisk -Z + dd zero on first/last 10MB) with nothing deployed → boots fine, F2 reaches Setup normally.
      • Drive with the Windows 11 image deployed (via FOG, resizable or not) → both normal boot AND F2/Setup entry hang.
      • Drive with a Windows 10 image deployed via the identical FOG/Hyper-V capture pipeline → boots fine, every time, F2 works.

      So it’s specific to whatever ends up on the drive after a Windows 11 deploy — and since F2/Setup-entry is affected too, this is firmware choking while scanning the disk itself, not a Windows boot-chain issue. I’d chased VBS/HVCI/TPM for a while (disabling the master switch reduced but did not eliminate the intermittent hang — went from failing most of the time to failing ~1 in 6-10 boots, which in hindsight might just be noise/coincidence rather than a real effect from that change).

      Given F2 itself is affected, I think this points back toward something written to the disk outside the partitions FOG restores — back to your original stale-metadata theory, just not RAID/mdadm specifically. Could this be something IMSM/Intel RST related that mdadm --examine doesn’t recognize, or a vendor-specific NVMe log/telemetry region that only Windows 11’s install process (vs. 10’s) writes to during setup? Genuinely unsure what else to check here — happy to run any other read-only diagnostics (nvme-cli, smartctl, anything) against the raw device to compare a “good” (Win10 or wiped) drive against a “bad” (Win11) one, sector-region by sector-region if needed.

      TThanks

      Tom ElliottT 1 Reply Last reply Reply Quote 0
      • Tom ElliottT
        Tom Elliott @servicedesk.pianezza
        last edited by

        @servicedesk-pianezza You are right and I was wrong. F2 hanging settles it. “Preparing to enter BIOS Setup” runs
        before any bootloader, so the firmware is the thing that stalls, and my reading of the iPXE
        test was bad. Ignore that whole post.

        Your Windows 10 control is the best evidence in this thread. Same capture pipeline, same
        FOG, same drive, same machine, and it boots. Together with “resizable and non-resizable both
        hang”, that clears our resize path and it clears FOG’s table writing as a general fault: we
        are laying down the captured disk faithfully, and this particular captured disk upsets this
        particular firmware.

        So the question is narrow now: which bytes? Two things, and the first one you can do from
        your desk.

        Diff the two tables. You have a good drive and a bad drive from the same pipeline. In a
        debug task on each, capture:

        sfdisk -d /dev/nvme0n1
        sgdisk -v /dev/nvme0n1
        gdisk -l /dev/nvme0n1
        

        Post both sets. Partition count, order, types, attributes and the end-of-disk figures are
        the things the firmware reads at enumeration, and a diff of Win10-good against Win11-bad
        names the difference without any guessing.

        Then bisect the disk, in this order. Start from a deployed drive that reproduces the
        hang, and after each step power off, power on, press F2, and note whether Setup opens.

        Destroy the partition table only, leaving every byte of partition data where it lies:

        sgdisk -Z /dev/nvme0n1
        

        If Setup now opens, the firmware is choking on the partition table — layout, types or
        attributes — and not on anything inside the partitions. If it still hangs, the data is the
        problem, so redeploy and zero the start of the ESP, which is the only partition the firmware
        reads:

        dd if=/dev/zero of=/dev/nvme0n1p1 bs=1M count=1
        

        If Setup opens after that, it is the ESP filesystem, and we can bisect it file by file from
        there.

        Three questions while you are at it. Which Windows 11 build is the image, and which build
        was the Windows 10 one? Was the Hyper-V VM Generation 2 with Secure Boot and a vTPM
        attached? And what BIOS version are the 3400s on — Dell’s last for that model is in the 1.3x
        range, and firmware is the component under suspicion now.

        Also worth saying plainly: if this turns out to be the Latitude 3400 firmware choking on a
        partition layout it does not like, there may be nothing for FOG to fix beyond documenting
        it. Two earlier reporters on this hardware ended up rebuilding the image from a clean install
        rather than finding a cause. I would rather find it, and your Win10 control is the first
        thing anyone has produced that makes that realistic.

        Please help us build the FOG community with everyone involved. It's not just about coding - way more we need people to test things, update documentation and most importantly work on uniting the community of people enjoying and working on FOG! Get in contact with me (chat bubble in the top right corner) if you want to join in.

        Web GUI issue? Please check apache error (debian/ubuntu: /var/log/apache2/error.log, centos/fedora/rhel: /var/log/httpd/error_log) and php-fpm log (/var/log/php*-fpm.log)

        Please support FOG if you like it: https://wiki.fogproject.org/wiki/index.php/Support_FOG

        1 Reply Last reply Reply Quote 0
        • S
          servicedesk.pianezza
          last edited by

          Hi @Tom-Elliott
          Here’s the table comparison you asked for:

          Win10 (23H2), SSSTC drive:
          label: gpt, protective MBR
          p1: start=2048 size=204800 EF00 EFI
          p2: start=206848 size=32768 0C01 MSR
          p3: start=239616 size=498573824 (237.7 GiB) 0700 Basic Data
          p4: start=498813440 size=1304576 (637.0 MiB) 2700 WinRE
          Free: 2157 sectors (1.1 MiB), no problems found

          Win11 (25H2), Toshiba drive:
          label: gpt, protective MBR
          p1: start=2048 size=204800 EF00 EFI
          p2: start=206848 size=32768 0C01 MSR
          p3: start=239616 size=498344448 (237.6 GiB) 0700 Basic Data
          p4: start=498584064 size=1533952 (749.0 MiB) 2700 WinRE
          Free: 2157 sectors (1.1 MiB), no problems found

          Type GUIDs, attributes, partition order, and start sectors for p1/p2/p3 are identical between both. The only difference is where the p3/p4 boundary falls — Win11’s WinRE partition is ~112MB larger, which just shifts that one boundary. No extra/hidden partitions, no reordering, no odd attributes on either side, no MBR anomaly (both show “protective” cleanly). I don’t see anything in the table itself that looks like a red flag.

          Given how clean this diff is, I’m not expecting the bisection’s first step (sgdisk -Z, table-only) to change anything, but running it now anyway as instructed. Will report whether F2/Setup opens after that, and if not, move to zeroing the ESP start next.

          Thanks

          1 Reply Last reply Reply Quote 0
          • S
            servicedesk.pianezza
            last edited by

            Hi @Tom-Elliott
            Update on the bisection: re-deploying and re-testing multiple times shows this is NOT deterministic even with what should be an identical table each time. Most attempts hang at F2, but occasionally one boots through — same image, same deploy process, same resulting table.

            So sgdisk -Z (empty table) → 100% reliable, boots/F2 every time.
            Populated table (Win11 deploy) → hangs most of the time, but not always — same exact deploy repeated shows different outcomes.

            Given that, I don’t think this is “a wrong byte value” in the table (that would be reproducible 100% of the time either way). This looks more like the Latitude 3400 firmware having a genuine timing/race issue enumerating a populated multi-partition GPT NVMe at POST — mostly failing, occasionally succeeding, regardless of the specific bytes. That would also explain why the two 2020 forum reports on this same hardware family “fixed” it by wiping on another machine: not because specific stale bytes were the cause, but because an empty/simple table happens to avoid triggering whatever race condition the firmware has with a fully populated one.

            If that’s right, this isn’t something byte-level bisection will resolve — it’s a firmware reliability issue with this NVMe controller + BIOS 1.39.0 combination when handling multi-partition GPT disks, independent of FOG, the OS, or the capture pipeline. I don’t have a way to test that hypothesis further without either a Dell hardware diagnostic tool or a firmware engineer’s tools.

            At this point, given BIOS 1.39.0 is the latest available for this model, I don’t think there’s a software fix on my end. Documenting this for anyone else who hits it on Latitude 3400 with an NVMe-heavy image: expect intermittent F2/boot hangs with populated GPT tables on this platform, and budget for multiple retry attempts as a practical workaround, since a full wipe-and-redeploy doesn’t reliably avoid it either (confirmed just now — it just probabilistically improves odds by luck, not a guaranteed fix).

            Thanks for pushing the bisection methodology — even though it landed on “flaky firmware” rather than a fixable root cause, it at least closes the loop definitively and rules out FOG/the image as the actual fault.

            1 Reply Last reply Reply Quote 0
            • S
              servicedesk.pianezza
              last edited by

              Hi @Tom-Elliott
              Finally have a clean, confirmed answer — and I apologize for the back-and-forth on disk labeling in my last few posts, that was entirely on me getting physical drives and PowerShell disk numbers crossed. This is now verified cleanly across two different physical NVMe drives (SSSTC and Toshiba), each tested in both states:

              FOG-deployed Windows 11 (either drive) → Protective MBR (LBA0) has:

              • Real x86 bootstrap code present (the classic “Invalid partition table” / “Missing operating system” stub)
              • Partition entry: End CHS = FF FF FF (sentinel), Size field = exact sector count (e.g. AF 32 CF 1D)
                → Hangs at F2/Setup and normal boot, on both drives, reproducibly.

              Same drive, wiped and repartitioned natively via Windows diskpart (clean, convert gpt, create partition efi/msr/primary), then had real Windows boot files copied onto the ESP via robocopy (so it’s not just an empty table — it has genuine bootmgfw.efi, CIPolicies, language resources, etc.) → Protective MBR has:

              • Boot code region entirely zeroed
              • Partition entry: End CHS = FE 7F 99 (specific computed value), Size field = FF FF FF FF (sentinel)
                → Boots/reaches F2 Setup reliably, every time, tested repeatedly on both drives.

              So the variable isn’t the drive brand, isn’t the GPT header (identical format on both), isn’t the partition table entries (identical GUIDs/types/order) — it’s specifically how the Protective MBR at LBA0 is written. sgdisk/gdisk (what FOG’s restore uses) writes a real bootstrap stub (likely copied forward from the source Windows install during Partclone’s capture, since GPT disks normally don’t need real MBR boot code) combined with an exact size field. diskpart writes it the opposite way: zeroed boot code, sentinel size, specific CHS.

              Given the 3400’s firmware hangs specifically at “Preparing to enter BIOS Setup” (confirmed earlier — this is firmware disk enumeration, before any OS involvement) when it encounters the sgdisk-style Protective MBR, my best guess is a legacy/CSM-compatibility path in this BIOS reads that boot code region and/or the exact-size field and gets stuck — possibly trying to validate or execute the bootstrap stub even though the system is UEFI-only, or choking on an exact size value in a field it expects to be a sentinel.

              Question for you: does FOG’s restore process (or the underlying sgdisk/gdisk call) preserve/write real x86 boot code into the Protective MBR, and is there a flag to zero that region and/or force the sentinel size convention instead? If sgdisk has an option for this (or if it’s something Partclone carries over from the source capture rather than something sgdisk actively writes), that would be great to know — happy to test a patched version if you can point me to where in the FOG scripts this happens.

              This feels like a genuinely actionable, reproducible finding now — thanks for sticking with the bisection methodology, it’s what got us here.

              Thanks

              Tom ElliottT 1 Reply Last reply Reply Quote 0
              • Tom ElliottT
                Tom Elliott @servicedesk.pianezza
                last edited by

                @servicedesk-pianezza Yes, FOG writes that boot code, and I can show you where. Your finding holds up against our
                source and against a reproduction here.

                Where it comes from. On capture, saveGRUB() in funcs.sh copies the first 1 MiB of the
                source disk into d1.mbr with dd — LBA0 included, boot code and all. On deploy,
                clearPartitionTables() runs sgdisk -Z, which does clear the MBR, and then restoreGRUB()
                writes d1.mbr straight back over it. The next two commands are sgdisk -z, which destroys
                the GPT structures only and leaves the MBR bytes alone, and sgdisk -gl, which rewrites only
                the protective partition entry. So the captured Windows bootstrap survives the whole
                sequence, and gdisk supplies the entry in its own convention: EndCHS ff ff ff and the exact
                sector count.

                I replayed that sequence here on a loop-backed disk with one of our own Windows 11 resizable
                images. The result matches what you found byte for byte in shape:

                LBA0  : 33 c0 8e d0 bc 00 7c 8e ...      (Windows MBR stub)
                0x1BE : 00 00 02 00 ee ff ff ff 01 00 00 00 af 32 cf 1d
                

                Why it is not simply a bug. On a BIOS-booted GPT Linux disk that same region is GRUB’s
                boot.img, and d1.grub.mbr exists precisely so we keep it. We cannot blanket-zero LBA0 on
                every GPT restore. On a Windows GPT image it is dead weight — Windows only boots GPT through
                UEFI — so a targeted change is available to us. But we should change the right bytes.

                Which byte is it? Your diskpart disk differs from ours in three places at once, so we do
                not yet know which one the firmware chokes on. Three one-liners settle it. Start from a fresh
                deploy that hangs, run one of them, power off, power on, press F2. Redeploy between tests so
                each one is measured on its own:

                # 1 - the bootstrap, nothing else
                dd if=/dev/zero of=/dev/nvme0n1 bs=1 count=446 conv=notrunc
                
                # 2 - the size field -> sentinel
                printf '\xff\xff\xff\xff' | dd of=/dev/nvme0n1 bs=1 seek=458 conv=notrunc
                
                # 3 - EndCHS -> the diskpart value
                printf '\xfe\xff\xff' | dd of=/dev/nvme0n1 bs=1 seek=451 conv=notrunc
                

                Each writes inside LBA0 only and leaves the GPT untouched. Test 1 is the one I expect to
                matter, on the theory you already stated — a CSM path in that firmware reading or validating
                a bootstrap it should be ignoring.

                Name the byte and the fix is small: for a GPT image whose OS is Windows, clear that field
                during the restore and leave Linux images alone. That also gives the two 2020 reports on this
                hardware an explanation, which is worth having on its own.

                Please help us build the FOG community with everyone involved. It's not just about coding - way more we need people to test things, update documentation and most importantly work on uniting the community of people enjoying and working on FOG! Get in contact with me (chat bubble in the top right corner) if you want to join in.

                Web GUI issue? Please check apache error (debian/ubuntu: /var/log/apache2/error.log, centos/fedora/rhel: /var/log/httpd/error_log) and php-fpm log (/var/log/php*-fpm.log)

                Please support FOG if you like it: https://wiki.fogproject.org/wiki/index.php/Support_FOG

                1 Reply Last reply Reply Quote 0
                • S
                  servicedesk.pianezza
                  last edited by

                  This post is deleted!
                  1 Reply Last reply Reply Quote 0
                  • S
                    servicedesk.pianezza
                    last edited by

                    Hi @Tom-Elliott,
                    follow-up on this thread: I never found the root cause of the intermittent firmware hang on the Latitude 3400 (16 hypotheses ruled out, including the Protective MBR fields we discussed, thanks again for the help with that), but deploying these machines with diskpart + DISM from a WinPE booted via wimboot (through a small FOG plugin) has worked reliably so far, while every other model stays on FOS/Partclone. The attached report covers the investigation, the workaround and a few side findings (WinPE deploys leaving the task queued, Secure Boot/MOK, a vsftpd home-directory pitfall). If anyone has seen the same hang or has ideas about the cause, I’d like to hear it.Report_Latitude3400_FOG_BootIssue_EN.md.txt

                    Thanks for your help

                    Tom ElliottT 1 Reply Last reply Reply Quote 0
                    • Tom ElliottT
                      Tom Elliott @servicedesk.pianezza
                      last edited by

                      @servicedesk-pianezza Follow-up on the findings from your report. Four changes are merged. Three came from your
                      report directly. The fourth came out of triaging it.

                      A failed snapin download now answers an HTTP status. Snapins::stream() set no status
                      on its error paths, so the client received 200 with the error text as the file body, hashed
                      that text, and reported the same wrong SHA-512 on every retry. A storage-node or FTP failure
                      therefore read as a permanently corrupt snapin. The download paths now set the status before
                      any body is written: 404 when the snapin does not exist, 503 when the storage node is
                      unreachable, when the file cannot be read, or when no storage group or node is assigned. No
                      client change is needed. Both clients already treat a non-2xx as a failed transfer, log the
                      reason, and check in without computing a hash. (GH-1817)

                      The installer now ensures the service account’s home directory on every run. This is the
                      root cause of your FTP failure. configureUsers() created /home/fogproject and set its
                      ownership only in the branch that runs when the account does not yet exist. On every later
                      install the account was found, the step printed “Skipped”, and the home directory was never
                      examined. So a home directory that was removed could not be restored by re-installing, which
                      is exactly what you observed. vsftpd answers “cannot change directory” with no home
                      directory, and that breaks every snapin transfer. A new _ensureSvcUserHome() runs for both
                      branches, before anything writes into that directory. A path that exists as a file now fails
                      with that cause named, instead of being retried on every install. (GH-1820)

                      wipefs is now built into FOS. This is larger than the diagnostic you were missing.
                      restoreLVM() in the FOS libraries has always called wipefs -a to clear stale filesystem
                      and RAID signatures before it recreates a physical volume. The binary was never in the image,
                      and busybox has no wipefs applet, so that call has always been “command not found” with its
                      output and exit status discarded. pvcreate -ff forced past the leftover signatures, so
                      nothing appeared broken. A build check now refuses a configuration that calls wipefs
                      without building it. (FOS GH-188)

                      An unchecked “remove file data” box deleted the file anyway. We found this while
                      triaging your report, and it is the one with no undo. On the image and snapin edit pages the
                      checkbox handler returned early when the box came off, leaving the previous state in place.
                      Ticking the box, changing your mind, and deleting still removed the files from the storage
                      node. The server also accepted an explicit andFile=0 as a yes on the single-item path,
                      while the bulk path required the value. Both are corrected. (GH-1819)

                      Three further items from your report are already corrected in 1.6, so they need no change:

                      • Task timestamps are stored in UTC, and the timezone setting became a display setting. Your
                        two-hour offset is the old behavior.
                      • The installer no longer rewrites its vhost file wholesale. It splices its own block between
                        markers and preserves everything outside them. Note the limit: an Alias or RewriteCond
                        placed inside FOG’s markers is still replaced. Put yours outside them and it survives.
                      • The certificate tree is split into zones. The snapin CA path points at the root certificate,
                        and the web CA is a separate file. The crossed symlink you found belongs to the older
                        layout.

                      All four changes are on the development branch and on the release-candidate branch, so they
                      reach you in the next release candidate.

                      On the Latitude 3400 hang itself, nothing has changed. It is not a FOG defect, your winpe3400
                      plugin is the right answer for those machines, and the thread is the best record of it that
                      exists. The one hypothesis still untested is the drive’s TRIM state: Partclone writes only used
                      blocks and never trims, while diskpart format and a full wipe both do. If you ever have a
                      hanging drive to spare, Optimize-Volume -DriveLetter C -ReTrim in the HP, then back in the
                      Dell, then F2, would settle it.

                      Thank you again. Your write-up found three real defects, and one of them could have destroyed
                      an image.

                      Please help us build the FOG community with everyone involved. It's not just about coding - way more we need people to test things, update documentation and most importantly work on uniting the community of people enjoying and working on FOG! Get in contact with me (chat bubble in the top right corner) if you want to join in.

                      Web GUI issue? Please check apache error (debian/ubuntu: /var/log/apache2/error.log, centos/fedora/rhel: /var/log/httpd/error_log) and php-fpm log (/var/log/php*-fpm.log)

                      Please support FOG if you like it: https://wiki.fogproject.org/wiki/index.php/Support_FOG

                      1 Reply Last reply Reply Quote 0
                      • S
                        servicedesk.pianezza
                        last edited by

                        Hi @Tom-Elliott
                        After moving from RC-2 to RC-4 the plugin, the WinPE image and the step that closes the task at the end of the WinPE flow all kept working. Only the Apache part needed a change.

                        With RC-4 the whole vhost sits inside FOG’s managed block, so my earlier Alias plus rewrite exception was gone and /fog/winpe_3400/wimboot started answering 308 (FOG’s rewrite rule sends it to the API). Instead of putting the exception back, I now serve the files from /winpe_3400/ (no /fog/ prefix), with an Alias and a <Directory> in a separate conf-available file enabled with a2enconf. FOG’s rewrite rule only matches /fog/, so it never sees these requests, and nothing in 001-fog.conf should need restoring after an update. The only change in the plugin is its WIMBOOT_HTTP_PATH constant.

                        Checked after the change: the Latitude fetches wimboot, BCD, boot.sdi and boot.wim from the new path (all 200), WinPE starts, and a full deploy on the Latitude 3400 and one on an HP with the normal FOS/Partclone flow both completed and closed their tasks.

                        1 Reply Last reply Reply Quote 0
                        • 1 / 1
                        • First post
                          Last post

                        67

                        Online

                        12.8k

                        Users

                        17.7k

                        Topics

                        157.2k

                        Posts
                        Copyright © 2012-2026 FOG Project