Windows 11 image captured from VM hangs indefinitely at Dell logo on Latitude 3400 (multiple units, multiple NVMe brands) — works fine on HP
-
Hi all, @Tom-Elliott
Update, and I want to push back on the VBS theory — I don’t think that’s actually it, despite what looked like supporting evidence.New, cleaner data point: with the deployed drive installed, the machine doesn’t just hang at the Dell logo during normal boot — it also hangs during “Preparing to enter BIOS Setup” when pressing F2. That’s firmware enumerating storage before Setup even opens, with no OS/bootloader/VBS involved at that stage. I confirmed this persists even after a full CMOS battery removal (physically pulled the coin cell, left it disconnected several minutes, reseated) — so NVRAM/CMOS state is ruled out too.
The clean isolation test: same physical NVMe drive, same machine —
- Drive wiped (sgdisk -Z + dd zero on first/last 10MB) with nothing deployed → boots fine, F2 reaches Setup normally.
- Drive with the Windows 11 image deployed (via FOG, resizable or not) → both normal boot AND F2/Setup entry hang.
- Drive with a Windows 10 image deployed via the identical FOG/Hyper-V capture pipeline → boots fine, every time, F2 works.
So it’s specific to whatever ends up on the drive after a Windows 11 deploy — and since F2/Setup-entry is affected too, this is firmware choking while scanning the disk itself, not a Windows boot-chain issue. I’d chased VBS/HVCI/TPM for a while (disabling the master switch reduced but did not eliminate the intermittent hang — went from failing most of the time to failing ~1 in 6-10 boots, which in hindsight might just be noise/coincidence rather than a real effect from that change).
Given F2 itself is affected, I think this points back toward something written to the disk outside the partitions FOG restores — back to your original stale-metadata theory, just not RAID/mdadm specifically. Could this be something IMSM/Intel RST related that mdadm --examine doesn’t recognize, or a vendor-specific NVMe log/telemetry region that only Windows 11’s install process (vs. 10’s) writes to during setup? Genuinely unsure what else to check here — happy to run any other read-only diagnostics (nvme-cli, smartctl, anything) against the raw device to compare a “good” (Win10 or wiped) drive against a “bad” (Win11) one, sector-region by sector-region if needed.
TThanks
-
@servicedesk-pianezza You are right and I was wrong. F2 hanging settles it. “Preparing to enter BIOS Setup” runs
before any bootloader, so the firmware is the thing that stalls, and my reading of the iPXE
test was bad. Ignore that whole post.Your Windows 10 control is the best evidence in this thread. Same capture pipeline, same
FOG, same drive, same machine, and it boots. Together with “resizable and non-resizable both
hang”, that clears our resize path and it clears FOG’s table writing as a general fault: we
are laying down the captured disk faithfully, and this particular captured disk upsets this
particular firmware.So the question is narrow now: which bytes? Two things, and the first one you can do from
your desk.Diff the two tables. You have a good drive and a bad drive from the same pipeline. In a
debug task on each, capture:sfdisk -d /dev/nvme0n1 sgdisk -v /dev/nvme0n1 gdisk -l /dev/nvme0n1Post both sets. Partition count, order, types, attributes and the end-of-disk figures are
the things the firmware reads at enumeration, and a diff of Win10-good against Win11-bad
names the difference without any guessing.Then bisect the disk, in this order. Start from a deployed drive that reproduces the
hang, and after each step power off, power on, press F2, and note whether Setup opens.Destroy the partition table only, leaving every byte of partition data where it lies:
sgdisk -Z /dev/nvme0n1If Setup now opens, the firmware is choking on the partition table — layout, types or
attributes — and not on anything inside the partitions. If it still hangs, the data is the
problem, so redeploy and zero the start of the ESP, which is the only partition the firmware
reads:dd if=/dev/zero of=/dev/nvme0n1p1 bs=1M count=1If Setup opens after that, it is the ESP filesystem, and we can bisect it file by file from
there.Three questions while you are at it. Which Windows 11 build is the image, and which build
was the Windows 10 one? Was the Hyper-V VM Generation 2 with Secure Boot and a vTPM
attached? And what BIOS version are the 3400s on — Dell’s last for that model is in the 1.3x
range, and firmware is the component under suspicion now.Also worth saying plainly: if this turns out to be the Latitude 3400 firmware choking on a
partition layout it does not like, there may be nothing for FOG to fix beyond documenting
it. Two earlier reporters on this hardware ended up rebuilding the image from a clean install
rather than finding a cause. I would rather find it, and your Win10 control is the first
thing anyone has produced that makes that realistic. -
Hi @Tom-Elliott
Here’s the table comparison you asked for:Win10 (23H2), SSSTC drive:
label: gpt, protective MBR
p1: start=2048 size=204800 EF00 EFI
p2: start=206848 size=32768 0C01 MSR
p3: start=239616 size=498573824 (237.7 GiB) 0700 Basic Data
p4: start=498813440 size=1304576 (637.0 MiB) 2700 WinRE
Free: 2157 sectors (1.1 MiB), no problems foundWin11 (25H2), Toshiba drive:
label: gpt, protective MBR
p1: start=2048 size=204800 EF00 EFI
p2: start=206848 size=32768 0C01 MSR
p3: start=239616 size=498344448 (237.6 GiB) 0700 Basic Data
p4: start=498584064 size=1533952 (749.0 MiB) 2700 WinRE
Free: 2157 sectors (1.1 MiB), no problems foundType GUIDs, attributes, partition order, and start sectors for p1/p2/p3 are identical between both. The only difference is where the p3/p4 boundary falls — Win11’s WinRE partition is ~112MB larger, which just shifts that one boundary. No extra/hidden partitions, no reordering, no odd attributes on either side, no MBR anomaly (both show “protective” cleanly). I don’t see anything in the table itself that looks like a red flag.
Given how clean this diff is, I’m not expecting the bisection’s first step (sgdisk -Z, table-only) to change anything, but running it now anyway as instructed. Will report whether F2/Setup opens after that, and if not, move to zeroing the ESP start next.
Thanks
-
Hi @Tom-Elliott
Update on the bisection: re-deploying and re-testing multiple times shows this is NOT deterministic even with what should be an identical table each time. Most attempts hang at F2, but occasionally one boots through — same image, same deploy process, same resulting table.So sgdisk -Z (empty table) → 100% reliable, boots/F2 every time.
Populated table (Win11 deploy) → hangs most of the time, but not always — same exact deploy repeated shows different outcomes.Given that, I don’t think this is “a wrong byte value” in the table (that would be reproducible 100% of the time either way). This looks more like the Latitude 3400 firmware having a genuine timing/race issue enumerating a populated multi-partition GPT NVMe at POST — mostly failing, occasionally succeeding, regardless of the specific bytes. That would also explain why the two 2020 forum reports on this same hardware family “fixed” it by wiping on another machine: not because specific stale bytes were the cause, but because an empty/simple table happens to avoid triggering whatever race condition the firmware has with a fully populated one.
If that’s right, this isn’t something byte-level bisection will resolve — it’s a firmware reliability issue with this NVMe controller + BIOS 1.39.0 combination when handling multi-partition GPT disks, independent of FOG, the OS, or the capture pipeline. I don’t have a way to test that hypothesis further without either a Dell hardware diagnostic tool or a firmware engineer’s tools.
At this point, given BIOS 1.39.0 is the latest available for this model, I don’t think there’s a software fix on my end. Documenting this for anyone else who hits it on Latitude 3400 with an NVMe-heavy image: expect intermittent F2/boot hangs with populated GPT tables on this platform, and budget for multiple retry attempts as a practical workaround, since a full wipe-and-redeploy doesn’t reliably avoid it either (confirmed just now — it just probabilistically improves odds by luck, not a guaranteed fix).
Thanks for pushing the bisection methodology — even though it landed on “flaky firmware” rather than a fixable root cause, it at least closes the loop definitively and rules out FOG/the image as the actual fault.
-
Hi @Tom-Elliott
Finally have a clean, confirmed answer — and I apologize for the back-and-forth on disk labeling in my last few posts, that was entirely on me getting physical drives and PowerShell disk numbers crossed. This is now verified cleanly across two different physical NVMe drives (SSSTC and Toshiba), each tested in both states:FOG-deployed Windows 11 (either drive) → Protective MBR (LBA0) has:
- Real x86 bootstrap code present (the classic “Invalid partition table” / “Missing operating system” stub)
- Partition entry: End CHS = FF FF FF (sentinel), Size field = exact sector count (e.g. AF 32 CF 1D)
→ Hangs at F2/Setup and normal boot, on both drives, reproducibly.
Same drive, wiped and repartitioned natively via Windows diskpart (clean, convert gpt, create partition efi/msr/primary), then had real Windows boot files copied onto the ESP via robocopy (so it’s not just an empty table — it has genuine bootmgfw.efi, CIPolicies, language resources, etc.) → Protective MBR has:
- Boot code region entirely zeroed
- Partition entry: End CHS = FE 7F 99 (specific computed value), Size field = FF FF FF FF (sentinel)
→ Boots/reaches F2 Setup reliably, every time, tested repeatedly on both drives.
So the variable isn’t the drive brand, isn’t the GPT header (identical format on both), isn’t the partition table entries (identical GUIDs/types/order) — it’s specifically how the Protective MBR at LBA0 is written. sgdisk/gdisk (what FOG’s restore uses) writes a real bootstrap stub (likely copied forward from the source Windows install during Partclone’s capture, since GPT disks normally don’t need real MBR boot code) combined with an exact size field. diskpart writes it the opposite way: zeroed boot code, sentinel size, specific CHS.
Given the 3400’s firmware hangs specifically at “Preparing to enter BIOS Setup” (confirmed earlier — this is firmware disk enumeration, before any OS involvement) when it encounters the sgdisk-style Protective MBR, my best guess is a legacy/CSM-compatibility path in this BIOS reads that boot code region and/or the exact-size field and gets stuck — possibly trying to validate or execute the bootstrap stub even though the system is UEFI-only, or choking on an exact size value in a field it expects to be a sentinel.
Question for you: does FOG’s restore process (or the underlying sgdisk/gdisk call) preserve/write real x86 boot code into the Protective MBR, and is there a flag to zero that region and/or force the sentinel size convention instead? If sgdisk has an option for this (or if it’s something Partclone carries over from the source capture rather than something sgdisk actively writes), that would be great to know — happy to test a patched version if you can point me to where in the FOG scripts this happens.
This feels like a genuinely actionable, reproducible finding now — thanks for sticking with the bisection methodology, it’s what got us here.
Thanks
-
@servicedesk-pianezza Yes, FOG writes that boot code, and I can show you where. Your finding holds up against our
source and against a reproduction here.Where it comes from. On capture,
saveGRUB()infuncs.shcopies the first 1 MiB of the
source disk intod1.mbrwithdd— LBA0 included, boot code and all. On deploy,
clearPartitionTables()runssgdisk -Z, which does clear the MBR, and thenrestoreGRUB()
writesd1.mbrstraight back over it. The next two commands aresgdisk -z, which destroys
the GPT structures only and leaves the MBR bytes alone, andsgdisk -gl, which rewrites only
the protective partition entry. So the captured Windows bootstrap survives the whole
sequence, and gdisk supplies the entry in its own convention: EndCHSff ff ffand the exact
sector count.I replayed that sequence here on a loop-backed disk with one of our own Windows 11 resizable
images. The result matches what you found byte for byte in shape:LBA0 : 33 c0 8e d0 bc 00 7c 8e ... (Windows MBR stub) 0x1BE : 00 00 02 00 ee ff ff ff 01 00 00 00 af 32 cf 1dWhy it is not simply a bug. On a BIOS-booted GPT Linux disk that same region is GRUB’s
boot.img, andd1.grub.mbrexists precisely so we keep it. We cannot blanket-zero LBA0 on
every GPT restore. On a Windows GPT image it is dead weight — Windows only boots GPT through
UEFI — so a targeted change is available to us. But we should change the right bytes.Which byte is it? Your diskpart disk differs from ours in three places at once, so we do
not yet know which one the firmware chokes on. Three one-liners settle it. Start from a fresh
deploy that hangs, run one of them, power off, power on, press F2. Redeploy between tests so
each one is measured on its own:# 1 - the bootstrap, nothing else dd if=/dev/zero of=/dev/nvme0n1 bs=1 count=446 conv=notrunc # 2 - the size field -> sentinel printf '\xff\xff\xff\xff' | dd of=/dev/nvme0n1 bs=1 seek=458 conv=notrunc # 3 - EndCHS -> the diskpart value printf '\xfe\xff\xff' | dd of=/dev/nvme0n1 bs=1 seek=451 conv=notruncEach writes inside LBA0 only and leaves the GPT untouched. Test 1 is the one I expect to
matter, on the theory you already stated — a CSM path in that firmware reading or validating
a bootstrap it should be ignoring.Name the byte and the fix is small: for a GPT image whose OS is Windows, clear that field
during the restore and leave Linux images alone. That also gives the two 2020 reports on this
hardware an explanation, which is worth having on its own. -
This post is deleted! -
Hi @Tom-Elliott,
follow-up on this thread: I never found the root cause of the intermittent firmware hang on the Latitude 3400 (16 hypotheses ruled out, including the Protective MBR fields we discussed, thanks again for the help with that), but deploying these machines with diskpart + DISM from a WinPE booted via wimboot (through a small FOG plugin) has worked reliably so far, while every other model stays on FOS/Partclone. The attached report covers the investigation, the workaround and a few side findings (WinPE deploys leaving the task queued, Secure Boot/MOK, a vsftpd home-directory pitfall). If anyone has seen the same hang or has ideas about the cause, I’d like to hear it.Report_Latitude3400_FOG_BootIssue_EN.md.txtThanks for your help
-
@servicedesk-pianezza Follow-up on the findings from your report. Four changes are merged. Three came from your
report directly. The fourth came out of triaging it.A failed snapin download now answers an HTTP status.
Snapins::stream()set no status
on its error paths, so the client received 200 with the error text as the file body, hashed
that text, and reported the same wrong SHA-512 on every retry. A storage-node or FTP failure
therefore read as a permanently corrupt snapin. The download paths now set the status before
any body is written: 404 when the snapin does not exist, 503 when the storage node is
unreachable, when the file cannot be read, or when no storage group or node is assigned. No
client change is needed. Both clients already treat a non-2xx as a failed transfer, log the
reason, and check in without computing a hash. (GH-1817)The installer now ensures the service account’s home directory on every run. This is the
root cause of your FTP failure.configureUsers()created/home/fogprojectand set its
ownership only in the branch that runs when the account does not yet exist. On every later
install the account was found, the step printed “Skipped”, and the home directory was never
examined. So a home directory that was removed could not be restored by re-installing, which
is exactly what you observed. vsftpd answers “cannot change directory” with no home
directory, and that breaks every snapin transfer. A new_ensureSvcUserHome()runs for both
branches, before anything writes into that directory. A path that exists as a file now fails
with that cause named, instead of being retried on every install. (GH-1820)wipefsis now built into FOS. This is larger than the diagnostic you were missing.
restoreLVM()in the FOS libraries has always calledwipefs -ato clear stale filesystem
and RAID signatures before it recreates a physical volume. The binary was never in the image,
and busybox has nowipefsapplet, so that call has always been “command not found” with its
output and exit status discarded.pvcreate -ffforced past the leftover signatures, so
nothing appeared broken. A build check now refuses a configuration that callswipefs
without building it. (FOS GH-188)An unchecked “remove file data” box deleted the file anyway. We found this while
triaging your report, and it is the one with no undo. On the image and snapin edit pages the
checkbox handler returned early when the box came off, leaving the previous state in place.
Ticking the box, changing your mind, and deleting still removed the files from the storage
node. The server also accepted an explicitandFile=0as a yes on the single-item path,
while the bulk path required the value. Both are corrected. (GH-1819)Three further items from your report are already corrected in 1.6, so they need no change:
- Task timestamps are stored in UTC, and the timezone setting became a display setting. Your
two-hour offset is the old behavior. - The installer no longer rewrites its vhost file wholesale. It splices its own block between
markers and preserves everything outside them. Note the limit: anAliasorRewriteCond
placed inside FOG’s markers is still replaced. Put yours outside them and it survives. - The certificate tree is split into zones. The snapin CA path points at the root certificate,
and the web CA is a separate file. The crossed symlink you found belongs to the older
layout.
All four changes are on the development branch and on the release-candidate branch, so they
reach you in the next release candidate.On the Latitude 3400 hang itself, nothing has changed. It is not a FOG defect, your
winpe3400
plugin is the right answer for those machines, and the thread is the best record of it that
exists. The one hypothesis still untested is the drive’s TRIM state: Partclone writes only used
blocks and never trims, whilediskpart formatand a full wipe both do. If you ever have a
hanging drive to spare,Optimize-Volume -DriveLetter C -ReTrimin the HP, then back in the
Dell, then F2, would settle it.Thank you again. Your write-up found three real defects, and one of them could have destroyed
an image. - Task timestamps are stored in UTC, and the timezone setting became a display setting. Your
-
Hi @Tom-Elliott
After moving from RC-2 to RC-4 the plugin, the WinPE image and the step that closes the task at the end of the WinPE flow all kept working. Only the Apache part needed a change.With RC-4 the whole vhost sits inside FOG’s managed block, so my earlier Alias plus rewrite exception was gone and /fog/winpe_3400/wimboot started answering 308 (FOG’s rewrite rule sends it to the API). Instead of putting the exception back, I now serve the files from /winpe_3400/ (no /fog/ prefix), with an Alias and a <Directory> in a separate conf-available file enabled with a2enconf. FOG’s rewrite rule only matches /fog/, so it never sees these requests, and nothing in 001-fog.conf should need restoring after an update. The only change in the plugin is its WIMBOOT_HTTP_PATH constant.
Checked after the change: the Latitude fetches wimboot, BCD, boot.sdi and boot.wim from the new path (all 200), WinPE starts, and a full deploy on the Latitude 3400 and one on an HP with the normal FOS/Partclone flow both completed and closed their tasks.