• Recent
    • Unsolved
    • Tags
    • Popular
    • Users
    • Groups
    • Search
    • Register
    • Login
    1. Home
    2. Tom Elliott
    3. Posts
    • Profile
    • Following 27
    • Followers 83
    • Topics 120
    • Posts 19,218
    • Groups 0

    Posts

    Recent Best Controversial
    • FOG 1.6.0-RC-3 Available

      The third release candidate for FOG 1.6.0 is available on the rc-1.6.0 branch. It reports version 1.6.0-RC-3.

      Test it on a lab or non-production server first. The upgrade from 1.5 to 1.6 is one-way: it changes the database schema, and there is no down-migration.

      Upgrade from RC-2 now if you use fog-agent, a wildcard web certificate, or more than one master node. RC-3 fixes failures in all three.

      Fixed since RC-2

      • Agent enrollment failed with a 503 after the root CA was replaced (#1810). The installer created the FOG Agent CA only when its file was missing. If the root CA changed later, for example when you copy an older server’s /opt/fog/snapins/ssl onto a new install, the Agent CA stayed signed by the old root. Every enrollment then failed, and the Apache log showed FOG agent enroll: signing for host <id> ... failed: the issued certificate does not verify against /opt/fog/snapins/ssl/CA/.fogCA.pem. The RC-3 installer re-creates the Agent CA when the current root did not sign it. It keeps the old pair beside the new one with a date suffix.
      • A wildcard web certificate broke the installer (#1800). The installer used the certificate’s commonName, for example *.example.org, as the server name. Apache refused the configuration, and the installer’s own calls to the server failed. --hostname was ignored. The installer now uses --hostname and checks that the certificate covers it. With no --hostname, it uses the server address.
      • Constant ssh connections between master nodes (#1806). The services opened a connection to the ssh port of every master node on each pass. sshd logged Connection closed by <server> every few seconds. The services no longer probe the nodes for this check.
      • fog-agent waited up to five minutes to see a task (#1802). The server told the agent to poll every 300 seconds. It now sends the client check-in interval, FOG_CLIENT_CHECKIN_TIME, which the legacy client already uses. With the default setting, agents poll every 60 seconds.

      New since RC-2

      • Windows activation through fog-agent (#1804). The server sends the host’s product key to the agent when the key is valid and the host enrolled as Windows. It needs a fog-agent release that supports activation. An older agent ignores it.

      Before you upgrade

      • Back up your database and /opt/fog/.fogsettings.
      • Read the release notes: https://github.com/FOGProject/fogproject/blob/rc-1.6.0/docs/release/1.6.0-release-notes.md
      • 1.6 removes Display Manager, Directory Cleaner, User Cleanup, Client Updater, Green FOG and the persistentgroups plugin. Their data is dropped. The site and accesscontrol plugins move into core.
      • PHP 7.4 or later is required. Plugins built for 1.5 do not load.

      Install or update (as root)

      • A 1.6 beta, RC-1 or RC-2 server: run bin/updatefog.sh --channel rc from your FOG checkout.
      • A 1.5 server installed from git: update to the current 1.5 stable, then run bin/updatefog.sh --channel rc from the checkout.
      • A new server, or a 1.5 server installed from a tarball:
        curl -fsSL https://raw.githubusercontent.com/FOGProject/fogproject/working-1.6/bin/bootstrap.sh | bash -s -- --channel rc

      Report problems

      Open an issue at https://github.com/FOGProject/fogproject/issues. Include your FOG version, your OS, and the installer log from bin/error_logs/. Report a security problem through “Report a vulnerability” on the same repository, not in a public issue.

      posted in Announcements
      Tom ElliottT
      Tom Elliott
    • RE: 1.6 Migration via new Install with no DB Transfer

      @Coolguy3289 Thanks for the log line. It points at a bug, not at your approach.

      The installer creates a “FOG Agent CA” under the server root CA. It only creates it when the file is missing. If the root CA changes later (for example, you copy the old server’s /opt/fog/snapins/ssl onto the new box so existing clients keep trusting it), the agent CA stays signed by the first root. Every enrollment then fails with the error you see, and the agent gets a 503.

      The fix is in PR #1810: the installer now re-creates the agent CA when the current root did not sign it.

      To fix your server now, without waiting for the PR:

      sudo grep PKI_AGENT_CA_CERT /opt/fog/.fog-pki
      

      Move the .fogAgentCA.pem and .fogAgentCA.key files in that directory to a backup location. Then re-run the installer. It creates a new agent CA under your current root, and enrollment works.

      You do not need your internal PKI for this. The FOG-generated root is fine for production.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Unable to capture image on NVME systems since 1.5.10

      @DiegoP Thanks, your workaround found the cause.

      FOS installs mdadm’s udev rule, and that rule assembles every RAID member at boot. It ignored mdraid=true. The arrays then held nvme0n1p2/p3, so capture could not read the partitions.

      FOS now assembles arrays only when mdraid=true is on the kernel command line. Without the flag, it leaves the member partitions alone. With the flag, nothing changes.

      The fix is in the experimental FOS release EXP_20261001-163131: https://github.com/FOGProject/fos/releases/tag/EXP_20261001-163131. Replace /var/www/html/fog/service/ipxe/init.xz on your FOG server with that release’s init.xz. Then remove the postinit workaround and capture again. Please tell us whether it works.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Task 0

      @kratkale I’m not sure what problem you’re trying to solve.

      if those 2 machines don’t have a valid image associated, what are you expecting to see for the 'Image Name"? In my head this is expected.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Task 0

      @kratkale Good, thank you for confirming.

      If a snap-in does not run, open a new topic for it, and post these three things there:

      • the snap-in task status from the web UI (Tasks, Active Tasks)
      • whether the PC runs the FOG Client or the FOG Agent
      • the client log from that PC: C:\fog.log for the FOG Client

      Your certificates are the original ones, restored from the old directory. The clients do not need to trust a new CA.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • FOG 1.6.0-RC-2 Available

      The second release candidate for FOG 1.6.0 is available on the rc-1.6.0 branch. It reports version 1.6.0-RC-2.

      Test it on a lab or non-production server first. The upgrade from 1.5 to 1.6 is one-way: it changes the database schema, and there is no down-migration.

      Upgrade from RC-1 now if your server was upgraded from 1.5. RC-1 left the certificate files of a 1.5 server in the wrong directory. RC-2 prevents this and repairs a server that RC-1 already affected.

      Fixed since RC-1

      • A 1.5 upgrade lost its certificates (#1797). FOG 1.5 made /etc/fog a link to /opt/fog/service/etc. The upgrade wrote the new CA and web certificates through that link, then replaced the link with an empty directory. The files stayed in /opt/fog/service/etc/pki, where FOG does not look. Apache kept running on the certificate it had loaded, so the upgrade looked clean. At the next restart or reboot, Apache failed with SSLCertificateFile: file '/opt/fog/pki/web/leaf/.webLeaf.pem' does not exist or is empty. RC-2 converts the link before it writes any certificate. On a server RC-1 already affected, the RC-2 installer copies the files back. It never overwrites a file that exists, and it keeps any differing copy beside the original as pki.recovered-<date>.
      • updatefog.sh --channel rc reported no release candidate when it could not reach the remote (#1793). A failed git ls-remote, for example as root with an SSH remote and no key, read as “No release candidate is currently published”. The installer now names the remote and says it could not reach it.
      • The installer summary showed the netboot protocol of the previous run (#1798). A re-run with --no-public-web-cert showed Netboot (PXE) protocol: https, and then correctly set up HTTP netboot. The summary now shows the protocol the run uses.
      • Versioning (#1796). The RC tag no longer changes the version that 1.5 branches compute.

      If Apache on your RC-1 server does not start

      Your certificates are most likely in /opt/fog/service/etc/pki. Update to RC-2 with bin/updatefog.sh --channel rc. The installer restores the files.

      Before you upgrade

      • Back up your database and /opt/fog/.fogsettings.
      • Read the release notes: https://github.com/FOGProject/fogproject/blob/rc-1.6.0/docs/release/1.6.0-release-notes.md
      • 1.6 removes Display Manager, Directory Cleaner, User Cleanup, Client Updater, Green FOG and the persistentgroups plugin. Their data is dropped. The site and accesscontrol plugins move into core.
      • PHP 7.4 or later is required. Plugins built for 1.5 do not load.

      Install or update (as root)

      • A 1.6 beta or RC-1 server: run bin/updatefog.sh --channel rc from your FOG checkout.
      • A 1.5 server installed from git: update to the current 1.5 stable, then run bin/updatefog.sh --channel rc from the checkout.
      • A new server, or a 1.5 server installed from a tarball:
        curl -fsSL https://raw.githubusercontent.com/FOGProject/fogproject/working-1.6/bin/bootstrap.sh | bash -s -- --channel rc

      Report problems

      Open an issue at https://github.com/FOGProject/fogproject/issues. Include your FOG version, your OS, and the installer log from bin/error_logs/. Report a security problem through “Report a vulnerability” on the same repository, not in a public issue.

      posted in Announcements
      Tom ElliottT
      Tom Elliott
    • RE: Task 0

      @kratkale That run is clean. Every step finished, the sudoers and certificate errors are gone, and all FOG services started.

      The summary near the top says “Netboot (PXE) protocol: https”. That line is wrong: it showed the setting from the previous run. The end of the run is correct: netboot uses HTTP. I am fixing the summary line in the installer.

      Please check these, in this order:

      grep chain /tftpboot/default.ipxe
      systemctl is-active FOGMulticastManager

      The first line must start with chain http://192.168.0.196/. The second must say active. Then PXE boot one PC. If it reaches the FOG menu, switch the other PCs back to PXE and try the multicast task again. If a step fails, post the output of that step only.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Task 0

      @kratkale Thank you, that output shows the cause, and your certificates are not lost.

      On FOG 1.5, /etc/fog was a link to /opt/fog/service/etc. The 1.6 upgrade moved your PKI files to /etc/fog/pki, so through that link they landed in /opt/fog/service/etc/pki. A later step of the same upgrade replaced the /etc/fog link with an empty directory. Your files are still in /opt/fog/service/etc/pki, but FOG no longer found them. This is an installer bug, and I am fixing it.

      Copy them back. cp -n never overwrites a file that already exists:

      cp -an /opt/fog/service/etc/pki/. /etc/fog/pki/
      [ -d /opt/fog/service/etc/customizations ] && cp -an /opt/fog/service/etc/customizations /etc/fog/
      ls -la /etc/fog/pki/root/ca /etc/fog/pki/web/leaf
      

      Both listings must show real files: .fogCA.key in the first, .webLeaf.pem and .webLeaf.key in the second.

      If you ran the Apache sed from my last post, put your original file back:

      [ -f /etc/apache2/sites-available/001-fog.conf.orig ] && cp /etc/apache2/sites-available/001-fog.conf.orig /etc/apache2/sites-available/001-fog.conf
      

      Then start Apache and run the installer, from your fogproject/bin directory:

      apachectl configtest && systemctl restart apache2
      ./installfog.sh -y --no-public-web-cert
      grep chain /tftpboot/default.ipxe
      

      The last line must start with chain http://. If any step fails, post the output of that step. Leave /opt/fog/service/etc/pki where it is for now.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Task 0

      @kratkale Apache stops because its certificate file is gone. The Secure Boot error last week has the same cause: files under FOG’s PKI directory are missing.

      I cannot find the cause without seeing what is left on disk. I asked for this on 2026-09-25, and every new error since then comes from the same missing files. So please run this one command first, before you change anything. It only lists file names and changes nothing:

      { ls -la /opt/fog /etc/fog /etc/fog/pki /etc/fog/pki/root/ca /etc/fog/pki/web /etc/fog/pki/web/leaf /opt/fog/snapins/ssl/CA; ls -ld /opt/fog/pki; find / -xdev \( -name '.fogCA.key' -o -name '.webLeaf.pem' -o -name '.fogWebCA.pem' \) -ls; } > /root/fog-pki-state.txt 2>&1
      

      Post the contents of /root/fog-pki-state.txt here.

      After that, this brings Apache back with a temporary certificate. Browsers will show a certificate warning, and PXE works again over http:

      apt-get install -y ssl-cert
      sed -i.orig --follow-symlinks -E \
        -e 's#^([[:space:]]*SSLCertificateFile)[[:space:]].*#\1 /etc/ssl/certs/ssl-cert-snakeoil.pem#' \
        -e 's#^([[:space:]]*SSLCertificateKeyFile)[[:space:]].*#\1 /etc/ssl/private/ssl-cert-snakeoil.key#' \
        -e 's#^([[:space:]]*)(SSLCertificateChainFile|SSLCACertificateFile|SSLVerifyClient|SSLVerifyDepth)#\1\# \2#' \
        /etc/apache2/sites-enabled/001-fog.conf
      apachectl configtest && systemctl restart apache2
      grep chain /tftpboot/default.ipxe
      

      The sed keeps your original file as /etc/apache2/sites-available/001-fog.conf.orig. The last line must start with chain http://.

      Do not run the installer again until we know where your CA files are. It would fail the same way.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • FOG 1.6.0-RC-1 Available

      FOG 1.6.0 Release Candidate 1 is available

      The first release candidate for FOG 1.6.0 is available on the rc-1.6.0 branch. It reports version 1.6.0-RC-1.

      Test it on a lab or non-production server first. The upgrade from 1.5 to 1.6 is one-way: it changes the database schema, and there is no down-migration.

      Before you upgrade

      • Back up your database and /opt/fog/.fogsettings.
      • Read the release notes: https://github.com/FOGProject/fogproject/blob/rc-1.6.0/docs/release/1.6.0-release-notes.md
      • 1.6 removes Display Manager, Directory Cleaner, User Cleanup, Client Updater, Green FOG and the persistentgroups plugin. Their data is dropped. The site and accesscontrol plugins move into core.
      • PHP 7.4 or later is required. Plugins built for 1.5 do not load.

      Install or update (as root)

      • A 1.6 beta server: run bin/updatefog.sh --channel rc from your FOG checkout.
      • A 1.5 server installed from git: update to the current 1.5 stable, then run bin/updatefog.sh --channel rc from the checkout.
      • A new server, or a 1.5 server installed from a tarball:
        curl -fsSL https://raw.githubusercontent.com/FOGProject/fogproject/working-1.6/bin/bootstrap.sh | bash -s -- --channel rc

      Report problems

      Open an issue at https://github.com/FOGProject/fogproject/issues. Include your FOG version, your OS, and the installer log from bin/error_logs/. Report a security problem through “Report a vulnerability” on the same repository, not in a public issue.

      posted in Announcements
      Tom ElliottT
      Tom Elliott
    • RE: Task 0

      @kratkale Two separate things. The first gets your PCs booting today.

      PXE boot, now. The installer stopped before it rewrote the boot file, so the file still says https. Change it by hand:

      sed -i 's#^chain https://#chain http://#' /tftpboot/default.ipxe
      grep chain /tftpboot/default.ipxe
      

      The line must now start with chain http://192.168.0.196/. Boot one PC to test it. The next complete installer run writes this file again, with http.

      The installer failure. Your certificates were not changed. The installer stopped before it created anything. The Secure Boot signing files that your settings name are not on disk, so it tried to create new ones. That needs the private key of your FOG root CA, and the key is not at /etc/fog/pki/root/ca/.fogCA.key. Your root certificate is still there, so your FOG clients are not affected.

      Please do not delete anything, and do not run the installer with --recreate-CA. That replaces the CA that all your FOG clients trust.

      Please post the output of these commands. They show only file names and paths, no key contents:

      ls -la /etc/fog/pki /etc/fog/pki/root/ca /etc/fog/pki/secureboot /etc/fog/pki/secureboot/ca /etc/fog/pki/secureboot/leaf /opt/fog/snapins/ssl/CA
      ls -ld /opt/fog/pki
      grep -E '^PKI_(root|sb)_' /opt/fog/.fogsettings
      find / -xdev -name '.fogCA.key' 2>/dev/null
      
      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Task 0

      @kratkale The PXE failure has one cause, and a re-run of the installer fixes it.

      On 2026-09-09 you ran the installer with --public-web-cert. The installer saved that setting. It is wrong for your server: your web certificate comes from FOG’s own CA, not from a public CA. Until yesterday, every upgrade stopped before the boot files. Yesterday the upgrade finished, and it applied the saved setting: iPXE now loads boot.php over HTTPS. iPXE cannot verify FOG’s own CA, so it stops with “Permission denied”.

      The sudoers errors have a second cause: the sudo package is not installed on your server. Those errors do not stop PXE boot.

      Update to 1.6.0-beta.5398 or newer. Then run this on the FOG server, from your fogproject/bin directory:

      ./installfog.sh -y --no-public-web-cert
      

      5398 installs sudo itself.

      Then check the boot file:

      grep chain /tftpboot/default.ipxe
      

      The line must start with chain http://192.168.0.196/. If it shows https://, post the output. Boot one PC before you switch the others back to PXE.

      The “Detected a web certificate managed outside FOG” message was wrong. 5398 fixes it. It does not affect PXE boot.

      5398 also warns if --public-web-cert is set on a certificate from FOG’s own CA.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Windows 11 Fog Client install failing with HTTPS

      @astrugatch I don’t think it matters really. The issue was more about the initial pinning of the certificate during the install process. After it is pinned I think this is the correct expectation (it should only communicate to the FOG server over HTTPS).

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Task 0

      @kratkale Thank you, that log shows the cause.

      Your web files are new, but the service files in /opt/fog/service are old. The new web code does not provide the name “FOGCore” that the old service code uses. So every FOG service stops at start, not only the multicast manager.

      Why: on each upgrade, the installer stopped at the certificate error before the schema step. It had already copied the web files, but it had not yet copied the service files.

      Both problems are now fixed in 1.6.0-beta.5396. The installer now accepts your FOG Web CA certificate. It also copies the service files immediately after the web files, so a failed step cannot leave them behind again.

      Please update to 1.6.0-beta.5396 or newer and run the installer again. It must finish without the “TLS verification failed” message. Then run:

      sudo systemctl restart FOGMulticastManager
      sleep 5
      sudo systemctl -l status FOGMulticastManager
      

      It must show “active (running)”. Then queue the multicast task, and post the output of tail -n 40 /opt/fog/log/multicast.log if it does not start.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Windows 11 image captured from VM hangs indefinitely at Dell logo on Latitude 3400 (multiple units, multiple NVMe brands) — works fine on HP

      @servicedesk-pianezza Yes, FOG writes that boot code, and I can show you where. Your finding holds up against our
      source and against a reproduction here.

      Where it comes from. On capture, saveGRUB() in funcs.sh copies the first 1 MiB of the
      source disk into d1.mbr with dd — LBA0 included, boot code and all. On deploy,
      clearPartitionTables() runs sgdisk -Z, which does clear the MBR, and then restoreGRUB()
      writes d1.mbr straight back over it. The next two commands are sgdisk -z, which destroys
      the GPT structures only and leaves the MBR bytes alone, and sgdisk -gl, which rewrites only
      the protective partition entry. So the captured Windows bootstrap survives the whole
      sequence, and gdisk supplies the entry in its own convention: EndCHS ff ff ff and the exact
      sector count.

      I replayed that sequence here on a loop-backed disk with one of our own Windows 11 resizable
      images. The result matches what you found byte for byte in shape:

      LBA0  : 33 c0 8e d0 bc 00 7c 8e ...      (Windows MBR stub)
      0x1BE : 00 00 02 00 ee ff ff ff 01 00 00 00 af 32 cf 1d
      

      Why it is not simply a bug. On a BIOS-booted GPT Linux disk that same region is GRUB’s
      boot.img, and d1.grub.mbr exists precisely so we keep it. We cannot blanket-zero LBA0 on
      every GPT restore. On a Windows GPT image it is dead weight — Windows only boots GPT through
      UEFI — so a targeted change is available to us. But we should change the right bytes.

      Which byte is it? Your diskpart disk differs from ours in three places at once, so we do
      not yet know which one the firmware chokes on. Three one-liners settle it. Start from a fresh
      deploy that hangs, run one of them, power off, power on, press F2. Redeploy between tests so
      each one is measured on its own:

      # 1 - the bootstrap, nothing else
      dd if=/dev/zero of=/dev/nvme0n1 bs=1 count=446 conv=notrunc
      
      # 2 - the size field -> sentinel
      printf '\xff\xff\xff\xff' | dd of=/dev/nvme0n1 bs=1 seek=458 conv=notrunc
      
      # 3 - EndCHS -> the diskpart value
      printf '\xfe\xff\xff' | dd of=/dev/nvme0n1 bs=1 seek=451 conv=notrunc
      

      Each writes inside LBA0 only and leaves the GPT untouched. Test 1 is the one I expect to
      matter, on the theory you already stated — a CSM path in that firmware reading or validating
      a bootstrap it should be ignoring.

      Name the byte and the fix is small: for a GPT image whose OS is Windows, clear that field
      during the restore and leave Linux images alone. That also gives the two 2020 reports on this
      hardware an explanation, which is worth having on its own.

      posted in General
      Tom ElliottT
      Tom Elliott
    • RE: Windows 11 image captured from VM hangs indefinitely at Dell logo on Latitude 3400 (multiple units, multiple NVMe brands) — works fine on HP

      @servicedesk-pianezza You are right and I was wrong. F2 hanging settles it. “Preparing to enter BIOS Setup” runs
      before any bootloader, so the firmware is the thing that stalls, and my reading of the iPXE
      test was bad. Ignore that whole post.

      Your Windows 10 control is the best evidence in this thread. Same capture pipeline, same
      FOG, same drive, same machine, and it boots. Together with “resizable and non-resizable both
      hang”, that clears our resize path and it clears FOG’s table writing as a general fault: we
      are laying down the captured disk faithfully, and this particular captured disk upsets this
      particular firmware.

      So the question is narrow now: which bytes? Two things, and the first one you can do from
      your desk.

      Diff the two tables. You have a good drive and a bad drive from the same pipeline. In a
      debug task on each, capture:

      sfdisk -d /dev/nvme0n1
      sgdisk -v /dev/nvme0n1
      gdisk -l /dev/nvme0n1
      

      Post both sets. Partition count, order, types, attributes and the end-of-disk figures are
      the things the firmware reads at enumeration, and a diff of Win10-good against Win11-bad
      names the difference without any guessing.

      Then bisect the disk, in this order. Start from a deployed drive that reproduces the
      hang, and after each step power off, power on, press F2, and note whether Setup opens.

      Destroy the partition table only, leaving every byte of partition data where it lies:

      sgdisk -Z /dev/nvme0n1
      

      If Setup now opens, the firmware is choking on the partition table — layout, types or
      attributes — and not on anything inside the partitions. If it still hangs, the data is the
      problem, so redeploy and zero the start of the ESP, which is the only partition the firmware
      reads:

      dd if=/dev/zero of=/dev/nvme0n1p1 bs=1M count=1
      

      If Setup opens after that, it is the ESP filesystem, and we can bisect it file by file from
      there.

      Three questions while you are at it. Which Windows 11 build is the image, and which build
      was the Windows 10 one? Was the Hyper-V VM Generation 2 with Secure Boot and a vTPM
      attached? And what BIOS version are the 3400s on — Dell’s last for that model is in the 1.3x
      range, and firmware is the component under suspicion now.

      Also worth saying plainly: if this turns out to be the Latitude 3400 firmware choking on a
      partition layout it does not like, there may be nothing for FOG to fix beyond documenting
      it. Two earlier reporters on this hardware ended up rebuilding the image from a clean install
      rather than finding a cause. I would rather find it, and your Win10 control is the first
      thing anyone has produced that makes that realistic.

      posted in General
      Tom ElliottT
      Tom Elliott
    • RE: Multicast: "No Active Task found for Host" on 1.5.10.2482 despite #1391 — msClients stays 0

      @Mikeee89 Thanks for the detailed report. It pointed straight at the cause.

      Cause: when udp-sender exits, the multicast manager closed the session and marked every host’s task Complete at the same time. The clients were still post-processing (UUID reset, NTFS flag, hostname). So when each client then reported back, its task was already gone: “No Active Task found for Host”. The imaging log row stayed open, which is where the “2026 years” duration comes from. This came in with the fix for #820.

      Fix: the manager still closes the session when the sender exits, but now leaves each host’s task for that host to close itself. PR #1776 (dev-branch), #1777 (working-1.6). I reproduced your exact error against a copy of a 1.5.10 database, and the patched code returns success and closes the imaging log.

      On your other observations:

      • msClients is reset to 0 when a session completes, so 0 after the run is expected. -2 is the marker for a session created from Image Management, not a counter underflow.
      • The multicast.log.udpcast.<id> file is deleted at the end of every session, and has been for years. That is a separate question, not part of this bug.
      • “has been killed” is printed even when the sender exited by itself. The wording is misleading, but nothing is actually killed.

      Once it merges, update to the latest dev-branch and the task should close normally.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Windows 11 Fog Client install failing with HTTPS

      @astrugatch Thanks, I can reproduce this.

      Cause: since 1.5.10.2253 the server signs its HTTPS certificate with a “FOG Web CA” intermediate under the FOG Server CA. The client installer only accepts an HTTPS certificate issued directly by “FOG Server CA” when it downloads ca.cert.der, so the download fails and the pin fails. The fix is in the client: https://github.com/FOGProject/zazzles/pull/48. It needs a new client release.

      Workaround until then, either one:

      • Trust the FOG CA before you install. Download http://<fog-server>/fog/management/other/ca.cert.der (plain HTTP works), then run certutil -addstore Root ca.cert.der as administrator. Then run the HTTPS install as before.
      • Or install with HTTPS off. The client downloads the CA over HTTP, and the server redirects it to HTTPS after that.
      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Task 0

      @kratkale 09-08-26 seems to me that the FOGMulticastManager service isn’t started or died somewhere.

      Can you run:

      sudo systemctl restart FOGMulticastManager
      sleep 5
      sudo systemctl -l status FOGMulticastManager
      

      On a separate window it might be helpful to see your php-fpm www-error logs (see my footer to see where to find that information)

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Windows 11 image captured from VM hangs indefinitely at Dell logo on Latitude 3400 (multiple units, multiple NVMe brands) — works fine on HP

      @servicedesk-pianezza Those are clean tests and they kill my theory. Correcting the record, and then I think your
      own last test moves this a long way.

      The stale-metadata idea is dead. A zeroed disk with a fresh deploy still hangs, and
      mdadm --examine found no superblock either side of the wipe. Drop it.

      I also replayed our GPT restore path locally against a Windows 11 resizable image, onto a
      disk deliberately smaller than the captured one — the same order FOS uses: dd of d1.mbr,
      sgdisk -z, sgdisk -gl, then the filldisk table through sfdisk. The result verifies clean:
      sgdisk -v reports no problems, the protective MBR is a single 0xEE entry spanning the whole
      device, and first/last usable sectors match the target. So a malformed partition table is not
      what we are looking at either.

      Now the part I think you undersold. You reached the FOG iPXE menu and chose “Boot from hard
      disk”, with the deployed drive fitted, and then it hung. That means the firmware finished
      POST, brought up the NIC, ran iPXE and drew a menu — all with that drive present. The
      firmware is not the thing that hangs.
      It hands off to bootmgfw.efi and the hang is after
      that point.

      Which reframes the symptom. “Stuck at the Dell logo” is not the firmware stalling. It is the
      Windows boot chain hanging before anything repaints the screen, so the OEM logo simply stays
      up. F12 being dead is expected there — the firmware gave up the keyboard at handoff. It also
      explains the missing Automatic Repair: Windows’ boot-failure counter is incremented by the
      boot manager, and a hang never reaches the code that does it.

      So the question is now why this image hangs in early Windows boot on a Whiskey Lake Latitude
      and not on a 12th-gen HP. Two things to do, in this order.

      Deploy the same image to the same Dell as Single Disk (Not Resizable). This splits the
      problem in half for the cost of one deploy. Both 2020 reports on this hardware said
      non-resizable worked where resizable did not — see banana123 in topic 14147, “Using Multiple
      Partition Image - Single Disk (Not Resizable) DOES work fine”. If non-resizable boots, the
      fault is in our resize path, it is ours, and I will want d1.minimum.partitions,
      d1.fixed_size_partitions and a full debug-deploy transcript. If it hangs too, the resize
      path is exonerated and it is the image.

      Make Windows tell you where it stops. Boot a Windows installer USB on the hung machine,
      Shift-F10 for a command prompt, find the ESP letter with diskpart, then:

      bcdedit /store X:\EFI\Microsoft\Boot\BCD /set {default} sos on
      bcdedit /store X:\EFI\Microsoft\Boot\BCD /set {globalsettings} bootmenupolicy legacy
      

      sos replaces the logo with the list of boot drivers as they load, so the screen names the
      last thing it got to instead of showing you a logo. bootmenupolicy legacy gives you the F8
      menu, and Safe Mode is itself a useful result.

      Last question, because you have not said it anywhere in the thread: was the image captured
      after sysprep /generalize /oobe /shutdown, or from a VM that had simply been shut down? An
      image captured without generalize carries the source machine’s driver and device state, and
      booting on one chipset but not another is the usual way that shows up.

      posted in General
      Tom ElliottT
      Tom Elliott
    • 1
    • 2
    • 3
    • 4
    • 5
    • 960
    • 961
    • 1 / 961