• Recent
    • Unsolved
    • Tags
    • Popular
    • Users
    • Groups
    • Search
    • Register
    • Login
    1. Home
    2. Tom Elliott
    3. Posts
    • Profile
    • Following 27
    • Followers 83
    • Topics 120
    • Posts 19,218
    • Groups 0

    Posts

    Recent Best Controversial
    • RE: Windows 11 image captured from VM hangs indefinitely at Dell logo on Latitude 3400 (multiple units, multiple NVMe brands) — works fine on HP

      @servicedesk-pianezza Thanks for the detail — that is a good writeup, and it saves a lot of back and forth.

      Two things before the questions.

      First, one thing you can rule out on our side: FOS never writes UEFI boot entries. There is
      no efibootmgr anywhere in the FOS image, so a deploy cannot leave a bad NVRAM entry
      behind. Nothing on the server writes the Host Primary Disk field either — not full
      registration, not Quick Registration, not inventory. The only writers are the host edit
      form, the group form and the API. So the /dev/md0 you found was set by a person or by an
      import, and it is a clue, not a side effect: /dev/md0 can only appear in FOS if an md
      array was assembled, and FOS only does that when the host boots with the mdraid=true
      kernel argument. That means one of those disks carried RAID (or Intel RST / IMSM) metadata.

      Second, your symptom has been reported twice before on this exact family and neither
      report reached a root cause:

      • https://forums.fogproject.org/topic/14191 — Latitude 3400, resizable, stuck at the Dell
        splash, “the only fix is to wipe the drive and re-format it”
      • https://forums.fogproject.org/topic/14147 — Latitude 3500, resizable, same, and the same
        cure

      In both, the drive booted fine once moved to another Dell model, and only a wipe on another
      machine made the 3400/3500 usable again. That points at the firmware stalling while it
      scans the disk, not at Windows. A wipe removes more than the partition table — it removes
      whatever is left outside the partitions FOG restores.

      So the working theory is stale RAID metadata in the tail of the drive. A resizable deploy
      writes the partition table and the partitions; it does not zero the rest of the disk, so
      anything the factory install left at the end of the drive survives. Your /dev/md0 is
      direct evidence that metadata was there.

      Could you run these? All of them in a debug deploy task (tick Debug on the task), at the
      shell, before you type fog:

      wipefs /dev/nvme0n1                 # lists signatures, changes nothing
      mdadm --examine /dev/nvme0n1        # and on each partition, e.g. nvme0n1p4
      sgdisk -v /dev/nvme0n1
      gdisk -l /dev/nvme0n1               # the header lines, including any warnings
      

      Then, on a machine that is already hanging, the one test that separates firmware from
      Windows: with the deployed drive in, does F12 reach the boot menu, or does the machine hang
      before that too? And with the drive wiped (sgdisk -Z /dev/nvme0n1; wipefs -a /dev/nvme0n1) but nothing deployed, does it boot to “no bootable device”?

      If the wipe-then-deploy sequence boots, that is the answer and we will make FOG clear those
      signatures itself.

      Two BIOS settings worth clearing on these while you are in there, both known to cause a
      logo hang independent of FOG: set SupportAssist OS Recovery / “Auto OS Recovery Threshold”
      to off, and set Fastboot to Thorough. Also clear NVRAM once, since a 3400 that has failed
      to boot several times will have accumulated stale boot entries.

      On your sanitize question: a VM capture needs sysprep /generalize /oobe /shutdown before
      the capture, and that is the whole of it for hardware differences. It does not explain a
      hang this early — generalize problems show up as a BSOD or a spinner, not as a freeze at
      the vendor logo.

      posted in General
      Tom ElliottT
      Tom Elliott
    • RE: [1.6.0-beta] productKeyIsValid() regex missing letter 'N' — rejects valid Windows Enterprise product keys

      @servicedesk-pianezza Confirmed, and your fix is the right one. Thanks for the clear report.

      The check was using the 24-character Base24 alphabet. That set is what you use to decode an older key into its binary form, and N is not one of its digits. It is not the set of characters a key is printed with. Windows 8 and later put N in the key itself, so any key carrying one was refused at entry.

      There was a second half of the same bug. The browser keeps its own copy of that character list, for the masked display of a saved key. So even where a key with an N got stored, the page showed it as five groups of dots instead of keeping the first and last group. Both copies now say the same thing, and a test compares them character for character so they cannot drift apart again.

      The test also runs Microsoft’s published KMS client keys through the validator, and checks that the characters Windows leaves out (A E I O U L S Z 0 1) are still rejected and the length is still exactly 25.

      It is in PR #1773 against working-1.6: https://github.com/FOGProject/fogproject/pull/1773

      It will be in the next beta build. If you would rather not wait, the one-character edit you already made is exactly what landed.

      posted in Bug Reports
      Tom ElliottT
      Tom Elliott
    • RE: Fogserver 1.6 - Agent 0.1.9

      @Jason89436 Thanks for the report. The agent had no snapin pack support, and on Windows it also dropped backslashes from snapin arguments. Both are fixed. Update the server to the latest working-1.6, then update the agent to 0.1.10. After that, [FOG_SNAPIN_PATH]\msoffice.ps1 resolves to the unzipped pack folder, as it did with the legacy client.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Upgraded to 1.5.10.2482 - Now problems with replication to nodes

      @mp12 Thanks for the logs. This is a bug in 1.5.10.2482, not your node passwords.

      A security change in 2482 removes the storage node password from the node data that the API returns. The image and snapin replicators read their node list from that same data. So they now send an empty password, and every node rejects the login. The Undefined property: stdClass::$pass warning is that missing field.

      The fix is merged to dev-branch: https://github.com/FOGProject/fogproject/pull/1770

      To get it now, update from dev-branch:

      cd /path/to/fogproject
      git checkout dev-branch
      git pull
      cd bin
      sudo ./installfog.sh -y
      

      Or wait for the next stable release. Your stored passwords are correct, so you do not need to change anything on the nodes.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Wake-On-LAN via fog agent with brand new PC's

      @rdr

      How FOG wakes a host through the fog-agent wake relay

      Summary

      FOG wakes a host in two ways at the same time. First, the FOG server and the storage nodes send the magic packet on their own subnets. Second, if the relay is on, the server asks up to three awake agents on the host’s subnet to send the packet. The server always picks the target and the senders. An agent only sends.

      The problem

      A Wake-on-LAN (WoL) magic packet is a broadcast. A broadcast stays on one subnet. Before the relay, only the FOG server and the storage nodes sent the packet. So FOG could not wake a host on a subnet with no FOG server and no storage node.

      To send a broadcast to a remote subnet (“directed broadcast”), the routers must forward it. Most networks disable that router feature, because attackers used it for amplification attacks. The wolbroadcast plugin depends on that feature.

      The relay removes the gap. Every such subnet has FOG hosts on it. When one of them is awake, it can send the packet for its neighbor.

      How FOG knows which hosts share a subnet

      Each fog-agent reports its network interfaces to the server on every poll. For each interface, the report gives:

      • the IPv4 address and the prefix length (for example, 10.1.5.23 and /24)
      • whether the interface is up and has a link (a NIC with no cable is not up)
      • whether the interface is wireless

      The server does not trust a network address from the agent. It calculates the network address and the broadcast address itself, from the address and the prefix. It stores one row for each address in the hostNetwork table.

      Example:

      Host Reported Network the server calculates
      Host 41 (asleep) 10.1.5.23/24 10.1.5.0/24
      Host 77 (awake) 10.1.5.80/24 10.1.5.0/24
      Host 90 (awake) 10.1.0.12/16 10.1.0.0/16

      Hosts 41 and 77 share a subnet: the network address and the prefix are both equal. Host 90 does not share it. Its network address is different, and a /16 and a /24 are never one subnet.

      To find senders, the server does one database lookup. It takes the rows of the sleeping host, and it finds other hosts with the same network address and the same prefix. It then keeps only the hosts that meet all of these conditions:

      • The interface is up and has a link.
      • The interface has a broadcast address. A /31 or /32 link has none.
      • The interface is not wireless. An access point does not pass a broadcast to a machine that is asleep, because that machine is no longer connected to the access point.
      • The host polled in the last 900 seconds, so it is probably awake.
      • The host is not the sleeping host.

      The server sorts the result by the most recent poll and keeps the first three.

      Limits of this method

      • The sleeping host must run fog-agent. The server uses the sleeping host’s own last report to find its subnet. A host with the legacy FOG Client, or with no client, has no rows, so the relay cannot help it. The old path still runs for it.
      • The sleeping host’s subnet is its last report. A laptop that moved to a different subnet while off is looked for on the old subnet.
      • Only IPv4. A magic packet uses an IPv4 broadcast.

      The flow, step by step

      Step: someone asks for a wake. The source is the Wake Up button, a group wake, or a scheduled wol task. Each of these calls Host::wakeOnLAN().

      Step: the old path runs first, and it does not change. The server sends a request to every storage node and to itself. Each of them sends the magic packet on its own subnets. If a storage node shares the host’s subnet, this path is enough.

      Step: the relay path runs as an addition. It does nothing unless the global setting FOG_AGENT_WAKE_RELAY_ENABLED is 1. The default is 0. The server finds up to three senders, as described above. It writes one agentWake row for each pair of sleeping host and sender. Each row expires after 600 seconds.

      Step: the sender agent receives the request on its next poll. The agent does not listen on a network port. The request is part of the normal poll answer:

      "wake": {"targets": [{"id": 41, "macs": ["00:11:22:33:44:55"]}]}
      

      The block contains no destination address. It contains only the host id and the MACs of that host. The server leaves out pending MACs that nobody approved.

      Step: the agent sends the packet.

      • The agent parses each MAC and builds it again. It refuses a MAC that is not valid.
      • It builds the 102-byte magic packet.
      • It sends the packet to UDP port 9, at 255.255.255.255 and at the broadcast address of each of its own interfaces.
      • It sends to at most 32 hosts per poll, and at most one packet for each MAC on each interface.

      Step: the agent reports the result. The result is sent with a packet count, or failed with a reason. The server accepts the result only if a pending agentWake row names this sender and this sleeping host. Otherwise it returns 404. So an agent cannot report on a host that the server did not ask it to wake.

      Why the design has this shape

      Choice Reason
      The server picks the target and the senders A magic packet has no authentication. The control must be on who can ask. Only the server knows which hosts are real FOG hosts.
      The request is part of the poll, with no network port on the agent An open port lets anyone who can reach it ask for a broadcast.
      The request has no address field An agent that accepts a destination can be used to send traffic at any address.
      Three senders, not one Extra packets cost almost nothing. With one sender, the wake fails silently if that sender goes to sleep.
      Requests expire after 600 seconds A laptop that comes back next week must not send an old wake.
      Off by default One customer machine sends traffic for another. The estate owner must choose that.
      Wireless interfaces are never senders The access point does not deliver the broadcast to a sleeping machine.

      The cost: delay

      A relayed wake goes out on the sender’s next poll. With the default interval, that is up to 5 minutes. Most WoL use in FOG is a scheduled overnight task, so the delay is acceptable. The old path still sends immediately.

      What the relay is not

      • It cannot wake a device that FOG does not manage.
      • It does not replace the storage-node path or the wolbroadcast plugin.
      • Agents do not talk to each other. Every decision is the server’s.
      • It does not wake a host across the internet. A magic packet stays on one subnet.

      Sources

      • Design: fog-agent/docs/design/0011-wake-relay.md
      • Wire format: fog-agent/docs/design/protocol-v1.md, section “Wake”
      • Server: packages/web/src/Agent/WakeRelay.php (senders()), packages/web/src/Agent/NetworkFacts.php, packages/web/src/Items/Host.php (wakeOnLAN())
      • Agent: internal/network/network.go, internal/provider/wake/
      posted in General
      Tom ElliottT
      Tom Elliott
    • RE: Wake-On-LAN via fog agent with brand new PC's

      @rdr I think I don’t fully understand either.

      Basically, behind the scenes as I understand the build:

      FOG Server wants to wake a machine. FOG Server tries to send a Magic packet to the MAC address in question. It also checks the ARP table of the network to see what FOG agents might know about the MAC being requested to wake, and if an agent is alive on the same subnet that MAC lives on, it will ask the FOG Agent to also try to wake it up.

      So I don’t think it knows it the way we are thinking about things. It’s a bit more nuanced than that, but that’s highly suspected from my understanding of things, not necessarily exactly how it knows.

      I’ll ask claude code to see if it can give me a quick rundown of how it works from a design and flow perspective.

      posted in General
      Tom ElliottT
      Tom Elliott
    • RE: Deploy task never marked complete on GPT/UEFI disks with "Single Disk - Resizable" + Partition: Everything — client reboots into infinite deploy loop

      @GRISLET The task is only marked complete by one request: POST /fog/service/Post_Stage3.php, sent by fog.imgcomplete as the last step of the deploy. Your log shows it never went out, so the script exited before reaching it.

      You see no error because S99fog prints * Rebooting system as task is complete and reboots whenever /bin/fog exits, for any reason. A silent early exit is indistinguishable on screen from a real completion.

      The 48-second gap points at where. fog.statusreporter posts progress.php every 3 seconds for the whole task and stops only when killStatusReporter kills it — which is the first line of completeTasking. After that line, only three things run before the completion POST. One of them is /images/postdownloadscripts/fog.postdownload, which is sourced into the imaging shell. An exit or a reboot in that script, or in any script it calls with ., ends the task before FOG is told about it.

      Two things would confirm it:

      1. cat /images/postdownloadscripts/fog.postdownload, plus any script it sources. Look for exit or reboot.
      2. The last ten lines on the client screen before the reboot. Do Stopping FOG Status Reporter, * Task Complete and Updating Database appear? If they do not, the run ended early and the image type is not involved.

      The image type is probably a red herring. Nothing between the end of the restore and the completion POST depends on Single Disk - Resizable.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: FOG 1.5.10.2482 iPXE 2.0.0 - intermittent UEFI boot failures on Realtek NICs (1.21.1+ and snponly.efi unaffected)

      @AUTH-IT-Center I do believe snponly would be the recommended, rather than iPXE’s driver.

      The developers at iPXE wrote the driver on their own (of course using documentation and stuff, but for all intents/purposes it is still a handrolled driver) so anything is possible.

      We shipped the native iPXE 2.0.0 mainly because of the feature it allows with actual Secureboot capabilities and instead of embedding everyfile with a custom script, a more dynamic approach for when iPXE releases new version we can upgrade more easily.

      For what it’s worth, I would almost want more people to default to snponly.efi (or secureboot/snponly-shimx64.efi if using/wanting secureboot after enrolling your machines of course) because this is supposed to be using the generic driver for EFI boot protocols on the NIC rather then attempting to discover the NIC using a driver loaded.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: Wake-On-LAN via fog agent with brand new PC's

      @rdr I don’t believe the FOG-Agent would be the right thing for this, and neither would the FOG Server.

      In very old versions of FOG the IP address was a distinguishing factor of what a “host” was.

      Relatively shortely after that became a norm, machines really started to come with multiple NICs, and/or with Wifi. That plus many systems prefering to move to DHCP for their networks (due to Wifi needs, and manual intervention to change an ethernet when a machine moved from one place to the other) FOG stopped trying to use the IP as a definition of what a host was.

      WOL Works on macs and generally locally to the subnet the FOG server (or local machines) run on, but with some finagling you can send WOL packets on different network subnets.

      Welcome the WOLBroadcast Plugin that already exists. I will admit I’ve not tested this plugin in quite some time now, so please see if that will be more what you’re looking for?

      The FOG Agent shouldn’t be simply spamming your network with WOL packets just because it doesn’t know what is on the same subnet (to your initial point), but the FOG Server should be able to especially if you have your network already done for this.

      I hope this helps with what you’re looking for. The FOG Agent, in my eyes, should only try to do things it knows about, instead of being a rogue “DDoS” vector.

      posted in General
      Tom ElliottT
      Tom Elliott
    • RE: Wake-On-LAN via fog agent with brand new PC's

      @rdr there is a feature in the fog agent (although untested at this point) that can tell the fog-agent to try to send a WOL packets. If an agent machine is up, and you need to wol a machine that’s on the same subnet, it can happen. It is disabled by default.

      I think that’s what you were attempting to ask? I hope I didn’t butcher it too badly.

      https://docs.fogproject.org/en/latest/management/web/fog-agent#waking-a-host-on-a-subnet-with-no-fog-server

      posted in General
      Tom ElliottT
      Tom Elliott
    • RE: Secureboot preventing booting into windows after imaging

      Glad you have a workaround. I think the cause is the Windows boot manager certificate change, not the image.

      Your golden Optiplex installed Windows with Secure Boot on. Windows servicing then added the “Windows UEFI CA 2023” certificate to that machine’s db, and switched the boot files to a boot manager signed with it. The other Optiplex 3000s only trust the 2011 Microsoft certificates, so they reject that boot manager. bcdboot works because it copies the older 2011-signed boot manager.

      Can you confirm with two checks, in admin PowerShell, on the golden machine and on one target?

      [Text.Encoding]::ASCII.GetString((Get-SecureBootUEFI db).bytes) -match 'Windows UEFI CA 2023'
      
      mountvol S: /s
      (Get-AuthenticodeSignature S:\EFI\Microsoft\Boot\bootmgfw.efi).SignerCertificate.Issuer
      

      If the golden machine says True and the target says False, that is the cause. A newer Dell BIOS may include the 2023 certificate in its default keys. I am also looking at having FOS add it during the Secure Boot enrollment task.

      posted in Windows Problems
      Tom ElliottT
      Tom Elliott
    • RE: Unable to Startup SFTP subsystem

      @Strahd Can you please get the output of the line:

      grep -P 'Subsystem.*sftp' /etc/ssh/sshd_config

      It should output something like:

      Subsystem       sftp    internal-sftp
      

      If it does not look like the above, edit the file (as root) look for the matched line and make it look like the above, then restart sshd service (systemctl restart sshd)

      Thank you,

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: FOG 1.5.10 - Problem with AD Join.

      @gmaurice The host that you provided the log from is most definitely going to have the “Reset Encryption Data” because that token existing in the first place is what produces the “Invalid security token”.

      If you’re saying this “same” message is happening on all machines in your fleet, I’m not sure how that is possible and suspect you’re seeing similar but very different messages?

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: FOG 1.6 with fog-agent 0.1.6, agent renames all PC's to the same golden image name when not using sysprep

      @rdr This issue you described, should be fixed.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: FOG 1.6 with fog-agent 0.1.6, agent renames all PC's to the same golden image name when not using sysprep

      @rdr It’s okay, though topics would probably be better for SEO reasons but we will answer the questions generally.

      This should be harmless to run with these bu tyou’re right. If you didn’t make any custom changes then why is it arguing about it?

      I will see what I can do about this.

      Thank you!

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: FOG 1.5.10 - Problem with AD Join.

      @gmaurice Invalid Security Token is the key indicator here:

      Find the host(s) and go to them within the UI and click the button for “Reset Encryption Data” then you may need to restart the machine(s) in question.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: FOG 1.6 with fog-agent 0.1.6, agent renames all PC's to the same golden image name when not using sysprep

      @rdr Both things should be fixed:

      The “empty” Last checkin you saw was from the Legacy Client checkin time which itself is relatively new.

      Added a specific Agent check in column.

      Also should fix the issue with selector drop down updating on snapin update with a new file.

      Please update and test.

      Thank you!

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: FOG 1.6 with fog-agent 0.1.6, agent renames all PC's to the same golden image name when not using sysprep

      @rdr 0.1.8 introduced rolling update patterns where 0.1.6/0.1.7 did not have them quite yet.

      Thanks for letting me know about the reboot issue you were seeing and that the updates seemingly fixed.

      You can unset your “desired version” on the hosts and configure your global settings:

      FOG Configuration -> FOG Settings -> FOG Agent.

      FOG_AGENT_DESIRED_VERSION is now meant to be a “Pinned” update version so you can maintain (in fleet, or per host - per host winning of course) what versions your environment is using.

      FOG_AGENT_UPDATE_MODE is a selector for Pinned, Latest, No update. By default it’s disabled just in case. Since the FOG Agent is still very new, this may be a good setting to ensure reporting and testing.

      Documentation on all of this is on docs.fogproject.org.

      https://docs.fogproject.org/en/latest/management/web/agent-self-update
      https://docs.fogproject.org/en/latest/kb/reference/fog-agent-reference
      https://docs.fogproject.org/en/latest/tags/agent

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • RE: You are not running the most current version of FOG!

      @scottrayr I’m not able to replicate the issue you’re seeing.

      posted in Bug Reports
      Tom ElliottT
      Tom Elliott
    • RE: FOG 1.6 with fog-agent 0.1.6, agent renames all PC's to the same golden image name when not using sysprep

      @rdr You found a real bug. Sysprep is not needed, and it would not have helped.

      Cause: the image carries the agent’s key and certificate from MMF-LAB2-00.
      The agent compared the machine’s SMBIOS identity with that key only when it
      had no certificate. A deployed copy has one, so every PC connected as
      MMF-LAB2-00 and took its name.

      fog-agent 0.1.7, released today, makes the check on every start.

      To fix the PCs you already deployed:

      • Go to FOG Configuration > FOG Settings > General Settings and set
        FOG_AGENT_DESIRED_VERSION to 0.1.7. Every enrolled agent updates itself,
        including each PC that thinks it is MMF-LAB2-00.
      • After the update, each PC sees that it is not MMF-LAB2-00, makes a new
        key, and enrolls as itself. FOG matches it to its host by MAC.
      • If FOG deployed to that host in the last 24 hours, the enrollment is
        approved automatically. If not, approve it under Hosts > Pending Agents.
        The agent then renames the PC to its name in FOG.

      Please post back whether the names come right.

      posted in FOG Problems
      Tom ElliottT
      Tom Elliott
    • 1 / 1