Categories

  • 13k Topics
    115k Posts
    Tom ElliottT

    @Coolguy3289 Thanks for the log line. It points at a bug, not at your approach.

    The installer creates a “FOG Agent CA” under the server root CA. It only creates it when the file is missing. If the root CA changes later (for example, you copy the old server’s /opt/fog/snapins/ssl onto the new box so existing clients keep trusting it), the agent CA stays signed by the first root. Every enrollment then fails with the error you see, and the agent gets a 503.

    The fix is in PR #1810: the installer now re-creates the agent CA when the current root did not sign it.

    To fix your server now, without waiting for the PR:

    sudo grep PKI_AGENT_CA_CERT /opt/fog/.fog-pki

    Move the .fogAgentCA.pem and .fogAgentCA.key files in that directory to a backup location. Then re-run the installer. It creates a new agent CA under your current root, and enrollment works.

    You do not need your internal PKI for this. The FOG-generated root is fine for production.

  • Get the latest news on what's happening.
    188 Topics
    829 Posts
    Tom ElliottT

    FOG 1.6.0-RC-4 is available

    The fourth release candidate for FOG 1.6.0 is on the rc-1.6.0 branch. It reports version 1.6.0-RC-4.

    Release page: https://github.com/FOGProject/fogproject/releases/tag/1.6.0-RC-4

    Upgrade from RC-3 now if you delete images or snapins from the web UI. In RC-3, the delete dialog can remove the files on the storage node even when the “remove file data” box is cleared.

    Fixed since RC-3

    A cleared “remove file data” box on the image or snapin delete dialog still deleted the files. It now keeps them. (#1819) When the storage node was unreachable, every snapin failed with “Hash does not match”. The server now returns an error status instead of a bad file. (#1817) The installer did not repair a missing service account home directory (by default /home/fogproject), which broke FTP and snapin downloads. It now repairs it on every run. (#1820) A failed schema update now shows the reason in the installer error log. (#1812) “On Server Size” now shows the image size right after a capture, not up to an hour later. This field is informational only. (#1814)

    Update (as root)

    A 1.6 beta or RC server: bin/updatefog.sh --channel rc from your FOG checkout. A 1.5 server installed from git: update to the current 1.5 stable, then run bin/updatefog.sh --channel rc. A new server, or a 1.5 server installed from a tarball:
    curl -fsSL https://raw.githubusercontent.com/FOGProject/fogproject/working-1.6/bin/bootstrap.sh | bash -s -- --channel rc

    Test on a lab or non-production server first. The upgrade from 1.5 to 1.6 changes the database schema, and there is no way back. Back up your database and /opt/fog/.fogsettings before you upgrade.

    Report problems at https://github.com/FOGProject/fogproject/issues with your FOG version, your OS, and the installer log from bin/error_logs/.

  • View tutorials or talk about FOG in general.
    2k Topics
    19k Posts
    8

    @Cpasjuste I just wanted to let you know that I’ve added in your method of using userland NFS support to my container. I have given you credit for the method in the README. I’d love for you to test it out if/when you get a chance. Thank you again for your work!

  • Report bugs, request features, or get the latest progress.
    2k Topics
    21k Posts
    Tom ElliottT

    @servicedesk-pianezza Follow-up on the findings from your report. Four changes are merged. Three came from your
    report directly. The fourth came out of triaging it.

    A failed snapin download now answers an HTTP status. Snapins::stream() set no status
    on its error paths, so the client received 200 with the error text as the file body, hashed
    that text, and reported the same wrong SHA-512 on every retry. A storage-node or FTP failure
    therefore read as a permanently corrupt snapin. The download paths now set the status before
    any body is written: 404 when the snapin does not exist, 503 when the storage node is
    unreachable, when the file cannot be read, or when no storage group or node is assigned. No
    client change is needed. Both clients already treat a non-2xx as a failed transfer, log the
    reason, and check in without computing a hash. (GH-1817)

    The installer now ensures the service account’s home directory on every run. This is the
    root cause of your FTP failure. configureUsers() created /home/fogproject and set its
    ownership only in the branch that runs when the account does not yet exist. On every later
    install the account was found, the step printed “Skipped”, and the home directory was never
    examined. So a home directory that was removed could not be restored by re-installing, which
    is exactly what you observed. vsftpd answers “cannot change directory” with no home
    directory, and that breaks every snapin transfer. A new _ensureSvcUserHome() runs for both
    branches, before anything writes into that directory. A path that exists as a file now fails
    with that cause named, instead of being retried on every install. (GH-1820)

    wipefs is now built into FOS. This is larger than the diagnostic you were missing.
    restoreLVM() in the FOS libraries has always called wipefs -a to clear stale filesystem
    and RAID signatures before it recreates a physical volume. The binary was never in the image,
    and busybox has no wipefs applet, so that call has always been “command not found” with its
    output and exit status discarded. pvcreate -ff forced past the leftover signatures, so
    nothing appeared broken. A build check now refuses a configuration that calls wipefs
    without building it. (FOS GH-188)

    An unchecked “remove file data” box deleted the file anyway. We found this while
    triaging your report, and it is the one with no undo. On the image and snapin edit pages the
    checkbox handler returned early when the box came off, leaving the previous state in place.
    Ticking the box, changing your mind, and deleting still removed the files from the storage
    node. The server also accepted an explicit andFile=0 as a yes on the single-item path,
    while the bulk path required the value. Both are corrected. (GH-1819)

    Three further items from your report are already corrected in 1.6, so they need no change:

    Task timestamps are stored in UTC, and the timezone setting became a display setting. Your
    two-hour offset is the old behavior. The installer no longer rewrites its vhost file wholesale. It splices its own block between
    markers and preserves everything outside them. Note the limit: an Alias or RewriteCond
    placed inside FOG’s markers is still replaced. Put yours outside them and it survives. The certificate tree is split into zones. The snapin CA path points at the root certificate,
    and the web CA is a separate file. The crossed symlink you found belongs to the older
    layout.

    All four changes are on the development branch and on the release-candidate branch, so they
    reach you in the next release candidate.

    On the Latitude 3400 hang itself, nothing has changed. It is not a FOG defect, your winpe3400
    plugin is the right answer for those machines, and the thread is the best record of it that
    exists. The one hypothesis still untested is the drive’s TRIM state: Partclone writes only used
    blocks and never trims, while diskpart format and a full wipe both do. If you ever have a
    hanging drive to spare, Optimize-Volume -DriveLetter C -ReTrim in the HP, then back in the
    Dell, then F2, would settle it.

    Thank you again. Your write-up found three real defects, and one of them could have destroyed
    an image.

49

Online

12.8k

Users

17.7k

Topics

157.2k

Posts