Update Flow
1. Overview
WHost ships two independent update channels:
- WHost software updates (vendor update service → signed tarball → atomic in-place apply). Driven by the agent's update service behind the
/system/update/*API routes and by the agent's daily update scheduler. - OS package updates (apt-get on Debian-family, dnf on RHEL-family). Driven by the same router's
/system/packages/*block for operator-triggered upgrades, plus the scheduler's daily refresh of the pending list that surfaces security counts.
Both channels run with root privilege, write to system paths, and persist a permanent history record. Both are gated behind admin authentication (session cookie or signed key). The WHost channel backs the agent code up before mutating it and rolls back on failure; the OS channel delegates to the package manager.
2. WHost Software Update Pipeline
sequenceDiagram
participant Operator
participant API as POST /system/update/install
participant US as update_service.run_install
participant WL as WLicense (check-update / download)
participant FS as filesystem
participant Systemd
Operator->>API: install request (after GET /system/update/check)
API->>API: refuse: install in flight (UPDATE_IN_PROGRESS), backup work running (BACKUP_IN_PROGRESS), no check yet (UPDATE_CHECK_REQUIRED), no newer version (UPDATE_ALREADY_LATEST), malformed version (UPDATE_INVALID_VERSION), older release (UPDATE_DOWNGRADE_REJECTED)
API->>US: asyncio task; audit row update_install_started
US->>US: validate the download URL host against the licence servers
US->>FS: backup (5%) — tar of the agent code only (/opt/whost/agent/whost_agent, no configuration) into /var/whost/backups/update_<time>/ when backup_before_update is on (the scheduler keeps the newest three such folders, §5); a backup that fails is logged and the run goes on
US->>WL: check-update again — a fresh download token from the node that answers; the run ends when the offer is no longer the release it was started for
US->>WL: download (15%) — from the node that issued the token (a node answering 403 INVALID_TOKEN, one that cannot be reached and one answering a server error are skipped; any other refusal ends the run with the vendor's code), X-WLicense-SHA256 header
US->>US: verifying (35%) — SHA-256 = check answer = header
US->>WL: fetch <package>.asc — same node; only a 404 means "no signature", any other refusal ends the run
US->>US: verifying (40%) — RSA-PSS-SHA256 detached signature (embedded release public key; a missing or invalid .asc aborts)
US->>FS: verifying (45%) — free space: the unpacked size (read from the package) once for the extraction, once more for the staged copies beside the agent and panel trees, plus 256 MB, on each filesystem involved; too little space ends the run before anything is touched
US->>FS: extracting (50%) — safe_extract (no traversal, no symlink escape, entry and size limits); the agent version the package carries must be the offered release, otherwise the run ends here with nothing replaced
US->>FS: migrating (60%) — migrations/{from}_to_{to}.py when the package carries one (WHost packages carry none: the changes a release makes on the host run at its first start)
US->>FS: migrating (65%) — Python dependencies: the package's requirements.lock (every distribution pinned with its digests; requirements.txt for a package without a lock) is checked against the installed distributions (no network); only when something is missing is the environment copied to venv.new and pip run inside the copy in hash-checking mode
US->>FS: applying (70%) — venv.new → rename when one was built (previous kept as venv.old); agent tree whost_agent.new → rename, previous kept as .old; service unit + daemon-reload; installed requirements.txt and requirements.lock replaced (previous kept as .old)
US->>FS: applying (80%) — panel.new → rename (dirs 755, files 644 whatever the package carried: the web server reads the panel as an unprivileged user), previous kept as panel.old; config.json rewritten
US->>US: health_check (85%) — every plain .py module of the installed tree compiled in memory (no bytecode file is written); integrity manifest swapped (.lock.old / .lock.new)
alt a step fails after files were replaced
US->>FS: rolling_back — what this run replaced goes back (Python environment, agent tree, service unit, requirements file, panel tree, manifest), staging copies removed; the release put back is the one this process runs, so nothing is restarted
US->>US: history entry status=rolled_back with the error; audit row update_failed
end
US->>FS: completed (100%) — history entry status=completed written BEFORE the restart; audit row update_installed
US->>Systemd: restarting (100%) — the post-update check is written to /var/lib/whost/update-recovery/ (with a copy of the version record) and started at once as a transient service that sleeps 20 s; then systemd-run --on-active=3s systemctl restart whost-agent; the progress stream closes on this stage (an open stream would hold the old process through its graceful shutdown; the unit caps that at 10 s and the stop at 30 s) and the periodic memory check is deferred until the new process is up
Note over US,Systemd: the restart timer fires after this process has answered; the check outlives it as a unit of its own
The post-update check. A package can pass the code check and still not start (a compiled module built for another Python, an import that fails at start-up). The check runs as a transient service of its own, started before the restart is queued; it sleeps 20 s, then asks the agent's health route up to 18 times, 5 s apart (about 90 s). The first answer ends it. (It is a service, not a timer: a transient timer armed by the agent was seen to keep postponing its elapse while the agent unit cycled through restarts, and it never fired.) When the new release never answers, the check stops the unit, puts back what the update replaced (whost_agent.old, panel.old, integrity.lock.old, whost-agent.service.old, venv.old when the update installed Python packages and requirements.txt.old when it changed that list — the refused agent tree, panel tree, manifest and environment are kept as .failed), restores the persistent version record from the copy taken before the restart, marks the history row rolled_back, starts the unit again and writes what it did — and whether the restored release answers 10 s later — to the agent log (whost.update_watch). A run that ends in a rollback never reaches the restart, so no check is queued for it.
Channels and the beta programme. A host asking for stable is offered stable releases only; a host asking for beta is offered betas and stable releases. A beta (X.Y.Z-beta.N) is published on the beta channel only and is never mandatory, so a beta is never installed by the security auto-update policy on its own (§5). The beta channel answers only licences admitted to the beta programme; admission is granted by WISECP. A licence that is not admitted is answered from the stable channel and the check says so: beta_denied is true (its channel field keeps naming the channel this host asks for), and the version card shows the programme note ("the stable channel was used") with the stable answer. A server installed from a beta starts on the beta channel (§8). A release can first be offered to listed hosts only; a host outside the list is not offered it until the release is widened.
Which version is offered. The vendor's answer names two versions: the package offered to this host — the next step of its update path — and the last release on that path (further along than the offered step between stable releases; the host's own version when nothing is reachable). The version card, the install, the history row and the downgrade gate all use the package offered; after extraction, the __version__ of the agent tree inside the package must be that version (the package is signed as a whole, so this is the version the new process starts with), and a package that carries another version ends the run before anything is replaced. Between stable releases a host is offered one release at a time; between betas of one line, the newest.
Python dependencies of a release. The package carries the agent's requirements.lock — every distribution the agent needs, the ones pulled in included, pinned to an exact version with the SHA-256 of every file the index serves for it — and the loose requirements.txt it was resolved from (scripts/release/lock_requirements.py). Before anything is replaced, the run reads the lock (the loose list for a package that carries none) against the distributions installed in the agent's virtual environment (/opt/whost/venv): name and version of every listed line, and the presence of what those distributions pull in, extras included. Nothing is downloaded for this check, so a release that adds no dependency updates a host without Internet access exactly as before. When something is missing, the environment is copied to /opt/whost/venv.new, pip install --require-hashes -r runs with the copy's interpreter (up to 15 minutes; this step does need the package index, and a file served with another digest than the lock names is refused), the console scripts pip wrote are pointed back at /opt/whost/venv, and the apply step swaps the copy in. The environment the running process imports from is never installed into: a pip failure or a full disk (a copy of the environment plus 256 MB must fit) ends the run as failed with the reason (pip's last lines, or the space needed) and the copy is removed; an agent stopped during this step leaves venv.new behind, and the next run removes it. Nothing else was touched. An agent that does not run from a virtual environment refuses such a release with the list of missing packages instead of changing the system interpreter's packages. The installed copies of both lists (/opt/whost/agent/requirements.lock, requirements.txt) follow the release, so /opt/whost/venv/bin/python -m pip install --require-hashes -r /opt/whost/agent/requirements.lock repairs an environment by hand. Hosts without Internet access: install the listed packages into the environment before the update; the check then finds nothing missing and pip is not run.
Where a run is recorded. The steps of a run are written to the agent log (/var/log/whost/agent.log) under whost.update_service (pre-update backup, manifest, service unit, the queued check and restart, the reason of a refusal) and whost.router.updates (how a run started from the panel or the API ended; whost.auto_update_scheduler for an automatic install); the post-update check writes under whost.update_watch. The outcome is also kept as the history row (/admin/updates → History, with the reason of a rollback) and as audit rows (update_install_started, then update_installed or update_failed; an automatic install writes only the outcome row). From the apply step until the restart lands — or until a rollback has put the previous release back — the agent's periodic self-checks are deferred, because the files on disk are not the ones the process has loaded; a check that falls into that window writes memory and manifest checks deferred (whost.main) instead of running.
Backups and an update never overlap. A run ends in a restart of the agent, and a restart cuts whatever backup work the process is doing — a backup that is copying or uploading, a restore half written, a scheduled backup before its retention step. So an install is refused with 409 BACKUP_IN_PROGRESS while such work runs; the message names each piece of work and when it started (the panel shows it in the error toast), and the install is started again once it has finished. Backup work that starts in the moment between the route's answer and the run's first step ends the run as failed with the same message; nothing was changed. The other way round, from the run's first step until the restart replaces the process, a backup or a restore started from the panel or the API is refused with 409 UPDATE_IN_PROGRESS, and a scheduled backup that falls due waits: it runs on the scheduler's first tick after the update. An automatic install that finds backup work running is deferred without a failure notification and tried again 30 minutes later instead of on the next daily check. A process stopped by other means (an operator's restart, needrestart, a crash) can still cut backup work; the next start then removes the orphaned working trees in the background — the service answers first — and a copy to a remote destination the stopped process had not finished is marked on the backup's row as failed ("The remote copy did not finish: the agent stopped during the upload.") rather than reading as copied.
An agent stopped during a run. systemctl restart whost-agent — typed by an operator, or issued by a package manager's restart helper such as needrestart — can arrive while a run is swapping the installed trees. The steps of a run live in a task of their own, so a caller that goes away (the scheduler being stopped, the event loop shutting down) does not cut them, and the agent's shutdown waits up to 25 s — below the unit's 30 s stop timeout — for a run that has reached its apply step; it writes shutdown requested while an update is being applied (whost.update_service) and the restart takes that much longer. A run that has not reached the apply step has changed nothing on disk and is cancelled. The wait cannot cover a process that is killed outright (SIGKILL, power loss) between the code swap and the manifest swap: the next start's integrity check refuses that pair, no post-update check has been queued yet, and the previous release has to be put back by hand from whost_agent.old (and panel.old, whost-agent.service.old, venv.old, requirements.txt.old when present).
Why deferred restart? A systemctl restart whost-agent issued from inside the agent waits for the unit to stop, but the running unit is the process making the call — systemd stops it mid-call and the run never finishes its own steps. Solution: schedule a transient systemd-run unit that fires 3 s later, after this process has answered, and let systemd own the restart cycle.
Why write history before restart? Same root cause: the process that writes the record is the one the restart replaces, and the new process keeps nothing of the run in memory; writing the success record first leaves a complete history whatever the restart does.
No marker file. The new process reads nothing the run left behind: the history row is the record of the install, and there is no update_marker.json. The only thing that follows the restart is the post-update check, which runs as a unit of its own.
nginx snippets on the first start of a release. The panel's location snippet (/etc/nginx/snippets/whost-panel-locations.conf) and the webmail snippet (/etc/nginx/snippets/roundcube.conf) are written by the installer; the agent rewrites them otherwise only when the hostname, the advertised IP or Force SSL is saved (panel) or when Roundcube is installed from the panel (webmail). So that an install that is only ever updated does not keep the location set it was installed with, every start compares both files (the webmail one when the web server is nginx or nginx + Apache) with the templates of the running release and brings a differing copy to them: the file is written, nginx -t decides, nginx is reloaded, and a copy nginx refuses (or an nginx that cannot be run) gets the previous file back. The panel refresh leaves the server block (sites-available/whost-panel.conf, which carries the host's names and certificate paths) alone and writes the rate-limit zone file only when it is missing; the webmail refresh keeps the docroot and the PHP-FPM socket the existing file names and leaves a file it cannot read those from as it is. A fresh install already carries the templates, so its starts change nothing. The agent log records a change as Panel nginx snippet brought to the current template (whost.system_router) and Nginx Roundcube snippet written (whost.webmail_service); a refused copy as … failed nginx -t — reverted.
3. Stage Progression
The panel subscribes to a Server-Sent Events stream (GET /api/v1/system/update/status/stream; exempt from the bundle-fingerprint header because a browser's EventSource cannot send one, still gated by the session cookie or a signed key) for live progress. The stages and the progress percentages the pipeline publishes are:
| Stage | % | Notes |
|---|---|---|
idle |
0 | No install in flight (the stream answers one event and closes) |
backup |
5 | tar of the agent code (skipped when backup_before_update is off) |
downloading |
15 | the vendor is asked again for a fresh token (a token lives five minutes and only on the node that issued it), then the package is fetched from that node |
verifying |
35 → 45 | SHA-256 against the check answer and the header, then the .asc release signature, then the free-space check |
extracting |
50 | safe_extract member walk, then the version the package carries |
migrating |
60 → 65 | migrations/{from}_to_{to}.py when the package carries one (WHost packages do not), then the Python dependency check (and, only when a package is missing, the pip run inside a copy of the environment — this can take minutes) |
applying |
70 → 80 | environment swap when one was built, atomic agent swap + service unit + requirements file, then the panel swap |
health_check |
85 | every plain .py module of the installed tree compiled in memory, integrity manifest swap |
completed |
100 | history flushed, deferred restart pending |
restarting |
100 | post-update check started, systemd-run restart queued; the stream closes on this stage |
rolling_back → rolled_back |
0 | a step failed after files were replaced; what the run replaced is put back |
failed |
0 | the run ended before replacing anything (or its rollback could not finish) |
Each event is the status snapshot (stage, progress, message, error, versions), one every 2 s. The stream closes on completed, failed, rolled_back, idle and restarting; while the old process is restarting the browser sees the stream drop and the page re-opens it every 3 s until the new process answers its first event with the new current_version, then reloads. After 60 s without that answer the page stops waiting and says the agent did not come back (the update may still have completed: reload and check the version).
4. Rollback
sequenceDiagram
participant US as update_service
participant Live as /opt/whost/agent + /var/www/whost/panel + /etc/whost/integrity.lock + whost-agent.service
Note over US: an apply step, the code check or the manifest swap failed
US->>US: stage = "rolling_back"
US->>Live: venv.old renamed back to venv (only when this run swapped the environment); requirements.txt and requirements.lock restored from their .old copies
US->>Live: whost_agent.old renamed back over the current agent tree, whost_agent.new removed
US->>Live: integrity manifest restored from .lock.old (only when this run swapped it)
US->>Live: service unit restored from its .old copy (only when this run replaced it)
US->>Live: panel.old renamed back over the panel tree, panel.new removed
US->>US: history entry status=rolled_back with error_message; audit row update_failed
The release put back is the one the running process was started from, so a rollback restarts nothing: the operator sees the failure card with the agent's reason while the service keeps answering.
The atomic-rename pattern means a failed install either left the .new staging dir in place (no impact, rollback removes it) or completed the rename (rollback flips it back). There is no half-applied agent tree, and the panel follows the agent: a rolled-back installation serves the previous panel again.
A rollback undoes what the failed run itself replaced and nothing else. Every apply step first removes the .old copy an earlier update left behind, so a .old tree or integrity.lock.old found during a rollback can only belong to the release that was running when the install started — a refused second update never puts the manifest, the panel or the code of an older release back under the running one. A package refused before the apply steps (checksum, signature, extraction) changes nothing and triggers no rollback and no restart.
A package that passes the code check and still cannot start (for example a compiled module built for another Python) is put back by the post-update check described above; the .old trees stay in place until the next install, and a restore leaves the refused trees beside them as .failed for diagnosis. Should the check itself fail to bring the previous release back, its log lines say so and recovery is manual from those copies. venv.old follows the same rule as the trees: every run removes the copy an earlier update left, whether or not it builds an environment itself, so the environment found by a rollback or by the post-update check is always the one the running release was started from.
A rollback does not touch /etc/whost/agent.conf or the data under /var/lib/whost/. One case leaves a visible trace: a release that keeps the whitelabel logos as files moves logos still held inline in agent.conf into /var/lib/whost/branding/ at its first start (configuration.md › whitelabel). When that release is followed by one that predates the files — the post-update check putting the previous release back after such a start, or a manual downgrade — the panel shows no custom logo until the original values kept in /var/lib/whost/branding-legacy-<timestamp>.yaml are copied back into the whitelabel block of agent.conf and the agent is restarted.
5. Auto-Update Scheduler
The scheduler is a background task that starts and stops with the agent:
GET /api/v1/system/update/status — the stored result of the last check — so an admin page load never queries the vendor; only the scheduler and the updates page's own check do/var/whost/backups/update_<time>/) beyond the newest three, (1) check_for_updates() — a live vendor query, (2) refresh the OS package cache (when os_packages_check_enabled), (3) push the notifications below (update available only when notify_available is on), (4) auto-install if auto_update is on and the version delta matches auto_update_type — deferred while a backup, a restore or a scheduled backup runs, and tried again after 30 minsecurity (vendor-mandatory only, and only with allow_vendor_critical_override; without that consent this mode only notifies), patch (same X.Y: a patch step), minor (same X: a minor or patch step), all (any step, a new major included). A beta (X.Y.Z-beta.N) is sized by its X.Y.Z: the next beta of the same number and its release are patch steps; a version with any other suffix is never installed on its own. The default (security) installs no beta on its own, because a beta release cannot be mandatory — the admin is notified and installs from the panelallow_vendor_critical_override — when true, vendor-flagged "mandatory" updates auto-install regardless of auto_update_type/var/whost/update_notify_state.json — the same upstream version never notifies twice across restarts; the watermark is cleared after an auto-install. With notify_available off nothing is sent and the watermark stays where it was, so switching it on again announces the release still on offer at the next checkThe scheduler is intentionally lazy: a single 24 h check is enough because every protected endpoint already runs through license enforcement on every request. There is no need to poll the upstream every minute.
6. OS Package Channel
A separate, simpler pipeline for system-level updates:
sequenceDiagram
participant Operator
participant API as POST /system/packages/upgrade
participant Apt as apt-get / dnf
participant Status as GET /system/packages/upgrade/status
Operator->>API: {security_only, package_names?}
API->>API: validate package names (model regex; a name not in the pending list → 422)
API->>API: claim the single upgrade slot (in-memory; a running job answers 200 with its live status)
API->>Apt: background task (argv list, no shell)
API-->>Operator: 200 {stage: "running", ...} (returns immediately, polling)
loop until done
Operator->>Status: GET status
Status-->>Operator: {stage, progress, message, upgraded, output_tail}
end
Apt->>Apt: write history /var/whost/os_packages_history.json (last 100 attempts), refresh the pending cache
Validation: package names are checked against ^[a-z0-9][a-z0-9.+\-]*$ (apt's own name shape) and against the pending list. Any mismatch returns 422 before reaching apt. There is no shell — the agent calls apt as an argv list.
Security-only mode: on Debian and Ubuntu, security_only=true builds the package list from apt-get --just-print upgrade, taking the packages whose source suite ends in -security; on the RHEL family it runs dnf upgrade --security, which reads dnf's own security metadata. When no security update is pending, the job ends as completed with "No security updates available." — not as an error.
Concurrency: one upgrade job at a time; a second POST while one runs answers the live status instead of starting another. The route accepts two starts per hour per address (a third answers 429). While a job runs, GET /system/packages/updates serves the cached list instead of running apt under dpkg's lock.
Cost of the list read: GET /system/packages/updates runs apt-get update and a dry-run upgrade on every call (about ten seconds on a host with ninety pending packages); the panel reads it when the System Updates page (/admin/updates) opens, whatever the tab (the OS Packages tab label carries the count), and at most every five minutes; the topbar badge reads the cached /system/packages/summary.
7. Notification Surface
| Event | Recipient | Channel | Dedupe |
|---|---|---|---|
Update available (scheduler check; only when notify_available is on) |
admin | panel notification + e-mail, per the notification settings (this event has its own row in the event matrix) | last_notified_version |
| OS security updates pending (scheduler refresh; only when the count grew) | admin | panel notification + e-mail, per the notification settings (update category) |
last_os_security_count |
Auto-install completed (only when notify_installed is on) |
admin | panel notification + e-mail, per the notification settings (update category) |
watermark cleared afterwards |
| Auto-install failed (whatever the two notify switches) | admin | panel notification + e-mail (critical: sent whenever the panel or e-mail channel is on) | — |
An install started from the panel writes an audit row when it starts (update_install_started) and one for its outcome (update_installed, or update_failed with the reason and whether the previous release was put back), plus the history entry; it does not push a notification — the operator who started it is watching the stream. The install route accepts six starts per hour per address; a refused start answers 429 with retry_after set to the seconds left in that window.
8. Settings
/etc/whost/update_settings.json (root, 0600) is the only source the agent reads; agent.conf carries no update: block:
{
"channel": "stable",
"auto_update": true,
"auto_update_type": "security",
"backup_before_update": true,
"notify_available": true,
"notify_installed": true,
"os_packages_check_enabled": true,
"os_packages_notify_security": true,
"allow_vendor_critical_override": true
}
The installer writes the file once, on a new server, with the channel alone: a server installed from a -beta.N release starts on beta, one installed from a release on stable. Every other field keeps the default above, and an existing file is left as it is — running the installer again does not reset the admin's choice. An absent file means the defaults above.
Operator changes are made through the panel (/admin/updates → Settings tab), which writes the changed field through PUT /api/v1/system/update/settings; a save that changes nothing writes neither the file nor an audit row, and the audit row of a real change names only the fields that changed. A file that exists but cannot be parsed is logged and read as the defaults with auto_update: false until it is repaired; a save from the panel writes that state to the file, so automatic updates stay off until the admin turns them on again. Defaults are friendly: a freshly-installed server auto-applies vendor-flagged critical security patches (configurable for compliance environments where every patch needs explicit approval); a beta is never flagged, so these defaults never install a beta on their own (§5).
9. API Routes and Files
| Concern | Route or file | Notes |
|---|---|---|
| WHost software updates | /system/update/* |
check, changelog, install, status and its progress stream, history, settings |
| OS package updates | /system/packages/* |
summary, pending updates, upgrade, upgrade status, upgrade history |
| History | /var/whost/update_history.json (WHost) and /var/whost/os_packages_history.json (OS) |
atomic-write JSON arrays |
| Settings | /etc/whost/update_settings.json |
the only source |
The install route and the OS package upgrade route need an active licence (or one still inside its grace window); otherwise they answer 403 with LICENSE_NOT_ACTIVATED (never activated) or LICENSE_SUSPENDED (every other state). Recover the licence before retrying an update.
Our support team is here around the clock for anything you can't find above.