Troubleshooting
Companion docs: admin-user-guide.md · installation.md
This is the operator's first-stop guide when something stops working. Cross-references to deeper architecture docs are inline where helpful.
1. Log paths cheatsheet
/var/log/whost/agent.log # Agent log, its stdout/stderr included (rotated by the agent: logging.max_size 100M, logging.keep 10)
/var/log/whost/frontend-debug.log # Browser-side debug beacon (opt-in; logrotate daily, 7 kept)
/var/log/whost/audit.jsonl # Audit trail, JSON lines (older entries move to audit.jsonl.1 / .2)
/var/log/nginx/{access,error}.log # Nginx — global
/home/<user>/logs/ # Per-account web logs (<domain>-access.log, <domain>-error.log)
/var/log/apache2/ # Debian/Ubuntu Apache backend
/var/log/httpd/ # RHEL Apache backend
/var/log/mail.log /var/log/maillog # Postfix / Dovecot
journalctl -u pdns -n 100 # PowerDNS (WHost configures no log file for it)
/var/log/fail2ban.log # Fail2Ban actions
/var/log/modsecurity/audit.log # Requests matched by ModSecurity rules
tail -f /var/log/whost/agent.log # Live tail of the agent log
journalctl -u whost-agent -n 50 # Agent start / stop / exit events recorded by systemd
journalctl -u nginx -n 100 # Nginx journal
tail -n 100 /var/log/whost/agent.log
The agent writes its output to this file, not to the journal: journalctl -u whost-agent shows systemd's start, stop and exit lines and the system-log lines of commands the agent runs (useradd, userdel and the like), never the agent's own log lines.
2. Agent lifecycle commands
systemctl status whost-agent # State; the lines it shows are systemd events only
systemctl restart whost-agent # Graceful restart (lifespan stops cleanly)
systemctl reload whost-agent # Do not use — see below; use restart
systemctl is-active whost-agent # Boolean
tail -f /var/log/whost/agent.log # Live tail
curl -k https://127.0.0.1:2000/health # Health probe (no auth required)
systemctl reload is not a reload: the unit maps it to a SIGHUP, which the
agent does not handle, so the process ends without a graceful stop and systemd
starts it again after 5 seconds. Use restart.
Config: /etc/whost/agent.conf (YAML, mode 0600, root-owned). After
editing, restart the agent. A YAML syntax error or an invalid value stops the
start: the error (for a syntax error with its line and column) is written to
/var/log/whost/agent.log, and systemd tries again every 5 seconds until the
file is fixed.
3. License-related issues
| Symptom | Diagnosis | Recovery |
|---|---|---|
| After signing in, the license page opens instead of the dashboard, or an existing admin session is redirected to the license page | State is NOT_ACTIVATED or SUSPENDED |
Restore the existing license's validity or connectivity and inspect the reason on the license page; the sign-in is not refused in these states, its session reaches the license page only. For a new installation, sign in with the installer's admin password and enter the key on the license page that opens. |
| Settings → License shows a license-server warning and hours remaining | License verification service unreachable (or rate-limiting the host, e.g. after many agent restarts); cached state still valid | Verify outbound HTTPS (port 443) reachability. Entering grace reschedules the agent's check to 9–18 minutes; recovery needs no restart. The remaining window is measured from the last successful verification, using license.grace_period. |
The activation form answers "This license key is already registered to an installation" (400 LICENSE_ACTIVATION_FAILED, details.vendor_code AUTH_FAILED; agent.log: app_secret required) |
The key was activated before — typically by this server before it was reinstalled — and the credentials that installation received are no longer on the disk. The form cannot supply them, so the license service refuses the key. The refusal changes nothing on either side | Reissue the license with your license provider (WISECP licenses: the service's management screen on wisecp.com) and enter the key again. If the earlier installation still runs, stop its agent first: a reissue goes to the first request that reaches the license service, and a running agent would collect it on its next check |
Activation returns 400 LICENSE_ACTIVATION_FAILED with another message |
The license service refused the key: unknown, cancelled or expired key, or a domain/IP lock on the license that does not match this server (the message carries the service's own reason, details.vendor_code its code; activation_secret plays no part in the activation form) |
Compare the key and the lock settings on the wisecp.com Service Management screen; update the license's IP/domain there or reissue it; if the message names the hardware fingerprint or the boot token (the server's hardware changed or WHost was reinstalled), choose Reissue License on the wisecp.com service page — it also releases the hardware lock — and activate again; WISECP support can reset the hardware lock as well |
agent.log shows CERTIFICATE_VERIFY_FAILED or a hostname mismatch for the license servers |
The license service's certificate chain or name did not validate against the host's CA store: a tampered /etc/hosts or DNS answer, an intercepting proxy, or an outdated ca-certificates package |
The agent never talks to an unverified endpoint; it counts that server as unreachable (grace) and keeps retrying. Fix the resolution path or the proxy, or refresh the CA store (update-ca-certificates on Debian/Ubuntu, update-ca-trust on EL9). No restart needed. |
Grace time expires and protected APIs return 403 LICENSE_SUSPENDED |
No successful verification within license.grace_period |
Restore license-service connectivity and verify through Settings → License. The agent restricts access when the window ends, without waiting for the next retry. Avoid repeated restarts, which consume the verification rate limit. |
Protected APIs such as /api/v1/accounts return 403 LICENSE_SUSPENDED |
Enforcement engaged | Check Settings → License for the reason and restore license validity. Health, auth, license and selected read-only system routes remain reachable; their normal authentication requirements still apply. |
Diagnostic:
The examples below use the default state path. If license.state_file was
changed in agent.conf, inspect that file and its .hmac companion instead.
cat /etc/whost/license.state # current state + token_valid_until (indented JSON)
grep -i license /var/log/whost/agent.log | tail -30
# Force a verify: Settings → License → Verify Now, or through the API with an
# admin session cookie or an HMAC-signed request. The agent listens on
# 127.0.0.1:2000 only; from another machine use the panel's HTTPS address.
curl -X POST -H "X-WHost-Key: ..." -H "X-WHost-Timestamp: ..." \
-H "X-WHost-Nonce: ..." -H "X-WHost-Signature: ..." \
https://<panel-host>/api/v1/license/verify
4. Authentication / 2FA
| Symptom | Action |
|---|---|
| "Invalid username or password" although the password is right | The admin sign-in accepts the admin user name (case-sensitive) or the admin e-mail address. grep 'Admin login failed' /var/log/whost/agent.log shows which part was refused (unknown identifier or wrong password). After five failed attempts the name is locked for 15 minutes and the form shows a countdown instead. An address banned by Fail2Ban is refused before it reaches the sign-in form, so it does not produce this message: check fail2ban-client status whost-agent (and recidive), lift a ban with fail2ban-client unban <ip> or from the Fail2Ban page. |
| 2FA code repeatedly rejected | The server accepts a code of the current 30-second step or the one before or after it, so a clock more than about 30 s off fails: run timedatectl status; enable NTP if needed. Each code is accepted once; after five wrong codes the sign-in is locked for 30 minutes. |
| Lost 2FA device | Use one of the backup codes issued at enrollment (each works once). If none is left, an operator with shell access sets enabled: false under admin: → two_factor: in /etc/whost/agent.conf, restarts the agent, signs in with the password and enrolls again from the Profile page. |
| All sessions logged out after settings save | Saving a new, non-empty Admin Panel IP Whitelist (Settings → Security) ends every admin session; a password change ends the other sessions. The admin panel also keeps a single session, so signing in elsewhere ends the earlier one. Expected — sign in again from an allowed address. |
401 AUTH_FAILED on every API request after replacing an API key |
The integration still signs with the old key: a revoked or deleted key stops working at once, and a key's secret is shown only when the key is created (keys are not rotated in place). Put the new key ID and secret into the integration; no restart is needed. agent.log names the refusal (Revoked API key used from …, Invalid API key from …, Invalid HMAC signature from …), and these lines count towards a Fail2Ban ban of the caller (whost-agent jail). |
| Sign-in answers "Security verification failed" (captcha) | Settings → Security → CAPTCHA Protection: the selected provider must match the site key / secret key pair, the key must be registered for the panel's host name, and a reCAPTCHA v3 score below the Score Threshold (default 0.5) is refused. If nobody can sign in, set enabled: false under admin: → recaptcha: in /etc/whost/agent.conf and restart the agent. |
5. Account provisioning
| Symptom | Likely cause | Fix |
|---|---|---|
POST /api/v1/accounts returns 409 ACCOUNT_EXISTS |
The username belongs to an existing account (/etc/whost/accounts/<username>.json), a Linux user of that name already exists (a system user, or one left by an interrupted create), or the domain is already served by another account — the message names the username or the domain |
Pick another username or domain. Remove a leftover Linux user only when no account record exists for it and it is not a system or other user: check id <username> and ls -la /home/<username> first and copy anything you need, then userdel -r <username> (it deletes the home directory) and retry |
422 VALIDATION_ERROR: "Invalid username … Must be 3-16 lowercase alphanumeric, starting with a letter." |
Username rule ^[a-z][a-z0-9]{2,15}$; reserved names (root, admin, mysql, test, …) are refused as well |
3–16 characters, lowercase letters and digits only, starting with a letter |
| Account uses more disk than its plan allows | Filesystem quota is not active on /: the plan's disk limit is recorded and shown, but setquota failed (agent.log: disk quota not applied for <username>) |
quotaon -p / and quota -u <username>; enable quota as described in Disk quota. Disk limits use filesystem quota on /; CPU, memory, process and I/O limits use the account's cgroup v2 slice (whost-<username>.slice) |
| Suspended account still receives mail | Suspension switches the account's mailboxes and forwarders off in the mail database, which Postfix and Dovecot read on every lookup — no reload is involved. If that step failed, agent.log shows Mailboxes of <username> not disabled: … |
Fix the cause named there, then unsuspend and suspend the account again (suspending an already suspended account changes nothing) |
Reseller can't create sub-account: 403 RESELLER_LIMIT |
A limit of the reseller (accounts, disk, bandwidth, domains, databases, e-mail or FTP accounts) would be exceeded; the message names it | Raise that limit on the reseller (Resellers page) or allow overselling |
Reset a broken account fully (operator): delete the account from the
Accounts page, then create it again. Deletion removes the Linux user and its
home directory, databases, FTP and mail accounts, DNS zones, SSL files, vhosts,
cron jobs, hosted apps and the account's local backups — copy any archive
you still need out of /var/whost/backups/<username>/ first. A cleanup step
that fails is logged (Failed to … for <username>) and the others still run.
6. Webserver (Nginx / Apache)
| Symptom | Action |
|---|---|
nginx -t fails after a settings save |
Vhost changes are tested before they stand: when the test fails, the agent puts the previous files back and the save reports the error. If nginx -t still fails, the cause is outside that change — its output names the file and line (vhosts live in /etc/nginx/sites-available/ on Debian/Ubuntu, /etc/nginx/conf.d/ on the RHEL family); fix it, then systemctl reload nginx |
| 502 Bad Gateway from PHP page | PHP-FPM pool down → systemctl status php<version>-fpm (Debian/Ubuntu, e.g. php8.3-fpm) or php<version without dot>-php-fpm (RHEL family, e.g. php83-php-fpm); restart that unit |
| 404 on a freshly-added subdomain | Vhost generated but Nginx not reloaded — happens if reload step in service code failed. Manual: systemctl reload nginx |
| ModSecurity blocking legitimate traffic | Find the rule ID in ModSecurity WAF → Audit Log (file: /var/log/modsecurity/audit.log). Switch the rule off server-wide in ModSecurity WAF → Rules, or exclude it for one account through the API (POST /api/v1/accounts/{username}/waf/rules/exclude, optionally for one URI) |
| LiteSpeed WebAdmin console won't load | The console listens on port 7080 (webserver.ols_admin_port), which the installer's firewall does not open — allow it for your own address only. LiteSpeed Web Server (Enterprise) also needs a valid trial or license (OpenLiteSpeed needs none): activate the serial key under Plugins → LiteSpeed Web Server |
A web-scan ban is listed in fail2ban-client status whost-<user>-webscan but the address still reaches the site |
fail2ban only queues the request; systemctl status whost-http-ban.path must be active and journalctl -u whost-http-ban-apply shows the apply run — a webserver that rejected the list (configtest) reverts the change and says so there. whost-http-ban list shows what is applied; whost-http-ban sync re-applies by hand. fail2ban-client get whost-<user>-webscan actions must name whost-http-deny (an agent restart repairs a jail that lost it). OpenLiteSpeed never denies the server's own address or 127.0.0.1. |
7. Databases (MariaDB)
agent.conf (mysql.root_password): the admin link uses it directly, the client link uses it to create a short-lived database user. If the root password was changed outside WHost, put the current one there and restart the agent. grep -i -e phpmyadmin -e pma /var/log/whost/agent.log shows the attempts.max_connections is reached. Raise Max Connections in Settings → General → Database Settings: the value is applied live and kept in 99-whost-tuning.cnf (/etc/mysql/mariadb.conf.d/ on Debian/Ubuntu, /etc/my.cnf.d/ on the RHEL family), which the next save from the panel rewrites — a hand edit there does not last.bind-address in 99-whost.cnf, same directory) and the firewall does not open 3306; the panel warns about this when the host is added. To allow remote connections, set bind-address to an address the client can reach, restart MariaDB, and open 3306 in the firewall only for the client addresses.The table is fullUsually a full disk: check df -h /var/lib/mysql and free space. WHost leaves innodb_data_file_path at the MariaDB default, which grows as needed.8. Email (Postfix / Dovecot / Rspamd)
mailq)mailq prints the reason next to each deferred message; the delivery log is /var/log/mail.log (Debian/Ubuntu) or /var/log/maillog (RHEL family) — usually a DNS / SPF / DKIM issue. Verify outbound port 25 is open (cloud providers often block).POST /api/v1/accounts/{username}/emails.whost-policyd); raise email_hourly_limit on the plan or on the account's own limits.9. DNS (PowerDNS)
| Symptom | Action |
|---|---|
POST /api/v1/accounts/{username}/dns/{domain}/records returns 500 DNS_ERROR |
PowerDNS API down or unreachable. systemctl status pdns ; journalctl -u pdns -n 50. |
Records added but dig returns NXDOMAIN |
Ask this server first: dig @<server-ip> <name>. If it answers, the domain's NS delegation at the registrar or a resolver's cached negative answer is the cause: wait for the negative TTL, or flush a resolver you run (rndc flush on BIND). A broken DNSSEC chain shows as SERVFAIL, not NXDOMAIN. |
| AXFR refused (intentional) | All zones refuse AXFR by default (WHost sets no allow-axfr-ips, so the PowerDNS default applies). Add specific peers via pdnsutil set-meta <zone> ALLOW-AXFR-FROM <ip>. |
| Zone serial not bumping after edit | PowerDNS moves the serial on each change the panel sends through its API; a save that changes nothing sends no change. To move it by hand: pdnsutil increase-serial <zone>. Zones are created as Native, so PowerDNS sends no NOTIFY to secondaries. |
10. SSL / Let's Encrypt
ssl_failed audit entry (webhook event ssl.failed)Check /var/log/letsencrypt/letsencrypt.log for the certbot trace. Usually DNS not pointing at the server.certificate (not PEM, expired or not yet valid — the first block must be the site's own certificate, not an intermediate), private_key (not PEM, passphrase-protected, or not the certificate's key) or ca_bundle. Check before upload: openssl verify -CAfile chain.pem cert.pem.systemctl status certbot.timer and journalctl -u certbot.service where the package ships the timer (Debian/Ubuntu); elsewhere (RHEL family) the installer adds a daily root cron line — crontab -l | grep certbot. Log: /var/log/letsencrypt/letsencrypt.log.11. Backups / restore
0 bytes sizeDisk pressure where backups are written. Check df -h /var/whost/backups (backup.local_path; the working copy is built inside the same folder, not in /tmp) and the run's warnings in agent.log.agent.log carry the reason. Fix the cause and run the same restore again. For single files, extract them from the archive by hand (tar -xzf /var/whost/backups/<username>/<archive> -C <empty directory>). Take a fresh backup before a restore when the current state may still be needed.agent.log says Backup scheduler started or Backup scheduler disabled by configuration; while an update installs, due runs wait (Scheduled backups wait: an update is being installed).12. Python / Node app hosting
pip install (or npm install) hangsThe install runs as the account's user from the agent, which stops it after 10 minutes (pip) or 15 minutes (npm); a server that cannot reach the package index waits until then. Check outbound HTTPS: curl -sI https://pypi.org/simple/, curl -sI https://registry.npmjs.org/./home/<user>/logs/<app>-error.log for Python, <app>-stderr.log for Node) → fix code or settings → Restart. systemctl status whost-pyapp-<user>-<app> (Node: whost-nodeapp-<user>-<app>) shows the unit's exit status.PYTHON_APP_LIMIT / NODE_APP_LIMITThe plan's max_python_apps / max_node_apps is reached (0 = unlimited); raise it or remove an unused app. Apps get no TCP port — each listens on a Unix socket under /run/whost/<user>/ — so there is no port pool to run out of.systemctl status whost-pyapp-<user>-<app> (Node: whost-nodeapp-<user>-<app>) and ls -l /run/whost/<user>/<app>.sockdmesg | grep -i kill. A Node app has its own memory limit (max_memory_mb on the app, capped by the plan's node_max_memory_mb); a Python app runs within the account's memory limit from the plan — raise the one that applies.Deep dive: docs/developer/python-apps.md.
13. Webhooks
connection refusedSubscriber URL stale or down. Confirm DNS, port, TLS. Use Send test event from the admin panel.HMAC-SHA256(secret, "{timestamp}.{raw_body}"), with the timestamp from X-WHost-Webhook-Timestamp, compared with X-WHost-Webhook-Signature (sha256=<hex>). Use SDK WHost\Http\WebhookVerifier or copy the recipe from docs/developer/webhooks.md.dead_letter and stays there until retried by hand.GET /api/v1/system/webhooks/events lists the catalog. If your event isn't there, it's intentional (read-only audit).Deep dive: docs/developer/webhooks.md.
14. Updates (WHost + OS packages)
UPDATE_INSTALL_FAILED after clickHistory tab shows error + auto-rollback completed. Read /var/log/whost/agent.log around that time for the root cause.UPDATE_IN_PROGRESS blocks re-installAn install is running: the flag lives in the agent's memory from the install's first step until the restart that ends it (backups and restores are refused meanwhile). There is no lock file to remove; wait for the install to finish — the Updates page shows its progress.UPDATE_ALREADY_LATEST for a known newer releaseInstalling acts on the last update check; re-check with Updates → Check for Updates (GET /api/v1/system/update/check). A release on the beta channel is offered only to a server set to that channel (Updates → Settings).apt / dnf, unattended-upgrades or dnf-automatic) may hold the package lock: ps -C apt,apt-get,dpkg,dnf -o pid,etime,cmd (Debian/Ubuntu also lsof /var/lib/dpkg/lock-frontend). Let it finish, then retry. Do not kill a running dpkg or rpm — an interrupted dpkg needs dpkg --configure -a — and do not restart the agent while its OS upgrade runs: the package manager runs inside the agent's service and would be stopped with it.agent.log says why a release was skipped (does not match auto_update_type) or put off (Auto-install of … deferred, e.g. while backups run).Deep dive: docs/operator/update-flow.md.
15. Frontend (admin / client panel)
BUNDLE_MISMATCH on every API callThe page carries the fingerprint of another agent build (typically a tab opened before an update); the panel reloads itself once when it gets this code. If it keeps coming, hard-refresh; if it still does, systemctl restart whost-agent — the agent writes its build's fingerprint into the panel's /config.json at every start. SDK / HMAC clients are exempt and never see this.whost_locale), not in a cookie, and a language chosen on the Profile page is applied again at every sign-in. Pick the language from the switcher, or set the profile language to "Automatic (keep the language I sign in with)".error_code + message JSON./var/log/whost/frontend-debug.log stays empty even with frontend.debug_log: true in agent.conf. Use the browser's DevTools (Console and Network tabs) instead.16. Common error codes → action
The error_code field is the stable identifier; map any handling on it (not the localised message).
| Code | HTTP | Operator action |
|---|---|---|
AUTH_FAILED |
401 | Verify API key + signature (a malformed timestamp is refused here too) |
AUTH_EXPIRED |
401 | HMAC: the request timestamp is more than 300 s from the server clock — timedatectl, enable NTP. Panel: the session ended (timeout, a newer sign-in, a password change) — sign in again |
BUNDLE_MISMATCH |
403 | The page was loaded for another agent build — reload (the panel does it once by itself); see Frontend |
CROSS_ORIGIN_DENIED |
403 | A cookie-authenticated write arrived with a foreign Origin. Expected when a page outside the panel posts to the API; if the panel itself is served under a second hostname, add it to security.trusted_origins |
PAYLOAD_TOO_LARGE |
413 | Body over the cap: 1 MB on the sign-in and sign-out routes under /api/v1/auth/, 10 MB elsewhere, 300 MB on upload endpoints |
INVALID_JSON |
400 | Request body was empty or not valid JSON |
LICENSE_NOT_ACTIVATED / LICENSE_SUSPENDED |
403 | Sign in: the license page (/admin/license) opens. Use Verify Now first; enter a key only for a new installation or after a reissue (see License-related issues) |
RATE_LIMIT |
429 | Honour Retry-After header; if persistent, raise the matching tier in the rate_limit block of agent.conf and restart the agent |
PERMISSION_DENIED |
403 | Caller authenticated but lacks scope — check reseller ACL |
PATH_TRAVERSAL |
403 | Client attempted .. or null byte in a path; usually a partner bug |
IDEMPOTENCY_KEY_REUSED |
409 | Same key sent with different body — partner regenerates UUID per logical op |
VALIDATION_ERROR |
422 | details carries field-level reasons; partner fixes payload |
UPDATE_IN_PROGRESS |
409 | Wait for the running install; it ends with the agent's restart (nothing to remove by hand) |
OS_UPGRADE_IN_PROGRESS |
409 | An OS package run started from the panel is still going; wait for it |
SYSTEM_ERROR |
500 | Generic — the trace is in /var/log/whost/agent.log (Unhandled exception) |
Full catalog: docs/developer/api-reference.md → Appendix A.
17. When to escalate
Escalate to WISECP support with a support ticket from your wisecp.com client area (or, when you cannot open one, an e-mail to [email protected]) and add the following bundle:
- What happened, when — UTC timestamp, the action you took, the visible symptom.
- Agent log slice —
tail -n 2000 /var/log/whost/agent.log > agent.log.txt, plusjournalctl -u whost-agent --since '30 minutes ago' > agent-journal.txtfor the start / stop events - License state — use the configured
license.state_filepath (default/etc/whost/license.state); redact secret and key fields before sharing any excerpt. - Version —
curl -k https://127.0.0.1:2000/health(returns version + ok status). - OS —
cat /etc/os-release | head -3. - Reproduction — minimal steps that re-trigger the symptom.
Never share /etc/whost/agent.conf raw — it contains decrypted secrets at rest. If a config snippet is needed, redact the password, secret, key fields.
For security-sensitive issues (suspected RCE, license bypass, multi-tenant break), email [email protected] and do not report them in public.
Our support team is here around the clock for anything you can't find above.