Troubleshooting & Operations Guide
This guide covers diagnostic procedures, common error resolutions, and operational maintenance for CertiStack on Proxmox VE.
1. Diagnostic CLI Commands
Inspect a Running Ephemeral VM
CertiStack includes an ad-hoc inspector to query the QMP hypervisor socket and QGA in-guest agent of any running VM:
Sample output:
Inspecting VM 9001...
QMP Socket: /var/run/qemu-server/9001.qmp ... CONNECTED (Status: running, Running: true)
QGA Socket: /var/run/qemu-server/9001.qga ... CONNECTED (Agent active)
Validate Plan Syntax Before Execution
For a plan set, pass its directory instead. CertiStack validates every
immediate .yaml/.yml file, shows each error, and returns a non-zero status
only after the complete set has been checked.
Capture CLI output
Run the CLI with --debug, a global flag on every command, and preserve the
plan ID, report ID, and state directory. The persistent daemon and its journal
are Enterprise Edition components. Every command and exit status is listed in
the CLI reference.
2. Common Issues & Resolutions
Issue 1: "admission denied: RAM usage XX% exceeds max YY%"
Symptom: The test run immediately halts or retries repeatedly with:
Cause: CertiStack’s admission controller detected that the hypervisor host is under heavy memory pressure. To prevent OOM-killing production instances, the DR simulation is deferred.
Remediation:
- Check current memory utilization:
free -morcat /proc/meminfo. - Ensure test plans configure memory ballooning for ephemeral VMs:
- If host memory is safely available, adjust the threshold in the plan:
Issue 2: "proxmox-backup-client map failed: out of loop devices"
Symptom: Mapping a backup image fails with loop device allocation errors.
Cause: The Linux kernel has exhausted available /dev/loopX device slots.
Remediation:
- Check existing mapped loop devices:
- Do not detach a loop merely because its number looks unused. Reconcile only the exact resources recorded in CertiStack's durable journal with the same controller credentials: Recovery verifies the recorded loop backing identity before it acts; it never sweeps or detaches arbitrary host loops.
certistack doctor(and every run's preflight) warns withstale_loopsabout a read-only loop whose backing file shows as/inlosetup -a: it is attached to a FUSE file that is no longer mounted, most likely a PBS map whose process died after something else detached its mount. Nothing can read such a loop, and nothing ties it to a run any more, so CertiStack names it and leaves it. Release it with the commands the warning gives, in their order.
When the warning also names a FUSE connection with requests waiting,
abort it first (echo 1 > /sys/fs/fuse/connections/<N>/abort). A map
daemon killed while one of its threads read its own file leaves that
connection open, and every open of the loop then hangs in the kernel:
losetup -d, udev, and PVE's vgs, which freezes the node's storage
status (pvestatd: status update time of minutes). Aborting the
connection fails those reads at once. Recovery does this itself for a
map it owns before it detaches the loop, and never waits more than 20 s
for an unmap or a detach that is stuck.
4. If the journal is already clean and capacity is genuinely insufficient,
have the PVE operator assess the kernel loop-device limit through normal
change control rather than altering a host during a validation run.
Issue 2a: LDAP base search returns noSuchObject
Cause: An explicit base_dn is a business assertion. The restored
directory is reachable, but that exact DN is absent from the snapshot or is
not anonymously readable. CertiStack's failure includes Root-DSE naming
contexts advertised by the restored server when they are available.
Remediation:
- If that DN is required for the application, retain the explicit assertion and investigate the snapshot or directory data; do not weaken the plan to make a failed business check pass.
- If the requirement is only that a heterogeneous directory service restored
correctly, use
base_dn: auto. CertiStack then verifies a real naming context advertised by that restored server. - For an Active Directory domain controller, where the failure is usually
operationsErrorbecause AD refuses anonymous searches, usenaming_context: DC=example,DC=cominstead ofbase_dn. It reads only the Root DSE, which AD lets anonymous clients read. - Omit
base_dnonly for an LDAP-v3 listener/bind check, where directory-data visibility is intentionally outside scope.
- type: ldap
target_ip: 192.0.2.10
port: 389
base_dn: auto
timeout_sec: 15
description: Verify an advertised restored directory naming context
Issue 2b: LDAP listener times out after a successful sandbox boot
Symptom: The signed report shows a timeout to TCP/389 after QMP readiness
and full mapped-image verification. This is different from noSuchObject: the
worker could not establish a connection to the restored listener at all.
Meaning: The guest booted, but the evidence does not yet distinguish a non-started directory daemon from a guest-local firewall or another local service dependency. A timeout is not proof that the isolated SDN boundary is wrong, and it must not be converted into a passing plan by weakening the LDAP assertion.
Remediation:
- Verify the failed JSON report with the trusted keyring and confirm
all_cleaned: truebefore examining any failed sandbox. Preserve the report, plan digest, and framebuffer artifact as the failure record. - Review the report's Failed-VM Network Diagnostic and forensic exhibit. When permitted by the guest policy, it includes interfaces, routes, listeners, failed systemd units, and NetworkManager state.
- If the guest agent prohibits diagnostic command execution, use an approved
lab-console or guest-operator session to inspect
systemctl status slapd,ss -lntup, and the guest firewall. Do not enable arbitrary QGA command execution or add a permissive probe merely to make the recovery pass. - Correct the service, dependency, or firewall configuration in the synthetic source fixture, take a new approved PBS snapshot, regenerate its plan digest, and repeat the individual case before returning it to a fleet run.
Issue 2c: A restored domain controller answers LDAP and Kerberos but not DNS
Symptom: A Windows domain controller passes its LDAP (389) and Kerberos (88) checks, but DNS queries on 53 (UDP, TCP or SRV) time out.
Meaning: The DNS Server service did not start in time. On a busy PBS, the sandbox's first reads are slow, and Windows' service manager gives up on services that miss its 30-second start deadline; see Issue 2e. LDAP and Kerberos come from the directory service, which started, so they answer. It is not Active Directory waiting for replication partners: a DC restored on its own from local disks starts DNS within a minute, without partners.
Remediation: As for Issue 2e. A domain's DCs can be restored in dependency order (the root DC first) or together.
A DC whose DNS server is set to listen only on its own IPv4 address
(ListeningIPAddress) answers only on that address. The sandbox VM keeps
the source's address unless the plan's network recovery gives it another
one, so plan a DC's sandbox address to be its own.
Issue 2d: A service is "inactive" because its data mount was skipped
Symptom: A service on a data disk is inactive (not failed), and the
failed-VM diagnostic's [mounts not active] and [dependency and device
failures] sections name its mount, for example Timed out waiting for
device /dev/disk/by-uuid/... and Dependency failed for /var/lib/redis.
Meaning: The sandbox VM boots straight from the backup: blocks are read
from PBS as the guest first touches them, rather than copied beforehand as
a restore does. When PBS is busy, the guest's first read of a disk (udev
probing its file system) can take longer than a mount's device timeout. An
/etc/fstab entry with nofail and a short x-systemd.device-timeout is
then skipped, and every service that requires the mount stays inactive.
A run's integrity scan reads every disk into the page cache before the VM starts, and the boot reads from there. The conditions under which a boot stays cold are the ones in Issue 2e.
Remediation: As for Issue 2e. Otherwise, re-run when PBS is less busy, or run fewer VMs in parallel against the same datastore. If the guest's timeout is shorter than its own restores need, lengthen it in the source: it would bite a slow restore in production too.
Issue 2e: A Windows guest's agent or services never start in the sandbox
Symptom: A restored Windows Server 2016 or later guest answers its wire
checks (IIS, for example), but its guest agent never does (QEMU guest agent
is not running), or a service such as SQL Server is STOPPED. The guest's
System log shows Service Control Manager events 7009 ("A timeout was
reached (30000 milliseconds) while waiting for the ... service to connect")
and 7000 for those services, early in the boot.
Meaning: The sandbox VM boots straight from the backup. A Windows boot makes many small reads scattered over the disk, and each block that is not in the node's page cache waits for a round trip through the backup map to PBS. In the lab, a cold boot kept its mapped disk over 90% busy while moving about 1 MB/s; the same boot from the cache kept it under 1% busy. A cold boot is slow enough that services miss the service manager's 30-second start deadline, and more so when PBS is busy. Windows does not retry them, so they stay stopped. A real restore copies the disks first and boots from local storage, so it does not hit this.
Before each VM starts, a run's integrity scan reads all of its disks through the page cache, and the sandbox boots from that cache. The boot is cold again only when:
- the run used the lab-only
--skip-integrity, which skips the scan; - the node can't keep the tier's disks cached until its VMs have booted. The scan spreads its cache over the host's NUMA nodes, so the limit is their free memory together, less the tier's VM memory. It is one node's, the smallest, when a cpuset confines the worker to a node. A run warns when a tier's scanned disks exceed it;
- the CertiStack version predates this fix: QEMU then read the backup with direct I/O, past the cache.
A run copies a VM's disks to local storage first, and boots from the
copies, when they would not stay cached (copy_before_boot: auto). It
can't when overlay storage is on the cluster database's disk or is too
full, and warns then.
Remediation: Upgrade, and don't use --skip-integrity for Windows
guests. Put overlay storage on a disk of its own, with room for the largest
VM's disks, so a VM that doesn't fit in memory is copied first. Leave free
memory for at least the total size of a tier's disks, or split a tier the
run warned about. Otherwise, run Windows-heavy plans when PBS is not busy
(outside the backup window) and with fewer VMs in parallel against the
same datastore.
Don't change the guest's service timeouts or restart services from the
plan: the evidence has to show the guest as a restore would find it.
Issue 2f: A Windows sandbox VM's clock is hours off
Symptom: Inside a restored Windows guest the time is off by the PVE host's UTC offset, so certificate validity checks or Kerberos from another client can fail.
Meaning: Proxmox VE starts Windows guests with a real-time clock in the
host's local time (-rtc base=localtime), and Windows corrects it from its
time source. The isolated sandbox has no time source, so the guest keeps the
clock it booted with. A real restore without network behaves the same way;
it is not a CertiStack or backup defect.
Remediation: Keep clock-sensitive checks relative to the guest's own clock, or restore the domain's time source (a DC) in the same tier so the guest can synchronize on the isolated VNet.
Issue 3: "SDN zone or VNet creation error (403 Forbidden / 400 Bad Request)"
Symptom: CertiStack fails during the network provisioning stage.
Cause: The Proxmox VE API token lacks the narrowly required permissions on the cluster SDN layer, or PVE SDN is not enabled.
Remediation:
- Verify that the PVE deployment has its supported SDN components available.
- Grant a dedicated CertiStack token only the PVE API permissions required by
the documented lifecycle operations in the deployment model. Do not grant a
blanket
Administratorrole merely to make a validation run work. - Ensure the
zone_idandvnet_idin your plan do not conflict with existing production network interfaces.
Issue 3a: "CertiStack does not own or reuse pre-existing resources"
Cause: The requested sandbox zone or VNet already exists. CertiStack did not create it in the current run and therefore refuses to reuse or delete it.
Remediation:
- Inspect the exact PVE SDN resource and determine whether it belongs to a customer, another tool, or an interrupted CertiStack run.
- If it is an interrupted CertiStack run, reconcile that run from its recorded state with the same controller credentials.
- If ownership cannot be established, leave the resource unchanged and resolve the collision with the operator responsible for it. Retain the original plan configuration for the retry.
Issue 3b: "cannot safely reconcile unconfirmed allocation"
Cause: PVE may have received an allocation request, but CertiStack was interrupted before it could durably confirm the response. The journal retains the exact pending resource name and intentionally refuses to guess whether it was created.
Remediation:
- Inspect the retained record before changing anything:
- Have a PVE operator review only the resource named by
allocation_pendingand confirm its ownership from the CertiStack run ID and PVE task/audit history. Do not delete similarly named resources. - With the same protected worker credentials, retry reconciliation:
- A new validation remains blocked until the journal is marked
cleaned. If ownership cannot be proven, retain the journal and escalate through the platform change-control process.
Issue 3c: "Lost contact with node …"
Cause: The node stopped answering the controller's SSH connection for 60 seconds. It crashed, rebooted, or lost its network, or the network between the controller and the node failed.
What the controller does: It retries the node for up to 20 minutes and cleans up the run's temporary resources as soon as the node answers. It says when the node answers again, and whether it rebooted during the run (its boot ID changed). The run ends with an error and no report: nothing was validated. While a run is still going, "Cannot reach node … to renew the worker's lease" warns of the same problem early.
If the node does not answer within 20 minutes, the controller keeps the run's protected workspace on the node and prints the exact recovery command. Then:
- If the node kept running and only the network failed, its worker notices that the controller's lease has expired (90 seconds), stops the run and cleans up by itself. Nothing is left to do.
- If the node rebooted or crashed, its next CertiStack run recovers this run's resources before it starts. To recover them sooner, run the printed command on the node:
- Check the node itself (
journalctl -b -1, the out-of-band console) before the next run. A node that failed under a validation can fail under production load.
Whoever runs a recovery, the node keeps what it did in
/var/lib/certistack/logs/<run>.log (the state directory's logs).
Issue 3d: "the cluster's SDN configuration has changes that have not been applied"
Cause: Someone changed the cluster's SDN (a zone, VNet, subnet or
controller) and has not applied the change yet. Applying SDN in PVE is all
or nothing, on every node: applying CertiStack's sandbox network would apply
that change too, into production networking. So CertiStack creates and
removes its sandbox only under PVE's SDN configuration lock, which PVE
refuses while changes are pending. The run stops before it creates
anything; plan reports the same. The message names the pending zones,
VNets and controllers.
Remediation: Have the owner of the change apply it or discard it in PVE (Datacenter > SDN), then run again. Don't apply someone else's change just to let a validation run.
The same can happen at the end of a run, when a change was made while it ran. The run then deletes its sandbox zone and VNet but does not apply the deletion, and the report says so: their bridges stay, isolated, until the cluster's next SDN apply. If another tool holds the SDN lock, a run waits up to 2 minutes for it.
Issue 4: "Stage 3 probe marked as SKIPPED"
Symptom: In the final report, a QGA probe shows [SKIPPED].
Cause: The qemu-guest-agent daemon is either not installed or not running inside the guest operating system.
Behavior: A QGA probe is optional only when require_qga is false. In that
case, a skipped Stage 3 check leaves the configured Stage 2 wire assertion as
the pass/fail gate. When require_qga: true is set, missing or unavailable QGA
is a failed validation, never a skip.
Remediation (if in-guest testing is desired):
- On Linux guests:
apt-get install qemu-guest-agent && systemctl enable --now qemu-guest-agent - On Windows guests: Install the VirtIO Guest Agent from the VirtIO Windows driver ISO.
Issue 4a: QGA is connected but the report says guest-exec is unavailable
Cause: A hardened guest-agent policy can expose safe inventory RPCs while
deliberately denying guest-exec. CertiStack records the available interface
inventory and continues to treat the configured Stage 2 wire probe as the
service pass/fail assertion; diagnostic-command denial does not turn a failed
service probe into a pass.
Remediation: Keep the restrictive policy unless an approved guest-security
review says otherwise. Use the signed report, framebuffer artifact, and a
controlled guest-console investigation for service-state triage. If a plan
requires Stage 3 evidence, declare a supported constrained QGA probe and set
require_qga: true; do not place arbitrary shell text in a plan.
Issue 5: "Forensic Screendump Examination"
Symptom: A test plan failed, and an audit report contains a forensic exhibit.
How to extract and view:
- The screenshot is embedded in the signed JSON as a base64-encoded PNG
(
failure_artifacton the failed VM record, and on a probe result when that probe captured it). Extract it without modifying the report: The node worker also writes the same PNG next to its temporary state while the run is active; the JSON copy is the retained evidence. - The Enterprise Edition daemon renders the same exhibit inline in its HTML compliance binder; the Community CLI does not render reports.
- Inspect the image for:
- Linux kernel panics or systemd service failure traces.
- Windows Blue Screen of Death (BSOD) stop error codes (e.g.
INACCESSIBLE_BOOT_DEVICE). - GRUB bootloader prompt or missing partition errors.
Issue 6: "another CertiStack operation holds the host lock"
Symptom: A run stops at once with an error such as
node worker failed before producing evidence: acquire node worker host lock:
another CertiStack operation holds the host lock, and no report is written.
Cause: Each node allows one CertiStack operation at a time, a run or a
recovery. It takes an exclusive lock on the node (/run/certistack.lock) for
its whole duration, and a second operation is refused rather than queued.
Another validation, a scheduled job that started early, or a recovery is still
active on that node.
Remediation:
- Find the other operation: a
certistack runfrom any controller that targets this node, a scheduled job, orcertistack recoveron the node. On the node,pgrep -af certistack-nodelists a running worker. - Wait for it to finish, then run again. Stagger scheduled jobs that target the same node; see Rules for scheduled runs.
- A lock cannot go stale: it is released when the process that holds it exits,
and
/runis cleared at boot. If nothing is running and the error persists, inspect the journal withcertistack inspect --state-dir /var/lib/certistackon the node, and do not delete the lock file or the journal by hand.
plan and init only read and do not take the lock.
Issue 7: The run cannot start, or cannot reach the node
Symptom: run, plan or init stops before it creates anything, with a
message about a setting or a line from ssh. The controller passes on ssh's own
words, so the message names the cause.
| Message | Cause and fix |
|---|---|
controller secret environment variable PVE_URL is required (or PVE_TOKEN_ID, PVE_TOKEN_SECRET, PBS_REPOSITORY, PBS_PASSWORD) |
A required connection variable is missing. With --env-file, only the file's values count: a variable exported in the shell is discarded. See the configuration reference. |
unsupported controller environment variable "NAME", duplicate controller environment variable, invalid controller environment line N |
The environment file holds a name CertiStack does not accept, a repeated name, or a line without =. Only the documented variables are allowed. |
controller environment file must not be accessible by group or other users (mode 0640) |
Run chmod 600 on the file. |
PVE node name is required, invalid PVE node name |
Set --node or CERTISTACK_NODE to the node's name as PVE shows it: letters, digits, dots and hyphens. |
an absolute SSH known_hosts path is required, read SSH known_hosts file: ... |
The file at --ssh-known-hosts (default ~/.ssh/known_hosts) must exist and be given as an absolute path. |
Host key verification failed. |
The node's host key is not in the known_hosts file, or it changed. Add it after verifying the fingerprint; see Set up SSH access to the node. If it changed unexpectedly, find out why before you trust it: a reinstalled node has a new key, but so does an impostor. |
Permission denied (publickey,...) |
The node did not accept the key. Check the public key is in the account's authorized_keys, that --ssh-user names that account, and that --ssh-identity is the matching private key. |
Could not resolve hostname, Connection timed out, Connection refused |
The controller cannot reach --ssh-host on --ssh-port. Check the name, the port and any firewall between them. |
sudo: a password is required, /usr/bin/sudo: No such file or directory |
--ssh-sudo is set but the account cannot run sudo without a password, or sudo is not installed at /usr/bin/sudo. Connect as root or fix the sudoers policy. |
reports architecture "aarch64", but this controller is a linux/amd64 build |
The controller uploads itself as the worker, so its build must match the node's architecture. Releases are amd64, which matches an x86_64 Proxmox VE node; an arm64 node is not a supported target. |
node preflight failed: ... |
The node refused the run before it created anything. The text names the failed check; run certistack doctor on the node for the full list. |
Reproduce a connection problem outside CertiStack with the exact ssh command
in Set up SSH access to the node;
it fails for the same reason.