Command-line reference
certistack is one static binary. The copy you run on the controller is also
the temporary worker that run uploads to the Proxmox VE node for the length
of one operation (plan and init do the same when you give them an SSH
host), so every command below ships in every install. This page lists each
command, its flags, and its exit status.
For what goes in a plan, see the test-plan reference.
For environment variables, the --env-file format and default paths, see the
configuration reference.
Where each command runs
| Command | Run it on | What it changes | What it is for |
|---|---|---|---|
validate |
anywhere | nothing | Check a plan, or a directory of plans, against the schema |
keygen |
anywhere | the key files you name | Create an Ed25519 signing key pair |
verify-report |
anywhere | nothing | Verify a signed report against a key you trust |
version |
anywhere | nothing | Print version and build information |
completion |
anywhere | nothing | Print a shell completion script |
init |
controller (with --ssh-host) or PVE node |
the plan file you name | Write a first plan from the VMs that have PBS backups |
plan |
controller (with --ssh-host) or PVE node |
nothing | Show what a run would do |
run |
controller | temporary resources on the node, removed at the end | Run a validation and sign the report |
doctor |
PVE node | nothing | Check the node's prerequisites |
inspect |
PVE node | nothing | Show a sandbox VM's sockets or the run journal |
recover |
PVE node | removes what the run journal records | Reconcile an interrupted run |
validate, keygen and verify-report have no dependency on Proxmox and
also run on macOS and Windows. Everything that touches a node needs the
Linux amd64 binary: the controller uploads its own executable to the node, so
it must match the node's architecture.
Three more commands, node-run, node-plan and node-discover, are hidden.
The controller runs them on the uploaded worker. They are not a supported
interface: do not script against them.
Global flags
| Flag | Meaning |
|---|---|
--debug |
Log at debug level on standard error. Use it, and keep the output, when you report a problem. |
-h, --help |
Show help for the command. |
Exit status
Scripts and schedulers should act on the exit status first and the output second.
| Code | Meaning |
|---|---|
0 |
Success. For run, the signed report records all_passed: true and the teardown was verified. |
1 |
The command ran and the answer is no. validate: a plan is invalid. plan: a run would refuse the plan. run: the validation finished, the signed report records a failure, and every temporary resource was verified removed. |
2 |
Anything else: a usage error or unknown flag, a missing file, an SSH, PVE or PBS failure, a failed doctor check, a failed verify-report (including an invalid signature), or a run whose cleanup could not be verified. |
130 |
The run was interrupted by SIGINT, SIGTERM or SIGHUP. |
A run that exits 2 after creating resources may have left a worker
workspace behind on purpose. The error message says how to recover it; see
Recover from an interrupted run.
The first Ctrl-C during a run stops it and waits for CertiStack to clean up.
A second Ctrl-C abandons the automatic cleanup and keeps the protected
worker workspace so you can run recover yourself.
Treat every non-zero code as a failure. Use 1 to tell "the recovery did not
pass" from 2, "the command itself did not complete".
Output streams
runprints a timeline of phases to standard error and, when it finishes,Controller report saved to: <path>to standard output.plan,doctor,inspect,version,keygenandverify-reportprint their result to standard output. A failure goes to standard error.validateprints a success summary to standard output and each invalid plan to standard error.
validate
Parses a plan against the v1 schema and checks every field constraint:
CIDR and address rules, probe requirements, VM ID uniqueness and framework
labels. It needs no credentials and touches nothing. Pass a directory to
check every immediate .yaml or .yml file in it (not subdirectories, and
never through a symbolic link); every invalid plan is reported before the
command exits non-zero.
| Flag | Default | Meaning |
|---|---|---|
-p, --plan |
Plan file or directory, as an alternative to the argument. Give one or the other. |
$ certistack validate my-first-plan.yaml
✓ Plan is valid
Plan ID: recovery-pve-node-01
Name: Recovery check of 1 VM(s) on pve-node-01
Tiers: 1
Ephemeral VMs: 1
Probes: 1
Network: simple (zone=csZone1, vnet=csVnet1)
Frameworks: []
Exit status: 0 when every plan is valid, 1 when any plan is invalid, 2
when a path cannot be read.
keygen
Creates an Ed25519 key pair for signing reports. Run it once on the controller and keep the private key there.
| Flag | Default | Meaning |
|---|---|---|
--private |
Where to write the private key. The positional argument is the same thing. | |
--public |
Where to write the public key. Give it to write the file you will use as a trusted-signer keyring. | |
--force |
off | Replace existing key files. |
The private key is written with mode 0600 and the public key with mode
0644. keygen refuses to overwrite an existing file without --force,
refuses a symbolic link, and checks both destinations before it writes
either. See Reports and verification for the file formats
and for how to rotate a key.
$ certistack keygen --private /secure/certistack/controller-signing.ed25519 \
--public /secure/certistack/trusted-signers.pub
✓ Ed25519 signing key pair generated successfully
Private Key: /secure/certistack/controller-signing.ed25519 (mode 0600)
Public Key: /secure/certistack/trusted-signers.pub (mode 0644)
Public Hex: c7c886ebaa4d7722003b19d141d622d635607097e838e9d1c88fe87424ba76a6
verify-report
Checks a report's signature against a key you name. A report never establishes trust in its own embedded public key, so exactly one trust source is required.
| Flag | Default | Meaning |
|---|---|---|
--keyring |
A trusted-signer keyring: a JSON keyring or a list of hex public keys. | |
--public-key |
A file holding one hex-encoded public key. | |
--trusted-key |
One 64-character hex public key given inline. | |
--format |
text |
text for people, or json for a small machine-readable summary. |
--allow-test-mode |
off | Accept a report produced with the lab-only --skip-integrity mode. Without it such a report is rejected as production evidence. |
Give exactly one of --keyring, --public-key and --trusted-key. A
missing trust source, more than one, an unsigned report, an untrusted signer
and a tampered report all fail with exit status 2.
Exit status 0 means the signature is valid and the signer is trusted. It
does not mean the validation passed: a correctly signed report of a failed
recovery verifies too. Read all_passed (and all_cleaned) in the output to
learn the result.
The json summary contains no credentials, screenshots or guest output, so a
pipeline can decide on it directly. Reports and
verification documents each field.
version
$ certistack version
certistack v0.1.0
commit: f362134a13d9
built: 2026-09-30T17:25:42Z
go: go1.26.8
os/arch: linux/amd64
A binary built with go install reports its module version but no commit, so
it cannot serve as release evidence. Use the signed release binary, or a build
of a clean tagged checkout, when a run must qualify.
completion
Prints a completion script for the shell you name. For example, for bash:
init
Reads what a first plan needs and changes nothing on the node or in PBS: the VMs that have a complete backup, what each VM's newest backup says about its guest (OS type, firmware, guest agent, NIC, disks), the node's storages for overlays, and sandbox VM IDs and SDN IDs that are free. It then writes a commented plan that validates.
Without --vm or --all, init only lists the VMs that have backups. With
--ssh-host it reads from the controller through a temporary worker, as run
does; without one it reads on the node itself, as root.
| Flag | Default | Meaning |
|---|---|---|
--vm |
Source VM ID to put in the plan. Repeat the flag or separate IDs with commas. | |
--all |
off | Put every VM that has a backup in the plan. |
-o, --output |
certistack-plan.yaml |
Plan file to write. |
--force |
off | Replace the plan file if it exists. |
--plan-id |
recovery-<node> |
The plan's plan_id. |
--name |
The plan's name. |
|
--key |
CERTISTACK_SIGN_KEY |
Controller-local signing key path to write into the plan. One is required to write a plan. |
--overlay-storage |
CERTISTACK_OVERLAY_STORAGE, else the node's roomiest usable storage |
PVE storage for the COW overlays. |
--namespace |
PBS_NAMESPACE |
PBS namespace to read. |
--first-vmid |
9100 |
First VM ID to consider for sandbox VMs. |
--node |
CERTISTACK_NODE, else this host's short name when run on the node |
PVE node name. |
--scratch |
Remote scratch directory. | |
--state-dir |
/var/lib/certistack |
Node state directory. |
--env-file |
Controller environment file. | |
| SSH flags | --ssh-host, --ssh-user, --ssh-port, --ssh-identity, --ssh-known-hosts, --ssh-sudo. |
A VM whose backup a run could not restore (for example an encrypted backup
without its key) is left out with the reason stated in the plan and in the
command's output. init also keeps the sandbox VM IDs and SDN IDs of the
other plans in the output directory out of the new plan, because PVE only
sees them while those plans run. See Write a first plan
for what it puts in the file.
plan
Works out what certistack run would do with the plan on the node and
changes nothing. It reads what a run reads before it creates anything, prints
every step in order, and exits non-zero when a run would refuse the plan.
With --ssh-host it runs from the controller through a temporary worker; with
none it runs on the node itself, as root.
| Flag | Default | Meaning |
|---|---|---|
-p, --plan |
Plan file, as an alternative to the argument. | |
--json |
off | Print the plan as JSON, to attach to a change request. |
--node |
CERTISTACK_NODE, else this host's short name when run on the node |
PVE node name. |
--scratch |
CERTISTACK_SCRATCH_DIR, else the plan's storage.scratch_dir, else /var/lib/certistack/scratch |
Scratch directory a run would use. |
--state-dir |
/var/lib/certistack |
Node state directory, whose journal a run would recover first. |
--overlay-storage |
CERTISTACK_OVERLAY_STORAGE, else the plan's storage.overlay_storage |
PVE storage for the COW overlays. |
--env-file |
Controller environment file. | |
| SSH flags | --ssh-host, --ssh-user, --ssh-port, --ssh-identity, --ssh-known-hosts, --ssh-sudo. |
Every step is one line: a symbol, an action, what it acts on, and a parenthesised reason.
| Symbol | Meaning |
|---|---|
✓ |
A check that passes |
< |
Reads the backup; never writes it |
+ |
Creates a resource the run owns; teardown deletes it |
~ |
Changes a resource the run created |
> |
Starts a VM or runs a probe |
- |
Deletes at teardown, or before the run |
= |
Evidence kept after the run |
! |
A warning |
✗ |
A refusal: the run would stop here |
· |
Not reached, because an earlier step would have stopped the run |
With --json, the output is one object: plan_id, node, mode (release,
or test_skip_integrity for the lab-only mode), steps (each with phase,
action, kind, target, and optionally detail, status and params),
warnings, and refusals. The JSON carries the same refusals as the text
output, so refusals being absent or empty means a run would not refuse the
plan.
Exit status: 0 when a run would proceed, 1 when a run would refuse the
plan, 2 when the plan could not be worked out (for example, SSH failed).
run
Runs a validation from the controller. It authenticates to the node over SSH with strict host-key checking, uploads a temporary worker to a private workspace under the node's state directory, and has the worker map the PBS snapshot read-only, create the overlay and the isolated network, boot the sandbox VM and run the probes. The worker then tears everything down and returns unsigned evidence. The controller signs the report with your key and saves it locally. Nothing persistent is installed on the node and the signing key never leaves the controller.
The five connection variables PVE_URL, PVE_TOKEN_ID, PVE_TOKEN_SECRET,
PBS_REPOSITORY and PBS_PASSWORD must be set, in the process environment or
in the --env-file. A node name, an SSH host, an existing known-hosts file and
a signing key (--key, or the plan's compliance.sign_key_path) are required
too.
| Flag | Default | Meaning |
|---|---|---|
-p, --plan |
Plan file, as an alternative to the argument. | |
--node |
CERTISTACK_NODE |
PVE node that hosts the temporary worker. Required. |
--key, --signing-key |
CERTISTACK_SIGN_KEY, else the plan's compliance.sign_key_path |
Controller path to the Ed25519 private key. The two flags are aliases. |
--output |
CERTISTACK_OUTPUT_DIR, else /var/lib/certistack/reports |
Controller directory for signed reports. Must be a clean absolute path. |
--state-dir |
CERTISTACK_STATE_DIR, else /var/lib/certistack |
Node directory for the recovery journal and worker workspaces. |
--scratch |
CERTISTACK_SCRATCH_DIR, else the plan's storage.scratch_dir, else /var/lib/certistack/scratch |
Node-local private work directory. |
--overlay-storage |
CERTISTACK_OVERLAY_STORAGE, else the plan's storage.overlay_storage |
PVE storage for the COW overlays. Required on stock Proxmox VE. |
--env-file |
Controller environment file. | |
| SSH flags | --ssh-host, --ssh-user, --ssh-port, --ssh-identity, --ssh-known-hosts, --ssh-sudo. |
$ certistack run my-first-plan.yaml --env-file /secure/certistack/controller.env
[00:00:01] INFO [Admission] Host admission check: RAM 41.2% (max 85.0%), IO wait 0.8% (max 12.0%) → ADMISSION GRANTED
[00:00:03] INFO [Storage] Mounted PBS snapshot vm/100/... (scsi0) read-only via /dev/loop2.
[00:00:18] INFO [Stage 1] QMP hypervisor handshake established → VM RUNNING.
[00:00:32] INFO [Stage 2] TCP 192.0.2.10:22 → PASSED (3ms).
[00:00:37] SUCCESS DR validation plan passed | duration: 34.1s | Ed25519 certificate: /var/lib/certistack/reports/recovery-pve-node-01/<report-id>.json
✓ Controller report saved to: /var/lib/certistack/reports/recovery-pve-node-01/<report-id>.json
Only one CertiStack operation can run on a node at a time. A second run is
refused with another CertiStack operation holds the host lock and exits 2;
see troubleshooting.
A failed validation still writes a signed report, and the command exits 1
when the teardown was verified. Verify it the same way as a passing report.
doctor
Runs read-only preflight checks on the machine it runs on. It is a
node-local diagnostic: run it on the PVE node, as root, not on the controller
or in the controller container. run performs the authoritative preflight on
the node by itself; use doctor to find out why a node is refused. With a
plan, it also reads each VM's backup as a run would, without creating
anything.
| Flag | Default | Meaning |
|---|---|---|
-p, --plan |
Plan to check as well, as an alternative to the argument. | |
--env-file |
Controller environment file, or a retained worker's runtime.env. |
|
--scratch |
/var/lib/certistack/scratch |
Scratch directory to check. |
--state-dir |
/var/lib/certistack |
State directory to check. |
--overlay-storage |
the plan's storage.overlay_storage |
PVE storage for the COW overlays. |
--json |
off | Print the checks as JSON: {"ready": bool, "checks": [{"name", "status", "detail"}]}. |
Each check has the status passed, warning, failed or skipped, shown as
✓, !, ✗ and -.
| Check | What it verifies |
|---|---|
plan |
The plan validates (only when you pass a plan). |
privileges |
The process runs as root. |
kvm |
/dev/kvm is accessible. |
qemu-img, proxmox-backup-client, ip |
The tool is installed and on the PATH. |
nft |
nft is installed, when the plan has wire probes or there is no plan. |
guestfish |
guestfish (package libguestfs-tools) is installed, when the plan uses guest network recovery. |
pbs_keyfile, pbs_key_passphrase |
The PBS encryption key is a private key file, and its passphrase is set when it needs one (only when CERTISTACK_PBS_KEYFILE is set). |
memory |
The host has the RAM the plan's largest concurrent tier needs, plus headroom, within the plan's RAM cap (only with a plan). |
scratch_path |
The scratch directory exists and is private: a clean absolute path that is not a shared temporary directory, with mode 0700, owned by the caller, and no symbolic link in the path. |
scratch_space |
The scratch directory has at least 5 GiB free. |
stale_loops |
No loop device is left attached to a FUSE file that is no longer mounted. |
pve_connectivity |
The PVE API answers and accepts the token (only when PVE_URL is set). |
pve_privileges |
The token holds the privileges the plan needs on every path it touches, against the documented CertiStackRole. |
overlay_disk |
The sandbox overlays do not share a disk with the cluster database. |
backup_sources, backup_vm_<id> |
Each VM's backup is readable and restorable as a run would restore it (with a plan, on a node with proxmox-backup-client). |
A node that has never hosted a run has no scratch directory yet, so
scratch_path fails there. Create it once (sudo install -d -m 0700
/var/lib/certistack/scratch), or run a validation first; the temporary worker
creates the directory when it runs.
To diagnose a blocked run on the node itself, run doctor with the protected
environment of the exact worker workspace; see
Validate and run a plan.
Exit status: 0 when every check passes or warns, 2 when any check fails.
inspect
A node-local diagnostic. With a VM ID, it connects to that VM's QMP and QGA
sockets and reports whether each answers; use it on a sandbox VM while a run
is active. Without one it prints the active run journal (active.json) as
JSON.
| Flag | Default | Meaning |
|---|---|---|
--qmp |
/var/run/qemu-server/<vmid>.qmp |
QMP socket path. |
--qga |
/var/run/qemu-server/<vmid>.qga |
QGA socket path. |
--state-dir |
/var/lib/certistack |
State directory whose active.json to print. |
$ certistack inspect 9001
Inspecting VM 9001...
QMP Socket: /var/run/qemu-server/9001.qmp ... CONNECTED (Status: running, Running: true)
QGA Socket: /var/run/qemu-server/9001.qga ... CONNECTED (Agent active)
recover
Reconciles an interrupted or crashed run on the node. It reads the run journal, confirms each recorded process by its kernel start time so a reused PID is never mistaken for the original, verifies that CertiStack owns each recorded VM and SDN object before it destroys it, stops a runner that is still hanging, and detaches only the loop devices the journal records. It never sweeps the host for things it did not record.
Run it on the affected node, with the protected environment of the worker that created the resources.
| Flag | Default | Meaning |
|---|---|---|
--state-dir |
/var/lib/certistack |
State directory holding active.json. |
--env-file |
The retained worker's runtime.env, which supplies the PVE connection. Only the worker's variables are accepted. |
certistack recover --state-dir /var/lib/certistack \
--env-file /var/lib/certistack/workers/<run-uuid>/runtime.env
The next controller-driven run on the node performs the same recovery
automatically before it changes anything, so recover is for when you need to
clear a node without starting a new run. Do not delete a retained worker
workspace before recovery has succeeded.
SSH flags
init, plan and run share these. The controller shells out to the OpenSSH
client (ssh and scp), which must be installed on the controller.
| Flag | Default | Meaning |
|---|---|---|
--ssh-host |
CERTISTACK_SSH_HOST |
The node's SSH host name or address. Required for run; for init and plan its absence means "run on this node". |
--ssh-user |
CERTISTACK_SSH_USER, else certistack |
The SSH account. |
--ssh-port |
CERTISTACK_SSH_PORT, else 22 |
The SSH port. |
--ssh-identity |
CERTISTACK_SSH_IDENTITY |
Absolute path to the private key. |
--ssh-known-hosts |
CERTISTACK_SSH_KNOWN_HOSTS, else ~/.ssh/known_hosts |
Absolute path to a known_hosts file that already holds the node's approved host key. Strict checking is always on; there is no way to turn it off. |
--ssh-sudo |
CERTISTACK_SSH_SUDO, else off |
Run the temporary worker through passwordless sudo, for a dedicated non-root account. |
Connections are non-interactive (BatchMode=yes): SSH never prompts for a
password or a passphrase, so the key must work unattended. The SSH identity
that uploads and starts the worker is root-equivalent on the node, even with
a sudoers rule that names the worker, so use a dedicated account on a lab or
staging node. See Controller deployment model.