Skip to content

Command-line reference

certistack is one static binary. The copy you run on the controller is also the temporary worker that run uploads to the Proxmox VE node for the length of one operation (plan and init do the same when you give them an SSH host), so every command below ships in every install. This page lists each command, its flags, and its exit status.

For what goes in a plan, see the test-plan reference. For environment variables, the --env-file format and default paths, see the configuration reference.

Where each command runs

Command Run it on What it changes What it is for
validate anywhere nothing Check a plan, or a directory of plans, against the schema
keygen anywhere the key files you name Create an Ed25519 signing key pair
verify-report anywhere nothing Verify a signed report against a key you trust
version anywhere nothing Print version and build information
completion anywhere nothing Print a shell completion script
init controller (with --ssh-host) or PVE node the plan file you name Write a first plan from the VMs that have PBS backups
plan controller (with --ssh-host) or PVE node nothing Show what a run would do
run controller temporary resources on the node, removed at the end Run a validation and sign the report
doctor PVE node nothing Check the node's prerequisites
inspect PVE node nothing Show a sandbox VM's sockets or the run journal
recover PVE node removes what the run journal records Reconcile an interrupted run

validate, keygen and verify-report have no dependency on Proxmox and also run on macOS and Windows. Everything that touches a node needs the Linux amd64 binary: the controller uploads its own executable to the node, so it must match the node's architecture.

Three more commands, node-run, node-plan and node-discover, are hidden. The controller runs them on the uploaded worker. They are not a supported interface: do not script against them.

Global flags

Flag Meaning
--debug Log at debug level on standard error. Use it, and keep the output, when you report a problem.
-h, --help Show help for the command.

Exit status

Scripts and schedulers should act on the exit status first and the output second.

Code Meaning
0 Success. For run, the signed report records all_passed: true and the teardown was verified.
1 The command ran and the answer is no. validate: a plan is invalid. plan: a run would refuse the plan. run: the validation finished, the signed report records a failure, and every temporary resource was verified removed.
2 Anything else: a usage error or unknown flag, a missing file, an SSH, PVE or PBS failure, a failed doctor check, a failed verify-report (including an invalid signature), or a run whose cleanup could not be verified.
130 The run was interrupted by SIGINT, SIGTERM or SIGHUP.

A run that exits 2 after creating resources may have left a worker workspace behind on purpose. The error message says how to recover it; see Recover from an interrupted run.

The first Ctrl-C during a run stops it and waits for CertiStack to clean up. A second Ctrl-C abandons the automatic cleanup and keeps the protected worker workspace so you can run recover yourself.

Treat every non-zero code as a failure. Use 1 to tell "the recovery did not pass" from 2, "the command itself did not complete".

Output streams

  • run prints a timeline of phases to standard error and, when it finishes, Controller report saved to: <path> to standard output.
  • plan, doctor, inspect, version, keygen and verify-report print their result to standard output. A failure goes to standard error.
  • validate prints a success summary to standard output and each invalid plan to standard error.

validate

certistack validate [plan.yaml | plan-directory] [flags]

Parses a plan against the v1 schema and checks every field constraint: CIDR and address rules, probe requirements, VM ID uniqueness and framework labels. It needs no credentials and touches nothing. Pass a directory to check every immediate .yaml or .yml file in it (not subdirectories, and never through a symbolic link); every invalid plan is reported before the command exits non-zero.

Flag Default Meaning
-p, --plan Plan file or directory, as an alternative to the argument. Give one or the other.
$ certistack validate my-first-plan.yaml
✓ Plan is valid
  Plan ID:        recovery-pve-node-01
  Name:           Recovery check of 1 VM(s) on pve-node-01
  Tiers:          1
  Ephemeral VMs:  1
  Probes:         1
  Network:        simple (zone=csZone1, vnet=csVnet1)
  Frameworks:     []

Exit status: 0 when every plan is valid, 1 when any plan is invalid, 2 when a path cannot be read.

keygen

certistack keygen [output_key_path] [flags]

Creates an Ed25519 key pair for signing reports. Run it once on the controller and keep the private key there.

Flag Default Meaning
--private Where to write the private key. The positional argument is the same thing.
--public Where to write the public key. Give it to write the file you will use as a trusted-signer keyring.
--force off Replace existing key files.

The private key is written with mode 0600 and the public key with mode 0644. keygen refuses to overwrite an existing file without --force, refuses a symbolic link, and checks both destinations before it writes either. See Reports and verification for the file formats and for how to rotate a key.

$ certistack keygen --private /secure/certistack/controller-signing.ed25519 \
    --public /secure/certistack/trusted-signers.pub
✓ Ed25519 signing key pair generated successfully
  Private Key: /secure/certistack/controller-signing.ed25519 (mode 0600)
  Public Key:  /secure/certistack/trusted-signers.pub (mode 0644)
  Public Hex:  c7c886ebaa4d7722003b19d141d622d635607097e838e9d1c88fe87424ba76a6

verify-report

certistack verify-report <report.json> [flags]

Checks a report's signature against a key you name. A report never establishes trust in its own embedded public key, so exactly one trust source is required.

Flag Default Meaning
--keyring A trusted-signer keyring: a JSON keyring or a list of hex public keys.
--public-key A file holding one hex-encoded public key.
--trusted-key One 64-character hex public key given inline.
--format text text for people, or json for a small machine-readable summary.
--allow-test-mode off Accept a report produced with the lab-only --skip-integrity mode. Without it such a report is rejected as production evidence.

Give exactly one of --keyring, --public-key and --trusted-key. A missing trust source, more than one, an unsigned report, an untrusted signer and a tampered report all fail with exit status 2.

Exit status 0 means the signature is valid and the signer is trusted. It does not mean the validation passed: a correctly signed report of a failed recovery verifies too. Read all_passed (and all_cleaned) in the output to learn the result.

The json summary contains no credentials, screenshots or guest output, so a pipeline can decide on it directly. Reports and verification documents each field.

version

certistack version
$ certistack version
certistack v0.1.0
  commit:    f362134a13d9
  built:     2026-09-30T17:25:42Z
  go:        go1.26.8
  os/arch:   linux/amd64

A binary built with go install reports its module version but no commit, so it cannot serve as release evidence. Use the signed release binary, or a build of a clean tagged checkout, when a run must qualify.

completion

certistack completion bash | zsh | fish | powershell

Prints a completion script for the shell you name. For example, for bash:

certistack completion bash | sudo tee /etc/bash_completion.d/certistack >/dev/null

init

certistack init [--vm ID]... | --all [flags]

Reads what a first plan needs and changes nothing on the node or in PBS: the VMs that have a complete backup, what each VM's newest backup says about its guest (OS type, firmware, guest agent, NIC, disks), the node's storages for overlays, and sandbox VM IDs and SDN IDs that are free. It then writes a commented plan that validates.

Without --vm or --all, init only lists the VMs that have backups. With --ssh-host it reads from the controller through a temporary worker, as run does; without one it reads on the node itself, as root.

Flag Default Meaning
--vm Source VM ID to put in the plan. Repeat the flag or separate IDs with commas.
--all off Put every VM that has a backup in the plan.
-o, --output certistack-plan.yaml Plan file to write.
--force off Replace the plan file if it exists.
--plan-id recovery-<node> The plan's plan_id.
--name The plan's name.
--key CERTISTACK_SIGN_KEY Controller-local signing key path to write into the plan. One is required to write a plan.
--overlay-storage CERTISTACK_OVERLAY_STORAGE, else the node's roomiest usable storage PVE storage for the COW overlays.
--namespace PBS_NAMESPACE PBS namespace to read.
--first-vmid 9100 First VM ID to consider for sandbox VMs.
--node CERTISTACK_NODE, else this host's short name when run on the node PVE node name.
--scratch Remote scratch directory.
--state-dir /var/lib/certistack Node state directory.
--env-file Controller environment file.
SSH flags --ssh-host, --ssh-user, --ssh-port, --ssh-identity, --ssh-known-hosts, --ssh-sudo.

A VM whose backup a run could not restore (for example an encrypted backup without its key) is left out with the reason stated in the plan and in the command's output. init also keeps the sandbox VM IDs and SDN IDs of the other plans in the output directory out of the new plan, because PVE only sees them while those plans run. See Write a first plan for what it puts in the file.

plan

certistack plan [plan.yaml] [flags]

Works out what certistack run would do with the plan on the node and changes nothing. It reads what a run reads before it creates anything, prints every step in order, and exits non-zero when a run would refuse the plan. With --ssh-host it runs from the controller through a temporary worker; with none it runs on the node itself, as root.

Flag Default Meaning
-p, --plan Plan file, as an alternative to the argument.
--json off Print the plan as JSON, to attach to a change request.
--node CERTISTACK_NODE, else this host's short name when run on the node PVE node name.
--scratch CERTISTACK_SCRATCH_DIR, else the plan's storage.scratch_dir, else /var/lib/certistack/scratch Scratch directory a run would use.
--state-dir /var/lib/certistack Node state directory, whose journal a run would recover first.
--overlay-storage CERTISTACK_OVERLAY_STORAGE, else the plan's storage.overlay_storage PVE storage for the COW overlays.
--env-file Controller environment file.
SSH flags --ssh-host, --ssh-user, --ssh-port, --ssh-identity, --ssh-known-hosts, --ssh-sudo.

Every step is one line: a symbol, an action, what it acts on, and a parenthesised reason.

Symbol Meaning
✓ A check that passes
< Reads the backup; never writes it
+ Creates a resource the run owns; teardown deletes it
~ Changes a resource the run created
> Starts a VM or runs a probe
- Deletes at teardown, or before the run
= Evidence kept after the run
! A warning
✗ A refusal: the run would stop here
· Not reached, because an earlier step would have stopped the run

With --json, the output is one object: plan_id, node, mode (release, or test_skip_integrity for the lab-only mode), steps (each with phase, action, kind, target, and optionally detail, status and params), warnings, and refusals. The JSON carries the same refusals as the text output, so refusals being absent or empty means a run would not refuse the plan.

Exit status: 0 when a run would proceed, 1 when a run would refuse the plan, 2 when the plan could not be worked out (for example, SSH failed).

run

certistack run [plan.yaml] [flags]

Runs a validation from the controller. It authenticates to the node over SSH with strict host-key checking, uploads a temporary worker to a private workspace under the node's state directory, and has the worker map the PBS snapshot read-only, create the overlay and the isolated network, boot the sandbox VM and run the probes. The worker then tears everything down and returns unsigned evidence. The controller signs the report with your key and saves it locally. Nothing persistent is installed on the node and the signing key never leaves the controller.

The five connection variables PVE_URL, PVE_TOKEN_ID, PVE_TOKEN_SECRET, PBS_REPOSITORY and PBS_PASSWORD must be set, in the process environment or in the --env-file. A node name, an SSH host, an existing known-hosts file and a signing key (--key, or the plan's compliance.sign_key_path) are required too.

Flag Default Meaning
-p, --plan Plan file, as an alternative to the argument.
--node CERTISTACK_NODE PVE node that hosts the temporary worker. Required.
--key, --signing-key CERTISTACK_SIGN_KEY, else the plan's compliance.sign_key_path Controller path to the Ed25519 private key. The two flags are aliases.
--output CERTISTACK_OUTPUT_DIR, else /var/lib/certistack/reports Controller directory for signed reports. Must be a clean absolute path.
--state-dir CERTISTACK_STATE_DIR, else /var/lib/certistack Node directory for the recovery journal and worker workspaces.
--scratch CERTISTACK_SCRATCH_DIR, else the plan's storage.scratch_dir, else /var/lib/certistack/scratch Node-local private work directory.
--overlay-storage CERTISTACK_OVERLAY_STORAGE, else the plan's storage.overlay_storage PVE storage for the COW overlays. Required on stock Proxmox VE.
--env-file Controller environment file.
SSH flags --ssh-host, --ssh-user, --ssh-port, --ssh-identity, --ssh-known-hosts, --ssh-sudo.
$ certistack run my-first-plan.yaml --env-file /secure/certistack/controller.env
[00:00:01] INFO [Admission] Host admission check: RAM 41.2% (max 85.0%), IO wait 0.8% (max 12.0%) → ADMISSION GRANTED
[00:00:03] INFO [Storage] Mounted PBS snapshot vm/100/... (scsi0) read-only via /dev/loop2.
[00:00:18] INFO [Stage 1] QMP hypervisor handshake established → VM RUNNING.
[00:00:32] INFO [Stage 2] TCP 192.0.2.10:22 → PASSED (3ms).
[00:00:37] SUCCESS DR validation plan passed | duration: 34.1s | Ed25519 certificate: /var/lib/certistack/reports/recovery-pve-node-01/<report-id>.json
✓ Controller report saved to: /var/lib/certistack/reports/recovery-pve-node-01/<report-id>.json

Only one CertiStack operation can run on a node at a time. A second run is refused with another CertiStack operation holds the host lock and exits 2; see troubleshooting.

A failed validation still writes a signed report, and the command exits 1 when the teardown was verified. Verify it the same way as a passing report.

doctor

certistack doctor [plan.yaml] [flags]

Runs read-only preflight checks on the machine it runs on. It is a node-local diagnostic: run it on the PVE node, as root, not on the controller or in the controller container. run performs the authoritative preflight on the node by itself; use doctor to find out why a node is refused. With a plan, it also reads each VM's backup as a run would, without creating anything.

Flag Default Meaning
-p, --plan Plan to check as well, as an alternative to the argument.
--env-file Controller environment file, or a retained worker's runtime.env.
--scratch /var/lib/certistack/scratch Scratch directory to check.
--state-dir /var/lib/certistack State directory to check.
--overlay-storage the plan's storage.overlay_storage PVE storage for the COW overlays.
--json off Print the checks as JSON: {"ready": bool, "checks": [{"name", "status", "detail"}]}.

Each check has the status passed, warning, failed or skipped, shown as ✓, !, ✗ and -.

Check What it verifies
plan The plan validates (only when you pass a plan).
privileges The process runs as root.
kvm /dev/kvm is accessible.
qemu-img, proxmox-backup-client, ip The tool is installed and on the PATH.
nft nft is installed, when the plan has wire probes or there is no plan.
guestfish guestfish (package libguestfs-tools) is installed, when the plan uses guest network recovery.
pbs_keyfile, pbs_key_passphrase The PBS encryption key is a private key file, and its passphrase is set when it needs one (only when CERTISTACK_PBS_KEYFILE is set).
memory The host has the RAM the plan's largest concurrent tier needs, plus headroom, within the plan's RAM cap (only with a plan).
scratch_path The scratch directory exists and is private: a clean absolute path that is not a shared temporary directory, with mode 0700, owned by the caller, and no symbolic link in the path.
scratch_space The scratch directory has at least 5 GiB free.
stale_loops No loop device is left attached to a FUSE file that is no longer mounted.
pve_connectivity The PVE API answers and accepts the token (only when PVE_URL is set).
pve_privileges The token holds the privileges the plan needs on every path it touches, against the documented CertiStackRole.
overlay_disk The sandbox overlays do not share a disk with the cluster database.
backup_sources, backup_vm_<id> Each VM's backup is readable and restorable as a run would restore it (with a plan, on a node with proxmox-backup-client).

A node that has never hosted a run has no scratch directory yet, so scratch_path fails there. Create it once (sudo install -d -m 0700 /var/lib/certistack/scratch), or run a validation first; the temporary worker creates the directory when it runs.

To diagnose a blocked run on the node itself, run doctor with the protected environment of the exact worker workspace; see Validate and run a plan.

Exit status: 0 when every check passes or warns, 2 when any check fails.

inspect

certistack inspect [vmid] [flags]

A node-local diagnostic. With a VM ID, it connects to that VM's QMP and QGA sockets and reports whether each answers; use it on a sandbox VM while a run is active. Without one it prints the active run journal (active.json) as JSON.

Flag Default Meaning
--qmp /var/run/qemu-server/<vmid>.qmp QMP socket path.
--qga /var/run/qemu-server/<vmid>.qga QGA socket path.
--state-dir /var/lib/certistack State directory whose active.json to print.
$ certistack inspect 9001
Inspecting VM 9001...
  QMP Socket: /var/run/qemu-server/9001.qmp ... CONNECTED (Status: running, Running: true)
  QGA Socket: /var/run/qemu-server/9001.qga ... CONNECTED (Agent active)

recover

certistack recover [flags]

Reconciles an interrupted or crashed run on the node. It reads the run journal, confirms each recorded process by its kernel start time so a reused PID is never mistaken for the original, verifies that CertiStack owns each recorded VM and SDN object before it destroys it, stops a runner that is still hanging, and detaches only the loop devices the journal records. It never sweeps the host for things it did not record.

Run it on the affected node, with the protected environment of the worker that created the resources.

Flag Default Meaning
--state-dir /var/lib/certistack State directory holding active.json.
--env-file The retained worker's runtime.env, which supplies the PVE connection. Only the worker's variables are accepted.
certistack recover --state-dir /var/lib/certistack \
  --env-file /var/lib/certistack/workers/<run-uuid>/runtime.env

The next controller-driven run on the node performs the same recovery automatically before it changes anything, so recover is for when you need to clear a node without starting a new run. Do not delete a retained worker workspace before recovery has succeeded.


SSH flags

init, plan and run share these. The controller shells out to the OpenSSH client (ssh and scp), which must be installed on the controller.

Flag Default Meaning
--ssh-host CERTISTACK_SSH_HOST The node's SSH host name or address. Required for run; for init and plan its absence means "run on this node".
--ssh-user CERTISTACK_SSH_USER, else certistack The SSH account.
--ssh-port CERTISTACK_SSH_PORT, else 22 The SSH port.
--ssh-identity CERTISTACK_SSH_IDENTITY Absolute path to the private key.
--ssh-known-hosts CERTISTACK_SSH_KNOWN_HOSTS, else ~/.ssh/known_hosts Absolute path to a known_hosts file that already holds the node's approved host key. Strict checking is always on; there is no way to turn it off.
--ssh-sudo CERTISTACK_SSH_SUDO, else off Run the temporary worker through passwordless sudo, for a dedicated non-root account.

Connections are non-interactive (BatchMode=yes): SSH never prompts for a password or a passphrase, so the key must work unattended. The SSH identity that uploads and starts the worker is root-equivalent on the node, even with a sudoers rule that names the worker, so use a dedicated account on a lab or staging node. See Controller deployment model.