Getting started with CertiStack Community
Community Edition runs from a dedicated controller host, not from a Proxmox VE hypervisor. Use a Linux management VM or the unprivileged controller container. Each run transfers a temporary node worker to the selected PVE node for read-only PBS mapping, QMP, sandbox probing, and cleanup; the controller keeps plans, credentials, signing keys, and reports.
If you want the shortest path to a first signed report, follow the quickstart and come back here for the details it skips.
Prerequisites
- Linux amd64 management host, or Docker/Podman for the controller image.
The controller uploads its own executable as the node worker, so it must
match the PVE node's architecture;
runrefuses a node it cannot execute on. Releases publish amd64 only; an arm64 build from source (make build-linux-arm64) runs the offline commands (validate,keygen,verify-report) but cannot drive an amd64 Proxmox VE node. - The OpenSSH client (
sshandscp) on the controller: it reaches the node by running them. The controller image includes them. - An amd64 Proxmox VE 9.x node with
qemu-img,proxmox-backup-clientandip, plusnftwhen a plan has wire probes, paired with PBS 4.x: the pair the release's signed matrix evidence covers. Proxmox VE 8.x with PBS 3.x, and any other pair, is not supported until it has its own signed matrix evidence, even where the tools install. libguestfs-toolson a node for isolated wire probes (the default COW-only Linux network recovery path; setnetwork_recovery.mode: preserveonly to opt out)- A file-based PVE storage for the disposable COW overlays, with
imagescontent and at least 5 GiB free on each node, named instorage.overlay_storage(or--overlay-storage). Stock Proxmox VE only attaches disks that belong to a storage, so without it the sandbox VM cannot be created. A node-local directory storage is enough:pvesm add dir certistack-scratch --path /var/lib/vz/certistack-scratch --content images --shared 0 - PVE API token and read-only PBS credential stored only on the controller
- SSH identity and an approved
known_hostsentry for every PVE node; see Set up SSH access to the node - Go 1.26.8+ to build from source
The 5 GiB storage check is a minimum reserve, not a per-run capacity estimate.
Allow additional room for guest writes, firmware copies, and any full disk
copies selected by the default copy_before_boot: auto policy. Auto copying
requires known disk separation from the cluster database and enough free
space for the disks plus that reserve. storage.max_copy_gib: 0 leaves copying
uncapped. Review the storage decisions with certistack plan before a run;
see the storage reference.
Build the CLI
git clone https://github.com/cliftcloud/certistack.git
cd certistack
make build-linux # bin/certistack-linux-amd64
make build-linux-arm64 # bin/certistack-linux-arm64 (offline commands only)
The standalone binary runs on the controller and is transferred only
temporarily to PVE during a job; it is not installed persistently on the
hypervisor. Build the unprivileged OCI image with make container, or use the
Packer VM-template recipe in deploy/appliance.
go install github.com/cliftcloud/certistack/cmd/certistack@latest also
works. A binary built that way reports its module version but no commit, so
it is fine for evaluation and cannot be used for release evidence; use the
signed release artifact or a Makefile build of a clean tagged checkout for
that.
Verify, install, upgrade, and roll back a release
For a published release, download certistack-linux-amd64,
CONTAINER-IMAGE.txt, SHA256SUMS-bin.txt, and SHA256SUMS-bin.txt.asc from the same official
release. First obtain the Clift Cloud release public
key through a separate trusted channel and compare its fingerprint out of
band. Then verify the signature and both assets before making the binary
executable:
gpgv --keyring /secure/certistack/release-signing-key.gpg \
SHA256SUMS-bin.txt.asc SHA256SUMS-bin.txt
sha256sum --strict --check SHA256SUMS-bin.txt
Both the binary and the image reference must pass. Use the verified contents
of CONTAINER-IMAGE.txt when installing the controller container or appliance.
Install each verified version in its own controller-local directory and point the command symlink at the active version. Do not change it while a campaign is running or while an interrupted-run journal requires reconciliation:
version=v0.1.0
sudo install -d -m 0755 "/opt/certistack/releases/$version"
sudo install -m 0755 certistack-linux-amd64 \
"/opt/certistack/releases/$version/certistack"
sudo ln -sfn "/opt/certistack/releases/$version/certistack" /usr/local/bin/certistack
certistack version
Before an upgrade, retain the previous binary, controller configuration,
signing key, trusted report keyring, reports, and recovery journal; finish or
recover every active run first. Verify the new artifact, update the symlink,
run certistack version, then execute a known-safe positive plan and verify
its signed report. If that check fails, stop new work, repoint the symlink to
the previously verified version, and use certistack recover with the exact
protected worker environment. Never delete a journal, report, or temporary
worker workspace to force a rollback.
Configure credentials for one run
Use a controller secret manager or a mode-0600 environment file outside the repository; do not commit credentials. The PVE token must have only the permissions needed for the sandbox scope, and the PBS credential must be read-only. Keep the report signing key on the controller.
export PVE_URL="https://pve-node-01.example.com:8006"
export PVE_TOKEN_ID="certistack@pam!automation"
export PVE_TOKEN_SECRET="your-generated-token-secret"
# Required when PVE uses a private CA and its CA is not in the host trust store.
# Do not use PVE_TLS_INSECURE=true in a production lab.
export CERTISTACK_PVE_CA_SOURCE="/secure/certistack/pve-ca.pem"
export PBS_REPOSITORY="automation@[email protected]:8007:datastore"
export PBS_PASSWORD="your-pbs-token-secret"
# The PBS certificate's SHA-256 fingerprint, as the PBS dashboard shows it
# (Show Fingerprint), when the certificate is not signed by a trusted CA.
export PBS_FINGERPRINT="aa:bb:cc:...:ff"
# Only when the backups are encrypted: the PBS encryption key, owner-only.
# For a PBS storage managed by Proxmox VE, copy /etc/pve/priv/storage/<id>.enc.
export CERTISTACK_PBS_KEYFILE="/secure/certistack/pbs-encryption.key"
# Only when that key is protected by a passphrase (its "kdf" is set);
# CertiStack refuses to start without it, and names a wrong one.
export PBS_ENCRYPTION_PASSWORD="your-key-passphrase"
export CERTISTACK_NODE="pve-node-01"
export CERTISTACK_SSH_HOST="pve-node-01.example.com"
export CERTISTACK_SSH_USER="certistack"
export CERTISTACK_SSH_IDENTITY="/secure/certistack/id_ed25519"
export CERTISTACK_SSH_KNOWN_HOSTS="/secure/certistack/known_hosts"
An encrypted backup needs its key for every restore. The controller copies
CERTISTACK_PBS_KEYFILE into the run's private worker workspace, which is
removed after the run, and the worker refuses a key other users can read.
Without a key, a run against an encrypted snapshot stops before restoring
anything and names the missing setting. The key is passed only for
encrypted or signed snapshots, so one key file serves a plan that mixes
encrypted and unencrypted backups. Each VM record in the report states the
weakest source_crypt_mode of its disks (encrypt, sign-only or none)
and whether the key verified the manifest (source_signature_verified).
PVE_URL must be an HTTPS endpoint in production. CertiStack permits plain
HTTP only for a loopback lab/mock endpoint; PVE_TLS_INSECURE=true is likewise
for disposable testing only and requires the explicit
CERTISTACK_ALLOW_INSECURE_TLS_FOR_LAB=true acknowledgement. The controller
does not request passwordless sudo by default; use --ssh-sudo (or
CERTISTACK_SSH_SUDO=true in the environment file) only for a dedicated PVE
worker account on a staging or lab node. Because the controller uploads and
executes the worker, that SSH/sudo identity is root-equivalent even when its
sudoers rule matches the worker command; treat command restrictions as an
operational safeguard, not a least-privilege boundary.
Alternatively, put only the allowlisted CertiStack variables above in a
mode-0600 dotenv file and pass it directly to the controller. Values are read
as literals—CertiStack never shell-sources the file, evaluates substitutions,
or accepts arbitrary environment variables. Lines beginning with # are
comments; an inline # remains part of the value. Command-line flags override
file values, and the file replaces the process environment for every supported
variable. Relative controller paths (CA, signing key, SSH key, known hosts,
and report output) are resolved from the dotenv file's directory. The
configuration reference lists every variable.
Set up SSH access to the node
The controller runs the OpenSSH client with a key used for nothing else and a host key you have verified. Do this once for each node.
-
Create a key pair on the controller. It is separate from the report signing key. A run never prompts (
BatchMode=yes), so the key must work unattended. -
Authorize the public key on the node, for the account you will name in
CERTISTACK_SSH_USER(defaultcertistack). From the controller, while you can still sign in another way:ssh-copy-id -i /secure/certistack/id_ed25519.pub [email protected] -
Decide how that account becomes root. The temporary worker maps PBS images and creates loop devices, so it runs as root. Either connect as
root, or use a dedicated account with passwordlesssudoand setCERTISTACK_SSH_SUDO=true(--ssh-sudo). PVE does not installsudoby default; it must be at/usr/bin/sudo. The controller runssudo -nwith/usr/bin/install,/usr/bin/systemd-run,/bin/sh,/usr/bin/rm,/usr/bin/cat,/usr/bin/touchand/usr/bin/test, and then the uploaded worker. A rule that allows a shell is full root, so treat the account as root-equivalent and use it on a lab or staging node first; see the access modes. -
Record the node's host key, and verify it. Strict host-key checking is always on, and there is no switch to disable it. Compare the fingerprint the node reports about itself with the one you collect over the network before you keep the entry:
# On the node: ssh-keygen -lf /etc/ssh/ssh_host_ed25519_key.pub # On the controller: ssh-keyscan -t ed25519 pve-node-01.example.com > /secure/certistack/known_hosts ssh-keygen -lf /secure/certistack/known_hostsKeep
known_hostsonly if the two fingerprints match. For an SSH port other than 22, add-p PORTtossh-keyscan, which records the[host]:PORTform thatsshlooks up. -
Try the connection the way CertiStack makes it. It should print the node's architecture (
x86_64) and nothing else;runrefuses a node whose architecture differs from the controller build's:ssh -o BatchMode=yes -o StrictHostKeyChecking=yes \ -o UserKnownHostsFile=/secure/certistack/known_hosts -o IdentitiesOnly=yes \ -i /secure/certistack/id_ed25519 [email protected] uname -m
Make the key, known_hosts and the environment file owner-only on the
controller; see the configuration reference.
Write a first plan: certistack init
init writes a first plan from what the node and its PBS repository hold,
and changes nothing. Without --vm or --all it lists the VMs that have a
complete backup. With them, it reads each chosen VM's newest backup as a
run does and writes a commented plan that validates:
certistack init --env-file /secure/certistack/controller.env
certistack init --env-file /secure/certistack/controller.env \
--vm 100 --vm 101 --key /secure/certistack/controller-signing.ed25519 \
-o plans/nightly.yaml
The plan it writes has:
- sandbox VM IDs from 9100 up (
--first-vmid) and SDN IDs (csZoneN,csVnetN) that PVE does not use. It also avoids the IDs of the other plans in the same directory, which PVE only sees while they run. - the node's usable overlay storage with the most room, preferring local
storage over shared (
--overlay-storageto choose); - the PBS namespace (
--namespace, elsePBS_NAMESPACE); - one tier, each VM restored from its newest backup with its backed-up configuration;
- a first probe per VM, chosen from its backup:
| Guest (from the backup) | First probe |
|---|---|
| Linux with a NIC | SSH (TCP 22) on the address the run gives it on the isolated VNet |
| Windows with the guest agent | sc query EventLog through the agent |
| Another guest with the agent | the agent answering (ping) |
| No guest agent | the VM running (qmp), with a comment on adding a probe to its own address |
A VM whose backup a run could not read (for example an encrypted backup
without its key) is left out, with the reason in the plan and in init's
output. So is a UEFI VM when the node has no overlay storage. init never
overwrites a file without --force.
The first probes prove the guest came up, not that it does its job. Change
each probe marked CHANGE to the service the VM is for, then preview the
run with certistack plan.
Validate and run a plan
Start from a plan init wrote, or from a copy of
examples/minimal-single-vm.yaml
with the marked values changed. Every accepted plan field is listed in the
test-plan reference.
certistack validate examples/minimal-single-vm.yaml
certistack run examples/minimal-single-vm.yaml \
--node pve-node-01 --ssh-host pve-node-01.example.com --ssh-user certistack \
--ssh-identity /secure/certistack/id_ed25519 \
--ssh-known-hosts /secure/certistack/known_hosts
To check a plan directory before a campaign, pass the directory in place of
the file. Immediate .yaml and .yml files are validated as one set.
run performs the authoritative node preflight automatically before it maps a
PBS image, creates an overlay, or changes SDN state. The doctor command is a
node-local diagnostic: it checks /dev/kvm, root access, PVE tooling, scratch
space, and optional PVE API connectivity. Given a plan, it also reads each
VM's backup as a run would, without creating anything. It reports per VM the
snapshot, firmware, restored disks, encryption and backup consistency. It
fails a VM the run would refuse: an encrypted backup without its key, the
wrong key, a UEFI VM without overlay_storage, or a plan disk the backup
lacks. A crash-consistent backup is a warning, or a failure for a VM whose plan
sets require_quiesced_backup: true. Do not run it in the unprivileged
controller container. If a run is blocked and you need to diagnose the target
host, run it on the affected PVE node with the same protected environment used
by the temporary worker:
sudo /path/to/certistack doctor /var/lib/certistack/workers/<run-uuid>/plan.yaml \
--env-file /var/lib/certistack/workers/<run-uuid>/runtime.env
Do not retain or copy that worker environment after the investigation; it contains per-run credentials and should remain inside the protected workspace.
Preview a run: certistack plan
plan shows, step by step, what run would do with a plan on a PVE node,
and changes nothing. Run it from the controller exactly as you run run,
before a first run or for a change review:
With an SSH host (--ssh-host or CERTISTACK_SSH_HOST in the environment
file), a temporary worker is staged on the node as for a run. It works the
plan out there, returns it, and is removed; nothing is kept on the node.
Without one, plan runs on the node itself, as root, like doctor.
It lists, in the order the run takes them:
- Before the run: what recovery of an earlier, interrupted run would delete first, and the retained worker workspaces whose credentials a run would remove.
- Checks: guest-network tools, overlay storage, and memory admission.
- Sandbox network: the SDN zone, VNet and subnet it creates (no gateway, no SNAT), the probe address, and its containment table.
- Per VM:
- the resolved snapshot, read-only;
- each disk's PBS map and the overlay it creates;
- the integrity scan, with the disk's size;
- whether the disks are copied before boot, and why;
- firmware copies and guest network recovery;
- the sandbox VM's exact PVE parameters.
- Probes, tier by tier.
- Teardown: everything it created, deleted in reverse, and the evidence it keeps.
Backups are only read. Source VMs and the production network are never
touched. It lists every warning the run would give, and every refusal, then
exits non-zero if the run would refuse the plan. --json prints the same
plan as JSON, to attach to a change request. File names that carry the run's
start time, and the run ID in the VM's tags, are fixed only when the run
starts.
During a run, the controller prints a concise elapsed-time timeline rather than raw node logs. Each line is emitted by an actual controller or worker phase; the final success line is printed only after the controller has saved and signed the local evidence.
[00:00:01] INFO [Admission] Host admission check: RAM 41.2% (max 85.0%), IO wait 0.8% (max 12.0%) → ADMISSION GRANTED
[00:00:03] INFO [Storage] Mounted PBS snapshot vm/100/... (scsi0) read-only via /dev/loop2.
[00:00:18] INFO [Stage 1] QMP hypervisor handshake established → VM RUNNING.
[00:00:32] INFO [Stage 2] HTTPS GET https://192.0.2.10:8443/healthz → PASSED (4ms; status=200).
[00:00:37] SUCCESS DR validation plan passed | duration: 34.1s | Ed25519 certificate: /secure/certistack/reports/...
Verify the report
The command writes signed JSON evidence under the report directory. That JSON is the deliverable: verify it against a trusted-signer keyring; a report cannot establish trust in its own embedded public key:
certistack verify-report \
--keyring /secure/certistack/trusted-signers.pub \
/var/lib/certistack/reports/<plan-id>/<report-id>.json
Automation can obtain a safe, machine-readable verified summary without parsing the human report output:
certistack verify-report --format json \
--keyring /secure/certistack/trusted-signers.pub \
/var/lib/certistack/reports/<plan-id>/<report-id>.json
verify-report exits 0 whenever the signature is valid and the signer is
trusted, including for a correctly signed report of a failed recovery. Read
all_passed and all_cleaned in the summary to learn the result. The
reports and verification guide documents every field, the key and
keyring formats, and how to verify a report without CertiStack.
The signed JSON is the machine-verifiable source of truth and is what an auditor should receive, together with the public key that verifies it. Human-readable HTML compliance binders and PDF certificates are rendered from that JSON by the Enterprise Edition daemon; the Community CLI does not render them. Framework labels in a report are declared scope, not a certification.
Recover from an interrupted run
The next controller-driven run automatically reconciles a prior node-worker
journal before it changes anything. The inspect and recover commands below
are node-local emergency diagnostics: execute them on the affected PVE node
with the same temporary worker credentials, not on the controller host.
They are intentionally not a way to inspect arbitrary VMs or sweep a host.
If a completed worker reports an incomplete teardown, the controller first replays that exact durable journal with the same temporary credentials. A successful retry is recorded in the signed teardown evidence before the workspace is removed. If the retry cannot be verified, CertiStack exits with an explicit cleanup error and retains only that protected worker workspace so the recovery command still has the credentials and journal it needs. Do not delete that workspace until recovery has succeeded. A run also exits non-zero if its exact temporary workspace or SSH staging directory cannot be proven removed; a passing service probe alone never creates a clean certificate.
# On the affected PVE node only.
certistack inspect --state-dir /var/lib/certistack
certistack recover --state-dir /var/lib/certistack \
--env-file /path/to/protected-worker/runtime.env
When a node worker dies, its supervisor on the node recovers the journal at
once. Every action recovery takes is shown in the controller's output with
the [Recovery] label: a sandbox VM destroyed, a PBS map process stopped, a
map released, a leftover file removed. The supervisor also keeps the log in
<state-dir>/logs/<workspace>.log on the node (mode 0700 directory, no
credentials), so an operator can still read it after the workspace is gone,
including when the controller had disconnected.
Next steps
- Run validations every night and get told when one fails: automation and scheduling.
- Look up a command, flag or variable: CLI reference and configuration reference.
- Copy a plan for a multi-tier Linux application or a Windows domain: example plans.
- Run many plans with signed, attributable evidence: lab campaigns.
- Qualify a lab before a release campaign: lab acceptance runbook.
- Deploy the controller as a container or appliance: controller deployment model.
For a licensed dashboard, API-driven runs, webhooks, and service-managed automation, use the separate Enterprise Edition.