Changelog
All notable changes to CertiStack Community Edition are documented in this file. Enterprise Edition changes are not tracked here.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[Unreleased]
Added
- Documentation: a command-line reference (every command, flag, exit status
and
doctorcheck), a configuration reference (every environment variable, the environment file, and the files CertiStack keeps on the controller and the node), a reports and verification guide with a script that verifies a report without CertiStack, an automation and scheduling guide with a tested wrapper script andsystemdunits, example plans, and a FAQ. The changelog is now part of the documentation site. examples/three-tier-app.yaml, a Linux database, application and web stack restored in dependency order, andexamples/windows-domain.yaml, a Windows domain controller, SQL Server and file server.- Getting started explains how to set up SSH access to a node, and the troubleshooting guide covers the node's host lock and the errors that stop a run before it starts.
Fixed
- Documentation:
compliance.frameworksandcompliance.retention_daysare optional, and the first fouradmission_controlkeys are required rather than defaulted. Theqmpandscreendumpprobe types, thenetwork_recovery.mode: preserverequirement for Windows guests with wire probes, the OpenSSH client andnftprerequisites, and the fact thatverify-reportexits0for a correctly signed report of a failed validation are now stated.
[0.1.0] - 2026-09-27
First public release of the Community Edition under the GNU Affero General Public License v3.0.
Added
storage.max_copy_gibcaps what copy-before-boot writes in one run, all VMs together.autochecked only the free space, which on a shared scratch disk is not all the run's to take. A single 48 GiB Windows copy nearly fills the lab's proposed 50 GiB share. Over the cap,autodoesn't copy and warns why, andalwaysfails the VM.planshows it the same way.admission_control.max_concurrent_bootslimits how many sandbox VMs boot at once on the node. A tier starts its VMs one after another, so their first minutes of boot, when each guest writes most to its new overlays, overlapped. On a node whose overlay storage shares the cluster database's disk, that can delay corosync (D36). Before starting a VM, the run waits while that many VMs started withinboot_settle_sec(default 180 s), across tiers.planshows each start it may hold. The run's warning about the shared disk suggests it. It is off unless set.certistack initwrites a first plan from what the node and its PBS repository hold, and changes nothing. Without--vmor--allit lists the VMs that have a complete backup. With them it reads each VM's newest backup as a run does (OS type, firmware, guest agent, NIC, disks) and writes a commented plan that validates, with:- sandbox VM IDs and SDN IDs that PVE and the other plans in the directory do not use;
- the node's roomiest usable overlay storage;
- a first probe per VM that fits its guest: SSH for Linux on its sandbox
address, the Event Log service through the agent for Windows, the agent
for other guests with one, and the VM running for guests without one.
VMs a run could not restore are left out with the reason. From the
controller it reads through a temporary worker on the node, as
plandoes. run,doctorandplancheck the PVE token's privileges on every path a run of the plan touches (each VM, the SDN zone and VNet, the overlay storage, the node). They list every missing one at once, against the documentedCertiStackRole. A run used to stop at the first missing privilege it met, for examplePermission check failed (/sdn/zones/csLab70, SDN.Allocate), and a token missing several needed several attempts. A privilege the run works around is a warning. A test keeps the checked set equal to the role interraform/main.tf.certistack plan <plan.yaml>shows, step by step, what a run would do on the node, and changes nothing. It lists:- what recovery of an interrupted run would delete first;
- the checks;
- every resource the run creates, in order: SDN zone, VNet and subnet, probe address and containment table, and per disk the read-only PBS map, the overlay, and any copy made before boot, with the reason;
- each sandbox VM's exact PVE create parameters;
- the probes;
- what teardown deletes, and the evidence it keeps.
It also gives the warnings and refusals a run would give, exits non-zero
when a run would refuse, and prints JSON with
--json. It shares the run's own code for resolving the plan, building the VM configuration and deciding what to copy, so it shows what the run will really do. It sends PVE only reads. From the controller (an SSH host configured, as forrun), it stages a temporary worker on the node the same way, has it work the plan out there, and removes it. - Copy-before-boot. When a VM's disks would not all stay cached until it boots, the integrity scan also copies them, sparsely, next to their overlays. The overlays are then re-pointed at the copies, and the sandbox boots from local storage as a restore does. It is the same single read of PBS, with the same digest. Large Windows servers no longer boot partly cold on hosts whose NUMA nodes can't hold their disks. The copy releases the map's cache and its own written pages as it goes, so it doesn't crowd out another VM's cached boot.
copy_before_bootper VM:auto(default),alwaysornever.autonever copies onto the cluster database's disk, or without room for the disks plus 5 GiB, and warns when it can't copy.- Reports record
copied_before_bootand each disk'sboot_source. - A worker workspace that a controller kept for recovery no longer keeps its credentials for ever. The next run on the node, holding the host lock with a clean journal, removes the PVE/PBS environment, the PBS key and the worker binary from any workspace whose controller has been silent for 15 minutes. It keeps the logs and results, and leaves a note saying when and why. Workspaces still in use or being staged are left alone.
require_quiesced_backup: trueon a plan VM fails that VM, before anything is restored, unless its backup wasquiescedorpowered-off. Anunknownconsistency fails too, because it proves nothing.doctorfails such a VM the same way. Without the setting, a crash-consistent backup is still only a warning.- The pre-boot scan spreads its page cache over the host's NUMA nodes. The kernel caches a read on the node of the CPU that issued it and evicts from that node once it is full, even while others are free. So a scan left on one node could cache only that node's free memory. In the lab, a 48 GiB Windows VM lost part of its cache during its own scan on a host with two 63 GiB nodes. The scan now moves, every GiB, to the node with the most free memory, within the CPUs the worker may use. The same VM's cache split 42/58 over both nodes, and nothing was evicted.
- A warning when the disks a tier has scanned exceed what the node can keep cached: its NUMA nodes' free memory together (one node's, the smallest, where the scan can't spread), less the tier's VM memory. Each sandbox boots from the cache its scan filled. When the cache overflows, the first-scanned disks are evicted, those reads go back to PBS at boot, and a Windows guest's services can miss their start deadline.
- An unknown key in a plan is reported with its line, its section (for
example
tiers[]), the key it most likely meant, and the keys that section accepts. It used to name a Go type:field grace not found in type config.Tier. - When a probe is refused ("connection refused") on a VM restored without
some of its NICs, the diagnosis names the missing NICs and suggests
network_recovery.attach_all_nics. A service that also listens on a missing NIC's address can fail to start and stop answering on every port. - A warning, in every run and in
doctor, when sandbox overlays share a physical disk with the cluster file system's database (/var/lib/pve-cluster). Partitions, LVM and RAID are resolved to their disks, and a ZFS dataset to the disks under the devices its imported pool lists (zpool status). When the disks can't be determined, the run treats them as possibly shared. In the lab, a Windows sandbox's first writes to its overlays on a node's single root HDD delayed corosync ("Token has not been received"). doctorwith a plan reads each VM's backup as a run would, on a node withproxmox-backup-client, and creates nothing. It reports the snapshot, firmware, restored disks, encryption and backup consistency per VM. It fails each VM a run would refuse (no key or the wrong key, UEFI withoutoverlay_storage, a disk the backup lacks) and warns about a crash-consistent backup.--overlay-storagematchesrun.- Backup consistency evidence. Each VM record says whether its backup was
quiesced(the guest agent froze the file systems),crash-consistent, takenpowered-off, orunknown, with the reason from the backup's own log (source_consistency,source_consistency_detail), and the run warns about a crash-consistent backup. In the lab this found 41 database VMs whose freeze SELinux refused and a Windows domain controller whose VSS freeze expired before the snapshot. The log is read through the PBS API with the configured credentials and certificate trust; the evidence never fails a run.
Fixed
- Teardown and recovery left an ephemeral sandbox's VNet and zone behind (D46). PVE refuses to delete a VNet that still has a subnet, and says only "Parameter verification failed", so the fallback that deleted the subnets on an error naming them never ran. The subnets are now always deleted first, then the VNet, then the zone. And when nothing could be deleted, teardown no longer applies the unchanged configuration, which only reloaded every node's network and turned forwarding back on for the leftover bridge: it releases the SDN lock and reports the failure.
- The default ephemeral sandbox network failed on stock PVE 9 (D44): the run refused its own VNet because IPv4 forwarding was on. PVE's ifupdown2 turns forwarding on for every SDN bridge at each network reload. And the run went on as soon as PVE's SDN apply task ended, although that task only forks each node's network reload, one node after another, without waiting for them: 78 s on a 12-node cluster, against a 90 s limit that a larger or slower cluster exceeds (D45). The run no longer waits for that task. It waits for its own node: the VNet's bridge up (and, at teardown, down) and ifupdown2's lock released (ifupdown2 runs as "python3", so the process name cannot tell), then turns forwarding off on its own bridge within a second or two of the node's reload; the containment table drops forwarded packets as well. A failed apply task is reported when the node's bridge does not settle. The zone is created for the run's node only. A pre-provisioned VNet that forwards is still refused, with what to do.
- An ephemeral sandbox network could apply an administrator's unfinished
SDN change. Applying SDN in PVE is all or nothing, on every node, and a
run applied the cluster's whole pending configuration along with its own
zone and VNet. A run now creates and removes them under PVE's SDN
configuration lock (PVE 9): it is refused, before it creates anything,
while the cluster has unapplied SDN changes, and the message names them.
planreports it too. At the end of a run over such a change, the sandbox is deleted but the deletion is not applied, and the report says so. The lock is held for seconds; a run that dies holding it leaves its token in the journal, and recovery releases it. - Automatic network recovery missed EL7 on a CentOS 7 template (D43): it
wrote a NetworkManager keyfile profile, which EL7's NetworkManager 1.18
does not read, so the guest kept its own address and the probes found no
one. CentOS 7's
/etc/os-releaseis a symlink, which the EL7 check read as absent, and a plan naming the interface skipped the check. The check now follows symlinks inside the guest, falls back to/usr/lib/os-release, and runs whether or not the interface is named. EL7 gets the MAC-bound ifcfg profile it reads. - A plan whose sandbox VMs need more memory than the node can ever hold
within
max_host_ram_percentis refused at once, byplan(✗, exit 1),doctorand the run, with what it needs, what the node can hold, and what to do: split the plan, or use a node with more memory. Before,plansaid "not admitted now; the run would wait and retry", and the run deferred ten times before failing. In the lab, an 84-VM tier needing 249 GiB was refused this way on a 126 GiB node. - The controller did not notice a node that crashed or rebooted under a run. The worker command is silent through a grace or soak wait, and without ssh keepalives only TCP keepalive would have noticed, after about two hours. A blade reseated in the lab went unnoticed for 9.5 minutes, until it came back. Now:
- Every ssh and scp to the node gives up after 60 s without an answer, and after 20 s when connecting.
- The controller says it lost contact, then retries the node through the 20-minute cleanup window and cleans up as soon as the node answers.
- It says so when the node rebooted during the run (its boot ID changed).
- If the node never answers, it says what happens next and prints the exact recover command to run on the node.
- Lease renewals that keep failing are reported while the run is still going, and so is the return of contact. Before, they were ignored.
- The final error ends with ssh's reason. Before, it carried the worker's whole stderr, 10 KB in one line.
- What the controller's own recovery does is shown in the run's output,
and the node keeps it in
<state-dir>/logs/<run>.log, whoever started the recovery. Before, only a recovery the worker's supervisor ran left a log on the node. - Release and CI builds use Go 1.26.8, including the digest-pinned controller builder, and check reachable vulnerabilities with the build toolchain.
- The signed release checksum manifest authenticates
CONTAINER-IMAGE.txtas well as the CLI binary. Container and appliance instructions verify it before using the image digest. - Controller PBS encryption keys must be regular, owner-only source files. Staging reads the checked file descriptor and rejects symlinks, including replacement during open on Linux.
- The explicit disposable-lab TLS acknowledgement now passes through controller environment files and the protected worker environment. An omitted acknowledgement clears any inherited value.
- Storage documentation describes automatic full-disk copies, temporary firmware copies, overlay growth, and the default uncapped copy policy.
plansaid a run would refuse when an unfinished run's sandbox VM still held the plan's VMID ("already exists in the PVE cluster"). The run recovers that VM first and goes on.plannow counts what its "before the run" recovery deletes, the unfinished run's VMs, zone and VNet, as gone, as the run finds them (S15-07b). It also lists the probe address's nftables containment table, which recovery removes with the address.- A run stopped by a signal now says which, and whether the node was shutting down. After a node reboot, the report and the controller used to say only "validation failed: context canceled", so an operator couldn't tell a node shutdown from a Ctrl-C or a bug (S15-07). The reason now reads, for example, "interrupted: the node began shutting down (SIGTERM from systemd); the run was at: …". The controller reports an interrupted run, exit status 130, not a failed validation.
- A sandbox VM that QEMU stopped went unnoticed until its probes timed out
as network errors. The case found was a guest stopped on a write to a
full overlay storage (S15-06e): PVE's status still says "running" there;
only
qmpstatussaysio-error. The run spent its whole grace period and probe timeouts, then reported "no route to host", and the report never named the full storage. The run now checks every 10 s, while it waits, probes and soaks, that its VMs have not stopped for good. A VM that has fails at once, with the reason: - a disk write error, with the room left where its overlays are;
- a QEMU internal error;
- a guest OS crash;
- a shutdown.
planrefused everycopy_before_boot: alwaysVM ("the free space where the copies go is unknown"), where the run copies. It measured the sandbox VM's image directory, which only the run creates, so its free space and disks came back unknown.autocould also plan a copy onto the cluster database's disk that the run would refuse.plannow measures the directory's storage (its nearest existing parent, never above the storage itself), less the copies of the VMs before it, as the run sees it.- Recovery could hang the node's management plane. A map daemon killed
while one of its threads read its own file (the loop's partition scan)
never exits, so its FUSE connection stays open. Every open of the loop
then waits in the kernel for ever: recovery's
losetup -d, udev, and PVE'svgs, which froze the node's storage status for minutes in the lab (S15-02f). Before it detaches a map it owns whose daemon is gone, recovery now aborts the map's FUSE connection, found through the mount table. It runs the unmap and the detach under a 20 s watchdog, aborts the connection when either hangs, and reports a detach that is still stuck instead of waiting for it.doctorand preflight never open a loop with reads in flight. They name FUSE connections with requests waiting and no mount, with theabortcommand to run beforelosetup -d. - An ephemeral sandbox, the default, could not start on PVE 9. Before
creating its zone and VNet, a run checked that neither existed by reading
each one. PVE answers a read of a missing zone with HTTP 500 ("sdn zone
'
' does not exist"), not 404, so every free ID looked like an error and the run and planstopped at the ownership check. The lab found it throughinit: its sandboxes are pre-provisioned, which skips the check. The check now reads PVE's zone and VNet lists. Teardown of a zone or VNet that is already gone accepts the same HTTP 500, but only when it names that object. - A run with a pre-provisioned sandbox read its zone at
/cluster/sdn/zones/<zone>, which PVE allows only withSDN.Allocate. A token without it passed the privilege check and then failed at that read (S14-18b). The run now reads the zone and VNet from PVE's lists, which need onlySDN.Audit, and the privilege check asks forSDN.Auditon the zone. planshowed the overlay storage check as passed where the run warns: overlay storage on the cluster database's disk (D36), and a PVE storage status that did not answer. Both are now a warning on that check and in the plan's warnings.- Recovery after a worker was killed while it mapped a disk could report
its journal clean and still leave the map behind (S15-02d, S15-02e). A
map daemon stopped before it wrote its pid file left a loop device or a
dead FUSE mount, and recovery found a map only through its pid file.
Recovery now also finds the map by its name in
/proc/self/mountinfoand/run/pbs-loopdev. It waits for killed map processes to exit, and detaches a dead mount withumount2(MNT_DETACH), which works whereumount -lgetsENOTCONN. A mount it cannot remove is reported, and recovery says incomplete, not clean.doctorand the run's preflight also warn (stale_loops) about a read-only loop attached to a FUSE file that is no longer mounted, which nothing ties to a run, and give thelosetup -dcommand that releases it. - After a node worker was killed, what its supervisor's recovery did went
unseen. The supervisor stopped streaming to the controller first, and the
log was deleted with the workspace. Recovery now reports every action as
a progress event (
[Recovery] Destroyed sandbox VM …,Stopping the PBS map process …,Releasing /dev/loopN …). The supervisor streams them to the controller until recovery ends, and keeps the log in<state-dir>/logs/<workspace>.logon the node, also when the controller had gone. - A run failed at its first check when PVE's status read of a local overlay
storage did not answer: "read storage ... status: PVE API error (HTTP
596)". PVE computes every storage's status on the node together, so a
slow PBS storage, as during its backup window, stalled it; in the lab for
2 min 19 s. The read now gets 45 s. When PVE doesn't answer, the run
checks the storage on the node itself: the directory
pvesm pathresolves for it must have room. It warns and names the cause. A definite answer from PVE, such as the wrong storage type, still refuses. - A worker killed while
proxmox-backup-client mapran, or a map that timed out after creating its loop device, left a read-only loop device and map daemon that no journal named. Each map is now journaled before it is made. Recovery, and the run itself when a map fails, releases the loop device only when all of these hold: - its PBS map file name ends with that snapshot and archive;
- its daemon started within the map's window;
- no journal already tracks it.
It first stops the map's own client process, since the daemon it forks
outlives a killed worker and can create its loop device after recovery
has looked. It sends SIGINT, as
proxmox-backup-client unmapdoes, and SIGKILL only after 20 s. The process must match by its arguments and start time. It then removes the stale FUSE mount and pid file a killed daemon leaves behind, which the client's own cleanup would trip over. - Sandbox overlay disks use
cache=unsafeand are created with lazy refcounts. Under PVE's default (cache=none), QEMU also read the backup through its qcow2 backing with O_DIRECT, so every read bypassed the pages the pre-boot integrity scan had cached and went back through the PBS map. Restored Windows Server 2016 and later could then miss the service manager's 30 s start deadline: the guest agent or a database service stayed stopped. The sandbox now boots from the cache the scan fills. In the lab, a Windows Server 2016 guest that had failed this way started its agent at the first check, with the node's I/O wait under 1% during boot. The overlays are disposable, so skipping the guest's flushes loses nothing. Runs with the lab-only--skip-integritystill boot cold. - The pre-boot integrity scan releases its own cached copy of each disk as it reads. A mapped disk is a loop device over the PBS map's FUSE file, so a buffered read cached every block twice. Only the FUSE file's copy survives the scan, and it is the copy the sandbox boots from. The other copy doubled the scan's memory use and could push the boot's copy out on a node with less free memory than twice the VM's disks.
- After the node worker was killed (SIGKILL, or with its SSH session), the controller refused to clean up: "recorded node worker PID no longer identifies this workspace binary" (exit 2). The sandbox VM, PBS maps and probe address then stayed live until the node's next run. A worker that no longer runs, or whose PID now belongs to another program, is now left alone and its journal is recovered at once. Recovery takes the host lock, which a live worker would hold, and acts only on the identities the journal recorded. If the controller died too, nothing recovered the journal before the node's next run. The worker's supervisor on the node now recovers it itself when the worker dies by a signal without a result.
- A Windows sandbox VM could run a newer machine version than a restore of
its backup. Proxmox VE starts a Windows guest with an unversioned machine
type on the version of the QEMU that created it (5.1 when that QEMU is
older than 9.1), which it reads from the VM's
metaline. PVE writes the sandbox VM's ownmetaand refuses to change it. So when a restore would start on another version the node can run, the sandbox VM is now pinned to it, and the run says so. Where both agree, it keeps the backup's unversioned setting. - The full mapped-image integrity scan ran while the sandbox VM booted, reading every byte of every disk through the same PBS map as the guest's first reads. The guest's boot starved, and Windows services can time out at start. The scan now completes before the VM starts, as a restore has all its data before it boots, and a damaged backup fails before anything boots from it. The run takes no longer: the guest's boot already fell inside the startup grace period that followed the scan.
- The failed-VM diagnostic listed only failed units. A data mount skipped
after its device timed out (
nofail), and the services left inactive by it, did not show at all. The diagnostic now lists mounts that are not active and the boot's dependency, device-timeout and mount failures. The live (guest agent) and offline (overlay) diagnostics both do. - The documented
CertiStackRolelackedVM.Config.CPUandVM.Config.Memory, which every VM create needs. A token with exactly that role could not run at all. The role now lists them, andVM.Config.CDROMfor the CD-ROM drives below. - The sandbox VM had no CD-ROM drives where its source has them (a
cloud-init drive, an ISO), which a restore keeps, so the guest saw a
storage controller disappear. They are now attached empty, never in place
of a restored disk. A token without
VM.Config.CDROMfalls back to root's localqmon the node, or creates the VM without them and warns. The signed record lists the drives the VM actually has (restored_hardware.empty_cdroms). - A passphrase-protected PBS key without
PBS_ENCRYPTION_PASSWORD, or with the wrong one, failed insideproxmox-backup-client, whose output CertiStack withholds, so the run showed only an exit status. A key whose file says it needs a passphrase is now refused before anything is staged or restored when none (or one shorter than PBS accepts) is set. A wrong passphrase is named in CertiStack's own words; the client's output stays withheld. - A clean run could end "cleanup unverified" (exit 2) and keep its worker
workspace, credentials included, when another run started on the same node
between the worker finishing and the controller's own check. That check
needed the host lock, and the node's journal might already belong to the
new run. The worker now records the run under an ID the controller chose.
The controller reads the journal without the lock and accepts it when it is
this run's and clean, absent, or a later run's: a run only replaces a clean
journal, after recovering it. Only this run's own unclean journal is
recovered, and a busy host lock is waited out instead of failing. With
--ssh-sudo, the root-only state directory is now checked through sudo. Before, every file in it looked absent. - The sandbox VM got a new SMBIOS identity, which a restore of the same
backup does not:
qmrestorekeepssmbios1(it regenerates the UUID only with--unique). The backup's SMBIOS UUID and vendor strings are now carried, each validated as Proxmox VE accepts them, and the signed record lists the UUID (restored_hardware.smbios_uuid). The VM generation ID is still new, as after any restore. - A Windows sandbox VM ran a newer machine version than a restore of the
same backup would. Proxmox VE pins a new Windows VM's machine version on
create (the backup's
q35becamepc-q35-11.0+pve2), whileqmrestorekeeps the backup's unversioned setting, which the same nodes resolve topc-q35-11.0+pve0. The backup's setting is now restored after create, and the signed record carries the version the VM ran (restored_hardware.running_machine). - The sandbox VM always had a memory balloon device, even when its source
had none (
balloon: 0, common on Windows guests), so the restored guest saw hardware it never had. The device now follows the backup unless the plan setshardware_overrides.balloon_mb, and the signed record says so (restored_hardware.no_balloon). - One transient PVE API failure failed the whole run: pveproxy answers HTTP 596 when a cluster node busy with clones or backups does not respond in time, and the same read succeeds a second later. Reads (GET) are now retried up to four times over about 17 seconds on HTTP 502-504 and 595-597 and on dropped connections, and every retry is logged. Other requests are retried only when they provably never arrived (a refused connection, or 595); timeouts are not retried. Errors for 595-597 now say what failed instead of an empty message.
- Network recovery refused a Debian ifupdown guest whose configuration
names more than one interface, which every cloud-init Debian image does
(cloud-init's
enp6s18beside the image's staleeth0), and asked fornetwork_recovery.interface_name. When the sandbox VM has only the one NIC, each name now gets anallow-hotplugrecovery stanza and only the one that exists comes up. Bridges, bonds, VLANs and tunnels are set aside in favor of the physical interface under them. VMs restored with several NICs still needinterface_name. sc querychecks copied sc.exe's state text into the report unbounded; a hostile guest could put about 600 KB, or terminal escape sequences, into a signed report. The state is now reported as its code and a fixed name. It is found by value, so translated labels ("STATUS" on German Windows) no longer fail a running service. All guest text in QGA results is now stripped of control characters as well as capped.- A DNS
srvcheck passed on an SRV record with target ".", which RFC 2782 defines as "service not available". Such records no longer count. Probe details list at most eight answers, so a reply packed with compressed records cannot inflate the report. ldapnaming_contextfailed when a domain controller refused the clear-text anonymous bind (strongAuthRequired or inappropriateAuthentication), although it still serves the Root DSE. The check now reads the Root DSE after such a refusal.proxmox-backup-clientran without the parent-death signal every other helper uses. A restore or map started just before the worker was killed could finish after crash recovery, leaving a TPM-state copy or a loop device behind. PBS commands now die with the worker.- A drive archive such as
drive-scsi0.img.fidx.fidxwas rewritten to a mangled name and passed to PBS; it is refused now. compliance.retention_daysis bounded at 36500 (100 years). Larger values are signed into reports and overflowed evidence stores' durations.- With
CERTISTACK_PBS_KEYFILEset, every run failed: the key was passed tosnapshot listandsnapshot files, which do not accept it. It also went to unsigned snapshots, whose restores then fail the manifest signature check. The key now goes only torestoreandmap, and only for encrypted or signed snapshots, so one plan can mix both. - SCSI disks got
iothread=1behind every controller. PVE honors it only withvirtio-scsi-single, and elsewhere its warning failed the VM start. A PVE task that ends withWARNINGS: Nnow counts as successful, as it does in PVE. - Without a boot order, CertiStack picked the first disk by ide, sata, scsi, virtio. PVE's order is ide, scsi, virtio, sata, so a VM with a SATA data disk booted and network-recovered the wrong disk.
- A source
machinewith options (q35,viommu=virtio) refused every plan. It is parsed as PVE's property string now, keeping allowlisted options. - A source without a
scsihworostypeline runs on PVE's defaults,lsiandother, and now restores on them rather than on virtio-scsi-single and l26. use_backup_configrefused a VM over any unusablenet1+ line. Those NICs are now reported as omitted (or refused only ifattach_all_nicsasks for them), and NIC errors name the real slot.- A plan could restore one archive into two slots, or an archive into another slot, and the omitted-disk evidence would be wrong. Each archive must now go to its own slot, once.
source_crypt_modewas the strongest mode of any file in the snapshot listing. It is now the weakest mode among the VM's disks, and the newsource_signature_verifiedsays whether a key verified the manifest. Disks excluded from backups (backup=0) are listed asexcluded_source_disksinstead of silently missing.- When the controller keeps a worker workspace for recovery, it now deletes the staged PBS encryption key from it; recovery never needs the key.
doctorand the node worker's preflight now check what the sandbox needs before anything starts:nftwhen the plan has wire probes,guestfishwhen it uses network recovery (with the package to install, ormode: preserveto opt out), and a configured PBS encryption key.- A wrong PBS key (another storage's key) failed inside the PBS client with
a withheld error. CertiStack now compares the key file's fingerprint with
the snapshot's key fingerprint from
snapshot listand refuses before any restore, naming both. A file that is not a PBS key file (a paperkey printout, a PEM master key) is refused at startup. - A run that missed
compliance.rto_target_secondssigned a failed report but exited 0; it now fails like any other validation failure, and the controller never reports success for a worker report that does not record a pass. - A QGA probe with
expected_outputon a guest that disablesguest-execsilently fell back to matching the host name, so a check such assystemctl is-active postgresqlcould pass without checking the service. Onlyhostnameprobes use that fallback now. - The
dnsprobe passed onNXDOMAIN, and withoutdomainit looked uplocalhost, which the worker answered from its own/etc/hostswithout contacting the guest. It now sends the query directly to the target and, withdomain, requires an A or AAAA answer. - Multi-tier plans reserved memory for the largest tier only, although every earlier tier keeps running until teardown; admission now reserves the sum.
- The CLI controller ignored
storage.scratch_dir; it now applies after--scratch/CERTISTACK_SCRATCH_DIRand before the built-in default. validatenow rejects plans the runtime would refuse: astorage_idother thanpbs, mixed PBS namespaces,vmidequal tosource_vmid, SDN zone or VNet IDs longer than 8 characters, andhttpprobes without a port. The shipped minimal example used 9-character SDN IDs and was corrected.- On stock Proxmox VE the sandbox VM could not be created: the API refuses
file paths for an API token, and even root's
qm createrefuses a disk outside a storage ("unable to associate path ... to any storage"). The newstorage.overlay_storagesetting (--overlay-storage) places overlays as VM-owned volumes of a file-based storage and attaches them by volume ID with the existing least-privilege token. The storage is validated before any mutation. Without it, the old file-path mode now explains how to fix the failure. - Restored Ubuntu 26.04 (netplan) guests started their services about two
minutes late, so Stage 2 probes timed out. netplan resolved the MAC-matched
recovery profile to the NIC's name at generator time (
ens18, after a dracut initrd rename), the guest's ownset-namethen renamed it back, and systemd-networkd-wait-online waited for a name that no longer existed. For netplan guests the overlay now carries a wait-online drop-in that waits for any online link. The profile also no longer marks the linkoptionaland disables IPv6 autoconfiguration, which the air-gapped sandbox never answers. - Network recovery silently did not take effect on two guest families, which
booted with their own address (and, on Debian, their own default gateway):
Debian-style ifupdown guests (auto picked systemd-networkd because
/etc/systemd/networkexists) and EL7 guests (NetworkManager 1.18 ignores keyfile profiles unless the keyfile plugin is configured). Auto now detects ifupdown and EL7 and writes a profile those guests actually apply. latestcould resolve to an aborted backup (partial files, noindex.json.blobmanifest), and every restore of that VM then failed with "restore PBS VM configuration: exit status 255". It now resolves to the newest complete snapshot and logs the incomplete ones it skipped; a group with only incomplete snapshots fails with an explicit message.- A probe that polled until its timeout reported only the final attempt,
which the deadline had cut off ("timed out waiting for guest process"), and
hid what the guest had answered on every completed attempt. It now reports
the last completed attempt. A failed QGA command also shows its stdout when
stderr is empty (
systemctl is-activeanswersinactiveon stdout). - guestfish failures lost guestfish's own error and always suggested installing libguestfs-tools, so an empty or non-Linux disk ("no operating system was found on this disk") read as a missing dependency. The first guestfish error line is now included, and the install hint appears only when guestfish is missing.
- Every restore failed with proxmox-backup-client 4.2, which reports the mapped device as "mapped on /dev/loopN" instead of "mapped to"; both wordings are accepted.
scripts/run-lab-campaign.shcould never credit a--skip-integritycampaign and did not recognise the engine's runtime admission denial.- An
httpprobe whose target answered with the wrong status on every attempt reported only "context deadline exceeded" when the probe deadline ended mid-attempt, hiding the status the target had returned. It now reports the last answer (for example "expected HTTP status 200, got 503 Service Unavailable") and says that later attempts ran out of time. Thetcpprobe had the same flaw: a stopped listener was reported as "context deadline exceeded" instead of "connection refused". - A scheduled backup job that selects all VMs backed up the sandbox VMs of
a running validation, copying restored guest data into the backup store
under the sandbox VMIDs and locking each VM for minutes. Overlays are now
attached with
backup=0, so such a job stores only the sandbox VM's configuration.docs/deployment-model.mdexplains how to exclude sandbox VMs from those jobs. - The documented Docker Compose deployment could not start a run: the root
filesystem is read-only and nothing pointed the controller's temporary
directory at the
/tmp/certistacktmpfs. The image, Compose file, and appliance wrapper now setTMPDIR=/tmp/certistack, and CI and the release start arunfrom the image under the published restrictions (scripts/controller-container-smoke.sh) and require it to reach its first node contact. - A worker whose recovery journal could not be written after it assigned the host probe address (and its nft containment table), or after it created the SDN zone or VNet, returned without a teardown step for that resource. Removal is now registered as soon as each resource exists, and the probe address is journaled as a pending allocation before it is assigned.
- An unreadable signing key or an unwritable report directory was found only
after the remote run, when the evidence was about to be removed with the
node workspace. The controller checks both before contacting the node, and
a report that still cannot be saved is returned and printed between
BEGIN/END CERTISTACK REPORTmarkers instead of being lost. - An arm64 controller uploaded itself as the worker to an x86_64 PVE node. The controller now checks the node's architecture before uploading, and releases publish amd64 only; arm64 remains a source build for the offline commands.
keygencould leave a new private key without its public half when the public path was refused, and--private k --public k --forcereplaced the private key with the public one. Both destinations are checked before either is written, the private key is restored if the public write fails, and paths naming one file are refused.scripts/generate-current-coverage-cases_test.sh, run by CI and the release, failed: the generator wrote its--storage-idvalue intostorage_id, which the engine now restricts topbs. The flag defaults topbsand rejects anything else.
Changed
- The supported pair is Proxmox VE 9.x with PBS 4.x, the pair every live campaign ran on. The docs used to list PVE 8.x and PBS 3.x as well; they are not supported until they have their own signed matrix evidence.
- On firewalld guests, network recovery keeps the guest's own default zone and adds the probe rules to it, instead of replacing the zone with the probe rules only. Restored guests on the same isolated VNet now reach each other as far as their own firewalls allow, as in production. Before, a multi-tier application failed in the sandbox because its Rocky tiers refused their peers even on ports their own zone opened. CertiStack still adds no rule between guests, and the sandbox stays air-gapped.
- Every plan now reads the backup's firmware settings (
bios,efidisk0,tpmstate0), not only plans withuse_backup_config. A plan with explicit drives used to restore a UEFI VM under SeaBIOS without its firmware state, which cannot boot. A plan that leavesbiosempty now gets OVMF for a UEFI backup, andbios: seabiosfor a UEFI backup is refused. - The source configuration parser stops at the first
[section]: snapshot and pending sections describe other states of the VM. - A plan that restores only some of a VM's backed-up disks now says so: the
disks it leaves out are listed in the VM's
omitted_source_disksin the signed report, and the run warns. A plan that lists a disk the backup does not have is refused before anything is created. Before, such a plan restored a partial VM without saying so. - The single-disk
source_pbs.driveshorthand uses the slot its archive names (drive-virtio0isvirtio0) instead of alwaysscsi0. - Plans without
use_backup_confignow also take the backup's machine type, SCSI controller and guest OS type when they do not set them. They used to get PVE's defaults (i440fx, virtio-scsi-single, l26). A guest installed on q35 or an LSI controller could fail to boot there, which says nothing about the backup.
Added
- Report schema 1.5: each VM record carries
source_vmidandrestored_hardware(the virtual hardware the sandbox VM was created with), and phase timings includeintegrity_verification_sec. network_recovery.mode: ifupdownfor Debian-style guests. The overlay's/etc/network/interfacesis replaced by one static stanza for the guest's single configured interface (no gateway, no DNS); the original is kept asinterfaces.certistack-original. Guests with several interfaces neednetwork_recovery.interface_name.- When a failed VM's guest agent shows the recovered NIC without its recovery address, the run says so ("network recovery did not take effect: guest NIC ... has ..., not the recovery address ...") in the progress output and the signed diagnostic.
ldapprobes takenaming_context: an anonymous Root DSE read that must advertise the given naming context, with no directory search. It asserts an Active Directory domain controller, which refuses anonymous searches by default, instead of settling for "port 389 answers a bind". The report records the DC'sdnsHostNameandisSynchronized.dnsprobes takerecord_type: srv, which requires SRV records fordomain(for example the DC locator_ldap._tcp.dc._msdcs.<domain>) and reports each target, port, priority and weight.qgaprobes acceptsc query <service>on Windows guests.sc.exeruns directly with literal arguments (no shell), and the probe passes only when the service's STATE is RUNNING. sc.exe exits 0 for a stopped service too, so the exit code alone is not used. Until now a Windows guest could only be checked with QGA ping and hostname, because in-guest commands ran through/bin/sh -c.mssqlandsmbprobes.mssqlsends a TDS PRELOGIN and passes on a well-formed PRELOGIN response, recording the server version and encryption setting.smbsends an SMB2 NEGOTIATE and passes on a successful response, recording the dialect and whether signing is required. Neither logs in, so both prove that the service answers its protocol, where atcpprobe only proves that a port accepts connections.- UEFI VMs restore with their firmware state. The UEFI variable store
(
efidisk0) and TPM state (tpmstate0) are restored from PBS as disposable raw copies onstorage.overlay_storageand attached in the source slots. The report records each copy's archive and SHA-256 undersource_disks, and its options under the newrestored_hardware.firmware_disks. Before, a backup with either disk was refused, so no UEFI VM, which includes most Windows Server 2022/2025 VMs, could be validated. - Disks on every bus restore: VirtIO block (
virtio0-virtio15), SATA (sata0-sata5) and IDE (ide0-ide3) as well as SCSI, each through the same read-only map and COW overlay. The sandbox VM boots the first restored disk of the source's boot order (recorded asrestored_hardware.boot_disk), so a VM that boots fromvirtio0orsata0restores as it runs. Before, such backups were refused, or restored without their non-SCSI disks. network_recovery.attach_all_nicsattaches a VM's other NICs (net1and up) with their MACs and models, only to the isolated VNet, so services bound to a secondary interface's address start. Withsandbox_ipset to net0's address, wire probes may also target those NICs' own addresses; every guest address in a plan must be unique. NICs left out are listed in the report'somitted_nics. All PVE NIC models are accepted, including the e1000 variants of VMs migrated from VMware (e1000-82545em), whichuse_backup_configused to refuse.- Encrypted PBS backups restore. Set
CERTISTACK_PBS_KEYFILEto the backup encryption key (andPBS_ENCRYPTION_PASSWORDfor a password-protected key). The controller stages the key in the run's private worker workspace, and the worker refuses a key that other users can read. Every PBS call passes it, and each VM record states the backup'ssource_crypt_mode. Before, the key was never passed, so an encrypted backup failed at its first restore with an error that did not say why. Now a missing key stops the run before anything is restored, with a message naming the setting.
Security
- A signed report with a duplicated JSON key could carry values its signature never covered: canonicalization keeps the last duplicate, while decoding merged the earlier one, so a failed probe could be turned into a skipped one. Verification now rejects duplicate keys at every level and decodes the result from the exact bytes whose signature it checked.
- Wire probes (
tcp,http,dns,ldap) used ordinary host sockets, so a plan with a loopback sandbox range passed against a listener on the worker host with no guest present. Probes are now sent from the host probe address and bound to the sandbox VNet interface, andvalidaterejects a sandbox range that overlaps loopback, link-local, multicast, or reserved space, and probe targets on the subnet's network or broadcast address. - With
PVE_TLS_INSECURE=truebut no lab acknowledgement,node-runrecovery and preflight anddoctorsent the PVE API token over an unverified connection before the runner refused the run. The PVE client itself now refuses to send any request in that case. - A bound wire probe could still pass against the worker itself: a host listener on another address of the sandbox interface (for example on a preprovisioned VNet) answered it with no guest present. A run now refuses any probe target that is an address of the worker host, before it creates anything and again on every connection attempt.
- When assigning the host probe address failed and removing its nftables
containment table also failed, the table stayed installed while the
report said
all_cleanedand the journal said cleaned. A failed rollback is now reported, the leftover table (or address) is journaled and removed at teardown and by recovery, and the evidence stays unclean until that succeeds. An address already on the interface is refused rather than treated as the run's. - A failing
httpprobe signed up to 200 characters of the response body, which can carry a stack trace or application data, into the report. Only the status code is recorded now.
Initial release capabilities
- Controller CLI:
validate,run,verify-report,keygen,doctor,inspect,recover, andversion. - Controller-to-node execution model: strict-SSH bootstrap of a short-lived
node worker under
/var/lib/certistack/workers/<uuid>, allowlisted literal environment files, and no persistent service on the hypervisor. - Mapped recovery: read-only
proxmox-backup-client mapof PBS snapshots with per-disk qcow2 COW overlays on private node scratch storage, full mapped-image integrity verification, and multi-SCSI-disk layouts. - Backup-configuration inheritance (
use_backup_config) for safe compute, firmware, machine, SCSI controller, MAC, and NIC model settings. - Isolated PVE SDN sandboxes (
simpleandvxlan, ephemeral or preprovisioned) with no gateway, no SNAT, run-scopednfthost containment, and refusal to adopt pre-existing zones or VNets. - COW-only guest network recovery for
ifcfg, NetworkManager, Netplan, and systemd-networkd guests, including a scoped firewalld probe rule; the source VM and backup are never modified. - Deterministic probes: QMP readiness, worker-origin TCP, HTTP/HTTPS, DNS, and LDAP wire probes, a constrained QEMU Guest Agent vocabulary, and forensic QMP screendumps on failure.
- Host admission control with capacity reservation for the largest concurrent tier.
- Durable run journal, LIFO teardown stack with signal and panic unwinding, and journal-scoped crash recovery with process-identity verification.
- Ed25519-signed RFC 8785 canonical JSON reports (schema 1.4) with pinned signer verification and machine-readable verification output. The JSON is the evidence; HTML/PDF binder rendering and outbound notifications are Enterprise Edition features and are not part of this repository.
- Lab campaign runner with resume, fault, and guard modes, plus the coverage-case generator and compatibility-matrix audit tools.
- Unprivileged controller OCI image, Docker Compose reference, Packer appliance template, and a Terraform example for least-privilege PVE RBAC.
- Documentation site (MkDocs Material), example plans, contributor license agreement, code of conduct, and security policy.
- Library packages for products built on the engine:
pkg/runner,pkg/recovery,pkg/attest, andpkg/config.
Security
- Chaos fault-injection hooks compile only with
-tags=chaosand are a constantfalsein every shipped binary. - TLS verification bypasses are rejected in plans;
PVE_TLS_INSECURErequires an explicit lab acknowledgement and is refused for release evidence. - PBS repository strings, credentials, and raw PBS-client output are withheld from logs, reports, and campaign bundles.
- Release binaries ship with a GPG-signed checksum manifest and SLSA build provenance.