Skip to content

Release-readiness and validation evidence

CertiStack must not be described as “battle tested” merely because its unit tests are green. A production release requires reproducible evidence from the exact PVE, PBS, storage, guest, and network combinations the release supports. This document defines that evidence, its boundaries, and the release decision rule for both the Community engine and products that embed it.

For the operator-facing path from a blocked inventory to an accepted release campaign, use the lab acceptance runbook. Use the release evidence checklist to retain the candidate identity and artifacts required for a final approval.

Supported recovery contract

Dimension Supported contract Explicit boundary
Controller Linux amd64 controller or unprivileged controller OCI image The controller is not a PVE host agent and does not need host mounts, KVM, a Docker socket, or a PVE shell.
Hypervisor Proxmox VE 9.x on amd64 with qemu-img, paired with PBS 4.x The pair the signed matrix evidence covers. Proxmox VE 8.x with PBS 3.x, or any other pair, needs its own signed matrix evidence before it is supported. A node worker performs node-local mapping, overlay, QMP/QGA, and cleanup with the documented PVE/PBS privileges.
Snapshot storage PBS snapshots mapped read-only; one COW overlay per recovered disk on any bus (SCSI, VirtIO, SATA, IDE); a UEFI VM's efidisk0 and tpmstate0 restored as disposable raw copies (needs overlay_storage) Boots the source's boot disk. Backed-up disks a plan leaves out are listed as omitted in the report. Cloud-Init and other non-backed-up media are not recovered.
VM configuration Backed-up safe CPU, memory, firmware, machine, SCSI controller, MAC, and NIC model may be inherited Source tags, bridge/VLAN/firewall settings, hooks, startup rules, arbitrary QEMU arguments, and device passthrough are never replayed.
Isolated networking One PVE simple or vxlan SDN VNet with no gateway, SNAT, routed path, or physical uplink A disconnected L2 domain may reuse source VLAN/IP addressing only with the retained isolation attestation. Restored guests on the one VNet reach each other as far as their own firewalls allow. Production-connected networks, DHCP fallback, and routed multi-VLAN topologies are unsupported.
Guest-network recovery Linux ifcfg, NetworkManager, Netplan, systemd-networkd, and ifupdown may be adapted only on the disposable COW overlay preserve is an explicit opt-out. Windows and any unrecognized guest-network backend require a separate supported plan or a documented exclusion.
Probes QMP readiness; worker-to-guest TCP, HTTP(S), DNS, LDAP; optional in-guest QGA checks Stage 2 wire probes prove worker-to-guest reachability, not guest-origin routing, DNS policy, or arbitrary east-west traffic. HTTPS certificates must chain to a worker-trusted CA and match the declared SNI/target; verification bypasses are unsupported.
Evidence Full mapped-image digest pass, signed JSON report, and verified teardown evidence --skip-integrity reports are lab-only and cannot satisfy production, compliance, or release evidence.

Required test matrix

The test objective is coverage of meaningful interactions, not every subset of the fleet. Exhaustively restoring every subset of even a modest fleet is both combinatorially infeasible and less useful than deliberate pairwise and failure-oriented coverage. The local campaign generator records the selected inventory and plan content so results are attributable to a particular matrix.

Coverage area Required evidence Passing condition
Healthy fleet One full healthy plan plus an individual restore for every supported source VM Every plan has a signed report with all_passed: true and teardown_evidence.all_cleaned: true.
OS/network backend At least one representative per supported guest OS and network backend Guest receives only its isolated recovery profile and configured probes pass.
Service behavior Every supported service family and every supported pairwise co-restore combination Each declared worker-to-guest protocol assertion passes with the execution origin recorded. This is not evidence of an application transaction. A guest-to-guest dependency is evidenced only where a guest reports it itself (for example a readiness endpoint that checks its peers), probed in tier order.
Storage layout Single-disk and every supported multi-SCSI-disk layout Every archive maps to its original SCSI slot; no source image is writable; service starts from its full layout.
Hardware/firmware BIOS/OVMF (with UEFI variable store, Secure Boot keys and TPM 2.0 state), machine type, SCSI controller, MAC/NIC model, CPU/RAM/balloon variants present in the supported inventory Inherited or explicit safe setting is visible in the report and recovery reaches the configured probe stage.
Network lifecycle ephemeral and, when offered, preprovisioned isolated SDN paths Ephemeral resources are exact-owned and removed; preprovisioned resources are read-only validated and retained.
Capacity Worst intended concurrent tier and a deliberate over-capacity plan A safe plan completes; the unsafe plan is rejected before mapping, overlay, VM, or SDN mutation.
Negative service checks One injected fault per supported probe family Exactly the declared target VM/probe fails; the signed report verifies; and teardown is clean. A generic non-zero exit, a wrong protocol, or collateral probe failures are not a pass.
Isolation Packet capture on relevant host and sandbox interfaces during a restore No sandbox packet reaches production interfaces; there is no gateway, SNAT, or physical uplink.
Interruption Controlled termination during mapping, overlay preparation, boot, and probe stages The durable journal permits exact recovery; no owned VM, loop, overlay, host address, or ephemeral SDN resource remains.
Upgrade Previous supported release to candidate release and rollback drill Plans, licenses/configuration, reports, and recovery state remain usable; a known-safe test and report verification pass after each transition.

How to run and retain evidence

Generate or obtain a reviewed synthetic campaign in a protected controller-local directory, then validate and review its exact membership before any mutation. The public repository deliberately does not contain site-specific VM IDs, addresses, or generated plans:

/secure/certistack/certistack validate /secure/certistack/lab-cases
scripts/run-lab-campaign.sh --node pve-node-01 --cases /secure/certistack/lab-cases --list-cases

Run healthy, fault, and admission suites separately. For a release candidate, do not use --skip-integrity, and provide a signer keyring so reports are both mathematically valid and identity-pinned. Resume is deliberately bound to the selected plans, validation mode, node, executable hash, and signer-keyring hash; changing any of them requires a new campaign:

scripts/run-lab-campaign.sh --node pve-node-01 --cases /secure/certistack/lab-cases \
  --env /secure/certistack/controller.env \
  --keyring /secure/certistack/trusted-signers.json \
  --certistack /secure/certistack/certistack \
  --mode positive --release-evidence

For each campaign retain the generated plans and manifest, campaign.meta, plans.tsv, cases.tsv, controller and node versions/commits, PVE/PBS versions, guest-image identifiers, environment-free logs, signed JSON reports, report-verification output, packet-capture result, fault-injection record, and interruption-recovery journal/report. campaign.meta binds the selected plans to the node, executable, and signer keyring; plans.tsv records each selected plan's SHA-256 digest; cases.tsv records the exact per-case verified outcome and report path. A report with the right plan ID but a different plans.tsv digest is not evidence for the reviewed plan revision. Keep secrets and raw credentials out of the evidence bundle.

CAMPAIGN-BLOCKED is a deliberate safety control. It means the generated directory has unmet prerequisites, such as missing snapshots, an unaccepted inventory, or an unverified recovery-network boundary. Do not delete or empty that sentinel to force a run. A separate air-gapped L2 VNet may reuse source VLAN/IP addressing only with the required retained isolation attestation; see lab acceptance. Regenerate the cases and retain the new manifest after every prerequisite is satisfied.

Before generating executable cases from an inventory handoff, every selected template must be accepted: its declared OS identity, cloud-init behavior, SSH access where required by the test, and QGA transport must match the contract. The current-coverage generator rejects an inventory whose test_boundary readiness is not ready; a powered-on VM or a plan file is not acceptance evidence. A partial deployed fleet is likewise not evidence for a broader published compatibility matrix.

For a fleet-derived campaign, make the approved population explicit rather than allowing a refreshed inventory to silently shrink. The coverage generator requires --expected-targets and rejects a count mismatch. For example, use the exact approved support-boundary count—not a count inferred from currently powered-on VMs—when regenerating cases:

CERTISTACK_SIGN_KEY=/secure/certistack/controller-signing.ed25519 \
  go run .lab/generate_current_coverage_cases.go \
  --targets /secure/certistack/approved-targets.yml \
  --expected-targets 155 \
  --out .lab/cases-release-candidate \
  --zone-id csZone01 --vnet-id csVnet01 \
  --vlan-tag 101 --ip-range 192.0.2.0/24 --probe-ip 192.0.2.1 \
  --allowed-nodes <pve-node>

The number 155 is only an example of an approved boundary; use the actual signed-off population for the candidate. Passing a smaller number is a documented narrowed support boundary, never proof of the larger one.

Then audit the accepted target inventory against the approved compatibility contract before running any mutation. The audit fails if an implemented OS/workload cell selected by the release suite has no accepted, running target or if the handoff is QGA-only/unaccepted. It also requires the inventory's qualification record to cover every selected PVE/PBS pair and every available storage, guest, backup, and restore profile. It deliberately reports planned recipes and dimensions as outside the supported claim rather than pretending they passed:

go run .lab/audit_compatibility_matrix.go \
  --contract /secure/certistack/compatibility-matrix.yml \
  --targets /secure/certistack/approved-targets.yml \
  --suite release_candidate

This is a coverage-accountability check, not a substitute for live evidence. The qualification inventory identifies only accepted preconditions; retain the signed campaign reports, packet captures, and interruption/cleanup evidence that proves the actual recovery results.

Release decision

A release candidate is not ready when any supported-matrix cell lacks the required evidence, an injected fault produces an ambiguous result, report signature/teardown verification fails, or an isolation/interruption drill is missing. A failure may narrow the published support boundary only when the removed combination is clearly documented, rejected before mutation where possible, and absent from all marketing and sales claims.

Approve a Community release only when all of the following are true:

  1. go test ./..., go vet ./..., race tests where supported, cross-builds, and binary/container build pass for the candidate commit. A developer workspace or a moving module branch is not reproducible release evidence.
  2. The required matrix above is complete for the exact versions and configurations being released, with signed, pinned, full-integrity evidence.
  3. Healthy, fault, admission, isolation, interruption/recovery, and upgrade drills all meet their stated pass conditions.
  4. The installation, configuration, compatibility, operations, troubleshooting, security, and rollback documentation has been exercised by an operator who did not author the implementation.
  5. Any residual limitation is published as a support boundary, not hidden in a test caveat.

This is intentionally a high bar. Passing it supports a defensible production-readiness claim; bypassing it does not.