Skip to content

Lab campaigns

A campaign runs many plans through the same controller-to-worker path and records signed, attributable evidence for each one. This page covers the campaign runner (scripts/run-lab-campaign.sh), the test-only shortcuts that exist for disposable labs, fault drills, and capacity guards. Running a single plan is covered in getting started; the acceptance and release process that consumes campaign evidence is in the lab acceptance runbook and release readiness.

Run a local lab campaign

Place reviewed, synthetic lab plans in a protected controller-local case directory. The campaign script runs on the management controller and invokes the same remote certistack run flow for every case; it is not installed or run directly on a PVE host. A release suite should cover individual replicas, each service and OS family, pairwise healthy service families, broader mixed tiers, and isolated fault cases. Keep site-specific VM IDs, networks, plans, and generated cases out of the public repository.

cases=/secure/certistack/lab-cases
certistack validate "$cases"
scripts/run-lab-campaign.sh \
  --node pve-node-01 \
  --cases "$cases" \
  --env /etc/certistack/lab.env \
  --keyring /etc/certistack/keyring.json \
  --certistack /usr/local/bin/certistack \
  --mode positive \
  --skip-integrity

The signing key must be controller-local and mode 0600; do not place it in a shared, Windows-mounted, or source-controlled workspace. Use CERTISTACK_SIGN_KEY during your own plan-generation process so unsigned plans cannot become campaign evidence.

The runner records each case's output beneath the report directory and accepts relative --output paths by normalizing them on the controller. Node preflight is mandatory for every case; --skip-doctor is retained only as a deprecated compatibility flag. The campaign completes the selected positive matrix and prints a pass/fail summary even after an unexpected case failure; it exits non-zero if any case failed. Use --fail-fast when immediate triage is more useful than coverage. Its --env file is passed directly to CertiStack's allowlisted literal dotenv parser; the campaign helper never shell-sources or executes its contents.

The default positive matrix excludes diagnostic-*.yaml plans. Those are explicit troubleshooting drills with known edge conditions; run one with --case so it cannot lower the healthy-suite completion ratio.

Review exactly what a campaign will run, without contacting PVE or creating an output directory, with:

scripts/run-lab-campaign.sh --node pve-node-01 --cases /secure/certistack/lab-cases --list-cases

For a long positive campaign, resume its printed log directory after a controller interruption:

scripts/run-lab-campaign.sh --node pve-node-01 --cases /secure/certistack/lab-cases --env /etc/certistack/lab.env \
  --certistack /usr/local/bin/certistack --mode positive \
  --resume /var/lib/certistack/reports/campaign-logs/20260918T120000Z-positive

Resume never trusts a PASS line by itself. It skips a case only when the stored signed report cryptographically verifies and proves the same plan and validation mode passed with owned resources cleaned; incomplete, changed, or unverifiable cases run again.

Every executable campaign requires --keyring to pin the report signer; a report cannot establish trust in its own embedded public key. For production, add --release-evidence. This mode rejects --skip-integrity, records the release-evidence claim in campaign metadata, rejects PVE_TLS_INSECURE (install or provide the PVE CA with CERTISTACK_PVE_CA_SOURCE instead), and refuses a binary whose certistack version reports a dev, unknown, or dirty identity. Use the signed release binary or a Makefile build of a clean tagged checkout.

Fast lab smoke tests

By default, CertiStack reads and verifies every mapped PBS disk image before probing the restored service. This is the required production behavior and can take several minutes per large disk. A disposable lab can explicitly skip only that full-image pass. Add the following exact line to the protected lab environment file before using the test-only flag; the controller requires this explicit opt-in and does not accept an inherited value when --env-file is loaded:

CERTISTACK_ALLOW_TEST_MODE=1
certistack run /secure/certistack/lab-cases/single-example.yaml \
  --env-file /secure/certistack/lab.env --skip-integrity

Keep the env file on the controller outside the repository and any shared or Windows-mounted workspace. Never place live PVE or PBS credentials in a dotenv file under the source tree.

--skip-integrity is test-only. It still maps source disks read-only, creates the isolated restore, and executes all configured probes, but it does not compute the full mapped-image checksum. The signed JSON result is labeled test_skip_integrity and must not be used as production recovery or compliance evidence. Without the scan, the sandbox boots cold from PBS: Windows Server 2016 and later guests can miss the service manager's start deadline (troubleshooting Issue 2e). Run Windows cases without it. Linux guests that mount a data disk with a short device timeout (for example nofail,x-systemd.device-timeout=30s) can likewise skip that mount, and the service on it stays inactive, when many sandboxes boot cold at once. Test mode reads each disk's partition tables and labels before boot, but the delay is the guest's cold root disk, which only the scan warms. Run those cases without it too.

Fault drills

For an injected fault, run exactly one matching fault plan after confirming the workload is in its expected broken state:

scripts/run-lab-campaign.sh --node pve-node-01 --env /etc/certistack/lab.env \
  --certistack /usr/local/bin/certistack --mode fault --case fault-9924.yaml

A fault plan describes the verification to run; it does not inject a fault into a workload. Apply the matching fault in the disposable lab first. A zero exit from a fault case means the restored snapshot was healthy and is an unexpected campaign result, not proof that fault detection worked.

The campaign runner also rejects a non-zero fault run unless its saved controller report contains both an overall failure and a failed probe result. That prevents a transport, worker, or cleanup error from being misreported as successful fault detection.

Capacity safety guard

The local generator may include admission-*-guard.yaml plans that deliberately request capacity the node cannot admit. Run these separately with --mode guard:

scripts/run-lab-campaign.sh --node pve-node-01 --env /etc/certistack/lab.env \
  --certistack /usr/local/bin/certistack --mode guard

A guard case passes only when CertiStack is denied at node preflight or admission control and the log proves that no snapshot mapping, COW overlay, or sandbox VM was created. It is a safety test, not a failed recovery validation.

Generated cases and the compatibility audit

The coverage-case generator and the compatibility-matrix audit that feed a release campaign are described step by step in the lab acceptance runbook and release readiness.