Lab campaigns
A campaign runs many plans through the same controller-to-worker path and
records signed, attributable evidence for each one. This page covers the
campaign runner (scripts/run-lab-campaign.sh), the test-only shortcuts that
exist for disposable labs, fault drills, and capacity guards. Running a single
plan is covered in getting started; the acceptance and
release process that consumes campaign evidence is in the
lab acceptance runbook and
release readiness.
Run a local lab campaign
Place reviewed, synthetic lab plans in a protected controller-local case
directory. The campaign script runs on the management controller and invokes
the same remote certistack run flow for every case; it is not installed or
run directly on a PVE host. A release suite should cover individual replicas,
each service and OS family, pairwise healthy service families, broader mixed
tiers, and isolated fault cases. Keep site-specific VM IDs, networks, plans,
and generated cases out of the public repository.
cases=/secure/certistack/lab-cases
certistack validate "$cases"
scripts/run-lab-campaign.sh \
--node pve-node-01 \
--cases "$cases" \
--env /etc/certistack/lab.env \
--keyring /etc/certistack/keyring.json \
--certistack /usr/local/bin/certistack \
--mode positive \
--skip-integrity
The signing key must be controller-local and mode 0600; do not place it in a
shared, Windows-mounted, or source-controlled workspace. Use
CERTISTACK_SIGN_KEY during your own plan-generation process so unsigned
plans cannot become campaign evidence.
The runner records each case's output beneath the report directory and accepts
relative --output paths by normalizing them on the controller. Node preflight
is mandatory for every case; --skip-doctor is retained only as a deprecated
compatibility flag. The campaign completes the selected positive matrix and
prints a pass/fail summary even after an unexpected case failure; it exits
non-zero if any case failed. Use --fail-fast when immediate triage is more
useful than coverage. Its --env file is passed directly to CertiStack's
allowlisted literal dotenv parser; the campaign helper never shell-sources or
executes its contents.
The default positive matrix excludes diagnostic-*.yaml plans. Those are
explicit troubleshooting drills with known edge conditions; run one with
--case so it cannot lower the healthy-suite completion ratio.
Review exactly what a campaign will run, without contacting PVE or creating an output directory, with:
For a long positive campaign, resume its printed log directory after a controller interruption:
scripts/run-lab-campaign.sh --node pve-node-01 --cases /secure/certistack/lab-cases --env /etc/certistack/lab.env \
--certistack /usr/local/bin/certistack --mode positive \
--resume /var/lib/certistack/reports/campaign-logs/20260918T120000Z-positive
Resume never trusts a PASS line by itself. It skips a case only when the
stored signed report cryptographically verifies and proves the same plan and
validation mode passed with owned resources cleaned; incomplete, changed, or
unverifiable cases run again.
Every executable campaign requires --keyring to pin the report signer; a
report cannot establish trust in its own embedded public key. For production,
add --release-evidence. This mode rejects --skip-integrity, records the
release-evidence claim in campaign metadata, rejects PVE_TLS_INSECURE
(install or provide the PVE CA with CERTISTACK_PVE_CA_SOURCE instead), and
refuses a binary whose certistack version reports a dev, unknown, or
dirty identity. Use the signed release binary or a Makefile build of a clean
tagged checkout.
Fast lab smoke tests
By default, CertiStack reads and verifies every mapped PBS disk image before
probing the restored service. This is the required production behavior and can
take several minutes per large disk. A disposable lab can explicitly skip only
that full-image pass. Add the following exact line to the protected lab
environment file before using the test-only flag; the controller requires this
explicit opt-in and does not accept an inherited value when --env-file is
loaded:
certistack run /secure/certistack/lab-cases/single-example.yaml \
--env-file /secure/certistack/lab.env --skip-integrity
Keep the env file on the controller outside the repository and any shared or Windows-mounted workspace. Never place live PVE or PBS credentials in a dotenv file under the source tree.
--skip-integrity is test-only. It still maps source disks read-only, creates
the isolated restore, and executes all configured probes, but it does not
compute the full mapped-image checksum. The signed JSON result is labeled
test_skip_integrity and must not be used as production recovery or
compliance evidence. Without the scan, the sandbox boots cold from PBS:
Windows Server 2016 and later guests can miss the service manager's start
deadline (troubleshooting Issue 2e). Run Windows cases without it.
Linux guests that mount a data disk with a short device timeout (for
example nofail,x-systemd.device-timeout=30s) can likewise skip that mount,
and the service on it stays inactive, when many sandboxes boot cold at once.
Test mode reads each disk's partition tables and labels before boot, but
the delay is the guest's cold root disk, which only the scan warms. Run
those cases without it too.
Fault drills
For an injected fault, run exactly one matching fault plan after confirming the workload is in its expected broken state:
scripts/run-lab-campaign.sh --node pve-node-01 --env /etc/certistack/lab.env \
--certistack /usr/local/bin/certistack --mode fault --case fault-9924.yaml
A fault plan describes the verification to run; it does not inject a fault into a workload. Apply the matching fault in the disposable lab first. A zero exit from a fault case means the restored snapshot was healthy and is an unexpected campaign result, not proof that fault detection worked.
The campaign runner also rejects a non-zero fault run unless its saved controller report contains both an overall failure and a failed probe result. That prevents a transport, worker, or cleanup error from being misreported as successful fault detection.
Capacity safety guard
The local generator may include admission-*-guard.yaml plans that deliberately
request capacity the node cannot admit. Run these separately with --mode guard:
scripts/run-lab-campaign.sh --node pve-node-01 --env /etc/certistack/lab.env \
--certistack /usr/local/bin/certistack --mode guard
A guard case passes only when CertiStack is denied at node preflight or admission control and the log proves that no snapshot mapping, COW overlay, or sandbox VM was created. It is a safety test, not a failed recovery validation.
Generated cases and the compatibility audit
The coverage-case generator and the compatibility-matrix audit that feed a release campaign are described step by step in the lab acceptance runbook and release readiness.