Automation and scheduling
A validation is worth most when it runs on its own, every night or every week,
and someone hears when it fails. The Community Edition is a command-line tool,
so the scheduler is yours: cron, a systemd timer, or your CI system. This
page gives a wrapper script, a timer, and the rules that keep scheduled runs
safe. Managed scheduling, dashboards, notifications and webhooks belong to the
Enterprise Edition.
What one scheduled job does
certistack runwith a plan and the protected environment file.certistack verify-reporton the report it produced, against your pinned keyring.- Exit non-zero unless the recovery passed and cleanup was verified.
Step 2 matters more than it looks. verify-report exits 0 for any validly
signed report, including one that records a failed recovery, so the job has to
read all_passed and all_cleaned itself. See
Reports and verification.
What the exit codes mean
run exit |
Meaning | What to do |
|---|---|---|
0 |
The recovery passed and cleanup was verified. | Nothing; keep the report. |
1 |
The recovery did not pass, the report is signed, and the node is clean. | Read the report's failed probes. This is a real finding, not an outage of the tool. |
2 |
The run could not complete: SSH or API trouble, a node busy with another operation, or cleanup that could not be verified. | Read the error. If it says cleanup could not be verified, recover the node before the next run. |
130 |
The run was interrupted: the job was stopped, the controller shut down, or a signal stopped the worker on the node (for example a node shutdown). | Re-run it. |
The full list is in the CLI reference.
The wrapper script
Save this as /usr/local/bin/certistack-validate and set the two paths at the
top. It runs one plan, forwards a stop request to certistack so the node is
cleaned up, refuses to overlap with a previous run of the same plan, verifies
the report, and sets its exit status from the verified result. It needs bash,
flock, sed and jq.
#!/usr/bin/env bash
# certistack-validate PLAN.yaml
# Runs one plan, verifies the signed report, and exits 0 only when the
# recovery passed and every temporary resource was verified removed.
set -uo pipefail
ENV_FILE=/secure/certistack/controller.env
KEYRING=/secure/certistack/trusted-signers.pub
plan=${1:?usage: certistack-validate PLAN.yaml}
# Do not overlap with a previous run of the same plan.
exec 9<"$plan"
flock -n 9 || { echo "a previous run of $plan is still active" >&2; exit 75; }
stdout=$(mktemp)
trap 'rm -f "$stdout"' EXIT
certistack run "$plan" --env-file "$ENV_FILE" >"$stdout" &
child=$!
# Forward a stop request so certistack cleans up the node before it exits.
trap 'kill -TERM "$child" 2>/dev/null' TERM INT HUP
status=
while [ -z "$status" ]; do
wait "$child"; result=$?
# Still running means a trapped signal ended the wait, not the child.
kill -0 "$child" 2>/dev/null || status=$result
done
report=$(sed -n 's/^.*Controller report saved to: //p' "$stdout" | tail -n 1)
if [ -z "$report" ]; then
echo "certistack run exited $status and produced no report" >&2
exit $(( status == 0 ? 2 : status ))
fi
if ! summary=$(certistack verify-report --keyring "$KEYRING" --format json "$report"); then
echo "report $report did not verify" >&2
exit 2
fi
echo "$summary"
[ "$status" -eq 0 ] && [ "$(jq -r '.all_passed and .all_cleaned' <<<"$summary")" = true ] && exit 0
exit $(( status == 0 ? 1 : status ))
The script prints the one-line JSON summary to standard output and leaves
certistack's progress timeline on standard error, so a scheduler's log holds
both. Its own exit status is 0 for a verified pass, 1 for a verified
failure, 75 when the previous run is still going, 2 when the report is
missing or does not verify, and otherwise certistack's own status.
Run it from systemd
Create the account that owns the controller's files, and put the plans where the unit expects them:
sudo useradd --system --no-create-home --shell /usr/sbin/nologin certistack
sudo install -d -m 0700 -o certistack /secure/certistack
sudo install -d -m 0700 -o certistack /var/lib/certistack/plans
The key, the environment file, known_hosts and the SSH identity under
/secure/certistack must be owned by certistack and unreadable to others.
Then add a service template and a timer template, so one pair of files serves
every plan:
# /etc/systemd/system/[email protected]
[Unit]
Description=CertiStack recovery validation of plan %i
Wants=network-online.target
After=network-online.target
[Service]
Type=oneshot
User=certistack
StateDirectory=certistack
StateDirectoryMode=0700
ExecStart=/usr/local/bin/certistack-validate /var/lib/certistack/plans/%i.yaml
# After a stop request certistack waits up to 20 minutes for the node to finish
# cleaning up; do not let systemd kill it sooner. KillMode=mixed signals only
# the wrapper, which forwards the stop to certistack.
KillMode=mixed
TimeoutStopSec=25min
# /etc/systemd/system/[email protected]
[Unit]
Description=Nightly CertiStack recovery validation of plan %i
[Timer]
OnCalendar=*-*-* 02:30:00
Persistent=true
[Install]
WantedBy=timers.target
Enable a timer for the plan /var/lib/certistack/plans/nightly.yaml:
sudo systemctl daemon-reload
sudo systemctl enable --now [email protected]
systemctl list-timers 'certistack-validate@*'
Run it once by hand, and read the result in the journal:
sudo systemctl start [email protected]
journalctl -u [email protected] -n 50
A failed run leaves the unit in the failed state. To be told, add an
OnFailure= line to the service that starts a unit of your own (mail, a chat
message, a ticket), or watch systemctl --failed from your monitoring. To give
another plan a different time, override the timer with
sudo systemctl edit [email protected] and set OnCalendar=
(clear it first with an empty OnCalendar= line).
Run it from cron
# /etc/cron.d/certistack — m h dom mon dow user command
PATH=/usr/local/bin:/usr/bin:/bin
30 2 * * * certistack /usr/local/bin/certistack-validate /var/lib/certistack/plans/nightly.yaml
cron mails a job's output to the crontab's owner (or to MAILTO), and the
wrapper writes its progress to standard error, so every run sends a mail.
Redirect the output to a log file, and alert on a non-zero status yourself, for
example by appending || your-alert-command to the job.
Run it from the controller container
The container image holds certistack only, so run run and verify-report as
two container invocations, with the plans mounted at /plans and the secrets
at /run/certistack-secrets as in the
controller README:
30 2 * * * cd /opt/certistack && docker compose -f deploy/controller/compose.yaml run --rm -T controller run /plans/nightly.yaml
Reports land in state/reports, next to the Compose file. Verify the newest
report from the host with the binary you installed, using the same
verify-report command as above. The -T flag stops Compose from allocating a
terminal, which cron does not have.
Rules for scheduled runs
- One operation per node at a time. A node takes an exclusive lock
(
/run/certistack.lock) for every run and recovery. A second run that starts while one is active is refused withanother CertiStack operation holds the host lockand exits2. Stagger schedules that target the same node, and leave room for the slowest plan.planandinitonly read and do not take the lock, butplanlists an active run's resources as work recovery would do first, so run it when nothing is active. - Distinct IDs for plans that overlap in time. The sandbox VM IDs,
zone_idandvnet_idare cluster-wide. Plans that never run at the same time may reuse them; plans that can overlap, even on different nodes, need their own. A plan whose IDs already exist on the cluster fails closed rather than reuse them. - Run after the backup window.
snapshot: latestresolves the newest backup when the run starts. A run during the backup window competes with the backup for PBS, and Windows guests in particular boot slowly then; see troubleshooting. - Keep sandbox VMs out of backup jobs. A job that selects all VMs can pick up a running sandbox VM. See the deployment model.
- Do not upgrade under a running job. Finish or recover every active run before you change the installed binary; see Verify, install, upgrade, and roll back a release.
- Keep every report. Copy the reports and the public keyring off the controller on the schedule your retention policy sets. CertiStack never deletes a report, so the directory grows until you prune it.
Test it before you trust it
Run the job by hand once and check the exit status. Then force the failure
paths once, on a staging cluster: point a probe at a port nothing listens on and
confirm the job exits 1 and your alert fires, and stop the job with systemctl
stop mid-run and confirm the node is left clean (certistack inspect on the
node shows a cleaned journal). An alert that has never fired is an
assumption.