Skip to content

Automation and scheduling

A validation is worth most when it runs on its own, every night or every week, and someone hears when it fails. The Community Edition is a command-line tool, so the scheduler is yours: cron, a systemd timer, or your CI system. This page gives a wrapper script, a timer, and the rules that keep scheduled runs safe. Managed scheduling, dashboards, notifications and webhooks belong to the Enterprise Edition.

What one scheduled job does

  1. certistack run with a plan and the protected environment file.
  2. certistack verify-report on the report it produced, against your pinned keyring.
  3. Exit non-zero unless the recovery passed and cleanup was verified.

Step 2 matters more than it looks. verify-report exits 0 for any validly signed report, including one that records a failed recovery, so the job has to read all_passed and all_cleaned itself. See Reports and verification.

What the exit codes mean

run exit Meaning What to do
0 The recovery passed and cleanup was verified. Nothing; keep the report.
1 The recovery did not pass, the report is signed, and the node is clean. Read the report's failed probes. This is a real finding, not an outage of the tool.
2 The run could not complete: SSH or API trouble, a node busy with another operation, or cleanup that could not be verified. Read the error. If it says cleanup could not be verified, recover the node before the next run.
130 The run was interrupted: the job was stopped, the controller shut down, or a signal stopped the worker on the node (for example a node shutdown). Re-run it.

The full list is in the CLI reference.

The wrapper script

Save this as /usr/local/bin/certistack-validate and set the two paths at the top. It runs one plan, forwards a stop request to certistack so the node is cleaned up, refuses to overlap with a previous run of the same plan, verifies the report, and sets its exit status from the verified result. It needs bash, flock, sed and jq.

#!/usr/bin/env bash
# certistack-validate PLAN.yaml
# Runs one plan, verifies the signed report, and exits 0 only when the
# recovery passed and every temporary resource was verified removed.
set -uo pipefail

ENV_FILE=/secure/certistack/controller.env
KEYRING=/secure/certistack/trusted-signers.pub

plan=${1:?usage: certistack-validate PLAN.yaml}

# Do not overlap with a previous run of the same plan.
exec 9<"$plan"
flock -n 9 || { echo "a previous run of $plan is still active" >&2; exit 75; }

stdout=$(mktemp)
trap 'rm -f "$stdout"' EXIT

certistack run "$plan" --env-file "$ENV_FILE" >"$stdout" &
child=$!
# Forward a stop request so certistack cleans up the node before it exits.
trap 'kill -TERM "$child" 2>/dev/null' TERM INT HUP

status=
while [ -z "$status" ]; do
  wait "$child"; result=$?
  # Still running means a trapped signal ended the wait, not the child.
  kill -0 "$child" 2>/dev/null || status=$result
done

report=$(sed -n 's/^.*Controller report saved to: //p' "$stdout" | tail -n 1)
if [ -z "$report" ]; then
  echo "certistack run exited $status and produced no report" >&2
  exit $(( status == 0 ? 2 : status ))
fi

if ! summary=$(certistack verify-report --keyring "$KEYRING" --format json "$report"); then
  echo "report $report did not verify" >&2
  exit 2
fi
echo "$summary"

[ "$status" -eq 0 ] && [ "$(jq -r '.all_passed and .all_cleaned' <<<"$summary")" = true ] && exit 0
exit $(( status == 0 ? 1 : status ))
sudo install -m 0755 certistack-validate /usr/local/bin/certistack-validate

The script prints the one-line JSON summary to standard output and leaves certistack's progress timeline on standard error, so a scheduler's log holds both. Its own exit status is 0 for a verified pass, 1 for a verified failure, 75 when the previous run is still going, 2 when the report is missing or does not verify, and otherwise certistack's own status.

Run it from systemd

Create the account that owns the controller's files, and put the plans where the unit expects them:

sudo useradd --system --no-create-home --shell /usr/sbin/nologin certistack
sudo install -d -m 0700 -o certistack /secure/certistack
sudo install -d -m 0700 -o certistack /var/lib/certistack/plans

The key, the environment file, known_hosts and the SSH identity under /secure/certistack must be owned by certistack and unreadable to others. Then add a service template and a timer template, so one pair of files serves every plan:

# /etc/systemd/system/[email protected]
[Unit]
Description=CertiStack recovery validation of plan %i
Wants=network-online.target
After=network-online.target

[Service]
Type=oneshot
User=certistack
StateDirectory=certistack
StateDirectoryMode=0700
ExecStart=/usr/local/bin/certistack-validate /var/lib/certistack/plans/%i.yaml
# After a stop request certistack waits up to 20 minutes for the node to finish
# cleaning up; do not let systemd kill it sooner. KillMode=mixed signals only
# the wrapper, which forwards the stop to certistack.
KillMode=mixed
TimeoutStopSec=25min
# /etc/systemd/system/[email protected]
[Unit]
Description=Nightly CertiStack recovery validation of plan %i

[Timer]
OnCalendar=*-*-* 02:30:00
Persistent=true

[Install]
WantedBy=timers.target

Enable a timer for the plan /var/lib/certistack/plans/nightly.yaml:

sudo systemctl daemon-reload
sudo systemctl enable --now [email protected]
systemctl list-timers 'certistack-validate@*'

Run it once by hand, and read the result in the journal:

sudo systemctl start [email protected]
journalctl -u [email protected] -n 50

A failed run leaves the unit in the failed state. To be told, add an OnFailure= line to the service that starts a unit of your own (mail, a chat message, a ticket), or watch systemctl --failed from your monitoring. To give another plan a different time, override the timer with sudo systemctl edit [email protected] and set OnCalendar= (clear it first with an empty OnCalendar= line).

Run it from cron

# /etc/cron.d/certistack — m h dom mon dow user command
PATH=/usr/local/bin:/usr/bin:/bin
30 2 * * * certistack /usr/local/bin/certistack-validate /var/lib/certistack/plans/nightly.yaml

cron mails a job's output to the crontab's owner (or to MAILTO), and the wrapper writes its progress to standard error, so every run sends a mail. Redirect the output to a log file, and alert on a non-zero status yourself, for example by appending || your-alert-command to the job.

Run it from the controller container

The container image holds certistack only, so run run and verify-report as two container invocations, with the plans mounted at /plans and the secrets at /run/certistack-secrets as in the controller README:

30 2 * * * cd /opt/certistack && docker compose -f deploy/controller/compose.yaml run --rm -T controller run /plans/nightly.yaml

Reports land in state/reports, next to the Compose file. Verify the newest report from the host with the binary you installed, using the same verify-report command as above. The -T flag stops Compose from allocating a terminal, which cron does not have.

Rules for scheduled runs

  • One operation per node at a time. A node takes an exclusive lock (/run/certistack.lock) for every run and recovery. A second run that starts while one is active is refused with another CertiStack operation holds the host lock and exits 2. Stagger schedules that target the same node, and leave room for the slowest plan. plan and init only read and do not take the lock, but plan lists an active run's resources as work recovery would do first, so run it when nothing is active.
  • Distinct IDs for plans that overlap in time. The sandbox VM IDs, zone_id and vnet_id are cluster-wide. Plans that never run at the same time may reuse them; plans that can overlap, even on different nodes, need their own. A plan whose IDs already exist on the cluster fails closed rather than reuse them.
  • Run after the backup window. snapshot: latest resolves the newest backup when the run starts. A run during the backup window competes with the backup for PBS, and Windows guests in particular boot slowly then; see troubleshooting.
  • Keep sandbox VMs out of backup jobs. A job that selects all VMs can pick up a running sandbox VM. See the deployment model.
  • Do not upgrade under a running job. Finish or recover every active run before you change the installed binary; see Verify, install, upgrade, and roll back a release.
  • Keep every report. Copy the reports and the public keyring off the controller on the schedule your retention policy sets. CertiStack never deletes a report, so the directory grows until you prune it.

Test it before you trust it

Run the job by hand once and check the exit status. Then force the failure paths once, on a staging cluster: point a probe at a port nothing listens on and confirm the job exits 1 and your alert fires, and stop the job with systemctl stop mid-run and confirm the node is left clean (certistack inspect on the node shows a cleaned journal). An alert that has never fired is an assumption.