Skip to content

Example plans

Three complete plans to copy from, and a guide to choosing probes. Every file on this page is in the repository's examples/ directory, is checked by the test suite against the current schema, and uses the RFC 5737 documentation range 192.0.2.0/24, so a copy can never point at a real network by accident.

The quickest route to a plan for your VMs is certistack init, which reads your PBS backups and writes a plan that already validates. Use these examples to go beyond its first probe, or to see how a plan for a particular kind of service is put together. Every field is defined in the test-plan reference.

To adapt an example:

cp three-tier-app.yaml my-app.yaml
# Change every value marked CHANGE.
certistack validate my-app.yaml
certistack plan my-app.yaml --env-file /secure/certistack/controller.env
certistack run my-app.yaml --env-file /secure/certistack/controller.env

validate needs no credentials. plan shows what a run would do, step by step, and changes nothing. Only then run.

A single VM

The smallest useful plan: boot the latest backup of one VM and prove that its SSH listener answers.

# SPDX-License-Identifier: AGPL-3.0-or-later
# Copyright (C) 2024-2026 Clift Cloud LLC
#
# The smallest useful CertiStack plan: boot the latest PBS backup of one VM in
# an isolated sandbox and prove that its SSH listener answers. Copy this file,
# change the values marked CHANGE, and run `certistack validate` on the copy.
#
# Addresses use the RFC 5737 documentation range 192.0.2.0/24. The sandbox is
# a disconnected L2 domain, so the range only has to be free of clashes with
# the guest's own expectations, not with your production network.
version: "v1"
plan_id: "minimal-single-vm"
name: "Minimal single-VM recovery check"
use_backup_config: true

compliance:
  frameworks: ["DORA-Article-12"]
  retention_days: 90
  # CHANGE: controller-local private key created with `certistack keygen`.
  sign_key_path: "/secure/certistack/controller-signing.ed25519"

admission_control:
  max_host_ram_percent: 85.0
  max_host_iowait_percent: 12.0
  defer_retry_interval_sec: 30
  max_defer_retries: 10

network:
  mode: "simple"
  zone_id: "csSbMin1"
  vnet_id: "csVnMin1"
  ip_range: "192.0.2.0/24"
  probe_ip: "192.0.2.1"

storage:
  # File-based PVE storage with "images" content for the disposable COW
  # overlays; stock Proxmox VE attaches only disks that belong to a storage.
  overlay_storage: "certistack-scratch"

tiers:
  - level: 1
    name: "Single VM"
    vms:
      # vmid is the temporary sandbox VM; it must not already exist on the node.
      - vmid: 9100
        # CHANGE: the protected VM whose latest PBS backup is being tested.
        source_vmid: 100
        name: "recovery-check-vm100"
        source_pbs:
          storage_id: "pbs"
          snapshot: "latest"
        probes:
          # The guest receives 192.0.2.10 through COW-only network recovery;
          # the source VM and its backup are never modified.
          - type: "tcp"
            target_ip: "192.0.2.10"
            port: 22
            timeout_sec: 60
            description: "SSH listener answers inside the sandbox"

A three-tier Linux application

A database, an application server and a web front end, restored in dependency order. Tier 2 starts only after every probe in tier 1 passed, and tier 3 after tier 2. It shows a TCP probe, a guest-agent service check, an HTTP health endpoint, verified HTTPS, and a recovery-time objective that fails the run if recovery takes too long.

# SPDX-License-Identifier: AGPL-3.0-or-later
# Copyright (C) 2024-2026 Clift Cloud LLC
#
# A three-tier Linux application restored in dependency order: database, then
# application server, then web front end. A tier starts only after every VM and
# probe in the tier before it passed. Copy this file, change the values marked
# CHANGE, and run `certistack validate` on the copy.
#
# Addresses use the RFC 5737 documentation range 192.0.2.0/24. The sandbox is
# a disconnected L2 domain, so the range only has to be free of clashes with
# the guests' own expectations, not with your production network. Each guest
# receives the address its probes target through COW-only network recovery.
version: "v1"
plan_id: "three-tier-app"
name: "Three-tier application recovery check"
use_backup_config: true

compliance:
  retention_days: 365
  # CHANGE: controller-local private key created with `certistack keygen`.
  sign_key_path: "/secure/certistack/controller-signing.ed25519"
  # Optional recovery-time objective: a run that takes longer is reported as failed.
  rto_target_seconds: 900

admission_control:
  max_host_ram_percent: 85.0
  max_host_iowait_percent: 12.0
  defer_retry_interval_sec: 30
  max_defer_retries: 10

network:
  mode: "simple"
  zone_id: "csSbApp1"
  vnet_id: "csVnApp1"
  ip_range: "192.0.2.0/24"
  probe_ip: "192.0.2.1"

storage:
  # CHANGE: file-based PVE storage with "images" content for the COW overlays.
  overlay_storage: "certistack-scratch"

tiers:
  - level: 1
    name: "Database"
    startup_grace_period_sec: 30
    vms:
      # vmid is the temporary sandbox VM; it must not already exist on the node.
      - vmid: 9101
        # CHANGE: the protected VM whose latest PBS backup is being tested.
        source_vmid: 201
        name: "recovery-check-db01"
        source_pbs:
          storage_id: "pbs"
          snapshot: "latest"
        # Uncomment to fail the VM unless its backup was taken with the guest
        # agent freezing the file systems (or while powered off).
        # require_quiesced_backup: true
        probes:
          - type: "tcp"
            target_ip: "192.0.2.21"
            port: 5432
            timeout_sec: 120
            description: "PostgreSQL accepts connections"
          - type: "qga"
            command: "systemctl is-active postgresql"
            expected_output: "active"
            timeout_sec: 60
            description: "PostgreSQL unit is active inside the guest"

  - level: 2
    name: "Application"
    startup_grace_period_sec: 30
    vms:
      - vmid: 9102
        source_vmid: 202
        name: "recovery-check-app01"
        source_pbs:
          storage_id: "pbs"
          snapshot: "latest"
        probes:
          - type: "http"
            target_ip: "192.0.2.22"
            port: 8080
            path: "/healthz"
            expected_status: 200
            timeout_sec: 120
            description: "Application health endpoint returns 200"

  - level: 3
    name: "Web"
    startup_grace_period_sec: 30
    vms:
      - vmid: 9103
        source_vmid: 203
        name: "recovery-check-web01"
        source_pbs:
          storage_id: "pbs"
          snapshot: "latest"
        probes:
          - type: "http"
            target_ip: "192.0.2.23"
            port: 443
            path: "/"
            tls: true
            # The certificate chain and this name are always verified; the
            # recovered certificate must chain to a CA the node trusts.
            server_name: "www.example.com"
            expected_status: 200
            timeout_sec: 120
            description: "Site answers over verified HTTPS"

Each guest receives the address its probes target through COW-only network recovery, so the three target_ip values must differ, stay inside network.ip_range, and differ from probe_ip.

A Windows domain

A domain controller first, then a SQL Server and a file server together. It shows an Active Directory Root DSE check, the DNS SRV record that locates a domain controller, the SQL Server and SMB protocol probes, and guest-agent sc query service checks.

# SPDX-License-Identifier: AGPL-3.0-or-later
# Copyright (C) 2024-2026 Clift Cloud LLC
#
# A Windows domain restored in dependency order: a domain controller first,
# then a SQL Server and a file server together. Copy this file, change the
# values marked CHANGE, and run `certistack validate` on the copy.
#
# Windows guests cannot have their network adapted offline, so each VM sets
# `network_recovery.mode: preserve`: the guest keeps its own address on the
# isolated VNet. Set `network.ip_range` to the subnet the guests use in
# production and each probe's `target_ip` to that guest's own address. The
# sandbox is a disconnected L2 domain, so reusing the production addressing
# cannot reach the production network. The 192.0.2.0/24 addresses here are
# documentation placeholders.
#
# A cold Windows boot from a backup is slow and its guest agent can take
# minutes to answer, so the probe budgets are generous. See the troubleshooting
# guide (Issues 2c, 2e and 2f) before shortening them.
version: "v1"
plan_id: "windows-domain"
name: "Windows domain recovery check"
use_backup_config: true

compliance:
  retention_days: 365
  # CHANGE: controller-local private key created with `certistack keygen`.
  sign_key_path: "/secure/certistack/controller-signing.ed25519"

admission_control:
  max_host_ram_percent: 85.0
  max_host_iowait_percent: 12.0
  defer_retry_interval_sec: 30
  max_defer_retries: 10

network:
  mode: "simple"
  zone_id: "csSbWin1"
  vnet_id: "csVnWin1"
  # CHANGE: the subnet the guests use in production.
  ip_range: "192.0.2.0/24"
  # A free address in ip_range that no guest uses.
  probe_ip: "192.0.2.250"

storage:
  # CHANGE: file-based PVE storage with "images" content. A UEFI guest needs
  # one, and Windows VMs are often UEFI.
  overlay_storage: "certistack-scratch"

tiers:
  - level: 1
    name: "Domain controller"
    startup_grace_period_sec: 60
    vms:
      - vmid: 9201
        # CHANGE: the protected domain controller.
        source_vmid: 301
        name: "recovery-check-dc01"
        source_pbs:
          storage_id: "pbs"
          snapshot: "latest"
        network_recovery:
          mode: "preserve"
        probes:
          - type: "ldap"
            target_ip: "192.0.2.10"
            port: 389
            # Reads the Root DSE only; AD refuses anonymous directory searches.
            naming_context: "DC=example,DC=com"
            timeout_sec: 300
            description: "Domain controller serves example.com"
          - type: "dns"
            target_ip: "192.0.2.10"
            port: 53
            domain: "_ldap._tcp.dc._msdcs.example.com"
            record_type: "srv"
            timeout_sec: 300
            description: "DC locator records are served"
          - type: "qga"
            command: "sc query NTDS"
            require_qga: true
            timeout_sec: 300
            description: "Active Directory Domain Services is running"

  - level: 2
    name: "Services"
    startup_grace_period_sec: 60
    vms:
      - vmid: 9202
        source_vmid: 302
        name: "recovery-check-sql01"
        source_pbs:
          storage_id: "pbs"
          snapshot: "latest"
        network_recovery:
          mode: "preserve"
        probes:
          - type: "mssql"
            target_ip: "192.0.2.20"
            port: 1433
            timeout_sec: 300
            description: "SQL Server answers its protocol"
          - type: "qga"
            command: "sc query MSSQLSERVER"
            require_qga: true
            timeout_sec: 300
            description: "SQL Server service is running"

      - vmid: 9203
        source_vmid: 303
        name: "recovery-check-fs01"
        source_pbs:
          storage_id: "pbs"
          snapshot: "latest"
        network_recovery:
          mode: "preserve"
        probes:
          - type: "smb"
            target_ip: "192.0.2.21"
            port: 445
            timeout_sec: 300
            description: "File server answers SMB"

Three things differ from a Linux plan:

  • network_recovery.mode: preserve. CertiStack cannot rewrite a Windows guest's network settings offline, so the guest keeps its own address. Each target_ip must be that guest's own address, inside network.ip_range, and sandbox_ip is not allowed. Without preserve, the run fails with an unsupported guest network layout error that names this setting.
  • Generous probe budgets. A cold Windows boot from a backup is slow. See troubleshooting before shortening them.
  • overlay_storage. Windows VMs are often UEFI, and a UEFI VM's firmware state is restored onto overlay storage, so the plan must name one.

Choosing probes

The first probe init writes proves the guest came up, not that it does its job. End every plan with a probe of the thing the VM exists for.

To prove that... Use Notes
The VM is running at all qmp Passes when the hypervisor reports the VM running. Says nothing about the guest OS.
A TCP service is listening (SSH, PostgreSQL, Redis, SMTP) tcp Proves the handshake, not that the service is healthy.
A web application or API answers http Asserts the status code. HTTPS always verifies the certificate chain and name; set server_name.
A DNS server answers, and serves a zone dns Set domain. Add record_type: srv to check SRV records.
An OpenLDAP or 389-ds directory serves data ldap base_dn: auto checks any advertised naming context; an explicit DN checks that exact entry.
An Active Directory domain controller serves its domain ldap with naming_context, plus dns with record_type: srv AD refuses anonymous searches, so naming_context reads only the Root DSE.
SQL Server answers mssql A protocol handshake; no login is attempted.
A file server answers smb A protocol negotiation; no share is opened.
A service is running inside the guest qga with systemctl is-active NAME or sc query NAME Needs the guest agent. require_qga: true makes a missing agent a failure instead of a skip.
Something on the screen, for the record screendump Passes when a capture succeeds. It asserts nothing about what is shown.

Wire probes (tcp, http, dns, ldap, mssql, smb) run on the node and connect to the guest over the isolated network; only qga runs inside the guest. Each signed result records which. When a probe fails, troubleshooting covers the common causes, and the signed report keeps the evidence.