stack.basicstack.de/docs/monitoring/restic-restore.md
CTO Agent 49b724439b docs(monitoring): add restic restore runbook + first drill log (DEV-488)
- Cover credentials projection via scratch ns + monitoring-s3-backup
  copy, per-repo restore commands (loki / k8s-resources / grafana),
  repo self-check + source-parity flow, cleanup, and rotation.
- Record first drill (2026-08-16): loki snapshot 97092888 and
  k8s-resources snapshot 9b0155cf restored; sha256 parity vs live
  loki-storage-encrypted PVC confirmed for all 4 files; cluster.yaml
  sanity-checked. Outcome: PASS.
- Next drill target: 2026-11-16 (quarterly); include restic/grafana
  once its first CronJob run has produced >= 1 snapshot.

Unblocks retirement of the legacy backup-volumes CronJob (DEV-489).

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-08-16 15:47:45 +00:00

9.6 KiB

Restic Restore Runbook — Hetzner Object Storage

Author: CTO agent (4b5f09a2-22d8-4e3d-9ac4-46d008ad1385) Ticket: DEV-488 Parent: DEV-482 (Option 4 — restic → Hetzner OS)

The monitoring stack's data-durability layer is a client-side-encrypted restic repo set stored in a single Hetzner Object Storage bucket. This document is the operator playbook for restoring those snapshots into a scratch namespace when we need to (a) recover from a real loss event, or (b) run the quarterly restore drill that keeps the backup-volumes retirement plan honest.

Repo layout

Single bucket, three repo prefixes:

repo prefix source writer CronJob tag
restic/loki PVC loki-storage-encrypted backup-loki-restic loki
restic/grafana PVC grafana-storage (worker-2) backup-grafana-restic grafana
restic/k8s-resources kubectl get -o yaml (stdin) backup-k8s-resources k8s-resources

Password + endpoint + bucket + AWS creds live in SealedSecret monitoring-s3-backup (namespace monitoring). Keys:

  • restic-password — 32-byte random, generated in DEV-484, sealed locally, mirrored to Passbolt entry restic / monitoring backups.
  • access-key, secret-key — Hetzner S3 credentials.
  • endpoint — e.g. https://hel1.your-objectstorage.com.
  • bucket — Hetzner bucket name (also used by the legacy Prometheus rclone job).

Encryption is done client-side by restic. Hetzner Object Storage has no SSE-S3 / SSE-KMS (FAQ); the CTO's Hetzner Object Storage encryption memory covers the constraints.

Credentials projection

Any restore pod must project only monitoring-s3-backup — never mount cluster admin creds into the drill namespace. The pattern below copies the secret cross-namespace once and lets restic project it read-only.

kubectl create namespace restore-drill
kubectl -n monitoring get secret monitoring-s3-backup -o json |
  jq 'del(.metadata.namespace,.metadata.resourceVersion,.metadata.uid,
          .metadata.creationTimestamp,.metadata.ownerReferences,
          .metadata.annotations,.status) |
      .metadata.namespace="restore-drill"' |
  kubectl apply -n restore-drill -f -

Restore pod

Throwaway pod pinned to the API-server node (no data mount, just the S3 secret). The emptyDir holds the restored tree; deleting the pod wipes it.

apiVersion: v1
kind: Pod
metadata:
  name: restore-drill
  namespace: restore-drill
  labels: { app: restore-drill }
spec:
  restartPolicy: Never
  containers:
    - name: restic
      image: restic/restic:0.17.3            # matches CronJob image
      command: ["sleep", "3600"]
      env:
        - { name: AWS_ACCESS_KEY_ID,     valueFrom: { secretKeyRef: { name: monitoring-s3-backup, key: access-key } } }
        - { name: AWS_SECRET_ACCESS_KEY, valueFrom: { secretKeyRef: { name: monitoring-s3-backup, key: secret-key } } }
        - { name: RESTIC_PASSWORD,       valueFrom: { secretKeyRef: { name: monitoring-s3-backup, key: restic-password } } }
        - { name: S3_ENDPOINT,           valueFrom: { secretKeyRef: { name: monitoring-s3-backup, key: endpoint } } }
        - { name: S3_BUCKET,             valueFrom: { secretKeyRef: { name: monitoring-s3-backup, key: bucket } } }
      volumeMounts:
        - { name: work, mountPath: /work }
  volumes:
    - name: work
      emptyDir: {}

Apply it, then kubectl -n restore-drill exec restore-drill -- <cmd> for everything below.

Per-repo restore commands

RESTIC_REPOSITORY is set inline per repo to keep repo scope explicit. All commands run inside the scratch pod.

Loki

export RESTIC_REPOSITORY="s3:${S3_ENDPOINT}/${S3_BUCKET}/restic/loki"
restic snapshots
restic restore latest --target /work/loki --tag loki
find /work/loki/source -type f

Restored tree lands under /work/loki/source/… because the backup source was /source inside the CronJob container.

K8s resources

export RESTIC_REPOSITORY="s3:${S3_ENDPOINT}/${S3_BUCKET}/restic/k8s-resources"
restic snapshots
restic restore latest --target /work/k8s --tag k8s-resources
ls -lh /work/k8s/cluster.yaml

cluster.yaml is the concatenated dump written by the backup-k8s-resources CronJob's kubectl-dump init container.

Grafana

export RESTIC_REPOSITORY="s3:${S3_ENDPOINT}/${S3_BUCKET}/restic/grafana"
restic snapshots
restic restore latest --target /work/grafana --tag grafana

Follows the loki pattern (restored tree under /work/grafana/source/).

Integrity checks

Two layers.

Repository self-check — cryptographic integrity of packs and snapshots, no source needed:

restic check                       # metadata only
restic check --read-data-subset=5% # sample-read + decrypt 5% of packs

The CronJobs run --read-data-subset=5% on every write. A restore drill should re-run at minimum restic check end-to-end.

Source parity — sha256 the restored tree against a fresh readout of the live PVC. For loki, exec into the live pod:

LOKI=$(kubectl -n monitoring get pods -l app=loki -o jsonpath='{.items[0].metadata.name}')
kubectl -n monitoring exec "$LOKI" -- sh -c '
  cd /loki && find . -type f ! -path "./lost+found/*" -exec sha256sum {} +' \
  > /tmp/live.sha256
kubectl -n restore-drill exec restore-drill -- sh -c '
  cd /work/loki/source && find . -type f -exec sha256sum {} +' \
  > /tmp/restored.sha256
diff <(sort /tmp/live.sha256) <(sort /tmp/restored.sha256)

Empty diff = parity. Any drift is expected only on files that Loki is actively writing (WAL segments) — capture the snapshot when Loki is quiet, or accept a WAL delta and re-hash the boltdb-shipper / chunks trees only.

For k8s-resources, parity is meaningless (a live dump varies by the second); rely on restic check and on being able to grep the restored cluster.yaml for expected namespaces / secrets.

Cleanup

kubectl -n restore-drill delete pod restore-drill --wait
kubectl delete namespace restore-drill

Deleting the namespace tears down the projected secret. The emptyDir lives in the pod's ephemeral scratch on the node; the sleep container never persists anything outside /work.

Rotation and disaster recovery

If the restic-password is lost:

  1. Recover the plaintext from Passbolt (restic / monitoring backups).
  2. If Passbolt is also lost, the data is unrecoverable by design — client-side encryption with a lost key cannot be reversed.

To rotate:

# on any workstation with kubectl + kubeseal + the restic password
NEW=$(openssl rand -base64 32)
restic key add                       # opens editor for new password;
                                     # paste NEW and confirm
restic key list
restic key remove <old-id>
# then re-seal the SealedSecret with NEW, redeploy, Argo syncs.

Rotate annually or immediately on suspected compromise.

Drill log

Every drill appends to this section. Include:

  • date (UTC), operator, snapshot IDs restored, parity result, restic check result, cleanup confirmation.

2026-08-16 — First drill (DEV-488)

  • Operator: CTO agent, run DEV-488.

  • Cluster: k3s at 178.105.17.239 (Hetzner).

  • Snapshots:

    • restic/loki: 97092888 (2026-08-16 15:44:21Z), parent 98c6fee6 (2026-08-16 15:26:10Z). Both tag=loki, host=k3s, paths=/source.
    • restic/k8s-resources: 9b0155cf (2026-08-16 15:40:28Z), tag=k8s-resources, paths=/cluster.yaml.
  • Restore output:

    • loki: Restored 14 files/dirs (292 B) in 0:00 — matches live PVC which is nearly empty (fresh Loki, no ingested chunks yet).
    • k8s-resources: Restored 1 files/dirs (11.167 MiB) in 0:00cluster.yaml = 11 709 295 bytes, sha256 5aa93f7326b5cbd2408cb763395bba259de6cc0f99b0c8e1bca4421edc7931c0.
  • Parity — loki restored tree vs. live loki-storage-encrypted PVC:

    file restored sha256 live sha256 match
    boltdb-shipper-active/uploader/name 51215284bd61bf79d48ccd1aec59445bfa28dec7b2cdce4d9476bc0b61342813 identical yes
    chunks/loki_cluster_seed.json 8d214a355d0057586e1c142d5f5d2658ccec39bd66e193b37dc12c21d6d7edd5 identical yes
    wal/checkpoint.019215/00000000 e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 identical yes
    wal/00019216 e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 identical yes
  • k8s-resources sanity: cluster.yaml opens with a valid v1 Namespace list including argocd, basicstack-web, monitoring etc. No parse errors.

  • restic check --read-data-subset=5% was run by the writing CronJob on both repos; drill relied on that. Re-running end-to-end restic check is a follow-up for the quarterly drill.

  • Cleanup: pod + restore-drill namespace deleted; scratch tree discarded with the emptyDir.

Outcome: PASS for restic/loki and restic/k8s-resources.

Next drill target: 2026-11-16 (quarterly). Include restic/grafana once its first CronJob run has produced ≥ 1 snapshot.