stack.basicstack.de/apps/monitoring/README.md
CTO Agent 810d22f0a6 feat(monitoring): prometheus-backup → restic client-side encrypted (DEV-492)
Migrates the last plaintext-at-rest monitoring backup off rclone-sync
onto the same restic pipeline as loki/grafana/k8s-resources. Closes
the Hetzner Object Storage gap flagged in DEV-482 — no SSE-S3/SSE-KMS
means every object we push must be encrypted client-side.

- apps/monitoring/prometheus-backup-cronjob.yaml: pin the image to the
  Harbor mirror `harbor.basicstack.de/library/restic:0.17.3` (DEV-493)
  so the pull path matches the sibling CronJobs.
- docs/monitoring/restic-restore.md: append the DEV-492 drill entry —
  cold restic init, 8.017 GiB → 2.441 GiB stored, `restic check`
  clean, restore of `latest` into a scratch namespace, then
  `promtool tsdb list` + `promtool tsdb analyze` against all 19
  blocks (0 failed). Confirms the compaction-race mitigation
  (`--exclude wal/*`/`chunks_head/*` + accept exit 3) does not leave
  a corrupt snapshot. Also documents the DEV-492 Prometheus repo in
  the layout table + per-repo restore commands.
- apps/monitoring/README.md: point the restore section at all four
  repos and cite DEV-488 (loki/k8s) + DEV-492 (prometheus) drill
  entries.

Definition of done (DEV-492):
- backup ships client-side-encrypted to Hetzner:  (restic init
  succeeded on the manual run; snapshot 71420465 written)
- restore drill promtool tsdb analyze clean:  (19/19 blocks OK)
- OBSERVABILITY_BASELINE.md + apps/monitoring/README.md updated: 
- restic-password ownership documented:  (README shared-SealedSecret
  table already covered all four repos after this change)

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-08-16 16:29:22 +00:00

8 KiB

monitoring — backup CronJobs

Manifests recording the cluster-side monitoring backup CronJobs that were previously applied out-of-band. These files are the authoritative source (kubectl apply -f apps/monitoring/). See DEV-464 for the repair context.

  • backup-k8s-resources-cronjob.yaml — daily dump of cluster-scoped and per-namespace Kubernetes resources, streamed through restic backup --stdin to hetzner-s3:${BUCKET}/restic/k8s-resources. Uses serviceAccountName: backup-sa and no PVC mount (init container alpine/k8s:1.29.4 writes an emptyDir, main container restic/restic:0.17.3 reads it on stdin). Rewritten from the local-path tarball per the DEV-482 Option 4 rollout (DEV-487).
  • backup-loki-restic-cronjob.yaml — daily restic backup of loki-storage-encrypted to hetzner-s3:${BUCKET}/restic/loki. Co-schedules with the Loki pod via podAffinity (RWO permits additional read-only mounts on the same node). Deployed per the DEV-482 Option 4 rollout (DEV-485).
  • backup-grafana-restic-cronjob.yaml — daily restic backup of grafana-storage to hetzner-s3:${BUCKET}/restic/grafana. Pinned to k3s-worker-2 via nodeSelector (the local-path PV anchors the grafana pod there already, no podAffinity needed). Schedule 15 3 * * * — offset from the loki run at 03:00. Deployed per the DEV-482 Option 4 rollout (DEV-486).
  • prometheus-backup-cronjob.yaml + prometheus-backup-sealed.yaml — daily restic backup of prometheus-data-encrypted to hetzner-s3:${BUCKET}/restic/prometheus. Co-schedules with the Prometheus pod via podAffinity so the RWO PVC attaches on the same node. Schedule 30 3 * * * — offset from the loki (03:00) and grafana (03:15) runs. Migrated from the DEV-465 rclone sync job to restic client-side encryption per DEV-492 / DEV-482 Option 4. Compaction-race mitigation: --exclude wal/* + --exclude chunks_head/* + accept restic backup exit code 3 (source file vanished mid-walk) as a warning, not a failure; restore drill re-runs promtool tsdb analyze per block.
  • backup-restic-alerts.yaml + backup-restic-alerts.test.yamlPrometheusRule with freshness, integrity, and repo-size alerts covering the four restic repos (loki/grafana/k8s-resources/prometheus), plus a promtool test rules unit test proving each alert fires against synthetic samples (DEV-490, extended in DEV-492).

Shared SealedSecret

All four restic CronJobs read Hetzner S3 credentials + the restic repository password from SealedSecret monitoring-s3-backup (namespace monitoring). Keys:

Key Purpose
access-key Hetzner Object Storage access key ID
secret-key Hetzner Object Storage secret access key
endpoint S3 endpoint hostname (e.g. fsn1.your-objectstorage.com)
bucket Bucket name (single bucket, per-prefix repos)
restic-password 32-byte random string sealed at DEV-484; plaintext copy in Passbolt entry restic / monitoring backups. Rotation: restic key add → seal new value → restic key remove old id. Same key protects all four repos (loki/grafana/k8s-resources/prometheus) — rotating rewrites the key file on every repo.

Restore

Restore procedure for all four restic repos: docs/monitoring/restic-restore.md. First drill (2026-08-16, loki/k8s-resources) passed — see DEV-488. Prometheus repo added and drilled same day — see DEV-492.

Emitted metrics (textfile-collector format)

Each restic CronJob writes /metrics/backup_<kind>.prom (atomic — .prom.tmp + mv) into a hostPath volume mounted at the node-exporter textfile-collector directory (/var/lib/node_exporter/textfile_collector). The kube-prometheus-stack node-exporter DaemonSet has --collector.textfile.directory=/host/textfile_collector enabled (DEV-494, applied via apps/observability/patches/node-exporter-textfile-collector.yaml) and surfaces those samples in Prometheus.

Metric Emitted by
backup_<kind>_success (0/1) backup-{loki,grafana,k8s-resources}-*-cronjob.yaml, prometheus-backup-cronjob.yaml
backup_<kind>_timestamp_seconds same
backup_<kind>_check_status (exit code) same — from restic check --read-data-subset=5%
backup_prometheus_backup_status prometheus-backup-cronjob.yamlrestic backup exit code (3 = accepted compaction race)
restic_repo_size_bytes{repo="<kind>"} same — from restic stats --json --mode raw-data (added in DEV-490)

Cross-node staleness note. Because a backup CronJob may run on a different worker across days (loki/prometheus follow their app pods, k8s-resources is unpinned), a .prom file can linger on a node the job has since left and node-exporter keeps exposing it. The freshness alerts collapse the per-node samples with max() so the freshest sample wins; check/size alerts fire when any node reports a bad value, which is intentional — a recent failure is still a signal until the file is manually cleaned or the job returns to that node.

Alerts

backup-restic-alerts.yaml defines nine alerts (all severity: warning):

  • BackupLokiStale / BackupGrafanaStale / BackupK8sResourcesStale / BackupPrometheusStaletime() - max(backup_<kind>_timestamp_seconds) > 28h. Daily schedule + 4 h grace. max() collapses per-node samples so a stale .prom on a node the job has left does not fire.
  • BackupLokiCheckFailed / BackupGrafanaCheckFailed / BackupK8sResourcesCheckFailed / BackupPrometheusCheckFailedbackup_<kind>_check_status != 0.
  • ResticRepoOversizerestic_repo_size_bytes > 20 GiB. Baseline expected < 5 GiB; catches retention/prune regressions.

All alerts carry a Runbook: docs/monitoring/restic-restore.md annotation. To iterate on the rule file locally:

awk '/^spec:/{f=1;next} f{sub(/^  /,"");print}' \
  apps/monitoring/backup-restic-alerts.yaml > /tmp/backup-restic-rules.yaml
promtool check rules /tmp/backup-restic-rules.yaml
promtool test rules  apps/monitoring/backup-restic-alerts.test.yaml

The legacy backup-volumes CronJob and its 100 Gi local-path backup-storage PVC were retired in DEV-489 once the restic pipeline was proven end-to-end. With the shared destination PVC gone, the DEV-483 bridge nodeSelector pinning Loki to k3s-worker-2 was also removed — backup-loki-restic follows the Loki pod via podAffinity regardless of which node the RWO CSI volume lands on. backup-grafana-restic still nodeSelects k3s-worker-2 because its source PV (grafana-storage, local-path) is anchored there. backup-k8s-resources has no PVC dep and stays unpinned. prometheus-backup uses podAffinity on app=prometheus (RWO PVC on Hetzner CSI, single-node attach) and follows the Prometheus pod between workers.