stack.basicstack.de/apps/monitoring/README.md
CTO Agent e2fd90a22b feat(monitoring): backup-loki-restic CronJob → Hetzner S3 (DEV-485)
Step 2 of the DEV-482 Option 4 rollout. Adds a restic-based Loki
backup that streams the loki-storage-encrypted PVC (mounted RO) to
s3:${S3_ENDPOINT}/${S3_BUCKET}/restic/loki, client-side encrypted by
restic. Co-schedules with the Loki pod via podAffinity so it lands on
whichever worker holds the RWO VolumeAttachment.

Verified with a manual run in-cluster:
- snapshot 98c6fee6 written, listable via a fresh restic pod
- restic check --read-data-subset=5% clean
- resource requests dropped from the plan's 200m/256Mi to 100m/128Mi
  because worker-2 has ~150m free CPU (Loki + Grafana + backup-* live
  there); limits stay generous for pack/check bursts.

Runs in parallel with the legacy backup-volumes CronJob — DEV-482
step 6 will retire that job only after the restore drill passes.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-08-16 15:27:40 +00:00

1.8 KiB

monitoring — backup CronJobs

Manifests recording the cluster-side monitoring backup CronJobs that were previously applied out-of-band. These files are the authoritative source (kubectl apply -f apps/monitoring/). See DEV-464 for the repair context.

  • backup-k8s-resources-cronjob.yaml — daily dump of Kubernetes resources into backup-storage PVC.
  • backup-volumes-cronjob.yaml — daily rsync/tar of Grafana + Loki PVCs into backup-storage. Prometheus data is NOT included here — it lives on a different node (see below). Being retired by the restic pipeline in DEV-482 — keep running until step 6 (restore drill passed).
  • backup-loki-restic-cronjob.yaml — daily restic backup of loki-storage-encrypted to hetzner-s3:${BUCKET}/restic/loki. Co-schedules with the Loki pod via podAffinity (RWO permits additional read-only mounts on the same node). Deployed in parallel with backup-volumes per the DEV-482 Option 4 rollout (DEV-485).
  • prometheus-backup-cronjob.yaml + prometheus-backup-sealed.yaml — dedicated Prometheus data backup that streams prometheus-data-encrypted to Hetzner S3 via rclone. Co-schedules with the Prometheus pod via podAffinity so the RWO PVC attaches on the same node (DEV-465).

The backup-storage PVC (100Gi, local-path, bound to k3s-worker-2) is the shared destination for backup-k8s-resources and backup-volumes.

backup-k8s-resources and backup-volumes pin themselves to k3s-worker-2 via nodeSelector because that is the node that holds all destination + source PVCs used there. prometheus-backup follows the Prometheus pod via podAffinity, writing to Hetzner S3 (hetzner-s3:basicstack-backup/prometheus/) so it stays independent of backup-storage.