stack.basicstack.de/apps/monitoring
CTO Agent 0c1c05fc50 feat(monitoring): retire legacy backup-volumes CronJob + backup-storage PVC (DEV-489)
The restic pipeline (loki/grafana/k8s-resources) is proven end-to-end
after the DEV-488 restore drill, so the legacy rsync/tar pipeline into
the 100 Gi local-path `backup-storage` PVC is removed.

- Delete `apps/monitoring/backup-volumes-cronjob.yaml`.
- Remove the DEV-483 bridge `nodeSelector: k3s-worker-2` from
  `apps/monitoring/loki-deployment.yaml`. Loki's data protection now
  runs via `backup-loki-restic`, which follows the pod via podAffinity
  regardless of which node the RWO CSI volume attaches on. The
  `Recreate` rollout strategy stays — it is unrelated (avoids the
  attach-deadlock during a rollout). Resolves the RWO/nodeSelector
  attach race that was blocking DEV-478 weekly OS updates.
- Update `apps/monitoring/README.md` to drop the `backup-volumes`
  section, link the restic restore runbook, and record the pin
  removal.
- Clean stale coexistence comments in the restic/prometheus CronJob
  manifests now that the legacy job is gone.

Cluster-side (already applied out-of-band, since these manifests are
`kubectl apply`-based, not Argo-managed):
- `kubectl -n monitoring delete cronjob backup-volumes` -> NotFound.
- `kubectl -n monitoring delete pvc backup-storage` -> gone; local-path
  PV `pvc-d0db0ba9-8f89-4f66-9e65-d573ebe1085a` reclaimed automatically
  (Delete policy). Two stale pre-DEV-487 `backup-k8s-resources` job
  pods that still referenced the PVC were deleted to release the
  `pvc-protection` finalizer.
- `kubectl -n monitoring apply -f loki-deployment.yaml` -> Recreate
  rollout, new pod Ready in ~60s, no nodeSelector on the new spec.
- No `VolumeAttachment` for the retired PV.
- Restic CronJobs (`backup-loki-restic`, `backup-grafana-restic`,
  `backup-k8s-resources`, `prometheus-backup`) intact.

Pre-delete snapshots retained on k3s-cp-1 under
`/root/dev489-snapshots-20260816T155411Z/` for post-mortem.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-08-16 15:58:43 +00:00
..
backup-grafana-restic-cronjob.yaml feat(monitoring): retire legacy backup-volumes CronJob + backup-storage PVC (DEV-489) 2026-08-16 15:58:43 +00:00
backup-k8s-resources-cronjob.yaml feat(monitoring): retire legacy backup-volumes CronJob + backup-storage PVC (DEV-489) 2026-08-16 15:58:43 +00:00
backup-loki-restic-cronjob.yaml feat(monitoring): retire legacy backup-volumes CronJob + backup-storage PVC (DEV-489) 2026-08-16 15:58:43 +00:00
loki-deployment.yaml feat(monitoring): retire legacy backup-volumes CronJob + backup-storage PVC (DEV-489) 2026-08-16 15:58:43 +00:00
prometheus-backup-cronjob.yaml feat(monitoring): retire legacy backup-volumes CronJob + backup-storage PVC (DEV-489) 2026-08-16 15:58:43 +00:00
prometheus-backup-sealed.yaml feat(monitoring): seal restic-password into monitoring-s3-backup (DEV-484) 2026-08-16 13:22:10 +00:00
README.md feat(monitoring): retire legacy backup-volumes CronJob + backup-storage PVC (DEV-489) 2026-08-16 15:58:43 +00:00

monitoring — backup CronJobs

Manifests recording the cluster-side monitoring backup CronJobs that were previously applied out-of-band. These files are the authoritative source (kubectl apply -f apps/monitoring/). See DEV-464 for the repair context.

  • backup-k8s-resources-cronjob.yaml — daily dump of cluster-scoped and per-namespace Kubernetes resources, streamed through restic backup --stdin to hetzner-s3:${BUCKET}/restic/k8s-resources. Uses serviceAccountName: backup-sa and no PVC mount (init container alpine/k8s:1.29.4 writes an emptyDir, main container restic/restic:0.17.3 reads it on stdin). Rewritten from the local-path tarball per the DEV-482 Option 4 rollout (DEV-487).
  • backup-loki-restic-cronjob.yaml — daily restic backup of loki-storage-encrypted to hetzner-s3:${BUCKET}/restic/loki. Co-schedules with the Loki pod via podAffinity (RWO permits additional read-only mounts on the same node). Deployed per the DEV-482 Option 4 rollout (DEV-485).
  • backup-grafana-restic-cronjob.yaml — daily restic backup of grafana-storage to hetzner-s3:${BUCKET}/restic/grafana. Pinned to k3s-worker-2 via nodeSelector (the local-path PV anchors the grafana pod there already, no podAffinity needed). Schedule 15 3 * * * — offset from the loki run at 03:00. Deployed per the DEV-482 Option 4 rollout (DEV-486).
  • prometheus-backup-cronjob.yaml + prometheus-backup-sealed.yaml — dedicated Prometheus data backup that streams prometheus-data-encrypted to Hetzner S3 via rclone. Co-schedules with the Prometheus pod via podAffinity so the RWO PVC attaches on the same node (DEV-465).

Restore procedure for all three restic repos: docs/monitoring/restic-restore.md. First drill (2026-08-16) passed — see DEV-488.

The legacy backup-volumes CronJob and its 100 Gi local-path backup-storage PVC were retired in DEV-489 once the restic pipeline was proven end-to-end. With the shared destination PVC gone, the DEV-483 bridge nodeSelector pinning Loki to k3s-worker-2 was also removed — backup-loki-restic follows the Loki pod via podAffinity regardless of which node the RWO CSI volume lands on. backup-grafana-restic still nodeSelects k3s-worker-2 because its source PV (grafana-storage, local-path) is anchored there. backup-k8s-resources and prometheus-backup remain unpinned.