From e2fd90a22b62046e8c963fbab995931ee86bdae3 Mon Sep 17 00:00:00 2001 From: CTO Agent Date: Sun, 16 Aug 2026 15:27:40 +0000 Subject: [PATCH] =?UTF-8?q?feat(monitoring):=20backup-loki-restic=20CronJo?= =?UTF-8?q?b=20=E2=86=92=20Hetzner=20S3=20(DEV-485)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Step 2 of the DEV-482 Option 4 rollout. Adds a restic-based Loki backup that streams the loki-storage-encrypted PVC (mounted RO) to s3:${S3_ENDPOINT}/${S3_BUCKET}/restic/loki, client-side encrypted by restic. Co-schedules with the Loki pod via podAffinity so it lands on whichever worker holds the RWO VolumeAttachment. Verified with a manual run in-cluster: - snapshot 98c6fee6 written, listable via a fresh restic pod - restic check --read-data-subset=5% clean - resource requests dropped from the plan's 200m/256Mi to 100m/128Mi because worker-2 has ~150m free CPU (Loki + Grafana + backup-* live there); limits stay generous for pack/check bursts. Runs in parallel with the legacy backup-volumes CronJob — DEV-482 step 6 will retire that job only after the restore drill passes. Co-Authored-By: Paperclip --- apps/monitoring/README.md | 3 +- .../backup-loki-restic-cronjob.yaml | 164 ++++++++++++++++++ 2 files changed, 166 insertions(+), 1 deletion(-) create mode 100644 apps/monitoring/backup-loki-restic-cronjob.yaml diff --git a/apps/monitoring/README.md b/apps/monitoring/README.md index f381914..5750b07 100644 --- a/apps/monitoring/README.md +++ b/apps/monitoring/README.md @@ -3,7 +3,8 @@ Manifests recording the cluster-side monitoring backup CronJobs that were previously applied out-of-band. These files are the authoritative source (`kubectl apply -f apps/monitoring/`). See [DEV-464](/DEV/issues/DEV-464) for the repair context. - `backup-k8s-resources-cronjob.yaml` — daily dump of Kubernetes resources into `backup-storage` PVC. -- `backup-volumes-cronjob.yaml` — daily rsync/tar of Grafana + Loki PVCs into `backup-storage`. Prometheus data is NOT included here — it lives on a different node (see below). +- `backup-volumes-cronjob.yaml` — daily rsync/tar of Grafana + Loki PVCs into `backup-storage`. Prometheus data is NOT included here — it lives on a different node (see below). Being retired by the restic pipeline in [DEV-482](/DEV/issues/DEV-482) — keep running until step 6 (restore drill passed). +- `backup-loki-restic-cronjob.yaml` — daily restic backup of `loki-storage-encrypted` to `hetzner-s3:${BUCKET}/restic/loki`. Co-schedules with the Loki pod via `podAffinity` (RWO permits additional read-only mounts on the same node). Deployed in parallel with `backup-volumes` per the [DEV-482](/DEV/issues/DEV-482) Option 4 rollout ([DEV-485](/DEV/issues/DEV-485)). - `prometheus-backup-cronjob.yaml` + `prometheus-backup-sealed.yaml` — dedicated Prometheus data backup that streams `prometheus-data-encrypted` to Hetzner S3 via rclone. Co-schedules with the Prometheus pod via `podAffinity` so the RWO PVC attaches on the same node ([DEV-465](/DEV/issues/DEV-465)). The `backup-storage` PVC (100Gi, local-path, bound to k3s-worker-2) is the shared destination for `backup-k8s-resources` and `backup-volumes`. diff --git a/apps/monitoring/backup-loki-restic-cronjob.yaml b/apps/monitoring/backup-loki-restic-cronjob.yaml new file mode 100644 index 0000000..40d53eb --- /dev/null +++ b/apps/monitoring/backup-loki-restic-cronjob.yaml @@ -0,0 +1,164 @@ +--- +# Loki data backup via restic to Hetzner Object Storage (DEV-485, +# DEV-482 Option 4). Step 2 of the Option 4 rollout. +# +# Streams the RWO PVC `loki-storage-encrypted` (mounted read-only) +# into `s3:${S3_ENDPOINT}/${S3_BUCKET}/restic/loki`, a client-side +# encrypted restic repository. Deployed in parallel with the legacy +# `backup-volumes` CronJob — do not retire that job until DEV-482 +# step 6 (restore drill passed). +# +# podAffinity co-schedules with the Loki pod (app=loki, topology +# kubernetes.io/hostname). RWO permits additional read-only mounts +# on the node that holds the PVC's VolumeAttachment, so this +# survives Loki being rescheduled to a different worker. +apiVersion: batch/v1 +kind: CronJob +metadata: + name: backup-loki-restic + namespace: monitoring + labels: + app: backup + type: loki + backend: restic +spec: + schedule: "0 3 * * *" + concurrencyPolicy: Forbid + successfulJobsHistoryLimit: 3 + failedJobsHistoryLimit: 3 + jobTemplate: + metadata: + labels: + app: backup + type: loki + backend: restic + spec: + backoffLimit: 2 + activeDeadlineSeconds: 3600 + template: + metadata: + labels: + app: backup + type: loki + backend: restic + spec: + restartPolicy: OnFailure + affinity: + podAffinity: + requiredDuringSchedulingIgnoredDuringExecution: + - labelSelector: + matchExpressions: + - key: app + operator: In + values: + - loki + topologyKey: kubernetes.io/hostname + containers: + - name: restic + image: restic/restic:0.17.3 + env: + - name: AWS_ACCESS_KEY_ID + valueFrom: + secretKeyRef: + name: monitoring-s3-backup + key: access-key + - name: AWS_SECRET_ACCESS_KEY + valueFrom: + secretKeyRef: + name: monitoring-s3-backup + key: secret-key + - name: RESTIC_PASSWORD + valueFrom: + secretKeyRef: + name: monitoring-s3-backup + key: restic-password + - name: S3_ENDPOINT + valueFrom: + secretKeyRef: + name: monitoring-s3-backup + key: endpoint + - name: S3_BUCKET + valueFrom: + secretKeyRef: + name: monitoring-s3-backup + key: bucket + - name: RESTIC_REPOSITORY + value: "s3:$(S3_ENDPOINT)/$(S3_BUCKET)/restic/loki" + command: + - /bin/sh + - -c + - | + set -eu + echo "=== backup-loki-restic started at $(date -u +%FT%TZ) ===" + echo "Repository: ${RESTIC_REPOSITORY}" + + # First-run tolerance: init if the repo isn't there yet. + # `restic cat config` is the tightest existence probe; use + # `snapshots` per plan spec — either exits 0 iff the repo + # is initialised. + if restic snapshots >/dev/null 2>&1; then + echo "Repo exists, skipping init." + else + echo "Repo missing, initialising..." + restic init + fi + + echo "--- restic backup /source ---" + restic backup /source \ + --tag loki \ + --host k3s \ + --exclude '*.tmp' + + echo "--- restic forget/prune ---" + restic forget --tag loki \ + --keep-daily 7 \ + --keep-weekly 4 \ + --keep-monthly 6 \ + --prune + + echo "--- restic check --read-data-subset=5% ---" + CHECK_STATUS=0 + restic check --read-data-subset=5% || CHECK_STATUS=$? + echo "restic check exit: ${CHECK_STATUS}" + + # Textfile-collector metrics. Written into an emptyDir per + # the current pattern used by backup-volumes and + # backup-k8s-resources. Once a node-exporter textfile + # collector path is wired up, these become scrapeable — + # see the follow-up notes in DEV-482. + { + echo "backup_loki_success $([ ${CHECK_STATUS} -eq 0 ] && echo 1 || echo 0)" + echo "backup_loki_timestamp_seconds $(date +%s)" + echo "backup_loki_check_status ${CHECK_STATUS}" + } > /metrics/backup_loki.prom + + echo "=== backup-loki-restic finished at $(date -u +%FT%TZ) ===" + exit ${CHECK_STATUS} + volumeMounts: + - name: loki-data + mountPath: /source + readOnly: true + - name: metrics + mountPath: /metrics + - name: cache + mountPath: /root/.cache/restic + resources: + # Requests deliberately lowered from the plan doc's + # 200m/256Mi — worker-2 (Loki node) has ~150m free CPU + # and podAffinity forces us onto it. 100m/128Mi mirrors + # the sibling backup CronJobs; limits stay generous so + # restic can burst during pack/check. + requests: + cpu: 100m + memory: 128Mi + limits: + cpu: 1500m + memory: 1Gi + volumes: + - name: loki-data + persistentVolumeClaim: + claimName: loki-storage-encrypted + - name: metrics + emptyDir: {} + - name: cache + emptyDir: {}