Mirrored docker.io/restic/restic:0.19.1 to harbor.basicstack.de/library/restic:0.19.1 (crane in-cluster Job). Updated all CronJob pins and the restore-drill/tag-bump docs. Digest: sha256:136600b6ff6843d61d355f7f71f460a166429f35de6fd11b568fece3c9a4d510 Co-Authored-By: Paperclip <noreply@paperclip.ing>
8 KiB
monitoring — backup CronJobs
Manifests recording the cluster-side monitoring backup CronJobs that were previously applied out-of-band. These files are the authoritative source (kubectl apply -f apps/monitoring/). See DEV-464 for the repair context.
backup-k8s-resources-cronjob.yaml— daily dump of cluster-scoped and per-namespace Kubernetes resources, streamed throughrestic backup --stdintohetzner-s3:${BUCKET}/restic/k8s-resources. UsesserviceAccountName: backup-saand no PVC mount (init containeralpine/k8s:1.29.4writes an emptyDir, main containerrestic/restic:0.19.1reads it on stdin). Rewritten from the local-path tarball per the DEV-482 Option 4 rollout (DEV-487).backup-loki-restic-cronjob.yaml— daily restic backup ofloki-storage-encryptedtohetzner-s3:${BUCKET}/restic/loki. Co-schedules with the Loki pod viapodAffinity(RWO permits additional read-only mounts on the same node). Deployed per the DEV-482 Option 4 rollout (DEV-485).backup-grafana-restic-cronjob.yaml— daily restic backup ofgrafana-storagetohetzner-s3:${BUCKET}/restic/grafana. Pinned tok3s-worker-2vianodeSelector(the local-path PV anchors the grafana pod there already, nopodAffinityneeded). Schedule15 3 * * *— offset from the loki run at03:00. Deployed per the DEV-482 Option 4 rollout (DEV-486).prometheus-backup-cronjob.yaml+prometheus-backup-sealed.yaml— daily restic backup ofprometheus-data-encryptedtohetzner-s3:${BUCKET}/restic/prometheus. Co-schedules with the Prometheus pod viapodAffinityso the RWO PVC attaches on the same node. Schedule30 3 * * *— offset from the loki (03:00) and grafana (03:15) runs. Migrated from the DEV-465rclone syncjob to restic client-side encryption per DEV-492 / DEV-482 Option 4. Compaction-race mitigation:--exclude wal/*+--exclude chunks_head/*+ acceptrestic backupexit code 3 (source file vanished mid-walk) as a warning, not a failure; restore drill re-runspromtool tsdb analyzeper block.backup-restic-alerts.yaml+backup-restic-alerts.test.yaml—PrometheusRulewith freshness, integrity, and repo-size alerts covering the four restic repos (loki/grafana/k8s-resources/prometheus), plus apromtool test rulesunit test proving each alert fires against synthetic samples (DEV-490, extended in DEV-492).
Shared SealedSecret
All four restic CronJobs read Hetzner S3 credentials + the restic repository password from SealedSecret monitoring-s3-backup (namespace monitoring). Keys:
| Key | Purpose |
|---|---|
access-key |
Hetzner Object Storage access key ID |
secret-key |
Hetzner Object Storage secret access key |
endpoint |
S3 endpoint hostname (e.g. fsn1.your-objectstorage.com) |
bucket |
Bucket name (single bucket, per-prefix repos) |
restic-password |
32-byte random string sealed at DEV-484; plaintext copy in Passbolt entry restic / monitoring backups. Rotation: restic key add → seal new value → restic key remove old id. Same key protects all four repos (loki/grafana/k8s-resources/prometheus) — rotating rewrites the key file on every repo. |
Restore
Restore procedure for all four restic repos: docs/monitoring/restic-restore.md. First drill (2026-08-16, loki/k8s-resources) passed — see DEV-488. Prometheus repo added and drilled same day — see DEV-492.
Emitted metrics (textfile-collector format)
Each restic CronJob writes /metrics/backup_<kind>.prom (atomic — .prom.tmp + mv) into a hostPath volume mounted at the node-exporter textfile-collector directory (/var/lib/node_exporter/textfile_collector). The kube-prometheus-stack node-exporter DaemonSet has --collector.textfile.directory=/host/textfile_collector enabled (DEV-494, applied via apps/observability/patches/node-exporter-textfile-collector.yaml) and surfaces those samples in Prometheus.
| Metric | Emitted by |
|---|---|
backup_<kind>_success (0/1) |
backup-{loki,grafana,k8s-resources}-*-cronjob.yaml, prometheus-backup-cronjob.yaml |
backup_<kind>_timestamp_seconds |
same |
backup_<kind>_check_status (exit code) |
same — from restic check --read-data-subset=5% |
backup_prometheus_backup_status |
prometheus-backup-cronjob.yaml — restic backup exit code (3 = accepted compaction race) |
restic_repo_size_bytes{repo="<kind>"} |
same — from restic stats --json --mode raw-data (added in DEV-490) |
Cross-node staleness note. Because a backup CronJob may run on a different worker across days (loki/prometheus follow their app pods, k8s-resources is unpinned), a .prom file can linger on a node the job has since left and node-exporter keeps exposing it. The freshness alerts collapse the per-node samples with max() so the freshest sample wins; check/size alerts fire when any node reports a bad value, which is intentional — a recent failure is still a signal until the file is manually cleaned or the job returns to that node.
Alerts
backup-restic-alerts.yaml defines nine alerts (all severity: warning):
BackupLokiStale/BackupGrafanaStale/BackupK8sResourcesStale/BackupPrometheusStale—time() - max(backup_<kind>_timestamp_seconds) > 28h. Daily schedule + 4 h grace.max()collapses per-node samples so a stale.promon a node the job has left does not fire.BackupLokiCheckFailed/BackupGrafanaCheckFailed/BackupK8sResourcesCheckFailed/BackupPrometheusCheckFailed—backup_<kind>_check_status != 0.ResticRepoOversize—restic_repo_size_bytes > 20 GiB. Baseline expected < 5 GiB; catches retention/prune regressions.
All alerts carry a Runbook: docs/monitoring/restic-restore.md annotation. To iterate on the rule file locally:
awk '/^spec:/{f=1;next} f{sub(/^ /,"");print}' \
apps/monitoring/backup-restic-alerts.yaml > /tmp/backup-restic-rules.yaml
promtool check rules /tmp/backup-restic-rules.yaml
promtool test rules apps/monitoring/backup-restic-alerts.test.yaml
The legacy backup-volumes CronJob and its 100 Gi local-path backup-storage PVC were retired in DEV-489 once the restic pipeline was proven end-to-end. With the shared destination PVC gone, the DEV-483 bridge nodeSelector pinning Loki to k3s-worker-2 was also removed — backup-loki-restic follows the Loki pod via podAffinity regardless of which node the RWO CSI volume lands on. backup-grafana-restic still nodeSelects k3s-worker-2 because its source PV (grafana-storage, local-path) is anchored there. backup-k8s-resources has no PVC dep and stays unpinned. prometheus-backup uses podAffinity on app=prometheus (RWO PVC on Hetzner CSI, single-node attach) and follows the Prometheus pod between workers.