- New PrometheusRule apps/monitoring/backup-restic-alerts.yaml with
seven warning-level rules:
* BackupLokiStale / BackupGrafanaStale / BackupK8sResourcesStale
-> time() - backup_<kind>_timestamp_seconds > 28h
* BackupLokiCheckFailed / BackupGrafanaCheckFailed /
BackupK8sResourcesCheckFailed -> backup_<kind>_check_status != 0
* ResticRepoOversize -> restic_repo_size_bytes > 20 GiB (per-repo)
- Emit restic_repo_size_bytes{repo="<kind>"} from all three restic
CronJobs (loki, grafana, k8s-resources) via `restic stats --json
--mode raw-data` so ResticRepoOversize has data to match once the
textfile-collector scrape path is wired.
- Companion promtool unit test backup-restic-alerts.test.yaml with
five scenarios (fresh/stale, check pass/fail, oversize) -- verified
locally with promtool 2.53.1: SUCCESS.
- README.md: document the three restic CronJobs, the shared
SealedSecret keys (access-key/secret-key/endpoint/bucket/
restic-password), the emitted textfile metrics, and the alert list.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
|
||
|---|---|---|
| .. | ||
| backup-grafana-restic-cronjob.yaml | ||
| backup-k8s-resources-cronjob.yaml | ||
| backup-loki-restic-cronjob.yaml | ||
| backup-restic-alerts.test.yaml | ||
| backup-restic-alerts.yaml | ||
| loki-deployment.yaml | ||
| prometheus-backup-cronjob.yaml | ||
| prometheus-backup-sealed.yaml | ||
| README.md | ||
monitoring — backup CronJobs
Manifests recording the cluster-side monitoring backup CronJobs that were previously applied out-of-band. These files are the authoritative source (kubectl apply -f apps/monitoring/). See DEV-464 for the repair context.
backup-k8s-resources-cronjob.yaml— daily dump of cluster-scoped and per-namespace Kubernetes resources, streamed throughrestic backup --stdintohetzner-s3:${BUCKET}/restic/k8s-resources. UsesserviceAccountName: backup-saand no PVC mount (init containeralpine/k8s:1.29.4writes an emptyDir, main containerrestic/restic:0.17.3reads it on stdin). Rewritten from the local-path tarball per the DEV-482 Option 4 rollout (DEV-487).backup-loki-restic-cronjob.yaml— daily restic backup ofloki-storage-encryptedtohetzner-s3:${BUCKET}/restic/loki. Co-schedules with the Loki pod viapodAffinity(RWO permits additional read-only mounts on the same node). Deployed per the DEV-482 Option 4 rollout (DEV-485).backup-grafana-restic-cronjob.yaml— daily restic backup ofgrafana-storagetohetzner-s3:${BUCKET}/restic/grafana. Pinned tok3s-worker-2vianodeSelector(the local-path PV anchors the grafana pod there already, nopodAffinityneeded). Schedule15 3 * * *— offset from the loki run at03:00. Deployed per the DEV-482 Option 4 rollout (DEV-486).prometheus-backup-cronjob.yaml+prometheus-backup-sealed.yaml— dedicated Prometheus data backup that streamsprometheus-data-encryptedto Hetzner S3 via rclone. Co-schedules with the Prometheus pod viapodAffinityso the RWO PVC attaches on the same node (DEV-465). Note: still usesrclone sync(plaintext at rest on Hetzner). Migration toresticorrclone cryptis filed as a follow-up.backup-restic-alerts.yaml+backup-restic-alerts.test.yaml—PrometheusRulewith freshness, integrity, and repo-size alerts covering the three restic repos, plus apromtool test rulesunit test proving each alert fires against synthetic samples (DEV-490).
Shared SealedSecret
All three restic CronJobs read Hetzner S3 credentials + the restic repository password from SealedSecret monitoring-s3-backup (namespace monitoring). Keys:
| Key | Purpose |
|---|---|
access-key |
Hetzner Object Storage access key ID |
secret-key |
Hetzner Object Storage secret access key |
endpoint |
S3 endpoint hostname (e.g. fsn1.your-objectstorage.com) |
bucket |
Bucket name (single bucket, per-prefix repos) |
restic-password |
32-byte random string sealed at DEV-484; plaintext copy in Passbolt entry restic / monitoring backups. Rotation: restic key add → seal new value → restic key remove old id. |
Restore
Restore procedure for all three restic repos: docs/monitoring/restic-restore.md. First drill (2026-08-16) passed — see DEV-488.
Emitted metrics (textfile-collector format)
Each restic CronJob writes to /metrics/backup_<kind>.prom inside the pod's emptyDir. Once node-exporter's textfile-collector path is wired up (follow-up), these become scrapeable and the alerts in backup-restic-alerts.yaml evaluate against live data.
| Metric | Emitted by |
|---|---|
backup_<kind>_success (0/1) |
backup-{loki,grafana,k8s-resources}-*-cronjob.yaml |
backup_<kind>_timestamp_seconds |
same |
backup_<kind>_check_status (exit code) |
same — from restic check --read-data-subset=5% |
restic_repo_size_bytes{repo="<kind>"} |
same — from restic stats --json --mode raw-data (added in DEV-490) |
Alerts
backup-restic-alerts.yaml defines seven alerts (all severity: warning):
BackupLokiStale/BackupGrafanaStale/BackupK8sResourcesStale—time() - backup_<kind>_timestamp_seconds > 28h. Daily schedule + 4 h grace.BackupLokiCheckFailed/BackupGrafanaCheckFailed/BackupK8sResourcesCheckFailed—backup_<kind>_check_status != 0.ResticRepoOversize—restic_repo_size_bytes > 20 GiB. Baseline expected < 5 GiB; catches retention/prune regressions.
All alerts carry a Runbook: docs/monitoring/restic-restore.md annotation. To iterate on the rule file locally:
awk '/^spec:/{f=1;next} f{sub(/^ /,"");print}' \
apps/monitoring/backup-restic-alerts.yaml > /tmp/backup-restic-rules.yaml
promtool check rules /tmp/backup-restic-rules.yaml
promtool test rules apps/monitoring/backup-restic-alerts.test.yaml
The legacy backup-volumes CronJob and its 100 Gi local-path backup-storage PVC were retired in DEV-489 once the restic pipeline was proven end-to-end. With the shared destination PVC gone, the DEV-483 bridge nodeSelector pinning Loki to k3s-worker-2 was also removed — backup-loki-restic follows the Loki pod via podAffinity regardless of which node the RWO CSI volume lands on. backup-grafana-restic still nodeSelects k3s-worker-2 because its source PV (grafana-storage, local-path) is anchored there. backup-k8s-resources and prometheus-backup remain unpinned.