- New PrometheusRule apps/monitoring/backup-restic-alerts.yaml with
seven warning-level rules:
* BackupLokiStale / BackupGrafanaStale / BackupK8sResourcesStale
-> time() - backup_<kind>_timestamp_seconds > 28h
* BackupLokiCheckFailed / BackupGrafanaCheckFailed /
BackupK8sResourcesCheckFailed -> backup_<kind>_check_status != 0
* ResticRepoOversize -> restic_repo_size_bytes > 20 GiB (per-repo)
- Emit restic_repo_size_bytes{repo="<kind>"} from all three restic
CronJobs (loki, grafana, k8s-resources) via `restic stats --json
--mode raw-data` so ResticRepoOversize has data to match once the
textfile-collector scrape path is wired.
- Companion promtool unit test backup-restic-alerts.test.yaml with
five scenarios (fresh/stale, check pass/fail, oversize) -- verified
locally with promtool 2.53.1: SUCCESS.
- README.md: document the three restic CronJobs, the shared
SealedSecret keys (access-key/secret-key/endpoint/bucket/
restic-password), the emitted textfile metrics, and the alert list.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
55 lines
6 KiB
Markdown
55 lines
6 KiB
Markdown
# monitoring — backup CronJobs
|
|
|
|
Manifests recording the cluster-side monitoring backup CronJobs that were previously applied out-of-band. These files are the authoritative source (`kubectl apply -f apps/monitoring/`). See [DEV-464](/DEV/issues/DEV-464) for the repair context.
|
|
|
|
- `backup-k8s-resources-cronjob.yaml` — daily dump of cluster-scoped and per-namespace Kubernetes resources, streamed through `restic backup --stdin` to `hetzner-s3:${BUCKET}/restic/k8s-resources`. Uses `serviceAccountName: backup-sa` and no PVC mount (init container `alpine/k8s:1.29.4` writes an emptyDir, main container `restic/restic:0.17.3` reads it on stdin). Rewritten from the local-path tarball per the [DEV-482](/DEV/issues/DEV-482) Option 4 rollout ([DEV-487](/DEV/issues/DEV-487)).
|
|
- `backup-loki-restic-cronjob.yaml` — daily restic backup of `loki-storage-encrypted` to `hetzner-s3:${BUCKET}/restic/loki`. Co-schedules with the Loki pod via `podAffinity` (RWO permits additional read-only mounts on the same node). Deployed per the [DEV-482](/DEV/issues/DEV-482) Option 4 rollout ([DEV-485](/DEV/issues/DEV-485)).
|
|
- `backup-grafana-restic-cronjob.yaml` — daily restic backup of `grafana-storage` to `hetzner-s3:${BUCKET}/restic/grafana`. Pinned to `k3s-worker-2` via `nodeSelector` (the local-path PV anchors the grafana pod there already, no `podAffinity` needed). Schedule `15 3 * * *` — offset from the loki run at `03:00`. Deployed per the [DEV-482](/DEV/issues/DEV-482) Option 4 rollout ([DEV-486](/DEV/issues/DEV-486)).
|
|
- `prometheus-backup-cronjob.yaml` + `prometheus-backup-sealed.yaml` — dedicated Prometheus data backup that streams `prometheus-data-encrypted` to Hetzner S3 via rclone. Co-schedules with the Prometheus pod via `podAffinity` so the RWO PVC attaches on the same node ([DEV-465](/DEV/issues/DEV-465)). Note: still uses `rclone sync` (plaintext at rest on Hetzner). Migration to `restic` or `rclone crypt` is filed as a follow-up.
|
|
- `backup-restic-alerts.yaml` + `backup-restic-alerts.test.yaml` — `PrometheusRule` with freshness, integrity, and repo-size alerts covering the three restic repos, plus a `promtool test rules` unit test proving each alert fires against synthetic samples ([DEV-490](/DEV/issues/DEV-490)).
|
|
|
|
## Shared SealedSecret
|
|
|
|
All three restic CronJobs read Hetzner S3 credentials + the restic repository password from **SealedSecret `monitoring-s3-backup`** (namespace `monitoring`). Keys:
|
|
|
|
| Key | Purpose |
|
|
|------------------|------------------------------------------------------------------------------------------------------------------|
|
|
| `access-key` | Hetzner Object Storage access key ID |
|
|
| `secret-key` | Hetzner Object Storage secret access key |
|
|
| `endpoint` | S3 endpoint hostname (e.g. `fsn1.your-objectstorage.com`) |
|
|
| `bucket` | Bucket name (single bucket, per-prefix repos) |
|
|
| `restic-password`| 32-byte random string sealed at [DEV-484](/DEV/issues/DEV-484); plaintext copy in Passbolt entry `restic / monitoring backups`. Rotation: `restic key add` → seal new value → `restic key remove` old id. |
|
|
|
|
## Restore
|
|
|
|
Restore procedure for all three restic repos: [`docs/monitoring/restic-restore.md`](../../docs/monitoring/restic-restore.md). First drill (2026-08-16) passed — see [DEV-488](/DEV/issues/DEV-488).
|
|
|
|
## Emitted metrics (textfile-collector format)
|
|
|
|
Each restic CronJob writes to `/metrics/backup_<kind>.prom` inside the pod's emptyDir. Once node-exporter's textfile-collector path is wired up (follow-up), these become scrapeable and the alerts in `backup-restic-alerts.yaml` evaluate against live data.
|
|
|
|
| Metric | Emitted by |
|
|
|--------------------------------------------|---------------------------------------------------------------------|
|
|
| `backup_<kind>_success` (0/1) | `backup-{loki,grafana,k8s-resources}-*-cronjob.yaml` |
|
|
| `backup_<kind>_timestamp_seconds` | same |
|
|
| `backup_<kind>_check_status` (exit code) | same — from `restic check --read-data-subset=5%` |
|
|
| `restic_repo_size_bytes{repo="<kind>"}` | same — from `restic stats --json --mode raw-data` (added in DEV-490)|
|
|
|
|
## Alerts
|
|
|
|
`backup-restic-alerts.yaml` defines seven alerts (all `severity: warning`):
|
|
|
|
- `BackupLokiStale` / `BackupGrafanaStale` / `BackupK8sResourcesStale` — `time() - backup_<kind>_timestamp_seconds > 28h`. Daily schedule + 4 h grace.
|
|
- `BackupLokiCheckFailed` / `BackupGrafanaCheckFailed` / `BackupK8sResourcesCheckFailed` — `backup_<kind>_check_status != 0`.
|
|
- `ResticRepoOversize` — `restic_repo_size_bytes > 20 GiB`. Baseline expected < 5 GiB; catches retention/prune regressions.
|
|
|
|
All alerts carry a `Runbook: docs/monitoring/restic-restore.md` annotation. To iterate on the rule file locally:
|
|
|
|
```bash
|
|
awk '/^spec:/{f=1;next} f{sub(/^ /,"");print}' \
|
|
apps/monitoring/backup-restic-alerts.yaml > /tmp/backup-restic-rules.yaml
|
|
promtool check rules /tmp/backup-restic-rules.yaml
|
|
promtool test rules apps/monitoring/backup-restic-alerts.test.yaml
|
|
```
|
|
|
|
The legacy `backup-volumes` CronJob and its 100 Gi local-path `backup-storage` PVC were retired in [DEV-489](/DEV/issues/DEV-489) once the restic pipeline was proven end-to-end. With the shared destination PVC gone, the DEV-483 bridge `nodeSelector` pinning Loki to `k3s-worker-2` was also removed — `backup-loki-restic` follows the Loki pod via `podAffinity` regardless of which node the RWO CSI volume lands on. `backup-grafana-restic` still nodeSelects `k3s-worker-2` because its source PV (`grafana-storage`, local-path) is anchored there. `backup-k8s-resources` and `prometheus-backup` remain unpinned.
|