- Create the missing forgejo/platform-backup-data PVC that forgejo-backup references (20Gi hcloud-volumes-encrypted). - Record monitoring/backup-k8s-resources with a k3s-worker-2 nodeSelector (backup-storage PVC is local-path pinned there), lower memory request (128Mi) so it fits worker-2 pressure, and switch to alpine/k8s image (bitnami/kubectl is no longer resolvable). - Rewrite monitoring/backup-volumes to only back up grafana + loki co-located with backup-storage on k3s-worker-2. Prometheus data lives on k3s-worker-1 and is intentionally excluded here; a dedicated Prometheus data backup follows in a separate ticket. The three CronJobs previously left Pending/ContainerCreating pods that blocked the OS-update health guard in DEV-463. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|---|---|---|
| .. | ||
| backup-k8s-resources-cronjob.yaml | ||
| backup-volumes-cronjob.yaml | ||
| README.md | ||
monitoring — backup CronJobs
Manifests recording the cluster-side monitoring backup CronJobs that were previously applied out-of-band. These files are the authoritative source (kubectl apply -f apps/monitoring/). See DEV-464 for the repair context.
backup-k8s-resources-cronjob.yaml— daily dump of Kubernetes resources intobackup-storagePVC.backup-volumes-cronjob.yaml— daily rsync/tar of Grafana + Loki PVCs intobackup-storage. Prometheus data backup is not included here; it needs a separate on-node backup (tracked as a follow-up because Prometheus is on a different node thanbackup-storage).
The backup-storage PVC (100Gi, local-path, bound to k3s-worker-2) is the shared destination for both jobs.
Both CronJobs pin themselves to k3s-worker-2 via nodeSelector because that is the node that holds all destination + source PVCs used here.