After the DEV-464 split, `monitoring/backup-volumes` only backs up grafana + loki (pinned to k3s-worker-2 with `backup-storage`), leaving Prometheus data unbacked. `prometheus-data-encrypted` is an RWO Hetzner Cloud volume attached to whichever node currently runs the Prometheus pod (typically k3s-worker-1), so it cannot join the shared backup-volumes job without provoking Multi-Attach errors. This introduces a dedicated `monitoring/prometheus-backup` CronJob that: - Streams `prometheus-data-encrypted` to Hetzner S3 via rclone (`basicstack-backup/prometheus/prometheus-<DATE>/`). - Uses `podAffinity` to co-schedule with the Prometheus pod so the RWO PVC always attaches on the same node. - Runs at 03:30 daily, `Forbid` concurrency, 60m hard deadline. - Retains 7 days of dated backups (rclone delete --min-age 7d). - Tolerates the expected TSDB compaction race (Prometheus deletes old block dirs mid-copy): rclone's non-zero exit from those transient errors is captured, then success is validated by comparing dest bytes to source bytes (>= 80% and > 100 MiB floor). S3 credentials are the same Hetzner Object Storage account used by `opencloud-backup` and `stalwart-backup`, resealed for the `monitoring` namespace as `SealedSecret monitoring-s3-backup`. Verified with a manual job on k3s-worker-1 (see DEV-465 for logs). Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|---|---|---|
| apps | ||
| docs | ||
| infrastructure | ||
| .gitignore | ||
| add-user-andreas.ldif | ||
| DEV-349-progress.md | ||
| README.md | ||
stack.basicstack.de
CD/CI deployment manifests and configurations for the basicstack.de Kubernetes cluster.
Repository Structure
stack.basicstack.de/
├── apps/ # Application deployments
│ ├── stalwart/ # Stalwart mail server (example)
│ └── forgejo/ # Forgejo Git service (placeholder)
├── infrastructure/ # Infrastructure-level configurations
│ ├── networking/ # Network policies, ingress, DNS
│ └── monitoring/ # Monitoring, logging, observability
└── docs/ # Documentation and guides
Purpose
This repository serves as the central source of truth for all deployment configurations targeting the basicstack.de Kubernetes cluster. It follows GitOps principles where infrastructure and application state is declaratively defined and version-controlled.
Directory Details
apps/
Contains deployment configurations for individual applications and services running on the cluster. Each application should have its own subdirectory with:
- Kubernetes manifests (Deployments, StatefulSets, Services, etc.)
- Helm values files
- Configuration files
- Application-specific documentation
Example: The stalwart/ directory contains the complete deployment configuration for the Stalwart mail server, including multiple deployment variants, monitoring setup, and operational guides.
infrastructure/
Contains cluster-wide infrastructure configurations:
- networking/: Ingress controllers, network policies, DNS configurations, load balancers
- monitoring/: Prometheus, Grafana, logging infrastructure, observability tools
docs/
General documentation including:
- Deployment procedures
- Cluster architecture
- Troubleshooting guides
- Best practices
GitOps with Argo CD
This repository is managed via Argo CD, the GitOps deployment platform for the cluster.
- Argo CD UI: https://argo.basicstack.de
- Authentication: Pocket ID SSO (https://auth.basicstack.de)
- Documentation: docs/argocd.md
All changes pushed to the main branch are automatically synchronized to the cluster. Applications are defined in apps/app-*.yaml files and reference subdirectories for their manifests.
For details on managing applications, repository credentials, troubleshooting, and emergency procedures, see the Argo CD documentation.
Getting Started
- Clone this repository
- Review the example Stalwart deployment in
apps/stalwart/ - Follow the pattern for new application deployments
- Ensure all manifests are tested before committing
- Argo CD will automatically sync changes to the cluster (or use manual sync for critical changes)
Contributing
All changes should be:
- Committed with clear, descriptive messages
- Tested in a development environment when possible
- Documented appropriately
- Reviewed before deployment to production
Cluster Information
- Cluster: basicstack.de
- Platform: K3s on Hetzner Cloud
- Namespace Strategy: One namespace per application (recommended)
- Ingress: Traefik (default K3s ingress controller)