The stalwart-backup CronJob has failed nightly since 2026-08-13, all with FailureTarget=DeadlineExceeded. Root cause: the backup pod had no scheduling constraint and got placed on a node different from stalwart-0. The hcloud CSI block volume is RWO and can only be attached to one node, so the backup pod stayed in ContainerCreating with FailedAttachVolume / Multi-Attach until the 600s active deadline killed it. - Add podAffinity requiredDuringScheduling on app=stalwart,statefulset.kubernetes.io/pod-name=stalwart-0 with topology key kubernetes.io/hostname so the backup pod always lands on the same node. Same-node co-location lets both pods share the already-attached block volume; the in-container script then scales stalwart-0 down, backs up, and scales it back up as before. - Raise activeDeadlineSeconds from 600s to 1800s as safety headroom (successful runs are ~86s; the extra budget covers prune growth). Verified: manual run of the patched CronJob completed in 86s and wrote restic snapshot f76f534e (2026-08-15 12:07:30) to s3://basicstack-backup/stalwart. stalwart-0 is back to Ready. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|---|---|---|
| apps | ||
| docs | ||
| infrastructure | ||
| .gitignore | ||
| add-user-andreas.ldif | ||
| DEV-349-progress.md | ||
| README.md | ||
stack.basicstack.de
CD/CI deployment manifests and configurations for the basicstack.de Kubernetes cluster.
Repository Structure
stack.basicstack.de/
├── apps/ # Application deployments
│ ├── stalwart/ # Stalwart mail server (example)
│ └── forgejo/ # Forgejo Git service (placeholder)
├── infrastructure/ # Infrastructure-level configurations
│ ├── networking/ # Network policies, ingress, DNS
│ └── monitoring/ # Monitoring, logging, observability
└── docs/ # Documentation and guides
Purpose
This repository serves as the central source of truth for all deployment configurations targeting the basicstack.de Kubernetes cluster. It follows GitOps principles where infrastructure and application state is declaratively defined and version-controlled.
Directory Details
apps/
Contains deployment configurations for individual applications and services running on the cluster. Each application should have its own subdirectory with:
- Kubernetes manifests (Deployments, StatefulSets, Services, etc.)
- Helm values files
- Configuration files
- Application-specific documentation
Example: The stalwart/ directory contains the complete deployment configuration for the Stalwart mail server, including multiple deployment variants, monitoring setup, and operational guides.
infrastructure/
Contains cluster-wide infrastructure configurations:
- networking/: Ingress controllers, network policies, DNS configurations, load balancers
- monitoring/: Prometheus, Grafana, logging infrastructure, observability tools
docs/
General documentation including:
- Deployment procedures
- Cluster architecture
- Troubleshooting guides
- Best practices
GitOps with Argo CD
This repository is managed via Argo CD, the GitOps deployment platform for the cluster.
- Argo CD UI: https://argo.basicstack.de
- Authentication: Pocket ID SSO (https://auth.basicstack.de)
- Documentation: docs/argocd.md
All changes pushed to the main branch are automatically synchronized to the cluster. Applications are defined in apps/app-*.yaml files and reference subdirectories for their manifests.
For details on managing applications, repository credentials, troubleshooting, and emergency procedures, see the Argo CD documentation.
Getting Started
- Clone this repository
- Review the example Stalwart deployment in
apps/stalwart/ - Follow the pattern for new application deployments
- Ensure all manifests are tested before committing
- Argo CD will automatically sync changes to the cluster (or use manual sync for critical changes)
Contributing
All changes should be:
- Committed with clear, descriptive messages
- Tested in a development environment when possible
- Documented appropriately
- Reviewed before deployment to production
Cluster Information
- Cluster: basicstack.de
- Platform: K3s on Hetzner Cloud
- Namespace Strategy: One namespace per application (recommended)
- Ingress: Traefik (default K3s ingress controller)