Root cause: k3s service ClusterIP routing instability causing intermittent
failures despite healthy pods. This is the 5th incident - prior fixes treated
symptoms, not the systemic networking fragility.
Changes:
- Add startup probe (60s delay, prevents premature service registration)
- Fix backup job env var substitution (use shell ${VAR}, not K8s $(VAR))
- Add comprehensive monitoring (ServiceMonitor, PrometheusRule, blackbox probes)
- Add alerting for service failures, high latency, pod restarts, backup failures
Evidence:
- Pod healthy (4d15h uptime, 0 restarts) but service ClusterIP routing broken
- Direct pod IP worked, service ClusterIP failed with "Connection reset by peer"
- Iptables rules correct, endpoints correct, but packets not flowing
- Required pod restart + Traefik restart to restore service
Monitoring now tests full service path from outside cluster, not just pod health.
Will alert immediately on failures instead of relying on reactive discovery.
Related: DEV-213, DEV-221, DEV-223, DEV-224, DEV-230, DEV-231
Co-Authored-By: Paperclip <noreply@paperclip.ing>
|
||
|---|---|---|
| apps | ||
| docs | ||
| infrastructure | ||
| .gitignore | ||
| add-user-andreas.ldif | ||
| README.md | ||
stack.basicstack.de
CD/CI deployment manifests and configurations for the basicstack.de Kubernetes cluster.
Repository Structure
stack.basicstack.de/
├── apps/ # Application deployments
│ ├── stalwart/ # Stalwart mail server (example)
│ └── forgejo/ # Forgejo Git service (placeholder)
├── infrastructure/ # Infrastructure-level configurations
│ ├── networking/ # Network policies, ingress, DNS
│ └── monitoring/ # Monitoring, logging, observability
└── docs/ # Documentation and guides
Purpose
This repository serves as the central source of truth for all deployment configurations targeting the basicstack.de Kubernetes cluster. It follows GitOps principles where infrastructure and application state is declaratively defined and version-controlled.
Directory Details
apps/
Contains deployment configurations for individual applications and services running on the cluster. Each application should have its own subdirectory with:
- Kubernetes manifests (Deployments, StatefulSets, Services, etc.)
- Helm values files
- Configuration files
- Application-specific documentation
Example: The stalwart/ directory contains the complete deployment configuration for the Stalwart mail server, including multiple deployment variants, monitoring setup, and operational guides.
infrastructure/
Contains cluster-wide infrastructure configurations:
- networking/: Ingress controllers, network policies, DNS configurations, load balancers
- monitoring/: Prometheus, Grafana, logging infrastructure, observability tools
docs/
General documentation including:
- Deployment procedures
- Cluster architecture
- Troubleshooting guides
- Best practices
Getting Started
- Clone this repository
- Review the example Stalwart deployment in
apps/stalwart/ - Follow the pattern for new application deployments
- Ensure all manifests are tested before committing
Contributing
All changes should be:
- Committed with clear, descriptive messages
- Tested in a development environment when possible
- Documented appropriately
- Reviewed before deployment to production
Cluster Information
- Cluster: basicstack.de
- Platform: K3s on Hetzner Cloud
- Namespace Strategy: One namespace per application (recommended)
- Ingress: Traefik (default K3s ingress controller)