CD/CI deployment manifests and configurations for basicstack.de cluster
Find a file
CTO Agent ef62dde67c feat(observability): wire node-exporter textfile collector for backup metrics (DEV-494)
Turns the four monitoring backup CronJobs' .prom output into scrapeable
Prometheus series so the DEV-490 alerts finally evaluate against live data.

- apps/observability/patches/node-exporter-textfile-collector.yaml:
  strategic-merge patch on the kube-prometheus-stack node-exporter DS
  that adds `--collector.textfile.directory=/host/textfile_collector`
  and mounts `/var/lib/node_exporter/textfile_collector` read-only.
  Chart isn't tracked in ArgoCD, so we keep the patch under version
  control and re-apply after any helm upgrade (see patches/README.md).
- apps/monitoring/backup-{loki,grafana,k8s-resources,prometheus}-*-cronjob.yaml:
  swap `emptyDir` /metrics for a `hostPath` on the same directory
  (`DirectoryOrCreate`). Write via `.prom.tmp` + `mv` so node-exporter
  never reads a truncated sample.
- apps/monitoring/backup-restic-alerts.yaml: `time() - max(...) > 28h`
  for all four freshness alerts so a stale `.prom` left on a node the
  job has since left does not fire the freshness pager.
- apps/monitoring/README.md: drop the "once wired" caveat; document
  the on-node directory, atomic write, cross-node staleness rationale.

Verified end-to-end with a synthetic `kubectl create job
--from=cronjob/backup-k8s-resources`: pod ran on k3s-worker-4,
.prom file materialized in /var/lib/node_exporter/textfile_collector,
and Prometheus returned all four metric families
(`backup_k8s_resources_success=1`, `..._timestamp_seconds`,
`..._check_status=0`, `restic_repo_size_bytes{repo="k8s-resources"}=13692516`).
`time() - max(backup_k8s_resources_timestamp_seconds)` returned ~84s
against a fresh run. `promtool check rules` + `promtool test rules`
still pass (9 rules, 5 scenarios).

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-08-16 16:21:56 +00:00
apps feat(observability): wire node-exporter textfile collector for backup metrics (DEV-494) 2026-08-16 16:21:56 +00:00
docs feat(monitoring): pull restic image from Harbor mirror (DEV-493) 2026-08-16 16:20:06 +00:00
infrastructure Add weekly rolling OS-update procedure for k3s nodes (DEV-462) 2026-08-09 15:55:39 +00:00
.gitignore Convert all secrets to SealedSecrets for enhanced security 2026-07-01 18:38:27 +00:00
add-user-andreas.ldif OpenCloud: Use native bash substitution in config file 2026-07-05 14:03:49 +00:00
DEV-349-progress.md Consolidate Forgejo backup into forgejo namespace 2026-07-26 09:35:31 +00:00
README.md Add comprehensive Argo CD documentation 2026-07-12 10:04:51 +00:00

stack.basicstack.de

CD/CI deployment manifests and configurations for the basicstack.de Kubernetes cluster.

Repository Structure

stack.basicstack.de/
├── apps/                    # Application deployments
│   ├── stalwart/           # Stalwart mail server (example)
│   └── forgejo/            # Forgejo Git service (placeholder)
├── infrastructure/          # Infrastructure-level configurations
│   ├── networking/         # Network policies, ingress, DNS
│   └── monitoring/         # Monitoring, logging, observability
└── docs/                   # Documentation and guides

Purpose

This repository serves as the central source of truth for all deployment configurations targeting the basicstack.de Kubernetes cluster. It follows GitOps principles where infrastructure and application state is declaratively defined and version-controlled.

Directory Details

apps/

Contains deployment configurations for individual applications and services running on the cluster. Each application should have its own subdirectory with:

  • Kubernetes manifests (Deployments, StatefulSets, Services, etc.)
  • Helm values files
  • Configuration files
  • Application-specific documentation

Example: The stalwart/ directory contains the complete deployment configuration for the Stalwart mail server, including multiple deployment variants, monitoring setup, and operational guides.

infrastructure/

Contains cluster-wide infrastructure configurations:

  • networking/: Ingress controllers, network policies, DNS configurations, load balancers
  • monitoring/: Prometheus, Grafana, logging infrastructure, observability tools

docs/

General documentation including:

  • Deployment procedures
  • Cluster architecture
  • Troubleshooting guides
  • Best practices

GitOps with Argo CD

This repository is managed via Argo CD, the GitOps deployment platform for the cluster.

All changes pushed to the main branch are automatically synchronized to the cluster. Applications are defined in apps/app-*.yaml files and reference subdirectories for their manifests.

For details on managing applications, repository credentials, troubleshooting, and emergency procedures, see the Argo CD documentation.

Getting Started

  1. Clone this repository
  2. Review the example Stalwart deployment in apps/stalwart/
  3. Follow the pattern for new application deployments
  4. Ensure all manifests are tested before committing
  5. Argo CD will automatically sync changes to the cluster (or use manual sync for critical changes)

Contributing

All changes should be:

  1. Committed with clear, descriptive messages
  2. Tested in a development environment when possible
  3. Documented appropriately
  4. Reviewed before deployment to production

Cluster Information

  • Cluster: basicstack.de
  • Platform: K3s on Hetzner Cloud
  • Namespace Strategy: One namespace per application (recommended)
  • Ingress: Traefik (default K3s ingress controller)