stack.basicstack.de/infrastructure
CTO Agent 34350a03bf os-update: exclude all control-plane nodes by role, not just cp-1 (DEV-513)
Since DEV-510 landed HA control plane (2026-08-22), the cluster runs cp-1
plus cp-2/cp-3. The previous exclusion in os-update.sh matched `k3s-cp-1`
by hard-coded name only, which would have caused cp-2 and cp-3 to be
treated as regular fsn1 workers and drained/rebooted without the
CP-specific procedure.

Fix: select the CP list from `kubectl get nodes -l
node-role.kubernetes.io/control-plane` and skip any of those nodes. This
covers all present and future CPs automatically.

Also updated OS_UPDATE_PROCEDURE.md topology table and order rule to
document that all three CPs exist and are excluded from the weekly
cycle. The HA-aware CP OS-update procedure is a separate follow-up.

Verified on the current cluster:
  [plan] EXCLUDING control-plane nodes: k3s-cp-1 k3s-cp-2 k3s-cp-3
  [plan] ordered nodes (6): k3s-worker-4 k3s-update-runner k3s-worker-1
                            k3s-worker-2 k3s-worker-3 k3s-worker-5

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-08-23 01:39:25 +00:00
..
k3s-upgrade Add k3s cluster management automation and documentation 2026-07-06 17:08:21 +00:00
networking Traefik: check in HelmChartConfig with DEV-457 changes for reproducibility 2026-08-08 14:52:29 +00:00
scripts os-update: exclude all control-plane nodes by role, not just cp-1 (DEV-513) 2026-08-23 01:39:25 +00:00
ADD_WORKER_NODE.md docs: add standalone worker node addition instruction 2026-07-26 16:07:52 +00:00
CLUSTER_ACCESS.md Update k3s cluster documentation 2026-07-06 17:54:53 +00:00
CP1_UPDATE_PROCEDURE.md docs(infra): design cp-1 OS update procedure (DEV-496) 2026-08-16 19:28:52 +00:00
K3S_OPERATIONS.md Document DNS configuration and service CIDR fix 2026-07-06 18:22:53 +00:00
OS_UPDATE_PROCEDURE.md os-update: exclude all control-plane nodes by role, not just cp-1 (DEV-513) 2026-08-23 01:39:25 +00:00
OS_UPDATE_ROUTINE.md Add weekly rolling OS-update procedure for k3s nodes (DEV-462) 2026-08-09 15:55:39 +00:00
README.md Add weekly rolling OS-update procedure for k3s nodes (DEV-462) 2026-08-09 15:55:39 +00:00

Infrastructure

This directory contains cluster-wide infrastructure configurations that support all applications.

Operational procedures (agent-facing)

Structure

networking/

Network-level configurations including:

  • Network Architecture - Comprehensive network architecture documentation
  • DNS Requirements - DNS records and configuration guide
  • Ingress controller configurations (Traefik)
  • Network policies
  • Load balancer configurations (k3s ServiceLB)
  • Certificate management (cert-manager, TLS)
  • Certificate reloading (Stakater Reloader)

monitoring/

Observability infrastructure:

  • Prometheus operator and configurations
  • Grafana dashboards and datasources
  • Logging stack (Loki, Promtail, etc.)
  • Alert rules and notification channels
  • Service monitors and pod monitors

Purpose

Infrastructure configurations in this directory are shared across all applications. Changes here can affect the entire cluster, so:

  1. Test thoroughly before applying
  2. Document all changes
  3. Consider the impact on existing deployments
  4. Coordinate with other team members

Adding Infrastructure Components

When adding new infrastructure components:

  1. Create appropriate subdirectories if needed
  2. Include clear documentation
  3. Define dependencies and prerequisites
  4. Provide rollback procedures
  5. Update this README with the new component