Since DEV-510 landed HA control plane (2026-08-22), the cluster runs cp-1
plus cp-2/cp-3. The previous exclusion in os-update.sh matched `k3s-cp-1`
by hard-coded name only, which would have caused cp-2 and cp-3 to be
treated as regular fsn1 workers and drained/rebooted without the
CP-specific procedure.
Fix: select the CP list from `kubectl get nodes -l
node-role.kubernetes.io/control-plane` and skip any of those nodes. This
covers all present and future CPs automatically.
Also updated OS_UPDATE_PROCEDURE.md topology table and order rule to
document that all three CPs exist and are excluded from the weekly
cycle. The HA-aware CP OS-update procedure is a separate follow-up.
Verified on the current cluster:
[plan] EXCLUDING control-plane nodes: k3s-cp-1 k3s-cp-2 k3s-cp-3
[plan] ordered nodes (6): k3s-worker-4 k3s-update-runner k3s-worker-1
k3s-worker-2 k3s-worker-3 k3s-worker-5
Co-Authored-By: Paperclip <noreply@paperclip.ing>
The weekly rolling OS-update cycle has been silently stripping docker.io from
worker nodes (DEV-498), breaking the Forgejo runner whose hostPath mount for
/var/run/docker.sock requires the socket to exist. Fix in two layers:
- Defensive pin: update-node.sh now runs `apt-mark manual docker.io` in the
apt phase (whenever it is installed) so `apt-get autoremove --purge` cannot
silently drop it during subsequent upgrades.
- Post-reboot reconciliation: new `ensure-node-docker.sh` installs docker.io
if missing, enables + starts the systemd unit, waits for /var/run/docker.sock,
and re-applies the `basicstack.de/docker=true` label. Wired into
update-node.sh between kubelet-Ready and uncordon. No-op on nodes without
the label (safe for cp-1 and the update runner).
Verified idempotent against all 5 labeled workers; `apt-mark manual docker.io`
now set on every worker (survived across reboots by design).
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Follow-up to the cp-1 design commit: update the two remaining places in
OS_UPDATE_PROCEDURE.md that still said "cp-1 last" / "planned as a distinct
issue" so they now name CP1_UPDATE_PROCEDURE.md + update-cp-1.sh.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
cp-1 is excluded from os-update.sh because a one-shot drain would trigger
the kine cascade documented in DEV-495. This adds:
- CP1_UPDATE_PROCEDURE.md: swap-add (Phase A), preflight, batched stateful
eviction (Phase C), drain+apt (Phase D), reboot with external livez
monitor (Phase E), uncordon+verify (Phase F), and rollback paths.
- scripts/os-update/update-cp-1.sh: subcommand-per-phase runner with the
same /tmp/os-update-cp-1-<ts>.log contract as update-node.sh; supports
--dry-run, --add-swap, --preflight, --drain-stateful, --apt, --reboot,
--finalize, --run.
- os-update.sh: explicitly excludes k3s-cp-1 with a pointer to the cp-1
script; kine thundering-herd guardrails preserved.
- OS_UPDATE_PROCEDURE.md: cross-reference to the cp-1 procedure.
Execution requires separate board approval; this change is design +
dry-run artifact only.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Add four hard rules to the OS_UPDATE_PROCEDURE:
1. Pre-plan drain order — no drain that evicts >3 StatefulSets at once
2. cp-1 must have swap before the cycle finishes (currently 0 swap, 3.7 GiB RAM)
3. Halt cycle if `kubectl get nodes` from cp-1 exceeds 5 s (kine slowness leading indicator)
4. cp-1 OS update is a separate design task, not part of standard os-update.sh cycle
Root cause reference: DEV-495 (worker-3 drain 2026-08-16 caused kine SQLite cascade
+ taint-eviction storm + near-OOM on cp-1; cluster self-recovered without operator
action after ~80 min).
Co-Authored-By: Paperclip <noreply@paperclip.ing>
- infrastructure/OS_UPDATE_PROCEDURE.md: agent-facing rolling update
procedure (drain -> apt -> reboot -> verify -> uncordon -> health ->
next). Explicit MUST NOT list around k3s config, PVs, and manifests.
- infrastructure/OS_UPDATE_ROUTINE.md: describes the weekly Paperclip
routine (Sun 03:00 Europe/Berlin) that fires this procedure.
- infrastructure/scripts/os-update/: cluster-health.sh, update-node.sh,
os-update.sh, README. Enforces the same guardrails in code:
workers-first-then-CP, one node at a time, no --force drains, halts on
reboot/kubelet/health failure, never touches k3s config.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Adds infrastructure/networking/traefik-helmchartconfig.yaml so the k3s
Traefik overrides (badger plugin, allowCrossNamespace, letsencrypt
resolver + persistent acme.json, non-root fsGroup) are tracked in git.
kube-system is not managed by ArgoCD in this cluster; kubectl apply of
this file is the manual reproducibility path.
Refs: DEV-455, DEV-457
Co-Authored-By: Paperclip <noreply@paperclip.ing>
- Add Hetzner Load Balancer (138.199.128.63) to network architecture diagram
- Update DNS configuration to show all domains pointing to Hetzner LB IP
- Update HTTP/HTTPS traffic flow to show traffic routing through Hetzner LB
- Update last modified date to 2026-08-02
Related to DEV-439.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Added comprehensive documentation of the two-tier load balancing setup:
- Hetzner Cloud Load Balancer (external layer, managed by Hetzner CCM)
- Kubernetes LoadBalancer services (internal layer, k3s ServiceLB)
Key points documented:
- Traffic flow from external client through both LB layers to pod
- Why LoadBalancer service type is required (CCM integration)
- Historical context of the migration from hostPort to Hetzner LB
- Service definitions and port configurations
Updated:
- apps/stalwart/README.md: Added Network Architecture section
- infrastructure/networking/NETWORK_ARCHITECTURE.md: Enhanced Stalwart
section with two-tier architecture details and updated traffic flows
Resolves documentation gap identified in DEV-439.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Created ADD_WORKER_NODE.md as a step-by-step checklist for adding
new worker nodes to the k3s cluster. This complements the existing
K3S_OPERATIONS.md with a focused, actionable guide.
Key features:
- Pre-flight checklist format
- Hetzner Cloud VM creation steps (console and CLI)
- SSH key requirements (cto-paperclip, silentmaxx)
- Firewall configuration verification
- Provisioning script usage
- Health verification steps
- Documentation update procedure
- Rollback instructions
- Troubleshooting guide
Created for DEV-408 (adding worker node to resolve Stalwart scheduling).
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Created apps/pangolin/ directory with namespace.yaml defining the pangolin namespace.
Configured DNS A record for pangolin.basicstack.de → 178.105.17.239 (cluster ingress IP).
Updated DNS_REQUIREMENTS.md to document the new Pangolin service.
This completes Phase 1 of the Pangolin deployment (DEV-390):
- Repository structure created with namespace definition
- DNS record configured and verified in Hetzner zone
- Documentation updated
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Created comprehensive network documentation for BasicStack k3s cluster:
- NETWORK_ARCHITECTURE.md: Complete network architecture with diagrams,
node configuration, CNI (Flannel) details, ingress/LoadBalancer setup,
DNS configuration, TLS certificates, network policies, traffic flows,
and troubleshooting procedures
- DNS_REQUIREMENTS.md: Complete DNS record requirements for all services
including A records, MX records, SPF, DKIM, DMARC, and PTR records
- NETWORK_VERIFICATION.md: Verification report documenting current state
of all network components with findings and recommendations
Updated infrastructure README with links to new network documentation.
Key findings:
- All worker nodes correctly configured with --node-ip set to private IPs
- Flannel VXLAN properly configured with public IP annotations
- Traefik ingress controller operational
- 16/17 TLS certificates valid (registry-tls needs investigation)
- 3 LoadBalancer services properly configured
- Network policies securing database services
Addresses DEV-225: Verify and document k3s cluster network configuration
Related: DEV-224 (node-ip configuration), DEV-223 (DNS issues)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
- Add CLUSTER_ACCESS.md with comprehensive cluster access guide
- Fix Service CIDR in K3S_OPERATIONS.md (10.43.0.0/16, not 10.96.0.0/12)
- Document API server instability fix (cluster-cidr configuration)
- Add troubleshooting section for CIDR mismatch issues
- Update change history with cluster update details
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Added Stakater Reloader to automatically restart Stalwart pods when
TLS certificates are renewed by cert-manager. This ensures seamless
certificate rotation without manual intervention.
Changes:
- Deploy Stakater Reloader in infrastructure/networking/
- Add Reloader annotation to Stalwart StatefulSet to watch stalwart-tls secret
- Document certificate renewal process and troubleshooting
The certificate is managed by cert-manager with Let's Encrypt and will
automatically renew 30 days before expiration (renewal date: 2026-08-20).
Reloader detects secret updates and triggers a rolling restart of the
Stalwart StatefulSet to load the new certificate.
Co-Authored-By: Paperclip <noreply@paperclip.ing>