Hetzner OS publishes 4 systemd-resolved upstreams and Kubernetes limits pod resolv.conf to 3 nameservers, so kubelet drops the 4th and fires a DNSConfigForming Warning event on every hostNetwork or dnsPolicy=Default pod restart. Silence the noise by pinning the pods to 3 explicit servers (same 3 kubelet was already picking). - apps/observability/patches/node-exporter-dns-config.yaml — strategic- merge patch adding dnsPolicy=None + dnsConfig to the kube-prometheus-stack node-exporter DaemonSet (Helm-managed, applied by hand) - apps/observability/patches/coredns-dns-config.yaml — companion patch for the k3s built-in CoreDNS Deployment. kubectl patch alone is not durable because the k3s addon controller reverts dnsPolicy; kept as a quick manual re-apply hook - infrastructure/k3s-manifests/coredns.yaml — the authoritative modified k3s addon manifest that must live at /var/lib/rancher/k3s/server/manifests/coredns.yaml on all 3 CP nodes - infrastructure/k3s-manifests/README-DEV-527.md — apply procedure, verification steps, and upgrade caveat Applied and verified on the live cluster: - node-exporter DaemonSet rolled with dnsPolicy=None; no DNSConfigForming events on current pods - coredns Deployment reconciled after pushing the modified manifest to all 3 CPs; new pod runs with dnsPolicy=None and 3-nameserver dnsConfig - internal + external DNS resolution still works Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|---|---|---|
| .. | ||
| k3s-manifests | ||
| k3s-upgrade | ||
| networking | ||
| scripts | ||
| ADD_WORKER_NODE.md | ||
| CLUSTER_ACCESS.md | ||
| CP1_UPDATE_PROCEDURE.md | ||
| CP_UPDATE_PROCEDURE.md | ||
| K3S_OPERATIONS.md | ||
| OS_UPDATE_PROCEDURE.md | ||
| OS_UPDATE_ROUTINE.md | ||
| README.md | ||
Infrastructure
This directory contains cluster-wide infrastructure configurations that support all applications.
Operational procedures (agent-facing)
- OS_UPDATE_PROCEDURE.md — weekly rolling Ubuntu OS update for all cluster nodes (drain → apt → reboot → verify → uncordon → cluster health → next). Never touches k3s config.
- OS_UPDATE_ROUTINE.md — the Paperclip routine that fires the above weekly.
- K3S_OPERATIONS.md — k3s version upgrades (separate concern from OS updates).
- ADD_WORKER_NODE.md — adding a worker node.
- CLUSTER_ACCESS.md — SSH / kubectl access.
Structure
networking/
Network-level configurations including:
- Network Architecture - Comprehensive network architecture documentation
- DNS Requirements - DNS records and configuration guide
- Ingress controller configurations (Traefik)
- Network policies
- Load balancer configurations (k3s ServiceLB)
- Certificate management (cert-manager, TLS)
- Certificate reloading (Stakater Reloader)
monitoring/
Observability infrastructure:
- Prometheus operator and configurations
- Grafana dashboards and datasources
- Logging stack (Loki, Promtail, etc.)
- Alert rules and notification channels
- Service monitors and pod monitors
Purpose
Infrastructure configurations in this directory are shared across all applications. Changes here can affect the entire cluster, so:
- Test thoroughly before applying
- Document all changes
- Consider the impact on existing deployments
- Coordinate with other team members
Adding Infrastructure Components
When adding new infrastructure components:
- Create appropriate subdirectories if needed
- Include clear documentation
- Define dependencies and prerequisites
- Provide rollback procedures
- Update this README with the new component