stack.basicstack.de/infrastructure/networking/DNS_CONFIGURATION.md
Paperclip CTO 27d26fb11b Document DNS configuration and service CIDR fix
- Fixed Service CIDR documentation (10.96.0.0/16, not 10.43.0.0/16)
- Added comprehensive DNS configuration guide
- Documented kubelet cluster-dns requirements
- Added CoreDNS forward configuration details
- Documented DNS CIDR mismatch troubleshooting (DEV-223)

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-07-06 18:22:53 +00:00

6.5 KiB

DNS Configuration

Overview

The cluster DNS configuration requires alignment between:

  1. The kube-dns service ClusterIP (10.96.0.10)
  2. The kubelet cluster-dns setting (must match the service IP)
  3. The CoreDNS upstream forwarders

Misalignment causes cluster-wide DNS failures, manifesting as HTTP 502 errors for ingresses and service-to-service communication failures.

Current Configuration

  • kube-dns Service IP: 10.96.0.10
  • Service CIDR: 10.96.0.0/16
  • CoreDNS Upstream: 8.8.8.8, 1.1.1.1

Kubelet DNS Configuration

All k3s nodes (control plane and workers) must be configured to use the correct DNS service IP.

Control Plane Configuration

File: /etc/rancher/k3s/config.yaml

# k3s server configuration
kube-apiserver-arg:
  - "encryption-provider-config=/etc/rancher/k3s/encryption-config.yaml"
# DNS configuration
kubelet-arg:
  - "cluster-dns=10.96.0.10"

Worker Node Configuration

File: /etc/rancher/k3s/config.yaml

kubelet-arg:
  - "cluster-dns=10.96.0.10"

Applying Configuration Changes

After updating the configuration file:

# Control plane
systemctl restart k3s

# Worker nodes
systemctl restart k3s-agent

CoreDNS Configuration

CoreDNS is managed by k3s through the auto-deploy manifest system.

Forward Configuration

File: /var/lib/rancher/k3s/server/manifests/coredns.yaml

The Corefile must use explicit upstream DNS servers instead of /etc/resolv.conf:

apiVersion: v1
kind: ConfigMap
metadata:
  name: coredns
  namespace: kube-system
data:
  Corefile: |
    .:53 {
        errors
        health
        ready
        kubernetes cluster.local in-addr.arpa ip6.arpa {
          pods insecure
          fallthrough in-addr.arpa ip6.arpa
        }
        hosts /etc/coredns/NodeHosts {
          ttl 60
          reload 15s
          fallthrough
        }
        prometheus :9153
        cache 30
        loop
        reload
        loadbalance
        import /etc/coredns/custom/*.override
        forward . 8.8.8.8 1.1.1.1
    }
    import /etc/coredns/custom/*.server

Critical: The forward . 8.8.8.8 1.1.1.1 line must NOT use /etc/resolv.conf because the host's resolv.conf points to 127.0.0.53 (systemd-resolved on localhost), which does not work from inside containers.

Applying CoreDNS Changes

K3s automatically reconciles the coredns ConfigMap from the manifest file. After editing:

# CoreDNS will reload automatically (reload plugin)
# Or restart CoreDNS for immediate effect:
kubectl rollout restart deployment -n kube-system coredns

Verification

Check Pod DNS Configuration

kubectl run test-dns --image=busybox:latest --rm -i --restart=Never -- cat /etc/resolv.conf

Expected output:

nameserver 10.96.0.10
search default.svc.cluster.local svc.cluster.local cluster.local
options ndots:5

Test DNS Resolution

# Test cluster DNS
kubectl run test-dns-resolve --image=busybox:latest --rm -i --restart=Never -- \
  nslookup kubernetes.default.svc.cluster.local

# Test external DNS
kubectl run test-external-dns --image=busybox:latest --rm -i --restart=Never -- \
  nslookup google.com

Both should resolve successfully.

Check CoreDNS Logs

kubectl logs -n kube-system -l k8s-app=kube-dns --tail=50

Should not show errors like:

  • Failed to watch: apiserver not ready
  • plugin/kubernetes: Failed to list

Troubleshooting

Symptom: Pods Cannot Resolve DNS

Check 1: Verify pod DNS configuration

kubectl run debug --image=busybox:latest --rm -i --restart=Never -- cat /etc/resolv.conf

If nameserver is wrong (e.g., 10.43.0.10 instead of 10.96.0.10):

  1. Check kubelet configuration on all nodes
  2. Restart k3s/k3s-agent services
  3. Delete and recreate test pods (existing pods keep old DNS config)

Check 2: Verify kube-dns service exists

kubectl get svc -n kube-system kube-dns

Should show ClusterIP 10.96.0.10

Check 3: Test CoreDNS directly

kubectl run test-dns-direct --image=busybox:latest --rm -i --restart=Never -- \
  nslookup kubernetes.default.svc.cluster.local 10.96.0.10

If this works but normal DNS doesn't, the issue is kubelet configuration.

Symptom: CoreDNS Returns NXDOMAIN

Check 1: Verify CoreDNS can reach Kubernetes API

kubectl logs -n kube-system -l k8s-app=kube-dns | grep -i error

Look for API connection errors.

Check 2: Verify CoreDNS configuration

kubectl get configmap -n kube-system coredns -o yaml | grep -A 5 "forward"

Should show forward . 8.8.8.8 1.1.1.1, NOT /etc/resolv.conf

Check 3: Restart CoreDNS

kubectl rollout restart deployment -n kube-system coredns
kubectl wait --for=condition=available deployment/coredns -n kube-system --timeout=60s

Symptom: HTTP 502 Bad Gateway for Ingresses

This is often caused by DNS failures. Traefik cannot resolve backend service names.

Fix:

  1. Verify DNS is working (see above checks)
  2. Restart Traefik to pick up DNS fixes:
    kubectl rollout restart deployment -n kube-system traefik
    

Adding New Nodes

When provisioning new worker nodes, ensure DNS configuration is included:

ssh root@<new-node-ip>

# Create k3s config directory
mkdir -p /etc/rancher/k3s

# Configure DNS
cat > /etc/rancher/k3s/config.yaml <<EOF
kubelet-arg:
  - "cluster-dns=10.96.0.10"
EOF

# Then proceed with k3s agent installation

Root Cause History

DEV-223: DNS Service CIDR Mismatch (2026-07-06)

Problem: Cluster-wide DNS failure causing HTTP 502 errors for all ingresses and service communication failures.

Root Cause:

  • The cluster was configured with service CIDR 10.96.0.0/16 (standard Kubernetes range)
  • k3s defaults to 10.43.0.0/16, and kubelet was never reconfigured
  • Pods received DNS server 10.43.0.10, but kube-dns service was at 10.96.0.10
  • All DNS queries timed out

Fix:

  1. Added kubelet-arg: ["cluster-dns=10.96.0.10"] to /etc/rancher/k3s/config.yaml on control plane and all workers
  2. Updated CoreDNS forward configuration to use 8.8.8.8 1.1.1.1 instead of /etc/resolv.conf
  3. Restarted k3s services cluster-wide

Verification:

  • HTTPS services returned HTTP 302/200 instead of 502
  • IMAP/SMTP ports accessible
  • DNS resolution working from all pods

Prevention: Always explicitly configure kubelet cluster-dns when using non-default service CIDR.


Change History

Date Change By
2026-07-06 Initial DNS configuration documentation, fixed DNS CIDR mismatch CTO Agent