# Weekly Rolling Ubuntu OS Update Procedure **Purpose:** Keep every k3s node's Ubuntu OS patched (kernel, security, package updates) with a **rolling, zero-downtime** update — one node at a time, drain → update → reboot → verify → uncordon → cluster health check → next node. **Audience:** Agents (currently CTO, optionally a dedicated ClusterOps agent). Also usable manually by an operator. **Scope — what this procedure does:** - Runs `apt-get update`, `apt-get -y upgrade`, `apt-get -y dist-upgrade`, `apt-get -y autoremove` on each cluster node. - Reboots the node if a reboot is required (kernel/libc updates). - Drains and cordons each node before touching it, uncordons after verification. - Verifies the node and cluster are healthy before moving on. **Scope — what this procedure MUST NOT do:** - **Never change the Kubernetes / k3s setup.** Do not touch `/etc/systemd/system/k3s*.service*`, `/etc/rancher/k3s/*`, k3s binary version, k3s config, kube-system manifests, network policies, or any manifest under `infrastructure/`, `apps/`, or applied via ArgoCD. - **Do not upgrade k3s** here. k3s upgrades are handled separately by system-upgrade-controller (see `K3S_OPERATIONS.md` and `k3s-upgrade/`). - **Do not delete PersistentVolumes, PVCs, or workload manifests.** - Do not "fix" application-level problems on a node — that's out of scope. Only fix problems in the OS/apt/reboot layer of the current node being updated. Anything else → stop, report, escalate. --- ## Cluster topology (context) | Role | Node | Private IP | Public IP | Datacenter | Notes | |------|------|-----------|-----------|------------|-------| | control-plane | k3s-cp-1 | 10.42.1.1 | 178.105.17.239 | fsn1 | update LAST | | worker | k3s-worker-1 | 10.42.1.2 | (via CP) | fsn1 | | | worker | k3s-worker-2 | 10.42.1.3 | (via CP) | fsn1 | | | worker | k3s-worker-3 | 10.42.1.5 | 167.233.121.121 | fsn1 | | | worker | k3s-worker-4 | 10.42.1.6 | 128.140.3.80 | nbg1 | | | worker | k3s-worker-5 | 10.42.1.7 | 167.233.192.86 | fsn1 | | | runner | k3s-update-runner | 167.233.79.65 | 167.233.79.65 | fsn1 | k3s-upgrade helper, still an updatable node | Always re-derive the live list before running — nodes may have been added/removed: ```bash ssh root@178.105.17.239 'kubectl get nodes -o wide' ``` **Order rule:** update ALL workers first, control plane LAST. Within workers, update in this order to protect stateful workloads: 1. runner + workers that host no PVs (safest — lowest disruption) 2. remaining workers 3. **Stalwart-hosting fsn1 workers last among workers** — Stalwart has hard fsn1 affinity, so draining a fsn1 worker while another fsn1 worker is also unavailable can leave Stalwart Pending. Never have two fsn1 workers cordoned/down at the same time. 4. **k3s-cp-1 last** — single control plane; the API server goes away during its reboot. **Concurrency:** exactly one node at a time. Never in parallel. --- ## Access Prereqs are the same as `CLUSTER_ACCESS.md`: - SSH key for `root` on every node (jump via control plane for private-IP workers). - `kubectl` available (either from the operator machine, or by SSH-ing to the control plane and using it there). - `hcloud` CLI configured (only needed for firewall-related recovery — not for normal runs). Environment variables the scripts expect: - `CONTROL_PLANE_HOST` — default `178.105.17.239` - `CONTROL_PLANE_PRIVATE` — default `10.42.1.1` - Drain timeout: `DRAIN_TIMEOUT_SECONDS` — default `600` - Reboot wait: `REBOOT_MAX_WAIT_SECONDS` — default `600` - Post-uncordon settle: `POST_UNCORDON_WAIT_SECONDS` — default `180` --- ## Preflight (run once per cycle, before touching any node) 1. **Cluster is currently healthy.** If any of the checks below fail, STOP and open an issue — do not start OS updates on an already-degraded cluster. ```bash ssh root@$CONTROL_PLANE_HOST bash -s <<'EOF' set -e kubectl get nodes echo "--- Not-ready nodes:" kubectl get nodes -o json | jq -r '.items[] | select(.status.conditions[] | select(.type=="Ready" and .status!="True")) | .metadata.name' echo "--- Pods not Running/Completed:" kubectl get pods -A --field-selector=status.phase!=Running,status.phase!=Succeeded | grep -v 'STATUS' || echo " (none)" echo "--- Non-Ready pods (Running but not Ready):" kubectl get pods -A -o json | jq -r '.items[] | select(.status.phase=="Running") | select([.status.conditions[]?|select(.type=="Ready")|.status]|contains(["False"])) | "\(.metadata.namespace)/\(.metadata.name)"' EOF ``` Only proceed if: all nodes `Ready`, no non-Running/Succeeded pods, no Running-but-not-Ready pods (small transient counts are OK — retry once and continue if it clears). 2. **Snapshot k3s datastore** (control plane only — this is a checkpoint you can restore etcd from if the control-plane reboot goes badly): ```bash ssh root@$CONTROL_PLANE_HOST 'k3s etcd-snapshot save --name pre-os-update-$(date +%Y%m%d)' ssh root@$CONTROL_PLANE_HOST 'k3s etcd-snapshot list | tail -5' ``` 3. **List the nodes to update** and derive the ordered plan. Persist to `/tmp/os-update-plan.txt` on the control plane for auditability. See `scripts/os-update/os-update.sh` for the reference implementation of the ordering rule. --- ## Per-node procedure Repeat for each node in the ordered plan. The reference implementation is `infrastructure/scripts/os-update/update-node.sh `; the manual steps below are what that script does. Throughout: every command's stdout+stderr goes to a per-run log at `/tmp/os-update--.log` on the control plane. Attach the log to the issue at the end. ### 1. Pre-check the node ```bash NODE= kubectl get node "$NODE" kubectl describe node "$NODE" | grep -A2 Conditions ``` Confirm: `Ready=True`, not already cordoned, no `DiskPressure`/`MemoryPressure`/`PIDPressure`. ### 2. Cordon and drain ```bash kubectl cordon "$NODE" kubectl drain "$NODE" \ --ignore-daemonsets \ --delete-emptydir-data \ --disable-eviction=false \ --timeout="${DRAIN_TIMEOUT_SECONDS:-600}s" ``` If drain fails on a PodDisruptionBudget: - Do NOT force-delete pods (breaks HA guarantees). - Log the blocking PDB, uncordon the node, mark the node as `SKIPPED_PDB` in the plan, and continue with the next node. Escalate the PDB conflict on the issue at the end. If drain fails on a lone pod without a controller: - Do NOT `--force`. Same as above — uncordon, mark `SKIPPED_ORPHAN_POD`, continue. ### 3. Update the OS On the node itself: ```bash ssh -o StrictHostKeyChecking=accept-new root@ bash -s <<'REMOTE' set -euo pipefail export DEBIAN_FRONTEND=noninteractive # Refresh package lists apt-get update # Configure apt to keep existing config files silently (no interactive prompts) APT_OPTS='-y -o Dpkg::Options::=--force-confdef -o Dpkg::Options::=--force-confold' apt-get $APT_OPTS upgrade apt-get $APT_OPTS dist-upgrade apt-get $APT_OPTS autoremove --purge apt-get clean # Report reboot need if [ -f /var/run/reboot-required ]; then echo "REBOOT_REQUIRED=yes" cat /var/run/reboot-required.pkgs 2>/dev/null || true else echo "REBOOT_REQUIRED=no" fi REMOTE ``` For private-IP workers, target them from the control plane (`ssh -J root@$CONTROL_PLANE_HOST root@10.42.1.X`) or run the whole block after SSH-ing to the control plane first. Common OS-only fixes the agent MAY perform on the node if apt fails: - `dpkg --configure -a` after an interrupted install - `apt-get -f install` to resolve broken deps - Free disk with `journalctl --vacuum-time=3d` or `apt-get clean` if `/` is full - Restart a system service that is stuck (`systemctl restart `) — but NOT `k3s`, `k3s-agent`, `containerd`, `flanneld`, or any container runtime **Never**: change k3s config, delete PVs, uninstall packages the OS didn't schedule, edit `/etc/rancher/`, or reinstall k3s. If a fix would touch any of those, stop and escalate. ### 4. Reboot if required If `REBOOT_REQUIRED=yes`: ```bash ssh root@ 'systemctl reboot' || true # Wait until SSH responds again (max REBOOT_MAX_WAIT_SECONDS) deadline=$(( $(date +%s) + ${REBOOT_MAX_WAIT_SECONDS:-600} )) while [ $(date +%s) -lt $deadline ]; do sleep 10 if ssh -o ConnectTimeout=5 -o StrictHostKeyChecking=accept-new root@ 'uptime' 2>/dev/null; then echo "Node back up"; break fi done ``` If the node doesn't come back within the timeout: escalate immediately. Do NOT rebuild the node or touch k3s — a rebuild requires the ADD_WORKER_NODE procedure and is a separate approved action. ### 5. Wait for k3s-agent / k3s to be ready again ```bash # Wait for kubelet Ready condition (control plane view) deadline=$(( $(date +%s) + 300 )) while [ $(date +%s) -lt $deadline ]; do READY=$(kubectl get node "$NODE" -o jsonpath='{.status.conditions[?(@.type=="Ready")].status}') [ "$READY" = "True" ] && break sleep 5 done kubectl get node "$NODE" ``` If Ready never returns True: escalate. Do NOT change k3s config. ### 6. Uncordon ```bash kubectl uncordon "$NODE" ``` ### 7. Post-node health verification Wait for pods to reschedule and settle, then verify: ```bash sleep "${POST_UNCORDON_WAIT_SECONDS:-180}" # All nodes Ready? kubectl get nodes kubectl get nodes -o json | jq -r '.items[] | select(.status.conditions[] | select(.type=="Ready" and .status!="True")) | .metadata.name' | grep . && { echo "NOT-READY NODES"; exit 1; } || true # Any pod not Running/Succeeded? BAD=$(kubectl get pods -A --field-selector=status.phase!=Running,status.phase!=Succeeded --no-headers 2>/dev/null | wc -l) [ "$BAD" -gt 0 ] && { echo "BAD PODS: $BAD"; kubectl get pods -A --field-selector=status.phase!=Running,status.phase!=Succeeded; exit 1; } || true # Any Deployment/StatefulSet under desired replicas? kubectl get deploy -A -o json | jq -r '.items[] | select((.status.readyReplicas // 0) < (.spec.replicas // 1)) | "deploy \(.metadata.namespace)/\(.metadata.name) \(.status.readyReplicas // 0)/\(.spec.replicas)"' kubectl get sts -A -o json | jq -r '.items[] | select((.status.readyReplicas // 0) < (.spec.replicas // 1)) | "sts \(.metadata.namespace)/\(.metadata.name) \(.status.readyReplicas // 0)/\(.spec.replicas)"' ``` Only proceed to the next node when ALL of the above are green. If not: - Give it another 3 minutes and re-check (workloads with large images may still be pulling). - If still not green: STOP the cycle. Uncordon everything, leave the cluster in a stable state, and open a follow-up issue with the failing workloads. Do NOT proceed to update more nodes. ### 8. Control plane special handling `k3s-cp-1` is the single control plane node. During its reboot: - The kube-apiserver is unavailable — kubectl commands from the operator machine will error out. - The `kubectl` waits in step 5/7 must run from a machine that is NOT the control plane, or must be scheduled after the control plane's SSH is back and `curl -k https://localhost:6443/healthz` returns `ok`. - Skip the drain for DaemonSet pods on the control plane (`--ignore-daemonsets` covers that), but hosted apps that scheduled onto CP (rare — verify with `kubectl get pods -A --field-selector spec.nodeName=k3s-cp-1`) will be evicted. --- ## Finalization After all nodes are done: 1. Final cluster health snapshot (same preflight commands as at the start). 2. Print apt history summary per node so the audit trail has "what changed": ```bash for host in ; do echo "=== $host ===" ssh root@$host 'zgrep -h "Commandline\|Install\|Upgrade\|Remove" /var/log/apt/history.log* 2>/dev/null | tail -60' done ``` 3. Optional: prune old etcd snapshots to keep disk in check: ```bash ssh root@$CONTROL_PLANE_HOST 'k3s etcd-snapshot list' # delete anything older than 14 days if desired (manual review) ``` 4. Attach the per-run logs to the Paperclip issue that triggered this run. 5. Update the issue with: - Nodes updated (list) - Nodes skipped (list + reason) - Any package that required manual intervention - Reboots performed - Final `kubectl get nodes` output If everything is clean → mark the issue `done`. If a node was skipped or errored → mark `blocked` with the unblock action, or open a child issue for the specific failure. --- ## Recovery / what to do when it goes wrong **Node stuck cordoned after failure:** `kubectl uncordon ` — always leave the cluster back in its normal scheduling state. **Node fails to boot after reboot:** - Check Hetzner console for boot errors (kernel panic, initramfs). - Try `hcloud server reboot ` from `hcloud` CLI. - If the node cannot recover: escalate. Do NOT delete or rebuild the Hetzner server without board approval; that is the ADD_WORKER_NODE flow. **k3s-agent won't start after reboot:** - Check `journalctl -u k3s-agent -n 100`. - Do NOT edit the k3s-agent unit file. Do NOT re-run the k3s installer. - Escalate. This is out of scope for the OS-update procedure. **Pods CrashLoopBackOff after node came back:** - Not an OS-update problem to fix — the node is healthy, apt succeeded. Leave the node uncordoned, stop the cycle, and open an application-level issue. **PDB blocked drain:** - Never `--force` the drain. Leave the node uncordoned, skip it, note in the report which PDB blocked and which workload owns it. **Datastore snapshot restore (last resort — CP only):** - Only if the control plane is broken beyond repair. See `K3S_OPERATIONS.md` for the `--cluster-reset --cluster-reset-restore-path=` procedure. Requires board approval — do NOT execute unattended. --- ## Automation entry points - `infrastructure/scripts/os-update/os-update.sh` — full cycle runner (preflight → per-node loop → finalization). Idempotent, resumable via `--start-from `. Use `--dry-run` to print the plan without touching anything. - `infrastructure/scripts/os-update/update-node.sh ` — single-node update (all 7 per-node steps). Callable standalone for retry. - `infrastructure/scripts/os-update/cluster-health.sh` — the preflight/post-node health check as a standalone command; exits non-zero on any failure. Read the script sources for the exact behavior before running them. They mirror this procedure step for step. --- ## Weekly schedule A Paperclip routine (see `infrastructure/OS_UPDATE_ROUTINE.md`) fires this procedure once per week. The routine creates a task whose description points here. The assigned agent reads this document and executes the automation. --- ## Change history | Date | Change | By | |------|--------|-----| | 2026-08-09 | Initial procedure + automation scripts | CTO agent (DEV-462) |