# HA k3s Control-Plane OS Update Procedure **Purpose:** apply Ubuntu OS updates (kernel, security, apt) to any control-plane node in the HA k3s cluster (`k3s-cp-1`, `k3s-cp-2`, `k3s-cp-3`) without losing etcd quorum and without a longer api-server outage than a normal reboot. **Audience:** the CTO agent, or an operator with root SSH to the cluster. Execution is board-approval-gated — see `## Approval gates` below. **Supersedes:** the earlier single-CP `CP1_UPDATE_PROCEDURE.md` (kept as a redirect stub). The cp-1-only procedure was written when cp-1 was the sole api-server and the datastore was embedded SQLite via kine. After [DEV-510](/DEV/issues/DEV-510) (2026-08-22), the cluster runs HA embedded etcd across cp-1/cp-2/cp-3 and all three CPs are `NoSchedule`-tainted, so: - **The kine cascade is no longer the driving risk.** The datastore is etcd 3.6.12 with 3-node quorum; a single CP drain no longer leaves a lone SQLite writer. The kubectl-latency guardrail stays as a sanity check but is now a much softer signal. - **CPs host no StatefulSet or single-replica Deployment workloads.** The pre-drain stateful eviction dance from the old cp-1 doc no longer applies — CPs only host DaemonSets + `metrics-server` (currently on cp-2) + occasional CronJob completions. - **The api-server is not lost on reboot.** During a CP reboot the other two apiservers keep serving. `/livez` is polled from the *other* apiservers, not from an external machine forced to poll the one being rebooted. --- ## Scope **In scope** - Add a durable ≥ 2 GiB swapfile on any CP that has zero swap (idempotent one-off — `--add-swap`). - k3s etcd snapshot as the restore point. - `kubectl cordon` + `kubectl drain` on the target CP. - `apt-get update/upgrade/dist-upgrade/autoremove` on the target CP. - Controlled reboot with etcd-quorum-aware liveness monitoring. - Uncordon + cluster health verify. **Out of scope — do NOT do here** - Any change to `/etc/rancher/k3s/*`, `/etc/systemd/system/k3s*.service*`, or the k3s binary version. k3s upgrades go through system-upgrade-controller (see `K3S_OPERATIONS.md`). - Any change to manifests under `apps/`, `infrastructure/`, or applied via ArgoCD. - Deleting PVs / PVCs. Any pod that gets rescheduled off the target CP stays on its new node. - Fixing application-level problems. - Rebuilding the node — if a CP does not return after reboot, escalate; the rebuild path is `ADD_WORKER_NODE.md` plus board approval, not this document. --- ## Approval gates This procedure has **two** independent gates. Neither happens without explicit board approval on the corresponding Paperclip issue: 1. **Add swap (Phase A).** Non-invasive, non-state-mutating, kubelet already runs with `failSwapOn=false`. Requires board approval per stateful-service safety rules because it modifies a CP node. 2. **Full OS update (Phases B–E).** Requires board approval because it drains + reboots one of the etcd members. Do NOT execute without an explicit `request_board_approval` acceptance on the execution ticket. Both gates are independent — swap can (and should) be added first, in a quiet window, before the full update is scheduled. --- ## CP ordering rule (multiple CPs in one cycle) - **One at a time. Never two CPs cordoned or draining at once.** Two of three CPs down = etcd quorum loss = api-server unavailability for the whole cluster. - **Leader last.** Query the current etcd leader before starting; update the two followers first (in any order), then the leader. Rationale: draining a follower is a no-op for the raft leader; draining the leader forces a re-election. Doing followers first minimises leader flapping. - **Health gate between CPs.** After each CP finishes Phase F (finalize) and cluster health is green, wait at least `POST_UNCORDON_WAIT_SECONDS` (default 180 s) and re-check etcd endpoint status before touching the next CP. This gives etcd time to fully re-sync the just-rebooted member. - **Halt on any yellow.** Any of the following aborts the cycle at the current CP (do NOT proceed to the next CP): etcd reports a member as not `started`, `kubectl get nodes` > 5 s, any node not `Ready`, any workload deployment/sts under desired replicas. Query the leader: ```bash ssh root@$CP1_HOST bash -c ' ETCDCTL_API=3 etcdctl \ --endpoints=https://127.0.0.1:2379 \ --cacert=/var/lib/rancher/k3s/server/tls/etcd/server-ca.crt \ --cert=/var/lib/rancher/k3s/server/tls/etcd/server-client.crt \ --key=/var/lib/rancher/k3s/server/tls/etcd/server-client.key \ endpoint status --cluster -w table' ``` --- ## Current CP workload (snapshot 2026-08-23) Re-derive live before any execution — this table is a design reference only. | Node | Kernel | Swap | StatefulSets | Single-replica Deployments | DaemonSets | |------|--------|------|--------------|-----------------------------|------------| | k3s-cp-1 | 6.8.0-137 | **0 B** | 0 | 0 | node-exporter, promtail, hcloud-csi-node, svclb-* | | k3s-cp-2 | 6.8.0-137 | **0 B** | 0 | metrics-server (evictable) | node-exporter, promtail, hcloud-csi-node, svclb-* | | k3s-cp-3 | 6.8.0-137 | **0 B** | 0 | 0 | node-exporter, promtail, hcloud-csi-node, svclb-* | etcd datastore live snapshot: 38 MiB per member, all three `started`, cp-3 is current leader. Because CPs are `NoSchedule`-tainted, no workload will end up on them via normal scheduling — the pre-drain stateful redistribution loop from the old cp-1 procedure is obsolete. --- ## Phase A — Add swap (one-off, idempotent) **Goal:** eliminate the "3.7 GiB, no swap" underlying constraint on any CP before ever draining it. All three CPs currently have zero swap; add swap once per CP. **Preconditions:** - kubelet in this k3s already runs with `failSwapOn=false` (confirmed on cp-1 via `/api/v1/nodes/k3s-cp-1/proxy/configz`) — enabling swap does NOT break the kubelet. - Target CP's `/` has ≥ `2 * SWAP_SIZE_MB` free. - Board approval on the swap-add ticket. **Sizing:** default **4 GiB** swap. Rationale — cp-1 baseline (k3s + hosted daemons) sits at ~2 GiB used; 4 GiB swap gives us headroom for the drain-eviction transient without inflating disk usage past ~10 % of `/`. Minimum acceptable per guardrail: 2 GiB. **Location:** `/swapfile` (root filesystem). Not a separate partition — reversible, no LVM changes. **Steps (encoded in `update-cp-node.sh --add-swap`):** identical to the old cp-1 procedure — `fallocate` → `mkswap` → `swapon` → `/etc/fstab` → `vm.swappiness=10` in `/etc/sysctl.d/99-k3s-swap.conf`. Idempotent: skips if a swapfile of the target size is already active. Full script source in `scripts/os-update/update-cp-node.sh`. **Phase A also installs `etcd-client`** on the target CP if `etcdctl` is missing. This is a hard requirement for the etcd-quorum probes used in Phase B/E when the CP is later itself the update target and one of the *other* CPs must query etcd cluster status. cp-1 already has `etcdctl` (installed pre-HA); cp-2 and cp-3 pick it up here. **Rollback for Phase A:** `swapoff /swapfile && rm /swapfile` and remove the fstab line. Safe at any time — swap is a soft resource. **Verification after Phase A:** - `free -h` shows `Swap: 4.0Gi` used ≈ 0. - `swapon --show` shows `/swapfile 4G`. - `sysctl vm.swappiness` returns `10`. - kubelet still Ready (`kubectl get node `). - No new `MemoryPressure` condition. --- ## Phase B — Preflight Runs from the operator machine (or from any healthy CP). Every command's stdout+stderr goes to `/tmp/os-update--.log` on the operator machine. Attach that log to the Paperclip execution ticket at the end. Run `update-cp-node.sh --preflight`. Checks: 1. **Cluster is currently healthy:** `cluster-health.sh` returns 0. 2. **Target node is a CP:** `kubectl get node ` carries the `node-role.kubernetes.io/control-plane` label. 3. **All 3 CPs are healthy etcd members:** `etcdctl endpoint status --cluster` shows every member `started`, no errors. 4. **Only one CP is being updated this cycle:** no other CP is currently cordoned. 5. **Target CP has swap on:** abort if `free -h` shows `Swap: 0B`. Run `--add-swap` first. 6. **kubectl latency probe:** `time kubectl get nodes` returns in ≤ 5 s. 7. **etcd snapshot** (restore point): `k3s etcd-snapshot save --name pre-cp-os-update--` executed on the target CP. 8. **Records the current etcd leader** — if the target *is* the leader, prints a warning ("prefer updating a follower first"). Does NOT auto-swap the target; the operator/agent makes the call per the CP ordering rule. If any preflight check fails: STOP. Do not proceed. --- ## Phase C — Cordon + drain the target CP CPs are `NoSchedule`-tainted, so cordon is mostly a belt-and-suspenders measure. Drain evicts the small tail of not-DaemonSet workloads (`metrics-server` currently lives on cp-2). ```bash kubectl cordon kubectl drain \ --ignore-daemonsets \ --delete-emptydir-data \ --timeout="${DRAIN_TIMEOUT_SECONDS:-600}s" ``` If drain reports a PDB block: do NOT `--force`. Uncordon the node, mark the run `blocked` on the PDB, and escalate. This is a workload PDB bug — fix separately. **Kine-cascade guardrail (softened for HA etcd):** during drain, `time kubectl get nodes` should still return in ≤ 5 s. Because we now have 3-member etcd (not lone SQLite kine), a single-CP drain does not create the write-amplification hazard from [DEV-495](/DEV/issues/DEV-495). But `> 5 s` still indicates something is wrong (etcd slow disk, leader flapping) — halt and investigate. --- ## Phase D — apt on the target CP Identical to the worker-node apt block (`update-node.sh` step 3). `update-cp-node.sh --apt` runs this via `ssh root@ bash -s`: ```bash export DEBIAN_FRONTEND=noninteractive APT_OPTS='-y -o Dpkg::Options::=--force-confdef -o Dpkg::Options::=--force-confold' uname -r > /root/pre-apt-kernel dpkg-query -W -f='${Package}\t${Version}\n' > /root/pre-apt-packages.tsv if dpkg --audit | grep -qE .; then dpkg --configure -a || true; fi apt-get update apt-get $APT_OPTS upgrade || { apt-get $APT_OPTS -f install; apt-get $APT_OPTS upgrade; } apt-get $APT_OPTS dist-upgrade apt-get $APT_OPTS autoremove --purge apt-get clean [ -f /var/run/reboot-required ] && echo REBOOT_REQUIRED=yes || echo REBOOT_REQUIRED=no ``` `pre-apt-kernel` and `pre-apt-packages.tsv` are the local rollback references — see Rollback below. --- ## Phase E — Reboot handling (etcd-quorum-aware) **Duration expectation:** 90–180 s of *this member's* api-server unavailability. The other two apiservers keep serving; the operator machine's kubectl continues to work via one of them. **Escalation trigger:** rebooted apiserver not back on `/livez` after **10 minutes** → escalate. First check the Hetzner console via `hcloud server describe ` for boot state; if kernel-panic / initramfs, use grub previous-kernel path (see Rollback). Do NOT rebuild the node. **Steps (executed by `update-cp-node.sh --reboot`):** ```bash # 1. Reboot the target CP (ssh will hang up mid-command — expected). ssh $SSH_OPTS root@$TARGET_HOST 'systemctl reboot' || true sleep 15 # 2. Poll the TARGET's /livez from the operator machine. It is served on port 6443. # Insecure (`-k`) because the server cert is self-signed by k3s. deadline=$(( $(date +%s) + ${REBOOT_MAX_WAIT_SECONDS:-600} )) while [ $(date +%s) -lt $deadline ]; do code=$(curl -sk -o /dev/null -w '%{http_code}' https://$TARGET_HOST:6443/livez 2>/dev/null || echo 000) [ "$code" = "200" ] && break sleep 5 done # 3. Verify etcd membership from a peer CP — the target should be back as `started`. ssh $SSH_OPTS root@$PEER_HOST ' ETCDCTL_API=3 etcdctl \ --endpoints=https://127.0.0.1:2379 \ --cacert=/var/lib/rancher/k3s/server/tls/etcd/server-ca.crt \ --cert=/var/lib/rancher/k3s/server/tls/etcd/server-client.crt \ --key=/var/lib/rancher/k3s/server/tls/etcd/server-client.key \ endpoint status --cluster -w table' # 4. Wait for kubelet Ready on the target from the api-server view. deadline=$(( $(date +%s) + 300 )) while [ $(date +%s) -lt $deadline ]; do ready=$(kubectl get node "$TARGET_NODE" -o jsonpath='{.status.conditions[?(@.type=="Ready")].status}' 2>/dev/null || echo Unknown) [ "$ready" = "True" ] && break sleep 5 done ``` **What to do if the api-server takes 3–10 minutes:** it may be normal for a large etcd on a slow disk. Wait it out. Do NOT restart k3s to "help" — an etcd member catching up on raft log is stateful and interrupting it can slow the whole cluster. **What to do if it takes > 10 minutes:** escalate; run the Hetzner-console diagnosis; if the console shows a bootable Ubuntu but the k3s service is failing, that's the boundary — this procedure stops here. Follow `K3S_OPERATIONS.md` for the k3s recovery path. --- ## Phase F — Uncordon + verify + finalize ```bash kubectl uncordon sleep "${POST_UNCORDON_WAIT_SECONDS:-180}" RETRY_ON_TRANSIENT=1 infrastructure/scripts/os-update/cluster-health.sh ssh root@ 'zgrep -h "Commandline\|Install\|Upgrade\|Remove" /var/log/apt/history.log* 2>/dev/null | tail -60' ``` Between CPs (when multiple CPs will be updated in the same cycle), also verify: ```bash ssh root@ ' ETCDCTL_API=3 etcdctl \ --endpoints=https://127.0.0.1:2379 \ --cacert=/var/lib/rancher/k3s/server/tls/etcd/server-ca.crt \ --cert=/var/lib/rancher/k3s/server/tls/etcd/server-client.crt \ --key=/var/lib/rancher/k3s/server/tls/etcd/server-client.key \ endpoint status --cluster -w table' ``` All three members must be `started`, no errors, DB sizes within 10 % of each other. Attach the `/tmp/os-update--.log` to the execution ticket. Mark the ticket `done` on green health, or `blocked` naming the specific residual issue. --- ## Rollback / recovery ### If apt broke a package - On the target CP, `dpkg --audit` to find half-configured packages. - `dpkg --configure -a`, then `apt-get -f install`. - If a specific package broke and you know the previous version from `/root/pre-apt-packages.tsv`, `apt-get install =`. ### If the new kernel does not boot - Hetzner cloud console: `hcloud server request-console ` → get VNC URL, watch boot. - If grub is up, select the previous-kernel entry. Ubuntu keeps ≥ 1 old kernel installed by default. - Once booted on the old kernel, `apt-get remove` the broken kernel and pin the working one: ```bash apt-mark hold linux-image- linux-headers- ``` - Escalate on the ticket regardless — a kernel rollback is a follow-up investigation, not a "done" outcome. ### If the node does not return at all - Do NOT `hcloud server delete`. Do NOT re-provision. - Escalate to the board with the Hetzner console output and `/tmp/os-update--*.log`. - The `--cluster-reset --cluster-reset-restore-path=` recovery path exists (see `K3S_OPERATIONS.md`) but requires board approval per stateful-service safety rules. The pre-flight snapshot from Phase B is the restore point. - **HA-etcd advantage:** if one CP is permanently lost but the other two are healthy, you have quorum. Do not rush a `cluster-reset`. Follow the "member remove + re-add" path in the k3s docs first. ### If etcd quorum degrades mid-cycle - Two of three CPs down = write outage until quorum returns. Do NOT attempt to update the third. - Uncordon all CPs, restore the just-updated one first, and confirm all three members `started` before doing anything else. --- ## Automation entry points - `infrastructure/scripts/os-update/update-cp-node.sh [phase-flag]` — the generalized HA-aware flow. - `--add-swap` — Phase A only (idempotent, safe standalone). - `--dry-run` — walk all phases printing what would be done, no state change. - `--preflight` — Phase B only. - `--drain` — Phase C only (cordon + drain). - `--apt` — Phase D apt commands only. - `--reboot` — Phase E only. - `--finalize` — Phase F only. - `--run` — all phases with a confirmation between each unless `ASSUME_YES=1`. - `infrastructure/scripts/os-update/update-cp-1.sh` — retained as a thin wrapper that calls `update-cp-node.sh k3s-cp-1 "$@"`. Historical callers keep working. The script follows the same log-dir contract as `update-node.sh` (`/tmp/os-update--.log`). --- ## Change history | Date | Change | By | |------|--------|-----| | 2026-08-16 | Initial cp-1-only procedure (DEV-496) | CTO agent | | 2026-08-23 | Generalized to HA (cp-1/cp-2/cp-3); dropped Phase C batched stateful eviction (CPs are NoSchedule-tainted now); softened kine cascade guardrail (embedded etcd 3-member quorum); added CP ordering rule (leader last, one at a time); introduced `update-cp-node.sh` (DEV-515) | CTO agent |