stack.basicstack.de/infrastructure/OS_UPDATE_PROCEDURE.md
CTO Agent 34350a03bf os-update: exclude all control-plane nodes by role, not just cp-1 (DEV-513)
Since DEV-510 landed HA control plane (2026-08-22), the cluster runs cp-1
plus cp-2/cp-3. The previous exclusion in os-update.sh matched `k3s-cp-1`
by hard-coded name only, which would have caused cp-2 and cp-3 to be
treated as regular fsn1 workers and drained/rebooted without the
CP-specific procedure.

Fix: select the CP list from `kubectl get nodes -l
node-role.kubernetes.io/control-plane` and skip any of those nodes. This
covers all present and future CPs automatically.

Also updated OS_UPDATE_PROCEDURE.md topology table and order rule to
document that all three CPs exist and are excluded from the weekly
cycle. The HA-aware CP OS-update procedure is a separate follow-up.

Verified on the current cluster:
  [plan] EXCLUDING control-plane nodes: k3s-cp-1 k3s-cp-2 k3s-cp-3
  [plan] ordered nodes (6): k3s-worker-4 k3s-update-runner k3s-worker-1
                            k3s-worker-2 k3s-worker-3 k3s-worker-5

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-08-23 01:39:25 +00:00

19 KiB

Weekly Rolling Ubuntu OS Update Procedure

Purpose: Keep every k3s node's Ubuntu OS patched (kernel, security, package updates) with a rolling, zero-downtime update — one node at a time, drain → update → reboot → verify → uncordon → cluster health check → next node.

Audience: Agents (currently CTO, optionally a dedicated ClusterOps agent). Also usable manually by an operator.

Scope — what this procedure does:

  • Runs apt-get update, apt-get -y upgrade, apt-get -y dist-upgrade, apt-get -y autoremove on each cluster node.
  • Reboots the node if a reboot is required (kernel/libc updates).
  • Drains and cordons each node before touching it, uncordons after verification.
  • Verifies the node and cluster are healthy before moving on.

Scope — what this procedure MUST NOT do:

  • Never change the Kubernetes / k3s setup. Do not touch /etc/systemd/system/k3s*.service*, /etc/rancher/k3s/*, k3s binary version, k3s config, kube-system manifests, network policies, or any manifest under infrastructure/, apps/, or applied via ArgoCD.
  • Do not upgrade k3s here. k3s upgrades are handled separately by system-upgrade-controller (see K3S_OPERATIONS.md and k3s-upgrade/).
  • Do not delete PersistentVolumes, PVCs, or workload manifests.
  • Do not "fix" application-level problems on a node — that's out of scope. Only fix problems in the OS/apt/reboot layer of the current node being updated. Anything else → stop, report, escalate.

Cluster topology (context)

Role Node Private IP Public IP Datacenter Notes
control-plane k3s-cp-1 10.42.1.1 178.105.17.239 fsn1 EXCLUDED from os-update.sh — see CP1_UPDATE_PROCEDURE.md
control-plane k3s-cp-2 188.245.85.199 fsn1 EXCLUDED — joined 2026-08-22 (DEV-510); needs HA-aware CP OS-update procedure
control-plane k3s-cp-3 49.13.92.162 fsn1 EXCLUDED — joined 2026-08-22 (DEV-510); needs HA-aware CP OS-update procedure
worker k3s-worker-1 10.42.1.2 (via CP) fsn1
worker k3s-worker-2 10.42.1.3 (via CP) fsn1
worker k3s-worker-3 10.42.1.5 167.233.121.121 fsn1
worker k3s-worker-4 10.42.1.6 128.140.3.80 nbg1
worker k3s-worker-5 10.42.1.7 167.233.192.86 fsn1
runner k3s-update-runner 167.233.79.65 167.233.79.65 fsn1 k3s-upgrade helper, still an updatable node

Always re-derive the live list before running — nodes may have been added/removed:

ssh root@178.105.17.239 'kubectl get nodes -o wide'

Order rule: update workers only. Within workers, update in this order to protect stateful workloads:

  1. runner + workers that host no PVs (safest — lowest disruption)
  2. remaining workers
  3. Stalwart-hosting fsn1 workers last among workers — Stalwart has hard fsn1 affinity, so draining a fsn1 worker while another fsn1 worker is also unavailable can leave Stalwart Pending. Never have two fsn1 workers cordoned/down at the same time.
  4. ALL control-plane nodes excludedos-update.sh selects nodes without the node-role.kubernetes.io/control-plane label, so k3s-cp-1, k3s-cp-2, and k3s-cp-3 are all skipped automatically. CPs need their own procedure (CP1_UPDATE_PROCEDURE.md covers cp-1 today; an HA-aware successor is a follow-up).

Concurrency: exactly one node at a time. Never in parallel.


Kine thundering-herd guardrails (added after 2026-08-16 incident, see DEV-495)

The k3s control plane uses embedded SQLite (kine) as its datastore. On a single-CP cluster with limited RAM and no swap (cp-1 currently: 3.7 GiB RAM, 0 swap), draining a worker with many StatefulSets triggers a self-amplifying eviction-and-slow-SQL cascade: kine backs up on writes → apiserver hangs → node-lease renewals fail → taint-eviction kicks in on more nodes → more writes → kine falls further behind → OOM risk on cp-1.

All four rules below MUST be observed on every DEV-478 fire. The reference implementation (os-update.sh) does not yet enforce (1)/(3) automatically; the operator must actively watch.

  1. Pre-plan drain order for StatefulSets. Before draining any worker, kubectl get pods -n <ns> -o wide against every namespace with StatefulSets and count how many will be evicted from the target node. If a single drain would evict more than 3 StatefulSets at once, redistribute first: cordon+delete individual StatefulSet pods one namespace at a time and wait for each reschedule to settle before draining the whole node.
  2. cp-1 must have swap before finishing the cycle. cp-1 has 0 swap. Add at least 2 GiB of swap on cp-1 before the cp-1 update step (and ideally before draining the last stateful-heavy worker). This is a one-off setup task; once done it is a durable capability.
  3. Halt on kine slowness. During any drain, keep a time kubectl get nodes running from cp-1. If it exceeds 5 s in real time, halt the cycle immediately (uncordon the current node, do not proceed), verify cluster health, and escalate. The 5 s threshold is the leading indicator that kine has fallen behind and the taint-eviction cascade is about to start.
  4. cp-1 update is a separate design task. Draining cp-1's own pods (Stalwart, coredns, harbor-database if there, etc.) plus rebooting the single apiserver is the highest-risk step of the whole cycle. It MUST be planned and approved as a distinct issue before it runs; it is NOT covered by the standard os-update.sh cycle. See CP1_UPDATE_PROCEDURE.md and scripts/os-update/update-cp-1.sh (DEV-496).

Access

Prereqs are the same as CLUSTER_ACCESS.md:

  • SSH key for root on every node (jump via control plane for private-IP workers).
  • kubectl available (either from the operator machine, or by SSH-ing to the control plane and using it there).
  • hcloud CLI configured (only needed for firewall-related recovery — not for normal runs).

Environment variables the scripts expect:

  • CONTROL_PLANE_HOST — default 178.105.17.239
  • CONTROL_PLANE_PRIVATE — default 10.42.1.1
  • Drain timeout: DRAIN_TIMEOUT_SECONDS — default 600
  • Reboot wait: REBOOT_MAX_WAIT_SECONDS — default 600
  • Post-uncordon settle: POST_UNCORDON_WAIT_SECONDS — default 180

Preflight (run once per cycle, before touching any node)

  1. Cluster is currently healthy. If any of the checks below fail, STOP and open an issue — do not start OS updates on an already-degraded cluster.

    ssh root@$CONTROL_PLANE_HOST bash -s <<'EOF'
    set -e
    kubectl get nodes
    echo "--- Not-ready nodes:"
    kubectl get nodes -o json | jq -r '.items[] | select(.status.conditions[] | select(.type=="Ready" and .status!="True")) | .metadata.name'
    echo "--- Pods not Running/Completed:"
    kubectl get pods -A --field-selector=status.phase!=Running,status.phase!=Succeeded | grep -v 'STATUS' || echo "  (none)"
    echo "--- Non-Ready pods (Running but not Ready):"
    kubectl get pods -A -o json | jq -r '.items[] | select(.status.phase=="Running") | select([.status.conditions[]?|select(.type=="Ready")|.status]|contains(["False"])) | "\(.metadata.namespace)/\(.metadata.name)"'
    EOF
    

    Only proceed if: all nodes Ready, no non-Running/Succeeded pods, no Running-but-not-Ready pods (small transient counts are OK — retry once and continue if it clears).

  2. Snapshot k3s datastore (control plane only — this is a checkpoint you can restore etcd from if the control-plane reboot goes badly):

    ssh root@$CONTROL_PLANE_HOST 'k3s etcd-snapshot save --name pre-os-update-$(date +%Y%m%d)'
    ssh root@$CONTROL_PLANE_HOST 'k3s etcd-snapshot list | tail -5'
    
  3. List the nodes to update and derive the ordered plan. Persist to /tmp/os-update-plan.txt on the control plane for auditability. See scripts/os-update/os-update.sh for the reference implementation of the ordering rule.


Per-node procedure

Repeat for each node in the ordered plan. The reference implementation is infrastructure/scripts/os-update/update-node.sh <node-name>; the manual steps below are what that script does.

Throughout: every command's stdout+stderr goes to a per-run log at /tmp/os-update-<node>-<timestamp>.log on the control plane. Attach the log to the issue at the end.

1. Pre-check the node

NODE=<node-name>
kubectl get node "$NODE"
kubectl describe node "$NODE" | grep -A2 Conditions

Confirm: Ready=True, not already cordoned, no DiskPressure/MemoryPressure/PIDPressure.

2. Cordon and drain

kubectl cordon "$NODE"
kubectl drain "$NODE" \
  --ignore-daemonsets \
  --delete-emptydir-data \
  --disable-eviction=false \
  --timeout="${DRAIN_TIMEOUT_SECONDS:-600}s"

If drain fails on a PodDisruptionBudget:

  • Do NOT force-delete pods (breaks HA guarantees).
  • Log the blocking PDB, uncordon the node, mark the node as SKIPPED_PDB in the plan, and continue with the next node. Escalate the PDB conflict on the issue at the end.

If drain fails on a lone pod without a controller:

  • Do NOT --force. Same as above — uncordon, mark SKIPPED_ORPHAN_POD, continue.

3. Update the OS

On the node itself:

ssh -o StrictHostKeyChecking=accept-new root@<node-ssh-target> bash -s <<'REMOTE'
set -euo pipefail
export DEBIAN_FRONTEND=noninteractive
# Refresh package lists
apt-get update
# Configure apt to keep existing config files silently (no interactive prompts)
APT_OPTS='-y -o Dpkg::Options::=--force-confdef -o Dpkg::Options::=--force-confold'
apt-get $APT_OPTS upgrade
apt-get $APT_OPTS dist-upgrade
apt-get $APT_OPTS autoremove --purge
apt-get clean
# Report reboot need
if [ -f /var/run/reboot-required ]; then
  echo "REBOOT_REQUIRED=yes"
  cat /var/run/reboot-required.pkgs 2>/dev/null || true
else
  echo "REBOOT_REQUIRED=no"
fi
REMOTE

For private-IP workers, target them from the control plane (ssh -J root@$CONTROL_PLANE_HOST root@10.42.1.X) or run the whole block after SSH-ing to the control plane first.

Common OS-only fixes the agent MAY perform on the node if apt fails:

  • dpkg --configure -a after an interrupted install
  • apt-get -f install to resolve broken deps
  • Free disk with journalctl --vacuum-time=3d or apt-get clean if / is full
  • Restart a system service that is stuck (systemctl restart <unit>) — but NOT k3s, k3s-agent, containerd, flanneld, or any container runtime

Never: change k3s config, delete PVs, uninstall packages the OS didn't schedule, edit /etc/rancher/, or reinstall k3s. If a fix would touch any of those, stop and escalate.

4. Reboot if required

If REBOOT_REQUIRED=yes:

ssh root@<node-ssh-target> 'systemctl reboot' || true
# Wait until SSH responds again (max REBOOT_MAX_WAIT_SECONDS)
deadline=$(( $(date +%s) + ${REBOOT_MAX_WAIT_SECONDS:-600} ))
while [ $(date +%s) -lt $deadline ]; do
  sleep 10
  if ssh -o ConnectTimeout=5 -o StrictHostKeyChecking=accept-new root@<node-ssh-target> 'uptime' 2>/dev/null; then
    echo "Node back up"; break
  fi
done

If the node doesn't come back within the timeout: escalate immediately. Do NOT rebuild the node or touch k3s — a rebuild requires the ADD_WORKER_NODE procedure and is a separate approved action.

5. Wait for k3s-agent / k3s to be ready again

# Wait for kubelet Ready condition (control plane view)
deadline=$(( $(date +%s) + 300 ))
while [ $(date +%s) -lt $deadline ]; do
  READY=$(kubectl get node "$NODE" -o jsonpath='{.status.conditions[?(@.type=="Ready")].status}')
  [ "$READY" = "True" ] && break
  sleep 5
done
kubectl get node "$NODE"

If Ready never returns True: escalate. Do NOT change k3s config.

5b. Reconcile Docker on labeled nodes (before uncordon)

For nodes carrying the basicstack.de/docker=true label (currently the workers that run the Forgejo runner), ensure docker.io is present + running before scheduling resumes. Without this the runner pod comes back ContainerCreating because hostPath requires the docker socket to exist (see DEV-498 and DEV-499).

scripts/os-update/ensure-node-docker.sh "$NODE"

This step is a no-op on nodes without the label. update-node.sh runs it automatically between step 5 and step 6; the script is also safe to run ad-hoc after any manual OS operation. As a defensive extra measure, the apt phase (step 3) now runs apt-mark manual docker.io on any node where it is installed, so apt-get autoremove --purge cannot silently strip it.

6. Uncordon

kubectl uncordon "$NODE"

7. Post-node health verification

Wait for pods to reschedule and settle, then verify:

sleep "${POST_UNCORDON_WAIT_SECONDS:-180}"

# All nodes Ready?
kubectl get nodes
kubectl get nodes -o json | jq -r '.items[] | select(.status.conditions[] | select(.type=="Ready" and .status!="True")) | .metadata.name' | grep . && { echo "NOT-READY NODES"; exit 1; } || true

# Any pod not Running/Succeeded?
BAD=$(kubectl get pods -A --field-selector=status.phase!=Running,status.phase!=Succeeded --no-headers 2>/dev/null | wc -l)
[ "$BAD" -gt 0 ] && { echo "BAD PODS: $BAD"; kubectl get pods -A --field-selector=status.phase!=Running,status.phase!=Succeeded; exit 1; } || true

# Any Deployment/StatefulSet under desired replicas?
kubectl get deploy -A -o json | jq -r '.items[] | select((.status.readyReplicas // 0) < (.spec.replicas // 1)) | "deploy \(.metadata.namespace)/\(.metadata.name) \(.status.readyReplicas // 0)/\(.spec.replicas)"'
kubectl get sts -A -o json | jq -r '.items[] | select((.status.readyReplicas // 0) < (.spec.replicas // 1)) | "sts \(.metadata.namespace)/\(.metadata.name) \(.status.readyReplicas // 0)/\(.spec.replicas)"'

Only proceed to the next node when ALL of the above are green. If not:

  • Give it another 3 minutes and re-check (workloads with large images may still be pulling).
  • If still not green: STOP the cycle. Uncordon everything, leave the cluster in a stable state, and open a follow-up issue with the failing workloads. Do NOT proceed to update more nodes.

8. Control plane special handling

k3s-cp-1 is the single control plane node and is NOT updated by the weekly os-update.sh cycle — it has its own dedicated procedure and script. See CP1_UPDATE_PROCEDURE.md and scripts/os-update/update-cp-1.sh. Reasons:

  • cp-1 hosts many stateful pods (currently 5 StatefulSets + 8 single-replica Deployments) — a one-shot kubectl drain would trigger the kine/SQLite cascade documented in the "Kine thundering-herd guardrails" section above (DEV-495).
  • cp-1 has 3.7 GiB RAM and (until Phase A of the cp-1 procedure is done) zero swap. The cp-1 procedure adds a durable 4 GiB swapfile before draining.
  • Rebooting cp-1 removes the entire kube-apiserver — the cp-1 procedure polls /livez from an external operator machine, not from cp-1 itself.

os-update.sh will SKIP cp-1 and print a pointer to CP1_UPDATE_PROCEDURE.md.


Finalization

After all nodes are done:

  1. Final cluster health snapshot (same preflight commands as at the start).
  2. Print apt history summary per node so the audit trail has "what changed":
    for host in <all-nodes>; do
      echo "=== $host ==="
      ssh root@$host 'zgrep -h "Commandline\|Install\|Upgrade\|Remove" /var/log/apt/history.log* 2>/dev/null | tail -60'
    done
    
  3. Optional: prune old etcd snapshots to keep disk in check:
    ssh root@$CONTROL_PLANE_HOST 'k3s etcd-snapshot list'
    # delete anything older than 14 days if desired (manual review)
    
  4. Attach the per-run logs to the Paperclip issue that triggered this run.
  5. Update the issue with:
    • Nodes updated (list)
    • Nodes skipped (list + reason)
    • Any package that required manual intervention
    • Reboots performed
    • Final kubectl get nodes output

If everything is clean → mark the issue done. If a node was skipped or errored → mark blocked with the unblock action, or open a child issue for the specific failure.


Recovery / what to do when it goes wrong

Node stuck cordoned after failure: kubectl uncordon <node> — always leave the cluster back in its normal scheduling state.

Node fails to boot after reboot:

  • Check Hetzner console for boot errors (kernel panic, initramfs).
  • Try hcloud server reboot <name> from hcloud CLI.
  • If the node cannot recover: escalate. Do NOT delete or rebuild the Hetzner server without board approval; that is the ADD_WORKER_NODE flow.

k3s-agent won't start after reboot:

  • Check journalctl -u k3s-agent -n 100.
  • Do NOT edit the k3s-agent unit file. Do NOT re-run the k3s installer.
  • Escalate. This is out of scope for the OS-update procedure.

Pods CrashLoopBackOff after node came back:

  • Not an OS-update problem to fix — the node is healthy, apt succeeded. Leave the node uncordoned, stop the cycle, and open an application-level issue.

PDB blocked drain:

  • Never --force the drain. Leave the node uncordoned, skip it, note in the report which PDB blocked and which workload owns it.

Datastore snapshot restore (last resort — CP only):

  • Only if the control plane is broken beyond repair. See K3S_OPERATIONS.md for the --cluster-reset --cluster-reset-restore-path=<snapshot> procedure. Requires board approval — do NOT execute unattended.

Automation entry points

  • infrastructure/scripts/os-update/os-update.sh — full cycle runner (preflight → per-node loop → finalization). Idempotent, resumable via --start-from <node>. Use --dry-run to print the plan without touching anything. Explicitly skips k3s-cp-1.
  • infrastructure/scripts/os-update/update-node.sh <node> — single-node update (all 7 per-node steps). Callable standalone for retry. Not for k3s-cp-1.
  • infrastructure/scripts/os-update/update-cp-1.sh — dedicated cp-1 update flow (add swap, batched stateful eviction, apt, reboot with external liveness monitor). See CP1_UPDATE_PROCEDURE.md.
  • infrastructure/scripts/os-update/cluster-health.sh — the preflight/post-node health check as a standalone command; exits non-zero on any failure.

Read the script sources for the exact behavior before running them. They mirror this procedure step for step.


Weekly schedule

A Paperclip routine (see infrastructure/OS_UPDATE_ROUTINE.md) fires this procedure once per week. The routine creates a task whose description points here. The assigned agent reads this document and executes the automation.


Change history

Date Change By
2026-08-09 Initial procedure + automation scripts CTO agent (DEV-462)
2026-08-16 Added kine thundering-herd guardrails after DEV-495 incident (worker-3 drain caused CP kine cascade + near-OOM on cp-1). Four new rules: pre-plan StatefulSet drain order, swap on cp-1, halt on kubectl >5 s, cp-1 update as separate design task. CTO agent (DEV-478/DEV-495)