Since DEV-510 landed HA control plane (2026-08-22), the cluster runs cp-1
plus cp-2/cp-3. The previous exclusion in os-update.sh matched `k3s-cp-1`
by hard-coded name only, which would have caused cp-2 and cp-3 to be
treated as regular fsn1 workers and drained/rebooted without the
CP-specific procedure.
Fix: select the CP list from `kubectl get nodes -l
node-role.kubernetes.io/control-plane` and skip any of those nodes. This
covers all present and future CPs automatically.
Also updated OS_UPDATE_PROCEDURE.md topology table and order rule to
document that all three CPs exist and are excluded from the weekly
cycle. The HA-aware CP OS-update procedure is a separate follow-up.
Verified on the current cluster:
[plan] EXCLUDING control-plane nodes: k3s-cp-1 k3s-cp-2 k3s-cp-3
[plan] ordered nodes (6): k3s-worker-4 k3s-update-runner k3s-worker-1
k3s-worker-2 k3s-worker-3 k3s-worker-5
Co-Authored-By: Paperclip <noreply@paperclip.ing>
|
||
|---|---|---|
| .. | ||
| cluster-health.sh | ||
| ensure-node-docker.sh | ||
| os-update.sh | ||
| README.md | ||
| update-cp-1.sh | ||
| update-node.sh | ||
OS Update Automation
Scripts that implement the weekly rolling Ubuntu OS-update procedure.
Authoritative doc: ../../OS_UPDATE_PROCEDURE.md — read it first. The scripts here mirror that procedure step-for-step.
Files
| Script | Purpose |
|---|---|
cluster-health.sh |
Non-zero if any node is not Ready, any pod is not Running/Ready, or any Deployment/StatefulSet is below its desired replica count. Used at preflight and after each node. |
update-node.sh <node> |
Drain, apt-update, reboot-if-required, wait for Ready, ensure Docker on labeled nodes, uncordon, post-node health check. Retriable per-node. |
ensure-node-docker.sh <node> |
Idempotent: on nodes labeled basicstack.de/docker=true, install docker.io if missing, apt-mark manual, enable+start the docker service, wait for /var/run/docker.sock, re-apply the label. No-op on nodes without the label. Called from update-node.sh between "kubelet Ready" and "uncordon"; also runnable ad-hoc for recovery. |
os-update.sh |
Full cycle runner: preflight → etcd snapshot → ordered per-node loop → finalization + apt history digest. |
Order of operations (encoded in os-update.sh)
- Workers with no stateful affinity concerns first.
fsn1workers (potential Stalwart hosts) last among workers — seestalwart-datacenter-affinitynote.k3s-cp-1last (single control plane).- One node at a time. Never in parallel.
Guardrails the scripts enforce
- Preflight cluster health failure → refuse to start.
- Drain with a PDB conflict → uncordon and mark the node
SKIPPED_DRAIN_FAILED, never--force. - Node reboot fails to come back within timeout → hard stop, escalate. Do NOT rebuild the node or touch k3s config.
- kubelet doesn't return
Ready→ hard stop, escalate. Do NOT touch k3s config. - Post-node cluster health check red → hard stop, do not proceed to the next node.
- Absolute rule: no k3s config, no manifests, no PVs, no service files touched — apt/dpkg only.
Typical invocations
# Dry-run: print the ordered plan, touch nothing.
./os-update.sh --dry-run
# Full cycle (weekly, triggered by the Paperclip routine).
./os-update.sh
# Retry a single node (after fixing a manual issue).
./update-node.sh k3s-worker-3
# Restart a partial cycle from a specific node onward.
./os-update.sh --start-from k3s-worker-4
# Ad-hoc: reconcile Docker on a single node (e.g. after emergency ops).
./ensure-node-docker.sh k3s-worker-3
Logs land in /tmp/os-update-<UTC-timestamp>/ on the machine that ran the cycle.