48 lines
2.1 KiB
Markdown
48 lines
2.1 KiB
Markdown
|
|
# OS Update Automation
|
||
|
|
|
||
|
|
Scripts that implement the weekly rolling Ubuntu OS-update procedure.
|
||
|
|
|
||
|
|
**Authoritative doc:** [`../../OS_UPDATE_PROCEDURE.md`](../../OS_UPDATE_PROCEDURE.md) — read it first. The scripts here mirror that procedure step-for-step.
|
||
|
|
|
||
|
|
## Files
|
||
|
|
|
||
|
|
| Script | Purpose |
|
||
|
|
|--------|---------|
|
||
|
|
| `cluster-health.sh` | Non-zero if any node is not Ready, any pod is not Running/Ready, or any Deployment/StatefulSet is below its desired replica count. Used at preflight and after each node. |
|
||
|
|
| `update-node.sh <node>` | Drain, apt-update, reboot-if-required, wait for Ready, uncordon, post-node health check. Retriable per-node. |
|
||
|
|
| `os-update.sh` | Full cycle runner: preflight → etcd snapshot → ordered per-node loop → finalization + apt history digest. |
|
||
|
|
|
||
|
|
## Order of operations (encoded in `os-update.sh`)
|
||
|
|
|
||
|
|
1. Workers with no stateful affinity concerns first.
|
||
|
|
2. `fsn1` workers (potential Stalwart hosts) last among workers — see [`stalwart-datacenter-affinity`](../../K3S_OPERATIONS.md) note.
|
||
|
|
3. `k3s-cp-1` last (single control plane).
|
||
|
|
4. One node at a time. Never in parallel.
|
||
|
|
|
||
|
|
## Guardrails the scripts enforce
|
||
|
|
|
||
|
|
- Preflight cluster health failure → refuse to start.
|
||
|
|
- Drain with a PDB conflict → uncordon and mark the node `SKIPPED_DRAIN_FAILED`, never `--force`.
|
||
|
|
- Node reboot fails to come back within timeout → hard stop, escalate. Do NOT rebuild the node or touch k3s config.
|
||
|
|
- kubelet doesn't return `Ready` → hard stop, escalate. Do NOT touch k3s config.
|
||
|
|
- Post-node cluster health check red → hard stop, do not proceed to the next node.
|
||
|
|
- Absolute rule: **no k3s config, no manifests, no PVs, no service files touched** — apt/dpkg only.
|
||
|
|
|
||
|
|
## Typical invocations
|
||
|
|
|
||
|
|
```bash
|
||
|
|
# Dry-run: print the ordered plan, touch nothing.
|
||
|
|
./os-update.sh --dry-run
|
||
|
|
|
||
|
|
# Full cycle (weekly, triggered by the Paperclip routine).
|
||
|
|
./os-update.sh
|
||
|
|
|
||
|
|
# Retry a single node (after fixing a manual issue).
|
||
|
|
./update-node.sh k3s-worker-3
|
||
|
|
|
||
|
|
# Restart a partial cycle from a specific node onward.
|
||
|
|
./os-update.sh --start-from k3s-worker-4
|
||
|
|
```
|
||
|
|
|
||
|
|
Logs land in `/tmp/os-update-<UTC-timestamp>/` on the machine that ran the cycle.
|