- Add k3s node provisioning script with version pinning - Add comprehensive K3S_OPERATIONS.md documentation - Add k3s system-upgrade-controller configuration This addresses DEV-221: prevents version skew issues by: 1. Enforcing version pinning when adding new nodes 2. Providing automated provisioning script 3. Setting up automated upgrades via upgrade controller 4. Documenting all cluster operations procedures Co-Authored-By: Paperclip <noreply@paperclip.ing>
3.4 KiB
3.4 KiB
k3s System Upgrade Controller
This directory contains configuration for automated k3s cluster upgrades using the system-upgrade-controller.
Installation
The system-upgrade-controller should already be installed. If not, install it with:
kubectl apply -f https://github.com/rancher/system-upgrade-controller/releases/latest/download/system-upgrade-controller.yaml
Verify installation:
kubectl get pods -n system-upgrade
kubectl get plans -n system-upgrade
Upgrade Plans
Two upgrade plans are configured:
k3s-server.yaml- Upgrades control plane nodesk3s-agent.yaml- Upgrades worker nodes
Triggering an Upgrade
To upgrade the cluster to a new k3s version:
-
Edit the plans to set the desired version:
kubectl edit plan k3s-server -n system-upgrade kubectl edit plan k3s-agent -n system-upgradeChange the
versionfield:spec: version: v1.37.0+k3s1 # Update to desired version -
Or apply updated plan files:
# Update version in the YAML files first kubectl apply -f infrastructure/k3s-upgrade/k3s-server.yaml kubectl apply -f infrastructure/k3s-upgrade/k3s-agent.yaml -
Monitor the upgrade:
# Watch nodes being upgraded kubectl get nodes -w # Check upgrade jobs kubectl get jobs -n system-upgrade # View controller logs kubectl logs -n system-upgrade -l upgrade.cattle.io/controller=system-upgrade-controller -f
How It Works
- The controller watches for Plan updates
- When a new version is detected, it creates Jobs for each matching node
- Nodes are cordoned and drained before upgrade
- The upgrade job runs
k3s-upgradescript to update k3s - Node is rebooted (if configured) and uncordoned
- Process repeats for next node (respecting concurrency limits)
Configuration Options
Key fields in the Plan spec:
- version: Target k3s version (e.g.,
v1.36.2+k3s1) - concurrency: Number of nodes to upgrade simultaneously (default: 1)
- nodeSelector: Which nodes to upgrade (e.g., control-plane, worker)
- cordon: Cordon nodes before upgrade (default: true)
- drain: Drain pods before upgrade (default: true)
- upgrade.cattle.io/tolerations: Allow upgrade pods on tainted nodes
Safety
- Upgrades are rolling - one node at a time by default
- Nodes are drained before upgrade to avoid pod disruption
- The controller respects Pod Disruption Budgets
- Failed upgrades can be rolled back by changing the version back
Troubleshooting
Upgrade stuck
Check job status:
kubectl get jobs -n system-upgrade
kubectl describe job <job-name> -n system-upgrade
kubectl logs -n system-upgrade job/<job-name>
Plan not triggering
Ensure the version changed:
kubectl get plan k3s-server -n system-upgrade -o yaml | grep version
Check controller logs:
kubectl logs -n system-upgrade -l upgrade.cattle.io/controller=system-upgrade-controller
Manual intervention needed
Delete stuck jobs:
kubectl delete job <job-name> -n system-upgrade
Uncordon nodes manually if needed:
kubectl uncordon <node-name>