Investigation confirmed all SMTP/IMAP ports working correctly as of 21:00 UTC. SMTP issue at 20:31 UTC coincided with LoadBalancer IP updates (10.42.1.1 → 178.105.17.239). External port tests verify correct protocol greetings on all ports. Co-Authored-By: Paperclip <noreply@paperclip.ing>
9.2 KiB
Stalwart Stability Check Procedure
CRITICAL OPERATIONAL RULES
READ THIS FIRST BEFORE ANY INVESTIGATION OR FIX
Single Point of Truth
This repository (https://forgejo.basicstack.de/basicstack/stack.basicstack.de) is the single point of truth for all Stalwart configuration.
- ✅ DO: Make all changes in this repository
- ✅ DO: Sync changes via Argo CD
- ❌ NEVER: Make direct kubectl changes without updating this repository first
- ❌ NEVER: Reset to recovery/default admin state as a "fix" (destroys working OIDC config)
Before Any Investigation
-
Read the git history:
git log --oneline --all -- apps/stalwart/ | head -20 git show <commit-hash> # for details on relevant fixes -
Check recent changes:
git log --since="7 days ago" -- apps/stalwart/ -
Read this file to understand what's been tried and what's known
-
Document your findings in this repository under issues, not just in task comments
Scope Discipline
- Web-UI working but SMTP/IMAP broken = Network/port issue, NOT auth/config issue
- Web-UI login broken = Auth/config issue, NOT network issue
- Do NOT "fix" the auth layer when the problem is transport/network, and vice-versa
Known Issues and History
Issue #1: k3s Service Networking Instability (Root Cause - FIXED 2026-07-11)
Commit: b0f2acf - "Stalwart reliability hardening: fix k3s service networking issues"
Symptom: Pod healthy but service ClusterIP routing breaks intermittently
Root Cause: k3s uses iptables-legacy, not iptables-nft (see memory file: k3s-iptables-backend.md)
- Service ClusterIP routing fails despite correct iptables rules and healthy endpoints
- Direct pod IP works, service ClusterIP fails with "Connection reset by peer"
- Required pod restart + Traefik restart to restore service
Fix Applied:
- Added startup probe (60s delay) to prevent premature service registration
- Added comprehensive monitoring (ServiceMonitor, PrometheusRule, blackbox probes)
- External health checks test full service path, not just pod health
- Alerting for: service failures, high latency, pod restarts, backup failures
Verification: 7-day stability window to confirm fix (2026-07-11 11:04 UTC → 2026-07-18 11:04 UTC)
Current Stability Check (DEV-300)
Original Window (DEV-235): 2026-07-11 11:04 UTC → 2026-07-18 11:04 UTC
Issue: LoadBalancer used incorrect IP (10.42.1.1 instead of 178.105.17.239) until 2026-07-13 19:10 UTC
Decision: Window restarted per CEO approval (2026-07-13 21:07 UTC)
Active Window: 2026-07-13 21:07 UTC → 2026-07-20 21:07 UTC
Rationale: Clean 7-day verification with working external access
Baseline Status: Web-UI SMTP and IMAP confirmed working by CEO (2026-07-13 21:07 UTC)
Investigation Results (2026-07-13)
Infrastructure Status: ✅ ALL VERIFIED WORKING (as of 21:00 UTC)
- SMTP (ports 25, 587, 465): Responding with correct ESMTP greeting
- IMAP (ports 143, 993): Responding with correct IMAP greeting
- LoadBalancer external IP: 178.105.17.239 (correct)
- Pod status: Running, started at 19:26 UTC
- Service configuration: Correct port mappings and selectors verified
- Monitoring: PrometheusRules and ServiceMonitors in place
SMTP Issue Timeline (2026-07-13):
- 20:16 UTC: LoadBalancer IP updates visible in events (10.42.1.1 → 178.105.17.239)
- 20:31 UTC: User reported SMTP port returning IMAP greeting (Thunderbird error)
- 21:00 UTC: External port tests confirm all ports working correctly with proper greetings
- Assessment: Transient issue during LoadBalancer IP transition, now resolved
- Action Required: User to retest SMTP sending with fresh client connection
Architecture Overview
External Traffic
↓
Hetzner LoadBalancer (SMTP: 25,587,465 / IMAP: 143,993)
↓
Kubernetes Service (LoadBalancer type)
↓
stalwart-0 Pod (StatefulSet)
↓
RocksDB Storage (PVC: hcloud-volumes-encrypted)
Web Traffic
↓
Traefik IngressRoute (mail.basicstack.de)
↓
stalwart-http Service (ClusterIP)
↓
stalwart-0 Pod (port 8080)
Passive Monitoring Checks
Run these after the 7-day window completes (2026-07-18 11:04 UTC):
- No service routing failures: Verify no manual restarts were needed
- Low pod restart count:
ssh root@178.105.17.239 'kubectl get pods -n stalwart' # RESTARTS column should be 0 or very low - No Prometheus alerts fired: Check alert history for:
StalwartDownStalwartExternalProbeFailureStalwartHighResponseTimeStalwartPodRestartingStalwartBackupFailing
Active Tests
Run these before closing the verification:
1. Pod Restart Test
# Delete the pod
ssh root@178.105.17.239 'kubectl delete pod stalwart-0 -n stalwart'
# Wait and verify it recovers within 3 minutes
ssh root@178.105.17.239 'kubectl get pods -n stalwart -w'
2. Node Drain Test
# Drain the worker node
ssh root@178.105.17.239 'kubectl drain k3s-worker-1 --ignore-daemonsets --delete-emptydir-data'
# Verify StatefulSet reschedules
ssh root@178.105.17.239 'kubectl get pods -n stalwart -o wide'
# Service should come back on another node
# Uncordon after test
ssh root@178.105.17.239 'kubectl uncordon k3s-worker-1'
3. Alert Delivery Test
# Temporarily stop Stalwart
ssh root@178.105.17.239 'kubectl scale statefulset stalwart -n stalwart --replicas=0'
# Wait 3-5 minutes for alert to fire
# Verify alert was delivered to the monitoring channel
# Scale back up
ssh root@178.105.17.239 'kubectl scale statefulset stalwart -n stalwart --replicas=1'
Diagnostic Playbook
Check Pod Health
ssh root@178.105.17.239 'kubectl get pods -n stalwart'
ssh root@178.105.17.239 'kubectl describe pod stalwart-0 -n stalwart'
ssh root@178.105.17.239 'kubectl logs stalwart-0 -n stalwart --tail=100'
Check Service and Endpoints
ssh root@178.105.17.239 'kubectl get svc -n stalwart'
ssh root@178.105.17.239 'kubectl get endpoints -n stalwart'
ssh root@178.105.17.239 'kubectl describe svc stalwart-smtp -n stalwart'
ssh root@178.105.17.239 'kubectl describe svc stalwart-imap -n stalwart'
Check LoadBalancer External IPs
# SMTP LoadBalancer
ssh root@178.105.17.239 'kubectl get svc stalwart-smtp -n stalwart -o jsonpath="{.status.loadBalancer.ingress[0].ip}"'
# IMAP LoadBalancer
ssh root@178.105.17.239 'kubectl get svc stalwart-imap -n stalwart -o jsonpath="{.status.loadBalancer.ingress[0].ip}"'
Test SMTP/IMAP Connectivity
From outside the cluster
# SMTP
telnet <smtp-lb-ip> 25
telnet <smtp-lb-ip> 587
# IMAP
telnet <imap-lb-ip> 143
telnet <imap-lb-ip> 993
From inside the cluster
ssh root@178.105.17.239 'kubectl run -it --rm debug --image=nicolaka/netshoot --restart=Never -- bash'
# Then inside the pod:
telnet stalwart-smtp.stalwart.svc.cluster.local 25
telnet stalwart-imap.stalwart.svc.cluster.local 143
Check Firewall Rules (Hetzner Cloud)
# List firewall rules
hcloud firewall list
hcloud firewall describe <firewall-id>
# Verify ports 25, 587, 465, 143, 993 are open to the LoadBalancer nodes
Check iptables (k3s nodes)
# On the node running the stalwart pod
ssh root@<node-ip> 'iptables-legacy -L -n -v | grep -A 20 stalwart'
ssh root@<node-ip> 'iptables-legacy -t nat -L -n -v | grep -A 20 stalwart'
Check Network Policies
ssh root@178.105.17.239 'kubectl get networkpolicies -n stalwart'
ssh root@178.105.17.239 'kubectl describe networkpolicy -n stalwart'
Issue Documentation Template
When you find an issue, create a file in this directory:
Filename: ISSUE-YYYY-MM-DD-<short-description>.md
Content:
# Issue: <Short Description>
**Date**: YYYY-MM-DD HH:MM UTC
**Discovered by**: <Name/Agent>
**Related Task**: DEV-XXX
## Symptom
What users are experiencing
## Investigation
Steps taken to diagnose
## Root Cause
What was actually broken
## Fix Applied
Changes made (commit hash if available)
## Verification
How to verify the fix worked
## Follow-up
Any remaining concerns or monitoring needed
Evidence Required for Completion
Before marking DEV-235 as done, document:
- Screenshot or log excerpt showing 7 days of uptime (pod age, restart count)
- Pod restart test: before/after pod status with timestamps
- Node drain test: before/after with timestamps showing pod migration
- Alert delivery: confirmation that alert fired and was received
- Summary: Any issues found during the 7-day window and how they were resolved
Related Documentation
- README.md - Main Stalwart deployment documentation
- CERTIFICATE-RENEWAL.md - TLS certificate automation
- manual_config_steps.md - Manual configuration procedures
- Memory files:
cluster-infrastructure.md- Cluster layout and node IPsk3s-iptables-backend.md- k3s uses iptables-legacy (root cause of past outages)
Emergency Contacts
DO NOT take destructive actions without approval:
- Resetting to recovery admin
- Dropping the RocksDB database
- Force-deleting the PVC
Instead: Post findings to the issue, request approval, and wait for explicit confirmation before proceeding.