incident: SFTPGo outage 2026-09-02 — CT299 tailscaled death + stale lxc-attach recovery
- New incident page documenting CT299 SFTPGo outage (02:18 AWST) - Root cause: swap exhaustion (43Gi, 100%) killed tailscaled; 5 stale lxc-attach PIDs blocked pct start - Resolution: killed PIDs 3037951/3334928/3352203/3402511/4004148, pct stop/start 299 - Updated current-state.md with Recent Changes (2026-09-02), bumped updated to 2026-09-04 - Updated log.md with incident entry - Updated index.md last-updated date No secrets written. Verified end-to-end: tailscale direct, SFTPGo WebAdmin HTTP 401, SFTP banner SSH-2.0-SFTPGo_2.7.0.
This commit is contained in:
+5
-2
@@ -3,7 +3,7 @@ title: Current State
|
||||
type: concept
|
||||
status: active
|
||||
created: 2026-07-22
|
||||
updated: 2026-08-14
|
||||
updated: 2026-09-04
|
||||
review_after: 2026-10-22
|
||||
confidence: high
|
||||
tags: [system]
|
||||
@@ -11,7 +11,10 @@ sources: []
|
||||
---
|
||||
|
||||
# Current State
|
||||
Updated: 2026-08-14
|
||||
Updated: 2026-09-04
|
||||
|
||||
## Recent Changes (2026-09-02)
|
||||
- **SFTPGo outage resolved (02:18 AWST):** Dashboard flagged "service down". CT299 swap fully exhausted (43Gi used, 100%); CT299's `tailscaled` had died, breaking tailnet reachability; PVE's `pct start` then failed because 5 stale `lxc-attach` probes from earlier diagnostics pinned CT299's cgroup (`/sys/fs/cgroup/lxc/299/ns/.lxc`). Killed those PIDs (3037951/3334928/3352203/3402511/4004148), `pct stop 299 --skiplock` + `pct start 299`, fresh PID 3471141, tailscaled reconnected direct. Verified end-to-end: `tailscale ping sftpgo` → direct 192.168.178.92:41641, SFTPGo WebAdmin HTTP 401, SFTP banner `SSH-2.0-SFTPGo_2.7.0`. See [[incidents/2026-09-02-sftpgo-outage]].
|
||||
|
||||
## Attested facts
|
||||
- **Config version:** 33 — Verified, `~/.hermes/config.yaml`, 2026-07-25
|
||||
|
||||
Reference in New Issue
Block a user