docs(lxc110): incident DNS/OOM 2026-07-31 + zram pe pve1 + limite TTS

LXC 110 (moltbot) a picat dupa un OOM local containerului urmat de reboot:
tailscaled a ramas delogat, iar containerul mostenea resolv.conf-ul Tailscale
de la host-ul pve1 (doar MagicDNS 100.100.100.100) -> rezolutie DNS zero desi
L3 era functional -> echo-core in crash-loop pe telegram.error.TimedOut.

Fixuri aplicate:
- pct set 110 --nameserver '10.0.20.1 1.1.1.1' (elimina dependenta de Tailscale)
- zram-tools pe pve1 (8G zstd, prio 100) - `swap: 4096` din config era fictiv,
  host-ul nu avea niciun swap; root pe ZFS deci zram, nu swapfile
- MemoryHigh/MemoryMax pe pocket-tts + supertonic-tts - toate serviciile user
  rulau cu limite `infinity`, de unde OOM-uri recurente (Apr 25, May 28 x4)
  cu victime aleatorii alese dupa oom_score_adj

Documentatie:
- nou post-mortem in cluster/incidents/, adaugat in indexuri
- README lxc110: host corectat pveelite -> pve1 (+ RAM/CPU/storage reale),
  comenzile pct redirectionate spre nodul corect
- README lxc171: tabel cu starea swap pe cele 3 noduri (pveelite are zvol,
  nu zram - nu necesita acelasi fix)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Claude Agent
2026-07-31 10:14:54 +00:00
parent 874698a242
commit 1df533edba
5 changed files with 292 additions and 17 deletions

View File

@@ -3,7 +3,7 @@
**Director:** `proxmox/lxc110-moltbot/`
**VMID:** 110
**IP:** 10.0.20.173 (intern) | 100.120.119.70 (Tailscale)
**Host Proxmox:** pveelite (10.0.20.202)
**Host Proxmox:** pve1 (10.0.20.200)
**Rol:** Bot AI pentru Telegram și WhatsApp cu Claude Opus 4.5
---
@@ -16,12 +16,13 @@
| Hostname | moltbot |
| IP intern | 10.0.20.173 |
| IP Tailscale | 100.120.119.70 |
| Host Proxmox | pveelite (10.0.20.202) |
| Host Proxmox | pve1 (10.0.20.200) |
| User serviciu | `moltbot` |
| Parola user | `Moltbot2026!` |
| Storage | local-zfs (8GB) |
| RAM | 4GB |
| CPU | 2 cores |
| Storage | local-zfs (50GB) |
| RAM | 8GB (+ 4GB swap) |
| CPU | 6 cores |
| DNS | `10.0.20.1 1.1.1.1` (explicit, vezi mai jos) |
| OS | Ubuntu 24.04 LTS |
## Componente Instalate
@@ -196,22 +197,85 @@ journalctl --user -u clawdbot-gateway -f
## Administrare via Proxmox
### De pe pvemini (sau alt nod cluster)
### De pe pve1 (sau alt nod cluster)
```bash
# Status container
ssh root@10.0.20.202 "pct status 110"
ssh root@10.0.20.200 "pct status 110"
# Exec comandă
ssh root@10.0.20.202 "pct exec 110 -- <comandă>"
ssh root@10.0.20.200 "pct exec 110 -- <comandă>"
# Stop/Start
ssh root@10.0.20.202 "pct stop 110"
ssh root@10.0.20.202 "pct start 110"
ssh root@10.0.20.200 "pct stop 110"
ssh root@10.0.20.200 "pct start 110"
# Console
ssh root@10.0.20.202 "pct enter 110"
ssh root@10.0.20.200 "pct enter 110"
```
## DNS & Swap (incident 2026-07-31)
### DNS — nu depinde de Tailscale
Containerul **nu** avea `nameserver` în `pct config`, deci moștenea `/etc/resolv.conf` de
la host-ul pve1 — unde fișierul e scris de Tailscale și conținea doar MagicDNS
(`100.100.100.100`). Când `tailscaled` din container s-a delogat, DNS-ul a murit complet
(deși L3 era OK: `ping 8.8.8.8` funcționa) → `echo-core` în crash-loop pe
`telegram.error.TimedOut`.
Fix aplicat — nameserver explicit:
```bash
ssh root@10.0.20.200 "pct set 110 --nameserver '10.0.20.1 1.1.1.1'"
```
> **Regula de diagnostic:** la „nu mai ajunge la internet", verifică întâi
> `getent hosts api.telegram.org` și `cat /etc/resolv.conf`, **nu** ping-ul.
> Ping-ul poate merge perfect cu DNS-ul complet mort.
### Swap — necesită zram pe host
`swap: 4096` din `pct config` e doar o limită cgroup (`memory.swap.max`), nu un backing
store. Cât timp pve1 nu avea swap deloc, în container `free -h` arăta `Swap: 0B`.
Rezolvat prin zram pe pve1 (root-ul e pe ZFS → swapfile exclus):
```bash
# pe pve1 (host)
apt install -y zram-tools
cat > /etc/default/zramswap <<'EOF'
ALGO=zstd
SIZE=8192
PRIORITY=100
EOF
systemctl restart zramswap.service
swapon --show # /dev/zram0 8G prio 100
```
Verificare în container: `free -h``Swap: 4.0Gi`.
### Limite de memorie pe TTS
Serviciile user systemd rulau toate cu `MemoryMax=infinity` / `MemoryHigh=infinity`, deci
oricare putea consuma singur cei 8 GB → OOM-uri recurente (Apr 25, May 28 ×4, Jul 31), cu
victime aparent aleatorii (`dbus-daemon`, `sd-pam`) alese după `oom_score_adj`.
Aplicat pe stiva TTS — `MemoryHigh` face throttle + reclaim în loc de kill:
| Serviciu | RSS repaus | MemoryHigh | MemoryMax |
|----------|-----------|------------|-----------|
| `pocket-tts` | ~945 MB | 1536M | 2G |
| `supertonic-tts` | ~500 MB | 1024M | 1536M |
Drop-in-uri: `~/.config/systemd/user/<serviciu>.service.d/limits.conf`
```bash
su - moltbot
export XDG_RUNTIME_DIR=/run/user/1000
systemctl --user show pocket-tts -p MemoryHigh,MemoryMax,MemoryCurrent
```
**Post-mortem complet:** `../cluster/incidents/2026-07-31-lxc110-dns-tailscale-oom.md`
## Troubleshooting
### OpenClaw gateway se restartează continuu (OOM kill)
@@ -223,10 +287,14 @@ journalctl --user -u openclaw-gateway | grep -i oom
free -h
systemctl --user status openclaw-gateway
# Soluție: Crește RAM-ul containerului pe Proxmox
# ssh root@pveelite
# Verifică dacă OOM-ul e local containerului sau al host-ului
# ssh root@10.0.20.200 "cat /sys/fs/cgroup/lxc/110/memory.events" # oom_kill > 0 => local
# ssh root@10.0.20.200 "free -h; swapon --show" # zram0 8G activ?
# Soluție: Crește RAM-ul containerului pe Proxmox (actual: 8192)
# ssh root@10.0.20.200
# pct stop 110
# pct set 110 --memory 4096 # 4GB (minim recomandat pentru OpenClaw)
# pct set 110 --memory 12288
# pct start 110
# Curăță sesiunile vechi pentru a reduce consumul
@@ -303,7 +371,7 @@ scp -r moltbot@10.0.20.173:~/.clawdbot ./backup-moltbot-$(date +%Y%m%d)/
### Backup complet LXC (via Proxmox)
```bash
ssh root@10.0.20.202 "vzdump 110 --storage local --compress zstd"
ssh root@10.0.20.200 "vzdump 110 --storage local --compress zstd"
```
## Provider AI - Anthropic
@@ -347,5 +415,5 @@ clawdbot onboard
---
**Data setup:** 2026-01-29
**Ultima actualizare:** 2026-01-29
**Ultima actualizare:** 2026-07-31 (host pve1, DNS explicit, zram — incident DNS/OOM)
**Autor:** Claude Code