Doua defecte, al doilea gasit la verificarea primului.
1. Procesul `claude` al unui fir isi porneste serverele MCP (npm/sh/node pentru
playwright), care nu se numesc `claude` si scapau cautarii. Masurat pe firul viu:
claude 300 MB + MCP 137 MB = ~440 MB per fir, nu 300. find_orphans merge acum pe
arbore, iar kill_orphans opreste copiii inaintea parintilor (altfel se reparenteaza
si scapa).
2. PERICULOS: definitia initiala a orfanului ("orice `claude` absent din state.json")
prindea sesiunile Claude interactive de pe container. Rularea seaca propunea 25 de
procese / ~3864 MB — sesiunile de lucru din tmux, inclusiv cea din care ar fi fost
data comanda. `/cleanup force:True` din Discord si-ar fi omorat propria sesiune.
Apartenenta la cgroup-ul `claude-discord.service` devine conditie necesara pentru
toate familiile, fail-closed la cgroup necitibil. Dupa fix: zero procese raportate.
Test de regresie pe instantaneul real (doua sesiuni in tmux-spawn-*.scope, una in
claude-discord.service): se raporteaza doar a treia. Verificat prin mutant ca testul
musca — cu filtrul scos pica 3 teste.
3. INTERFACES.md cerea pid_start_time in ticks, dar session_store scrie secunde (si
state.json viu contine secunde). Contractul era imprecis, nu codul: un consumator
care compara ticks nu s-ar potrivi niciodata, iar cum acea comparatie protejeaza
turul in desfasurare, esecul ar fi fost tacut si ar fi facut eligibil exact ce
trebuia protejat. Documentat, cu avertismentul explicit.
Bug colateral prins de agent: raportul cu rezultate de omorare ajungea la 2289 caractere,
peste limita Discord — mesajul ar fi fost respins exact la `/cleanup force:True`. Buget
de caractere adaugat.
Suita: 322 passed cu discord.py, 319 passed + 3 skipped fara. Rulare seaca pe container: curat.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B29CApsP1JkSdjYaGaHpE7
694 lines
26 KiB
Python
694 lines
26 KiB
Python
#!/usr/bin/env python3
|
|
"""cleanup.py — comanda `!cleanup`: procese lasate in urma de tururi anterioare (T13).
|
|
|
|
De ce exista (Codex #8): `KillMode=control-group` opreste arborele serviciului la
|
|
restart, dar NU prinde procesele care s-au desprins — daemoni pornite cu `&` sau
|
|
`nohup`, servere de dezvoltare lansate intr-un tur si uitate acolo, joburi lungi
|
|
reparentate la init. Fiecare proces `claude` are ~406 MB RSS masurat, iar
|
|
containerul are istoric de OOM: acumularea lor nu e cosmetica.
|
|
|
|
Contract (INTERFACES.md, granita A <-> C):
|
|
|
|
find_orphans(state) -> [{"pid", "cmdline", "age_s", "rss_mb"}]
|
|
kill_orphans(orphans, dry_run=True) -> [{...}]
|
|
|
|
`dry_run=True` e implicit si e intentionat: omul vede intai ce s-ar omori.
|
|
|
|
Un fir NU e un singur proces. Masurat pe un fir viu in productie:
|
|
|
|
78848 python 41 MB bot.py
|
|
79423 claude 300 MB --resume 0a3f9f33...
|
|
79454 npm exec @playwright/mcp@latest 52 MB
|
|
79482 sh -c playwright-mcp 1 MB
|
|
79483 node .../playwright-mcp 84 MB
|
|
|
|
Adica ~440 MB pe fir, nu 300, iar copiii (serverele MCP) NU se numesc `claude`.
|
|
Daca procesul `claude` moare izolat, copiii lui raman si se reparenteaza — exact
|
|
scurgerea lenta pentru care exista comanda asta, si exact ce scapa unui filtru pe
|
|
nume. De aceea cautarea merge pe ARBORE, iar criteriul pentru copiii deja
|
|
reparentati (care si-au pierdut parintele) e apartenenta la cgroup-ul serviciului.
|
|
|
|
Ce NU e considerat orfan:
|
|
* procesele `claude` inregistrate in state.json si toti descendentii lor
|
|
(sunt turul care ruleaza chiar acum, cu tot cu serverele lui MCP);
|
|
* procesul curent, parintii lui si botul insusi (bot.py);
|
|
* orice proces al altui utilizator;
|
|
* infrastructura sesiunii (systemd --user, sshd, tmux, code-server, dbus...).
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import os
|
|
import pathlib
|
|
import signal
|
|
import sys
|
|
import time
|
|
|
|
try: # pragma: no cover - config.py apartine Lane A, poate lipsi la momentul asta
|
|
import config as _config # type: ignore
|
|
except Exception: # pragma: no cover
|
|
_config = None
|
|
|
|
PROC = pathlib.Path("/proc")
|
|
CLOCK_TICKS = float(os.sysconf("SC_CLK_TCK"))
|
|
PAGE_SIZE = float(os.sysconf("SC_PAGE_SIZE"))
|
|
|
|
# Cgroup-ul serviciului. E CONDITIA NECESARA pentru orice candidat: pe containerul
|
|
# asta ruleaza permanent sesiuni Claude interactive (tmux, ttyd, VS Code) care nu au
|
|
# nicio legatura cu puntea. Cgroup-urile le separa curat:
|
|
#
|
|
# sesiune de lucru: 0::/user.slice/user-1000.slice/user@1000.service/tmux-spawn-<uuid>.scope
|
|
# procesele puntii: 0::/user.slice/user-1000.slice/user@1000.service/app.slice/claude-discord.service
|
|
#
|
|
# Fara filtrul asta, `/cleanup force:True` dat din Discord si-ar omori propria sesiune
|
|
# si tot ce ruleaza omul in tmux. Nu e ipotetic: prima versiune propunea exact asta.
|
|
SERVICE_CGROUP = "claude-discord.service" # rezerva, cand nu se poate deriva
|
|
|
|
# Unitati care NU sunt niciodata "serviciul nostru", oricat de mult ar semana:
|
|
# user@1000.service e managerul intregii sesiuni de utilizator, deci ar cuprinde TOT.
|
|
CGROUP_NICIODATA = ("user@", "user-", "init.scope", "session-")
|
|
|
|
# Numele executabilului CLI-ului. Se compara pe BASENAME-ul fiecarui argument, nu ca
|
|
# subsir: utilizatorul containerului se numeste tot `claude`, deci orice cale din
|
|
# home-ul lui contine cuvantul si un match pe subsir ar da fals pozitiv pe tot.
|
|
CLAUDE_BASENAMES = ("claude", "claude.exe")
|
|
CLAUDE_PATH_MARKERS = ("claude-code/bin/claude", "node_modules/.bin/claude")
|
|
|
|
# Procese care nu se ating niciodata, chiar daca sunt in cgroup si reparentate.
|
|
NEVER_KILL = (
|
|
"systemd", "sshd", "dbus-daemon", "tmux", "code-server", "vscode-server",
|
|
"bot.py", "claude-discord", "login", "bash -l", "(sd-pam)", "pytest",
|
|
)
|
|
|
|
GRACE_S = 3.0 # cat asteptam intre SIGTERM si SIGKILL
|
|
MAX_REPORT_ROWS = 15 # cate procese se listeaza cel mult
|
|
MAX_REPORT_CHARS = 1800 # limita unui mesaj Discord e 2000; lasam loc de subsol
|
|
|
|
|
|
# --- citire /proc ----------------------------------------------------------
|
|
|
|
def _read(path: pathlib.Path) -> str:
|
|
try:
|
|
return path.read_text(errors="replace")
|
|
except Exception:
|
|
return ""
|
|
|
|
|
|
def _boot_epoch() -> float:
|
|
"""Momentul (epoch) fata de care e masurat `starttime` din /proc/<pid>/stat.
|
|
|
|
ATENTIE, capcana de container: pe LXC 171 /proc/uptime e virtualizat de lxcfs
|
|
(arata uptime-ul CONTAINERULUI), in timp ce `starttime` ramane raportat la boot-ul
|
|
GAZDEI. Scaderea celor doua da varste negative (verificat: -561s pentru un proces
|
|
pornit acum 10 minute; si `ps -o etimes` arata acolo 4123168064). `btime` din
|
|
/proc/stat e coerent cu `starttime`, deci ea e referinta corecta.
|
|
"""
|
|
for line in _read(PROC / "stat").splitlines():
|
|
if line.startswith("btime"):
|
|
try:
|
|
return float(line.split()[1])
|
|
except Exception:
|
|
break
|
|
# ultima varianta: boot dedus din uptime (poate fi gresit in container)
|
|
try:
|
|
return time.time() - float(_read(PROC / "uptime").split()[0])
|
|
except Exception:
|
|
return 0.0
|
|
|
|
|
|
def _parse_stat(pid: int) -> tuple[int, float] | None:
|
|
"""(ppid, starttime_ticks) din /proc/<pid>/stat.
|
|
|
|
comm-ul e intre paranteze si poate contine spatii, deci se taie dupa ultimul ')'.
|
|
"""
|
|
raw = _read(PROC / str(pid) / "stat")
|
|
if not raw:
|
|
return None
|
|
try:
|
|
rest = raw[raw.rindex(")") + 2:].split()
|
|
# dupa comm, campul 1 e state; ppid e campul 4 global = rest[1]
|
|
ppid = int(rest[1])
|
|
starttime = float(rest[19]) # campul 22 global
|
|
return ppid, starttime
|
|
except Exception:
|
|
return None
|
|
|
|
|
|
def _start_time_matches(ticks: float, expected) -> bool:
|
|
"""Compara `starttime` cu ce a notat state.json, in AMBELE conventii.
|
|
|
|
INTERFACES.md spune "campul 22 din /proc/<pid>/stat", adica ticks. Masurat pe
|
|
firul viu 89112, state.json noteaza insa 69065.87, iar /proc da 6906587 ticks —
|
|
exact de 100 de ori mai mult, adica secunde (ticks / SC_CLK_TCK).
|
|
Acceptam ambele: e o comparatie care PROTEJEAZA un proces, iar o nepotrivire
|
|
aici inseamna ca declaram orfan un fir viu si ii omoram serverele MCP la
|
|
mijlocul turului. In caz de dubiu, protejam.
|
|
"""
|
|
try:
|
|
expected = float(expected)
|
|
except (TypeError, ValueError):
|
|
return False
|
|
if abs(ticks - expected) <= 1.0: # ticks (conform INTERFACES)
|
|
return True
|
|
return abs(ticks / CLOCK_TICKS - expected) <= 1.0 # secunde (ce scrie Lane A)
|
|
|
|
|
|
def _is_zombie(pid: int) -> bool:
|
|
"""Un proces terminat dar nereaped are inca /proc/<pid>/stat; e mort, nu viu."""
|
|
raw = _read(PROC / str(pid) / "stat")
|
|
try:
|
|
return raw[raw.rindex(")") + 2:].split()[0] in ("Z", "X", "x")
|
|
except Exception:
|
|
return False
|
|
|
|
|
|
def _rss_mb(pid: int) -> float:
|
|
try:
|
|
pages = float(_read(PROC / str(pid) / "statm").split()[1])
|
|
return round(pages * PAGE_SIZE / (1024 * 1024), 1)
|
|
except Exception:
|
|
return 0.0
|
|
|
|
|
|
def _cmdline(pid: int) -> str:
|
|
raw = _read(PROC / str(pid) / "cmdline")
|
|
if raw:
|
|
parts = [p for p in raw.split("\0") if p]
|
|
if parts:
|
|
return " ".join(parts)
|
|
# proces de kernel sau cmdline gol: cadem pe comm
|
|
comm = _read(PROC / str(pid) / "comm").strip()
|
|
return f"[{comm}]" if comm else ""
|
|
|
|
|
|
def _uid(pid: int) -> int | None:
|
|
for line in _read(PROC / str(pid) / "status").splitlines():
|
|
if line.startswith("Uid:"):
|
|
try:
|
|
return int(line.split()[1])
|
|
except Exception:
|
|
return None
|
|
return None
|
|
|
|
|
|
def _cgroup(pid: int) -> str:
|
|
return _read(PROC / str(pid) / "cgroup").strip()
|
|
|
|
|
|
def _service_cgroup_marker() -> str:
|
|
"""Numele unitatii in al carei cgroup ne aflam, ca sa nu fie codat rigid.
|
|
|
|
Se citeste frunza din propriul cgroup:
|
|
* daca e un `.service` (rulam ca serviciu, adica suntem chiar puntea) — aia e,
|
|
deci o redenumire a unitului se propaga singura;
|
|
* daca e un `.scope` (rulam dintr-o sesiune interactiva: tmux, ttyd, ssh) —
|
|
NU o folosim, fiindca ar insemna sa tintim chiar sesiunea omului. Cadem pe
|
|
constanta.
|
|
Se urca doar peste felii (`.slice`): mai sus se afla `user@1000.service`, care ar
|
|
cuprinde intreaga sesiune de utilizator, de unde si lista CGROUP_NICIODATA.
|
|
"""
|
|
try:
|
|
propriu = _cgroup(os.getpid())
|
|
except Exception:
|
|
return SERVICE_CGROUP
|
|
for linie in propriu.splitlines():
|
|
cale = linie.rsplit(":", 1)[-1]
|
|
for componenta in reversed([c for c in cale.split("/") if c]):
|
|
if componenta.endswith(".scope"):
|
|
return SERVICE_CGROUP # sesiune interactiva, nu serviciu
|
|
if componenta.endswith(".service"):
|
|
if any(componenta.startswith(x) for x in CGROUP_NICIODATA):
|
|
return SERVICE_CGROUP
|
|
return componenta
|
|
if componenta.endswith(".slice"):
|
|
continue # felie intermediara, urcam
|
|
break
|
|
return SERVICE_CGROUP
|
|
|
|
|
|
def scan_processes() -> dict[int, dict]:
|
|
"""Instantaneu al proceselor utilizatorului curent, indexat pe pid."""
|
|
me = os.getuid()
|
|
boot = _boot_epoch()
|
|
now = time.time()
|
|
out: dict[int, dict] = {}
|
|
try:
|
|
entries = [e for e in os.listdir(PROC) if e.isdigit()]
|
|
except Exception:
|
|
return out
|
|
for entry in entries:
|
|
pid = int(entry)
|
|
st = _parse_stat(pid)
|
|
if st is None:
|
|
continue # procesul a disparut intre listare si citire
|
|
uid = _uid(pid)
|
|
if uid is None or uid != me:
|
|
continue
|
|
ppid, starttime = st
|
|
age = now - (boot + starttime / CLOCK_TICKS) if boot else 0.0
|
|
out[pid] = {
|
|
"pid": pid,
|
|
"ppid": ppid,
|
|
"cmdline": _cmdline(pid),
|
|
"age_s": int(max(age, 0)),
|
|
"rss_mb": _rss_mb(pid),
|
|
"start_time": starttime,
|
|
"cgroup": _cgroup(pid),
|
|
}
|
|
return out
|
|
|
|
|
|
# --- clasificare -----------------------------------------------------------
|
|
|
|
def _known_pids(state: dict) -> set[int]:
|
|
"""Pid-urile pe care state.json le declara vii, cu verificare de PID reuse."""
|
|
known: set[int] = set()
|
|
threads = (state or {}).get("threads") or {}
|
|
if not isinstance(threads, dict):
|
|
return known
|
|
for entry in threads.values():
|
|
if not isinstance(entry, dict):
|
|
continue
|
|
pid = entry.get("pid")
|
|
if not isinstance(pid, int) or pid <= 0:
|
|
continue
|
|
expected = entry.get("pid_start_time")
|
|
if expected is not None:
|
|
st = _parse_stat(pid)
|
|
# Pid reciclat: alt proces poarta acum acelasi numar. Nu-l protejam,
|
|
# dar nici nu-l omoram automat — intra pe filtrele obisnuite.
|
|
if st is None or not _start_time_matches(st[1], expected):
|
|
continue
|
|
known.add(pid)
|
|
return known
|
|
|
|
|
|
def _descendants(roots: set[int], procs: dict[int, dict]) -> set[int]:
|
|
"""Toti descendentii pid-urilor date (turul care ruleaza acum e intocmai asta)."""
|
|
children: dict[int, list[int]] = {}
|
|
for pid, info in procs.items():
|
|
children.setdefault(info["ppid"], []).append(pid)
|
|
seen: set[int] = set()
|
|
stack = list(roots)
|
|
while stack:
|
|
pid = stack.pop()
|
|
for child in children.get(pid, ()):
|
|
if child not in seen:
|
|
seen.add(child)
|
|
stack.append(child)
|
|
return seen
|
|
|
|
|
|
def _ancestors_of_self(procs: dict[int, dict]) -> set[int]:
|
|
out: set[int] = set()
|
|
pid = os.getpid()
|
|
while pid and pid not in out:
|
|
out.add(pid)
|
|
info = procs.get(pid)
|
|
if not info:
|
|
break
|
|
pid = info["ppid"]
|
|
return out
|
|
|
|
|
|
def _is_claude(cmdline: str) -> bool:
|
|
for token in cmdline.split():
|
|
if token.rsplit("/", 1)[-1] in CLAUDE_BASENAMES:
|
|
return True
|
|
return any(m in cmdline for m in CLAUDE_PATH_MARKERS)
|
|
|
|
|
|
def _is_protected(cmdline: str) -> bool:
|
|
low = cmdline.lower()
|
|
return any(marker.lower() in low for marker in NEVER_KILL)
|
|
|
|
|
|
def _nearest_in(pid: int, multime: set[int], procs: dict[int, dict]) -> int | None:
|
|
"""Cel mai apropiat stramos al lui `pid` care se afla in `multime`.
|
|
|
|
Se opreste la pid 1: un proces reparentat la init nu mai are legatura reala
|
|
cu parintele lui original.
|
|
"""
|
|
info = procs.get(pid)
|
|
vazute = {pid}
|
|
while info:
|
|
parinte = info["ppid"]
|
|
if parinte <= 1 or parinte in vazute:
|
|
return None
|
|
if parinte in multime:
|
|
return parinte
|
|
vazute.add(parinte)
|
|
info = procs.get(parinte)
|
|
return None
|
|
|
|
|
|
def find_orphans(state: dict, min_age_s: int = 0) -> list[dict]:
|
|
"""Procese ramase in urma, care nu apar in state.json — ARBORI intregi.
|
|
|
|
DOMENIUL e cgroup-ul serviciului si numai el. Sesiunile Claude interactive ale
|
|
omului (tmux, ttyd, VS Code) traiesc in `tmux-spawn-<uuid>.scope`, nu in
|
|
`claude-discord.service`, deci nu apar niciodata aici — nici macar in dry-run.
|
|
|
|
Inauntrul cgroup-ului, e orfan orice proces care nu e protejat si nu se leaga de
|
|
state.json:
|
|
1. procese `claude` neinregistrate;
|
|
2. descendentii lor — serverele MCP (npm/sh/node) si orice altceva au pornit;
|
|
se gasesc prin arborele din /proc, desi nu se numesc `claude`;
|
|
3. copii deja reparentati la init, care si-au pierdut parintele si pe care
|
|
arborele singur nu-i mai poate atribui nimanui.
|
|
|
|
Intoarce o lista PLATA (contractul din INTERFACES.md), dar ordonata pe familii:
|
|
fiecare radacina, imediat urmata de copiii ei. Campurile suplimentare
|
|
`parent_pid`, `root_pid`, `depth` si `family_rss_mb` descriu ierarhia, pentru
|
|
raport si pentru ordinea de omorare.
|
|
"""
|
|
procs = scan_processes()
|
|
known = _known_pids(state or {})
|
|
protejate = set(known) | _descendants(known, procs) | _ancestors_of_self(procs)
|
|
marker = _service_cgroup_marker()
|
|
|
|
def in_serviciu(info: dict) -> bool:
|
|
"""Apartenenta la cgroup-ul serviciului. FAIL-CLOSED.
|
|
|
|
Cgroup necitibil (proces disparut intre listare si citire, /proc
|
|
restrictionat) inseamna NU. Mai bine ratam un orfan decat sa omoram
|
|
sesiunea cuiva.
|
|
"""
|
|
cgroup = info.get("cgroup") or ""
|
|
return bool(cgroup) and marker in cgroup
|
|
|
|
def eligibil(pid: int, info: dict) -> bool:
|
|
"""Filtrele care se aplica oricarui candidat, indiferent de familie."""
|
|
if not in_serviciu(info):
|
|
return False # conditie NECESARA, si pentru radacini
|
|
if pid in protejate or pid == os.getpid():
|
|
return False
|
|
cmdline = info["cmdline"]
|
|
if not cmdline or _is_protected(cmdline):
|
|
return False # bot.py, systemd, sshd, tmux, code-server...
|
|
return not _is_zombie(pid)
|
|
|
|
# Domeniul e cgroup-ul serviciului si atat. Inauntru, orice proces care nu e
|
|
# protejat si nu se leaga de state.json e un rest — indiferent cum se numeste.
|
|
# Asa intra si serverele MCP (npm/sh/node), care nu se numesc `claude`.
|
|
candidati = {pid for pid, info in procs.items() if eligibil(pid, info)}
|
|
if not candidati:
|
|
return []
|
|
|
|
# radacinile `claude`, doar pentru textul motivului
|
|
radacini = {pid for pid in candidati if _is_claude(procs[pid]["cmdline"])}
|
|
|
|
# ierarhia in interiorul multimii de candidati
|
|
parinti: dict[int, int | None] = {
|
|
pid: _nearest_in(pid, candidati, procs) for pid in candidati
|
|
}
|
|
|
|
def radacina_lui(pid: int) -> int:
|
|
vazute = {pid}
|
|
cur = pid
|
|
while True:
|
|
urmator = parinti.get(cur)
|
|
if urmator is None or urmator in vazute:
|
|
return cur
|
|
vazute.add(urmator)
|
|
cur = urmator
|
|
|
|
def adancime(pid: int) -> int:
|
|
d = 0
|
|
cur = pid
|
|
vazute = {pid}
|
|
while True:
|
|
urmator = parinti.get(cur)
|
|
if urmator is None or urmator in vazute:
|
|
return d
|
|
vazute.add(urmator)
|
|
cur = urmator
|
|
d += 1
|
|
|
|
orphans: list[dict] = []
|
|
for pid in candidati:
|
|
info = procs[pid]
|
|
rad = radacina_lui(pid)
|
|
parinte = parinti[pid]
|
|
if pid in radacini and parinte is None:
|
|
motiv = "proces claude neinregistrat in state.json"
|
|
elif parinte is not None:
|
|
motiv = f"copil al procesului orfan {rad}"
|
|
elif info["ppid"] <= 1:
|
|
motiv = "copil reparentat la init, ramas in cgroup-ul serviciului"
|
|
else:
|
|
motiv = "proces ramas in cgroup-ul serviciului, nelegat de state.json"
|
|
|
|
orphans.append({
|
|
"pid": pid,
|
|
"cmdline": info["cmdline"],
|
|
"age_s": info["age_s"],
|
|
"rss_mb": info["rss_mb"],
|
|
"ppid": info["ppid"],
|
|
"start_time": info["start_time"],
|
|
"reason": motiv,
|
|
"parent_pid": parinte,
|
|
"root_pid": rad,
|
|
"depth": adancime(pid),
|
|
})
|
|
|
|
# Varsta se judeca pe FAMILIE, dupa radacina: un server MCP pornit acum un
|
|
# minut sub un `claude` orfan de o ora tot rest e, si nu are sens sa taiem
|
|
# familia in doua.
|
|
if min_age_s > 0:
|
|
varsta_radacinii = {o["pid"]: o["age_s"] for o in orphans if o["depth"] == 0}
|
|
orphans = [o for o in orphans
|
|
if varsta_radacinii.get(o["root_pid"], o["age_s"]) >= min_age_s]
|
|
|
|
# RSS-ul intregii familii, pus pe fiecare membru: omul trebuie sa vada ca
|
|
# sterge 440 MB, nu 300.
|
|
familie_mb: dict[int, float] = {}
|
|
for o in orphans:
|
|
familie_mb[o["root_pid"]] = familie_mb.get(o["root_pid"], 0.0) + o["rss_mb"]
|
|
for o in orphans:
|
|
o["family_rss_mb"] = round(familie_mb.get(o["root_pid"], 0.0), 1)
|
|
|
|
# Familiile grele primele; in interiorul unei familii, radacina apoi copiii.
|
|
orphans.sort(key=lambda o: (
|
|
-familie_mb.get(o["root_pid"], 0.0), o["root_pid"],
|
|
o["depth"], -o["rss_mb"], o["pid"],
|
|
))
|
|
return orphans
|
|
|
|
|
|
# --- oprire ----------------------------------------------------------------
|
|
|
|
def _still_same_process(orphan: dict) -> bool:
|
|
"""Aparare impotriva PID reuse intre `find_orphans` si `kill_orphans`."""
|
|
pid = orphan.get("pid")
|
|
if not isinstance(pid, int) or pid <= 0:
|
|
return False
|
|
st = _parse_stat(pid)
|
|
if st is None or _is_zombie(pid):
|
|
return False
|
|
expected = orphan.get("start_time")
|
|
if expected is None:
|
|
return True
|
|
return _start_time_matches(st[1], expected)
|
|
|
|
|
|
def _kill_order(orphans: list[dict]) -> list[dict]:
|
|
"""Copiii inaintea parintilor.
|
|
|
|
Daca omori intai parintele, copiii lui se reparenteaza la init si scapa din
|
|
aceeasi trecere — fix scurgerea pe care comanda ar trebui s-o opreasca.
|
|
Adancimea vine din `find_orphans`; pentru intrari construite de mana se
|
|
recalculeaza din /proc, ca ordonarea sa fie corecta si atunci.
|
|
"""
|
|
lista = list(orphans or [])
|
|
in_set = {o.get("pid") for o in lista if isinstance(o.get("pid"), int)}
|
|
|
|
def adancime(orphan: dict) -> int:
|
|
d = orphan.get("depth")
|
|
if isinstance(d, int):
|
|
return d
|
|
pid = orphan.get("pid")
|
|
if not isinstance(pid, int):
|
|
return 0
|
|
d = 0
|
|
vazute = {pid}
|
|
cur = pid
|
|
while True:
|
|
st = _parse_stat(cur)
|
|
if st is None:
|
|
return d
|
|
parinte = st[0]
|
|
if parinte <= 1 or parinte in vazute:
|
|
return d
|
|
if parinte in in_set:
|
|
d += 1
|
|
vazute.add(parinte)
|
|
cur = parinte
|
|
|
|
# stabil: la aceeasi adancime pastram ordinea primita
|
|
return sorted(lista, key=lambda o: -adancime(o))
|
|
|
|
|
|
def kill_orphans(orphans: list[dict], dry_run: bool = True, grace_s: float = GRACE_S) -> list[dict]:
|
|
"""Opreste orfanii, copiii inaintea parintilor. Implicit NU omoara nimic.
|
|
|
|
SIGTERM, apoi SIGKILL dupa `grace_s` daca procesul inca traieste.
|
|
Intoarce cate un rezultat per intrare: {"pid", "cmdline", "action", "detail"}.
|
|
action: "dry-run" | "terminated" | "killed" | "gone" | "skipped" | "error"
|
|
"""
|
|
results: list[dict] = []
|
|
for orphan in _kill_order(orphans):
|
|
pid = orphan.get("pid")
|
|
entry = {"pid": pid, "cmdline": orphan.get("cmdline", ""), "action": "", "detail": ""}
|
|
|
|
if not isinstance(pid, int) or not _still_same_process(orphan):
|
|
entry["action"] = "gone"
|
|
entry["detail"] = "procesul nu mai exista sau pid-ul a fost reciclat"
|
|
results.append(entry)
|
|
continue
|
|
|
|
if pid == os.getpid():
|
|
entry["action"] = "skipped"
|
|
entry["detail"] = "e chiar procesul curent"
|
|
results.append(entry)
|
|
continue
|
|
|
|
# Plasa de siguranta, independenta de cine a construit lista: chiar daca
|
|
# cineva ne pasaza botul sau un serviciu de sesiune, nu-l atingem.
|
|
# Se verifica AMBELE: linia de comanda vie din /proc (adevarul de acum) si
|
|
# cea din intrare (ce credea apelantul). Oricare dintre ele protejata = refuz.
|
|
if _is_protected(_cmdline(pid)) or _is_protected(str(orphan.get("cmdline", ""))):
|
|
entry["action"] = "skipped"
|
|
entry["detail"] = "proces protejat (bot.py / infrastructura sesiunii)"
|
|
results.append(entry)
|
|
continue
|
|
|
|
if dry_run:
|
|
entry["action"] = "dry-run"
|
|
entry["detail"] = (
|
|
f"s-ar trimite SIGTERM (varsta {orphan.get('age_s')}s, "
|
|
f"{orphan.get('rss_mb')}MB)"
|
|
)
|
|
results.append(entry)
|
|
continue
|
|
|
|
try:
|
|
os.kill(pid, signal.SIGTERM)
|
|
except ProcessLookupError:
|
|
entry["action"] = "gone"
|
|
entry["detail"] = "disparut inainte de SIGTERM"
|
|
results.append(entry)
|
|
continue
|
|
except Exception as exc:
|
|
entry["action"] = "error"
|
|
entry["detail"] = f"SIGTERM a esuat: {exc}"
|
|
results.append(entry)
|
|
continue
|
|
|
|
deadline = time.time() + grace_s
|
|
while time.time() < deadline:
|
|
if not _still_same_process(orphan):
|
|
break
|
|
time.sleep(0.1)
|
|
|
|
if not _still_same_process(orphan):
|
|
entry["action"] = "terminated"
|
|
entry["detail"] = "oprit cu SIGTERM"
|
|
results.append(entry)
|
|
continue
|
|
|
|
try:
|
|
os.kill(pid, signal.SIGKILL)
|
|
entry["action"] = "killed"
|
|
entry["detail"] = f"nu a raspuns la SIGTERM in {grace_s:.0f}s, SIGKILL"
|
|
except ProcessLookupError:
|
|
entry["action"] = "terminated"
|
|
entry["detail"] = "oprit cu SIGTERM"
|
|
except Exception as exc:
|
|
entry["action"] = "error"
|
|
entry["detail"] = f"SIGKILL a esuat: {exc}"
|
|
results.append(entry)
|
|
|
|
return results
|
|
|
|
|
|
# --- randare pentru Discord ------------------------------------------------
|
|
|
|
def format_report(orphans: list[dict], results: list[dict] | None = None) -> str:
|
|
"""Raport ierarhic pentru raspunsul comenzii `/cleanup` (sub 2000 caractere).
|
|
|
|
Radacina pe prima linie cu totalul familiei, copiii indentati sub ea. Fara
|
|
ierarhie, un `claude` de 300 MB pare tot ce se sterge, cand de fapt pleaca
|
|
440 MB cu tot cu serverele MCP.
|
|
"""
|
|
if not orphans:
|
|
return "Niciun proces orfan. Nimic de curatat."
|
|
|
|
total_mb = sum(o.get("rss_mb", 0) or 0 for o in orphans)
|
|
lines = [f"**{len(orphans)} procese orfane** (~{total_mb:.0f} MB RSS in total)", "```"]
|
|
by_pid = {r.get("pid"): r for r in (results or [])}
|
|
|
|
randuri = 0
|
|
lungime = sum(len(l) + 1 for l in lines)
|
|
for orphan in orphans:
|
|
if randuri >= MAX_REPORT_ROWS:
|
|
break
|
|
adancime = orphan.get("depth") or 0
|
|
pid = orphan.get("pid")
|
|
rss = orphan.get("rss_mb", 0) or 0
|
|
|
|
if adancime == 0:
|
|
familie = orphan.get("family_rss_mb")
|
|
cmd = str(orphan.get("cmdline", ""))[:64]
|
|
linie = (f"pid={pid:<7} {rss:>7.1f} MB {orphan.get('age_s', 0):>7}s {cmd}")
|
|
# totalul familiei se arata doar cand chiar are copii
|
|
if familie is not None and round(familie, 1) != round(rss, 1):
|
|
linie += f"\n familie: {familie:.1f} MB in total"
|
|
else:
|
|
indent = " " * adancime
|
|
cmd = str(orphan.get("cmdline", ""))[:60 - len(indent)]
|
|
linie = f"{indent}`-- pid={pid:<7} {rss:>7.1f} MB {cmd}"
|
|
|
|
rezultat = by_pid.get(pid)
|
|
if rezultat:
|
|
linie += f"\n{' ' * adancime} -> {rezultat.get('action')}: {rezultat.get('detail')}"
|
|
# Bugetul de caractere, nu doar numarul de randuri: cu rezultatele de
|
|
# omorare atasate un rand poate fi de trei ori mai lung.
|
|
if randuri and lungime + len(linie) + 1 > MAX_REPORT_CHARS:
|
|
break
|
|
lines.append(linie)
|
|
lungime += len(linie) + 1
|
|
randuri += 1
|
|
|
|
if len(orphans) > randuri:
|
|
lines.append(f"... si inca {len(orphans) - randuri}")
|
|
lines.append("```")
|
|
if not results:
|
|
lines.append("Rulare seaca. `/cleanup force:True` le opreste efectiv "
|
|
"(copiii inaintea parintilor).")
|
|
return "\n".join(lines)
|
|
|
|
|
|
def _load_state() -> dict:
|
|
"""Citeste state.json pentru rularea din linia de comanda. Tolerant la lipsa."""
|
|
if _config is not None and getattr(_config, "STATE_FILE", None):
|
|
path = pathlib.Path(_config.STATE_FILE)
|
|
else:
|
|
path = pathlib.Path(os.path.expanduser("~/.claude-discord/state.json"))
|
|
try:
|
|
import json
|
|
with open(path, "r", encoding="utf-8") as fh:
|
|
data = json.load(fh)
|
|
return data if isinstance(data, dict) else {}
|
|
except Exception:
|
|
return {}
|
|
|
|
|
|
if __name__ == "__main__":
|
|
# Uz manual: python3 cleanup.py -> doar raporteaza
|
|
# python3 cleanup.py --force -> opreste efectiv
|
|
force = "--force" in sys.argv[1:]
|
|
found = find_orphans(_load_state())
|
|
res = kill_orphans(found, dry_run=not force)
|
|
print(format_report(found, res))
|