feat(steering): mesaje mid-tur + /stop pe turul în zbor

Un al doilea mesaj trimis cât Claude încă lucra aștepta până se termina
turul 1 — corecția „stai, nu în master" ajungea după ce greșeala era gata.
Verificat în producție înainte de commit: mesajul 2 stătea 25s blocat în
lock, apoi pornea ca tur separat.

Acum canalele de chat pot ține un proces `claude` viu per canal, cu stdin
deschis, și al doilea mesaj intră în ACELAȘI tur.

- `src/claude_runner.py` — ClaudeProcess (steering, respawn cu --resume,
  drenare stderr, respawn la comutarea OpenRouter) + RunnerRegistry
  (max_live, reaper pe inactivitate, stop_all la shutdown)
- `src/stream_json.py` — parser stream-json partajat cu `_run_claude`;
  pur, nu aruncă niciodată pe is_error (PlanningSession retrimite pe
  error_max_turns și depinde de asta)
- `src/sentinels.py` — un singur loc pentru __AUDIO__/__STEERED__, în loc
  de 4 verificări copiate; repară și bug-ul preexistent prin care
  WhatsApp posta literal `__AUDIO__:/cale`
- dispecer în `send_message`: lock.acquire(blocking=False) — eșecul de a
  lua lock-ul ESTE „rulează un tur", ceea ce elimină flagul inflight din
  decizie și cursa TOCTOU odată cu el
- `/stop` oprește turul, nu sesiunea — active.json rămâne valid
- rate limit prin proces persistent vine ca result.is_error, nu ca exit
  code; convertit înapoi în același RuntimeError, altfel fallback-ul
  local nu s-ar mai declanșa niciodată, în tăcere

Steering-ul nu face niciodată cross-adapter (un mesaj text nu intră
într-un tur voice: împart același channel_id). Mesajele steered dintr-un
tur care pică sunt re-livrate, nu pierdute.

Testat live cu CLI-ul real: corecție la secunda 10 dintr-un tur de 24s,
un singur result, num_turns=2. Notă: mesajele steered sunt împachetate în
[EXTERNAL CONTENT], deci o corecție formulată ca override agresiv poate
fi refuzată ca prompt injection — pentru oprire folosește /stop.

Suită: 1199 passed, 12 failed (toate pre-existente pe HEAD curat).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiJGsZVSEGjRHZEJiXaxCC
This commit is contained in:
2026-09-02 11:05:58 +00:00
parent 2d8ee9b581
commit 747afbaf9d
23 changed files with 4183 additions and 95 deletions

View File

@@ -24,7 +24,10 @@ from src.claude_session import (
rate_limit_detail as _rate_limit_detail,
RATE_LIMIT_RE as _RATE_LIMIT_RE,
VALID_MODELS,
stop_turn,
pop_pending_steers as _pop_pending_steers,
)
from src.sentinels import is_steered as _is_steered
from src.jsonlock import read_locked, write_locked
from src.planning_orchestrator import PlanningOrchestrator
from src.planning_session import (
@@ -566,6 +569,11 @@ def route_message(
if text.lower() == "/status":
return _status(channel_id), True
if text.lower() == "/stop":
if stop_turn(channel_id):
return "⏹ Oprit.", True
return "Nu rulează nimic pe canalul ăsta.", True
if text.lower().startswith("/model"):
return _model_command(channel_id, text), True
@@ -602,15 +610,29 @@ def route_message(
try:
response = send_message(
session_key, claude_text, model=model, on_text=on_text,
voice_mode=voice_mode,
voice_mode=voice_mode, adapter_name=adapter_name,
)
if _is_steered(response):
# Same pattern as the existing __AUDIO__: sentinel — no
# _set_last_response, the adapter reacts instead of replying.
return response, False
_set_last_response(channel_id, response)
return response, False
except Exception as e:
log.error("Claude error for channel %s: %s", channel_id, e)
# C3: texts steered into this same turn before it failed — their own
# request threads already returned __STEERED__ and are gone, so the
# only way left to answer them is `on_text`, the same real-time
# channel already used for intermediate assistant text.
pending_steers = _pop_pending_steers(channel_id)
if _is_rate_limit_error(e):
log.warning("Rate limit detected for channel %s — trying local fallback", channel_id)
fallback = _local_fallback_reply(text, channel_id=channel_id)
for steered_text in pending_steers:
_redeliver_steered_reply(
steered_text, channel_id, on_text,
_local_fallback_reply(steered_text, channel_id=channel_id),
)
if fallback is not None:
_set_last_response(channel_id, fallback)
return fallback, False
@@ -622,9 +644,37 @@ def route_message(
"⚠️ Claude e la limită, iar modelul local nu a răspuns.\n"
f"{_rate_limit_detail(e)}"
), False
for steered_text in pending_steers:
_redeliver_steered_reply(steered_text, channel_id, on_text, f"Error: {e}")
return f"Error: {e}", False
def _redeliver_steered_reply(
steered_text: str,
channel_id: str,
on_text: Callable[[str], None] | None,
reply: str | None,
) -> None:
"""C3 — a message steered into a turn that then failed must still get
an answer. Its own request thread already returned `__STEERED__` and is
gone, so the only way left to reach the user is `on_text` (the same
real-time channel adapters already use for intermediate assistant
text). Logs instead of dropping silently when there's no `on_text` to
push through (T12 — "it ignored my message" must stay diagnosable)."""
if reply is None:
reply = "⚠️ Claude e la limită — mesajul tău steered nu a primit răspuns."
if on_text is None:
log.warning(
"channel=%s: steered message lost — no on_text to redeliver it: %r",
channel_id, steered_text[:80],
)
return
try:
on_text(reply)
except Exception:
log.exception("channel=%s: failed to redeliver steered reply via on_text", channel_id)
def _status(channel_id: str) -> str:
"""Build status message for a channel."""
session = get_active_session(channel_id)