docs(cluster): procedura de oprire planificata + corectii stare HA reala

Runbook nou pentru oprirea controlata a celor 3 noduri (lucrari electrice,
mutare rack, mentenanta UPS) fara ca HA sa relocheze resursele si fara ca
watchdog-ul sa reseteze hard nodul ramas fara quorum.

Doua capcane documentate:
- /etc/pve/datacenter.cfg nu exista => shutdown_policy implicit `conditional`
  => poweroff pe nod declanseaza FAILOVER, nu freeze
- watchdog-ul face reset hard dupa ~60s pe nodul care pierde quorumul cu
  servicii HA inca active

Corectii in failover/README.md, verificate pe clusterul live:
- VM 201 ESTE in HA (grup ha-prefer-pvemini), docul spunea ca nu e
- CT 108 e in ha-prefer-pvemini (pvemini:100, pve1:50, pveelite:10),
  nu in ha-group-main cu pveelite:50/pve1:33
- replicare CT 108 si VM 201 la */15 min, nu */5 => RPO real 15 min

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RRyaDj39hPQ89SZS6URRpS
This commit is contained in:
Marius
2026-08-27 12:28:54 +03:00
parent 3ba36b6119
commit ff7e7da6d1
4 changed files with 316 additions and 6 deletions

View File

@@ -39,6 +39,7 @@ input/ # Oracle DMP files for import
## Key Documentation Entry Points
- **Infrastructure overview**: `proxmox/README.md`
- **Oprire planificată a clusterului (lucrări electrice) — fără failover HA / fence**: `proxmox/cluster/docs/oprire-planificata-cluster.md`
- **SQL migration guidelines**: `system_instructions/system_prompt.md` (always read before generating migration SQL)
- **Oracle database setup**: `proxmox/lxc108-oracle/README.md`
- **Migration orchestration**: `proxmox/lxc108-oracle/migration/00-MASTER-MIGRATION.sh`