Testul DR din 2026-08-08 a raportat "Restore failed" dupa 11 secunde, fara niciun log RMAN. Cauza nu a fost restore-ul: KB5101001 fusese descarcat in timpul testului din 2026-08-01 (singurul moment in care VM 109 e pornit), a ramas staged dupa qm stop si s-a finalizat la boot-ul testului urmator. Cronologie din Event Log-ul guest-ului: 06:00:58 RestartManager 10010 - nu poate reporni powershell.exe (restore-ul) 06:01:04 SCM 7034 - OpenSSH SSH Server terminat neasteptat 06:01:06 pveelite: client_loop: send disconnect: Broken pipe -> FAILED 06:01:38 VM-ul se reboteaza singur Fereastra testului (Sambata 06:00) era in afara Active Hours (08:00-17:00), deci pentru Windows era fereastra de mentenanta valida - iar VM 109 fiind pornit doar in timpul testului, aceea era singura fereastra posibila. Agravant: sshd nu avea acsiuni de recovery (RESET_PERIOD 0), deci dupa ce a murit a ramas mort si au esuat si colectarea logului si shutdown-ul gratios. Masuri: - NoAutoUpdate=1 + AUOptions=2 pe VM 109 (aplicat direct in registry) - actiuni de recovery pentru sshd: restart la 5s/10s/30s, reset=86400 - guard "STEP 3b: Windows servicing" inainte de restore (check_servicing.ps1): asteapta idle 300s, consuma controlat un reboot in asteptare, altfel abandoneaza cu "ABORTED - Windows servicing" in loc de un "Restore failed" inselator. Fail-open daca checkul lipseste - nu are voie sa pice testul. - fereastra lunara de patching (vm109-patch-window.sh + install_updates.ps1), prima duminica 03:00, cu re-armare NoAutoUpdate=1 indiferent de rezultat Adaugat si .gitattributes: cu core.autocrlf=true scripturile .sh ajungeau in working tree cu CRLF, iar ele se deployeaza prin scp direct pe Proxmox. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BhQBTegE4PiMPPaapLHjkc
VM 109 - Oracle DR System (Windows Standby)
Director Proxmox: proxmox/vm109-windows-dr/
VMID: 109
Rol: Disaster Recovery pentru Oracle Database (backup RMAN de pe server Windows extern)
⚠️ Important — Topologie după 2026-04-25
VM 109 trăiește pe pveelite (10.0.20.202), co-located cu storage-ul NFS Oracle backups. Configurația post-incident 04-20:
- VM 109 în HA, grup
ha-prefer-pveelite(pveelite=100, pvemini=50, pve1=10),state=stopped,nofailback=1— HA face failover dacă pveelite cade dar nu repornește VM 109 automat (rămâne stopped, scriptul DR îl pornește săptămânal). - Apărări împotriva incident 04-20:
trap cleanup_vm EXITîn scriptul DR (commit8a0c557) cu guardDR_VM_STARTED_BY_US(commit2e8cd9c) — oprește VM 109 doar dacă scriptul l-a pornit.vm109-watchdog.shcron pe ambele pveelite + pvemini (cluster-aware) — oprește forțat VM 109 dacă rulează > 60 min în afara ferestrei test (Sâmbătă 05:55-07:30). Debug exempt:touch /var/run/vm109-debug.flag.- Pre-flight check în DR script: refuză
qm start 109dacă cluster degraded sau memorie disponibilă < (VM 109 mem + 1 GB margin). max_restart=3, max_relocate=2pe toate serviciile HA — cap pe restart loops la OOM.
Verificare status:
ssh root@10.0.20.201 "ha-manager status | grep -E '109|201|108'"
ssh root@10.0.20.202 "qm status 109" # trebuie stopped între teste
🔄 Storage Failover (pveelite → pvemini)
/mnt/pve/oracle-backups e dataset ZFS replicat pveelite → pvemini la 15 min (zfs-replicate-oracle-backups.sh) + mirror nightly pe pve1 backup-ssd (nightly-backup-mirror.sh). La pveelite down:
- Email automat din
pveelite-down-alert.sh(cron pe pvemini, prag 5 min) cu instrucțiuni de failover copy-paste. - Operator rulează pe pvemini:
/opt/scripts/failover-dr-to-pvemini.sh— promote ZFS readonly → off, configurează NFS export, patch primary Oracle scheduled task IP via SSH. - Când pveelite revine:
/opt/scripts/failback-dr-to-pveelite.sh— invers, cu zfs send incremental + restaurare config.
Script-urile refuză să ruleze dacă cealaltă parte e accesibilă (anti-split-brain).
🩹 Windows Update pe VM 109 — incident 2026-08-08
Simptom: testul DR din 2026-08-08 a raportat FAILED, Database Restore: Restore failed după 11 secunde, Tables restored: 0, iar raportul nu conținea niciun log RMAN („No restore logs or RMAN scripts found"). Testul din 2026-08-01 trecuse în 18 minute pe exact același cod (scripturile nu s-au modificat din 2026-04-25).
Cauza reală: RMAN nici nu a apucat să pornească. Windows Update a rebootat VM-ul în mijlocul testului.
Cronologia din Event Log-ul guest-ului:
| Ora | Eveniment |
|---|---|
| 06:00:26 | boot VM 109 |
| 06:00:53 | scriptul DR confirmă SSH + PowerShell |
| 06:00:55 | STEP 4 pornește rman_restore_from_zero.ps1 |
| 06:00:58 | RestartManager 10010 — nu poate reporni powershell.exe (pid 5964) = procesul de restore |
| 06:00:59 | Winlogon 6004 — TrustedInstaller a eșuat o notificare critică |
| 06:01:02 | SCM 7023 — Update Orchestrator Service terminat |
| 06:01:04 | SCM 7034 — OpenSSH SSH Server terminat neașteptat |
| 06:01:06 | pveelite vede client_loop: send disconnect: Broken pipe → restore marcat FAILED |
| 06:01:38 | VM-ul se reboteaza singur (al doilea eveniment Kernel-General 12) |
| 06:02:20 | scriptul renunță la SSH și forțează qm stop 109 |
Update-ul vinovat: KB5101001, InstalledOn 8/8/2026.
De ce s-a întâmplat exact atunci — problemă structurală, nu ghinion:
- VM 109 e pornit doar în timpul testului DR. Deci fereastra testului era singura fereastră în care Windows Update putea rula vreodată.
- Update-ul fusese descărcat în timpul testului precedent — evenimentele
WindowsUpdateClient id=44sunt la 08-01 06:13–06:17, adică în interiorul rulării din 1 august. A rămas staged dupăqm stopși s-a finalizat la boot-ul următor = testul următor. - Active Hours erau 08:00–17:00, deci 06:00 era pentru Windows o fereastră de mentenanță validă.
NoAutoRebootWithLoggedOnUsers=1nu ajută: o sesiune SSH nu e logon interactiv.
Factor agravant: sc qfailure sshd nu avea nicio acțiune de recovery (RESET_PERIOD: 0). După ce sshd a murit a rămas mort, deci și STEP 6 (colectare log) și STEP 7 (shutdown grațios) au eșuat — de aici raportul fără loguri.
Ce NU a fost: host-ul e curat — fără OOM, fără erori de storage, cluster quorate, replicarea ZFS a rulat normal la 06:00. Backup-urile nu au fost implicate în niciun fel.
Măsuri aplicate (2026-08-08)
| # | Măsură | Unde |
|---|---|---|
| 1 | NoAutoUpdate=1 + AUOptions=2 — Windows Update nu mai pornește singur |
registry VM 109, HKLM\SOFTWARE\Policies\Microsoft\Windows\WindowsUpdate\AU |
| 2 | Acțiuni de recovery pentru sshd: restart la 5s / 10s / 30s, reset=86400 |
serviciu sshd pe VM 109 |
| 3 | Guard „STEP 3b: Windows servicing" înainte de restore | weekly-dr-test-proxmox.sh + check_servicing.ps1 |
| 4 | Fereastră lunară dedicată de patching | vm109-patch-window.sh + install_updates.ps1 |
3 — guard-ul de servicing. Între verificarea NFS și restore, scriptul rulează check_servicing.ps1 pe guest, care întoarce STATE=IDLE sau STATE=BUSY REBOOT_PENDING=... REASONS=... (CBS RebootPending, WU RebootRequired, PendingFileRename, TiWorker, TrustedInstaller, UsoSvc). Comportament:
- așteaptă până la 300s ca servicing-ul să devină idle;
- dacă e reboot în așteptare, îl consumă controlat (un singur reboot) și reia verificarea;
- dacă tot nu e curat, abandonează înainte de restore cu
ABORTED - Windows servicingîn loc de un „Restore failed" înșelător, care arată ca o problemă de backup; - dacă
check_servicing.ps1lipsește sau nu răspunde, guard-ul se auto-dezactivează cu un warning — nu are voie să pice testul din cauza propriei absențe.
4 — fereastra de patching. Prima duminică din lună, 03:00, pe nodul care găzduiește VM 109:
0 3 1-7 * * /opt/scripts/vm109-patch-window.sh > /dev/null 2>&1 # garda „duminică" e în script
Scriptul pornește VM 109 (cu vm109-debug.flag setat, fiindcă patching-ul depășește limita de 60 min a watchdog-ului), permite temporar Windows Update (NoAutoUpdate=0), aplică update-urile prin API-ul COM Microsoft.Update.Session, tolerează până la 3 cicluri de reboot, re-armează NoAutoUpdate=1 indiferent de rezultat, verifică starea finală și oprește VM-ul. Trimite mail doar când rezultatul nu e curat.
Rulare manuală (ignoră garda de calendar):
ssh root@10.0.20.202 "/opt/scripts/vm109-patch-window.sh --now"
tail -f /var/log/oracle-dr/patch-window.log
Verificare stare Windows Update pe VM 109:
ssh root@10.0.20.202 "ssh -p 22122 romfast@10.0.20.37 \
'powershell -ExecutionPolicy Bypass -File D:\\oracle\\scripts\\check_servicing.ps1'"
# Așteptat între ferestre: STATE=IDLE
Zgomot preexistent, fără legătură cu incidentul: la fiecare boot apar VSS 8213 + SCM 7023 pentru OracleVssWriterROA („General access denied"). Serviciul VSS writer al Oracle nu are drepturi suficiente. Nu afectează restore-ul RMAN (care nu folosește VSS) — de rezolvat separat.
🛡️ Oracle DR System - Complete Architecture
📊 System Overview
┌─────────────────────────────────────────────────────────────────┐
│ PRODUCTION ENVIRONMENT │
├─────────────────────────────────────────────────────────────────┤
│ PRIMARY SERVER (10.0.20.36) │
│ Windows Server + Oracle 19c │
│ ┌──────────────────────────────┐ │
│ │ Database: ROA │ │
│ │ Size: ~80 GB │ │
│ │ Tables: 42,625 │ │
│ └──────────────────────────────┘ │
│ │ │
│ ▼ Backups (Daily) │
│ ┌──────────────────────────────┐ │
│ │ 02:30 - FULL backup (6-7 GB) │ │
│ │ 13:00 - CUMULATIVE (200 MB) │ │
│ │ 18:00 - CUMULATIVE (300 MB) │ │
│ └──────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
│
│ SSH Transfer (Port 22)
▼
┌─────────────────────────────────────────────────────────────────┐
│ DR ENVIRONMENT │
├─────────────────────────────────────────────────────────────────┤
│ PROXMOX HOST (10.0.20.202 - pveelite) │
│ ┌──────────────────────────────┐ │
│ │ Backup Storage (NFS Server) │◄─────── Monitoring Scripts │
│ │ /mnt/pve/oracle-backups/ │ /opt/scripts/ │
│ │ └── ROA/autobackup/ │ │
│ └──────────────────────────────┘ │
│ │ │
│ │ NFS Mount (F:\) │
│ ▼ │
│ ┌──────────────────────────────┐ │
│ │ DR VM 109 (10.0.20.37) │ │
│ │ Windows Server + Oracle 19c │ │
│ │ Status: OFF (normally) │ │
│ │ Starts for: Tests or Disaster │ │
│ └──────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
🎯 Quick Actions
⚡ Emergency DR Activation (Production Down!)
# 1. Start DR VM
ssh root@10.0.20.202 "qm start 109"
# 2. Connect to VM (wait 3 min for boot)
ssh -p 22122 romfast@10.0.20.37
# 3. Run restore (takes ~10-15 minutes)
D:\oracle\scripts\rman_restore_from_zero.cmd
# 4. Database is now RUNNING - Update app connections to 10.0.20.37
🔄 Failback DR → PRIMARY (when production is repaired)
Procedura inversă, pentru când serverul de producție a fost reparat sau reinstalat și
trebuie mutată producția înapoi pe 10.0.20.36:
⚠️ Pe PRIMARY instalează Oracle 19c (NU 21c) pentru failback acut. Backup-urile sunt 19.3. 21c poate restore tehnic, dar cere upgrade-of-dictionary suplimentar (~30-60 min în plus) — risc inutil în fereastra de criză. Migrarea la 21c se face separat după failback. Detalii în
FAILBACK_PROCEDURE.md.
➡️ Vezi docs/FAILBACK_PROCEDURE.md — pași end-to-end:
- Backup final pe DR (cu DB în read-only / restricted)
- Restore pe PRIMARY nou cu
scripts/rman_restore_to_primary.ps1 - Switch connection strings + reactivare scheduled tasks RMAN
- Stop VM 109, revenire la state normal
🧪 Weekly Test (Every Saturday)
# Automatic at 06:00 via cron, or manual:
ssh root@10.0.20.202 "/opt/scripts/weekly-dr-test-proxmox.sh"
# What it does:
# ✓ Starts VM → Restores DB → Tests → Cleanup → Shutdown
# ✓ Sends email report with results
📊 Check Backup Health
# Manual check (runs daily at 09:00 automatically)
ssh root@10.0.20.202 "/opt/scripts/oracle-backup-monitor-proxmox.sh"
# Output:
# Status: OK
# FULL backup age: 11 hours ✓
# CUMULATIVE backup age: 2 hours ✓
# Disk usage: 45% ✓
🗂️ Component Locations
📁 PRIMARY Server (10.0.20.36)
D:\rman_backup\
├── rman_backup_full.txt # RMAN script for FULL backup
├── rman_backup_incremental.txt # RMAN script for CUMULATIVE
└── transfer_backups.ps1 # UNIFIED: Transfer ALL backups to Proxmox
Scheduled Tasks:
├── 02:30 - Oracle RMAN Full Backup
├── 03:00 - Transfer backups to DR (transfer_backups.ps1)
├── 13:00 - Oracle RMAN Cumulative Backup
├── 14:45 - Transfer backups to DR (transfer_backups.ps1)
└── 18:00 - Oracle RMAN Cumulative Backup
📁 PROXMOX Host (10.0.20.202)
/opt/scripts/
├── oracle-backup-monitor-proxmox.sh # Daily backup monitoring
├── weekly-dr-test-proxmox.sh # Weekly DR test
├── vm109-patch-window.sh # Monthly Windows patching (incident 08-08)
├── vm109-watchdog.sh # Force-stop VM 109 outside test window
└── PROXMOX_NOTIFICATIONS_README.md # Documentation
/mnt/pve/oracle-backups/ROA/autobackup/
├── FULL_20251010_023001.BKP # Latest FULL backup
├── INCR_20251010_130001.BKP # CUMULATIVE 13:00
└── INCR_20251010_180001.BKP # CUMULATIVE 18:00
Cron Jobs:
0 9 * * * /opt/scripts/oracle-backup-monitor-proxmox.sh
0 6 * * 6 /opt/scripts/weekly-dr-test-proxmox.sh
0 3 1-7 * * /opt/scripts/vm109-patch-window.sh # prima duminică (gardă în script)
* * * * * /opt/scripts/vm109-watchdog.sh
📁 DR VM 109 (10.0.20.37) - When Running
D:\oracle\scripts\
├── rman_restore_from_zero.cmd # Main restore script ⭐
├── cleanup_database.cmd # Cleanup after test
├── check_servicing.ps1 # Servicing guard (incident 08-08)
├── install_updates.ps1 # Used only by vm109-patch-window.sh
└── mount-nfs.bat # Mount F:\ at startup
F:\ (NFS mount from Proxmox)
└── ROA\autobackup\ # All backup files
🔄 How It Works
Backup Flow (Daily)
PRIMARY PROXMOX
│ │
├─02:30─FULL─Backup─────────────►
│ (6-7 GB) │
├─03:00─Transfer ALL────────────► Skip duplicates
│ (transfer_backups.ps1) │
│ │
├─13:00─CUMULATIVE──────────────►
│ (200 MB) │
├─14:45─Transfer ALL────────────► Skip duplicates
│ (transfer_backups.ps1) │ (only new files)
│ │
└─18:00─CUMULATIVE──────────────►
(300 MB) Storage
│
┌──────────┐
│ Monitor │ 09:00 Daily
│ Check Age│ Alert if old
└──────────┘
Restore Process
Start VM → Mount F:\ → Copy Backups → RMAN Restore → Database OPEN
2min Auto 2min 8min Ready!
Total Time: ~15 minutes
🔧 Manual Operations
Test Individual Components
# 1. Test backup transfer (on PRIMARY)
powershell -ExecutionPolicy Bypass -File "D:\rman_backup\transfer_backups.ps1"
# 2. Test NFS mount (on VM 109)
mount -o rw,nolock,mtype=hard,timeout=60 10.0.20.202:/mnt/pve/oracle-backups F:
dir F:\ROA\autobackup
# 3. Test notification system
ssh root@10.0.20.202 "touch -d '2 days ago' /mnt/pve/oracle-backups/ROA/autobackup/*FULL*.BKP"
ssh root@10.0.20.202 "/opt/scripts/oracle-backup-monitor-proxmox.sh"
# Should send WARNING notification
# 4. Test database restore (on VM 109)
D:\oracle\scripts\rman_restore_from_zero.cmd
Force Actions
# Force backup now (on PRIMARY)
rman cmdfile=D:\rman_backup\rman_backup_incremental.txt
# Force cleanup VM (on VM 109)
D:\oracle\scripts\cleanup_database.cmd
# Force VM shutdown
ssh root@10.0.20.202 "qm stop 109"
🐛 Troubleshooting
⚡ Simptome frecvente — unde te uiți întâi
| Simptom în raportul DR | Cauză probabilă |
|---|---|
Restore failed în < 60s, fără log RMAN, Broken pipe în log-ul bash |
Nu e problemă de backup — sesiunea SSH a murit. Vezi Windows Update pe VM 109. Verifică SCM 7034 sshd în Event Log-ul guest-ului. |
ABORTED - Windows servicing |
Guard-ul STEP 3b a oprit testul intenționat: Windows Update era activ. Backup-urile nu sunt implicate. Rulează fereastra de patching manual. |
Restore failed după 10+ minute, cu log RMAN |
Problemă reală de restore — continuă cu secțiunile de mai jos. |
🔍 Debugging Restore Tests
Check Backup Files on Proxmox (10.0.20.202)
# 1. List all backup files with size and date
ssh root@10.0.20.202 "ls -lht /mnt/pve/oracle-backups/ROA/autobackup/*.BKP"
# 2. Count backup files
ssh root@10.0.20.202 "ls /mnt/pve/oracle-backups/ROA/autobackup/*.BKP | wc -l"
# 3. Check latest backups (last 24 hours)
ssh root@10.0.20.202 "find /mnt/pve/oracle-backups/ROA/autobackup -name '*.BKP' -mtime -1 -ls"
# 4. Show backup files grouped by type (with new naming convention)
ssh root@10.0.20.202 "ls -lh /mnt/pve/oracle-backups/ROA/autobackup/ | grep -E '(L0_|L1_|ARC_|SPFILE_|CF_|O1_MF)'"
# 5. Check disk space usage
ssh root@10.0.20.202 "df -h /mnt/pve/oracle-backups"
ssh root@10.0.20.202 "du -sh /mnt/pve/oracle-backups/ROA/autobackup/"
# 6. Verify newest backup timestamp
ssh root@10.0.20.202 "stat /mnt/pve/oracle-backups/ROA/autobackup/L0_*.BKP 2>/dev/null | grep Modify || echo 'No L0 backups with new naming'"
Verify Backup Files on DR VM (when running)
# 1. Check NFS mount is accessible
Test-Path F:\ROA\autobackup
# 2. List all backup files
Get-ChildItem F:\ROA\autobackup\*.BKP | Format-Table Name, Length, LastWriteTime
# 3. Count backup files
(Get-ChildItem F:\ROA\autobackup\*.BKP).Count
# 4. Show total backup size
"{0:N2} GB" -f ((Get-ChildItem F:\ROA\autobackup\*.BKP | Measure-Object -Property Length -Sum).Sum / 1GB)
# 5. Check latest Level 0 backup
Get-ChildItem F:\ROA\autobackup\L0_*.BKP -ErrorAction SilentlyContinue | Sort-Object LastWriteTime -Descending | Select-Object -First 1
# 6. Check what was copied during last restore
Get-Content D:\oracle\logs\restore_from_zero.log | Select-String "Copying|Copied"
Check DR Test Results
# 1. View latest DR test log
ssh root@10.0.20.202 "ls -lt /var/log/oracle-dr/dr_test_*.log | head -1 | awk '{print \$9}' | xargs cat | tail -100"
# 2. Check test status (passed/failed)
ssh root@10.0.20.202 "grep -E 'PASSED|FAILED|Database Verification' /var/log/oracle-dr/dr_test_*.log | tail -5"
# 3. See backup selection logic output
ssh root@10.0.20.202 "grep -A5 'TEST MODE: Selecting' /var/log/oracle-dr/dr_test_*.log | tail -20"
# 4. Check how many files were selected
ssh root@10.0.20.202 "grep 'Total files selected' /var/log/oracle-dr/dr_test_*.log | tail -1"
# 5. View RMAN errors (if any)
ssh root@10.0.20.202 "grep -i 'RMAN-\|ORA-' /var/log/oracle-dr/dr_test_*.log | tail -20"
Simulate Test Locally (on DR VM)
# 1. Start Oracle service manually
Start-Service OracleServiceROA
# 2. Run cleanup to prepare for restore
D:\oracle\scripts\cleanup_database.ps1 /SILENT
# 3. Run restore in test mode
D:\oracle\scripts\rman_restore_from_zero.ps1 -TestMode
# 4. Verify database opened correctly
sqlplus / as sysdba @D:\oracle\scripts\verify_db.sql
# 5. Check what backups were used
Get-Content D:\oracle\logs\restore_from_zero.log | Select-String "backup piece"
# 6. View database verification output
Get-Content D:\oracle\logs\restore_from_zero.log | Select-String -Pattern "DB_NAME|OPEN_MODE|TABLES" -Context 0,1
Common Restore Test Issues
| Issue | Check | Fix |
|---|---|---|
| Test reports FAILED but DB is open | Check log for "OPEN_MODE: READ WRITE" | Already fixed in latest version |
| Missing datafiles in restore | Count backup files: should be 15-40+ | Wait for next full backup or copy all files |
| "No backups found" error | Verify NFS mount: Test-Path F:\ |
Remount NFS or check Proxmox NFS service |
| Restore takes > 30 min | Check backup size: should be ~5-8 GB | Normal for first restore after format change |
| RMAN-06023 errors | Check for L0_*.BKP files on F:\ | Old format: need new backup with naming convention |
Verify Naming Convention is Active
# Check if new naming convention is being used (after Oct 11, 2025)
ssh root@10.0.20.202 "ls /mnt/pve/oracle-backups/ROA/autobackup/ | grep -E '^(L0_|L1_|ARC_|SPFILE_|CF_)' | wc -l"
# Should return > 0 if active
# If 0, backups are still using old format (O1_MF_ANNNN_*)
# Wait for next scheduled backup (02:30 daily) or run manual backup
Manual Test Run with Verbose Output
# Run test with full output visible
ssh root@10.0.20.202
cd /opt/scripts
./weekly-dr-test-proxmox.sh 2>&1 | tee /tmp/dr_test_manual.log
# Watch in real-time what's happening
# Look for these key stages:
# - "TEST MODE: Selecting latest backup set"
# - "Total files selected: XX"
# - "RMAN restore completed successfully"
# - "OPEN_MODE: READ WRITE"
❌ Backup Monitor Not Sending Alerts
# 1. Check templates exist
ssh root@10.0.20.202 "ls /usr/share/pve-manager/templates/default/oracle-*"
# 2. Reinstall templates
ssh root@10.0.20.202 "/opt/scripts/oracle-backup-monitor-proxmox.sh --install"
# 3. Check Proxmox notifications work
ssh root@10.0.20.202 "pvesh create /nodes/$(hostname)/apt/update"
# Should receive update notification
❌ F:\ Drive Not Accessible in VM
# On VM 109:
# 1. Check NFS Client service
Get-Service | Where {$_.Name -like "*NFS*"}
# 2. Manual mount
mount -o rw,nolock,mtype=hard,timeout=60 10.0.20.202:/mnt/pve/oracle-backups F:
# 3. Check Proxmox NFS server
ssh root@10.0.20.202 "showmount -e localhost"
# Should show: /mnt/pve/oracle-backups 10.0.20.37
❌ Restore Fails
# 1. Check backup files exist
dir F:\ROA\autobackup\*.BKP
# 2. Check Oracle service
sc query OracleServiceROA
# 3. Check PFILE exists
dir C:\Users\oracle\admin\ROA\pfile\initROA.ora
# 4. View restore log
type D:\oracle\logs\restore_from_zero.log
❌ VM Won't Start
# Check VM status
ssh root@10.0.20.202 "qm status 109"
# Check VM config
ssh root@10.0.20.202 "qm config 109 | grep -E 'memory|cores|bootdisk'"
# Force unlock if locked
ssh root@10.0.20.202 "qm unlock 109"
# Start with console
ssh root@10.0.20.202 "qm start 109 && qm terminal 109"
📈 Monitoring & Metrics
Key Metrics
| Metric | Target | Alert Threshold |
|---|---|---|
| FULL Backup Age | < 24h | > 25h |
| CUMULATIVE Age | < 6h | > 7h |
| Backup Size | ~7 GB/day | > 10 GB |
| Restore Time | < 15 min | > 30 min |
| Disk Usage | < 80% | > 80% |
Check Logs
# Backup logs (on PRIMARY)
Get-Content D:\rman_backup\logs\backup_*.log -Tail 50
# Transfer logs (on PRIMARY) - UNIFIED script
Get-Content D:\rman_backup\logs\transfer_*.log -Tail 50
# Monitoring logs (on Proxmox)
tail -50 /var/log/oracle-dr/*.log
# Restore logs (on VM 109)
type D:\oracle\logs\restore_from_zero.log
🔐 Security & Access
SSH Keys Setup
PRIMARY (10.0.20.36) ──────► PROXMOX (10.0.20.202)
SSH Key
Port 22
LINUX WORKSTATION ─────────► PROXMOX (10.0.20.202)
SSH Key
Port 22
LINUX WORKSTATION ─────────► VM 109 (10.0.20.37)
SSH Key
Port 22122
Required Credentials
- PRIMARY: Administrator (for scheduled tasks)
- PROXMOX: root (for scripts and VM control)
- VM 109: romfast (user), SYSTEM (Oracle service)
📅 Maintenance Schedule
| Day | Time | Action | Duration | Impact |
|---|---|---|---|---|
| Daily | 02:30 | FULL Backup | 30 min | None |
| Daily | 09:00 | Monitor Backups | 1 min | None |
| Daily | 13:00 | CUMULATIVE Backup | 5 min | None |
| Daily | 18:00 | CUMULATIVE Backup | 5 min | None |
| Saturday | 06:00 | DR Test | 30 min | None |
🚨 Disaster Recovery Procedure
When PRIMARY is DOWN:
-
Confirm PRIMARY is unreachable
ping 10.0.20.36 # Should fail -
Start DR VM
ssh root@10.0.20.202 "qm start 109" -
Wait for boot (3 minutes)
-
Connect to DR VM
ssh -p 22122 romfast@10.0.20.37 -
Run restore
D:\oracle\scripts\rman_restore_from_zero.cmd -
Verify database
sqlplus / as sysdba SELECT name, open_mode FROM v$database; -- Should show: ROA, READ WRITE -
Update application connections
- Change from: 10.0.20.36:1521/ROA
- Change to: 10.0.20.37:1521/ROA
-
Monitor DR system
- Database is now production
- Do NOT run cleanup!
- Keep VM running
📝 Quick Reference Card
╔══════════════════════════════════════════════════════════════╗
║ DR QUICK REFERENCE ║
╠══════════════════════════════════════════════════════════════╣
║ PRIMARY DOWN? ║
║ ssh root@10.0.20.202 ║
║ qm start 109 ║
║ # Wait 3 min ║
║ ssh -p 22122 romfast@10.0.20.37 ║
║ D:\oracle\scripts\rman_restore_from_zero.cmd ║
╠══════════════════════════════════════════════════════════════╣
║ TEST DR? ║
║ ssh root@10.0.20.202 "/opt/scripts/weekly-dr-test-proxmox.sh"║
╠══════════════════════════════════════════════════════════════╣
║ CHECK BACKUPS? ║
║ ssh root@10.0.20.202 "/opt/scripts/oracle-backup-monitor-proxmox.sh"║
╠══════════════════════════════════════════════════════════════╣
║ SUPPORT: ║
║ Logs: /var/log/oracle-dr/ ║
║ Docs: proxmox/vm109-windows-dr/docs/ ║
╚══════════════════════════════════════════════════════════════╝
📂 Structură Director
vm109-windows-dr/
├── README.md # Acest fișier
├── docs/
│ ├── PLAN_TESTARE_MONITORIZARE.md # Plan testare și monitorizare DR
│ ├── PROXMOX_NOTIFICATIONS_README.md # Configurare notificări Proxmox
│ ├── FAILBACK_PROCEDURE.md # Failback DR → PRIMARY (procedura inversă)
│ └── archive/ # Planuri și statusuri anterioare
│ ├── DR_UPGRADE_TO_CUMULATIVE_PLAN.md
│ ├── DR_VM_MIGRATION_GUIDE.md
│ ├── DR_WINDOWS_VM_IMPLEMENTATION_PLAN.md
│ └── DR_WINDOWS_VM_STATUS_2025-10-09.md
└── scripts/
├── oracle-backup-monitor-proxmox.sh # Monitorizare zilnică (Proxmox)
├── weekly-dr-test-proxmox.sh # Test săptămânal DR (Proxmox)
├── rman_backup.bat # RMAN full backup (Windows)
├── rman_backup_incremental.bat # RMAN incremental (Windows)
├── transfer_backups.ps1 # Transfer backup-uri (Windows)
├── rman_restore_from_zero.ps1 # Restore PRIMARY → DR (disaster activation)
├── rman_restore_to_primary.ps1 # Restore DR → PRIMARY (failback)
├── cleanup_database.ps1 # Cleanup după test (Windows DR)
└── *.ps1 # Alte scripturi configurare
Last Updated: 2026-01-27 Version: 2.2 - Unified transfer script (transfer_backups.ps1) Status: ✅ Production Ready
📋 Changelog
v2.2 (Oct 31, 2025)
- ✨ Unified transfer script: Replaced
transfer_to_dr.ps1andtransfer_incremental.ps1with singletransfer_backups.ps1 - 🎯 Smart duplicate detection: Automatically skips files that exist on DR
- ⚡ Flexible scheduling: Can run after any backup type or manually
- 🔧 Simplified maintenance: One script to maintain instead of two
v2.1 (Oct 11, 2025)
- Added restore test debugging guide
- Implemented new backup naming convention