HostAgentics Restores
Last updated: 2026-08-06
This document describes how a runtime restore works on the HostAgentics platform: selecting a provider-confirmed snapshot set, restoring it, verifying health, and auditing the outcome. The restore gate is strict: **a restore is never marked successful before its health check passes.** The operational runbook lives in `docs/recovery-runbook.md`; this page is the design reference.
1. When restores happen
In both cases the same flow and the same gates apply.
2. Restore flow
1. **Select a confirmed backup.** Only backups in `completed` status with an intact snapshot manifest and inside the retention window (7 days standard, 14 days with Resource Boost) are restore-eligible. The list shows backup time and provider-reported size; the worker rechecks the manifest before restore.
2. **Snapshot restore.** The chosen backup artifact is restored onto the runtime's volume through the provider abstraction (`restoreVolumeBackup`). For n8n, both the n8n data volume and the dedicated PostgreSQL volume are restored — a partial restore (data without database) is not a valid n8n restore.
3. **Restart.** The runtime service is restarted with the restored volume attached.
4. **Health verification.** A real health probe against the runtime's own health endpoint must pass (n8n: `/healthz`; OpenClaw: gateway response on 18789; Hermes: `/api/status` on the dashboard port). **The restore operation is not marked `completed` until this probe passes** — `restore_operations.healthVerifiedAt` records the moment, and `completedAt` follows it. If health does not recover, the restore is marked `failed`, the runtime is left in a recoverable state, and on-call is alerted.
5. **Audit.** The restore is recorded as a `restore_operations` row linked to a `runtime_operations` entry (type `restore`) with a correlation ID, plus an audit event.
3. Completion gate (never simulated)
The invariant: `restore_operations.status = completed` implies data restored **and** runtime started **and** health check passed **and** the event audited. There is no code path that sets `completed` without the health verification timestamp. A restore whose health check cannot be confirmed is reported as failed or in-progress — never as successful.