HostAgentics Docs

HostAgentics Restores

Last updated: 2026-08-06

This document describes how a runtime restore works on the HostAgentics platform: selecting a provider-confirmed snapshot set, restoring it, verifying health, and auditing the outcome. The restore gate is strict: **a restore is never marked successful before its health check passes.** The operational runbook lives in `docs/recovery-runbook.md`; this page is the design reference.

1. When restores happen

  • **Self-service restore**: a customer restores their runtime from a provider-confirmed snapshot set (e.g. after an accidental destructive change inside their runtime).
  • **Operational restore**: during recovery from infrastructure incidents or shard evacuation (see `docs/recovery-runbook.md`).
  • In both cases the same flow and the same gates apply.

    2. Restore flow

    1. **Select a confirmed backup.** Only backups in `completed` status with an intact snapshot manifest and inside the retention window (7 days standard, 14 days with Resource Boost) are restore-eligible. The list shows backup time and provider-reported size; the worker rechecks the manifest before restore.

    2. **Snapshot restore.** The chosen backup artifact is restored onto the runtime's volume through the provider abstraction (`restoreVolumeBackup`). For n8n, both the n8n data volume and the dedicated PostgreSQL volume are restored — a partial restore (data without database) is not a valid n8n restore.

    3. **Restart.** The runtime service is restarted with the restored volume attached.

    4. **Health verification.** A real health probe against the runtime's own health endpoint must pass (n8n: `/healthz`; OpenClaw: gateway response on 18789; Hermes: `/api/status` on the dashboard port). **The restore operation is not marked `completed` until this probe passes** — `restore_operations.healthVerifiedAt` records the moment, and `completedAt` follows it. If health does not recover, the restore is marked `failed`, the runtime is left in a recoverable state, and on-call is alerted.

    5. **Audit.** The restore is recorded as a `restore_operations` row linked to a `runtime_operations` entry (type `restore`) with a correlation ID, plus an audit event.

    3. Completion gate (never simulated)

    The invariant: `restore_operations.status = completed` implies data restored **and** runtime started **and** health check passed **and** the event audited. There is no code path that sets `completed` without the health verification timestamp. A restore whose health check cannot be confirmed is reported as failed or in-progress — never as successful.

    4. RPO/RTO: design targets, not guarantees

  • **Recovery point objective (RPO)**: bounded by the daily backup cadence — the design target is at most one day of data loss for a full restore. This is a design target, not a guarantee.
  • **Recovery time objective (RTO)**: the design target for a full restore is the time to restore the snapshot, restart the service, and pass the health check — measured in minutes for a healthy restore path. This is a design target, not a guarantee; actual time depends on artifact size, region, and concurrent load.
  • HostAgentics does not yet offer a formal SLA; recovery framing on customer-facing surfaces must use "design targets", never commitments (see `docs/pricing.md`).
  • 5. Restore eligibility and limitations

  • A backup older than the retention window is not restore-eligible; retention windows are shown in the dashboard.
  • Restores restore the runtime's own storage to the point of the backup. Anything the customer's workflows sent to third-party services is outside the backup and is not restored.
  • A restore replaces the current state of the runtime's volume with the backup's state. Restoring is a deliberate action: the UI requires explicit confirmation, and the operation records who initiated it and from which backup.
  • During an active restore, the runtime may be briefly unavailable; the dashboard shows restore progress rather than implying the runtime is healthy mid-restore.
  • 6. Related documents

  • `docs/backups.md` — how backups become restore-eligible
  • `docs/backup-runbook.md`, `docs/recovery-runbook.md` — operational procedures
  • `docs/incident-response.md` — when restores are part of incident handling