HostAgentics Docs

HostAgentics Backups

Last updated: 2026-08-06

Every HostAgentics runtime is backed up daily. This document describes the provider-snapshot pipeline: scheduling, artifact confirmation, manifest integrity, retention, restore validation, and audit. The operational runbook lives in `docs/backup-runbook.md`; this page is the design reference.

1. What is backed up

  • **n8n**: two artifacts — the n8n data volume (`/home/node/.n8n`) and the dedicated PostgreSQL volume. Both must succeed for the runtime's backup to be complete.
  • **OpenClaw**: the state volume (`/home/node/.openclaw`, configuration + workspaces + agent state).
  • **Hermes Agent**: the data volume (`/opt/data`, memories, skills, profiles).
  • Backups are volume snapshots taken through the provider abstraction's backup primitives (`createVolumeBackup`), attributable to exactly one runtime — a backup artifact is never shared between runtimes.

    2. Schedule and retention

  • **Daily** backups run on a schedule driven by the worker. Retention: **7 days** standard, **14 days** with Resource Boost (`backupRetentionDays` entitlement).
  • Expired artifacts are deleted by the retention worker; a backup record moves to `expired` and the artifact is removed.
  • Backup artifacts are stored in the same region as the runtime they belong to (region affinity).
  • 3. Storage protection boundary

  • Backups use the infrastructure provider's native volume-backup facility and remain inside that provider boundary. HostAgentics does not download or re-encrypt snapshot bytes.
  • `backups.encryptionKeyEnvelope` is reserved for a future HostAgentics-managed artifact format and is not populated for provider snapshots. Customer-facing copy must therefore not claim application-managed backup encryption.
  • 4. Confirmation before "completed"

    A backup is **not** marked `completed` until all of the following hold (spec §27):

    1. The artifact **exists** (the snapshot is actually present, not merely requested).

    2. The complete **snapshot manifest is hashed** — `checksumSha256` protects the provider backup ids, resource types, and creation metadata from control-plane corruption. It is a manifest checksum, not a byte-level checksum of provider-owned snapshot data.

    3. **Metadata is persisted** — artifact manifest, reported size, manifest checksum, and expiry are written to the `backups` row.

    4. The completion event is **audited**.

    If any step fails, the backup is marked `failed` and alerts on-call. There is no path that marks a backup completed without verification — a backup that cannot be verified is a failed backup, not a completed one.

    5. Artifact validation and restore eligibility

  • A backup becomes **restore-eligible** only in `completed` status with a valid manifest checksum and inside its retention window. The restore flow (`docs/restores.md`) refuses to use anything else and rechecks the manifest before calling the provider.
  • `packages/backups` encodes retention selection and restore completion gates; the worker confirms provider artifacts and validates the stored manifest.
  • Restore operations reference a specific provider-confirmed backup (`restore_operations.backupId`); the operation is not marked completed until data is restored, the runtime starts, and its health check passes.
  • 6. Audit and observability

  • Every backup run is a `runtime_operations` entry (type `backup`) with a correlation ID linking the operation to its provider calls and logs.
  • Completion, failure, and expiry transitions are audit events; backup failures page on-call via the incident process (`docs/incident-response.md`).
  • Backups are counted toward internal cost accounting (`internalCostSnapshots` category `backup_storage`) so the fixed-price model stays sustainable.
  • 7. Known limitations

  • Backups are point-in-time snapshots at the daily cadence; the design target for data loss is bounded by that cadence (see `docs/restores.md` for RPO/RTO framing — targets, not guarantees).
  • A backup cannot protect against data the customer's workflows send to third-party services; it protects the runtime's own storage.
  • Snapshot creation is confirmed by the provider, but no backup system can guarantee restorability under every failure mode; restore drills are part of the recovery process (see `docs/recovery-runbook.md`).
  • 8. Related documents

  • `docs/backup-runbook.md` — operational procedure
  • `docs/restores.md` — restore flow and eligibility
  • `docs/recovery-runbook.md` — evacuation and disaster recovery