Hermes Agent Deployment (Provisioning Flow)
Last updated: 2026-08-06
This document describes how a Hermes Agent runtime is provisioned on the HostAgentics platform: the container model, dashboard, persistence, gateway, safe mode, and scheduled work. It is an internal engineering document; customer-facing claims must stay within what this flow actually delivers.
1. Container model: single container, official entrypoint
A Hermes runtime is a **single container** from the official Nous Research Hermes Agent image (`nousresearch/hermes-agent:v2026.8.3` at the tested default, digest-pinned via the version catalog). The official container uses s6-overlay supervision: `/init` is PID 1, the gateway runs as the main program, and the dashboard is a supervised service enabled via `HERMES_DASHBOARD=1`. The adapter **never bypasses this architecture** — the default entrypoint is kept intact, so the runtime behaves exactly like an officially supported Hermes Agent deployment.
2. Dashboard (port 9119, always authenticated)
3. Persistence: memories, skills, profiles
All customer data — memories, skills, profiles, and any files the agent writes — persists under **`/opt/data`** (`HERMES_HOME=/opt/data`, `HERMES_WRITE_SAFE_ROOT=/opt/data`) on the runtime's own persistent volume. The write-safe root is confined to that directory, so agent file writes stay inside the runtime's storage. Storage is never shared between runtimes, and the platform does not read agent memory.
4. Gateway and API access
5. Safe mode
Safe mode is the default posture: agent tool use and outbound actions are restricted unless the customer explicitly configures otherwise, and the platform enforces the plan's concurrency and browser-session caps regardless of mode (see `docs/sandboxing.md`). There is no host Docker socket, no cloud metadata access, and no shared agent memory across runtimes.
6. Cron jobs and scheduled work
Hermes runtimes support cron jobs and scheduled work, persisted with the runtime's data. Scheduled jobs run within the runtime's own resource envelope and concurrency caps; they are subject to the same limit warnings and enforcement as on-demand work (see `docs/resource-limits.md`).