HostAgentics Docs

HostAgentics Monitoring

Last updated: 2026-08-06

This document describes how the HostAgentics platform monitors runtimes and the control plane: health checks, metric and usage snapshots, limit warnings, the status page, and reference codes. A core rule applies everywhere: **when a value is not available, surfaces say "Data unavailable" — they never show an invented number.**

1. Runtime health checks

  • Every runtime has a health endpoint that the monitoring engine probes (`apps/worker/src/monitoring.ts`): n8n `/healthz` on port 5678, OpenClaw gateway on 18789 (any HTTP response, including 401), Hermes `/api/status` on the dashboard port 9119 (OK or 401).
  • **Consecutive failures → degraded → auto-restart**: a single failed probe is not an incident. The engine counts consecutive failures; after the threshold, the runtime is marked `degraded` and restarted automatically (with backoff and alerting). Restarts are recorded (`metricSnapshots.restartCount`, `healthChecks` rows).
  • If health does not recover after restart with backoff, the runtime enters the incident process (`docs/incident-response.md`); the customer is notified with a safe message and a reference code, never raw diagnostics.
  • The same health probes gate provisioning (`checking_health` before `running`) and restore completion (see `docs/restores.md`).
  • 2. Metric snapshots

    `metric_snapshots` capture CPU percent, memory used, storage used, and restart counts per runtime at intervals. **All metric fields are nullable**: a null value means data was not available at capture time. The dashboard renders nulls as "Data unavailable" — there is no interpolation, no placeholder zero, and no fabricated trend. Charts may show gaps; they never show invented points.

    3. Usage snapshots and limit warnings

    `usage_snapshots` track monthly outbound transfer, workflow executions, and agent task counts per period (unique per runtime + period start, so the limit engine upserts safely). The limit engine compares measured usage against the plan's entitlements (`docs/resource-limits.md`) and emits warnings at **70%, 85%, 95%, and 100%** of storage and of monthly outbound transfer:

  • Each crossing is recorded once per runtime per threshold in `limit_events` (with `notified` to prevent spam) and surfaced as dashboard notifications and email.
  • At 100%: data is preserved, new work is paused/rate-limited, and an upgrade is offered — **never an overage charge** (see `docs/fixed-pricing.md`).
  • 4. Status page

    `status.hostagentics.com` reflects platform components (`status_components`: dashboard, provisioning, n8n hosting, agent hosting, backups, updates, API, Relay) with statuses operational / degraded / partial_outage / major_outage / maintenance. Incidents (`incidents`) carry public updates (timestamp, status, message) published through the incident process. A machine-readable aggregate is exposed at `GET /status/health.json` for monitors. Component status is derived from real checks — a component is never marked operational while its checks are failing.

    5. HA-* reference codes

    Errors that reach customers carry a safe message plus an internal **HA-*** diagnostic reference (e.g. `runtime_operations.diagnosticCode`). Customers can quote the code; support and on-call look it up to find the full internal context (correlation ID, provider error, steps). The codes are stable identifiers in the platform's error taxonomy — never free-form text, and never containing provider identifiers.

    6. Control-plane health

  • API and worker expose `/healthz` (liveness) and `/readyz` (readiness: database and Redis reachable, migrations current).
  • Web/docs expose `/healthz`; a failing readiness check prevents traffic and triggers restart with backoff.
  • Observability: structured logs with redaction (`packages/observability`), Sentry for errors (scrubbed), PostHog for product analytics (filtered, no secrets, no provider identifiers).
  • 7. Related documents

  • `docs/resource-limits.md` — entitlements the limit engine enforces
  • `docs/incident-response.md` — what happens when monitoring catches a real failure
  • `docs/backups.md`, `docs/updates.md` — monitored operations