Skip to content

Incident runbook

This runbook helps the incident lead protect customer service, establish which subsystem owns the failure, and choose a reversible recovery path. The incident lead owns priorities and communication; the platform operator owns Gateway and managed-node evidence; application and database owners validate their services. Do not let the person typing commands become the only person deciding risk.

Success means customer impact is understood, further change is controlled, the authoritative runtime state is reconciled, and recovery is verified from the user-facing path. A green Dashboard alone is not recovery evidence.

  1. Record start time, affected customer path, and recent changes.
  2. Check Gateway, PostgreSQL, Redis, and Relay health.
  3. Identify whether impact is control plane, ingress, one node, one workload, or a shared dependency.
  4. Freeze unrelated changes.
  5. Use maintenance mode or a status-page incident when customer impact is confirmed.

Inspect the Relay service, identity volume, database reachability, image version, and public 9443/tcp. Do not move the port back to the application or bypass authorization with a new public listener.

Existing workloads may continue. Check host power/network, daemon service, time, certificate identity, and outbound Relay reachability. Do not mutate from stale inventory.

For Deployments, roll back to the previous healthy slot. For Pages, move the Tag. For Git-built containers, redeploy the last approved digest. Preserve logs and operation history.

Check Relay, both nodes, database health, binding desired state, the target daemon’s listener status, and the engine principal. Retry reconciliation before manual identity changes; there is no per-binding database connector container to restart or replace.

Use the layer-by-layer checks in Application database bindings and preserve the failed Task before retrying.

Confirm recovery from an external client, clear maintenance intentionally, document the root cause and detection gap, and create follow-up work for every temporary mitigation.

Check the Gateway container/process, PostgreSQL, Redis, persistent volumes, disk pressure, migrations, and application logs. Preserve the existing encryption keys and database before attempting a replacement instance. If customer workloads continue, avoid broad host changes while restoring the control plane.

First determine whether Relay is still reachable. If Relay is healthy, already applied nginx configuration, containers, databases, and permitted existing private streams can continue while new management and authorization decisions wait. If the only Relay is local to the failed Gateway host, Secure Links and managed-node sessions lose their transport as well; ordinary host-local workloads may still run, but existing Relay-dependent streams are not guaranteed. Independent external Relay members reduce this shared failure only when Nodes can actually reach and use them.

After Gateway returns, verify sign-in, decryption, Relay health, background jobs, update discovery, notifications, and fresh node snapshots. Do not assume daemons reconciled merely because the Dashboard loads.

Use Availability, compatibility, and limits to classify the affected paths and Updates and backups before restoring a replacement instance or changing persistent state.

Follow Ingress troubleshooting from DNS through network, TLS, nginx apply, Route health, and upstream. Use maintenance mode when the virtual host should remain stable during repair. Preserve generated-config validation errors and external request evidence.

Inspect the durable Task, owning Node connection, daemon operation history, runtime resource, and request/operation ID. A browser timeout does not prove the command failed. Reconcile the first operation before starting a duplicate create, recreate, migration, or delete.

If Force Cancel is available, understand whether it cancels only Gateway tracking or also dispatches cancellation to the owner. After cancellation, refresh runtime state and clear any interrupted-operation marker only through the supported reconciliation path.

Use Tasks, events, and audit to distinguish accepted, running, cancelled, and completed work, then continue with the owning Docker resource.

Separate engine health, storage, direct publication, managed binding, and application-query failures. Verify the database Node, engine container, storage image/mount, monitoring snapshot, Relay path, target Docker Node, daemon-owned listener, binding identity, and application configuration.

Do not delete the database or engine principal as a diagnostic step. Preserve storage and operation history, then retry the narrow failed reconciliation.

Continue with Database operations for engine, storage, backup, and restore checks, or Application database bindings when only the private application path is affected.

Maintain one incident timeline with UTC timestamps, affected resources, user-visible impact, changes, task IDs, request IDs, decisions, and verification. Public updates describe impact and progress without exposing topology or security details.

Recovery requires:

  • the customer-facing path works from outside the managed network;
  • Gateway desired state matches current owner state;
  • alerts recover for the right reason;
  • queued operations and deliveries are understood;
  • temporary exposure or bypasses are removed;
  • a follow-up owner and deadline exist for every remaining risk.

Prefer actions that preserve evidence and reduce blast radius. Pause unrelated automation before restarting shared services. If the control plane is unavailable but workloads still serve traffic, restore management without recreating healthy resources. If ingress is unhealthy, keep the hostname stable with maintenance mode or a tested rollback rather than changing DNS repeatedly. If data integrity is uncertain, stop writes before optimizing availability.

Escalate immediately when the incident involves possible credential disclosure, damaged storage, unknown schema migration state, loss of Relay identity, inconsistent database binding ownership, or a release artifact that cannot be verified. These conditions can turn a routine restart into permanent loss or unauthorized access.

Evidence Why it matters
External DNS, TLS, and application response Confirms actual customer impact
Gateway, PostgreSQL, Redis, and Relay health Separates control-plane and data-plane failure
Node last seen, capability state, and daemon logs Identifies the runtime owner and stale inventory
Task, operation, and request IDs Prevents duplicate or ambiguous mutations
Last approved version and artifact digest Provides a known rollback target
Database and volume backup status Defines safe recovery options

After recovery, perform a short review while evidence is still available. Record the initiating change, why detection did or did not work, which manual steps were needed, and which verification would have caught the problem earlier. Convert temporary instructions into a tested runbook change rather than preserving undocumented shell history.