Incident runbook
This runbook helps the incident lead protect customer service, establish which subsystem owns the failure, and choose a reversible recovery path. The incident lead owns priorities and communication; the platform operator owns Gateway and managed-node evidence; application and database owners validate their services. Do not let the person typing commands become the only person deciding risk.
Success means customer impact is understood, further change is controlled, the authoritative runtime state is reconciled, and recovery is verified from the user-facing path. A green Dashboard alone is not recovery evidence.
First five minutes
Section titled “First five minutes”- Record start time, affected customer path, and recent changes.
- Check Gateway, PostgreSQL, Redis, and Relay health.
- Identify whether impact is control plane, ingress, one node, one workload, or a shared dependency.
- Freeze unrelated changes.
- Use maintenance mode or a status-page incident when customer impact is confirmed.
Relay unavailable
Section titled “Relay unavailable”Inspect the Relay service, identity volume, database reachability, image version, and public 9443/tcp. Do not move the port back to the application or bypass authorization with a new public listener.
Node offline
Section titled “Node offline”Existing workloads may continue. Check host power/network, daemon service, time, certificate identity, and outbound Relay reachability. Do not mutate from stale inventory.
Bad release
Section titled “Bad release”For Deployments, roll back to the previous healthy slot. For Pages, move the Tag. For Git-built containers, redeploy the last approved digest. Preserve logs and operation history.
Database binding failure
Section titled “Database binding failure”Check Relay, both nodes, database health, binding desired state, the target daemon’s listener status, and the engine principal. Retry reconciliation before manual identity changes; there is no per-binding database connector container to restart or replace.
Use the layer-by-layer checks in Application database bindings and preserve the failed Task before retrying.
Closeout
Section titled “Closeout”Confirm recovery from an external client, clear maintenance intentionally, document the root cause and detection gap, and create follow-up work for every temporary mitigation.
Control-plane unavailable
Section titled “Control-plane unavailable”Check the Gateway container/process, PostgreSQL, Redis, persistent volumes, disk pressure, migrations, and application logs. Preserve the existing encryption keys and database before attempting a replacement instance. If customer workloads continue, avoid broad host changes while restoring the control plane.
First determine whether Relay is still reachable. If Relay is healthy, already applied nginx configuration, containers, databases, and permitted existing private streams can continue while new management and authorization decisions wait. If the only Relay is local to the failed Gateway host, Secure Links and managed-node sessions lose their transport as well; ordinary host-local workloads may still run, but existing Relay-dependent streams are not guaranteed. Independent external Relay members reduce this shared failure only when Nodes can actually reach and use them.
After Gateway returns, verify sign-in, decryption, Relay health, background jobs, update discovery, notifications, and fresh node snapshots. Do not assume daemons reconciled merely because the Dashboard loads.
Use Availability, compatibility, and limits to classify the affected paths and Updates and backups before restoring a replacement instance or changing persistent state.
Ingress outage
Section titled “Ingress outage”Follow Ingress troubleshooting from DNS through network, TLS, nginx apply, Route health, and upstream. Use maintenance mode when the virtual host should remain stable during repair. Preserve generated-config validation errors and external request evidence.
Workload operation stuck
Section titled “Workload operation stuck”Inspect the durable Task, owning Node connection, daemon operation history, runtime resource, and request/operation ID. A browser timeout does not prove the command failed. Reconcile the first operation before starting a duplicate create, recreate, migration, or delete.
If Force Cancel is available, understand whether it cancels only Gateway tracking or also dispatches cancellation to the owner. After cancellation, refresh runtime state and clear any interrupted-operation marker only through the supported reconciliation path.
Use Tasks, events, and audit to distinguish accepted, running, cancelled, and completed work, then continue with the owning Docker resource.
Database unavailable
Section titled “Database unavailable”Separate engine health, storage, direct publication, managed binding, and application-query failures. Verify the database Node, engine container, storage image/mount, monitoring snapshot, Relay path, target Docker Node, daemon-owned listener, binding identity, and application configuration.
Do not delete the database or engine principal as a diagnostic step. Preserve storage and operation history, then retry the narrow failed reconciliation.
Continue with Database operations for engine, storage, backup, and restore checks, or Application database bindings when only the private application path is affected.
Communication and evidence
Section titled “Communication and evidence”Maintain one incident timeline with UTC timestamps, affected resources, user-visible impact, changes, task IDs, request IDs, decisions, and verification. Public updates describe impact and progress without exposing topology or security details.
Recovery gate
Section titled “Recovery gate”Recovery requires:
- the customer-facing path works from outside the managed network;
- Gateway desired state matches current owner state;
- alerts recover for the right reason;
- queued operations and deliveries are understood;
- temporary exposure or bypasses are removed;
- a follow-up owner and deadline exist for every remaining risk.
Decision framework for the incident lead
Section titled “Decision framework for the incident lead”Prefer actions that preserve evidence and reduce blast radius. Pause unrelated automation before restarting shared services. If the control plane is unavailable but workloads still serve traffic, restore management without recreating healthy resources. If ingress is unhealthy, keep the hostname stable with maintenance mode or a tested rollback rather than changing DNS repeatedly. If data integrity is uncertain, stop writes before optimizing availability.
Escalate immediately when the incident involves possible credential disclosure, damaged storage, unknown schema migration state, loss of Relay identity, inconsistent database binding ownership, or a release artifact that cannot be verified. These conditions can turn a routine restart into permanent loss or unauthorized access.
Operator evidence table
Section titled “Operator evidence table”| Evidence | Why it matters |
|---|---|
| External DNS, TLS, and application response | Confirms actual customer impact |
| Gateway, PostgreSQL, Redis, and Relay health | Separates control-plane and data-plane failure |
| Node last seen, capability state, and daemon logs | Identifies the runtime owner and stale inventory |
| Task, operation, and request IDs | Prevents duplicate or ambiguous mutations |
| Last approved version and artifact digest | Provides a known rollback target |
| Database and volume backup status | Defines safe recovery options |
After recovery, perform a short review while evidence is still available. Record the initiating change, why detection did or did not work, which manual steps were needed, and which verification would have caught the problem earlier. Convert temporary instructions into a tested runbook change rather than preserving undocumented shell history.
