Database operations
Database operations combine engine health with the infrastructure that makes the engine usable: database-node availability, storage, binding listeners, workload connectivity, and access identity. Monitor all of those layers before treating a failed application query as an engine incident.
The operating goal is to protect data and restore the customer path with the smallest justified change. The database owner owns backup, integrity, and engine decisions; the platform owner owns Node and storage availability; the application owner verifies real queries and permission boundaries. A successful recovery restores the original client path, explains the failure, preserves evidence, and removes temporary access or exposure.
Agree on recovery point objective (RPO), recovery time objective (RTO), escalation ownership, and restore evidence before the database becomes critical. Monitoring and restart controls reduce diagnosis time, but they do not replace a tested data-recovery plan.
What to monitor
Section titled “What to monitor”For a running managed database, monitor engine health, database-node availability, storage growth, latency, connection pressure, recent operation results, and each binding’s state. The database detail view also exposes the engine-specific metrics, logs, and eligible Explorer or Console functions.
A paused instance intentionally disables normal health, metrics, Explorer, and Console behavior until it is unpaused. Do not alert on that expected absence as an engine failure; record the pause window and verify that the service returns to Ready before re-enabling dependent workloads.
Logs, Explorer, and Console
Section titled “Logs, Explorer, and Console”Use logs to classify startup, storage, authentication, permission, and recovery errors. Read an operation’s stage before making a corrective change, especially after provisioning, resizing, or deletion. Use Explorer or Console only with the required read, write, or administrative permission, and avoid unbounded production queries through an operator interface.
Gateway configuration backups do not replace database data backups. Use an engine-appropriate backup method, retain recovery evidence, and test restores against a separate target. Plan a backup before destructive lifecycle work, substantial resize, or a migration that involves persistent data.
The interactive Console is intentionally bounded by an execution budget shared across statements. It is an operator tool, not a batch migration runner. Prefer versioned migration tooling for schema changes and long-running maintenance, and use a read-only or narrowly privileged identity whenever possible. A browser disconnect must not be treated as proof that a query was cancelled; check the resulting engine state before retrying.
Monitoring collection is background-owned. Gateway starts collection after bootstrap and when a managed database becomes ready; opening the database detail page is not the trigger. On a first visit, a chart may wait for the first real sample or historical query, but the UI must distinguish absent data from a healthy sample.
Recover a failed workload binding
Section titled “Recover a failed workload binding”Check a failed binding in this order:
- Confirm that the managed database is Ready and that its database node is online.
- Confirm the target Docker node and workload are online and reporting current state.
- Review the durable binding desired state and its most recent operation.
- Check that the daemon-owned listener is reconciled on the dedicated bridge network.
- Verify the binding’s distinct engine principal and intended permissions.
- Check the workload’s received configuration and application logs.
Prefer reconciliation or a targeted retry after correcting the root cause. Do not remove Gateway-owned network objects or engine identities manually; doing so can turn a recoverable listener issue into a state mismatch or orphaned cleanup problem.
Operational changes and deletion
Section titled “Operational changes and deletion”Pause/unpause, restart, resize, credential rotation, certificate rotation, direct publication, binding deletion, and database deletion should each be followed by the relevant verification: engine health, client connection, route or listener state, logs, metrics, and recorded Task result. Direct publication is opt-in; monitor it as an external-client path separately from private bindings.
When deleting a database, remove or migrate bindings first and wait for their listener and identity cleanup. A failed delete should remain in Gateway’s operation history for recovery. Do not force-remove storage or users at the daemon or engine layer unless an explicit recovery procedure calls for it.
Backup and restore runbook
Section titled “Backup and restore runbook”Define an engine-specific recovery point objective and recovery time objective before production use. Keep backups off the Database Node, encrypt them, record the engine version and restore command, and test them on an isolated target. A successful backup command without a successful restore test is not recovery evidence.
For a restore exercise:
- Create a separate target with compatible engine and storage capacity.
- Restore the backup without overwriting the active instance.
- Run integrity checks and representative application queries.
- Recreate access with new bindings or narrowly scoped client identities.
- Compare row/key/table counts and application-level invariants.
- Record duration, manual steps, and any version constraint discovered.
Gateway does not provide point-in-time recovery merely because it manages the container lifecycle. WAL, Redis persistence, ClickHouse backup strategy, replication, and off-site retention remain database-operations responsibilities.
Incident classification
Section titled “Incident classification”Classify before changing state: control-plane failures affect Tasks or reconciliation; node failures affect daemon freshness and local runtime; engine failures appear in database logs and health; storage failures affect mount or capacity; binding failures affect listener or principal state; application failures appear after connectivity succeeds. This ordering prevents a healthy engine from being restarted to solve an application permission error.
After recovery, verify the original customer path, not only the administrative screen. Confirm a real query, expected permissions, stable health collection, and that no temporary direct publication, elevated role, debug token, or manual network change remains.
