Skip to content

Updates, backups, and restore

Updates and backups protect two different outcomes: updates move the platform forward, while backups make it possible to recover identity, desired state, service credentials, and application data after failure. The platform owner owns the update sequence; the security owner owns recovery keys; database and application owners own engine-native data backups and validation.

Define the recovery point objective (how much data may be lost) and recovery time objective (how long restoration may take) before choosing backup frequency or update windows. Success is a tested restore, not the existence of archive files.

Before a Gateway update, read release notes, verify compatibility floors, back up PostgreSQL and persistent files, preserve encryption/master keys, and record the running image references. Update through the supported installer or generated Compose configuration.

Daemon updates are signed and managed per node. Keep mixed-version behavior within the documented compatibility window and verify reconnect, inventory, bindings, logs, and operations after each role update.

Back up:

  • PostgreSQL;
  • persistent Gateway artifact/configuration volumes;
  • encryption and PKI master keys;
  • Relay identity volume;
  • engine-native database backups;
  • external DNS, identity-provider, email, registry, and SIEM configuration held outside Gateway.

A restore test must prove sign-in, decryption, node reconnection, Relay authorization, Routes, certificates, managed database bindings, and audit continuity. A backup that has never been restored is not a recovery plan.

External storage connections on every plan and managed storages with Secure Links on Personal and higher are expected in 2.11. Built-in managed-database backup/restore on Personal and higher, and Gateway configuration export for transfer to another instance on every plan, are expected in 2.12. These are roadmap estimates, not available operations today.

Planned configuration export is distinct from existing container archive export. Until these features ship, use the manual recovery set below and engine-native database backups; configuration transfer alone does not replace application-data protection.

Document the owner, location, retention, encryption method, and restore order for every backup component. Keep master keys and recovery credentials outside the Gateway host and outside the same failure domain as the database backup.

PostgreSQL is the authoritative store for identities, desired state, relationships, operations, and audit metadata. Persistent volumes hold artifacts and service identity material that cannot be reconstructed from PostgreSQL alone. Engine-native managed database backups protect application data; a control-plane backup is not a substitute.

The manual installer uses services named app, relay, registry, postgres, and redis, plus these persistent volumes:

Volume Recovery purpose
postgres_data users, resources, desired state, Tasks, and audit metadata
gateway_data uploaded artifacts, generated configuration, TLS, and application state
gateway_relay_identity local Relay identity shared by app and relay
gateway_relay_state local Relay runtime state
gateway_registry_data private registry blobs and manifests
gateway_registry_auth registry token signing material
redis_data persisted Redis state; useful for a coordinated full recovery but not a substitute for PostgreSQL

Also preserve .env and the exact docker-compose.manual.yml used by the installation. The .env file contains recovery-sensitive secrets; encrypt its backup separately and never attach it to a support ticket.

The following example creates a short offline consistency window. Run it from the directory containing the manual Compose file:

Terminal window
mkdir -p gateway-recovery/{gateway_data,gateway_relay_identity,gateway_relay_state,gateway_registry_data,gateway_registry_auth,redis_data}
cp .env docker-compose.manual.yml gateway-recovery/
chmod 700 gateway-recovery
chmod 600 gateway-recovery/.env
docker compose -f docker-compose.manual.yml stop app relay registry
docker compose -f docker-compose.manual.yml exec -T postgres \
pg_dump -U gateway -d gateway -Fc > gateway-recovery/postgres.dump
docker compose -f docker-compose.manual.yml stop redis
docker compose -f docker-compose.manual.yml cp --archive app:/var/lib/gateway/. gateway-recovery/gateway_data/
docker compose -f docker-compose.manual.yml cp --archive app:/var/lib/gateway-relay/. gateway-recovery/gateway_relay_identity/
docker compose -f docker-compose.manual.yml cp --archive relay:/var/lib/gateway-relay/state/. gateway-recovery/gateway_relay_state/
docker compose -f docker-compose.manual.yml cp --archive registry:/var/lib/registry/. gateway-recovery/gateway_registry_data/
docker compose -f docker-compose.manual.yml cp --archive app:/var/lib/gateway-registry-auth/. gateway-recovery/gateway_registry_auth/
docker compose -f docker-compose.manual.yml cp --archive redis:/data/. gateway-recovery/redis_data/
docker compose -f docker-compose.manual.yml start redis registry app relay

Copy gateway-recovery to encrypted off-host storage and verify its checksum there. This example does not back up databases managed on Database Nodes; use each engine’s native backup procedure for application data.

To restore into a clean compatible host, first restore .env and the same Compose file, then create stopped containers, copy the persistent data, start PostgreSQL, and import the dump before starting the rest of the control plane:

Terminal window
docker compose -f docker-compose.manual.yml create
docker compose -f docker-compose.manual.yml cp --archive gateway-recovery/gateway_data/. app:/var/lib/gateway
docker compose -f docker-compose.manual.yml cp --archive gateway-recovery/gateway_relay_identity/. app:/var/lib/gateway-relay
docker compose -f docker-compose.manual.yml cp --archive gateway-recovery/gateway_relay_state/. relay:/var/lib/gateway-relay/state
docker compose -f docker-compose.manual.yml cp --archive gateway-recovery/gateway_registry_data/. registry:/var/lib/registry
docker compose -f docker-compose.manual.yml cp --archive gateway-recovery/gateway_registry_auth/. app:/var/lib/gateway-registry-auth
docker compose -f docker-compose.manual.yml cp --archive gateway-recovery/redis_data/. redis:/data
docker compose -f docker-compose.manual.yml up -d postgres
docker compose -f docker-compose.manual.yml cp gateway-recovery/postgres.dump postgres:/tmp/gateway.dump
docker compose -f docker-compose.manual.yml exec -T postgres \
pg_restore -U gateway -d gateway --clean --if-exists /tmp/gateway.dump
docker compose -f docker-compose.manual.yml up -d redis registry app relay

Use the same release first. Verify sign-in, secret decryption, Relay identity, and fresh Node reconnects before attempting an upgrade. File ownership can differ across hosts; if a service reports permission errors, compare ownership inside the original and restored containers rather than applying broad writable permissions.

  1. Read the release notes and required intermediate versions.
  2. Confirm the active update channel and target version.
  3. Verify recent PostgreSQL and persistent-volume backups.
  4. Record current Gateway, Relay, and daemon versions and image digests.
  5. Confirm PostgreSQL, Redis, Relay, and disk health.
  6. Place customer-facing services into maintenance only when the update requires it.
  7. Apply the supported update without replacing persistent volumes or master keys.
  8. Wait for Gateway readiness and schema migration completion.
  9. Verify Relay, sign-in, Dashboard bootstrap, node reconnects, Routes, and database bindings.
  10. Remove maintenance only after external verification.

Do not treat a running container as a successful update. The application must decrypt existing secrets, read prior state, reconcile background services, and expose healthy customer paths.

Update one failure domain at a time. For Relay Pool members, drain, update, verify, and return each member before moving to the next. For ordinary Nodes, verify role-specific resources after reconnect. Do not update every ingress, database, or Relay member simultaneously unless the environment has an independently tested recovery path.

  1. Provision a clean compatible Gateway environment.
  2. Restore PostgreSQL and persistent volumes from the same recovery point.
  3. Restore encryption, PKI, and Relay identity material with original permissions.
  4. Start internal dependencies, Gateway, and Relay in the documented order.
  5. Verify sign-in and decryption before permitting mutations.
  6. Allow managed nodes to reconnect with their existing identities.
  7. Reconcile Routes, certificates, workloads, database bindings, notifications, and integrations.
  8. Restore application databases through engine-native procedures where required.
  9. Confirm audit continuity and record the recovery event.

If only part of the recovery set is available, stop and assess the consequence. Generating replacement master keys or identities can permanently orphan encrypted state or managed nodes.

Decide the rollback trigger before starting an update: failed migration, inability to decrypt existing secrets, Relay authorization failure, incompatible daemon state, broken customer Routes, or another measurable condition. Keep the prior approved image references and the matching backup until the new release completes its observation window.

Application rollback and data rollback are separate. Reverting a Gateway image does not reverse a completed PostgreSQL schema migration, and moving a workload back to an older artifact does not revert mutable volume or database data. Follow the release-specific compatibility guidance and restore data only from a coordinated recovery point when integrity requires it.

Layer Required proof
Identity Existing users sign in; MFA, OAuth, and decryption work
Control plane Migrations complete; background schedulers and audit continue
Relay and Nodes Versions are compatible; Nodes reconnect with fresh inventory
Customer traffic External DNS, TLS, Route health, and application response succeed
Data Managed databases are healthy and application bindings can query
Automation Builds, webhooks, notifications, and source connectors used by the installation work

Keep the update record with start and end time, versions, image digests, backup identifiers, operator, exceptions, and final verification. This makes the next update a controlled procedure instead of a reconstruction from shell history.