Skip to content

Relay and Relay Pool

Relay is the authenticated transport that connects Gateway with managed Nodes and carries supported private traffic such as Secure Links. A Relay Pool adds members so transport capacity and availability do not depend on one failure domain. It is not a general VPN or an application load balancer.

Add pool members when the business requires more connection capacity, planned maintenance without losing all transport, or resilience across hosts or sites. The network owner must provide reachable endpoints; the platform owner controls enrollment, assignment, drain, and updates. Success means managed Nodes can establish new sessions through the intended members and private application paths continue through a member outage or drain.

The local Relay is a required long-lived data-plane service and the sole public owner of 9443/tcp. It authenticates managed-node sessions and supported private tunnel traffic.

A Relay Pool adds supervisor/worker members and fault-domain-aware placement. Operators verify distinct physical domains, advertise reachable worker endpoints, rebalance explicitly, drain members, and apply signed rolling updates one member at a time.

Gateway does not create firewall rules, NAT traversal, or a general overlay network. Every assigned managed host must reach its worker data endpoint. During a Relay incident, preserve the identity volume and avoid introducing an alternate unauthenticated path.

Every installation has a local Relay service associated with Gateway. Additional Relay Nodes run a supervisor that enrolls through the standard Node identity flow and manages a Relay worker. Gateway distributes signed policy and workload assignments; managed nodes establish authenticated outbound connections to their assigned Relay endpoints.

Relay is transport, not workload ownership. Ingress, Docker, database, monitoring, and builder daemons continue to own their host operations. Relay carries authenticated control and supported private streams without becoming a general network tunnel.

  1. Prepare a dedicated host in the intended physical or network failure domain.
  2. Ensure its advertised address is reachable by participating nodes on TCP 9443.
  3. Open Settings → Relay and choose Add relay node, or create a Relay from Nodes.
  4. Enter the display name and reachable Relay Address.
  5. Run the generated one-time installer command on the target host.
  6. Wait for the enrollment dialog to close after that Node becomes online.
  7. Confirm the Relay instance reports ready before assigning or rebalancing traffic.

A pending Relay Node is not pool capacity. Do not treat a created database record or installed supervisor as ready until Gateway has verified the worker endpoint and health state.

Place members in distinct failure domains when resilience matters. Confirm every managed host can reach every Relay that may be assigned to it. Assignment spread controls how many ready Relays receive new workload connections; increasing spread improves redundancy but consumes additional connections and memory.

Rebalance is explicit. Use it after adding capacity, changing placement, or recovering a member; do not assume existing assignments move automatically merely because a new Relay is online.

  1. Drain the member so new tunnels stop selecting it.
  2. Allow existing streams to finish or disconnect them explicitly when the incident requires it.
  3. Apply the signed supervisor/worker update.
  4. Verify worker readiness, version, and reconnect.
  5. Return the member to service and observe new assignments.
  6. Remove a remote Relay identity only after it is drained and no assignments depend on it.

Removing the Gateway record does not replace normal host decommissioning. Uninstall the supervisor and remove host material through the documented operational process.

For a Relay warning, inspect the local Relay or remote worker process, identity volume, PostgreSQL authorization path, advertised endpoint, listener on 9443, policy freshness, and assigned-node reachability. Preserve logs and identity before restart. After recovery, verify node reconnects, Secure Links, database bindings, active streams, assignment state, and the Dashboard warning.

Pool size alone does not prove resilience. Members must occupy independent failure domains and managed nodes must be able to reach the endpoints they may receive. Test connectivity from representative Ingress, Docker, Databases, Monitoring, and Build Worker networks before increasing assignment spread. A second Relay behind the same host, power domain, firewall, or failed route may add capacity without adding availability.

Observe connection count, active streams, memory, reconnect rate, and admission failures during normal load. Establish enough headroom for a remaining member to accept reassigned nodes when one member is drained or unavailable. Rebalancing during an incident can increase reconnect load, so recover the failed dependency first unless assignment concentration is itself the problem.

If a member update fails, keep it drained and restore the signed known-good supervisor and worker release through the supported update path. Do not copy binaries or identity material from another member. Verify its advertised endpoint and readiness before returning it to assignment.

If the pool control state is healthy but one network cannot connect, correct routing, DNS, firewall, or NAT for that advertised endpoint rather than replacing Relay identities. If the identity volume is lost, treat the member as a replacement enrollment and remove the old record only after its assignments are safe. For a local Relay incident, preserve the configured identity and data volume; recreating the container without them creates a different and unusable transport identity.

After any recovery, test a new managed-node session and a real private path such as a Secure Link or database binding. Existing long-lived streams alone can hide a failure to admit new connections.

For the customer-impact matrix when Gateway, the local Relay, or every Relay is unavailable, continue with Availability, compatibility, and limits.