Skip to content

Ingress troubleshooting

The incident objective is to restore the public service without destroying the state that explains the failure. Ingress crosses several ownership boundaries—DNS, network, certificate, nginx configuration, Route policy, private transport, and application runtime—so changing several layers at once usually increases outage time.

Assign one incident owner, record the last known-good change, and verify from the outside in. Success means external behavior is restored, Gateway desired and reported state agree, and the failed layer and recovery action are recorded well enough to prevent recurrence.

Use this order to avoid changing the wrong layer:

  1. DNS: resolve the hostname from an external client and confirm it points to the assigned ingress node.
  2. Network: confirm public 80/443 reachability and host firewall rules.
  3. TLS: inspect certificate hostname, chain, expiry, and distribution status.
  4. nginx config: verify the latest revision applied and configuration validation succeeded.
  5. Route health: inspect expected status/body and maintenance state.
  6. Upstream: test the application from the nginx node or inspect Secure Link health.
  7. Logs: correlate nginx access/error logs with workload logs and request IDs.

Common mistakes include cross-node Domain/Route placement, stale external DNS after migration, HTTP-01 on a node without public port 80, a WebSocket application without WebSocket forwarding, and selecting a workload port that is not actually published or linked.

Symptom Most likely layer First evidence
Hostname does not resolve DNS External A/AAAA lookup and Domain placement
Connection timeout network Public address, firewall, ports 80/443, node availability
Certificate warning TLS Certificate hostname, chain, expiry, and assigned Route
Immediate 404 Route matching Hostname, path prefix, enabled state, raw config
Managed 503 maintenance or unhealthy upstream Maintenance state and health history
502/504 upstream transport Secure Link, target port, workload health, timeout settings
WebSocket disconnects protocol forwarding WebSocket option, application path, proxy/read timeouts
Changes never appear apply/reconciliation task state, nginx revision, node connection, validation error

Record the failing hostname, path, time, client IP class, expected response, actual response, assigned node, Route ID, latest task ID, and relevant request ID. Preserve the generated-config validation error and bounded logs before retrying; repeated edits can erase the most useful evidence.

Compare three views:

  1. Gateway desired state in the Route and Domain details;
  2. latest acknowledged state from the nginx node;
  3. externally observed DNS, TLS, and HTTP behavior.

A mismatch between these views identifies whether the failure is control-plane state, node apply, or external infrastructure.

  • Correct DNS only after the intended node placement is confirmed.
  • Correct a certificate assignment rather than disabling TLS globally.
  • Fix config validation before forcing another apply.
  • Use maintenance mode when the application must remain unavailable during repair.
  • Retry reconciliation only after the dependency that caused the failure is healthy.
  • Roll back the application or Route configuration when a known-good revision exists.

Do not delete and recreate the Domain, certificate, or Route as a first response. That destroys relationship history and may introduce a second outage.

Test from an external client, then confirm the canonical hostname, certificate chain, status code, body, redirects, WebSockets, health history, and nginx logs. Close the incident only when public behavior and Gateway state agree.

Rollback should reverse the smallest change that introduced the outage. For DNS, restore the recorded address values and account for resolver caches. For TLS, restore the previous valid certificate assignment without disabling HTTPS. For nginx configuration, return to the last known-good managed settings and wait for acknowledgement. For an application regression, roll back the owning Deployment or Compose revision rather than rewriting the Route around a broken release.

Secure Link failures require restoring Relay and Node connectivity or the target runtime; changing DNS or certificates cannot repair them. Access-policy failures require restoring the previous Access List or trusted-proxy interpretation, not opening the workload port. Maintenance mode is useful while repairing these layers because it keeps the hostname and TLS path explicit while preventing accidental traffic to a partially recovered application.

When the first recovery attempt does not explain the failure, collect a bounded evidence bundle before escalating:

  • Domain and Route identifiers, assigned Ingress Node, and latest acknowledged config revision;
  • external DNS answers and TLS certificate details observed at the incident time;
  • Route health history and the exact failing path, method, status, and timestamp;
  • relevant nginx access/error entries and the owning workload or Secure Link operation;
  • Node and Relay versions, connectivity state, and the last failed Task message;
  • the last known-good application, Route, and DNS revision.

Redact credentials, cookies, authorization headers, private keys, and sensitive query values. The goal is to preserve causality without turning an incident record into a secret store.