Production readiness checklist
Use this checklist as an approval record, not as a list one operator silently ticks. The service owner defines business-critical paths, the platform owner proves infrastructure recovery, the security owner approves identity and exposure, and application owners prove data and rollback. Assign every section before the readiness review.
A production-ready installation has known owners, tested recovery, observable customer paths, and evidence that common failures do not require improvising on the host. Passing a build or seeing green cards in the UI is necessary but not sufficient.
Control plane
Section titled “Control plane”- Canonical HTTPS URL, cookies, WebSockets, and OAuth redirects work.
- PostgreSQL, Redis, Relay identity, encryption keys, and uploaded artifacts are persistent and backed up.
- Relay owns public
9443/tcpand recovers after restart. - License state and required entitlements are healthy.
Identity
Section titled “Identity”- Two independent administrator recovery paths exist.
- MFA and group scopes are tested with real non-admin users.
- API/OAuth/MCP credentials are least-privilege and inventoried.
Managed nodes
Section titled “Managed nodes”- Every role reports expected capabilities and update channel.
- Offline behavior, daemon restart, node restart, and reconnect are tested.
- Firewalls allow only documented traffic.
Workloads and data
Section titled “Workloads and data”- Deploy, health check, rollback, logs, backup, and restore are proven.
- Managed database bindings survive Gateway, Relay, daemon, node, and workload recreation tests.
- No owner credentials are present in workload configuration.
Customer traffic
Section titled “Customer traffic”- DNS, TLS renewal, ingress migration, maintenance mode, and upstream failure are tested.
- Notifications and status-page communication reach their intended recipients.
Operations
Section titled “Operations”- Update and rollback runbooks have owners.
- Disk, certificate, database, queue, Relay, and node alerts are active.
- An incident drill has been completed without relying on undocumented shell changes.
Evidence package
Section titled “Evidence package”Do not approve production from visual inspection alone. Preserve evidence for:
- installation version and image digests;
- backup completion and a recent restore test;
- external DNS/TLS/HTTP verification;
- non-admin permission tests;
- Node restart and reconnect tests for every deployed role;
- Gateway, Relay, daemon, and workload restart behavior;
- database binding query results before and after recreation;
- alert open/recovery and notification delivery;
- update and rollback rehearsal;
- unresolved exceptions with owner and expiry.
Go/no-go gate
Section titled “Go/no-go gate”Deployment is no-go when administrator recovery is untested, persistent keys are not backed up, Relay has a single unknown recovery path, node capability errors are ignored, customer traffic cannot be verified externally, or a database binding requires manual container/role repair.
An accepted temporary exception must state the customer impact, compensating control, owner, deadline, and rollback trigger. “Works on the current host” is not a production-readiness argument.
First-day observation
Section titled “First-day observation”After enabling business traffic, watch Route health, TLS, Relay pressure, node freshness, workload restarts, database connections, disk growth, notification queues, and audit activity. Keep the rollback owner available through the observation window and avoid unrelated platform changes.
Readiness review format
Section titled “Readiness review format”Hold the review against a named release and installation. For each area, record pass, accepted exception, or no-go, plus an evidence link and owner. Do not carry evidence forward from an older release when the relevant lifecycle, daemon, Relay, database, or ingress component changed.
Use one row for every acceptance check so a less experienced reviewer can see what was tested and a power user can reproduce it:
| Area | Owner | Check performed | Result | Evidence | Exception or follow-up |
|---|---|---|---|---|---|
| Example: public application | Application team | External DNS, TLS, and primary user path | Pass | Task, request, or report link | None |
Do not write only “works” or “checked.” The evidence should identify the exact resource, version, test location, time, and expected result.
Start with the three journeys that would create immediate business impact: administrator recovery, public customer traffic, and application access to persistent data. Then review supporting capabilities such as builds, Pages, notifications, AI, and optional integrations according to what the installation will actually use. Untested unused features do not block launch; enabled critical features do.
Minimum acceptance scenarios
Section titled “Minimum acceptance scenarios”- A new authorized user signs in, completes MFA, and is denied an out-of-scope resource.
- A supported workload is deployed, observed, restarted, and rolled back without host edits.
- A public Route is verified from an external network with the expected certificate and health behavior.
- A managed Node disconnects and reconnects without losing desired state.
- Gateway restarts while customer workloads remain in their documented state.
- A backup is restored into a clean compatible environment and required secrets decrypt.
- Alerts open and recover during a controlled failure, and the intended recipient receives them.
The approval should state the scope that was tested. Do not claim that the whole product is validated when only one Node role, database engine, runtime profile, or update path was exercised.
