Summary
Azure readiness admits HTTP server processes before their lazy database initialization completes. A successful warm-up request does not establish readiness across replicas. This caused the main-branch deployment smoke test to fail after PR #472 merged at 4fe935e3dee6c0a7464cda279fbc308cbca728d5.
This is separate from #472's AWS recommendation completeness change. The startup, health, workflow, and Azure compute files traced here are unchanged between reviewed head ed80d762cd269b45482f57296306611d966beb8b and that merge.
Current behavior and evidence
Azure run 37243528487, Test Deployment records:
2026-10-04T23:29:33.116831115Z: the health gate returns healthy, including auth, config, and migrations, after an API warm-up.
2026-10-04T23:29:33.283074595Z: the first smoke request returns degraded, with auth store uninitialized, config connection pending, and migrations not yet run. The interval is approximately 167 ms.
- Deployment succeeded; smoke verification failed. The diagnostic log step was skipped.
Confirmed source behavior:
.github/workflows/deploy-azure.yml:446 warms /api/auth/check-admin; :469 smoke requests only /health and fails on the first degraded response.
internal/server/http.go:31 routes /health without ensureDB; API requests initialize through :162.
internal/server/app.go:485-493,530 creates the uninitialized state. :605-666 initializes it; successful initialization does not reset it to pristine state.
internal/server/health.go:58-64 returns HTTP 200 while degraded.
terraform/modules/compute/azure/container-apps/main.tf:165-173 uses /health for readiness, admitting those processes.
Inference: the second response came from another cold replica or a restarted process. The logs lack replica identity, so they do not distinguish these cases. This is not evidence of a failed migration: that state reports failed, not pending.
Reproduction and expected behavior
Locally, route requests across two independent cold application instances. Warm one through the API, then send /health to the other. The current design admits the second instance and returns degraded HTTP 200. This proposed local reproduction has not yet been executed.
Every instance eligible for traffic must have completed required initialization. Deployment verification must continue rejecting degraded responses.
Proposed fix and verification
Initialize each HTTP-serving process independently of user traffic in internal/server/http.go and the existing initialization path. Add a dedicated readiness handler in internal/server/health.go, returning non-2xx until required checks pass. Preserve liveness separately. Wire Azure readiness and workflow verification to that handler.
Test the two-instance scenario, permanent initialization failures, failed migrations, and bounded smoke-test failure. Returning non-2xx alone needs an initialization trigger; adding sleeps or warm-up requests cannot guarantee every replica initializes.
Related work and severity
#438 covers script tests and failure logs. #31 covers deploy-time migrations. #217 established the healthy-body assertion; preserve it. A fresh open-issue search for readiness OR "cold replica" returned no matches; these related issues do not explicitly cover this readiness contract.
Severity: high. The default-branch Azure deployment is red, and readiness does not represent dependency readiness. Proposed labels, subject to existing-label verification: priority/p0, type/bug, severity/high, urgency/now, impact/many, effort/m, triaged.
Summary
Azure readiness admits HTTP server processes before their lazy database initialization completes. A successful warm-up request does not establish readiness across replicas. This caused the main-branch deployment smoke test to fail after PR #472 merged at
4fe935e3dee6c0a7464cda279fbc308cbca728d5.This is separate from #472's AWS recommendation completeness change. The startup, health, workflow, and Azure compute files traced here are unchanged between reviewed head
ed80d762cd269b45482f57296306611d966beb8band that merge.Current behavior and evidence
Azure run 37243528487, Test Deployment records:
2026-10-04T23:29:33.116831115Z: the health gate returnshealthy, including auth, config, and migrations, after an API warm-up.2026-10-04T23:29:33.283074595Z: the first smoke request returnsdegraded, with auth store uninitialized, config connection pending, and migrations not yet run. The interval is approximately 167 ms.Confirmed source behavior:
.github/workflows/deploy-azure.yml:446warms/api/auth/check-admin;:469smoke requests only/healthand fails on the first degraded response.internal/server/http.go:31routes/healthwithoutensureDB; API requests initialize through:162.internal/server/app.go:485-493,530creates the uninitialized state.:605-666initializes it; successful initialization does not reset it to pristine state.internal/server/health.go:58-64returns HTTP 200 while degraded.terraform/modules/compute/azure/container-apps/main.tf:165-173uses/healthfor readiness, admitting those processes.Inference: the second response came from another cold replica or a restarted process. The logs lack replica identity, so they do not distinguish these cases. This is not evidence of a failed migration: that state reports
failed, notpending.Reproduction and expected behavior
Locally, route requests across two independent cold application instances. Warm one through the API, then send
/healthto the other. The current design admits the second instance and returns degraded HTTP 200. This proposed local reproduction has not yet been executed.Every instance eligible for traffic must have completed required initialization. Deployment verification must continue rejecting degraded responses.
Proposed fix and verification
Initialize each HTTP-serving process independently of user traffic in
internal/server/http.goand the existing initialization path. Add a dedicated readiness handler ininternal/server/health.go, returning non-2xx until required checks pass. Preserve liveness separately. Wire Azure readiness and workflow verification to that handler.Test the two-instance scenario, permanent initialization failures, failed migrations, and bounded smoke-test failure. Returning non-2xx alone needs an initialization trigger; adding sleeps or warm-up requests cannot guarantee every replica initializes.
Related work and severity
#438 covers script tests and failure logs. #31 covers deploy-time migrations. #217 established the healthy-body assertion; preserve it. A fresh open-issue search for
readiness OR "cold replica"returned no matches; these related issues do not explicitly cover this readiness contract.Severity: high. The default-branch Azure deployment is red, and readiness does not represent dependency readiness. Proposed labels, subject to existing-label verification:
priority/p0,type/bug,severity/high,urgency/now,impact/many,effort/m,triaged.