Skip to content

fix(server/azure): exclude uninitialized replicas from readiness #488

Description

@cristim

Summary

Azure readiness admits HTTP server processes before their lazy database initialization completes. A successful warm-up request does not establish readiness across replicas. This caused the main-branch deployment smoke test to fail after PR #472 merged at 4fe935e3dee6c0a7464cda279fbc308cbca728d5.

This is separate from #472's AWS recommendation completeness change. The startup, health, workflow, and Azure compute files traced here are unchanged between reviewed head ed80d762cd269b45482f57296306611d966beb8b and that merge.

Current behavior and evidence

Azure run 37243528487, Test Deployment records:

  • 2026-10-04T23:29:33.116831115Z: the health gate returns healthy, including auth, config, and migrations, after an API warm-up.
  • 2026-10-04T23:29:33.283074595Z: the first smoke request returns degraded, with auth store uninitialized, config connection pending, and migrations not yet run. The interval is approximately 167 ms.
  • Deployment succeeded; smoke verification failed. The diagnostic log step was skipped.

Confirmed source behavior:

  • .github/workflows/deploy-azure.yml:446 warms /api/auth/check-admin; :469 smoke requests only /health and fails on the first degraded response.
  • internal/server/http.go:31 routes /health without ensureDB; API requests initialize through :162.
  • internal/server/app.go:485-493,530 creates the uninitialized state. :605-666 initializes it; successful initialization does not reset it to pristine state.
  • internal/server/health.go:58-64 returns HTTP 200 while degraded.
  • terraform/modules/compute/azure/container-apps/main.tf:165-173 uses /health for readiness, admitting those processes.

Inference: the second response came from another cold replica or a restarted process. The logs lack replica identity, so they do not distinguish these cases. This is not evidence of a failed migration: that state reports failed, not pending.

Reproduction and expected behavior

Locally, route requests across two independent cold application instances. Warm one through the API, then send /health to the other. The current design admits the second instance and returns degraded HTTP 200. This proposed local reproduction has not yet been executed.

Every instance eligible for traffic must have completed required initialization. Deployment verification must continue rejecting degraded responses.

Proposed fix and verification

Initialize each HTTP-serving process independently of user traffic in internal/server/http.go and the existing initialization path. Add a dedicated readiness handler in internal/server/health.go, returning non-2xx until required checks pass. Preserve liveness separately. Wire Azure readiness and workflow verification to that handler.

Test the two-instance scenario, permanent initialization failures, failed migrations, and bounded smoke-test failure. Returning non-2xx alone needs an initialization trigger; adding sleeps or warm-up requests cannot guarantee every replica initializes.

Related work and severity

#438 covers script tests and failure logs. #31 covers deploy-time migrations. #217 established the healthy-body assertion; preserve it. A fresh open-issue search for readiness OR "cold replica" returned no matches; these related issues do not explicitly cover this readiness contract.

Severity: high. The default-branch Azure deployment is red, and readiness does not represent dependency readiness. Proposed labels, subject to existing-label verification: priority/p0, type/bug, severity/high, urgency/now, impact/many, effort/m, triaged.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions