feat(release): scope 9.1.0 qualification to GPT-6 Sol - #131
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 18e4a05a57
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| echo "::notice::No eval model matrix or provider credentials configured; skipping." | ||
| echo "models=" >> "$GITHUB_OUTPUT" | ||
| else | ||
| FLOW_EVAL_MODEL="$models" bun run scripts/check-release-models.ts |
There was a problem hiding this comment.
Require the OpenAI credential before authorizing the matrix
When only ANTHROPIC_API_KEY is configured, the condition above enters this branch and the new model check accepts the required openai/gpt-6-sol matrix. The workflow then authorizes the dispatch budget and starts the eval even though the run exports no OpenAI credential and explicitly disables credential copying, so the scheduled campaign cannot call its sole provider. Update the preflight to require OPENAI_API_KEY for this version-bound OpenAI-only profile.
Useful? React with 👍 / 👎.
|
Independent current-head verification at All seven GitHub checks pass. Fifty-five focused release-policy, exact-artifact qualification/regrade, patch-baseline, and cancellation tests pass. The local release CLI rejects wrong models, the old two-model grid, narrowed scenarios, and unsupported concurrency before inference. The 9.1.0 plan is 38 primary attempts plus eight reserves; historical and future two-provider behavior remains covered. The branch is clean and mergeable. Note: This verifies the one-provider gate and candidate merge readiness. No paid GPT-6 Sol matrix or exact-artifact canary has run under this policy. Its source and policy digest require a fresh artifact and new paid authorization before a release can be tagged. |
Why
The direct
xai/grok-4.6OAuth route reported a team limit of 0/0 requests per minute during the 9.1.0 release preflight. No release scenario was scored. The operator chose OpenAI-only evidence for this version. The existing two-provider gate correctly refuses that campaign, so this PR makes the reduced scope explicit and version-bound.Scope
For 9.1.0 only, require the canonical direct
openai/gpt-6-solidentity and schedule 38 primary attempts with eight environment reserves. Keep every case's pass rate, review, false-completion, transcript, artifact, canary, and regrade gates. Other versions retain two distinct providers. Release metadata and the scheduled eval workflow use the same versioned rule. The patch-release exception refuses a one-provider baseline. Release guides and notes disclose that 9.1.0 will have no cross-provider reliability claim.Tradeoffs
This release will measure one OpenAI route. It cannot establish behavior on xAI, OpenCode Zen, or other providers. The exception is not reusable for later versions. Delegated Jev recovery remains unqualified and unavailable.
Blast Radius
Evaluation policy, qualification verification, release metadata, the optional model-eval workflow, and release guidance change. The production Flow runtime and Session v5 schema do not change. A new source commit and policy digest require a fresh 9.1.0 artifact and new paid authorization. The failed two-provider probe provides no scored evidence for this path.
Verification
bun run checkpassed with 1,589 tests and no failures. The focused exact-artifact CLI test sealed and independently regraded the 38-cell report. The four reserve-cancellation tests passed with 428 assertions. A forged gateway, family, revision, or variant is rejected.bun run replaymatched all 13 gated cassettes. The packed plugin loaded in a model-free OpenCode 1.18.31 smoke with 22 passing checks. No paid calls ran for this PR.