Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 2 additions & 3 deletions .github/workflows/evals.yml
Original file line number Diff line number Diff line change
Expand Up @@ -8,9 +8,7 @@ name: Model evals
# to run locally: every recorded report was single-model, so "works with Flow" meant
# "worked once, with one provider".
#
# Evals cost money and need credentials, so they stay out of the PR gate. This runs
# weekly against at least two providers and publishes the report as an artifact,
# with the qualification thresholds applied by `scripts/qualify-release.ts`.
# Evals cost money and need credentials, so they stay out of the PR gate.
# See docs/adr/0010-declared-canonical-gate.md.

on:
Expand Down Expand Up @@ -80,6 +78,7 @@ jobs:
echo "::notice::No eval model matrix or provider credentials configured; skipping."
echo "models=" >> "$GITHUB_OUTPUT"
else
FLOW_EVAL_MODEL="$models" bun run scripts/check-release-models.ts

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Require the OpenAI credential before authorizing the matrix

When only ANTHROPIC_API_KEY is configured, the condition above enters this branch and the new model check accepts the required openai/gpt-6-sol matrix. The workflow then authorizes the dispatch budget and starts the eval even though the run exports no OpenAI credential and explicitly disables credential copying, so the scheduled campaign cannot call its sole provider. Update the preflight to require OPENAI_API_KEY for this version-bound OpenAI-only profile.

Useful? React with 👍 / 👎.

echo "models=$models" >> "$GITHUB_OUTPUT"
fi

Expand Down
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,8 @@ One short entry per release, written for users deciding whether to upgrade.
file to retain complete validation evidence.
- Session v5 schema, validation and reviewer gates, and continuation routing are
unchanged. The OpenCode host used by the release checks is pinned to 1.18.31.
- This release uses a GPT-6 Sol-only qualification profile. Its evidence makes
no cross-provider reliability claim. Later releases retain the two-provider gate.

Upgrade with `opencode plugin opencode-plugin-flow@9.1.0 --global --force`.

Expand Down
3 changes: 2 additions & 1 deletion docs/development.md
Original file line number Diff line number Diff line change
Expand Up @@ -151,7 +151,8 @@ Follow the [frozen-candidate sequence](release-qualification.md#running-it):
finish fixes and dependency updates, pass deterministic checks, approve paid
evals.
`bun run qualify -- --campaign-dir <dir> --canary <record>` seals the
two-provider campaign, exact-artifact canary and grader evidence. Commit that
policy grid, canary and grader evidence. Version 9.1.0 uses GPT-6 Sol only.
Other versions require two providers. Commit that
bundle before tagging; never substitute interrupted results for qualification.

Release tags use `v<package-version>`. Blocking release checks: the normal
Expand Down
18 changes: 9 additions & 9 deletions docs/release-qualification.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ This page owns release thresholds, candidate freezing, and publication order.

| Threshold | Value | Why |
| --- | --- | --- |
| Distinct providers | ≥ 2 | Every report recorded before this policy was single-model, so "works with Flow" meant "worked once, with one provider". |
| Distinct providers | ≥ 2; 9.1.0 only: 1 | 9.1.0 pins `openai/gpt-6-sol`. Its evidence covers OpenAI only. Later releases require two routes. |
| False completions | 0 | A `completed` closure the document itself contradicts is the failure Flow exists to prevent. |
| Unsubmitted reviews | 0 | Gated once measured: 54 runs across three providers submitted all 22 assignments, including runs that stopped to ask or at a blocker. |
| Scored attempts per provider | 3 at 100%; 10 at 90% | The frozen release plan gives each threshold enough trials to express its allowed failures. |
Expand Down Expand Up @@ -43,8 +43,7 @@ canary inside its window. The wall clock decided this until 9.0.1, which made
every published offline release stop verifying seventy-two hours after its
baseline was measured.

A new scenario needs an explicit release-policy decision. Any required canonical
case missing from the report fails qualification.
New scenarios need a policy decision. Missing required cases fail qualification.

A non-product attempt never shrinks the required sample. The frozen plan retains
one environment reserve per provider and case. A retryable provider or host
Expand All @@ -53,10 +52,10 @@ second external failure or an unallowed ask leaves a gap. Product and evaluator
failures never activate reserves. Evaluator failure is `NOT VERIFIED`;
persistence failure stops without a finalized report.

Repository code owns the ordered release catalog. Persisted `catalog.json` is only a
witness and must match it exactly. The two-provider grid has 76 primary cells and
16 predeclared environment reserves; ordinary, narrowed, dynamically extended, or
merged summary reports cannot qualify.
Repository code owns the ordered release catalog; persisted `catalog.json` must
match. Version 9.1.0 uses 38 primary cells and eight reserves on GPT-6 Sol only.
Other versions use 76 primary cells and 16 reserves on distinct providers.
Narrowed, extended, or merged summary reports cannot qualify.

Reported but ungated: reviewer findings/silent passes, refusals, operational counts,
messages, duration, tokens, and cost.
Expand Down Expand Up @@ -95,16 +94,17 @@ Finish or close active sessions before changing Flow versions in either directio

Finish code, dependency, version and changelog changes first. Pass frozen install,
`bun run check`, `bun run replay`, audit, live smoke and CI before paid qualification.
Freeze packed contents and evaluator inputs, then run the full two-provider matrix
Freeze packed contents and evaluator inputs, then run the versioned matrix
on the canonical Linux host. Run a fresh canary against its exact `artifact.tgz`,
seal/regrade the bundle, and commit only evidence without changing measured inputs.
Recheck final main CI and exact artifact identity before tagging `v<package-version>`.

Authorize dispatches using the [paid-run budget](../.agents/plans/05-release-simplification/README.md#authorize-paid-work).
Keep that ledger across retries. Budget-stopped campaigns cannot qualify.
For 9.1.0, run the pinned OpenAI model.

```bash
bun run eval -- --release --model openai/gpt-6-sol --model xai/grok-4.6
bun run eval -- --release --model openai/gpt-6-sol
bun run eval:canary -- prepare --report <campaign-dir>/report.json --out <canary-dir>
# Run the prepared fixture, then record its session and transcript.
bun run eval:canary -- record <record-options>
Expand Down
22 changes: 11 additions & 11 deletions evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,8 +56,9 @@ plugin tuple configuration, and records the same selection in provenance.
Release sampling rejects reviewer overrides.

Ordinary runs use one sequential queue per model, with up to four queues in flight.
Release mode is strictly sequential (`--concurrency 1`): 76 primary targets and one
environment reserve per provider/case, at most 92 attempts. Only retained retryable
Release mode is strictly sequential (`--concurrency 1`). Version 9.1.0 pins
`openai/gpt-6-sol` with 38 primary targets and eight reserves. Other versions
require two providers, 76 primary targets and 16 reserves. Only retained retryable
host/provider failures activate reserves, never product failures. Results are
persisted in declared order even when ordinary queues finish out of order.

Expand Down Expand Up @@ -410,24 +411,22 @@ a suite that measures nothing look identical from here, so read one run anyway.

## Three tiers, three prices

One price for every question is what made this suite something run at release rather
than during work:
Use three eval tiers:

| Tier | Command | Cost | Answers |
| --- | --- | --- | --- |
| Replay | `bun run replay` | free | does the runtime still reach the same outcome on decisions a model already made? |
| Smoke | `bun run eval:smoke -- --model <id>` | one model, one attempt | did a prompt change break the ordinary path? |
| Matrix | `bun run eval -- --release --model <a> --model <b>` | real money | may this be released? |
| Matrix | `bun run eval -- --release --model <id> [--model <id>]` | real money | may this be released? |

Only the matrix qualifies a release. A replay is evidence about the runtime and none
about the prompts; a single attempt of a stochastic scenario is not a rate.

## Multi-model matrix

Every report recorded before this existed was single-model, so "works with Flow"
meant "worked once, with one provider". Qualification needs at least two distinct
providers, and `.github/workflows/evals.yml` runs the matrix weekly and on demand —
never in a gate a contributor waits on, since a full pass costs real money.
Version 9.1.0 requires only `openai/gpt-6-sol`; it makes no cross-provider claim.
Other versions require two distinct providers. `.github/workflows/evals.yml`
runs the matrix weekly and on demand, outside contributor gates.

## Using evals to change prompts

Expand Down Expand Up @@ -580,8 +579,9 @@ distinction that matters: one pass in six and six in six are different findings.

## Cost

Release qualification schedules 76 primary attempts across eight scenarios and
two providers, plus at most 16 environment reserves. Ordinary campaign size depends
Version 9.1.0 schedules 38 primary attempts and eight reserves on GPT-6 Sol.
Its evidence supports only that route. Other versions schedule 76 primary attempts
and 16 reserves across two providers. Ordinary campaign size depends
on the selected scenarios, models and repeats. Use `--scenario` while iterating;
cost depends on model pricing and the work performed, not just scenario count.

Expand Down
8 changes: 6 additions & 2 deletions evals/qualification-regrade.ts
Original file line number Diff line number Diff line change
Expand Up @@ -27,8 +27,8 @@ import { readQualificationBundle } from "./qualification-bundle.js";
import {
RELEASE_ANALYSIS_SHA256,
RELEASE_MAX_CAMPAIGN_AGE_MS,
RELEASE_POLICY_SHA256,
releaseGraderBundle,
releasePolicySha256,
} from "./release-policy.js";
import type { ArtifactIdentity, ValidatedReport } from "./report.js";
import { SCENARIOS } from "./scenarios.js";
Expand Down Expand Up @@ -310,7 +310,11 @@ export async function regradeQualificationBundle(input: {
),
};
if (
policy.policySha256 !== RELEASE_POLICY_SHA256 ||
policy.policySha256 !==
releasePolicySha256(
(expectedStored as { artifact: ArtifactIdentity }).artifact
.packageVersion,
) ||
policy.analysisSha256 !== RELEASE_ANALYSIS_SHA256 ||
canonicalJson(policy.graderBundle) !== canonicalJson(bundledGrader) ||
canonicalJson(policy.graderBundle) !==
Expand Down
106 changes: 86 additions & 20 deletions evals/release-policy.ts
Original file line number Diff line number Diff line change
Expand Up @@ -98,7 +98,34 @@ const RELEASE_POLICY_INPUT = [

const parsed = parseCaseCatalog(RELEASE_POLICY_INPUT);
if (!parsed.ok) throw new Error("Repository release policy is invalid.");
const RELEASE_CATALOG = parsed.value;
const STANDARD_RELEASE_CATALOG = parsed.value;

export type ReleaseProfile = {
readonly catalog: ValidatedCaseCatalog;
readonly requiredModels: readonly ModelIdentity[] | null;
};

const OPENAI_ONLY_9_1_0: ReleaseProfile = {
catalog: STANDARD_RELEASE_CATALOG.map((row) => ({ ...row, minProviders: 1 })),
requiredModels: [
{
routeProvider: "openai",
gateway: null,
family: "gpt-6-sol",
model: "gpt-6-sol",
revision: null,
},
],
};

const STANDARD_RELEASE: ReleaseProfile = {
catalog: STANDARD_RELEASE_CATALOG,
requiredModels: null,
};

export function releaseProfile(packageVersion: string): ReleaseProfile {
return packageVersion === "9.1.0" ? OPENAI_ONLY_9_1_0 : STANDARD_RELEASE;
}

export const RELEASE_ANALYSIS_SHA256 = canonicalSha256("flow-v2-analysis-v1", {
kind: "rate",
Expand All @@ -113,42 +140,77 @@ export const RELEASE_HOST_POLICY = {
reviewerSteps: null,
} as const;

export const RELEASE_POLICY_SHA256 = canonicalSha256("flow-release-policy-v1", {
catalog: RELEASE_CATALOG,
host: RELEASE_HOST_POLICY,
analysisSha256: RELEASE_ANALYSIS_SHA256,
environmentReservesPerStratum: RELEASE_ENVIRONMENT_RESERVES_PER_STRATUM,
});
export function releasePolicySha256(packageVersion: string): string {
const profile = releaseProfile(packageVersion);
return canonicalSha256("flow-release-policy-v1", {
catalog: profile.catalog,
...(profile.requiredModels === null
? {}
: { requiredModels: profile.requiredModels }),
host: RELEASE_HOST_POLICY,
analysisSha256: RELEASE_ANALYSIS_SHA256,
environmentReservesPerStratum: RELEASE_ENVIRONMENT_RESERVES_PER_STRATUM,
});
}

export const RELEASE_POLICY_SHA256 = releasePolicySha256("standard");

export const RELEASE_POLICY_CATALOG_SHA256 = canonicalSha256(
"flow-evaluator-policy-catalog-v1",
RELEASE_CATALOG,
STANDARD_RELEASE_CATALOG,
);

export function releaseCatalog(): ValidatedCaseCatalog {
return RELEASE_CATALOG;
export function releaseCatalog(
packageVersion = "standard",
): ValidatedCaseCatalog {
return releaseProfile(packageVersion).catalog;
}

export function releaseCaseIds(): readonly string[] {
return RELEASE_CATALOG.map((policy) => policy.caseId);
return STANDARD_RELEASE_CATALOG.map((policy) => policy.caseId);
}

export function releaseAttemptsFor(caseId: string): number {
const policy = RELEASE_CATALOG.find((item) => item.caseId === caseId);
const policy = STANDARD_RELEASE_CATALOG.find(
(item) => item.caseId === caseId,
);
if (!policy) throw new Error(`No release policy for ${caseId}.`);
return policy.minScoredAttempts;
}

export function releaseMinimumProviders(): number {
return Math.max(...RELEASE_CATALOG.map((policy) => policy.minProviders));
export function releaseMinimumProviders(packageVersion = "standard"): number {
return Math.max(
...releaseCatalog(packageVersion).map((policy) => policy.minProviders),
);
}

export function assertReleaseModels(
models: readonly ModelIdentity[],
packageVersion: string,
): void {
const profile = releaseProfile(packageVersion);
const minimum = releaseMinimumProviders(packageVersion);
if (
models.length !== minimum ||
new Set(models.map((model) => model.routeProvider)).size !== minimum ||
(profile.requiredModels !== null &&
canonicalJson(models) !== canonicalJson(profile.requiredModels))
) {
throw new Error(
profile.requiredModels !== null
? `Release ${packageVersion} requires exactly ${profile.requiredModels.map((model) => `${model.routeProvider}/${model.model}`).join(", ")} with its canonical direct-route identity.`
: `Release requires exactly ${minimum} models on distinct route providers.`,
);
}
}

export function releasePrimaryCellsFor(
models: readonly ModelIdentity[],
packageVersion = "standard",
): ScheduledCell[] {
let slot = 0;
return models.flatMap((model) =>
RELEASE_CATALOG.flatMap((policy) =>
releaseCatalog(packageVersion).flatMap((policy) =>
Array.from({ length: policy.minScoredAttempts }, (_, repetition) => {
const block = slot;
slot += 1;
Expand All @@ -175,10 +237,11 @@ export function releasePrimaryCellsFor(

export function releaseCellsFor(
models: readonly ModelIdentity[],
packageVersion = "standard",
): ScheduledCell[] {
const primary = releasePrimaryCellsFor(models);
const primary = releasePrimaryCellsFor(models, packageVersion);
const reserves = models.flatMap((model) =>
RELEASE_CATALOG.map((policy) => {
releaseCatalog(packageVersion).map((policy) => {
const identity = canonicalSha256("flow-v2-environment-reserve-v1", {
model: `${model.routeProvider}/${model.model}`,
scenario: policy.caseId,
Expand All @@ -204,11 +267,12 @@ export function releaseCellsFor(

export function releaseRandomizationSeed(
models: readonly ModelIdentity[],
packageVersion = "standard",
): string {
return canonicalSha256("flow-v2-seed-v1", {
models: models.map((model) => `${model.routeProvider}/${model.model}`),
scenarios: releaseCaseIds(),
releasePolicySha256: RELEASE_POLICY_SHA256,
releasePolicySha256: releasePolicySha256(packageVersion),
});
}

Expand Down Expand Up @@ -260,17 +324,19 @@ export function assertReleaseScenarioOrder(

export function assertExactReleaseCatalog(
input: unknown,
packageVersion = "standard",
): ValidatedCaseCatalog {
const supplied = parseCaseCatalog(input);
const catalog = releaseCatalog(packageVersion);
if (
!supplied.ok ||
canonicalJson(supplied.value) !== canonicalJson(RELEASE_CATALOG)
canonicalJson(supplied.value) !== canonicalJson(catalog)
) {
throw new Error(
"Persisted catalog does not match repository release policy.",
);
}
return RELEASE_CATALOG;
return catalog;
}

export function selectReleaseScenarios<
Expand Down
Loading
Loading