Skip to content

Docs: a proven new test starts at tier B; tier C is for unproven tests - #840

Open
lmoresi wants to merge 1 commit into
developmentfrom
docs/test-tier-policy
Open

lmoresi wants to merge 1 commit into
developmentfrom
docs/test-tier-policy

Conversation

@lmoresi

@lmoresi lmoresi commented Oct 10, 2026

Copy link
Copy Markdown
Member

Tier C is virtually obsolete: it exists to keep an unproven test from derailing the code (decided 2026-10-11). The testing document's guideline "all new tests begin at tier C" contradicted that. It also contradicted the document's own C1/C2 sections, which call characterisations rare by design and give tier C its purpose: stopping a freshly written test, which may simply be wrong, from steering the implementation.

Changes:

  • docs/developer/TESTING-RELIABILITY-SYSTEM.md: a new test that is proven starts at tier B and gates. Proven means a regression test written first and shown to fail without its fix, or a test asserting a hard baseline whose own correctness has been checked. Tier C stays for tests not yet proven.
  • CLAUDE.md (Tests): the same rule, in two sentences.
  • tests/test_0025_jit_compile_flags.py moves to tier B. It is the -fno-math-errno pin from JIT: compile kernels with -fno-math-errno by default (#834) #835, shown failing on development before its fix.

The style charter's tier table and its characterisation paragraph are unchanged; they do not say where new tests start.

Underworld development team with AI support from Claude Code

Louis, 2026-10-11: tier C is virtually obsolete, there to keep an unproven test
from derailing the code. The testing document's guideline "all new tests begin
at tier C" contradicted that and its own C1/C2 sections: a regression test
written first and shown to fail, or a test asserting a checked hard baseline,
now starts at tier B and gates. CLAUDE.md's tests section says the same, and
test_0025 (the -fno-math-errno pin, shown failing before its fix) moves to tier B.

Underworld development team with AI support from Claude Code
Copilot AI balanced review requested due to automatic review settings October 10, 2026 21:02

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@lmoresi

lmoresi commented Oct 10, 2026

Copy link
Copy Markdown
Member Author

Review (docs and one test marker, before merge):

  • test_0025 is proven. It failed on development before JIT: compile kernels with -fno-math-errno by default (#834) #835 for the right reason (the flag was missing), passed after it, and asserts the exact argument list. It now gates, since CI runs -m "not tier_c", and it still matches the tests/test_00[0-4]*py glob.
  • Consistency: the guideline now agrees with the document's own C1/C2 sections and with the charter's tier table. Nothing else in docs/ tells a new test to start at tier C (searched for "Start at Tier C" and "begin as experimental").
  • No library code changed.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants