Repository navigation
Conversation
Louis, 2026-10-11: tier C is virtually obsolete, there to keep an unproven test from derailing the code. The testing document's guideline "all new tests begin at tier C" contradicted that and its own C1/C2 sections: a regression test written first and shown to fail, or a test asserting a checked hard baseline, now starts at tier B and gates. CLAUDE.md's tests section says the same, and test_0025 (the -fno-math-errno pin, shown failing before its fix) moves to tier B. Underworld development team with AI support from Claude Code
Member
Author
|
Review (docs and one test marker, before merge):
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Tier C is virtually obsolete: it exists to keep an unproven test from derailing the code (decided 2026-10-11). The testing document's guideline "all new tests begin at tier C" contradicted that. It also contradicted the document's own C1/C2 sections, which call characterisations rare by design and give tier C its purpose: stopping a freshly written test, which may simply be wrong, from steering the implementation.
Changes:
docs/developer/TESTING-RELIABILITY-SYSTEM.md: a new test that is proven starts at tier B and gates. Proven means a regression test written first and shown to fail without its fix, or a test asserting a hard baseline whose own correctness has been checked. Tier C stays for tests not yet proven.CLAUDE.md(Tests): the same rule, in two sentences.tests/test_0025_jit_compile_flags.pymoves to tier B. It is the-fno-math-errnopin from JIT: compile kernels with -fno-math-errno by default (#834) #835, shown failing ondevelopmentbefore its fix.The style charter's tier table and its characterisation paragraph are unchanged; they do not say where new tests start.
Underworld development team with AI support from Claude Code