The idea: the toolkit arm's tests were stronger:
- 149 unit tests and 33 Playwright/OQL scripts, all committed.
- Timer compression, which found an escalation bug.
- Fresh data for each variant.
The other arm did three things better, and those are the gaps to close:
- Permission refusal was checked only client-side. Add a server-side probe that expects access denied.
- Unexpected 400/422 responses were treated as informational. Make them failures.
- 20/20 was stitched together from 3 runs. Require one clean-database final run that includes the timers.
Target: testing-shape.md, test-result-audit.md.
Field evidence: contrib/inbox/2026-09-29-acceptance-test-strength.md (#165).
The idea: the toolkit arm's tests were stronger:
The other arm did three things better, and those are the gaps to close:
Target:
testing-shape.md,test-result-audit.md.Field evidence:
contrib/inbox/2026-09-29-acceptance-test-strength.md(#165).