fix(campaign): preserve artifacts and traces for every attempt - #908
Conversation
tangletools
left a comment
There was a problem hiding this comment.
✅ Auto-approved PR — cd53ce02
Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.
tangletools · auto-approval · reason: blanket_auto_approve · 2026-10-01T17:21:18Z
|
Independent review completed at cd53ce0. Verified all160 retained public operational files (512,603 bytes), five source files and474 built files by SHA256 and size. Source has one execution-attempt path helper, one maintained artifact writer, and the existing stored-cell reader. This is actual SDK filesystem use through the public API. Fetched main and checked merge-tree. |
Change
Published Eval's native retry and resume overwrite earlier consumer artifacts and traces under the same stable cell directory.
Each actual execution now writes its own attempt directory containing identity, artifacts, trace, result and failure receipt.
The mutable latest-attempt pointer makes current readback direct.
Native success caching and scheduling retain their existing behavior.
Receipt tags bind every agent call and judge input to runAttemptId + attemptNumber; retry totals remain cumulative.
The existing optimization reader follows the current attempt.
Historical root evidence stays read-only.
Artifact writes cannot escape their attempt's artifacts directory.
Actual consumer proof
The public /campaign runEval entrypoint reproduced both artifact and trace loss against installed, published 0.203.0.
The candidate public consumer retains both attempts for retry and resume.
Success-cache reuse makes no new callback.
Four existing historical cell files remain byte-identical.
A parent-path write is refused without changing the attempt identity.
Interrupted runOptimization reopens its native ledger and recovers the exact earlier failed candidate and error.
Four explicitly non-inference filesystem meter receipts verify attempt attribution without provider billing claims.
Judge input tags were independently exercised through the public callback.
Evidence: docs/proofs/campaign-attempt-retention-20261001 (160 files, 512,603 bytes).
Exact source and built module hashes, both failed instruments, complete records and execution receipts are retained.
Target: drew-gtr-pro SDK filesystem operations. Zero provider requests, model inference or research units.
Discovery Lab #1115 retains zero confirmation and its original >=200 paired independent source-unit acceptance.
Build, typecheck and package verification passed.
No unit tests added or run.
No package version or publisher changes.
Risk and rollback
Consumers that hardcode stable artifact or trace filenames must use returned paths or the attempt pointer.
Custom trace factories receive the attempt's trace directory.
The sole maintained stored-cell reader is updated.
Historical caches and files remain intact.
Reverting the source change restores prior write behavior while preserving already captured attempt directories.