Skip to content

new(bin/context-audit.sh): which files fill the context window, from real transcripts - #181

Merged
MendixMau merged 2 commits into
masterfrom
new/context-audit
Sep 30, 2026
Merged

MendixMau merged 2 commits into
masterfrom
new/context-audit

Conversation

@MendixMau

Copy link
Copy Markdown
Owner

What changed and why (one paragraph)

bin/token-burn.sh says how many tokens a project burned, but not which files burned them. Deciding what to split, trim or stop routing needs that per-file number from real runs, not wc on the skill files. bin/context-audit.sh reads the Claude Code transcripts on a machine and reports: (1) the context size before any work, from the first call's billed input + cache, split into sessions and subagents; (2) the instruction files loaded every run and their size; (3) every file read, by the Read tool or a simple shell read (cat, sed -n, git show REV:path, following cd and VAR=), with reads, sessions, total and per-read size, and re-reads within a session (paging through a file is not a re-read, asking for the same part again is); (4) other tool output by tool; (5) compactions, with the files read before them. Project names are masked by default (project-1/architecture/modules/*.md), so the output can be pasted; --names shows them locally. It is read-only and exits 0, and prints NOT AVAILABLE when there is no transcript tree.

Field evidence

This cloud session, plus one pipeline-start subagent run here (read CLAUDE.local.md, the runbook, two baseline skills, one again):

Measure Value
Subagent context before reading anything 52,503 tokens
Runbook 6 paged reads, 2 of them repeats
learned-mdl-preflight.md 9,846 ≈tok in one read
Instruction files per run project CLAUDE.md 7,342, toolkit CLAUDE.md 6,326, CLAUDE.local.md 2,033
Share of file reads that were toolkit files 90%

Fixture: tests/wave2/test-context-audit.sh bin/context-audit.sh over tests/wave2/fixtures/context-audit/, which is that capture scrubbed. The scrubber keeps every field the script reads, swaps all text for filler of the same length, rewrites paths to /srv/example/ and replaces model ids (see CAPTURE.md). Its assertions were checked by hand against the output.

Checklist

  • No client data (leak guard, denylist grep; fixture text is filler, paths generic)
  • Size cap
  • Test tier: new wave2 fixture from a real capture
  • Instrument rules: golden input captured (rule 1), producer named in the header (rule 5), field run cited (rule 4); read-only, never blocks
  • Routing row: n/a (a bin tool, like token-burn.sh)
  • CHANGELOG line
  • Bug entries: n/a

🤖 Generated with Claude Code

https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw


Generated by Claude Code

…real transcripts

token-burn.sh counts tokens per project; this attributes them to files:
start context per session and per subagent, auto-loaded instruction files,
every file read (Read tool and simple shell reads) with re-reads, other tool
output, and compactions with what was read before them. Project names are
masked by default so the output can be pasted.

Field run: this cloud session plus one pipeline-start subagent captured here;
the subagent starts at 52,503 tokens and pages the runbook 6 times, 2 repeats.
Fixture is that capture, scrubbed to same-length filler.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
@MendixMau
MendixMau merged commit 7c63127 into master Sep 30, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants