Skip to content

Enumerate all three TEST_CMD substitution sites in the Step 4 row - #151

Merged
dmccoystephenson merged 1 commit into
mainfrom
fix/test-cmd-row-enumeration
Sep 27, 2026
Merged

dmccoystephenson merged 1 commit into
mainfrom
fix/test-cmd-row-enumeration

Conversation

@dmccoystephenson

Copy link
Copy Markdown
Member

Summary

  • The Step 4 TEST_CMD row is corrected to enumerate all three places {{TEST_CMD}} appears in the template body. PR Require unique scratch filenames and define the structurally-red-anchor path #121 added a third occurrence (the Phase 3 "Confirm the anchor actually executed tests" rule, create-dev-loop.md:297) without revisiting the row, which still named only the Phase 3 verify block and the Phase 8 rebase fence.
  • The same Phase 3 rule is added to the list of test-naming passages that Step 4's no-test-suite note asks the generator to reword. Without that, a repo whose TEST_CMD is a shell comment would receive a rule reading "an executed-test count from # re-run the manual checklist" — the issue asked for the missing site to be checked against that clause.
  • The other multi-site rows were checked at the same time and are accurate: COMPILE_CMD ("both places") has two occurrences, LINT_CMD ("both") has two sites. Only TEST_CMD had drifted.

Closes #139

Research grounding

No RESEARCH.md finding applies. This is a documentation-accuracy correction to the substitution table with no empirical claim behind it — the issue body says the same.

Doc sync check

  • Every {{placeholder}} added or changed has a corresponding Step 4 substitution-table row — no placeholders were added or removed; python3 scripts/check_docs.py passes.

README.md is unaffected (no Step was added, removed, or renamed). No finding is shipped, so no RESEARCH.md entry is needed.

Test plan

  • python3 scripts/check_docs.py — "Doc consistency check passed."
  • python3 -m unittest discover -s tests -v — Ran 13 tests, OK
  • Diff inspected for literal backslash-quotes after inserting quoted text — none present
  • /create-dev-loop behavioral run — not performed. The change edits Step 4 instruction prose only (a parenthetical and a rewording list); it adds no placeholder, changes no template-body text, and touches no Step 3/5/6 path. The text a generated skill receives is unchanged for repos with a test suite, and for repos without one the change only widens which passages the generator is told to reword. Reviewers should treat the behavioral half as UNVERIFIED in the ordinary sense for this repo — the doc-consistency CI job covers mechanical parity only.

Skipped issues (recorded per Phase 1)

The rest of the open backlog was deferred this cycle for the following reasons:

This PR description was drafted during a Gardener session (https://github.com/Stephenson-Software/gardener).

🤖 Generated with Claude Code


drafted by Claude on behalf of Daniel Stephenson

PR #121 added a third `{{TEST_CMD}}` occurrence to the template (the Phase 3
"Confirm the anchor actually executed tests" rule) without updating the
Step 4 row, which still listed only the Phase 3 verify block and the Phase 8
rebase fence. Name the third site in the row, and add that rule to the list
of test-naming passages Step 4 rewords when the repo has no test suite, so
the shell-comment substitution does not land in prose that reads as an
"executed-test count from" a comment.

Closes #139

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@dmccoystephenson

Copy link
Copy Markdown
Member Author

Self-review rubric (posted as a plain comment; no independent reviewer is implied):

Universal items

  • Scope: PASS — git diff --name-only origin/main...HEAD lists only create-dev-loop.md; two lines changed, both named in Step 4's TEST_CMD row enumerates two substitution sites, but the template has three #139.
  • Tests-new: no signal — no function or public surface was added; scripts/check_docs.py is untouched.
  • Tests-fix: FAIL (structural, not a defect in this PR) — the fix is to instruction prose in the Step 4 table, which no automated test can reach. check_docs.py verifies that a row exists per placeholder, not that a row's prose enumerates every site, so no stash-and-run experiment is possible. The evidence that the old text was wrong is the occurrence count itself: the template body contains three {{TEST_CMD}} tokens (lines 290, 297, 581) and the row named two.
  • Sibling structure: no signal — no new file.
  • Sibling renames: no signal — no identifier renamed.
  • Docs: PASS — a Phase 7 pass was run against the implementation: README.md has no Step change; SECURITY.md:31-35 lists the sites where commands are executed (Phase 3 verify, Phase 4 anchor, Phase 8 rebase) and the newly enumerated site is a prose mention, not an execution, so that list is still accurate; RESEARCH.md, CONTRIBUTING.md, the PR template, and the issue templates do not restate the row; CLAUDE.md is unaffected.
  • Issue resolution: PASS — Step 4's TEST_CMD row enumerates two substitution sites, but the template has three #139 asks for the third site to be added to the parenthetical and for it to be checked against the no-test-suite clause; both are in the diff. Its "confirm whether other multi-site rows drifted" request was checked (COMPILE_CMD: 2 occurrences vs "both places"; LINT_CMD: 2 sites vs "both") and recorded in the PR body.
  • CI: PASS — doc-consistency green on head c163fd1.

Repo-specific items

  • Placeholder parity: PASS — no placeholder added or removed; python3 scripts/check_docs.py → "Doc consistency check passed."
  • README Step parity: PASS — no Step added, removed, or renamed.
  • Fence escaping: PASS — no fence added; grep -n '^```' create-dev-loop.md shows only the outer fences (123, 606) inside the template range.
  • Phase numbering: PASS — no phase renumbered; the diff references existing Phase 3 and Phase 8 by their current numbers.
  • Version comments: no signal — the generated skill's header is untouched.
  • Research grounding: PASS — the PR body states that no finding applies, matching the issue.
  • No back-ported specifics: PASS — the diff introduces no build command, reviewer name, or branch prefix.
  • gardener interface: PASS — Steps 3, 5, and 6 untouched.
  • Retrofit note: no signal — no rule promoted into the template body; the change is to Step 4 generator instructions.
  • Behavioral validation: UNVERIFIED — no /create-dev-loop run was performed (headless dispatch; Step 6 creates a repo and Steps 3/5 write outside the checkout). The change alters no text a generated skill receives for repos with a test suite; for repos without one it only widens the list of passages the generator is told to reword. Green doc-consistency is not read as behavioral verification.

Out-of-diff observation: the COMPILE_CMD and LINT_CMD rows say "both places"/"both" rather than naming the sites, so they will drift silently the same way TEST_CMD did if a later PR adds an occurrence. Left as-is — rewording them is outside #139's scope — but a check_docs.py case that compares each row's stated site count to the template's occurrence count would make this class of drift mechanical. Not filed as an issue this cycle; noted for a future Stage B pass.

Merge-readiness judgment: the Tests-fix FAIL is the structural no-test condition this repo's skill documents for prose changes, not missing regression coverage for a code fix; Behavioral validation is UNVERIFIED on a change that alters no generated-skill text for the common case. Both are stated in the PR body.

This review was drafted during a Gardener session (https://github.com/Stephenson-Software/gardener).


drafted by Claude on behalf of Daniel Stephenson

@dmccoystephenson

Copy link
Copy Markdown
Member Author

Merge-readiness re-assessment (continuation of an interrupted dispatch):

  • Branch state: head c163fd1, up to date with main (a61398d); doc-consistency green; locally python3 scripts/check_docs.py → "Doc consistency check passed." and python3 -m unittest discover -s tests → Ran 13 tests, OK.
  • Occurrence count re-confirmed: {{TEST_CMD}} appears at create-dev-loop.md:290, :297, :581 — the three sites the corrected row now names.
  • Do-not-auto-merge paths: none matched (only create-dev-loop.md, Step 4 prose; Steps 3/5/6 untouched).

The PR is left open rather than merged. The skill's own merge bar is not met: the change edits create-dev-loop.md's Steps, so the Behavioral-validation rubric item requires a named /create-dev-loop run or UNVERIFIED + no auto-merge, and the Phase 8 regression gate requires either a test or a recorded manual validation run for a correction issue. Neither could be produced in a headless session (Step 6 creates a GitHub repo; Steps 3/5 write outside the checkout). The structural cause is already tracked in #109; the related site-count drift observation from the self-review is tracked in #152.

A maintainer merge is expected to be low-risk: the diff changes no text a generated skill receives for repos with a test suite.

This comment was drafted during a Gardener session (https://github.com/Stephenson-Software/gardener).


drafted by Claude on behalf of Daniel Stephenson

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Step 4's TEST_CMD row enumerates two substitution sites, but the template has three

1 participant