diff --git a/.claude/CLAUDE.md b/.claude/CLAUDE.md index 44934dcf8..34e8bee6f 100644 --- a/.claude/CLAUDE.md +++ b/.claude/CLAUDE.md @@ -11,43 +11,24 @@ Rust library for NP-hard problem reductions. Implements computational problems w - **Tests over implementation.** Spend more time designing tests than implementing code. Well-designed tests catch bugs early and document intended behavior. -## Agent Preferences -- For tasks with complex logic (implementing reductions, fixing CI, issue-to-pr, check-issue, etc.), prefer invoking the `codex` CLI (if installed) with the latest model: `codex exec -c model="gpt-5.4" ""`. - ## Skills -These repo-local skills live under `.claude/skills/*/SKILL.md`. - -- [run-pipeline](skills/run-pipeline/SKILL.md) -- Pick a Ready issue from the GitHub Project board, move it through In Progress -> issue-to-pr -> Review pool. One issue at a time; forever-loop handles iteration. -- [issue-to-pr](skills/issue-to-pr/SKILL.md) -- Convert a GitHub issue into a PR with an implementation plan. Default rule: one item per PR. Exception: a `[Model]` issue that explicitly claims direct ILP solvability should implement the model and its direct ` -> ILP` rule together; `[Rule]` issues still require both models to exist on `main`. -- [add-model](skills/add-model/SKILL.md) -- Add a new problem model. Can be used standalone (brainstorms with user) or called from `issue-to-pr`. -- [add-rule](skills/add-rule/SKILL.md) -- Add a new reduction rule. Runs mathematical verification by default (via `/verify-reduction`); pass `--no-verify` to skip for trivial reductions. Can be used standalone or called from `issue-to-pr`. -- [review-structural](skills/review-structural/SKILL.md) -- Project-specific structural completeness check: model/rule checklists, build, semantic correctness, issue compliance. Read-only, no code changes. Called by `review-pipeline`. -- [review-quality](skills/review-quality/SKILL.md) -- Generic code quality review: DRY, KISS, cohesion/coupling, test quality, HCI. Read-only, no code changes. Called by `review-pipeline`. -- [fix-pr](skills/fix-pr/SKILL.md) -- Resolve PR review comments, fix CI failures, and address codecov coverage gaps. Uses `gh api` for codecov (not local `cargo-llvm-cov`). -- [write-model-in-paper](skills/write-model-in-paper/SKILL.md) -- Write or improve a problem-def entry in the Typst paper (standalone, for improving existing entries). Core instructions are inlined in `add-model` Step 6. -- [write-rule-in-paper](skills/write-rule-in-paper/SKILL.md) -- Write or improve a reduction-rule entry in the Typst paper (standalone, for improving existing entries). Core instructions are inlined in `add-rule` Step 6. -- [release](skills/release/SKILL.md) -- Create a new crate release. Determines version bump from diff, verifies tests/clippy, then runs `make release`. -- [check-issue](skills/check-issue/SKILL.md) -- Quality gate for `[Rule]` and `[Model]` issues. Checks usefulness, non-triviality, correctness of literature, and writing quality. Posts structured report and adds failure labels. -- [fix-issue](skills/fix-issue/SKILL.md) -- Fix quality issues found by check-issue — auto-fixes mechanical problems, brainstorms substantive issues with human, then re-checks and moves to Ready. -- [topology-sanity-check](skills/topology-sanity-check/SKILL.md) -- Run sanity checks on the reduction graph: detect orphan (isolated) problems and redundant reduction rules. - - `topology-sanity-check orphans` -- Detect isolated problem types (runs `examples/detect_isolated_problems.rs`) - - `topology-sanity-check np-hardness` -- Verify NP-hardness proof chains from 3-SAT (runs `examples/detect_unreachable_from_3sat.rs`) - - `topology-sanity-check redundancy [source target]` -- Check for dominated reduction rules -- [review-pipeline](skills/review-pipeline/SKILL.md) -- Agentic review for PRs in Review pool: runs structural check, quality check, and agentic feature tests (no code changes), posts combined verdict, always moves to Final review. -- [propose](skills/propose/SKILL.md) -- Interactive brainstorming to help domain experts propose a new model or rule. Asks one question at a time, uses mathematical language (no programming jargon), and files a GitHub issue. -- [final-review](skills/final-review/SKILL.md) -- Interactive maintainer review for PRs in "Final review" column. Merges main, walks through agentic review bullets with human, then merge or hold. -- [dev-setup](skills/dev-setup/SKILL.md) -- Interactive wizard to install and configure all development tools for new maintainers. -- [verify-reduction](skills/verify-reduction/SKILL.md) -- Standalone mathematical verification of a reduction rule: Typst proof, constructor Python (≥5000 checks), adversary Python (≥5000 independent checks). Reports verdict, no artifacts saved. Also called as a subroutine by `/add-rule` (default behavior). -- [update-papers](skills/update-papers/SKILL.md) -- Update research paper collection: download new papers from references.bib, retry failed downloads, sync to Google Drive, regenerate index.md. -- [find-solver](skills/find-solver/SKILL.md) -- Interactive guide: match a real-world problem to a library model, explore reduction paths, recommend solvers (built-in + external), and generate a solution doc. -- [find-problem](skills/find-problem/SKILL.md) -- Reverse of find-solver: given a solver for a model, discover what other problems it can handle via incoming reductions, ranked by effective complexity. - -## Codex Compatibility -- Claude slash commands such as `/issue-to-pr 42 --execute` are aliases for the matching repo-local skill files under `.claude/skills/`. -- In Codex, read the relevant `SKILL.md` directly and follow it; do not assume slash-command support exists. -- The Makefile targets `run-plan`, `run-issue`, `run-pipeline`, and `run-review` already translate these workflows into explicit `SKILL.md` prompts for Codex. -- The default Codex model in the Makefile is `gpt-5.4`. Override it with `CODEX_MODEL=` if needed. -- The Step 0/Step 1 packet builders under `scripts/pipeline_skill_context.py` and `scripts/pipeline_checks.py` are expensive GitHub-backed calls. Per top-level skill invocation, generate each packet at most once and reuse the resulting text/JSON for all later steps unless the skill explicitly requires a fresh rerun. +Repo-local skills live under `.claude/skills/*/SKILL.md`; any agent can read and follow them. There is no project board or pipeline: agents pick up the `how-to-*` guides automatically while working, and humans merge. + +Guides (auto-invoked): +- [how-to-code](skills/how-to-code/SKILL.md) -- Implement or modify a problem model or reduction rule. +- [how-to-verify](skills/how-to-verify/SKILL.md) -- Certify a reduction (type gate, Typst proof, constructor + adversary checks, PR verification certificate) and check reduction-graph topology. +- [how-to-write-manual](skills/how-to-write-manual/SKILL.md) -- Write or audit Typst manual entries (`docs/paper/reductions.typ`) and mdBook docs. +- [how-to-review](skills/how-to-review/SKILL.md) -- Fresh-context PR review: structural, quality, and `pred` feature test; posts an Agentic Review Report. +- [how-to-triage-issue](skills/how-to-triage-issue/SKILL.md) -- Quality-check `[Model]`/`[Rule]` issues, label them, fix mechanical problems, discuss substantive ones. +- [how-to-ship](skills/how-to-ship/SKILL.md) -- Issue to merge-ready PR: branch, gates, review subagent, CI/comments/codecov fixes. One item per PR, except a `[Model]` claiming direct ILP solvability ships its ` -> ILP` rule too. + +Tools (invoked on request): +- [propose](skills/propose/SKILL.md) -- Help a domain expert turn an idea into a well-formed model/rule issue. +- [find-solver](skills/find-solver/SKILL.md) -- Match a real-world problem to a model, route, and solver; writes a doc to `docs/solutions/`. +- [find-problem](skills/find-problem/SKILL.md) -- Given a solver for a model, find source problems it handles via incoming reductions. +- [dev-setup](skills/dev-setup/SKILL.md) -- Install and configure development tools. +- [release](skills/release/SKILL.md) -- Guarded crate release via `make release`. +- [update-papers](skills/update-papers/SKILL.md) -- Refresh the research paper collection. ## Commands ```bash @@ -73,22 +54,12 @@ make cli # Build the pred CLI tool (without MCP, fast) make mcp # Build the pred CLI tool with MCP server support make cli-demo # Run closed-loop CLI demo (exercises all commands) make mcp-test # Run MCP server tests (unit + integration) -make run-plan # Execute a plan with Codex or Claude -make run-issue N=42 # Run issue-to-pr --execute for a GitHub issue -make run-pipeline # Pick next Ready issue from project board, implement, move to Review pool -make run-pipeline N=97 # Process a specific issue from the project board -make run-pipeline-forever # Drain eligible Ready issues forever; poll only while idle -make run-review # Pick next PR from Review pool column, run agentic review, move to Final review -make run-review N=570 # Process a specific PR from the Review pool column -make run-review-forever # Drain eligible Review pool PRs forever; poll only while idle make copilot-review # (Optional) Request Copilot code review on current PR make release V=x.y.z # Tag and push a new release (CI publishes to crates.io) make papers # Full paper fetch: lookup + download + scihub make papers-status # Show research paper collection stats make papers-push # Push PDFs to shared remote (requires rclone + PAPERS_REMOTE) make papers-pull # Pull PDFs from shared remote -# Set RUNNER=claude to use Claude instead of Codex (default: codex) -# Default Codex model: CODEX_MODEL=gpt-5.4 # Set PAPERS_REMOTE=gdrive:folder for paper sync (requires rclone) ``` @@ -116,7 +87,6 @@ make papers-pull # Pull PDFs from shared remote - `tests/main.rs` - Integration tests (modules in `tests/suites/`); example tests use `include!` for direct invocation (no subprocess) - `tests/data/` - Ground truth JSON for integration tests - `scripts/` - Python test data generation scripts (managed with `uv`) -- `docs/plans/` - Implementation plans ### Trait Hierarchy @@ -212,7 +182,7 @@ Reduction graph nodes use variant key-value pairs from `Problem::variant()`: - Same-name variant relations are explicit `#[reduction]` registrations - Each primitive reduction is determined by the exact `(source_variant, target_variant)` endpoint pair - Reduction edges carry `EdgeCapabilities { witness, aggregate, turing }`; graph search defaults to witness mode, aggregate mode is available through `ReductionMode::Aggregate`, and Turing (multi-query) mode via `ReductionMode::Turing` -- `#[reduction]` requires one `transform = exact`, `transform = upper_bound`, or `transform = unavailable` declaration and currently registers witness/config reductions; aggregate-only and Turing edges require manual `ReductionEntry` registration +- `#[reduction]` requires one `transform = exact`, `transform = upper_bound`, or `transform = unavailable` declaration and registers witness/config reductions; aggregate value mappings register through `#[aggregate_reduction]` / `register_aggregate_reduction!` (see Key Patterns); Turing edges (generated by `register_decision_variant!`) and metadata-only edges without an executor use a manual `ReductionEntry` - `Decision

→ P` supports both mappings: compare the exact optimum to the bound, and recover a witness only if it meets the bound. `P → Decision

` is a Turing edge (binary search over decision bound). ### Extension Points @@ -222,10 +192,10 @@ Reduction graph nodes use variant key-value pairs from `Problem::variant()`: - **Each construction input has one name and one concrete type per variant.** Do not add compatibility aliases or infer types from flag names. `CreateSpec` field names render as `snake_case → kebab-case` in CLI and remain `snake_case` in MCP. Add a reusable codec only for a genuinely new transport representation, never a model-name parser branch. - **Random generation is optional and variant-owned.** Not every model has a useful, well-defined random-instance distribution. Add `RandomGenerate` only when the generator has clear semantics and a concrete use (for example, testing or examples); never invent arbitrary bounds or distributions merely to make every model support `--random`. Implement it beside the model (normally through `impl_random_generate!` and a typed `CreateSpec` input DTO), then add `random` only to the applicable `declare_variants!` entries. CLI and MCP discover the exact variant's inputs and callback; never add a model-name random dispatch or advertise random generation on an unsupported variant. - **Decision variants** of optimization problems use `Decision

` wrapper. Add via: (1) `decision_problem_meta!` for the inner type, (2) inherent methods on `Decision`, (3) `register_decision_variant!` with `dims`, `fields`, `parameter_getters`. The generated construction spec accepts flat inner fields plus `bound`; persisted JSON remains `{inner: {...}, bound}`. `Decision

` delegates canonical parameters to `P`; its objective bound is semantic instance data, not a problem parameter. -- Aggregate-only and Turing reduction edges still need manual `ReductionEntry` wiring because `#[reduction]` only registers solution-mapping reductions today; this edge capability does not imply that a problem may solve successfully without a `Solution` +- Aggregate value mappings register with `#[aggregate_reduction]` (or `register_aggregate_reduction!` for generic results), not manual `ReductionEntry` wiring. Manual `ReductionEntry` remains only for Turing edges and metadata-only edges without an executor. An aggregate edge does not imply that a problem may solve successfully without a `Solution` - Exact registry dispatch lives in `src/registry/`; alias resolution and partial/default variant resolution live in `problemreductions-cli/src/problem_name.rs` - `pred create` schema-driven dispatch lives in `problemreductions-cli/src/commands/create.rs` (`create_schema_driven()`) -- Canonical model examples live in `src/example_db/model_builders.rs`; rule examples live beside their rules and are collected by `src/rules/mod.rs` +- Canonical model examples live beside each model in `canonical_model_example_specs()` (collected by `src/example_db/model_builders.rs`); rule examples live beside their rules and are collected by `src/rules/mod.rs` ## Conventions @@ -250,7 +220,7 @@ fields to issue templates. Changes to issue templates require user approval. ### File Naming - Reduction files: `src/rules/_.rs` (e.g., `maximumindependentset_qubo.rs`) - Model files: `src/models//.rs` — category is by input structure: `graph/` (graph input), `formula/` (boolean formula/circuit), `set/` (universe + subsets), `algebraic/` (matrix/linear system/lattice), `misc/` (other) -- Canonical examples: model builders in `src/example_db/model_builders.rs`; rule-local `canonical_rule_example_specs()` functions collected by `src/rules/mod.rs` +- Canonical examples: model-local `canonical_model_example_specs()` functions collected by `src/example_db/model_builders.rs`; rule-local `canonical_rule_example_specs()` functions collected by `src/rules/mod.rs` - Example binaries in `examples/`: utility/export tools and pedagogical demos only (not per-reduction files) - Test naming: `test__to__closed_loop` diff --git a/.claude/skills/add-model/SKILL.md b/.claude/skills/add-model/SKILL.md deleted file mode 100644 index 6cf14b947..000000000 --- a/.claude/skills/add-model/SKILL.md +++ /dev/null @@ -1,336 +0,0 @@ ---- -name: add-model -description: Use when adding a new problem model to the codebase, either from an issue or interactively ---- - -# Add Model - -Step-by-step guide for adding a new problem model to the codebase. - -## Step 0: Gather Required Information - -Before any implementation, collect all required information. If called from `issue-to-pr`, the issue should already provide these. If used standalone, brainstorm with the user to fill in every item below. - -### Required Information Checklist - -| # | Item | Description | Example | -|---|------|-------------|---------| -| 1 | **Problem name** | Struct name with optimization prefix | `MaximumClique`, `MinimumDominatingSet` | -| 2 | **Mathematical definition** | Formal definition with objective/constraints | "Given graph G=(V,E), find max-weight subset S where all pairs in S are adjacent" | -| 3 | **Problem type** | Objective (`Max`/`Min`), witness (`bool`), or aggregate-only (`Sum`/`And`/custom `Aggregate`) | Objective (Maximize) | -| 4 | **Type parameters** | Graph type `G`, weight type `W`, or other | `G: Graph`, `W: WeightElement` | -| 5 | **Struct fields** | What the struct holds | `graph: G`, `weights: Vec` | -| 6 | **Configuration space** | What `dims()` returns | `vec![2; num_vertices]` for binary vertex selection | -| 7 | **Feasibility check** | How to validate a configuration | "All selected vertices must be pairwise adjacent" | -| 8 | **Per-configuration value** | How `evaluate()` computes the aggregate contribution | "Return `Max(Some(total_weight))` for feasible configs" | -| 9 | **Best known exact algorithm** | Complexity with variable definitions | "O(1.1996^n) by Xiao & Nagamochi (2017), where n = \|V\|" | -| 10 | **Solving strategy** | How it can be solved | "BruteForce works; ILP reduction available" | -| 11 | **Category** | Which sub-module under `src/models/` | `graph`, `formula`, `set`, `algebraic`, `misc` | -| 12 | **Expected outcome from the issue** | Concrete outcome for the issue's example instance | Objective: one optimal solution + optimal value. Witness: one valid/satisfying solution + why it is valid. Aggregate-only: the final aggregate value and how it is derived | - -If any item is missing, ask the user to provide it. Do NOT proceed until the checklist is complete. - -The issue's **Expected Outcome** section is the source of truth for the implementation-facing example. -- For optimization problems, use the issue's optimal solution and optimal objective value. -- For satisfaction problems, use the issue's valid / satisfying solution and its justification. -- Do not invent or replace the expected outcome during implementation unless the issue is corrected first. - -### Associated Rule Check - -Before implementation, verify that at least one reduction rule exists or is planned for this problem — otherwise it will be an orphan node in the reduction graph. - -**Check both directions:** - -1. **Outbound (this issue → rule issues):** Look for rule issue numbers in the model issue's "Reduction Rule Crossref" section. -2. **Inbound (rule issues → this problem):** Search open rule issues that reference this problem as source or target: - ```bash - gh issue list --label rule --state open --limit 500 --json number,title | \ - jq '[.[] | select(.title | test(""; "i"))]' - ``` - -**If no associated rules are found:** -- Warn the user: "This model has no associated rule issues. It will be an orphan node in the reduction graph and will be flagged during review." -- Ask whether to proceed anyway or file a companion rule issue first (via `/propose rule`). -- If proceeding, add a visible `` comment in the PR description. - -**If associated rules are found:** List them and continue. - -**If the issue explicitly claims ILP solvability in "How to solve":** -- One associated rule MUST be a direct `[Rule] to ILP` -- Treat that direct ILP rule as part of the same implementation scope -- Do NOT split the model and its direct ILP rule into separate PRs - -## Reference Implementations - -Read these first to understand the patterns: -- **Optimization problem:** `src/models/graph/maximum_independent_set.rs` -- **Satisfaction problem:** `src/models/formula/sat.rs` -- **Model tests:** `src/unit_tests/models/graph/maximum_independent_set.rs` -- **Trait definitions / aggregate types:** `src/traits.rs` (`Problem`), `src/types.rs` (`Aggregate`, `Max`, `Min`, `Sum`, `Or`, `And`, `Extremum`) -- **Registry dispatch boundary:** `src/registry/mod.rs`, `src/registry/variant.rs` -- **CLI and MCP construction:** discovered from the model's registry entry; no frontend model-name dispatch -- **Canonical model examples:** `src/example_db/model_builders.rs` - -## Pre-review Checklist - -Before implementing, make sure the plan explicitly covers these items that structural review checks later: -- Follow `docs/src/design.md#numeric-types-and-arithmetic`: `usize` is for in-memory indices, collection lengths, and brute-force dimensions; registered problem size parameters use `u64`; signed mathematical integers use `i64`; Boolean data uses `bool`; and approximate real or rational data uses finite `f64`. Use another format only when required by the mathematical problem or schema, such as `BigUint` for arbitrary-precision problems or `One` for unit weights; implementation convenience is not sufficient, and there is no `i32` model or I/O format. Implementation-local values are outside this contract. -- Keep failure phases explicit: fallible constructors, create specs, serde-facing validation, and random generation return `ConstructionError`; `evaluate()` returns `EvaluationError`; no public model path returns `Result<_, String>`. Stored-field arithmetic and evaluation arithmetic are checked and reported in their own phase. -- Serde/CLI construction uses the same validation as `new`/`try_new`, and boundary tests cover the supported maximum without requiring impractical allocation. -- `ProblemSchemaEntry` metadata is complete (`display_name`, `aliases`, `dimensions`, explicit `category`, and construction `fields`) -- `Problem::Value` uses the correct aggregate wrapper and witness support is intentional -- `declare_variants!` is present with exactly one `default` variant when multiple concrete variants exist -- CLI discovery and `pred create ` support are included where applicable -- A canonical model example is registered for example-db / `pred create --example` -- If the issue explicitly claims direct ILP solving, the plan also includes the direct ` -> ILP` rule with exact overhead metadata, feature-gated registration, strong regression tests, and ILP-enabled verification -- `docs/paper/reductions.typ` adds both the display-name dictionary entry and the `problem-def(...)` - -## Step 1: Determine the category - -Choose the appropriate sub-module under `src/models/`: -- `graph/` -- problems defined on graphs (vertex/edge selection, SpinGlass, etc.) -- `formula/` -- logical formulas and circuits (SAT, k-SAT, CircuitSAT) -- `set/` -- set-based problems (set packing, set cover) -- `algebraic/` -- matrices, linear systems, lattices (QUBO, ILP, CVP, BMF) -- `misc/` -- unique input structures that don't fit other categories (BinPacking, PaintShop, Factoring) - -Declare the same structural choice explicitly in `ProblemSchemaEntry.category`. This is required metadata and is never inferred from `module_path!()` or the file location. - -## Step 1.5: Infer problem size getters - -From the **best known exact algorithm** complexity (item 9), infer what problem size getter methods the struct should expose. The variables used in the complexity expression define the natural size metrics. - -**How to infer:** -- Parse the complexity expression for variable names (e.g., `O(1.1996^n)` where `n = |V|` → `num_vertices`) -- Each variable that measures a distinct dimension of the input becomes a getter method -- Common mappings: - - `n = |V|` → `num_vertices()` - - `m = |E|` → `num_edges()` - - `n` (number of variables) → `num_vars()` - - `m` (number of clauses) → `num_clauses()` - - `k` (number of sets) → `num_sets()` - -These getters are used by the overhead system for reduction overhead expressions. Implement them as inherent methods on the struct. - -## Step 2: Implement the model - -Create `src/models//.rs`: - -```rust -// Required structure: -// 1. inventory::submit! for ProblemSchemaEntry -// 2. Struct definition with #[derive(Debug, Clone, Serialize, Deserialize)] -// 3. Constructor (new) + accessor methods -// 4. Problem trait impl (NAME, Value, dims, evaluate, variant) -// 5. #[cfg(test)] #[path = "..."] mod tests; -``` - -Key decisions: -- **Schema metadata:** `ProblemSchemaEntry` must include the explicit structural `category` and reflect the construction interface through `display_name`, `aliases`, `dimensions`, and `fields` -- **Objective problems:** use `type Value = Max<_>`, `Min<_>`, or `Extremum<_>` when the model should expose optimization-style witness helpers -- **Witness problems:** use `type Value = Or` for existential feasibility problems -- **Aggregate-only problems:** use a value-only aggregate such as `Sum<_>`, `And`, or a custom `Aggregate` when witnesses are not meaningful -- **Weight management:** use inherent methods (`weights()`, `set_weights()`, `is_weighted()`), NOT traits -- **`dims()`:** returns the configuration space dimensions (e.g., `vec![2; n]` for binary variables) -- **`evaluate()`:** must return `Result`. Invalid configurations remain the aggregate's invalid/false contribution; arithmetic overflow and non-finite computed values are errors. -- **`variant()`:** use the `variant_params!` macro — e.g., `crate::variant_params![G, W]` for `Problem`, or `crate::variant_params![]` for problems with no type parameters. Each type parameter must implement `VariantParam` (already done for standard types like `SimpleGraph`, `i64`, `One`). See `src/variant.rs`. -- **Solve surface:** `Solver::solve()` always computes the aggregate value. `pred solve problem.json` prints a `Solution` only when a witness exists; `pred solve bundle.json` and `--solver ilp` remain witness-only workflows - -## Step 2.5: Register variant complexity - -Add `declare_variants!` at the bottom of the model file (after the trait impls, before the test link). Each line declares a concrete type instantiation with its best-known worst-case complexity: - -```rust -crate::declare_variants! { - ProblemName => "1.1996^num_vertices", - default ProblemName => "1.1996^num_vertices", -} -``` - -- Mark exactly one concrete variant `default` when the problem has multiple registered variants -- The complexity string references the getter method names from Step 1.5 (e.g., `num_vertices`) — variable names are validated at compile time against actual getters, so typos cause compile errors -- One entry per supported `(graph, weight)` combination -- The string is parsed as an `Expr` AST — supports `+`, `-`, `*`, `/`, `^`, `exp()`, `log()`, `sqrt()` -- Use only concrete numeric values (e.g., `"1.1996^num_vertices"`, not `"(2-epsilon)^num_vertices"`) -- A compiled `complexity_eval_fn` plus registry-backed load/serialize/solve dispatch metadata are auto-generated alongside the symbolic expression -- See `src/models/graph/maximum_independent_set.rs` for the reference pattern - -`declare_variants!` now handles objective, witness-capable, and aggregate-only models uniformly. Use manual `VariantEntry` wiring only for unusual dynamic-registration work, not for ordinary models. - -## Step 3: Register the model - -Update these files to register the new problem type: - -1. `src/models//mod.rs` -- add `pub(crate) mod ;` and `pub use ::;` -2. `src/models/mod.rs` -- add to the appropriate re-export line -3. `src/lib.rs` or `prelude` -- if the type should be in `prelude::*`, add it there - -## Step 4: Register for CLI discovery - -The CLI now loads, serializes, and brute-force solves problems through the core registry. Do **not** add manual match arms in `problemreductions-cli/src/dispatch.rs`. - -1. **Registry-backed dispatch comes from `declare_variants!`:** - - Make sure every concrete variant you want the CLI to load is listed in `declare_variants!` - - Mark the intended default variant with `default` when applicable - - Declare well-established problem aliases in `ProblemSchemaEntry.aliases` and variant-specific aliases in `declare_variants!`; CLI and MCP discover both from the registry - -## Step 4.5: Add construction support - -CLI and MCP construction are registry-driven. Do not edit either frontend to recognize a model name. - -1. If user-facing inputs exactly match persisted JSON fields, do nothing. The ordinary `declare_variants!` entry uses `ProblemSchemaEntry.fields` as required construction inputs and deserializes the model directly. - -2. If construction has derived fields, renamed inputs, defaults depending on other inputs, or a composite value assembled from multiple inputs, define a model-local DTO with `#[derive(Deserialize, CreateSpec)]`. Its named fields are the complete public construction contract. Use `Option` only for genuinely optional inputs, doc comments for help text, and `#[create(codec = "...")]` when the transport syntax cannot be inferred from the Rust type. Set `ProblemSchemaEntry.fields` to `LocalCreateSpec::FIELDS` so the catalog and executable constructor share the derived metadata. - -3. Implement `TryFrom for Model` with `ConstructionError`. The direct constructor and serde path must share validation. Use `Conversion` for contract violations, `IntegerOverflow` for stored integer arithmetic, and `NonFiniteFloat` for stored non-finite values; do not call a panicking constructor. - -4. Register the spec on each applicable variant: `default Model => "..." create LocalCreateSpec`. Both frontends then discover the inputs automatically and serialize the constructed typed model back to canonical persisted JSON. - -5. A new reusable external syntax may add one transport codec. It must dispatch by codec/type, never by canonical model name. Unknown or missing inputs are rejected by the core construction contract. - -### Optional random generation - -Random generation is an optional model capability, not a model-completeness requirement. Many models do not have a natural or useful probability distribution over instances; leave random generation unregistered for those models. Do not invent arbitrary size limits, value ranges, or distributions merely to make `--random` available. - -When the model does have a well-defined generator with a concrete testing or example use, random generation is registry-driven and belongs beside the model. Do not edit CLI or MCP dispatch code. - -1. Define a typed random input DTO with `#[derive(Deserialize, CreateSpec)]`, or reuse a matching shared spec from `crate::random`. -2. Implement `RandomGenerate` with `crate::impl_random_generate!(ConcreteModel, RandomSpec, |spec| { ... })`. Validate values and return `Result`; do not round, clamp, or silently replace invalid inputs. -3. Add `random` only to the exact `declare_variants!` entries that implement the trait: `default Model => "..." create LocalCreateSpec random`. -4. The generated problem must have the same canonical name and variant as the selected registry entry. Use the concrete variant's actual graph and numeric types instead of attaching requested metadata to a different concrete instance. - -## Step 4.6: Add canonical model example to example_db - -Add a builder function in `src/example_db/model_builders.rs` that constructs a small, canonical instance for this model. Register it in `build_model_examples()`. - -Also add `canonical_model_example_specs()` **in the model file itself** (gated by `#[cfg(feature = "example-db")]`), and register it in the category `mod.rs` example chain (e.g., `specs.extend(::canonical_model_example_specs());`). See any existing model in `src/models/graph/` for the pattern. - -This example is now the canonical source for: -- `pred create --example ` -- paper/example exports via `load-model-example()` in `reductions.typ` -- example-db invariants tested in `src/unit_tests/example_db.rs` - -## Step 4.7: Implement Direct ILP Rule When Claimed - -If the issue explicitly says the model is solvable by reducing **directly** to ILP, implement `src/rules/_ilp.rs` in the **same PR** as the model. This is the one exception to the normal "one item per PR" policy: the direct ` -> ILP` rule is part of the model feature, not optional follow-up work. - -Completeness bar: -- Feature-gate the rule under `ilp-solver` and register it normally -- Add exact overhead expressions and any required size-field getters; metadata must match the constructed ILP exactly -- Add strong tests in `src/unit_tests/rules/_ilp.rs`: structure/metadata, closed-loop semantics vs the source problem or brute force, extraction, `solve_reduced()` or ILP path coverage when appropriate, and weighted/infeasible/pathological regressions whenever the model semantics admit them -- Update CLI/example-db/paper paths so the claimed ILP solver route is actually usable and documented -- Verify with ILP-enabled workspace commands, not just non-ILP unit tests - -A direct ILP rule shipped with a model issue must match the completeness bar of a standalone production ILP reduction. Do not add a stub just to satisfy the issue text. - -## Step 5: Write unit tests - -Create `src/unit_tests/models//.rs`: - -Every model needs **at least 3 test functions** (the structural reviewer enforces this). Choose from the coverage areas below — pick whichever are relevant to the model: - -- **Creation/basic** — exercise constructor inputs, key accessors, `dims()` / `num_variables()`. -- **Evaluation** — valid and invalid configs so the feasibility boundary or aggregate contribution is explicit. -- **Direction / sense** — verify runtime optimization sense only for models that use `Extremum<_>`. -- **Solver** — brute-force `solve()` returns the correct aggregate value; if witnesses are supported, verify `find_witness()` / `find_all_witnesses()` as well. -- **Serialization** — round-trip serde (when the model is used in CLI/example-db flows). -- **Paper example** — verify the worked example from the paper entry (see below). - -If Step 4.7 applies, also add a dedicated ILP rule test file under `src/unit_tests/rules/_ilp.rs`. Use strong direct-to-ILP reductions in the repo as the reference bar: the tests should validate the actual formulation semantics, not just that an ILP file exists. - -When you add `test__paper_example`, it should: -1. Construct the same instance shown in the paper's example figure -2. Evaluate the solution from the issue's **Expected Outcome** section as shown in the paper and assert it is valid (and optimal for optimization problems) -3. Use `BruteForce` to confirm the claimed optimum/satisfying solution count when the instance is small enough for unit tests - -This test is usually written **after** Step 6 (paper entry), once the example instance and expected outcome are finalized. If writing tests before the paper, use the issue's Example Instance + Expected Outcome as the source of truth and come back to verify consistency. - -Link the test file via `#[cfg(test)] #[path = "..."] mod tests;` at the bottom of the model file. - -## Step 6: Document in paper - -Write a `problem-def` entry in `docs/paper/reductions.typ`. **Reference example:** search for `problem-def("MaximumIndependentSet")` to see the gold-standard entry — use it as a template. - -### 6a. Register display name - -Add to the `display-name` dictionary near the top of `reductions.typ`: -```typst -"ProblemName": [Display Name], -``` - -### 6b. Write formal definition (`def` parameter) - -```typst -#problem-def("ProblemName")[ - Given [inputs with domains], find [solution] [maximizing/minimizing] [objective] such that [constraints]. -][ -``` -Requirements: introduce all inputs first, state the objective, define all notation before use. - -### 6c. Write body (background + example) - -The body goes AFTER auto-generated sections (complexity table, reductions, schema). Four parts: - -**Background (1-3 sentences):** Historical context, applications, structural properties. - -**Best known algorithms:** Integrate naturally into prose with citations. Every complexity claim MUST have `@citation`. If best known is brute-force, add `#footnote[No algorithm improving on brute-force is known for ...]`. - -**Example with visualization:** A concrete small instance with a CeTZ diagram. For graph problems, use `g-node()` and `g-edge()` helpers — see the MaximumIndependentSet entry. Highlight solution with `graph-colors.at(0)`. - -**Evaluation:** Show the objective/verifier computed on the example solution (can be woven into example text). - -**Reproducibility:** The example section must include a `pred-commands()` call showing the create/solve/evaluate pipeline. The `pred create --example ...` spec must be derived from the loaded canonical example data via the helper pattern in `write-model-in-paper`; do not hand-write a bare alias and assume the default variant matches. - -### 6d. Build and verify - -```bash -make paper # Must compile without errors -``` - -Checklist: display name registered, notation self-contained, background present, algorithms cited, example with diagram present, evaluation shown, paper compiles. - -## Step 7: Verify - -For ordinary model-only work: -```bash -make test clippy # Must pass -``` - -If Step 4.7 applied, run ILP-enabled workspace verification instead: -```bash -cargo clippy --all-targets --features ilp-highs -- -D warnings -cargo test --features "ilp-highs example-db" --workspace --verbose -``` - -Structural and quality review is handled by the `review-pipeline` stage, not here. The run stage just needs to produce working code. - -## Naming Conventions - -- Struct names use explicit optimization prefixes: `MaximumX`, `MinimumX` -- No prefix for problems without clear min/max direction: `QUBO`, `Satisfiability`, `KColoring` -- File names use snake_case: `maximum_independent_set.rs` -- See CLAUDE.md "Problem Names" section for the full list - -## Common Mistakes - -| Mistake | Fix | -|---------|-----| -| Implementing weight management as a trait | Use inherent methods: `weights()`, `set_weights()`, `is_weighted()` | -| Forgetting `inventory::submit!` | Every problem needs a `ProblemSchemaEntry` registration | -| Omitting or inferring the model category | Set the required `ProblemSchemaEntry.category` explicitly to one of `Algebraic`, `Formula`, `Graph`, `Misc`, or `Set`; never parse `module_path!()`. | -| Missing `#[path]` test link | Add `#[cfg(test)] #[path = "..."] mod tests;` at file bottom | -| Wrong `dims()` | Must match the actual configuration space (e.g., `vec![2; n]` for binary) | -| Using the wrong aggregate wrapper | Objective models use `Max` / `Min` / `Extremum`, witness models use `bool`, aggregate-only models use a fold value like `Sum` / `And` | -| Not registering in `mod.rs` | Must update both `/mod.rs` and `models/mod.rs` | -| Forgetting `declare_variants!` | Required for variant complexity metadata and registry-backed load/serialize/solve dispatch | -| Wrong aggregate wrapper | Use `Max` / `Min` / `Extremum` for objective problems, `Or` for existential witness problems, and `Sum` / `And` (or a custom aggregate) for value-only folds | -| Wrong `declare_variants!` syntax | Entries no longer use `opt` / `sat`; one entry per problem may be marked `default` | -| Adding aliases in CLI code | Declare problem aliases in `ProblemSchemaEntry.aliases` and variant aliases in `declare_variants!` | -| Adding a hand-written decision model | Use `Decision

` wrapper instead — see `decision_problem_meta!` + `register_decision_variant!` in `src/models/graph/minimum_vertex_cover.rs` for the pattern | -| Inventing short aliases | Only use well-established literature abbreviations (MIS, SAT, TSP); do NOT invent new ones | -| Adding frontend model-name branches | Construction is model-owned. Use a local `CreateSpec` and register it with `declare_variants!`; CLI and MCP must discover it. | -| Hand-maintaining custom construction fields twice | Derive `CreateSpec`, use `LocalCreateSpec::FIELDS` in `ProblemSchemaEntry`, and register the same type in `declare_variants!`. | -| Calling a panicking constructor from `TryFrom` | Share a fallible constructor and preserve its `ConstructionError`. | -| Missing canonical model example | Add a builder in `src/example_db/model_builders.rs` and keep it aligned with paper/example workflows | -| Paper example not tested | Must include `test__paper_example` that verifies the exact instance, solution, and solution count shown in the paper | -| Claiming direct ILP solving but leaving ` -> ILP` for later | If the issue promises a direct ILP path, implement that rule in the same PR with exact overhead metadata and production-level ILP tests | diff --git a/.claude/skills/add-rule/SKILL.md b/.claude/skills/add-rule/SKILL.md deleted file mode 100644 index 3097be526..000000000 --- a/.claude/skills/add-rule/SKILL.md +++ /dev/null @@ -1,313 +0,0 @@ ---- -name: add-rule -description: Use when adding a new reduction rule to the codebase, either from an issue or interactively ---- - -# Add Rule - -Step-by-step guide for adding a new reduction rule (A -> B) to the codebase. By default, every rule goes through mathematical verification (via `/verify-reduction`) before implementation. Pass `--no-verify` to skip verification for trivial reductions. - -## Invocation - -``` -/add-rule # interactive, with verification (default) -/add-rule --no-verify # interactive, skip verification -``` - -When called from `/issue-to-pr`, the `--no-verify` flag is passed through if present. - -## Step 0: Gather Required Information - -Before any implementation, collect all required information. If called from `issue-to-pr`, the issue should already provide these. If used standalone, brainstorm with the user to fill in every item below. - -### Required Information Checklist - -| # | Item | Description | Example | -|---|------|-------------|---------| -| 1 | **Source problem** | The problem being reduced FROM (must already exist) | `MinimumVertexCover` | -| 2 | **Target problem** | The problem being reduced TO (must already exist) | `MaximumIndependentSet` | -| 3 | **Reduction algorithm** | How to transform source instance to target | "Copy graph and weights; IS on same graph as VC" | -| 4 | **Solution extraction** | How to map target solution back to source | "Complement: `1 - x` for each variable" | -| 5 | **Correctness argument** | Why the reduction preserves optimality | "S is independent set iff V\S is vertex cover" | -| 6 | **Parameter transform** | How target size relates to source size | `num_vertices = "num_vertices", num_edges = "num_edges"` | -| 7 | **Concrete example** | A small worked-out instance (tutorial style, clear intuition) | "Triangle graph: VC={0,1} -> IS={2}" | -| 8 | **Solving strategy** | How to solve the target problem | "BruteForce, or existing ILP reduction" | -| 9 | **Reference** | Paper, textbook, or URL for the reduction | URL or citation | - -If any item is missing, ask the user to provide it. Put a high standard on item 7 (concrete example): it must be in tutorial style with clear intuition and easy to understand. Do NOT proceed until the checklist is complete. - -## Step 0.5: Type Compatibility Gate - -Check source/target `Value` types before any work: - -```bash -grep "type Value = " src/models/*/.rs src/models/*/.rs -``` - -**Compatible pairs for `ReduceTo` (witness-capable):** -- `Or`->`Or`, `Min`->`Min`, `Max`->`Max` (same type) -- `Or`->`Min`, `Or`->`Max` (feasibility embeds into optimization) - -**Incompatible — STOP if any of these:** -- `Min`->`Or` or `Max`->`Or` — optimization source has no threshold K; needs a decision-variant source model -- `Max`->`Min` or `Min`->`Max` — opposite optimization directions; needs `ReduceToAggregate` or a decision-variant wrapper -- `Or`->`Sum` or `Min`->`Sum` — Sum is aggregate-only; needs `ReduceToAggregate` -- Any pair involving `And` or `Sum` on the target side - -If incompatible, STOP and comment on the issue explaining the type mismatch and options. Do NOT proceed. - -## Numeric Safety Gate - -Read `docs/src/design.md#numeric-types-and-arithmetic`. Derive implementation -types, supported ranges, and checked conversions from the mathematical source, -target, and reduction algorithm. Use `usize` for in-memory indices, collection -lengths, and brute-force dimensions; `u64` for registered problem size -parameters; `i64` for signed mathematical integers; `bool` for Boolean data; -and finite `f64` for real or rational data. Another format needs mathematical -or target-schema justification; there is no `i32` boundary format. -Temporary reduction calculations are outside this format contract, but fields -written into the target must use the target model's format. - -Ask the contributor only when a mathematical domain or constraint is ambiguous; -do not ask them to choose Rust types. Do not use `as` for range/sign changes. -Check target-size arithmetic and auxiliary identifiers before constructing the -target, verify serde/CLI uses the same ranges, and add focused boundary tests. -The public reduction returns `ReductionError`: preserve a target constructor's -`ConstructionError` as `ReductionError::Construction`, and report reduction -arithmetic directly as the corresponding `ReductionError`; do not stringify or -silently handle either error. Convert model-derived `i64` values to `f64` only -through `i64_to_exact_f64`. - -## Reference Implementations - -Read these first to understand the patterns: -- **Reduction rule:** `src/rules/minimumvertexcover_maximumindependentset.rs` -- **Reduction tests:** `src/unit_tests/rules/minimumvertexcover_maximumindependentset.rs` -- **Paper entry:** search `docs/paper/reductions.typ` for `MinimumVertexCover` `MaximumIndependentSet` -- **Traits:** `src/rules/traits.rs` (`ReduceTo`, `ReduceToAggregate`, `ReductionResult`, `AggregateReductionResult`) - -## Step 1: Mathematical Verification (default, skip with `--no-verify`) - -**If `--no-verify` was passed, skip to Step 2.** - -Invoke the `/verify-reduction` skill to mathematically verify the reduction before writing Rust code. This runs the full verification pipeline: Typst proof, constructor Python script (>=5000 checks), adversary subagent (>=5000 independent checks), and cross-comparison. - -All verification artifacts are ephemeral — they exist only in conversation context and temp files. Nothing is committed to the repository. - -**If verification FAILS: STOP. Report to user. Do NOT proceed to implementation.** - -If verification passes, the verified Python `reduce()` and `extract_solution()` functions, along with the YES/NO instances, carry forward in conversation context to inform Steps 2-5. Use them as the canonical spec for the Rust implementation. - -## Step 2: Implement the reduction - -Create `src/rules/_.rs` (all lowercase, no underscores between words within a problem name): - -```rust -// Required structure: -// 1. ReductionResult struct (holds the target problem + mapping state) -// 2. ReductionResult trait impl (target_problem + extract_solution) -// 3. #[reduction(overhead = { ... })] on ReduceTo impl -// 4. ReduceTo trait impl (reduce_to method) -// 5. #[cfg(test)] #[path = "..."] mod tests; -``` - -Key elements: - -**ReductionResult struct:** -```rust -#[derive(Debug, Clone)] -pub struct ReductionXToY { - target: TargetType, - // any additional mapping state needed for extract_solution -} -``` - -**ReductionResult trait impl:** -```rust -impl ReductionResult for ReductionXToY { - type Source = SourceType; - type Target = TargetType; - fn target_problem(&self) -> &Self::Target { &self.target } - fn extract_solution( - &self, - target_solution: &[usize], - ) -> crate::rules::ExtractionResult> { - crate::rules::traits::validate_target_solution(self.target_problem(), target_solution)?; - let source_solution = /* translate the verified mathematical mapping exactly */; - Ok(source_solution) - } -} -``` - -Every direct extractor must call `validate_target_solution()` once before decoding. It checks only length and value domains, not feasibility, optimality, or rule-specific structure; reject malformed structure with `ExtractionError`. - -**ReduceTo with `#[reduction]` macro** (overhead is **required**): -```rust -#[reduction(overhead = { - field_name = "source_field", -})] -impl ReduceTo for SourceType { - type Result = ReductionXToY; - fn reduce_to(&self) -> Self::Result { - // If Step 1 ran: translate the verified Python reduce() logic - } -} -``` - -Each primitive reduction is determined by the exact source/target variant pair. Keep one primitive registration per endpoint pair and use only the `overhead` form of `#[reduction]`. - -**Aggregate-only reductions:** when the rule preserves aggregate values but cannot recover a source witness from a target witness, implement `AggregateReductionResult` + `ReduceToAggregate` instead of `ReductionResult` + `ReduceTo`. Those edges are not auto-registered by `#[reduction]` yet; register them manually with `ReductionEntry { reduce_aggregate_fn: ..., capabilities: EdgeCapabilities::aggregate_only(), ... }`. See `src/unit_tests/rules/traits.rs` and `src/unit_tests/rules/graph.rs` for the reference pattern. - -## Step 3: Register in mod.rs - -Add to `src/rules/mod.rs`: -- `mod _;` -- If feature-gated (e.g., ILP): wrap with `#[cfg(feature = "ilp-solver")]` - -## Step 4: Write unit tests - -Create `src/unit_tests/rules/_.rs`: - -**Required: closed-loop test** (`test__to__closed_loop`): -```rust -// 1. Create source problem instance -// 2. Reduce: let reduction = ReduceTo::::reduce_to(&source); -// 3. Solve target: solver.find_all_witnesses(reduction.target_problem()) -// 4. Extract: reduction.extract_solution(&target_sol) -// 5. Verify: extracted solution is valid and optimal for source -``` - -If Step 1 ran, use the verified YES/NO instances from conversation context to construct test cases. Include both a feasible (closed-loop) and infeasible (no witnesses) test. - -Additional recommended tests: -- Verify target problem structure (correct size, edges, constraints) -- Edge cases (empty graph, single vertex, etc.) -- Weight preservation (if applicable) - -Test every malformed representation distinguished by the decoder (for example, zero or multiple one-hot selections, or duplicate permutation entries). The canonical example supplies shared wrong-length and out-of-domain tests. - -For aggregate-only reductions, replace the closed-loop witness test with value-chain tests: -- Solve the target with `Solver::solve()` -- Map the aggregate value back with `extract_value()` -- If testing a path, use `ReductionGraph::reduce_aggregate_along_path(...)` - -Link via `#[cfg(test)] #[path = "..."] mod tests;` at the bottom of the rule file. - -## Step 5: Add canonical example - -Define `canonical_rule_example_specs()` in the rule module and include it from `src/rules/mod.rs::canonical_rule_example_specs()`. This enrolls the rule in shared round-trip, wrong-length, and out-of-domain extraction tests. - -## Step 6: Document in paper (MANDATORY — DO NOT SKIP) - -**This step is NOT optional.** Every reduction rule MUST have a corresponding `reduction-rule` entry in the paper. Skipping documentation is a blocking error — the PR will be rejected in review. Do not proceed to Step 6 until the paper entry is written and `make paper` compiles. - -Write a `reduction-rule` entry in `docs/paper/reductions.typ`. **Reference example:** search for `reduction-rule("KColoring", "QUBO"` to see the gold-standard entry — use it as a template. For a minimal example, see MinimumVertexCover -> MaximumIndependentSet. - -If Step 1 ran, adapt the verified Typst proof into the paper's macros. Do not rewrite the proof from scratch — reformat it. - -### 6a. Write theorem body (rule statement) - -```typst -#reduction-rule("Source", "Target", - example: true, - example-caption: [Description ($n = ...$, $|E| = ...$)], -)[ - This $O(...)$ reduction @citation constructs [target structure] ... ($n k$ variables indexed by ...). -] -``` - -Three parts: complexity with citation, construction summary, overhead hint. - -### 6b. Write proof body - -Use these subsections with italic labels: - -```typst -][ - _Construction._ [Full mathematical construction — enough detail to reimplement] - - _Correctness._ ($arrow.r.double$) If ... ($arrow.l.double$) If ... - - _Variable mapping._ [Only if non-trivial mapping] - - _Solution extraction._ [How to convert target solution back to source] -] -``` - -Must be self-contained (all notation defined) and reproducible. - -### 6c. Write worked example (extra block) - -Step-by-step walkthrough with concrete numbers from JSON data. Required steps: -1. Show source instance (dimensions, structure, graph visualization if applicable) -2. Walk through construction with intermediate values -3. Verify a concrete solution end-to-end -4. Witness semantics: state that the fixture stores one canonical witness; if multiplicity matters mathematically, explain it from the construction rather than from `solutions.len()` - -Use `graph-colors`, `g-node()`, `g-edge()` for graph visualization — see reference examples. - -**Reproducibility:** The `extra:` block must start with a `pred-commands()` call showing the create/reduce/solve/evaluate pipeline. The source-side `pred create --example ...` spec must be derived from the loaded canonical example data via the helper pattern in `write-rule-in-paper`; do not hand-write a bare alias and assume the default variant matches. - -### 6d. Build and verify - -```bash -make paper # Must compile without errors -``` - -Checklist: notation self-contained, complexity cited, overhead consistent, example uses JSON data (not hardcoded), solution verified end-to-end, witness semantics respected, paper compiles. - -## Step 7: Regenerate exports and verify - -```bash -cargo run --example export_graph # Generate reduction_graph.json for docs/paper builds -cargo run --example export_schemas # Generate problem schemas for docs/paper builds -cargo run --features "example-db" --example export_examples -make test clippy # Must pass -``` - -`export_examples` refreshes the gitignored `docs/paper/data/examples.json` used by the paper. - -Structural and quality review is handled by the `review-pipeline` stage, not here. The run stage just needs to produce working code. - -## Solver Rules - -- If the target problem already has a solver, use it directly. -- If the solving strategy requires ILP, implement the ILP reduction rule alongside (feature-gated under `ilp-solver`). -- A direct-to-ILP rule is a production reduction, not a stub. Match the completeness bar used by strong ILP reductions in this repo: exact overhead metadata, structure + closed-loop + extraction tests, weighted/infeasible/pathological regressions whenever the semantics require them, and ILP-enabled workspace verification. -- When this rule is the companion to a `[Model]` issue that explicitly claims ILP solvability, it belongs in the same PR as the model. -- If a custom solver is needed, implement in `src/solvers/` and document. - -## CLI Impact - -Adding a witness-preserving reduction rule does NOT require CLI changes -- the reduction graph is auto-generated from `#[reduction]` macros and the CLI discovers paths dynamically. However, both source and target models must already be fully registered through their model files (`ProblemSchemaEntry` and `declare_variants!`), including any aliases and `pred create` construction contract (see `add-model` skill). - -`ExtractionError` already propagates through `pred extract` and bundle `pred solve`; add a rule-specific CLI test only when the CLI surface changes. - -Aggregate-only reductions currently have a narrower CLI surface: -- `pred solve ` can still compute direct aggregate values for aggregate-only problems -- `pred reduce` and `pred solve bundle.json` remain witness-only workflows and reject aggregate-only paths -- Manual aggregate-edge registration affects runtime graph search and internal value extraction, but not bundle solving - -## File Naming - -- Rule file: `src/rules/_.rs` -- no underscores within a problem name - - e.g., `maximumindependentset_qubo.rs`, `minimumvertexcover_maximumindependentset.rs` -- Test file: `src/unit_tests/rules/_.rs` -- Canonical example: `canonical_rule_example_specs()` in the rule module, included from `src/rules/mod.rs` - -## Common Mistakes - -| Mistake | Fix | -|---------|-----| -| Forgetting `#[reduction(...)]` macro | Required for compile-time registration in the reduction graph | -| Using `#[reduction]` for an aggregate-only rule | `#[reduction]` currently registers witness/config edges only; aggregate-only rules need manual `ReductionEntry` wiring with `reduce_aggregate_fn` | -| Wrong overhead expression | Must accurately reflect the size relationship | -| Adding extra reduction metadata or duplicate primitive endpoint registration | Keep one primitive registration per endpoint pair and use only the `overhead` form of `#[reduction]` | -| Missing `extract_solution` mapping state | Store any index maps needed in the ReductionResult struct | -| Permissive extraction | Validate first, then map exactly or return `ExtractionError` | -| Not adding a canonical example | Add the rule-local spec and include it from `src/rules/mod.rs` | -| Not regenerating reduction graph | Run `cargo run --example export_graph` after adding a rule | -| Skipping Step 6 (paper documentation) | **Every rule MUST have a `reduction-rule` entry in the paper. This is mandatory, not optional. PRs without documentation will be rejected.** | -| Source/target model not fully registered | Both problems must already have `ProblemSchemaEntry`, `declare_variants!`, registry aliases as needed, and a construction contract -- use `add-model` skill first | -| Treating a direct-to-ILP rule as a toy stub | Direct ILP reductions need exact overhead metadata and strong semantic regression tests, just like other production ILP rules | -| Skipping verification for complex reductions | Verification is default for a reason — `--no-verify` is for trivial identity/complement reductions only | diff --git a/.claude/skills/auto-pipeline/SKILL.md b/.claude/skills/auto-pipeline/SKILL.md deleted file mode 100644 index 271e3f08e..000000000 --- a/.claude/skills/auto-pipeline/SKILL.md +++ /dev/null @@ -1,389 +0,0 @@ ---- -name: auto-pipeline -description: Use when you want to take a Backlog issue all the way to Final review without manual orchestration — chains check-issue, fix-issue, add-model/add-rule, run-pipeline, and review-pipeline; substantive issue-quality problems are sent to a rewrite subagent; algorithmically unsalvageable issues are parked on OnHold ---- - -# Auto Pipeline - -Take **one** Backlog issue all the way from quality gate to **Final review** without human intervention. The merge step itself is still left to the human (see `/final-review`). - -This skill is an **orchestrator**: it never runs the heavy work itself. Each phase is delegated to a fresh-context subagent. Most phases invoke an existing skill (`check-issue`, `fix-issue`, `run-pipeline`, `review-pipeline`); Phase 3 is owned by the orchestrator and runs raw `cargo test --workspace` + `make paper` to catch breakage the per-item sub-skills cannot see. The only thing the main agent does directly is: - -1. pick the issue, -2. read structured reports from subagents, -3. decide whether to retry, dispatch a rewrite subagent for substantive issues, or park the issue on OnHold, -4. move the project board card forward. - -## Invocation - -- `/auto-pipeline` — pick the highest-priority Backlog issue (Good label first, then lowest issue number) -- `/auto-pipeline 123` — run on a specific Backlog issue number - -## Board states this skill writes - -Only three transitions happen here directly (the rest are owned by sub-skills): - -| Symbolic name passed to `pipeline_board.py move` | When | -|---|---| -| `ready` | Step 1d (quality gate passed) | -| `on-hold` | Step 1e (fundamental flaw or substantive retry cap hit) | - -The orchestrator reads the Backlog column in Step 0 and never writes to it. ID constants for all other columns live in [`run-pipeline`](../run-pipeline/SKILL.md) / [`review-pipeline`](../review-pipeline/SKILL.md) — sub-skills move the card through In Progress → Review pool → Final review. - -## Autonomous Mode - -Runs **fully autonomously** — no confirmation prompts, no clarifying questions. All sub-skills called from here must also auto-approve. The human only gets involved at `/final-review`, or when the issue is parked on OnHold with a diagnostic comment. - -## Subagent Contract - -Every subagent dispatched by this skill operates under the same contract. Each per-step prompt below references this contract by name and only adds the step-specific scope + JSON shape. - -**Output:** the subagent's LAST message must be a single fenced ```json``` block matching the shape given by the dispatching step. No prose before or after. The orchestrator parses only that block. - -**Don'ts:** -- Do NOT modify any source files unless the step's prompt explicitly says so. -- Do NOT move the project board card. The orchestrator owns all board transitions. -- Do NOT open pull requests or invoke `/issue-to-pr`, `gh pr create`, etc. unless the step is `run-pipeline` (which manages its own PR via the existing skill). -- Do NOT brainstorm with a human or wait for input. - -**Severity vocabulary** (used by every Phase-1 step that reports findings): -- `mechanical` — issue-body-fixable without changing the claim (typo, missing G&J number, wrong alias, malformed example, wrong heading). -- `substantive` — the claim is wrong or unsupported (incorrect complexity, broken overhead, mis-cited paper, flawed proof sketch) but a public reference probably exists. -- `fundamental` — algorithm/reduction is mathematically unsound AND your literature search found no public reference that would salvage it. Only assign after a genuine search. - -**Malformed JSON:** if the subagent's reply is missing the fenced JSON block, re-dispatch once with the prompt prefixed by "Your previous reply did not contain a parseable JSON block. Run the skill again from scratch and return ONLY the JSON block." If the second attempt also fails, park the issue on OnHold with reason `subagent contract violation in `. - -## Architecture - -```dot -digraph auto_pipeline { - rankdir=TB; - "Pick issue from Backlog" [shape=box]; - "Phase 1: check-issue (subagent)" [shape=box, style=filled, fillcolor="#cce0ff"]; - "Classify report" [shape=diamond]; - "Phase 1b: auto-fix (subagent)" [shape=box, style=filled, fillcolor="#cce0ff"]; - "Phase 1c: rewrite (subagent)" [shape=box, style=filled, fillcolor="#ffe0cc"]; - "Apply revised issue body" [shape=box]; - "Substantive loop counter" [shape=diamond]; - "Move to OnHold + comment" [shape=box, style=filled, fillcolor="#ffcccc"]; - "Move to Ready" [shape=box]; - "Phase 2: run-pipeline (subagent)" [shape=box, style=filled, fillcolor="#cce0ff"]; - "Phase 3: integration gate (subagent)" [shape=box, style=filled, fillcolor="#cce0ff"]; - "Phase 4: review-pipeline (subagent)" [shape=box, style=filled, fillcolor="#cce0ff"]; - "Final report" [shape=box, style=filled, fillcolor="#ccffcc"]; - - "Pick issue from Backlog" -> "Phase 1: check-issue (subagent)"; - "Phase 1: check-issue (subagent)" -> "Classify report"; - "Classify report" -> "Move to Ready" [label="pass"]; - "Classify report" -> "Phase 1b: auto-fix (subagent)" [label="mechanical only"]; - "Classify report" -> "Phase 1c: rewrite (subagent)" [label="substantive"]; - "Classify report" -> "Move to OnHold + comment" [label="fundamental + no reference"]; - "Phase 1b: auto-fix (subagent)" -> "Phase 1: check-issue (subagent)"; - "Phase 1c: rewrite (subagent)" -> "Apply revised issue body"; - "Apply revised issue body" -> "Substantive loop counter"; - "Substantive loop counter" -> "Phase 1: check-issue (subagent)" [label="< 2 retries"]; - "Substantive loop counter" -> "Move to OnHold + comment" [label=">= 2 retries"]; - "Move to Ready" -> "Phase 2: run-pipeline (subagent)"; - "Phase 2: run-pipeline (subagent)" -> "Phase 3: integration gate (subagent)" [label="success"]; - "Phase 2: run-pipeline (subagent)" -> "Final report" [label="fail (stop)"]; - "Phase 3: integration gate (subagent)" -> "Phase 4: review-pipeline (subagent)" [label="all pass"]; - "Phase 3: integration gate (subagent)" -> "Move to OnHold + comment" [label="any fail"]; - "Phase 4: review-pipeline (subagent)" -> "Final report"; -} -``` - -## Step 0: Pick the Issue - -`scripts/pipeline_board.py backlog` accepts only `model` or `rule` (NOT `all`), returns `{"issue_type": ..., "items": [{number, title, item_id, labels, has_good}, ...]}`, and **exits with code 1 when the queried kind is empty** even though it prints valid JSON — so the picker queries both kinds and ignores subprocess return codes. - -### 0a. Pick - -Set `ISSUE` to the requested number, or leave empty to auto-pick the top of Backlog (Good label first, then lowest number): - -```bash -ISSUE="${ISSUE:-}" # set this to a specific number, or leave empty to auto-pick - -PICK_JSON=$(ISSUE="$ISSUE" python3 <<'PY' -import json, os, subprocess -target = int(os.environ["ISSUE"]) if os.environ.get("ISSUE") else None -items = [] -for kind in ("model", "rule"): - out = subprocess.run( - ["uv", "run", "--project", "scripts", "scripts/pipeline_board.py", - "backlog", kind, "--format", "json"], - capture_output=True, text=True, - ) - try: - items.extend(json.loads(out.stdout)["items"]) - except Exception: - pass -if target is not None: - hit = next((i for i in items if i["number"] == target), None) - print(json.dumps(hit) if hit else "") -elif items: - items.sort(key=lambda i: (not i["has_good"], i["number"])) - print(json.dumps(items[0])) -else: - print("") -PY -) - -if [ -z "$PICK_JSON" ]; then - if [ -n "$ISSUE" ]; then - echo "Issue #$ISSUE is not in the Backlog column." - else - echo "Backlog is empty." - fi - exit 0 -fi -``` - -### 0b. Extract fields - -```bash -ISSUE=$(printf '%s' "$PICK_JSON" | python3 -c "import sys,json; print(json.load(sys.stdin)['number'])") -ITEM_ID=$(printf '%s' "$PICK_JSON" | python3 -c "import sys,json; print(json.load(sys.stdin)['item_id'])") -TITLE=$(printf '%s' "$PICK_JSON" | python3 -c "import sys,json; print(json.load(sys.stdin)['title'])") -LABELS=$(printf '%s' "$PICK_JSON" | python3 -c "import sys,json; print(','.join(json.load(sys.stdin)['labels']))") - -echo "Auto-pipeline starting on issue #$ISSUE — $TITLE" -echo " item_id: $ITEM_ID" -echo " labels: $LABELS" -``` - -### 0c. Initialise loop counter - -```bash -SUBSTANTIVE_RETRIES=0 -MAX_SUBSTANTIVE_RETRIES=2 -``` - -## Step 1: Quality Gate (check-issue + fix loop) - -### 1a. Dispatch `check-issue` subagent - -Use the `Agent` tool with `subagent_type=general-purpose`. The subagent must run the existing `check-issue` skill (force re-check) and report back **structured JSON only**. - -**Prompt template** (subagent follows the Subagent Contract above for everything else): - -``` -Run /check-issue on issue # in CodingThrust/problem-reductions -(--force re-check). Follow .claude/skills/check-issue/SKILL.md exactly. - -For [Rule] issues, Rule Check 5 (Completeness) is the most important -check — find and quote the cited theorem, enumerate corner cases the -source model allows via `pred show --json` and existing -src/rules/ implementations, and hand-trace the algorithm on >= 2 -non-canonical corner cases. A cited precondition the issue ignores is -"substantive"; a cited reference that does not contain the reduction -at all is "fundamental" (set fundamental_no_reference: true). - -You may post the check-issue comment and apply failure/Good labels per -the skill. Do NOT close the issue. - -Return ONLY this JSON shape: -{ - "verdict": "pass" | "fail", - "errors": [{"check": "...", "label": "...", "summary": "...", "severity": "mechanical|substantive|fundamental"}], - "warnings": [{"check": "...", "summary": "...", "severity": "mechanical|substantive"}], - "fundamental_no_reference": true | false, - "comment_url": "" -} -``` - -### 1b. Classify the report - -Parse the JSON. Then branch: - -| Condition | Action | -|---|---| -| `verdict == "pass"` | → Step 1d (move to Ready) | -| `fundamental_no_reference == true` | → Step 1e (OnHold) | -| all `errors`/`warnings` have `severity == "mechanical"` | → Step 1c-mech | -| any `severity == "substantive"` | → Step 1c-sub | - -### 1c-mech. Dispatch auto-fix subagent (mechanical only) - -``` -Run /fix-issue on issue # in auto-fix-only mode: -- Apply only the mechanical auto-fixes from fix-issue's auto-fix step - (the one that runs before the human-brainstorm step). -- Edit the issue body via `gh issue edit` as the skill instructs. -- Skip the re-check and the project-card move (orchestrator handles - re-check by re-dispatching Phase 1). - -Return ONLY this JSON shape: -{ - "applied": [""], - "skipped_substantive": [""], - "errors": [""] -} -``` - -Loop back to **Step 1a** (re-check). Do not increment `SUBSTANTIVE_RETRIES` — mechanical fixes don't count toward the cap. - -### 1c-sub. Rewrite subagent (substantive) - -If `SUBSTANTIVE_RETRIES >= MAX_SUBSTANTIVE_RETRIES` → jump to Step 1e (OnHold) with reason `"substantive issues persist after $MAX_SUBSTANTIVE_RETRIES rewrites"`. - -Otherwise, fetch the current issue body and the latest check-issue comment: - -```bash -ISSUE_BODY=$(gh issue view "$ISSUE" --json body --jq .body) -CHECK_REPORT=$(gh issue view "$ISSUE" --json comments --jq '[.comments[] | select(.body | startswith("## Issue Quality Check"))] | last | .body') -``` - -Dispatch a subagent (`subagent_type=general-purpose`) to research and rewrite: - -``` -Issue # failed /check-issue with substantive findings. Read the -current issue body and the latest check-issue report (both pasted in -the prompt), research public literature with WebSearch / WebFetch, and -either rewrite the body grounded in citations or report that no public -reference can salvage the proposal. - -Issue body: -$ISSUE_BODY - -Latest check-issue report: -$CHECK_REPORT - -Return ONLY one of these JSON shapes: - {"outcome": "revised", "new_body": ""} - {"outcome": "fundamental_flaw", "reason": ""} -``` - -When the subagent returns: - -- **`outcome == "fundamental_flaw"`** → Step 1e (OnHold) with the reason. -- **`outcome == "revised"`** → orchestrator applies the new body (the subagent must NOT edit GitHub itself — keep all edits in the orchestrator for a clean audit trail): - - ```bash - printf '%s' "$NEW_BODY" > /tmp/auto-pipeline-issue-$ISSUE.md - gh issue edit "$ISSUE" --body-file /tmp/auto-pipeline-issue-$ISSUE.md - gh issue comment "$ISSUE" --body "auto-pipeline: issue body rewritten (substantive retry $((SUBSTANTIVE_RETRIES + 1)))" - rm /tmp/auto-pipeline-issue-$ISSUE.md - ``` - - Increment: `SUBSTANTIVE_RETRIES=$((SUBSTANTIVE_RETRIES + 1))` and loop back to **Step 1a**. - -### 1d. Move card to Ready - -```bash -uv run --project scripts scripts/pipeline_board.py move "$ITEM_ID" ready -gh issue comment "$ISSUE" --body "auto-pipeline: quality check passed — moving to Ready." -``` - -Continue to Step 2. - -### 1e. Park on OnHold - -```bash -REASON="" -gh issue comment "$ISSUE" --body "auto-pipeline: parked on OnHold — $REASON. Human triage needed." -uv run --project scripts scripts/pipeline_board.py move "$ITEM_ID" on-hold -``` - -Print the final report and STOP: - -``` -Auto-pipeline halted at quality gate: - Issue: # - Reason: - Board: Backlog -> OnHold -``` - -## Step 2: Implementation (`run-pipeline` subagent) - -Dispatch the existing `run-pipeline` skill against the same issue: - -**Prompt template** (this is the one step the Subagent Contract's no-board-moves rule does NOT apply to — run-pipeline owns its worktree, PR, and board transitions from Ready to Review pool): - -``` -Run /run-pipeline on issue # (already in Ready). Follow -.claude/skills/run-pipeline/SKILL.md exactly — it handles the -worktree, issue-to-pr invocation, and the Ready -> In Progress -> -Review pool transitions, including moving to OnHold on failure. - -Return ONLY this JSON shape: -{ - "outcome": "success" | "failure", - "pr_number": , - "board_status": "Review pool" | "OnHold" | "", - "summary": "" -} -``` - -When the subagent returns: - -- **`outcome == "success"`** → continue to Step 3. -- **`outcome == "failure"`** → STOP. The `run-pipeline` skill already moves the card to OnHold and posts a diagnostic comment, so we do not duplicate. Print: - - ``` - Auto-pipeline halted at implementation: - Issue: # - PR: # - Reason:

- Board: - ``` - - Implementation failures need human eyes — `run-pipeline` already moves the card to OnHold and posts a diagnostic, so the orchestrator just stops here. - -## Step 3: Integration Gate (orchestrator-owned) - -The per-item sub-skills only test the new item in isolation, so cross-crate regressions (e.g. a relaxed model validator breaking pre-existing CLI tests) and paper-compile errors (orphan bib keys, math-mode typos like `intersect` vs Typst's `inter`) slip through Phase 2 and the per-item structural review. Running this gate after Phase 2 catches them locally instead of waiting for CI. - -Dispatch a fresh subagent (`subagent_type=general-purpose`, not invoking any existing skill): - -``` -Run the auto-pipeline integration gate on PR #. Check out the PR -branch in a fresh worktree, run `make check` then `make paper`, clean up. -Do not modify files. Return ONLY: - -{"tests": "pass" | "fail", "paper": "pass" | "fail", - "first_failure": ""} -``` - -- Both `pass` → continue to Step 4. -- Either `fail` → dispatch a fresh subagent (`subagent_type=general-purpose`) with the `first_failure` string and write access to the PR branch, asking it to fix the failure directly (CI-class problems are usually small: deleting a stale test, fixing a typo'd bib key, swapping `intersect` for `inter`). After it returns, re-run Step 3 once. If still failing, park on OnHold. - -## Step 4: Agentic Review (`review-pipeline` subagent) - -Dispatch the existing `review-pipeline` skill against the PR: - -**Prompt template** (board transitions to Final review are owned by review-pipeline; that's its contract): - -``` -Run /review-pipeline on PR #. Follow -.claude/skills/review-pipeline/SKILL.md exactly; it always moves the -PR to Final review at the end. - -Return ONLY this JSON shape: -{ - "outcome": "success" | "failure", - "board_status": "Final review" | "", - "review_verdicts": {"structural": "...", "quality": "...", "agentic": "..."}, - "summary": "" -} -``` - -Whatever the outcome, the PR is now either in Final review (success) or stuck somewhere the review skill left it (failure). Print the final report: - -``` -Auto-pipeline complete: - Issue: # - PR: # - Board: - Verdicts: structural=<...> quality=<...> agentic=<...> - Next: human runs /final-review -``` - -## Common Mistakes - -| Mistake | Fix | -|---------|-----| -| Calling sub-skills directly in the main agent | Always dispatch via `Agent` tool — keeps the orchestrator context clean | -| Letting the rewrite subagent edit GitHub | The orchestrator owns all `gh issue edit` calls — subagents only return text | -| Treating implementation failures as substantive issue problems | Step 2 failures go straight to a stop; the orchestrator does not attempt to auto-fix `run-pipeline` output | -| Picking from a non-Backlog column when no issue number is given | Auto-pick must read from Backlog only — never from OnHold, Ready, or elsewhere | -| Skipping Step 3 because Phase 2 reported `success` | Phase 2 success is scoped to the new item's own tests; workspace-wide regressions and paper-compile bugs are only visible from `make check` + `make paper`. | diff --git a/.claude/skills/check-issue/SKILL.md b/.claude/skills/check-issue/SKILL.md deleted file mode 100644 index e63957476..000000000 --- a/.claude/skills/check-issue/SKILL.md +++ /dev/null @@ -1,571 +0,0 @@ ---- -name: check-issue -description: Use when reviewing a [Rule] or [Model] GitHub issue for quality before implementation — checks usefulness, non-triviality, correctness of literature claims, and writing quality ---- - -# Check Issue - -Quality gate for `[Rule]` and `[Model]` GitHub issues. Runs 4 checks (adapted per issue type) and posts a structured report as a GitHub comment. Adds labels for failures but does NOT close issues. - -## Invocation - -``` -/check-issue -``` - -## Process - -```dot -digraph check_issue { - rankdir=TB; - "Fetch issue" [shape=box]; - "Detect issue type" [shape=diamond]; - "Run Rule checks" [shape=box]; - "Run Model checks" [shape=box]; - "Unknown type: stop" [shape=box, style=filled, fillcolor="#ffcccc"]; - "Compose report" [shape=box]; - "Post comment + add labels" [shape=box]; - - "Fetch issue" -> "Detect issue type"; - "Detect issue type" -> "Run Rule checks" [label="[Rule]"]; - "Detect issue type" -> "Run Model checks" [label="[Model]"]; - "Detect issue type" -> "Unknown type: stop" [label="other"]; - "Run Rule checks" -> "Compose report"; - "Run Model checks" -> "Compose report"; - "Compose report" -> "Post comment + add labels"; -} -``` - -### Prerequisites - -This skill uses the `pred` CLI tool. If `pred` is not available, build it first: - -```bash -pred --version 2>/dev/null || make cli -``` - -### Step 0: Fetch and Parse Issue - -```bash -gh issue view --json title,body,labels,comments -``` - -- Detect issue type from title: `[Rule]` or `[Model]` -- If neither, stop with message: "This skill only checks [Rule] and [Model] issues." - -#### Duplicate Check Detection - -Before running checks, scan existing comments for a previous `## Issue Quality Check` heading. If found: -- **Default:** Skip and report: "Already checked (comment from YYYY-MM-DD). Use `/check-issue --force` to re-check." -- **`--force` flag:** Proceed with re-check. Post the new report as "Re-check" and note any changes from the previous report (e.g., "Previously: 2 warnings → Now: 0 warnings after issue edits"). - ---- - -# Part A: Rule Issue Checks - -Applies when the title contains `[Rule]`. - -## Rule Check 1: Usefulness (fail label: `Useless`) - -**Goal:** Is this reduction novel or does it improve on an existing one? - -1. Parse **Source** and **Target** problem names from the issue body. - -2. Resolve problem aliases (the issue may say "MIS" but `pred` needs "MaximumIndependentSet"): - ```bash - pred show --json 2>/dev/null - pred show --json 2>/dev/null - ``` - If either fails, try the full name. If still fails, **Warn** (unknown problem — may be a new model). - -3. Check existing path: - ```bash - pred path --json - ``` - -4. Decision (principle: new rule must reduce the reduction overhead): - - **No path exists** → **Pass** (novel reduction) - - **Path exists** → run `/topology-sanity-check redundancy ` to perform a full overhead dominance analysis against all composite paths. Use its verdict: - - **Not Redundant** → **Pass** ("improves existing reduction — not dominated by any composite path") - - **Redundant** (dominated by a composite path) → **Fail** — include the dominating path from the redundancy report - - **Inconclusive** → **Warn** with the details from the redundancy report - -5. Check **Motivation** field: if empty, placeholder, or just "enables X" without explaining *why this path matters* → **Warn** - ---- - -## Rule Check 2: Non-trivial (fail label: `Trivial`) - -**Goal:** Does the reduction involve genuine structural transformation? - -Read the "Reduction Algorithm" section and flag as **Fail** if: - -- **Variable substitution only:** The mapping is a 1-to-1 relabeling (e.g., `x_i → 1 - x_i` for complement problems). A valid reduction must construct new constraints, objectives, or graph structure. -- **Subtype coercion:** The reduction merely casts to a more general type within an existing variant hierarchy (e.g., UnitDiskGraph → SimpleGraph) with no structural change to the problem instance. -- **Same-problem identity:** Reducing between variants of the same problem with no insight (e.g., `MIS` → `MIS` by setting all weights to 1). -- **Insufficient detail:** The algorithm is a hand-wave ("map variables accordingly", "follows from the definition") — not a step-by-step procedure a programmer could implement. This is also a **Fail**. - -If the construction involves gadgets, penalty terms, auxiliary variables, or non-trivial structural transformation → **Pass**. - -Note: if a trivial reduction make original disconnected problems connected, it is also **Pass**. - ---- - -## Rule Check 3: Correctness (fail label: `Wrong`) - -**Goal:** Are the cited references real and do they support the claims? - -### 3a: Extract References - -Parse all literature citations from the issue body — paper titles, author names, years, DOIs, arxiv IDs, textbook references. - -### 3b: Cross-check Against Project Knowledge Base - -First, read `references.md` (in this skill's directory) for quick fact-checking — it contains known complexity bounds, key results, and established reductions with BibTeX keys. If the issue's claims contradict known facts in this file → **Fail** immediately. - -Then read `docs/paper/references.bib` for the full bibliography. If a cited paper is already in the bibliography: -- Verify the claim in the issue matches what the paper actually says -- If the issue cites a result from a known paper but misquotes it → **Fail** - -### 3c: External Verification (best-effort fallback chain) - -For each reference NOT in the bibliography: - -1. **Try arxiv MCP** (if available): search by title/author/arxiv ID -2. **Try Semantic Scholar MCP** (if available): search by title/DOI -3. **Fall back to WebSearch + WebFetch**: search for the paper, fetch abstract - -For each reference, verify: -- Paper exists with matching title, authors, and year -- The specific theorem/result cited actually appears in the paper -- Cross-check the claim against at least one other source (survey paper, textbook, or independent reference) - -If any cited fact **cannot be verified** (paper not found, claim not in paper) → **Fail** with specifics. - -### 3d: Better Algorithm Discovery (not fatal) - -While searching, if you find: -- A **more recent paper** that supersedes the cited reference -- A **lower overhead** construction for the same reduction -- A **different approach** that achieves better bounds - -→ Include in the report as a **Recommendation** (not a failure). Example: "Note: Smith et al. (2024) improve on this with O(n) overhead instead of O(n^2)." - ---- - -## Rule Check 4: Well-written (fail label: `PoorWritten`) - -**Goal:** Is the issue clear, complete, and implementable? - -### 4a: Information Completeness - -Check all template sections are present and substantive (not placeholder text): - -| Section | Required content | -|---------|-----------------| -| Source | Valid problem name | -| Target | Valid problem name | -| Motivation | Why this reduction matters | -| Reference | At least one literature citation | -| Reduction Algorithm | Complete step-by-step procedure | -| Size Overhead | Table with code metric names and formulas | -| Validation Method | How to verify correctness | -| Example | Concrete worked instance | - -Missing or placeholder sections → list them as **Fail** items. - -### 4b: Algorithm Completeness - -The reduction algorithm must be a **complete, step-by-step procedure** detailed enough for a programmer to implement: -- Every step numbered and unambiguous -- All intermediate values defined -- No gaps ("similarly for the remaining variables") -- Solution extraction must be clear from the variable mapping - -If the algorithm is a high-level sketch rather than an implementable procedure → **Fail**. - -### 4c: Symbol and Notation Consistency - -- **All symbols must be defined before first use.** E.g., if the algorithm references `n`, there must be a prior line "let n = |V|". -- **Symbols must be consistent** between sections. If the algorithm defines `G = (V, E)` but the overhead table uses `N` for vertex count without defining it → **Fail**. -- **Code metric names** in the overhead table must match actual getter methods on the target problem (e.g., `num_vertices`, `num_edges`, `num_vars`, `num_clauses`). Check with `pred show --json` → `size_fields`. - -### 4d: Example Quality - -- **Non-trivial**: Must have enough structure to exercise the reduction meaningfully (not just 2 vertices) -- **Brute-force solvable**: Small enough to verify by hand or with `pred solve` -- **Fully worked**: Shows the source instance, the reduction construction step by step, and the target instance — not just "apply the reduction to get..." -- **Round-trip testable**: The example must be complex enough to validate correctness via a closed-loop test: reduce the source instance → solve the target → extract the solution back → verify it is optimal for the source. A too-simple example (e.g., a single edge, a trivially satisfiable formula) can pass the round trip even with a buggy reduction. The example should have multiple feasible solutions with different objective values so that only a correct reduction maps to the true optimum. Rule of thumb: the source instance should have at least 2 suboptimal feasible solutions in addition to the optimal one. - ---- - -## Rule Check 5: Completeness (fail label: `Incomplete`) - -**Goal:** Does the proposed mapping work for *every* instance of the source problem, not just the canonical case in the example? - -A reduction that only handles a subset of source instances (e.g., "assumes connected graph", "assumes no duplicate elements", "works only when k is even") is **not a valid polynomial-time reduction** unless the issue explicitly: -- restricts the source to a sub-variant that the codebase actually exposes as a distinct model, AND -- the restriction is part of the algorithm statement, not a hidden assumption. - -This check is **mandatory** for every `[Rule]` issue and must be backed by **explicit literature and codebase research** — not vibes. - -### 5a: Literature research (mandatory) - -For the cited construction(s): - -1. Find the original paper or textbook section that defines the reduction. -2. Read the **statement of the theorem**: what does it claim the reduction handles? Look for phrases like: - - "for any instance of P" → covers all instances (good) - - "for a P-instance such that ..." → has a precondition (must be flagged) - - "in the special case where ..." → only a special case (must be flagged) -3. Read the **proof**: are there steps that silently assume something about the source (no isolated vertices, no zero weights, integer-valued capacities, ...)? - -Use the same fallback chain as Check 3c. - -If the cited paper is **not actually a reduction from the full source problem** but from a restricted variant → **Fail** with the precise restriction quoted from the paper. The fix is one of: (a) add a preprocessing step that reduces the full source to the restricted variant, (b) split into a `[Rule]` issue from the actual restricted source, or (c) drop the reduction. - -### 5b: Codebase corner-case research (mandatory) - -Check the actual codebase to see what shape the source problem can take: - -```bash -pred show --json -``` - -Read the `size_fields` and any variant getters, then enumerate corner cases the issue's algorithm must handle: - -| Class | Example corner cases the reduction must accept | -|---|---| -| Graph-input problems | empty graph, single vertex, isolated vertices, self-loops if the model allows them, parallel edges if allowed, disconnected components, complete graph | -| Weighted problems | all weights equal, all weights zero, mixed signs (if the weight type allows), one weight dominating the rest | -| Formula/circuit | empty clause set, single-literal clauses, tautological clauses, repeated variables in a clause | -| Set systems | empty universe, empty subsets, identical subsets, universe element appearing in no subset | -| Algebraic | zero matrix, identity, singular matrix | - -Then trace the **issue's** algorithm by hand against at least 2 corner cases that are not the worked example: - -1. Pick a corner case from the table above that the source model actually allows. -2. Simulate the issue's construction step by step. -3. Check: is the target problem well-defined? Does solution extraction still work? - -Also grep the codebase for any existing rule whose source has the same problem name — if it already handles certain corner cases, the new rule should at least match that coverage: - -```bash -grep -rl "impl.*ReduceTo.*for " src/rules/ -``` - -Read 1–2 of those existing rules for how they handle edge inputs. - -If the issue's algorithm **crashes, produces an invalid target instance, or loses information** on a legitimate corner case → **Fail** with the corner case spelled out. - -If the issue's algorithm appears to handle corner cases correctly but the issue body doesn't *state* this explicitly → **Warn** ("works on tested corner cases, but the algorithm description does not address edge inputs — please document"). - -### 5c: Verdict - -| Finding | Verdict | -|---|---| -| Literature explicitly covers all instances AND traced corner cases work | **Pass** | -| Literature explicitly covers all instances but issue is silent on corner cases | **Warn** | -| Literature has a precondition the issue ignores | **Fail** | -| Traced corner case breaks the algorithm | **Fail** | - -(If the cited reference doesn't actually contain the reduction at all, Check 3c already catches it — don't double-flag here.) - -Report the literature evidence and the corner cases you traced in the comment — this is the most expensive check and reviewers will want to see your work. - ---- - -# Part B: Model Issue Checks - -Applies when the title contains `[Model]`. - -## Model Check 1: Usefulness (fail label: `Useless`) - -**Goal:** Does this problem add value to the reduction graph? - -1. Parse the **problem name** from the issue body ("Name" field under Definition). - -2. Check if the problem already exists: - ```bash - pred show --json 2>/dev/null - ``` - If it succeeds, the problem **already exists** → **Fail** ("Problem already implemented"). - -3. Check **planned reductions** — the issue must mention at least one concrete reduction rule connecting this problem to the existing graph: - - Look for explicit statements like "reduces to/from X", "interreducible with Y", or references to planned `[Rule]` issues - - If **no reduction is mentioned at all** → **Fail** ("Orphan node — a problem without any planned reduction rule has no value in the reduction graph. Add at least one planned reduction to/from an existing problem.") - - If reductions are mentioned but vague ("can be connected to other problems") → **Warn** - -4. Check **Motivation** field: - - Is there a concrete use case? (quantum computing, network design, scheduling, etc.) - - If motivation is empty, placeholder, or vague → **Warn** - -5. Check **How to solve** section: - - At least one solver method must be checked (brute-force, ILP reduction, or other) - - If no solver path is identified → **Warn** ("No solver means reduction rules can't be verified") - - If direct ILP solving is claimed, the issue must link a direct `[Rule] to ILP` companion issue in the "Reduction Rule Crossref" section; otherwise → **Fail** - ---- - -## Model Check 2: Non-trivial (fail label: `Trivial`) - -**Goal:** Is this genuinely a distinct problem, not a repackaging of an existing one? - -Flag as **Fail** if: - -- **Isomorphic to existing problem:** The definition is mathematically equivalent to a problem already in the codebase under a different name (e.g., proposing "Maximum Weight Clique" when `MaximumClique` with weights already exists). -- **Trivial variant:** The proposed problem is just an existing problem restricted to a specific graph type or weight type that could be handled by adding a variant to the existing model (e.g., "MIS on bipartite graphs" is a variant, not a new problem). -- **Trivial renaming:** Same feasibility constraints and objective, different name. - -Check against existing problems: -```bash -pred list --json -``` - -If the problem has a genuinely different feasibility constraint or objective function from all existing problems → **Pass**. - ---- - -## Model Check 3: Correctness (fail label: `Wrong`) - -**Goal:** Are the definition, complexity claims, and references accurate? - -### 3a: Definition Correctness - -- Verify the formal definition is mathematically well-formed -- Check that feasibility constraints and objective are clearly separated -- Verify the variable domain matches the problem semantics (binary for selection, k-ary for coloring, etc.) - -### 3e: Representation Feasibility - -Verify that the proposed data types in the Schema can represent the stated problem domain: -- If the Schema proposes a data type but the Definition or Variants mention domains that exceed that type's range (e.g., proposing integer coefficients for a finite field larger than any fixed-width integer can hold) → **Fail** ("Proposed data type cannot represent the stated domain") -- If multiple variants are listed, check that the proposed schema handles all of them or explicitly restricts scope -- If the issue acknowledges a limitation and restricts scope (e.g., "initial implementation targets small fields only"), this is acceptable → **Pass** with a note - -### 3b: Complexity Verification - -The issue claims a best-known exact algorithm with a specific time bound. Verify: -- The cited paper/algorithm actually exists (use same fallback chain as Rule Check 3) -- The time bound matches what the paper claims -- For polynomial-time problems: verify they are indeed polynomial (not NP-hard) -- For NP-hard problems: verify the exponential base is correct (e.g., 1.1996^n for MIS, not 2^n) -- **If a better algorithm is found** during search → add as a **Recommendation** in the comment - -### 3c: Cross-check Against Project Knowledge Base - -Same process as Rule Check 3b — first read `references.md` (in this skill's directory) for quick fact-checking against known complexity bounds and results. Then read `docs/paper/references.bib` and verify claims against known papers. - -### 3d: External Verification - -Same fallback chain as Rule Check 3c: -1. arxiv MCP → 2. Semantic Scholar MCP → 3. WebSearch + WebFetch - -Verify each reference exists and supports the claims made in the issue. - ---- - -## Model Check 4: Well-written (fail label: `PoorWritten`) - -**Goal:** Is the issue clear, complete, and implementable? - -### 4a: Information Completeness - -Check all template sections are present and substantive: - -| Section | Required content | -|---------|-----------------| -| Motivation | Concrete use case and graph connectivity | -| Name | Valid problem name following naming conventions | -| Reference | At least one literature citation | -| Definition | Formal: input, feasibility constraints, objective | -| Variables | Count, per-variable domain, semantic meaning | -| Schema | Type name, variants, field table | -| Complexity | Best known algorithm with citation **and** a concrete complexity expression in terms of problem parameters (e.g., `q^n`, `2^{0.8765n}`) | -| How to solve | At least one solver method checked; if ILP is claimed, a direct `[Rule] to ILP` issue must be linked | -| Example Instance | Concrete instance that exercises the core structure | -| Expected Outcome | Satisfaction: one valid / satisfying solution with brief justification. Optimization: one optimal solution with the optimal objective value | - -Missing or placeholder sections → list them as **Fail** items. - -### 4b: Definition Completeness - -The formal definition must be **precise and implementable**: -- Input structure clearly specified (graph, formula, matrix, etc.) -- Feasibility constraints stated as mathematical conditions -- Objective (if optimization) stated as what to maximize/minimize -- All quantifiers explicit ("for all edges (u,v) in E" not "adjacent vertices don't share colors") - -### 4c: Symbol and Notation Consistency - -- **All symbols defined before first use.** If the definition uses `G = (V, E)`, the Variables section should reference `V` consistently. -- **Symbols consistent across sections.** The Schema field descriptions must match symbols in the Definition. -- **Naming conventions:** optimization problems must use `Maximum`/`Minimum` prefix. Check against CLAUDE.md naming rules. - -### 4d: Example Quality - -- **Non-trivial**: Enough vertices/variables to exercise constraints meaningfully (not just a triangle) -- **Exercises core structure**: Examples must use the defining features of the problem. For instance, a "MultivariateQuadratic" example that only has linear terms does not exercise the quadratic structure → **Fail**. If the problem's name or definition highlights a specific structural feature (quadratic, k-colorable, bipartite, etc.), at least one example must exercise that feature. -- **Expected outcome provided**: - - Satisfaction problems must include a concrete valid / satisfying solution and say why it is valid - - Optimization problems must include a concrete optimal solution and the optimal objective value -- **Detailed enough for paper**: This example will appear in the paper — it needs to be illustrative -- **Round-trip testable**: The example must be complex enough that a round-trip test (construct instance → solve → verify) can catch implementation bugs. A too-simple instance (e.g., 2 vertices, a single clause) may have a trivially correct solution that passes even with a wrong implementation. The example should have multiple feasible configurations with different objective values (for optimization) or a mix of satisfying and non-satisfying configurations (for satisfaction problems), so that correctness is meaningfully tested. Rule of thumb: the instance should have at least 2 suboptimal feasible solutions in addition to the optimal one. -- **ILP-testable when claimed**: If the issue advertises a direct ILP path, the example should be rich enough to support strong ILP closed-loop tests rather than a degenerate "any formulation passes" case. - -### 4e: Representation Feasibility - -Same check as Correctness 3e — if the proposed data types cannot represent the stated domain, this is also a **Fail** here (the schema is not implementable as written). - ---- - -# Step 2: Compose and Post Report - -### Report Format - -Post a single GitHub comment. The table adapts to the issue type: - -**For [Rule] issues:** - -````markdown -## Issue Quality Check — Rule - -| Check | Result | Details | -|-------|--------|---------| -| Usefulness | ✅ Pass | No existing direct reduction Source → Target | -| Non-trivial | ✅ Pass | Gadget construction with penalty terms | -| Correctness | ❌ Fail | Paper "Smith 2020" not found on arxiv or Semantic Scholar | -| Completeness | ⚠️ Warn | Algorithm correct on traced corner cases but issue body silent on edge inputs | -| Well-written | ⚠️ Warn | Symbol `m` used in overhead table but not defined in algorithm | - -**Overall: 2 passed, 1 failed, 2 warnings** - ---- - -### Usefulness -[Detailed explanation] - -### Non-trivial -[Detailed explanation] - -### Correctness -[Per-reference verification results, any better algorithms found] - -### Completeness -[Literature passages cited (with quote + section/theorem number) showing whether the construction covers all source instances, and the corner cases you traced by hand with the algorithm — including any that broke or any preconditions you discovered] - -### Well-written -[Specific items to fix] - -#### Recommendations -- [Better algorithms or papers discovered] -- [Suggestions for improving the issue] -```` - -**For [Model] issues:** - -````markdown -## Issue Quality Check — Model - -| Check | Result | Details | -|-------|--------|---------| -| Usefulness | ✅ Pass | Novel problem not yet in reduction graph | -| Non-trivial | ✅ Pass | Distinct feasibility constraints from existing problems | -| Correctness | ⚠️ Warn | Complexity bound not independently verified | -| Well-written | ❌ Fail | Missing Variables section; symbol `K` undefined | - -**Overall: 2 passed, 1 warning, 1 failed** - ---- - -### Usefulness -[Detailed explanation] - -### Non-trivial -[Detailed explanation] - -### Correctness -[Per-reference verification, complexity check results, better algorithms found] - -### Well-written -[Specific items to fix] - -#### Recommendations -- [Better complexity bounds discovered] -- [Suggestions for improving the issue] -```` - -### Label Application - -```bash -# Add labels for FAILED checks (not warnings) -gh issue edit --add-label "Useless" # if Check 1 failed -gh issue edit --add-label "Trivial" # if Check 2 failed -gh issue edit --add-label "Wrong" # if Check 3 failed -gh issue edit --add-label "PoorWritten" # if Check 4 failed -gh issue edit --add-label "Incomplete" # if Rule Check 5 failed - -# "Good" label requires: zero failures AND zero warnings on Usefulness, Correctness, or Completeness. -# Warnings on Non-trivial or Well-written alone do NOT block "Good". -gh issue edit --add-label "Good" - -# If re-checking after fixes, remove stale failure labels and add "Good" if now passing -gh issue edit --remove-label "Useless,Trivial,Wrong,PoorWritten,Incomplete" 2>/dev/null -gh issue edit --add-label "Good" -``` - -**Never close the issue.** Labels and comments only. - -### Comment Posting - -```bash -gh issue comment --body "$(cat <<'EOF' - -EOF -)" -``` - ---- - -## Tool Fallback Chain for Literature - -| Priority | Tool | Use for | -|----------|------|---------| -| 1 | arxiv MCP | arxiv papers (search by ID, title, author) | -| 2 | Semantic Scholar MCP | DOI lookup, citation graphs, abstracts | -| 3 | WebSearch | General paper search, cross-referencing | -| 4 | WebFetch | Fetch specific paper pages for claim verification | - -If an MCP tool is not available, skip to the next in the chain. All checks should be possible with just WebSearch + WebFetch as a baseline. - ---- - -## Step 3: Offer to Fix (optional) - -After posting the report, if there are **fixable failures** (not just warnings), ask the user: - -> "Would you like me to help fix the issues found? I can update the issue body to address: [list fixable items]" - -**Auto-fixable items** (if the user agrees): -- Missing or placeholder sections → fill with templates from the issue template -- Incorrect DOI format → reformat to standard `https://doi.org/...` form -- Inconsistent notation → standardize symbols across sections -- Missing symbol definitions → add definitions based on context - -**NOT auto-fixable** (require the contributor's input): -- Missing reduction algorithm or proof details -- Incorrect mathematical claims -- Missing references (need the contributor to provide them) - -If the user agrees, edit the issue body with `gh issue edit --body "..."` and re-run the checks to verify the fixes. - ---- - -## Common Mistakes - -- **Don't fail on warnings.** Only add labels for definitive failures. Ambiguous cases get warnings. -- **Don't close issues.** This skill labels and comments only. -- **Don't hallucinate paper content.** If you can't find a paper, say "not found" — don't guess what it might contain. -- **Don't hallucinate issue references.** Do NOT reference other GitHub issues unless you have fetched them with `gh issue view` and verified their content. Do NOT reference file paths unless you have verified they exist. -- **Match problem names carefully.** Issues may use aliases (MIS, MVC, SAT) that need resolution via `pred show`. -- **Check the right template.** `[Rule]` and `[Model]` issues have different sections — don't check for "Reduction Algorithm" on a Model issue. diff --git a/.claude/skills/dev-setup/SKILL.md b/.claude/skills/dev-setup/SKILL.md index 333a693e0..8ed7f506f 100644 --- a/.claude/skills/dev-setup/SKILL.md +++ b/.claude/skills/dev-setup/SKILL.md @@ -1,187 +1,38 @@ --- name: dev-setup -description: Interactive wizard to install and configure all development tools for new maintainers +description: Use when a new maintainer or contributor needs a working development environment for this repo — detects missing tools, installs them after one confirmation, authenticates gh, and verifies with make check --- # Dev Setup -Interactive wizard that helps new maintainers install and configure all tools needed for the problemreductions project. - -## Step 1: Dependencies Checklist - -Check if `skills/dev-setup/dependencies.md` exists and has content. - -- **If it exists**, ask the user: - > "Found existing dependencies checklist. Use it as-is, or rescan project files for changes?" - - **Use existing** → read `dependencies.md` and proceed to Step 2 - - **Rescan** → scan project files (see Scan Targets below), overwrite `dependencies.md`, then proceed - -- **If it does not exist**, scan project files and generate `dependencies.md` with the format shown in the existing file. Then proceed. - -### Scan Targets - -When scanning, read these files for tool references: - -- `Makefile` — tool invocations (cargo, mdbook, typst, uv, julia, python3, jq, gh, claude) -- `.claude/skills/*/SKILL.md` — CLI references (gh, git, make, cargo, claude, pred) -- `.github/workflows/*.yml` — installed tools, rustup components, and actions -- `scripts/pyproject.toml` — Python tooling (uv) -- `scripts/jl/Project.toml` — Julia dependency -- `Cargo.toml` / `problemreductions-cli/Cargo.toml` — feature flags and build deps - -Organize tools into three tiers in `dependencies.md`: -- **Core** — needed to build, test, and generate docs -- **Skill** — needed for the AI-assisted pipeline (gh, claude, pred) -- **Optional** — nice to have but not required (julia) - -Each tool needs: name, check command, install command (macOS), install command (Linux), purpose. - -## Step 2: Detect Platform - -```bash -uname -s -``` - -- `Darwin` → use macOS install commands -- `Linux` → use Linux install commands - -## Step 3: Install Core Tools - -For each tool in the **Core Tools** table of `dependencies.md`: - -1. Run the check command -2. **If found** → print `[tool] installed` and continue -3. **If missing** → print the install command for the detected platform, then execute it - -After all core tools are done, ask: -> "Core tools are installed. Do you also want to set up the AI pipeline tools (gh, claude, pred)?" - -- **Yes** → proceed to Step 4 -- **No** → skip to Step 6 - -## Step 4: Install Skill Tools - -For each tool in the **Skill Tools** table: - -1. Run the check command -2. **If found** → print `[tool] installed` and continue -3. **If missing** → print the install command, then execute it - -Note: `pred` is built from the local workspace. Use `cargo install --path problemreductions-cli`. - -After skill tools, ask: -> "Want to install optional tools (julia)?" - -- **Yes** → install optional tools using the same check/install pattern -- **No** → continue - -## Step 5: Auth and Configuration - -Skip this step if the user declined skill tools in Step 3. - -### 5a: GitHub CLI auth - -```bash -gh auth status -``` - -If not authenticated, run `gh auth login`. - -### 5b: Repo access - -```bash -gh repo view --json name -``` - -If this fails, the user needs repo access. Explain how to request it. - -### 5c: Project board access - -```bash -gh project list --owner -``` - -If this fails with permission errors, run: -```bash -gh auth refresh -s read:project,project -``` - -Explain that the `project-pipeline` and `review-pipeline` skills require these OAuth scopes. - -## Step 6: Verification - -Run the full check: - -```bash -make check -``` - -This runs `fmt-check + clippy + test`. Print a pass/fail summary for each stage. - -### Troubleshooting Common Failures - -| Failure | Fix | -|---------|-----| -| `fmt-check` fails | Run `make fmt` to auto-fix | -| Linker errors in clippy/test | Missing C/C++ toolchain for `ilp-highs` feature. Install Xcode CLT (`xcode-select --install` on macOS) or `build-essential` (`sudo apt install build-essential` on Linux) | -| "HiGHS not found" or cmake errors | Install cmake: `brew install cmake` (macOS) or `sudo apt install cmake` (Linux) | -| `cargo llvm-cov` fails with "missing llvm-profdata" | `rustup component add llvm-tools-preview` | - -If `make check` passes and the user declined skill tools, print: -> "Setup complete! All core tools installed and verified. You're ready to contribute." - -If it fails, walk through the troubleshooting table and offer to run the fix commands. - -### Pipeline Verification (skill tier only) - -If the user installed skill tools, also verify the autonomous pipeline works: - -**6b: Test `make run-pipeline` prerequisites** - -```bash -# Verify gh can access the project board -gh project item-list 8 --owner CodingThrust --format json --limit 1 -``` - -If this fails, the user likely needs org-level project scopes: -```bash -gh auth refresh -s read:project,project -``` - -**6c: Test claude CLI** - -```bash -claude --version -``` - -If all pipeline checks pass, explain the project-based contribution pipeline: - -> **Setup complete!** All tools installed and verified. -> -> ## How the Project Pipeline Works -> -> This project uses a [GitHub Project board](https://github.com/orgs/CodingThrust/projects/8/views/1) to track issues through an automated pipeline. Issues flow through these columns: -> -> ``` -> Ready → In Progress → Review pool → Under review → Final review → Done -> ``` -> -> Two `make` commands drive this pipeline: -> -> ### `make run-pipeline` (issue → PR) -> Picks the next **Ready** issue, moves it to **In Progress**, implements it (using `/issue-to-pr` → `/add-model` or `/add-rule`), creates a PR, then moves it to **Review pool**. -> -> ### `make run-review` (PR → review) -> Picks the next **Review pool** PR, runs agentic review (structural + quality + feature tests), then moves it to **Final review** for human approval. -> -> ### Targeting specific items -> - `make run-pipeline N=42` — process issue #42 -> - `make run-review N=570` — process PR #570 -> -> ### Available skills for manual work -> You can also invoke individual skills directly: -> - `/issue-to-pr 42` — convert a specific issue into a PR -> - `/add-model` — interactively add a new problem model -> - `/add-rule` — interactively add a new reduction rule -> - `/fix-pr` — fix review comments and CI on the current PR -> - `/release` — prepare a new crate release +Get a machine to the point where `make check` passes and the contributor tools work. +The tool list lives in `.claude/skills/dev-setup/dependencies.md` (core, contributor workflow, +optional). + +## Flow + +1. **Detect.** `uname -s` selects the macOS or Linux install commands (on Linux also check the package + manager; the table assumes apt). Run every check command in `dependencies.md` and collect + what is missing. +2. **Confirm once.** Show one list of missing tools with the exact install command for each, + grouped by tier, and mark which need `sudo`. Ask whether to install all, core only, or a + subset. Never run `sudo` or a `curl | sh` installer before this confirmation. +3. **Install** the approved set without further prompts, then re-run the checks and report + anything that still fails. +4. **Auth.** If `gh` is in scope: `gh auth status`, else `gh auth login`; then + `gh repo view CodingThrust/problem-reductions --json name` to confirm access. +5. **Verify.** `make check` (fmt-check + clippy + test). If contributor tools were installed, + also `pred list | head -3` and `pred-sym big-o "2^n"`. + +## Non-obvious failures + +| Symptom | Cause / fix | +|---------|-------------| +| `highs-sys` build fails: cmake not found, C++ errors | `highs` is an unconditional dependency built from source; needs cmake and a C/C++ toolchain | +| `highs-sys`: "Unable to find libclang" | bindgen needs libclang (`libclang-dev`; on macOS the Xcode CLT) | +| `make fmt-check` fails | `make fmt` | +| `make doc` fails at `npm ci` | node/npm missing or too old; CI uses Node 22 | +| `make coverage` missing llvm-profdata | `rustup component add llvm-tools-preview` | +| `pred` behaves differently from the docs | stale install; rerun `make cli` (reinstalls from the workspace) | + +Finish with a short summary: installed, already present, skipped, and the `make check` result. diff --git a/.claude/skills/dev-setup/dependencies.md b/.claude/skills/dev-setup/dependencies.md index 8f5b34dbc..307556de3 100644 --- a/.claude/skills/dev-setup/dependencies.md +++ b/.claude/skills/dev-setup/dependencies.md @@ -1,35 +1,38 @@ # Development Dependencies -Auto-generated tool checklist for problemreductions maintainers. -Rescan by running `/dev-setup` and choosing "rescan". +Static tool list for problemreductions maintainers. Update it by hand when the Makefile, CI +workflows, or `Cargo.toml` start needing a new tool. -## Core Tools (build, test, docs) +## Core (build, test, docs, paper) -| Tool | Check Command | Install (macOS) | Install (Linux) | Purpose | -|------|--------------|-----------------|-----------------|---------| -| git | `git --version` | preinstalled | `sudo apt install git` | Version control | -| rust | `rustc --version` | `curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs \| sh` | same | Compiler and toolchain | -| clippy | `cargo clippy --version` | `rustup component add clippy` | same | Linting | -| rustfmt | `rustfmt --version` | `rustup component add rustfmt` | same | Code formatting | -| llvm-tools-preview | `rustup component list --installed \| grep llvm-tools` | `rustup component add llvm-tools-preview` | same | Required by cargo-llvm-cov | -| make | `make --version` | preinstalled | `sudo apt install make` | Build orchestration | -| python3 | `python3 --version` | `brew install python` | `sudo apt install python3` | Scripts and data generation | -| uv | `uv --version` | `curl -LsSf https://astral.sh/uv/install.sh \| sh` | same | Python package manager | -| jq | `jq --version` | `brew install jq` | `sudo apt install jq` | JSON processing (cli-demo, compare) | -| mdbook | `mdbook --version` | `cargo install mdbook` | same | Documentation site | -| typst | `typst --version` | `brew install typst` | `cargo install typst-cli` | Paper and diagrams | -| cargo-llvm-cov | `cargo llvm-cov --version` | `cargo install cargo-llvm-cov` | same | Coverage reports | +| Tool | Check | Install (macOS) | Install (Linux, Debian/Ubuntu) | Needed by | +|------|-------|-----------------|--------------------------------|-----------| +| git | `git --version` | `xcode-select --install` | `sudo apt install git` | everything | +| C/C++ toolchain | `cc --version && c++ --version` | `xcode-select --install` | `sudo apt install build-essential` | `highs` (HiGHS built from source) | +| cmake | `cmake --version` | `brew install cmake` | `sudo apt install cmake` | `highs` build | +| libclang | `ldconfig -p \| grep libclang` (Linux) | `xcode-select --install` | `sudo apt install libclang-dev` | `highs-sys` bindgen | +| rust (stable) | `rustc --version` | `curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs \| sh` | same | build, test | +| clippy, rustfmt | `cargo clippy --version && rustfmt --version` | `rustup component add clippy rustfmt` | same | `make check` | +| make | `make --version` | `xcode-select --install` | `sudo apt install make` | all targets | +| python3 | `python3 --version` | `brew install python` | `sudo apt install python3` | `scripts/`, papers targets | +| uv | `uv --version` | `curl -LsSf https://astral.sh/uv/install.sh \| sh` | same | `make qubo-testdata`, script tests | +| jq | `jq --version` | `brew install jq` | `sudo apt install jq` | `make cli-demo`, `make compare` | +| node + npm (22) | `node --version && npm --version` | `brew install node@22` | NodeSource or `nvm install 22` | `make doc` / `mdbook` / `website` (`npm ci`) | +| mdbook | `mdbook --version` | `cargo install mdbook` | same | `make doc`, `make mdbook` | +| typst | `typst --version` | `brew install typst` | `cargo install --locked typst-cli` | `make paper`, `make diagrams` | +| cargo-llvm-cov | `cargo llvm-cov --version` | `cargo install cargo-llvm-cov && rustup component add llvm-tools-preview` | same | `make coverage` | -## Skill Tools (AI-assisted pipeline) +## Contributor workflow -| Tool | Check Command | Install (macOS) | Install (Linux) | Purpose | -|------|--------------|-----------------|-----------------|---------| -| gh | `gh --version` | `brew install gh` | `sudo apt install gh` | GitHub CLI | -| claude | `claude --version` | `npm install -g @anthropic-ai/claude-code` | same | AI-assisted pipeline | -| pred | `pred --version` | `make cli` or `cargo install --path problemreductions-cli` | same | Project CLI (check-issue, topology-sanity-check, propose) | +| Tool | Check | Install (macOS) | Install (Linux) | Needed by | +|------|-------|-----------------|-----------------|-----------| +| gh | `gh --version` | `brew install gh` | `sudo apt install gh` | issues, PRs, `propose` | +| pred, pred-sym | `pred --version && pred-sym --help` | `make cli` | same | `propose`, `find-solver`, `find-problem` | -## Optional Tools +## Optional -| Tool | Check Command | Install (macOS) | Install (Linux) | Purpose | -|------|--------------|-----------------|-----------------|---------| -| julia | `julia --version` | `brew install julia` | `sudo apt install julia` | Julia parity tests | +| Tool | Check | Install (macOS) | Install (Linux) | Needed by | +|------|-------|-----------------|-----------------|-----------| +| julia | `julia --version` | `brew install julia` | `curl -fsSL https://install.julialang.org \| sh` | `make jl-testdata` | +| rclone | `rclone version` | `brew install rclone` | `sudo apt install rclone` | `make papers-push` / `papers-pull` | +| claude | `claude --version` | `npm install -g @anthropic-ai/claude-code` | same | running repo skills | diff --git a/.claude/skills/final-review/SKILL.md b/.claude/skills/final-review/SKILL.md deleted file mode 100644 index ced146b7d..000000000 --- a/.claude/skills/final-review/SKILL.md +++ /dev/null @@ -1,456 +0,0 @@ ---- -name: final-review -description: Interactive maintainer review for PRs in "Final review" column — assess usefulness, safety, completeness, quality ranking, then merge or hold ---- - -# Final Review - -Interactive review with the maintainer for PRs in the `Final review` column on the [GitHub Project board](https://github.com/orgs/CodingThrust/projects/8/views/1). The goal is to decide whether to **merge**, put **OnHold** (with reason), or **quick fix** before merging. - -**Rules:** -- Every `AskUserQuestion` must include your recommendation (e.g., "My recommendation: **Merge** — clean implementation with full coverage"). -- **Skip questions when no issues found.** If a check (usefulness, safety, completeness) finds no concerns, report the positive result and continue to the next step without asking the reviewer. Only use `AskUserQuestion` when there are findings that need the reviewer's input or when the recommendation is not clearly positive. -- **Issue presentation format** — whenever reporting an issue (in any step), use this format: - > **N. [Short title]** (`file:line`) - > ```rust - > // 5-15 lines showing the problem - > ``` - > - **Why**: What's wrong and why it matters — assume the reviewer hasn't seen the code - > - **Suggested fix**: Concrete action or code sketch - -## Invocation - -- `/final-review` -- pick the first PR from "Final review" column -- `/final-review 42` -- review a specific PR number - -## Constants - -GitHub Project board IDs (for `gh project item-edit`): - -| Constant | Value | -|----------|-------| -| `PROJECT_ID` | `PVT_kwDOBrtarc4BRNVy` | -| `STATUS_FIELD_ID` | `PVTSSF_lADOBrtarc4BRNVyzg_GmQc` | -| `STATUS_FINAL_REVIEW` | `51a3d8bb` | -| `STATUS_ON_HOLD` | `48dfe446` | -| `STATUS_DONE` | `6aca54fa` | - -## Workflow - -### Step 0: Select PR and Create Worktree - -**0a. Select the PR** from the Final review column: - -```bash -REPO=$(gh repo view --json nameWithOwner --jq .nameWithOwner) -``` - -If a specific PR number was provided (`/final-review 42`), use it directly. Otherwise, find the first item from the Final review column: - -```bash -# List items in Final review column (returns issues, not PRs) -python3 scripts/pipeline_board.py list final-review --repo "$REPO" --format json -``` - -The board tracks **issues**, not PRs. For each issue in the list, find the associated **open** PR: - -```bash -gh pr list --repo "$REPO" --state open --search "Fix #" --json number,title,headRefName -``` - -**Skip stale items**: if an issue has no open PR (the PR was already merged or doesn't exist), move it to Done (`python3 scripts/pipeline_board.py move done`) and try the next issue. Do not spend time investigating why the board item is stale. - -Extract the `PR` number from the first matched open PR and its `ITEM_ID` from the board list. - -**Claim the item** by moving it to OnHold immediately, so no other reviewer picks it up: - -```bash -python3 scripts/pipeline_board.py move on-hold -``` - -If the review completes successfully the item moves to Done; if abandoned, move it back to Final review (`python3 scripts/pipeline_board.py move final-review`). - -**0b. Create worktree and check out the PR branch:** - -```bash -REPO_ROOT=$(pwd) -WORKTREE_JSON=$(python3 scripts/pipeline_worktree.py enter --name "final-review-pr-$PR" --format json) -WORKTREE_DIR=$(printf '%s\n' "$WORKTREE_JSON" | python3 -c "import sys,json; print(json.load(sys.stdin)['worktree_dir'])") -cd "$WORKTREE_DIR" -gh pr checkout "$PR" -``` - -**0c. Resolve conflicts with main** (commit only, do not push yet — push happens in Step 8 after the review): - -```bash -git fetch origin main -git merge origin/main --no-edit -``` - -- **Conflicted** (common case) — most conflicts in this codebase are "both sides added new entries in ordered lists" (in `mod.rs`, `lib.rs`, `create.rs`, `dispatch.rs`, `reductions.typ`). These are mechanical: keep both sides, maintain alphabetical order. Delegate to a subagent for resolution, then run `make check && make paper` to verify. For `.bib` conflicts, also check for duplicate BibTeX keys (`grep '^@' docs/paper/references.bib | sed 's/@[a-z]*{//' | sed 's/,$//' | sort | uniq -d`) and remove duplicates. Continue with the review; if conflicts are too complex, decide whether to hold in Step 5. -- **Clean merge** — merge commit is ready locally. Run `make check && make paper` to verify the merge didn't break anything (API incompatibilities, formatting, test failures, paper compilation). If any step fails, fix the errors before proceeding. -- **Merge failed** — note the error and continue. - -**0d. Sanity check**: verify the diff touches `src/models/` or `src/rules/` (for model/rule PRs). If the diff only contains unrelated files, STOP and flag the mismatch. - -### Step 1: Gather Context - -Use `gh` commands to get the PR's actual data — always from GitHub, never from local git state: - -```bash -# PR diff (always correct, regardless of local state) -gh pr diff "$PR" - -# PR metadata -gh pr view "$PR" --json title,body,author,headRefName,baseRefName,labels,comments,reviews - -# Linked issue (if any — extract from PR body "Fix #N" / "Close #N") -gh issue view --json title,body,labels -``` - -Also run the review-implementation context inside the worktree for deterministic checks: - -```bash -IMPL_REPORT=$(python3 scripts/pipeline_skill_context.py review-implementation --repo-root . --format text) -printf '%s\n' "$IMPL_REPORT" -``` - -If any deterministic check shows `fail` or `skipped` without clear explanation, fall back to `gh pr diff "$PR"` for the actual PR file list and perform the corresponding check manually against those files only (not the full merge range). The `review-implementation` script may scope to base→HEAD across the full merge, which includes unrelated changes from main. Do not silently accept partial results. - -### Step 1b: Walk Through Agentic Review Findings - -The `review-pipeline` skill already posted a structured **Review Pipeline Report** as a PR comment (structural check, quality check, agentic feature tests). Read it: - -```bash -gh pr view "$PR" --json comments --jq '.comments[] | select(.body | contains("Review Pipeline Report")) | .body' -``` - -**Do not re-evaluate from scratch.** Walk through each finding from the agentic review report with the reviewer: -- For each issue flagged: is it still present? Was it already fixed? Is it a false positive? -- For each "Remaining issues for final review" item: disposition it explicitly. - -Prepare a short summary: - -> **Agentic Review Walk-Through** -> -> [N findings reviewed: X addressed, Y still open, Z false positive] -> -> Still open: [list each using the issue presentation format] - -If no agentic review report exists (or the report is poorly structured / missing key sections), note this to the reviewer and perform the checks in Steps 2–5 yourself from scratch — read the full diff, verify correctness, check completeness, and assess quality as if no prior review had been done. - -### Step 2: Usefulness assessment - -Think critically about whether this model/rule is genuinely useful. Consider: - -- **For models**: Is this problem well-known in the literature? Does it connect to existing problems via reductions? Is it a trivial variant of something already implemented? Would researchers or practitioners actually use this? -- **For rules**: Is this reduction well-known? Is it non-trivial (not just a relabeling)? Does it strengthen the reduction graph connectivity? Is the overhead reasonable? - -Present your assessment to the reviewer: - -> **Usefulness Assessment** -> -> [Your reasoning — 2-3 sentences with specific justification] -> -> Verdict: [Useful / Marginal / Not useful] - -Use `AskUserQuestion` with your recommendation: - -> My recommendation: **[Useful / Marginal / Not useful]** — [one-sentence justification] -> -> **Do you agree with this usefulness assessment?** -> - "Agree" — continue review -> - "Not useful, hold" — keep OnHold, post reason (skip remaining steps, go to Step 8) -> - Reviewer provides their own reasoning — agent updates verdict accordingly and continues -> - "Skip" — skip this check - -### Step 3: Safety check - -Scan the PR diff for dangerous actions: - -- **Blacklisted files**: If the diff touches `docs/src/reductions/reduction_graph.json` or `docs/src/reductions/problem_schemas.json`, **block merge**. These files are auto-generated and must not be committed in PRs — they are rebuilt by CI/`make doc`/`make paper`. Flag immediately and recommend OnHold. -- **Removed features**: Any existing model, rule, test, or example deleted? -- **Unrelated changes**: Files modified that don't belong to this PR (e.g., changes to unrelated models/rules, CI config, Cargo.toml dependency changes not needed for this PR) -- **Force push indicators**: Any sign of history rewriting -- **Broad modifications**: Changes to core traits, macros, or shared infrastructure that could affect other features -- **No committed `examples.json`**: The example database is generated on demand by `make paper` (via `export_examples`). Do not commit the gitignored `docs/paper/data/examples.json` build artifact. - -Report findings with fix options for each concern: - -> **Safety Check** -> -> [If no concerns: "No safety issues found."] -> -> [If concerns found, for each one:] -> **1. [Short title]** (`file:line`) -> - **What**: [Describe the concern and why it matters] -> - **Suggested fix**: [Concrete action — e.g., "revert this file", "split into separate PR", "remove unrelated hunk"] -> - **Recommendation**: [Block merge / Quick fix / Acceptable — with reasoning] - -Use `AskUserQuestion` with your recommendation: - -> My recommendation: **[Safe / Fix needed before merge]** — [one-sentence justification] -> -> **Do you agree with the safety assessment?** -> - "Agree" — continue -> - "Unsafe, hold" — keep OnHold, post reason (skip remaining steps, go to Step 8) -> - Reviewer flags additional concerns — agent adds them to the fix plan -> - "Skip" — skip this check - -### Step 3b: File whitelist check - -Use the `IMPL_REPORT`'s deterministic checks section. If files fall outside the whitelist, flag it: - -> **File Whitelist Check** -> -> Found N file(s) outside expected whitelist: -> - `path/to/file` — [what it does, why it may not belong] - -If all files are whitelisted, report "All files within expected whitelist" and continue. - -### Step 4: Completeness and correctness check - -Use the `IMPL_REPORT`'s deterministic checks as the baseline checklist. Then apply maintainer judgment. - -Verify the PR includes all required components: - -**For [Model] PRs:** -- [ ] Model implementation (`src/models/...`) -- [ ] Unit tests (`src/unit_tests/models/...`) -- [ ] Variant registration exists: usually `declare_variants!` with the intended default variant, or justified manual `VariantEntry` wiring for unusual dynamic-registration work -- [ ] Schema / registry entry for CLI-facing model creation (`ProblemSchemaEntry`) -- [ ] `canonical_model_example_specs()` function in the model file (gated by `#[cfg(feature = "example-db")]`) and registered in the category `mod.rs` example chain -- [ ] Paper section in `docs/paper/reductions.typ` (`problem-def` entry) -- [ ] `display-name` entry in paper -- [ ] Aliases: if provided, verify they are standard literature abbreviations (not made up); if empty, confirm no well-known abbreviation is missing; check no conflict with existing aliases -- [ ] After merge-with-main, verify no stale manual dispatch arms remain in `dispatch.rs` or `problem_name.rs` (main uses registry-based dispatch via `load_dyn` / `find_problem_type_by_alias`) - -**For [Rule] PRs:** -- [ ] Reduction implementation (`src/rules/...`) -- [ ] `src/rules/mod.rs` registration -- [ ] Unit tests (`src/unit_tests/rules/...`) -- [ ] `#[reduction(overhead = {...})]` with correct expressions -- [ ] Uses only the `overhead` form of `#[reduction]` -- [ ] Canonical rule example function in the rule file -- [ ] Paper section in `docs/paper/reductions.typ` (`reduction-rule` entry) - -**Paper-example consistency check (both Model and Rule PRs):** - -The paper example must use data from the canonical example database (generated on demand by `make paper` via `export_examples`), not hand-written data. To verify: -1. If the PR changes example specs, run `make paper` to regenerate `docs/paper/data/examples.json`. -2. For **[Rule] PRs**: the paper's `reduction-rule` entry must call `load-example(source, target, ...)` (defined in `reductions.typ`) to load the canonical example from `examples.json`, and derive all concrete values from the loaded data using Typst array operations — no hand-written instance data. -3. For **[Model] PRs**: run the export and read the problem's entry in the generated `examples.json` under `models`, compare its `instance` field against the paper's `problem-def` example. The paper example must use the same instance (allowing 0-indexed JSON vs 1-indexed math notation). If they differ, flag: "Paper example does not match `example_db` canonical instance." - -**Issue–test round-trip consistency check (both Model and Rule PRs):** - -The unit test's example instance and expected solution must match the issue's example: - -1. **Instance match**: The unit test's `example_instance()` (or equivalent setup) must construct the same graph/weights/parameters as described in the issue's "Example Instance" section. -2. **Solution match**: The expected optimal value in the test must equal the issue's stated optimal. -3. **Brute-force verification**: A brute-force test must exist that independently confirms the expected optimum, not just assert a hardcoded value. - -Report missing items: - -> **Completeness Check** -> -> [Checklist with pass/fail for each item] -> Missing: [list missing items, or "None — all complete"] - -For each missing item, describe what's missing, why it matters, and propose a concrete fix (e.g., "add a `test_evaluate_optimal` test", "register in `mod.rs`"). - -Use `AskUserQuestion` with your recommendation: - -> My recommendation: **[Complete / Fixable during review / Incomplete]** — [one-sentence justification] -> -> **Is the completeness acceptable?** -> - "Agree" — continue -> - "Incomplete, hold" — keep OnHold, post reason (skip remaining steps, go to Step 8) -> - Reviewer flags additional gaps or overrides — agent updates the fix plan accordingly -> - "Skip" — skip this check - -### Step 4b: Issue consistency check - -**This is a critical step.** Read the linked issue and the implementation side by side, verifying each of the following. Do NOT delegate this to a subagent — read the actual PR files yourself. - -| Check | What to compare | -|-------|-----------------| -| **Mathematical definition** | Does `evaluate()` implement exactly the condition stated in the issue? Watch for subtle differences (e.g., "at least 2 of 3 edges" vs "all 3 edges"). | -| **Variable encoding** | Does `dims()` match the issue's `dims()` specification? | -| **Complexity string** | Does `declare_variants!` use the same complexity expression as the issue? | -| **Test YES instance** | Does a unit test construct the same graph/weights/parameters as the issue's YES example? | -| **Test NO instance** | Does a unit test construct the same instance as the issue's NO example (if provided)? | -| **Solution match** | Does the test verify the same partition/solution as the issue's stated answer? | -| **Paper definition** | Does the paper's `problem-def` match the issue's mathematical definition? | -| **Brute-force verification** | Does a brute-force test independently confirm the expected answer? | - -Report results: - -> **Issue Consistency Check** -> -> | Check | Result | -> |-------|--------| -> | Mathematical definition | [Match / Mismatch — details] | -> | Variable encoding | [Match / Mismatch] | -> | Complexity string | [Match / Mismatch] | -> | Test YES instance | [Match / Mismatch] | -> | Test NO instance | [Match / Mismatch / N/A] | -> | Solution match | [Match / Mismatch] | -> | Paper definition | [Match / Mismatch] | -> | Brute-force verification | [Present / Missing] | - -If any mismatch is found, use `AskUserQuestion` with your recommendation and the specific discrepancy. If all checks pass, report the positive result and continue. - -### Step 5: Quality review - -Review the PR's code quality. Focus on issues that matter, not percentile scores. - -**Do not assume the reviewer knows the context.** The reviewer may not have looked at this PR before. Each issue must be self-contained and actionable. - -Present to reviewer: - -> **Quality Review — Merge confidence: [High / Medium / Low]** -> -> **Reason**: [1-2 sentences explaining why this confidence level, referencing specific strengths or concerns] -> -> Based on: mathematical/algorithmic correctness, code quality, and writing quality (paper, docs, comments). -> - **High**: Correct, clean, well-written — ready to merge as-is or with minor follow-ups. -> - **Medium**: Mostly correct but has quality or writing issues fixable during this review. -> - **Low**: Correctness concerns or significant quality problems that may need rework. -> -> Strengths: -> - [bullet points] -> -> Issues: [list each using the issue presentation format, plus for each:] -> - **Pros/cons**: Tradeoffs of fixing now vs deferring -> - **Recommendation**: Quick fix / Record for follow-up / Informational only -> -> Notable observations: [optional — unusual design choices, clever techniques, or patterns that diverge from the codebase. Omit if nothing stands out.] - -### Step 6: Triage and fix - -Classify each issue from Steps 1b–5 into two categories: - -**Mechanical fixes** — issues with an obvious, unambiguous fix (e.g., alphabetical ordering, `cargo fmt`, removing blacklisted files, updating a renamed API call, fixing a merge conflict). Apply these immediately without asking. Commit and report what was done. - -**Uncertain fixes** — issues where the correct action is unclear, involves design judgment, or could change behavior (e.g., incorrect algorithm, missing edge case handling, questionable mathematical correctness, ambiguous issue compliance). Present these to the reviewer using `AskUserQuestion`. - -Summarize all findings from Steps 1b–5: - -| Aspect | Result | -|--------|--------| -| Agentic review | [N open / all addressed] | -| Usefulness | [Useful/Marginal/Not useful] | -| Safety | [Safe/Concerns found] | -| Completeness | [Complete/Missing: X, Y] | -| Merge confidence | [High/Medium/Low] | -| PR URL | [link] | - -Then report: - -> **Mechanical fixes applied:** -> - [list each fix already committed, or "None"] -> -> **Issues requiring reviewer input:** -> For each uncertain issue, present fix options with pros/cons and a recommendation: -> **N. [Short title]** (`file:line`) -> - **Option A**: [description] — Pros: ... / Cons: ... -> - **Option B**: [description] — Pros: ... / Cons: ... -> - **Recommendation**: Option [X] — [one-sentence justification] -> -> If no uncertain issues: "No uncertain issues." - -If there are uncertain issues, use `AskUserQuestion` with your recommendations for the reviewer to confirm or override. - -If all issues were mechanical and already fixed, skip the question and proceed to Step 7. - -### Step 7: Final decision - -If all issues were mechanical and already fixed (no uncertain issues remain), skip the question and proceed directly to Step 8 (Push and fix CI). Only ask the reviewer when there are uncertain issues or when the recommendation is OnHold. - -Use `AskUserQuestion` only when needed: - -> My recommendation: **[Push and fix CI / OnHold]** — [one-sentence justification] -> -> **Final decision:** -> - **1** — Push and fix CI: push all commits, fix any CI failures, then present the merge link -> - **2** — OnHold: keep in OnHold column, post reason - -### Step 8: Execute decision - -**If Push and fix CI:** -1. **Pre-push verification** — run `make check && make paper` locally before pushing. Both must pass. Fix any failures, commit, then push: - ```bash - cd - make check && make paper - git push - ``` -2. Wait for CI — but **do not poll excessively**. Check once after ~60 seconds. If CI hasn't triggered or is still pending, note "CI pending, local checks pass" when presenting the merge link (since pre-push already verified). The reviewer can admin-merge or wait for CI at their discretion. Do not spend more than 1–2 CI poll attempts. If CI does run and fails, fix the issues, commit, and push again. -3. If any follow-up items were noted during the review, post them as a comment: - ```bash - COMMENT_FILE=$(mktemp) - cat > "$COMMENT_FILE" <<'EOF' - **Follow-up items** (recorded during final review): - - [item 1] - - [item 2] - EOF - python3 scripts/pipeline_pr.py comment --repo "$REPO" --pr "" --body-file "$COMMENT_FILE" - rm -f "$COMMENT_FILE" - ``` -4. Approve the PR (may fail if you are the PR author — that's OK): - ```bash - gh pr review --approve || true - ``` -5. Post a community call validation checklist as a comment on the **linked issue** (not the PR). Substitute actual problem names from the PR diff (no angle-bracket placeholders). Example for a rule PR adding `Satisfiability` → `MaximumIndependentSet`: - ````bash - COMMENT_FILE=$(mktemp) - cat > "$COMMENT_FILE" <<'EOF' - Please kindly check the following items (PR #123): - - [ ] **Paper** ([PDF](https://github.com/CodingThrust/problem-reductions/blob/main/docs/paper/reductions.pdf)): check definition, proof sketch, example figure, and reproducible `pred` commands - - [ ] **Implementation (Optional)**: spot-check the source files changed in this PR for correctness - - 💬 Join the discussion on [Zulip](https://julialang.zulipchat.com/#narrow/channel/365542-problem-reductions) — feel free to ask questions or leave feedback there. - EOF - gh issue comment --body-file "$COMMENT_FILE" - rm -f "$COMMENT_FILE" - ```` - If there is no linked issue, post the checklist as a PR comment instead. -6. Present the PR link for the reviewer to merge: - > CI green, commits pushed, PR approved. Community call checklist posted on #. Please merge when ready: - > **** -7. After the reviewer merges, use `AskUserQuestion` to confirm: - > **Merged? (continue to move card & cleanup worktree)** Once confirmed, I will move the board item to Done and clean up the worktree. - > - "Yes" — proceed with cleanup -8. Move the project board item to `Done` and clean up: - ```bash - python3 scripts/pipeline_board.py move done - cd "$REPO_ROOT" - python3 scripts/pipeline_worktree.py cleanup --worktree "$WORKTREE_DIR" - ``` - -**If OnHold:** -The item is already in the OnHold column (claimed in Step 0a), so no board move is needed. -1. Ask the reviewer for the reason (use `AskUserQuestion` with free text). -2. Post a comment on the PR (or linked issue) with the reason: - ```bash - COMMENT_FILE=$(mktemp) - printf '**On Hold**: %s\n' "" > "$COMMENT_FILE" - python3 scripts/pipeline_pr.py comment --repo "$REPO" --pr "" --body-file "$COMMENT_FILE" - rm -f "$COMMENT_FILE" - ``` -3. Clean up the worktree: - ```bash - cd "$REPO_ROOT" - python3 scripts/pipeline_worktree.py cleanup --worktree "$WORKTREE_DIR" - ``` - -## Pipeline Script Subcommands - -Only use subcommands that exist. Available subcommands per script: - -| Script | Subcommands | -|--------|-------------| -| `pipeline_board.py` | `next`, `claim-next`, `ack`, `list`, `move`, `backlog` | -| `pipeline_pr.py` | `context`, `current`, `snapshot`, `comments`, `ci`, `wait-ci`, `codecov`, `linked-issue`, `create`, `comment`, `edit-body` | -| `pipeline_worktree.py` | `enter`, `create-issue`, `prepare-issue-branch`, `checkout-pr`, `prepare-review`, `merge-main`, `cleanup` | -| `pipeline_skill_context.py` | `review-pipeline`, `final-review`, `review-implementation`, `project-pipeline` | -| `pipeline_checks.py` | `detect-scope`, `file-whitelist`, `completeness`, `review-context`, `issue-guards`, `issue-context` | diff --git a/.claude/skills/find-problem/SKILL.md b/.claude/skills/find-problem/SKILL.md index d1695524a..80ea4fcb5 100644 --- a/.claude/skills/find-problem/SKILL.md +++ b/.claude/skills/find-problem/SKILL.md @@ -1,198 +1,87 @@ --- name: find-problem -description: Reverse of find-solver — given a solver for a model, discover what other problems it can handle via incoming reductions, ranked by effective complexity +description: Use when someone has (or plans) a solver for one model and wants to know which other problems it can handle through incoming reductions — ranks those problems by effective complexity against their best-known bounds and writes a solution doc --- # Find Problem -Given a solver for a specific model, discover what other problems it can handle by exploring the reduction graph in the incoming direction. Produces a solution doc ranking all reachable problems by effective complexity. +Reverse of `find-solver`: given a solver for model `M` with complexity `T(M's parameters)`, +find the problems that reduce to `M`, compute what the solver costs on each, and document the +worthwhile ones in `docs/solutions/`. -## Invocation - -``` -/find-problem — start from Step 1 (identify solver) -/find-problem — skip model identification, ask for complexity -``` +Invocation: `/find-problem` or `/find-problem `. -Do NOT modify project source files, write Rust code, or create PRs. -Only outputs: `pred` CLI commands executed live, web searches, conversational commentary, and one solution doc in `docs/solutions/`. -If the user asks about contributing code, point them to `/add-model`, `/add-rule`, or `/propose`. +No source edits, no Rust, no PRs. The only file written is the solution doc, after the user +confirms its filename. For contributions, point to `/propose`. -## Audience - -Users who have built or have access to a solver for a specific problem model and want to understand the full scope of problems their solver can handle through reductions. - -## Flow Overview - -``` -Step 1: Identify Solver (user provides model + complexity) -Step 2: Discover Reachable Problems (pred from --hops 3, compute effective complexity) -Step 3: Rank and Present (table ranked by effective complexity, web search for applications) -Step 4: Generate Solution Doc (docs/solutions/.md) -``` - -## Prerequisites - -Build the CLI tools before starting: `make cli` (builds `pred` and `pred-sym`). All commands below assume `pred` and `pred-sym` are available via `cargo run -p problemreductions-cli --bin pred --` and `cargo run -p problemreductions-cli --bin pred-sym --` respectively. - -## CRITICAL: Output Visibility - -Bash tool results are hidden from the user in the Claude Code UI. **After every `pred` / `pred-sym` command, you MUST copy-paste the full stdout/stderr into your response as text.** The pattern for every command is: - -1. Announce the command and why: "Let me run `pred to MIS --hops 3` to discover all problems that can reduce to MIS:" -2. Run the command via the Bash tool -3. Copy the COMPLETE output into your text response inside a fenced code block -4. Then add your brief explanation - -Never skip step 1 or 3. - ---- - -## Step 1: Identify Solver - -**Goal:** Get the user's model name and solver complexity. - -**If invoked as `/find-problem `:** validate with `pred show `. If it exists, show the output (including parameters), then ask for solver complexity. - -**If invoked as `/find-problem`:** ask using `AskUserQuestion`: "Which problem model does your solver handle?" Validate the answer with `pred show`. - -**Ask for complexity** using `AskUserQuestion`: "What is your solver's time complexity? Use the size field names from the output above (e.g., `O(1.1996^num_vertices)`, `O(2^(num_variables/3))`)." - -- Variable names should match the model's parameters shown in `pred show` output -- If the user gives informal notation (e.g., "exponential in n"), help them formalize it using the model's actual size field names - -**Exit condition:** Validated model name + complexity expression with variables matching the model's parameters. Proceed to Step 2. - ---- - -## Step 2: Discover Reachable Problems - -**Goal:** Find all problems that can reduce to the user's model and compute effective complexity for each. - -**Actions:** - -1. **Run `pred to --hops 3`** to find all problems that can reduce to the user's model within 3 hops (incoming direction). Copy-paste the full output. - -2. **For each discovered problem**, run: - - `pred path ` — get the cheapest witness-capable reduction path - - **IMPORTANT:** Use the exact variant-qualified name from `pred to` output (e.g., `SpinGlass/SimpleGraph/f64`, not bare `SpinGlass`). Bare names resolve to the default variant, which may differ from the reachable variant and cause false "no path" errors. - - `pred show ` — get best-known brute-force complexity +**Output visibility:** Bash output is hidden from the user. For every `pred` / `pred-sym` +command, say why, paste the full output in a fenced block, then interpret briefly. Never invent +output. Ask one question at a time and mark a recommendation. -3. **Compute effective complexity** for each source problem: - - Take the user's solver complexity expression (e.g., `O(1.1996^num_vertices)`) - - Substitute the overhead expressions from the reduction path into the solver's variables - - Example: if MVC→MIS has overhead `num_vertices = num_vertices`, then solving MVC via MIS costs `O(1.1996^num_vertices)` — same as MIS - - Example: if overhead is `num_vertices = num_clauses * 3`, then effective complexity is `O(1.1996^(3 * num_clauses))` - - **Use `pred-sym` to verify:** after manual substitution, run `pred-sym big-o ""` to normalize the expression. Use `pred-sym eval --vars ""` at a concrete size (e.g., n=20) to numerically verify the simplification. +`make cli` installs both `pred` and `pred-sym`; run it only if they are missing. -4. **Compare to best-known**: for each source, compare effective complexity to the source's own best-known complexity from `pred show`. Classify as: - - **Better** — effective complexity has a smaller base or exponent than best-known - - **Similar** — comparable asymptotic behavior - - **Worse** — effective complexity exceeds best-known (reduction overhead makes it impractical) - - **When effective and best-known use different variables** (e.g., `O(1.5^num_subsets)` vs `O(2^universe_size)`): this happens when a problem has multiple independent parameters and the best-known algorithm's dominant variable differs from the reduction overhead's. In this case, use `pred-sym eval` at representative concrete values to determine the comparison. State the result conditionally: "Better when num_subsets ≤ c·universe_size" with the crossover ratio. +## 1. Solver and complexity -5. **Web search** only the **Better** and **Similar** candidates for real-world applications (not the Worse ones). Use `WebSearch` tool with query " real-world applications". +Validate the model with `pred show ` and show its `parameters`. Get the solver's time +complexity written in those parameter names (e.g. `1.1996^num_vertices`); help formalize +informal answers ("exponential in n"). -**If `--hops 3` returns more than 15 results:** present only the top 10 by effective complexity and mention the rest are available if the user wants to see them. +## 2. Discover and score sources -**Proceed to Step 3.** - ---- - -## Step 3: Rank and Present - -**Goal:** Show all discovered problems ranked by practical usefulness. - -Present a ranked table (most practical first). **Mark a recommendation** — highlight the "Better" entries as the most valuable discoveries: - -| # | Problem | Hops | Overhead | Effective Complexity | vs Best-Known | Applications | -|---|---------|------|----------|---------------------|---------------|--------------| -| 1 | **MinimumVertexCover** | 1 | same size | O(1.1996^n) | **Better** | Network monitoring | -| 2 | **MaximumClique** | 2 | complement graph | O(1.1996^n) | **Better** | Social network cliques | -| 3 | GraphColoring | 3 | n^2 vars | O(1.1996^(n^2)) | Worse | Register allocation | - -Ask using `AskUserQuestion`: "Which problems would you like included in the solution doc? Pick numbers, or 'all practical' for only the Better/Similar ones." - -**Proceed to Step 4 with the selected problems.** - ---- - -## Step 4: Generate Solution Doc - -**Goal:** Write a static reference document listing all selected problems and how to solve them via the user's model. - -**File path:** `docs/solutions/problems-solvable-via--.md` - -Where: -- `` is the library model name (e.g., `MIS`, `QUBO`) -- `` is a short label for the user's solver (e.g., `custom-1.1996`, `ILP`) - -Ask the user to confirm the filename before writing. - -**Before writing the doc**, run `pred create --help` for each selected problem to verify the correct CLI flag names. Use the flags exactly as shown in the help output. - -**Doc template — write all sections:** +```bash +pred to --hops 3 # incoming neighbors (default is 1 hop) +pred path --limit 5 # candidate paths with per-step and Overall: formulas +pred show # the source's own best-known complexity +``` -```markdown -# Problems Solvable via () +- Use the exact variant-qualified names from `pred to` (e.g. `SpinGlass/SimpleGraph/f64`); bare + names resolve to the default variant and can give false "no path" results. +- `pred path` returns a set of paths, not a single cheapest one; pick the path with the best + composed overhead and say which. `unavailable` overhead formulas mean the effective + complexity cannot be computed symbolically; report that rather than guess. +- **Effective complexity** = the solver's expression with each `M` parameter replaced by the + path's `Overall:` formula in source parameters. Check every substitution: -## Overview +```bash +pred-sym big-o "1.1996^(3 * num_clauses)" # normal form +pred-sym eval --vars num_clauses=20 "1.1996^(3 * num_clauses)" +pred-sym compare "" "" # exits 1 when not Big-O equal +``` - +- Classify against the source's best-known bound: **Better**, **Similar**, or **Worse**. When + the two bounds use different variables (e.g. `1.5^num_subsets` vs `2^universe_size`), evaluate + both with `pred-sym eval` at representative sizes and state the crossover condition + ("better when num_subsets ≤ c·universe_size"). +- WebSearch real-world applications only for Better and Similar sources. -## Summary Table +If more than ~15 sources come back, show the top 10 by effective complexity and offer the rest. -| Problem | Hops | Overhead | Effective Complexity | vs Best-Known | Applications | -|---------|------|----------|---------------------|---------------|--------------| -| ... | ... | ... | ... | ... | ... | +## 3. Rank and choose -## -> +Present one table (problem, hops, overhead, effective complexity, vs best-known, applications), +Better entries bolded. The user picks which go into the doc. -- **What it is:** -- **Reduction path:** -> ... -> -- **Overhead:** -- **Effective complexity:** -- **vs best-known:** +## 4. Solution doc -### CLI Commands +Propose `docs/solutions/problems-solvable-via--.md` (`` a short label such +as `custom-1.1996` or `ILP`) and write it only after confirmation. Contents: overview and ranking +method, summary table, then per problem: what it is and where it appears, path, overhead, +effective complexity, comparison, and commands. Run `pred create --help` for each +selected source and use its real flags. Command pattern, verified against the current CLI: ```bash -# Create a source problem instance pred create -o input.json - -# Reduce to your solver's model -pred reduce input.json --to -o bundle.json - -# Solve (built-in ILP or your external solver) -pred solve bundle.json --solver ilp --timeout 60 -``` - -## -> - -... +pred path --json | jq '.paths[0]' > route.json # the entry you chose +pred reduce input.json --via route.json -o bundle.json # bundle.target is the M instance +pred solve bundle.json --timeout 30 # built-in solver, mapped back ``` -**After writing the doc:** - -1. Show the user the generated filename and a brief summary of what's in it. -2. **If a built-in solver covers the model** (brute-force or ILP), offer to run a live demo with one of the "Better" problems: "Want me to run an example end-to-end so you can see it in action?" -3. Ask if they want to make any changes before finishing. - ---- +For an external solver, feed it `bundle.target` and map its answer back with +`pred extract bundle.json --config ''`. Whether a built-in solver exists for +`M` is shown by `pred inspect` on an instance (`solver_capabilities`); reachability alone does +not imply one. -## Key Behaviors - -- **One question at a time.** Never ask multiple questions in one message. Use `AskUserQuestion` for every decision point. -- **Web search only Better/Similar candidates.** In Step 2, web search only the problems classified as Better or Similar for real-world use cases. Skip Worse ones unless the user asks for all. Never guess applications from internal knowledge alone. -- **Show full output.** After every Bash tool call, copy-paste the COMPLETE output into your text response as a fenced code block. Bash tool results are hidden in the UI. -- **Announce every command.** Before running, say what command you're using and why. -- **Always use variant-qualified names in `pred path`.** When `pred to` returns names like `SpinGlass/SimpleGraph/f64`, use that exact string in subsequent `pred path` calls. Bare names (e.g., `SpinGlass`) resolve to the default variant, which may differ from the reachable variant and cause false "no path" errors. -- **Recommend, don't just list.** When presenting the ranked table in Step 3, bold the "Better" entries as the most valuable discoveries. The user can still pick freely. -- **Compact formatting.** Write explanations as plain paragraphs. Do not use blockquote `>` syntax for explanations. Keep tight: command announcement, code block output, 1-3 sentence explanation. -- **Conversational tone.** Guided consultation, not a lecture. -- **Live execution.** Every `pred` command runs for real. No fake output. -- **Graceful fallbacks.** If `pred to` returns no results (no incoming reductions), suggest trying with more hops or a different model. If `pred path` fails for a specific source, skip it and note it in the table. -- **Help with complexity notation.** If the user gives informal complexity, show `pred show ` parameters and help them write a formal expression. -- **Cap results at 10.** If discovery returns many problems, show top 10 by effective complexity and offer to show more. +After writing, summarize the doc, offer a live demo on one Better source when a built-in solver +applies (`--timeout 30`), and ask for changes. diff --git a/.claude/skills/find-solver/SKILL.md b/.claude/skills/find-solver/SKILL.md index 704c76ced..bb5bb8217 100644 --- a/.claude/skills/find-solver/SKILL.md +++ b/.claude/skills/find-solver/SKILL.md @@ -1,317 +1,99 @@ --- name: find-solver -description: Interactive guide — match a real-world problem to a library model, explore reduction paths, recommend solvers (built-in + external), and generate a solution doc +description: Use when someone has a real-world or formally named problem and wants to know how to solve it with this library — matches it to a model, explores reduction paths, checks which built-in solvers actually apply, surveys external solvers, and writes a solution doc --- # Find Solver -Guide users from a real-world algorithmic problem to a concrete solving strategy. Produces a static solution doc in `docs/solutions/`. +Take the user from a problem ("assign jobs to machines, minimize makespan", or "MIS on unit-disk +graphs") to a concrete solving recipe, ending in one doc under `docs/solutions/`. -## Invocation - -``` -/find-solver — start from Step 1 (clarify problem) -/find-solver — skip to Step 3 (explore reductions for a known model) -``` +Invocation: `/find-solver` (start by clarifying the problem) or `/find-solver ` +(validate with `pred show` and go straight to reductions). -Do NOT modify project source files, write Rust code, or create PRs. -Only outputs: `pred` CLI commands executed live, web searches, conversational commentary, and one solution doc in `docs/solutions/`. -If the user asks about contributing code, point them to `/add-model`, `/add-rule`, or `/propose`. +No source edits, no Rust, no PRs. The only file written is the solution doc, after the user +confirms its filename. For contributions, point to `/propose`. -## Audience - -This skill serves two types of users: - -- **Practitioners** who have a fuzzy real-world problem (e.g., "I need to assign tasks to machines minimizing makespan") and need help identifying which NP-hard problem it maps to -- **Researchers** who already know the formal problem name and want to find the best reduction path + solver - -Adapt the flow: if the user provides a formal problem name, validate it with `pred show` and skip directly to Step 3. - -## Flow Overview - -``` -Step 1: Clarify Problem (skip if user knows the formal name) -Step 2: Match to Library Models (web search + pred list) -Step 3: Explore Reduction Paths (auto-explore via pred from --hops 3) -Step 4: Recommend Solvers (web search + pred solve options) -Step 5: Generate Solution Doc (docs/solutions/.md) -``` - -## CRITICAL: Output Visibility - -Bash tool results are hidden from the user in the Claude Code UI. **After every `pred` command, you MUST copy-paste the full stdout/stderr into your response as text.** The pattern for every command is: - -1. Announce the command and why: "Let me run `pred to MIS` to see what MIS can be reduced to:" -2. Run the command via the Bash tool -3. Copy the COMPLETE output into your text response inside a fenced code block -4. Then add your brief explanation +**Output visibility:** Bash output is hidden from the user. For every `pred` command, say why +you run it, then paste its full output in a fenced block, then 1-3 sentences of interpretation. +Every command runs for real; never invent output. -Never skip step 1 or 3. +Ask one question at a time and always mark a recommended choice. ---- - -## Step 1: Clarify Problem - -**Goal:** Understand the user's problem well enough to form a search query for matching models. - -**Researcher shortcut:** If the user invoked `/find-solver ` or describes their problem using a formal name (e.g., "maximum independent set", "graph coloring"), run `pred show ` to validate it exists. If it does, skip directly to Step 3 with that model. If it doesn't exist in the library, tell the user and fall through to Step 2. - -**For practitioners, ask one question at a time:** - -1. "Describe your problem in plain language — what are you trying to optimize, decide, or compute?" - -2. Based on the answer, if the input structure is ambiguous, ask: - "Is your input a graph/network? A set of items? A boolean formula? Numbers/matrices?" - -3. If the objective is unclear: - "What's the objective — minimize something, maximize something, or check feasibility?" - -4. "Roughly how large are your instances? (e.g., 10 nodes, 1000 variables)" - -Use `AskUserQuestion` for each question. Format options as **(a)**/**(b)**/**(c)** when multiple choice is natural. - -**Exit condition:** You have enough context to form a search query like "task scheduling minimize makespan NP-hard" or "maximum weight independent set unit disk graph". Proceed to Step 2. - ---- +## 1. Pin down the problem -## Step 2: Match to Library Models +For a fuzzy description, clarify only what is missing: input structure (graph, sets, formula, +numbers), objective (min, max, feasibility), typical instance size. Build `pred` first if +absent (`command -v pred || make cli`). -**Goal:** Identify which problem model(s) in the library match the user's problem. +## 2. Match to models -**Key rule:** Web search happens BEFORE presenting options. Never guess model matches from internal knowledge alone. - -**Actions:** - -1. **Web search** the clarified problem description together with terms like "NP-hard", "computational complexity", or "reduction" to find formal problem names and known relationships in the literature. Use `WebSearch` tool. - -2. **Search the catalog** with `pred list `. Use `pred list --json` when exhaustive machine-readable discovery is needed. Do not paste the full catalog into the response. - -3. **Cross-reference** the web search results against the catalog. For each candidate model that exists in the library (3-5 max), present a table: - -| # | Model | Why it might match | Caveat | -|---|-------|--------------------|--------| -| 1 | ... | ... | ... | -| 2 | ... | ... | ... | -| 3 | ... | ... | ... | - -**Include a recommendation:** Bold or mark the option you think is the best fit, with a brief reason why. - -4. For each candidate, run `pred show ` and show the output — fields, complexity, available reductions. This helps the user see what data they would need to provide. - -5. **Check optimization vs decision mismatch.** If the user's goal is "minimize X" or "maximize X" but the matched model is a decision/feasibility problem (Value = `Or`, fields include a `deadline`/`bound`), explain the gap: - - "This model checks feasibility ('can it be done within bound D?'), not optimization directly." - - "To find the optimum, we'll binary search on the bound parameter." - - This is common for scheduling problems (deadline), knapsack (bound), etc. - -6. **Ask the user to pick one** using `AskUserQuestion`. If none fit, ask the user for more detail and re-run the web search with refined keywords. - -**Proceed to Step 3 with the chosen model.** - ---- - -## Step 3: Explore Reduction Paths - -**Goal:** Discover all solver-ready targets reachable from the chosen model and present them ranked. - -**Actions:** - -1. **Run `pred from --hops 3`** to find all problems reachable via outgoing reductions within 3 hops. Copy-paste the full output. - -2. **For each reachable problem**, gather info: - - Run `pred path ` to get the cheapest witness-capable reduction path and composed overhead - - **IMPORTANT:** Use the exact variant-qualified name from `pred from` output (e.g., `SpinGlass/SimpleGraph/f64`, not bare `SpinGlass`). Bare names resolve to the default variant, which may differ from the reachable variant and cause false "no path" errors. - - Run `pred show ` to get its best-known complexity - - Check if it's a solver-ready target (ILP, QUBO, SAT) or has a path to one via `pred path ILP` - -3. **Present a ranked table** (most practical paths first — fewest hops, lowest overhead). **Mark a recommendation** for the most practical path: - - | # | Target | Hops | Composed Overhead | Target Complexity | Solver-Ready? | - |---|--------|------|-------------------|-------------------|---------------| - | 1 | ILP | 2 | num_vars = 2*n + m | O(2^num_vars) | Yes (is ILP) | - | 2 | QUBO | 1 | num_vars = n | O(2^num_vars) | Yes (is QUBO) | - | 3 | MaxSetPacking | 1 | num_sets = n | O(2^num_sets) | Yes (ILP in 2 steps) | - - When overhead grows significantly between options (e.g., linear vs quadratic), note the practical implication: "QUBO adds quadratic variable blowup — prefer this only if targeting quantum/annealing hardware." - -4. **Ask the user** using `AskUserQuestion`: "Which reduction path would you like to use? Pick a number." - -**If `pred from --hops 3` returns more than 15 results:** present only the top 10 by overhead and mention the rest are available. - -**Proceed to Step 4 with the chosen path.** - ---- +WebSearch the clarified problem ("… NP-hard", "… reduction") before proposing matches; don't +guess from memory. Search the catalog with `pred list ` (`pred list --category graph`, +`pred list --all` for more). Present 3-5 candidates with why/caveat, run `pred show ` on +each so the user sees the required input fields and complexity, and let them pick. -## Step 4: Recommend Solvers +**Optimization vs decision mismatch.** If the user wants to minimize/maximize but the model is a +feasibility problem (value `Or`, a `bound`/`deadline` field), say so: the optimum needs a binary +search over the bound, and the doc gets a "Finding the optimum" section. -**Goal:** Find the best solver options — both built-in and external — for the target problem. - -**Key rule:** Web search happens BEFORE presenting solver options. Do not recommend solvers from internal knowledge alone. - -**Actions:** - -1. **Web search** the final target problem + "solver" + "benchmark" + "library" to find state-of-the-art external tools. Use `WebSearch` tool. Example queries: - - "QUBO solver open source benchmark" - - "integer linear programming solver comparison" - - "maximum independent set practical solver" - -2. **Check built-in solver availability:** - - Run `pred path ILP` — if a path exists, `pred solve --solver ilp` is available - - `pred solve --solver brute-force` is always available (feasible for small instances, ~25 variables) - -3. **Present solver options** in a table: - - | # | Solver | Type | How to Use | Notes | - |---|--------|------|------------|-------| - | 1 | pred solve --solver ilp | Built-in | pred solve reduced.json | HiGHS backend | - | 2 | pred solve --solver brute-force | Built-in | pred solve problem.json --solver brute-force | Exact, small instances only | - | 3 | (from web search) | External | (brief setup) | (strengths/limitations) | - | 4 | (from web search) | External | (brief setup) | (strengths/limitations) | - -4. **Ask the user** which solver(s) to include in the solution doc using `AskUserQuestion`. - -**Proceed to Step 5 with the selected solver(s).** - ---- - -## Step 5: Generate Solution Doc - -**Goal:** Write a static reference document with everything the user needs to solve their problem. - -**File path:** `docs/solutions/-via--.md` - -Where: -- `` is a short kebab-case description of the real-world problem -- `` is the library model name (e.g., `MIS`, `QUBO`) -- `` is the primary solver (e.g., `ILP`, `brute-force`, `gurobi`) - -Ask the user to confirm the filename before writing. - -**Doc template — write all sections:** - -```markdown -# via -> - -## Problem Description - - - -## Matched Model - -- **Name:** {variant} -- **Why this model:** -- **Best-known complexity:** O(...) - -## Input Schema - -`, with field explanations.> - -Example instance: - -```json -` and paste the JSON output> -``` - -## Reduction Path - - - -### Step N: -> - -- **Overhead:** +## 3. Explore reduction paths ```bash -pred reduce input.json --to -o step_N.json +pred from --hops 3 # reachable targets +pred path --limit 5 # a set of candidate paths with composed parameters +pred show # target complexity ``` -## Solving +- Use the exact variant-qualified names printed by `pred from` (e.g. `SpinGlass/SimpleGraph/f64`). + Bare names resolve to the default variant and can produce false "no path" results. +- `pred path` returns several paths, not a ranked best. Compare them yourself by hops and the + `Overall:` parameter formulas; some formulas are `unavailable`, say so rather than guess. +- If more than ~15 targets come back, show the top 10 and offer the rest. -```bash -pred solve step_N.json --solver ilp --timeout 60 -``` +Present a table (target, hops, composed overhead, target complexity) with a recommendation and +note blowups (e.g. quadratic QUBO variables only pay off on annealing hardware). - +## 4. Solvers -## Finding the Optimum (decision models only) - - - -The model checks feasibility ("can it be done within bound D?"), not optimization directly. -To find the minimum/maximum, binary search on the bound parameter: +**Reachability does not imply a built-in solver.** Solver dispatch uses only registered +customized solvers and fixed ILP pipelines. Check with a concrete instance: ```bash -# Binary search for minimum deadline -# Upper bound: sum of all task lengths (trivially feasible with 1 processor) -# Lower bound: max(longest task, ceil(total / num_processors)) -# Try midpoint, narrow based on Or(true)/Or(false) +pred create --example -o ex.json # or: pred create +pred inspect ex.json # solver_capabilities: brute_force, customized, ilp (with its fixed path) ``` -## Solution Extraction - - +`pred solve --solver` accepts `customized`, `ilp`, `brute-force`; omit it for the default +(customized, else ILP, else brute force). Brute force is exact but only for small instances +(~25 variables). ILP uses HiGHS. Also inspect the final target of the chosen path. -Using the reduction bundle workflow (recommended): - -```bash -pred reduce input.json --to -o bundle.json -pred solve bundle.json --solver ilp --timeout 60 -``` +Then WebSearch " solver benchmark open source" for external tools, and present built-in +and external options together; the user picks which go into the doc. -The solver automatically extracts the solution back to the original problem space. +## 5. Solution doc -## External Solver Alternatives +Propose `docs/solutions/-via--.md` (kebab-case problem description) and +write it only after confirmation. Sections: problem description; matched model (variant, why, +complexity); input schema from `pred show --json` plus an example instance; the chosen +reduction path with per-step overhead; solving and interpreting the output (`Max(3)`, +`Or(true)`); finding the optimum (decision models only); external alternatives (if chosen); a +quick-reference command block. Use real flags from `pred create --help`. - - -### - -- **What:** -- **When to prefer:** -- **How to use:** - -## Quick Reference - -All commands in sequence: +The explicit-path workflow, verified against the current CLI: ```bash -# 1. Create your problem instance pred create -o input.json - -# 2. Reduce to solver-ready form -pred reduce input.json --to -o bundle.json - -# 3. Solve -pred solve bundle.json --solver ilp --timeout 60 - -# 4. Verify (optional) -pred evaluate input.json --config -``` +pred path --json | jq '.paths[0]' > route.json # pick the entry you chose +pred reduce input.json --via route.json -o bundle.json +pred solve bundle.json --timeout 30 # solves the target, maps the solution back +pred evaluate input.json --config '' ``` -**After writing the doc:** - -1. Show the user the generated filename and a brief summary of what's in it. -2. **If a built-in solver covers the chosen path** (brute-force or ILP), offer to run a live demo with the example instance: "Want me to run the example end-to-end so you can see it in action?" -3. Ask if they want to make any changes before finishing. - ---- - -## Key Behaviors +`--via` takes exactly one path entry, not the whole `{paths, truncated}` set. If the model has a +registered ILP pipeline, `pred solve input.json --solver ilp` is the one-step alternative. -- **One question at a time.** Never ask multiple questions in one message. Use `AskUserQuestion` for every decision point. -- **Web search before recommendations.** In Step 2 (model matching) and Step 4 (solver recommendation), always web search first. Never rely on internal knowledge alone. -- **Show full output.** After every Bash tool call, copy-paste the COMPLETE output into your text response as a fenced code block. Bash tool results are hidden in the UI. -- **Announce every command.** Before running, say what command you're using and why. -- **Always use variant-qualified names in `pred path`.** When `pred from` returns names like `SpinGlass/SimpleGraph/f64`, use that exact string in subsequent `pred path` calls. Bare names (e.g., `SpinGlass`) resolve to the default variant, which may differ from the reachable variant and cause false "no path" errors. -- **Recommend, don't just list.** When presenting options (models in Step 2, paths in Step 3, solvers in Step 4), always bold or mark your recommended choice with a brief reason. The user can still pick freely. -- **Compact formatting.** Write explanations as plain paragraphs. Do not use blockquote `>` syntax for explanations. Keep tight: command announcement, code block output, 1-3 sentence explanation. -- **Conversational tone.** Guided consultation, not a lecture. -- **Live execution.** Every `pred` command runs for real. No fake output. -- **Graceful fallbacks.** If a path doesn't exist or a command fails, explain what happened and suggest alternatives (try another model, use brute-force, backtrack). -- **Adapt to user level.** If the user gives a formal problem name, skip clarification. If they describe a fuzzy real-world problem, ask follow-ups one at a time. -- **Use `--timeout 30`** with `pred solve` in any live demos during the session. -- **Doc template sections are conditional.** "Finding the Optimum" only applies to decision models. "External Solver Alternatives" only applies when external solvers were chosen. "Solution Extraction" can be folded into "Solving" when the bundle workflow handles it automatically. +After writing, summarize the doc, offer a live end-to-end demo when a built-in solver applies +(always `--timeout 30`), and ask for changes. diff --git a/.claude/skills/fix-issue/SKILL.md b/.claude/skills/fix-issue/SKILL.md deleted file mode 100644 index fe1d1d777..000000000 --- a/.claude/skills/fix-issue/SKILL.md +++ /dev/null @@ -1,376 +0,0 @@ ---- -name: fix-issue -description: Fix quality issues found by check-issue — auto-fixes mechanical problems, brainstorms substantive issues with human, then re-checks and moves to Ready ---- - -# Fix Issue - -Fix errors and warnings from a `check-issue` report. Auto-fixes mechanical issues, brainstorms substantive ones with the human, edits the issue body, re-checks once, then asks the human to approve or iterate. - -## Invocation - -``` -/fix-issue -``` - -- `/fix-issue model` or `/fix-issue rule` — pick next from Backlog -- `/fix-issue 207` — fix a specific issue by number (skip Step 1a/1b, go directly to 1c) - -## Constants - -GitHub Project board IDs: - -| Constant | Value | -|----------|-------| -| `PROJECT_ID` | `PVT_kwDOBrtarc4BRNVy` | -| `STATUS_FIELD_ID` | `PVTSSF_lADOBrtarc4BRNVyzg_GmQc` | -| `STATUS_BACKLOG` | `ab337660` | -| `STATUS_ON_HOLD` | `48dfe446` | -| `STATUS_READY` | `f37d0d80` | - -## Process - -```dot -digraph fix_issue { - rankdir=TB; - "Pick issue from Backlog" [shape=box]; - "Move card to OnHold (claim lock)" [shape=box]; - "Fetch issue + check comment" [shape=box]; - "Parse failures & warnings" [shape=box]; - "Auto-fix mechanical issues" [shape=box]; - "Present auto-fixes to human" [shape=box]; - "Brainstorm substantive issues" [shape=box]; - "Re-check locally" [shape=box]; - "Ask human" [shape=diamond]; - "Edit GitHub + update labels + move to Ready" [shape=box, style=filled, fillcolor="#ccffcc"]; - "Ask what to change (free-form)" [shape=box]; - "Apply changes + re-check locally" [shape=box]; - - "Pick issue from Backlog" -> "Move card to OnHold (claim lock)"; - "Move card to OnHold (claim lock)" -> "Fetch issue + check comment"; - "Fetch issue + check comment" -> "Parse failures & warnings"; - "Parse failures & warnings" -> "Auto-fix mechanical issues"; - "Auto-fix mechanical issues" -> "Present auto-fixes to human"; - "Present auto-fixes to human" -> "Brainstorm substantive issues"; - "Brainstorm substantive issues" -> "Re-check locally"; - "Re-check locally" -> "Ask human"; - "Ask human" -> "Edit GitHub + update labels + move to Ready" [label="1: looks good"]; - "Ask human" -> "Ask what to change (free-form)" [label="2: modify again"]; - "Ask what to change (free-form)" -> "Apply changes + re-check locally"; - "Apply changes + re-check locally" -> "Ask human"; -} -``` - ---- - -## Step 1: Pick and Claim the Issue - -The argument is `model`, `rule`, or a specific issue number. - -- If a **number** is given, skip to Step 1c with that issue. -- If `model` or `rule`, pick from the Backlog as below. - -### 1a: Fetch candidate list from project board - -```bash -uv run --project scripts scripts/pipeline_board.py backlog --format json -``` - -Returns all Backlog issues of the requested type, sorted by `Good` label first then by issue number: - -```json -{ - "issue_type": "rule", - "items": [ - {"number": 246, "title": "[Rule] A → B", "item_id": "PVTI_xxx", "has_good": true, "labels": ["Good", "rule"]}, - {"number": 91, "title": "[Rule] C to D", "item_id": "PVTI_yyy", "has_good": false, "labels": ["rule"]} - ] -} -``` - -### 1b: Pick the top issue - -Pick the first item from the list and retain both `` and ``. If the list is empty, STOP with message: "No `[Model]`/`[Rule]` issues in Backlog." - -If the top issue already has the `Good` label and its check report has **0 failures and 0 warnings**, skip to Step 8 (just move it to Ready — no edits needed). If it has warnings, proceed normally. - -### 1c: For a specific issue number, locate the board item and validate status - -When `/fix-issue ` is used, do **not** start editing immediately. First look up the card: - -```bash -uv run --project scripts scripts/pipeline_board.py find -``` - -Returns JSON: `{"item_id": "PVTI_xxx", "status": "Backlog", "number": 207, "title": "..."}` or an error if not found. - -- If the status is `Backlog`, continue and claim it in Step 1d. -- If the status is `OnHold`, continue only as a **resume** of in-progress `fix-issue` work. -- If the status is anything else (`Ready`, `In progress`, `Review pool`, etc.), STOP with message: "Issue # is not available for `/fix-issue` because its board status is ." -- If an error is returned, STOP with message: "Issue # is not on the project board." - -### 1d: Move the card to OnHold immediately - -Claim the work item **before any further action** (before `gh issue view`, before parsing comments, before drafting edits): - -```bash -uv run --project scripts scripts/pipeline_board.py move OnHold -``` - -This is a temporary lock to avoid two agents or humans editing the same issue concurrently. Do not leave the card in Backlog while you inspect or modify it. - -Use `OnHold` as the in-progress state for `/fix-issue`. Once claimed, do not move the issue back to `Backlog` automatically; keep it in `OnHold` until Step 8 or explicit human re-triage. - -If the issue is already in `OnHold` because you are resuming previously started `fix-issue` work, do not move it again; just continue. - -### 1e: Fetch the chosen issue and related context - -```bash -gh issue view --json title,body,labels,comments -``` - -- Find the **most recent** comment that starts with `## Issue Quality Check` — this is the check-issue report -- If no check comment found, run `/check-issue ` first, then re-fetch the issue - -**Cross-reference lookup (required):** - -- **For `[Model]` issues:** Fetch **all** associated `[Rule]` issues that reference this problem as source or target. Search broadly — the problem name may appear in different forms (CamelCase, spaces, abbreviations): - ```bash - gh issue list --search " in:title label:rule" --state open --json number,title,body --limit 50 - ``` - Read the body of each matching rule issue to understand how the model is used (what fields the rule constructs, whether it relies on decision vs optimization framing, what overhead expressions reference). This context is essential for making informed decisions about schema, framing, and naming. - -- **For `[Rule]` issues:** Check whether the source and target models exist in the codebase or as issues: - ```bash - # Check if models exist in pred - pred show 2>&1 - pred show 2>&1 - # Check for model issues if not in pred - gh issue list --search " in:title label:model" --state open --json number,title,body --limit 10 - gh issue list --search " in:title label:model" --state open --json number,title,body --limit 10 - ``` - Read relevant model issue bodies to understand schema fields, variants, and framing. This informs whether the rule's overhead expressions, field names, and algorithm are consistent with the models. - ---- - -## Step 2: Parse Failures and Warnings - -Extract from the check comment's summary table: - -| Field | How to extract | -|-------|---------------| -| Check name | First column (Usefulness, Non-trivial, Correctness, Well-written) | -| Result | Second column (Pass / Fail / Warn) | -| Details | Third column (one-line summary) | - -Then parse the detailed sections below the table for specifics: -- Each `### ` section contains the full explanation -- The `#### Recommendations` section (if present) contains suggestions - -Build a structured list of **all issues to fix** — include both `Fail` **and** `Warn` results. Warnings are not ignorable; they must be resolved before moving to Ready. - -Tag each issue as: -- `mechanical` — can be auto-fixed without human input -- `substantive` — requires human brainstorming - -### Classification Rules - -**Mechanical** (auto-fixable): - -| Issue pattern | Fix strategy | -|--------------|-------------| -| Undefined symbol in overhead/algorithm | Add definition derived from context (e.g., "let n = \|V\|") | -| Inconsistent notation across sections | Standardize to the most common usage in the issue | -| Missing/wrong code metric names | Look up correct names via `pred show --json` → `size_fields`. If `size_fields` is empty, use the plain-text `pred show ` output which lists Fields directly | -| Formatting issues (broken tables, missing headers) | Reformat to match issue template | -| Incomplete `(TBD)` in fields derivable from other sections | Fill from context | -| Incorrect DOI format | Reformat to `https://doi.org/...` | - -**Substantive** (brainstorm with human, ask for human's input): - -| Issue pattern | Why human input needed | -|--------------|----------------------| -| Naming decisions (optimization prefix, CamelCase choice, too long name) | Codebase convention judgment call | -| Missing or incorrect complexity bounds | Requires literature verification | -| Missing type dependencies | Architectural decision about codebase | -| Incorrect mathematical claims | Domain expertise needed | -| Incomplete reduction algorithm | Core technical content | -| Incomplete or trivial example | Present **3 concrete example options** with pros/cons (use `AskUserQuestion` with previews showing vertex/edge counts, optimal values, and suboptimal cases). Prefer examples that match the model issue's example when a companion model exists. | -| Decision vs optimization framing | **Default to objective-style models** unless evidence points otherwise. In the current aggregate-value architecture, that usually means `type Value = Max<_>`, `Min<_>`, or `Extremum<_>` when the sense is runtime data. Check associated `[Rule]` issues (`gh issue list --search " in:title label:rule"`) to see how rules use this model — if rules only need the decision version (e.g., reducing to SAT with a bound), an objective model still works because the bound can be read from the optimal aggregate value. Use `Or` for inherently existential feasibility problems (SAT, KColoring) where there is no natural objective. Use aggregate-only values such as `Sum<_>` or `And` only when the answer is genuinely a fold over all configurations and there is no representative witness. If switching to an objective model, add the appropriate `Minimum`/`Maximum` prefix per codebase conventions. | -| Ambiguous overhead expressions | Requires understanding the reduction | - ---- - -## Step 3: Auto-Fix Mechanical Issues - -For each `mechanical` issue: - -1. Identify the exact section in the issue body that needs editing -2. Apply the fix -3. Record what was changed (for presenting to human in Step 4) - -Use `cargo run -p problemreductions-cli --bin pred -- show ` (or `./target/debug/pred show ` after `make cli`) to look up: -- Valid problem names and aliases -- `size_fields` for correct metric names (via `--json`); if empty, use plain-text output which lists Fields directly -- Existing variants and fields - -**Do NOT edit the issue on GitHub yet** — collect all fixes (mechanical + substantive) first. - ---- - -## Step 4: Present Full Context and Auto-Fixes to Human - -**IMPORTANT: Show all context BEFORE asking for any decisions.** - -Present everything in one block so the human has full visibility: - -1. **Full problem definition** — quote the issue's Definition and Schema sections verbatim so the human can see exactly what is being discussed without switching to the browser. For rule issues, include the Reduction Algorithm and Size Overhead sections instead. -2. **Check report summary** — the parsed table from Step 2 (Check / Result / Details). For `Fail` or `Warn` results, include the specific sub-issues from the detailed section, not just the one-liner. -3. **Related issues** — for model issues, list all associated rule issues found in Step 1e with a one-line summary of how each rule uses this model (fields referenced, framing assumed). For rule issues, state whether the source and target models exist in `pred` or as open issues, and note any schema/framing mismatches. -4. **Research results** — any web searches, `pred show` lookups, or other research gathered during Steps 2–3. -5. **Auto-fixes applied** — table of mechanical fixes (Section / Issue / Fix). -6. **Substantive issues requiring input** — numbered list of issues that need human judgment, with enough context for the human to evaluate each one. - ---- - -## Step 5: Brainstorm Substantive Issues - -For each substantive issue, present it to the human **one at a time**: - -1. **Show the evidence first** — quote the relevant check report section, web research results, `pred show` output, or related rule/model issue findings that inform this decision. The human should be able to evaluate the options based on the evidence shown, not just the option labels. If the issue's Definition or Schema is relevant to the decision, quote the relevant sections inline. -2. State the problem clearly -3. Offer 2-3 concrete options when possible (with your recommendation) -4. Wait for the human's response -5. Apply the chosen fix to the draft issue body - -**Example instances require special handling:** If the example is flagged as incorrect, incomplete, or trivial, you MUST present **3 concrete example options** using `AskUserQuestion` with previews. Each preview should show the full graph/instance specification, optimal value, suboptimal cases, and an invalid configuration. Do NOT silently reuse or fix the existing example without offering alternatives — the human must choose. - -Use web search if needed to help resolve issues: -- Literature search for correct complexity bounds -- Verify algorithm claims -- Find better references - -After all substantive issues are resolved, show the human the complete updated issue body (or a diff summary if the body is long). - ---- - -## Step 6: Re-Check Locally - -Re-run the 4 quality checks (Usefulness, Non-trivial, Correctness, Well-written) from `check-issue` against the **draft issue body** (not yet pushed to GitHub). Use `pred show`, `pred path`, web search as needed. Do NOT post a GitHub comment. - -Print results to the human as a summary table (Check / Result / Details). - -If the issue cannot be completed in this session because the check report is missing, the issue body is malformed, required context is unavailable, or the proposed fix turns out to be wrong, STOP and tell the human exactly what blocked completion. Leave the card in `OnHold`. - ---- - -## Step 7: Ask Human for Decision - -Show the human the draft issue body. - -Use `AskUserQuestion` to present the options: - -> The issue has been re-checked locally. What would you like to do? -> -> 1. **Looks good** — I'll push the edits to GitHub, update labels, and move it from OnHold to Ready -> 2. **Modify again** — tell me what else you'd like to change - -### If human picks 2: Modify Again - -Ask the human (free-form) what they want to change: - -> What would you like to modify? Describe the changes you want. - -Apply the requested changes to the draft issue body, re-check locally (Step 6), then ask again (Step 7). Repeat until the human picks "Looks good". - -If the session pauses without approval, leave the card in `OnHold`. Do **not** move it back to Backlog automatically; `OnHold` is the conflict-avoidance state for partially completed `fix-issue` work. - ---- - -## Step 8: Finalize (If human picks 1 "Looks good") - -Only reached when the human approves. Now push everything to GitHub. - -### 8a: Edit the issue body and title - -Use the Write tool to save the updated body to `/tmp/fix_issue_body.md`, then: - -```bash -gh issue edit --body-file /tmp/fix_issue_body.md -``` - -If the problem name was changed (e.g., renamed to add `Minimum`/`Maximum` prefix), also update the issue **title**: - -```bash -gh issue edit --title "[Model] NewProblemName" -``` - -Then find and update **all related issues** that reference the old name in their title: - -```bash -gh issue list --search "OldName in:title" --state open --json number,title -# For each related issue, update the title: -gh issue edit --title "" -``` - -### 8b: Comment on the issue with a changelog - -Post a comment summarizing what was changed, so reviewers can see the diff at a glance: - -```bash -gh issue comment --body "$(cat <<'EOF' -## Fix-issue changelog - -- -- ... - -Applied by `/fix-issue`. -EOF -)" -``` - -### 8c: Update labels - -```bash -gh issue edit --remove-label "Useless,Trivial,Wrong,PoorWritten" 2>/dev/null -gh issue edit --add-label "Good" -``` - -### 8d: Move from OnHold to Ready on project board - -Use the `item_id` obtained from Step 1b/1c: - -```bash -uv run --project scripts scripts/pipeline_board.py move Ready -``` - -### 8e: Confirm - -```text -Done! Issue #: - - Body updated on GitHub - - Labels: removed failure labels, added "Good" - - Board: moved OnHold -> Ready -``` - ---- - -## Common Mistakes - -| Mistake | Fix | -|---------|-----| -| Pushing to GitHub before human approves | All edits stay local until human picks "Looks good" | -| Hallucinating paper content for complexity bounds | Use web search; if not found, say so and ask human | -| Using `pred show` on a problem that doesn't exist yet | Check existence first; for new problems, skip metric lookup | -| Overwriting human's original content | Preserve original text; only modify the specific sections flagged | -| Not preserving `` markers | Keep existing provenance markers; add new ones for AI-filled content | -| Running check-issue more than once per iteration | Re-check exactly once after edits, then ask human | -| Leaving the card in Backlog while you inspect/edit | Move it to OnHold before `gh issue view` or any drafting, and keep blocked work there | -| Closing the issue | Never close. Labels and board status only | -| Force-pushing or modifying git | This skill only edits GitHub issues via `gh`. No git operations | -| Inventing `pipeline_board.py` subcommands | Only `next`, `claim-next`, `ack`, `list`, `move`, `backlog`, `find` exist | -| Forgetting to update the issue title | If the problem name changed, update the title with `gh issue edit --title "..."` and find all related issues referencing the old name | -| Asking questions without context | Before every `AskUserQuestion`, show the relevant issue content (definition, example, constraints) so the human can make an informed decision | -| Not showing the full problem definition in Step 4 | Always quote the Definition and Schema (or Reduction Algorithm and Size Overhead for rules) verbatim in Step 4 so the human has the full picture without switching to the browser | -| Skipping cross-reference lookup for models | For model issues, fetch and read **all** associated rule issues to understand how the model is used before making framing/schema decisions | -| Skipping cross-reference lookup for rules | For rule issues, check whether source and target models exist in `pred` or as open issues, and read their schema/framing to ensure consistency | diff --git a/.claude/skills/fix-pr/SKILL.md b/.claude/skills/fix-pr/SKILL.md deleted file mode 100644 index 755cef9c9..000000000 --- a/.claude/skills/fix-pr/SKILL.md +++ /dev/null @@ -1,150 +0,0 @@ ---- -name: fix-pr -description: Use when a PR has review comments to address, CI failures to fix, or codecov coverage gaps to resolve ---- - -# Fix PR - -Resolve PR review comments, fix CI failures, and address codecov coverage gaps for the current branch's PR. - -## Step 1: Gather PR State - -Step 1 should be a single report-generation step. Use the shared scripted helper to produce one skill-readable PR context packet. Do not rebuild this logic inline with `gh api | python3 -c` unless you are debugging the helper itself. - -```bash -REPORT=$(python3 scripts/pipeline_pr.py context --current --format text) -printf '%s\n' "$REPORT" -``` - -The report should already include: -- repo, PR number, title, URL, head SHA -- comment counts -- CI summary -- Codecov summary -- linked issue context - -Use the values printed in that report for the rest of this skill. If you absolutely need raw structured data for a corner case, rerun the same command with `--format json`, but do not rebuild Step 1 manually. - -### 1a. Fetch Review Comments - -**Check ALL four sources.** User inline comments are the most commonly missed — do not skip any. - -Start from the report's `Comment Summary`. It should tell you whether any source is non-empty before you inspect raw threads. - -If you need the raw comment arrays for detailed triage, rerun `python3 scripts/pipeline_pr.py context --current --format json` and inspect: -- `comments["inline_comments"]` -- `comments["reviews"]` -- `comments["human_issue_comments"]` -- `comments["human_linked_issue_comments"]` -- `comments["codecov_comments"]` - -### 1b. Check CI Status - -Read the report's `CI Summary`. The structured JSON fallback includes: -- `state` — `pending`, `failure`, or `success` -- `runs` — normalized check-run details -- `pending` / `failing` / `succeeding` counts - -### 1c. Check Codecov Report - -Read the report's `Codecov` section. The structured JSON fallback includes: -- `found` — whether a Codecov comment is present -- `patch_coverage` -- `project_coverage` -- `filepaths` — deduplicated paths referenced by Codecov links -- `body` — the raw latest Codecov comment body - -## Step 2: Triage and Prioritize - -Categorize all findings: - -| Priority | Type | Action | -|----------|------|--------| -| 1 | CI failures (clippy/test/coverage) | Fix immediately — blocks merge | -| 2 | Inline/review comments | Address each one — evaluate validity, fix if correct | -| 3 | Codecov coverage gaps | Add tests for uncovered lines | - -## Step 3: Fix CI Failures - -For each failing check: - -1. **Clippy**: Run `make clippy` locally, fix warnings -2. **Test**: Run `make test` locally, fix failures (build errors surface here too) -3. **Code Coverage**: See Step 5 (codecov-specific flow) - -## Step 4: Address Review Comments - -For each review comment: - -1. Read the comment and the code it references -2. Evaluate if the suggestion is correct -3. If valid: make the fix, commit -4. If debatable: fix it anyway unless technically wrong -5. If wrong: prepare a response explaining why - -**Do NOT respond on the PR** -- just fix and commit. The user will push and respond. - -### Handling Suggestions - -Review suggestions with `suggestion` blocks contain exact code. Evaluate each: -- **Correct**: Apply the suggestion -- **Partially correct**: Apply the spirit, adjust details -- **Wrong**: Skip, note why in commit message - -## Step 5: Fix Codecov Coverage Gaps - -**IMPORTANT: Do NOT run `cargo-llvm-cov` locally.** Use the `gh api` to read the codecov report instead. - -### 5a. Identify Uncovered Lines - -From the `CODECOV` JSON (fetched in Step 1c), extract: -- Files with missing coverage -- Patch coverage percentage -- Specific uncovered files referenced in `filepaths` - -Then read the source files and identify which new/changed lines lack test coverage. - -### 5b. Add Tests for Uncovered Lines - -1. Read the uncovered file and identify the untested code paths -2. Write tests targeting those specific paths (error branches, edge cases, etc.) -3. Run `make test` to verify tests pass -4. Commit the new tests - -### 5c. Verify Coverage Improvement - -After pushing, CI will re-run coverage. Check the updated codecov comment on the PR. - -## Step 6: Commit and Report - -After all fixes: - -```bash -# Verify everything passes locally -make check # fmt + clippy + test -``` - -Commit with a descriptive message referencing the PR: - -```bash -git commit -m "fix: address PR #$PR review comments - -- [summary of fixes applied] -" -``` - -Report to user: -- List of review comments addressed (with what was done) -- CI fixes applied -- Coverage gaps filled -- Any comments left unresolved (with reasoning) - -## Integration - -### With review-implementation - -Run `/review-implementation` first to catch issues before push. Then `/fix-pr` after push to address CI and reviewer feedback. - -### With executing-plans / finishing-a-development-branch - -After creating a PR, use `/fix-pr` to address review feedback and CI failures. diff --git a/.claude/skills/how-to-code/SKILL.md b/.claude/skills/how-to-code/SKILL.md new file mode 100644 index 000000000..93c53afe6 --- /dev/null +++ b/.claude/skills/how-to-code/SKILL.md @@ -0,0 +1,134 @@ +--- +name: how-to-code +description: Use when implementing or modifying a problem model (src/models/) or a reduction rule (src/rules/) in this repo, including Decision

variants, aggregate reductions, and direct -> ILP rules. +--- + +# How to Code a Model or Rule + +`.claude/CLAUDE.md` is the architecture reference (traits, macros, registry, numeric contract). +This guide adds the checklists and the traps that are easy to miss. + +## Principles + +- Copy the closest existing model/rule, not a template from memory. APIs here change often. +- Implement exactly the issue's mathematics. Missing definitions, domains, or constraints are an + issue problem (how-to-triage-issue), not something to invent. +- Rules: verify the math first with how-to-verify, then translate its verified `reduce()` / + `extract_solution()` into Rust. Skip only for trivially mechanical rules (identity, complement, + variant cast). +- One item per PR. A rule needs both models already on `main`. Exception: a `[Model]` issue that + explicitly claims direct ILP solvability ships its ` -> ILP` rule in the same PR, held to + the full production bar. + +## Required inputs (from the issue) + +Model: name (with `Maximum`/`Minimum` prefix only when there is a direction), formal definition +with input domains, objective or feasibility condition, best-known exact algorithm with citation +(concrete numbers only, e.g. `1.1996^num_vertices`), a small example with its expected outcome +(optimal solution + value, or a satisfying witness + why). The issue's expected outcome is the +source of truth for tests, example-db, and paper. + +Rule: source/target variants, construction, extraction, correctness argument, parameter transform, +worked example, reference. + +A model with no existing or planned rule becomes an orphan in the reduction graph — say so in the +PR and link or file a companion rule issue. + +## Model checklist + +Reference: `src/models/graph/maximum_independent_set.rs` + `src/unit_tests/models/graph/maximum_independent_set.rs`. +Decision variant reference: `src/models/graph/minimum_vertex_cover.rs` (`DecisionProblemMeta` impl or +`decision_problem_meta!`, inherent getters on `Decision

`, `register_decision_variant!`, +`decision_canonical_model_example_specs()`). Never hand-write a decision model. + +1. `src/models//.rs`: + - `inventory::submit! { ProblemSchemaEntry { .. } }` with `display_name`, well-established + `aliases` only (never invent), `dimensions`, explicit `category`, `fields`. + - `impl Problem`: `NAME`, `type Solution` (mathematical witness, e.g. `Vec`), + `type Value` (`Max`/`Min`/`Extremum` for objectives, `Or` for feasibility, value-only + aggregates like `Sum`/`And` when no witness exists), `crate::problem_parameters![(..)]` + (names used by complexity strings and rule transforms), `variant()` via `crate::variant_params!`, + `evaluate() -> Result` (infeasible = `Max(None)`/`Or(false)`; malformed + input and overflow = `Err`). + - Fallible `try_new` shared by `new`, serde `Deserialize`, and `TryFrom`. + - `crate::declare_variants!` — one `default`, `create ` when construction differs from + persisted JSON (model-local `#[derive(Deserialize, crate::CreateSpec)]` DTO, `FIELDS` used in + the schema entry), `random` only where a meaningful generator exists (`impl_random_generate!`). + - Brute force (when a finite enumeration exists): `impl BruteForceProblem { fn dimensions() }` + plus `crate::register_brute_force! { Model<..> decode |_, indices| .. }` per variant. + - `#[cfg(feature = "example-db")] canonical_model_example_specs()` in the model file. + - `#[cfg(test)] #[path = "../../unit_tests/models//.rs"] mod tests;` +2. `src/models//mod.rs`: `pub(crate) mod`, `pub use`, and + `specs.extend(::canonical_model_example_specs())`. Re-export in `src/models/mod.rs`. +3. Tests: ≥3 functions (creation, evaluate valid/invalid/malformed, brute-force solve, serde + round-trip, `test__paper_example` reproducing the issue/paper example exactly). +4. Paper `problem-def` + `display-name` entry — how-to-write-manual. + +Never add model-name branches, aliases, or parsers in CLI (`problemreductions-cli/`) or MCP code; +construction, aliases, random, and solving are all discovered from the registry. + +## Rule checklist + +Reference: `src/rules/minimumvertexcover_maximumindependentset.rs` + `src/unit_tests/rules/minimumvertexcover_maximumindependentset.rs`. +Traits and helpers: `src/rules/traits.rs`, `src/rules/test_helpers.rs`, `src/rules/ilp_helpers.rs`. + +1. `src/rules/_.rs` (lowercase, no underscores inside a name): + - Result struct holding the target + any index maps. `extract_solution(&self, &Target::Solution) + -> ExtractionResult` starts with exactly one + `crate::rules::traits::validate_target_solution(self.target_problem(), target_solution)?` + (composed extractors delegate to the first direct decoder). + - `#[reduction(transform = exact { .. })]`, `upper_bound { .. }`, or `unavailable { .. }`, with an + auxiliary `unavailable = { param = "reason" }` block for unrepresentable target parameters. + Every target parameter appears exactly once. There is no `overhead =` form. + - `reduce_to(&self) -> Result`; wrap target construction failures + with `Self::target_construction(e)`, never stringify. + - Aggregate value mapping: `#[crate::aggregate_reduction] impl AggregateReductionResult` on the + same result type (`extract_value`); generic results need `crate::register_aggregate_reduction!` + per concrete instance (see `src/rules/kcoloring_casts.rs`). + - `#[cfg(feature = "example-db")] canonical_rule_example_specs()` in the rule file. +2. `src/rules/mod.rs`: `pub(crate) mod` + `specs.extend(::canonical_rule_example_specs())`. + ILP rules are not feature-gated (the `ilp-solver`/`ilp-highs` features no longer exist). +3. Exactly one primitive registration per exact (source variant, target variant) pair; share helpers, + don't duplicate endpoints. +4. Tests in `src/unit_tests/rules/_.rs`: `test__to__closed_loop` + (use `assert_*_round_trip_*` from `test_helpers.rs`), an infeasible instance, target structure + and parameter counts vs the transform, and one test per malformed representation the decoder + rejects (zero/multiple one-hot bits, duplicate permutation entries, ...). Aggregate edges: test + `extract_value` against `BruteForce::solve` on both sides. +5. Paper `reduction-rule` entry — how-to-write-manual (adapt the how-to-verify proof, don't rewrite). + +## Traps + +- Type gate before coding: resolve concrete `Value` types (see how-to-verify). `Min` and + `Min` are different; `Max`->`Min`, `Min`->`Or`, and anything targeting `Sum`/`And` cannot + be a witness `ReduceTo`. +- Numeric contract (`docs/src/design.md#numeric-types-and-arithmetic`): `usize` only for indices, + lengths, brute-force dimensions; `u64` parameters; `i64` signed values; finite `f64`. `TryFrom` + at every range/sign boundary (`ReduceTo::exact_i64` for counts), checked arithmetic for derived + totals, `i64_to_exact_f64` for i64->f64. No `as` casts that change range or sign. +- Error phases stay typed: `ConstructionError` / `EvaluationError` / `ReductionError` / + `ExtractionError`. No `Result<_, String>`, no panics on user-reachable paths. +- Extraction decodes only the defined mapping: never truncate, clamp, default, or "repair". +- Solver API: `BruteForce::new().solve(&p) -> Result, SolveError>` and + `find_all_witnesses`. There is no `Solver` trait, `find_witness`, or `dims()`. +- Complexity strings: best-known algorithm with citation; polynomial problems must not get + exponential bounds; variable names must be declared parameters. +- Direct ILP rules are production rules: exact transform where possible, closed-loop plus + infeasible/weighted/pathological tests, `assert_bf_vs_ilp` where cheap. + +## Gates + +- `make check` (fmt-check + clippy `-D warnings` + full test suite) and `make paper` pass. +- New code coverage >95% (`make coverage` locally; codecov on the PR per how-to-ship). +- Every test <5 s — shrink instances rather than slow the suite. +- Never commit generated JSON: `*.json` is gitignored (`docs/src/reductions/reduction_graph.json`, + `problem_schemas.json`, `docs/paper/data/`); don't `git add -f` them. + +## Helper commands + +```bash +pred show [--json] # schema, variants, reductions +pred path --limit all # existing paths and parameter transforms +pred create --example | pred solve - # smoke-test the canonical example +cargo run --example export_graph # refresh reduction_graph.json (also run by make paper) +``` diff --git a/.claude/skills/how-to-review/SKILL.md b/.claude/skills/how-to-review/SKILL.md new file mode 100644 index 000000000..f0756cc30 --- /dev/null +++ b/.claude/skills/how-to-review/SKILL.md @@ -0,0 +1,127 @@ +--- +name: how-to-review +description: Use when reviewing a pull request or a finished, uncommitted diff in this repo — as a fresh-context reviewer dispatched after implementation, or when a human asks to "review PR N". +--- + +# How to Review + +You are a read-only reviewer: evaluate and report, never edit, commit, push, or merge. Run from a +checkout of the PR head (`gh pr checkout N`); the scope helper diffs `HEAD` against +`merge-base origin/main`. Architecture and conventions are in `.claude/CLAUDE.md`; judge against +them, and against the current code rather than memory. + +## Gather scope once + +```bash +REPO=$(gh repo view --json nameWithOwner -q .nameWithOwner) +python3 scripts/pipeline_skill_context.py review-implementation --repo-root . --format text +python3 scripts/pipeline_pr.py context --repo "$REPO" --pr "$PR" --format text # comments, CI, linked issue +gh pr diff "$PR" +``` + +The review-implementation packet gives review type (model/rule/generic), subject, whitelist and +completeness checks, and changed files. Treat a `fail` there as a lead, not a verdict: confirm it +against the actual PR file list. Pass `--kind/--name/--source/--target` when auto-detection is wrong. +Read every changed file in full. + +## (a) Structural and semantic + +**Hard fails, whatever the PR type** +- Generated exports in the diff: `docs/src/reductions/reduction_graph.json`, + `docs/src/reductions/problem_schemas.json`, `docs/paper/data/examples.json`. They are gitignored + build outputs of `make doc` / `make paper`. +- Unrelated edits, deleted models/rules/tests, or core-trait/macro changes the issue does not need. + +**Model** (`src/models//.rs`) +- `Problem` impl with `type Value`; `declare_variants!` with one `default` and a complexity string + whose base matches the cited best-known algorithm (polynomial problems must not be exponential). +- `ProblemSchemaEntry` with an explicit category; aliases only in `ProblemSchemaEntry.aliases` / + `declare_variants!`, and only standard literature abbreviations that do not collide. +- Brute force via `register_brute_force!`; `BruteForceProblem::dimensions()` is the real + configuration space. +- Infeasible configurations evaluate to `Max(None)` / `Min(None)` / `Extremum(None)`, or `Or(false)` + for feasibility problems. +- `pred create --help` flags come from the schema fields or a model-local `CreateSpec` — + no model-name branches in `problemreductions-cli/` or MCP code, and no leftover manual dispatch + arms in `dispatch.rs` / `problem_name.rs` after merging main. +- `#[cfg(feature = "example-db")] canonical_model_example_specs()` in the model file, chained from + the category `mod.rs` (`src/example_db/model_builders.rs` only aggregates categories — grepping it + for the model name proves nothing). +- Test file `src/unit_tests/models//.rs` linked via `#[path]`, with at least 3 tests. + +**Rule** (`src/rules/_.rs`) +- `#[reduction(transform = exact|upper_bound|unavailable { ... })]` covering every target parameter + exactly once (there is no `overhead =` form any more); registered in `src/rules/mod.rs`. +- `canonical_rule_example_specs()` in the rule file, collected in `src/rules/mod.rs`. +- Every direct `extract_solution()` calls `validate_target_solution()` once, then decodes only the + mathematical mapping; malformed input yields `ExtractionError` — no panic, clamp, truncation, + invented default, or recovery branch — and each rejection path is tested. +- Closed-loop test `test__to__closed_loop` in `src/unit_tests/rules/`. + +**Both** +- Error contracts: construction returns `ConstructionError`, `evaluate()` returns `EvaluationError`, + reductions return `ReductionError` (target construction failures kept as + `ReductionError::Construction`); no public `Result<_, String>` in new code. +- Numeric contract per `docs/src/design.md#numeric-types-and-arithmetic`: `u64` parameters, `TryFrom` + at range boundaries, checked arithmetic for derived sizes, same range in Rust/serde/CLI/MCP. +- Paper: `display-name` + `problem-def`, or `reduction-rule` per directed edge. Example values must be + loaded (`load-model-example` / `load-example`) from `examples.json`, never hand-written; see + how-to-write-manual. + +**Semantic checks — do these by hand, do not trust green CI** +- Rule: pick a small source instance, trace `reduce_to()` line by line, and confirm the target + encodes the same question and `extract_solution` inverts it. Count every target size by hand and + compare with the `transform` block; `pred path instance.json` measures the constructed + sizes on a real instance. Check the paper proof is sound, not just present. +- Model: `evaluate()` implements exactly the stated definition (watch "at least" vs "all"), including + empty/zero/infeasible edge cases. + +**Issue compliance** (linked issue from the context packet) +- Issue comments override the body; read all of them first. +- Definition, framing (objective vs feasibility), configuration space, complexity, reduction + algorithm, extraction, and transform formulas match the issue. +- Round trip: the issue's example instance appears in a test and in the canonical example, with the + issue's stated optimum confirmed by brute force rather than a hardcoded assertion. + +## (b) Quality — project-specific only + +- Trivial instances (single edge, 2 vertices) hide bugs; expect at least one instance with 5+ vertices + (or equivalent size). +- Closed-loop tests must check the extracted source solution is optimal against brute force on the + source, with the target also solved by brute force (see `src/rules/test_helpers.rs` + `assert_*_round_trip_*`). +- Tests that recompute the implementation's formula prove nothing; tests that only check `is_some()` + or shapes are too weak; missing infeasible/boundary cases is Important. +- Duplicated logic that an existing helper already covers. +- HCI only if `problemreductions-cli/` changed: actionable error messages, `--help` examples, + consistency with sibling commands, no silent data loss. + +## (c) Feature test as a user + +Install the PR's CLI with `make cli` (`cargo install` of `problemreductions-cli`), then exercise the new item from a scratch directory, as a user would: + +```bash +pred list # or: pred list --rules +pred show +pred create --example -o inst.json # rule: --example --to +pred inspect inst.json +pred solve inst.json # also --solver brute-force +pred path -o paths.json && jq '.paths[0]' paths.json > route.json +pred reduce inst.json --via route.json -o bundle.json && pred solve bundle.json +``` + +Check outputs against the issue's example. Reproduce every problem before reporting it. Do not fix. + +## Report + +Classify findings as **Critical** (wrong results, broken contract, blacklisted file), **Important** +(missing tests/components, weak verification, issue deviation), **Minor** (style, docs). Give +`file:line` and a concrete suggested fix for each. Post one comment: + +```bash +python3 scripts/pipeline_pr.py comment --repo "$REPO" --pr "$PR" --body-file review.md +``` + +The body starts with `## Agentic Review Report`, then sections for structural, quality, and feature +test, then the severity-sorted findings. For an uncommitted diff with no PR, return the same report +to the caller instead of posting. diff --git a/.claude/skills/how-to-ship/SKILL.md b/.claude/skills/how-to-ship/SKILL.md new file mode 100644 index 000000000..86e3a35a8 --- /dev/null +++ b/.claude/skills/how-to-ship/SKILL.md @@ -0,0 +1,97 @@ +--- +name: how-to-ship +description: Use when turning a GitHub issue into a pull request, opening or updating a PR, fixing CI failures, review comments or codecov gaps on a PR, resolving merge conflicts with main, or preparing a PR for a human to merge. +--- + +# How to Ship + +Carry one item from issue to a merge-ready PR. Implementation itself follows how-to-code, +how-to-verify, and how-to-write-manual; this guide covers the gates and GitHub mechanics around it. +A human merges — never run `gh pr merge`. Never force push (`--force`, `-f`, `--force-with-lease`). + +```bash +REPO=$(gh repo view --json nameWithOwner -q .nameWithOwner) +``` + +## Issue → PR + +1. Preflight once and reuse the JSON: + `python3 scripts/pipeline_checks.py issue-context --repo "$REPO" --issue N --format json`. + It returns title/body/labels/comments, `kind`, `source_problem`/`target_problem`, + `checks.{good_label,source_model,target_model}`, `existing_prs`, `resume_pr`, `action`. +2. Refuse to start when: + - `checks.good_label` fails — the issue has not passed triage; use how-to-triage-issue. + - a `[Rule]`'s source or target model is missing on main — post + `gh issue comment N --body "Blocked: model does not exist on main yet. Implement it first (or file a [Model] issue)."` + and stop. Never implement the missing model inside the rule PR. +3. Read every issue comment: maintainer and contributor comments override the body. +4. Branch: if `action == "resume-pr"` and `resume_pr.head_ref_name` contains `issue-N`, check out + that branch and continue the existing PR. Otherwise branch from fresh `origin/main` with a name + containing `issue-N` (e.g. `issue-N-`). +5. Plan in an uncommitted working note (scratch dir); never commit plan files or add `docs/plans/`. +6. Implement. Scope is one item per PR. Only exception: a `[Model]` issue that explicitly claims + direct ILP solvability ships the model and its ` -> ILP` rule together, with the rule held + to the full production bar (exact transform, closed-loop/infeasible tests, example, paper entry). +7. Gate locally: `make check && make paper`. `make paper` regenerates ignored exports — check + `git status --short` and never stage `docs/src/reductions/*.json` or `docs/paper/data/`. +8. Push and open the PR titled `Fix #N: ` with `Fixes #N` in the body (the body ends + with the attribution line from the session instructions): + `python3 scripts/pipeline_pr.py create --repo "$REPO" --title "Fix #N: ..." --body-file body.md --base main --head `. +9. Mandatory review: dispatch a fresh-context subagent told to follow how-to-review for this PR; it + posts `## Agentic Review Report`. Fix every Critical and Important finding, then continue below. + +## Fixing feedback + +Start from one packet: `python3 scripts/pipeline_pr.py context --repo "$REPO" --pr "$PR" --format text` +(`--format json` for raw arrays). Triage in this order: + +1. **CI failures** — `pipeline_pr.py ci`; reproduce with `make clippy` / `make test` / `make fmt-check`. +2. **Review comments** — check all five sources, listing each explicitly: + `inline_comments` (user inline comments are the most often missed), `reviews`, + `human_issue_comments` (PR conversation), `human_linked_issue_comments`, `codecov_comments`. + Evaluate each on its merits: apply correct suggestions, apply the spirit of partial ones, and + explain in the reply why a wrong one was not applied. +3. **Coverage gaps** — `python3 scripts/pipeline_pr.py codecov --repo "$REPO" --pr "$PR"` lists + patch coverage and files; read the uncovered paths and add tests for them. For PR gaps, trust + this report over rerunning `make coverage`. + +Then `make check` (plus `make paper` if the paper or examples changed), commit, push, and reply on +each addressed inline thread saying what changed and in which commit: +`gh api repos/$REPO/pulls/$PR/comments//replies -f body="..."` +(ids from `gh api repos/$REPO/pulls/$PR/comments`). Answer non-inline feedback with one PR comment +via `pipeline_pr.py comment`. + +## Merging main + +`git fetch origin && git merge origin/main` (merge, not rebase — no history rewrites). Conflicts are +almost always both sides appending to ordered lists in `mod.rs`, `lib.rs`, `create.rs`, +`dispatch.rs`, or `reductions.typ`: keep both entries in order. After a `.bib` conflict, check for +duplicate keys: + +```bash +grep '^@' docs/paper/references.bib | sed 's/@[a-z]*{//; s/,$//' | sort | uniq -d +``` + +After merging main, remove any manual dispatch arm the registry now makes redundant, and re-run +`make check && make paper` before pushing. + +## Merge-ready + +1. Local gate passes and the branch is pushed. +2. CI: `python3 scripts/pipeline_pr.py wait-ci --repo "$REPO" --pr "$PR"` (defaults: 900 s timeout, + 30 s interval). Check once or twice at most; if still pending, report "CI pending, local checks + pass" instead of babysitting. Fix and push if it fails. +3. No open Critical/Important findings and no unanswered reviewer comments. +4. Approve: `gh pr review "$PR" --approve` (fails harmlessly when you authored the PR). +5. Post the community-call checklist on the linked issue (on the PR if there is none), with real + problem names and PR number substituted: + + ```markdown + Please kindly check the following items (PR #): + - [ ] **Paper** ([PDF](https://github.com/CodingThrust/problem-reductions/blob/main/docs/paper/reductions.pdf)): check definition, proof sketch, example figure, and reproducible `pred` commands + - [ ] **Implementation (Optional)**: spot-check the source files changed in this PR for correctness + + Join the discussion on [Zulip](https://julialang.zulipchat.com/#narrow/channel/365542-problem-reductions) — feel free to ask questions or leave feedback there. + ``` + +6. Hand the PR URL to the human to merge. Record any deferred follow-ups as a PR comment. diff --git a/.claude/skills/how-to-triage-issue/SKILL.md b/.claude/skills/how-to-triage-issue/SKILL.md new file mode 100644 index 000000000..d8423ed2f --- /dev/null +++ b/.claude/skills/how-to-triage-issue/SKILL.md @@ -0,0 +1,108 @@ +--- +name: how-to-triage-issue +description: Use when a [Model] or [Rule] GitHub issue needs a quality check before implementation, or when problems found by such a check need fixing. +--- + +# How to Triage an Issue + +Gate `[Model]` / `[Rule]` issues before anyone implements them: check, report once, label, and fix +what can be fixed. Never close issues. + +Fetch with `gh issue view N --json title,body,labels,comments`. If a `## Issue Quality Check` comment +already exists, stop and say so unless asked to re-check; a re-check notes what changed since the +previous report. Build `pred` with `make cli` if missing; resolve aliases (MIS, MVC, SAT) with +`pred show --json`. + +## Checks + +Each check yields Pass / Warn / Fail. Only Fail adds a label. Unsure → Warn, never a guessed Fail. + +**1. Usefulness** (label `Useless`) +- Rule: no existing `pred path Src Tgt` → Pass. If a path exists, the new rule must not be dominated + by a composite path — apply the redundancy heuristic from how-to-verify; dominated → Fail naming + the dominating path, inconclusive → Warn. +- Model: already in `pred show` → Fail. No concrete planned reduction to/from an existing problem + (orphan) → Fail; vague ("can connect to others") → Warn. Direct ILP solvability claimed without a + linked `[Rule] to ILP` issue → Fail. No solver path at all → Warn. +- Empty or hand-wavy Motivation → Warn. + +**2. Non-triviality** (label `Trivial`) +- Rule Fails on: pure relabeling / complement substitution, subtype coercion (e.g. UnitDiskGraph → + SimpleGraph), same-problem variant identity, or a hand-waved algorithm ("map variables + accordingly"). A trivial rule that connects otherwise disconnected components of the graph still + Passes. +- Model Fails if isomorphic to an existing problem (`pred list --json`), a mere graph/weight variant + of one, or a renaming. + +**3. Correctness** (label `Wrong`) +- Check claims against `references.md` (this directory) and `docs/paper/references.bib`, then the + literature itself (arXiv / Semantic Scholar tools if present, else WebSearch + WebFetch). +- Read the actual construction and quote the theorem you rely on. Paper not found or claim not in it + → Fail; say "not found", never reconstruct what a paper "probably" says. +- Model: definition well-formed, feasibility vs objective separated, variable domain fits the + semantics; complexity bound matches the cited paper, polynomial problems not given exponential + bounds, exponential base correct (1.1996^n for MIS, not 2^n). +- Better algorithms or lower-overhead constructions found along the way are Recommendations, not + failures. + +**4. Completeness** (rules only, label `Incomplete`) — does the construction handle every source +instance? +- Literature: does the theorem say "for any instance" or carry a precondition ("connected", + "no isolated vertices", "k even")? A hidden restriction the issue ignores → Fail, quoting it. + Fixes: preprocess to the restricted form, re-target the rule at an existing restricted model, or drop. +- Codebase: read `parameters` and inputs from `pred show --json`, then hand-trace the issue's + algorithm on ≥ 2 corner cases other than the worked example (empty/single-vertex/disconnected + graphs, zero or equal weights, empty or unit clauses, empty/duplicate subsets, zero or singular + matrices — whatever the model allows). Compare with existing rules from the same source + (`grep -rl "ReduceTo.*for " src/rules/`). A break → Fail; works but issue is silent on edge + inputs → Warn. +- Show the quoted passages and the traces in the report — this is the expensive check. + +**5. Writing quality** (label `PoorWritten`) +- All template sections of `.github/ISSUE_TEMPLATE/rule.md` / `problem.md` present and substantive. +- Rule algorithm is an implementable step-by-step procedure with solution extraction; no "similarly + for the rest". +- Symbols defined before use and consistent across sections; overhead metric names match the + target's `parameters` in `pred show --json`. +- Model: precise definition with explicit quantifiers; `Maximum`/`Minimum` prefix for optimization; + complexity given as a concrete expression with citation; expected outcome given (a satisfying + solution, or an optimal solution with its value). +- Example: small enough to brute-force, fully worked, exercises the defining structure (a + "quadratic" model with only linear terms fails), and has ≥ 2 suboptimal feasible solutions besides + the optimum so a buggy round trip cannot pass by accident. +- Do not fail on implementation data types in the Schema section — contributors state the math, not + Rust types. + +**Value choice for models:** default to objective style (`Max` / `Min`, or `Extremum` when the sense +is runtime data). `Or` only for inherently existential problems (SAT, KColoring). `Sum` / `And` only +for genuine folds over all configurations with no representative witness. + +## Report and labels + +Post one comment headed `## Issue Quality Check — Rule` or `## Issue Quality Check — Model`: a +summary table (check, result, one-line detail), an overall count, a section per check with evidence, +then Recommendations. Never cite an issue you have not fetched or a path you have not verified. + +Labels: add `Useless`, `Trivial`, `Wrong`, `Incomplete`, `PoorWritten` per failed check. Add `Good` +only with zero failures and no warning on Usefulness, Correctness, or Completeness. On re-check, +remove stale failure labels. + +## Fixing + +Gather context first: for a model, every open `[Rule]` issue mentioning it +(`gh issue list --search " in:title" --state open`) to see which fields and framing rules rely +on; for a rule, whether source and target exist in `pred show` or as open `[Model]` issues. + +- **Mechanical** (undefined or inconsistent symbols, wrong metric names, broken formatting, fields + derivable from other sections, DOI format): fix directly. Write the new body to a file in the + scratchpad, `gh issue edit N --body-file `, post a `## Fix-issue changelog` comment listing + each change, re-check, update labels (remove the fixed failure labels including `Incomplete`; add + `Good` once the re-check passes). +- **Substantive** (wrong or incomplete reduction, bad or trivial example, missing/unread reference, + complexity claim, naming, value framing): stop and discuss with the human, one issue at a time. + Quote the relevant issue text and evidence first, then offer 2–3 concrete options with a + recommendation (for examples: concrete candidate instances with optimum and suboptimal solutions). +- Preserve the author's text outside flagged sections and keep `` provenance + markers; mark newly AI-filled content the same way. +- On rename, update the title (`gh issue edit N --title ...`) and propagate to related issues whose + titles or bodies use the old name. diff --git a/.claude/skills/check-issue/references.md b/.claude/skills/how-to-triage-issue/references.md similarity index 87% rename from .claude/skills/check-issue/references.md rename to .claude/skills/how-to-triage-issue/references.md index bc42d77c1..6c0a49c95 100644 --- a/.claude/skills/check-issue/references.md +++ b/.claude/skills/how-to-triage-issue/references.md @@ -1,6 +1,6 @@ # Quick Reference: Known Facts for Issue Fact-Checking -Use this file to cross-check claims in `[Rule]` and `[Model]` issues against established results. Built from the Related Projects in README.md. Each entry includes source URLs for traceability. +Cross-check `[Rule]` and `[Model]` issue claims against established results. Each entry has source URLs. --- @@ -84,21 +84,6 @@ SAT **Source:** [complexityzoo.net](https://complexityzoo.net/) — Comprehensive catalog of 550+ complexity classes (Scott Aaronson). -### Key Complexity Classes - -| Class | Description | Source | -|-------|-------------|--------| -| **P** | Deterministic polynomial time | [P](https://complexityzoo.net/Complexity_Zoo:P#p) | -| **NP** | Nondeterministic polynomial time; "yes" certificates verifiable in poly time | [NP](https://complexityzoo.net/Complexity_Zoo:N#np) | -| **co-NP** | Complements of NP problems | [co-NP](https://complexityzoo.net/Complexity_Zoo:C#conp) | -| **PSPACE** | Polynomial space (contains NP) | [PSPACE](https://complexityzoo.net/Complexity_Zoo:P#pspace) | -| **EXP** | Exponential time | [EXP](https://complexityzoo.net/Complexity_Zoo:E#exp) | -| **BPP** | Bounded-error probabilistic polynomial time | [BPP](https://complexityzoo.net/Complexity_Zoo:B#bpp) | -| **BQP** | Bounded-error quantum polynomial time | [BQP](https://complexityzoo.net/Complexity_Zoo:B#bqp) | -| **PH** | Polynomial hierarchy | [PH](https://complexityzoo.net/Complexity_Zoo:P#ph) | -| **APX** | Problems with constant-factor approximation | [APX](https://complexityzoo.net/Complexity_Zoo:A#apx) | -| **MAX SNP** | Syntactically defined optimization class | [MAX SNP](https://complexityzoo.net/Complexity_Zoo:M#maxsnp) | - ### Canonical NP-Complete Problems (from Complexity Zoo) Source: [Complexity Zoo: NP](https://complexityzoo.net/Complexity_Zoo:N#np) @@ -110,13 +95,6 @@ Source: [Complexity Zoo: NP](https://complexityzoo.net/Complexity_Zoo:N#np) - **Maximum Clique** — Do k mutually-adjacent vertices exist? - **Subset Sum** — Does a subset sum to exactly x? -### Key Class Relationships - -- P vs NP: Open problem; unequal relative to random oracles -- NP = co-NP iff PH collapses -- NP ⊆ PSPACE (Savitch's theorem) -- If NP ⊆ P/poly then PH collapses to Σ₂P - --- ## 3. Compendium of NP Optimization Problems @@ -242,16 +220,16 @@ Uses same category codes as the Compendium (GT, ND, SP, SS, MP, AN, GP, LO, AL, | GT20 | Maximum Independent Set | MaximumIndependentSet | | GT21 | Maximum Clique | MaximumClique | | GT24 | Maximum Cut | MaxCut | -| GT34 | Hamiltonian Circuit | — | -| GT39 | Feedback Vertex Set | — | +| GT34 | Hamiltonian Circuit | HamiltonianCircuit | +| GT39 | Feedback Vertex Set | MinimumFeedbackVertexSet | | GT46 | Traveling Salesman | TravelingSalesman | -| ND5 | Steiner Tree in Graphs | — | -| SP1 | 3-Dimensional Matching | — | -| SP2 | Partition | — | +| ND5 | Steiner Tree in Graphs | SteinerTree | +| SP1 | 3-Dimensional Matching | ThreeDimensionalMatching | +| SP2 | Partition | Partition | | SP5 | Set Covering | MinimumSetCovering | | SP3 | Set Packing | MaximumSetPacking | | SP13 | Bin Packing | BinPacking | -| SS1 | Multiprocessor Scheduling | — | +| SS1 | Multiprocessor Scheduling | MultiprocessorScheduling | | MP1 | Integer Programming | ILP | | LO1 | Satisfiability (SAT) | Satisfiability | | LO2 | 3-Satisfiability | KSatisfiability | @@ -303,4 +281,4 @@ Source: Chapter 3, "Proving NP-Completeness Results" (pp. 45-89). | BMF | | | | | | BicliqueCover | | | [GT: Covering](https://www.csc.kth.se/tcs/compendium/node9.html) | | | MaximalIS | | | | | -| CVP | | | | | +| ClosestVectorProblem | | | | | diff --git a/.claude/skills/how-to-verify/SKILL.md b/.claude/skills/how-to-verify/SKILL.md new file mode 100644 index 000000000..e59afb28a --- /dev/null +++ b/.claude/skills/how-to-verify/SKILL.md @@ -0,0 +1,117 @@ +--- +name: how-to-verify +description: Use when a reduction rule (or a model's mathematical claim) needs verification before Rust implementation, or when checking reduction-graph topology — orphan problems, NP-hardness reachability from 3-SAT, or whether a proposed rule is dominated by existing paths. +--- + +# How to Verify + +Two jobs: (1) prove a reduction correct before anyone writes Rust, (2) check where a rule or model +sits in the reduction graph. A FAILED verification blocks implementation — report it on the issue +or PR and stop. + +## Principles + +- The math is verified independently twice (constructor + adversary), then cross-checked. +- Nothing is committed. Proofs and scripts live in the session scratchpad directory (never `/tmp`, + never the repo). Durable evidence goes in the PR as a certificate comment. +- Everything is variant-qualified (`MinimumDominatingSet/SimpleGraph/One`), never base names only. + +## 1. Type-resolution gate (before any math) + +Resolve the concrete `Problem::Value` of both endpoints: substitute the rule's concrete generics, +follow aliases and associated types to their `impl`, record the chain with file evidence. If +anything stays unresolved, compile a probe printing +`std::any::type_name::<::Value>()` from a scratchpad crate with a path +dependency on the repo. Never infer the type from the problem's name or from Python integers. + +```text +TYPE RESOLUTION: + Source: Min, W = One, ::Sum = i64 -> Min + Target: Min + Full-domain compatibility: FAILED +``` + +- Witness-compatible: `Or->Or`, `Or->Min`, `Or->Max`, `Min->Min`, `Max->Max` (identical V). +- `Min->Min` with `S != T` is not automatically compatible: proceed only with a declared bound + covering every legal source instance and a proven total, order-preserving conversion. +- Stop on `Min/Max->Or` (needs a `Decision

` source), `Max<->Min` (aggregate or decision + wrapper), and any `Sum`/`And` target (aggregate only). +- Regression: `MinimumDominatingSet` is `Min`, `MinimumHittingSet` is + `Min`. The classical reduction is correct and small cases pass, yet the gate FAILS. + +## 2. Typst proof + +Theorem, then proof with: numbered _Construction_ (every symbol defined before use); _Correctness_ +with genuinely independent (⇒) and (⇐) paragraphs; _Solution extraction_; an overhead table +(target parameter -> formula); a YES example and a NO example (showing why no solution exists), +each with ≥3 variables/vertices and fully worked numbers. Banned: "clearly", "obviously", "it is +easy to see", "straightforward", "similarly for the converse", and any scratch work. + +## 3. Constructor script + +One Python script, 0 failures, ≥5000 checks (≥10000 for identity and algebraic reductions), +exhaustive over all instances with n ≤ 5 (n ≤ 6 for identity reductions; otherwise ≥300 samples +per (n, m) where full enumeration is infeasible). Seven sections, none empty: + +1. symbolic (sympy) check of every overhead formula — "trivial" is no excuse; +2. exhaustive forward + backward: source feasible ⇔ target feasible (optimum preserved); +3. extraction from every feasible target witness (the most skipped section); +4. measured target size vs formula; +5. structural well-formedness of the target (gadget invariants, no degenerate cases); +6. YES example reproduced number-for-number; +7. NO example reproduced, both sides infeasible. + +Print a check-count audit per section, and map every claim in the proof to the section that tests +it; add tests for uncovered claims or state why a claim is untestable. + +## 4. Adversary + +Dispatch a fresh-context subagent given ONLY the Typst proof. It writes its own `reduce`, +`extract_solution`, and source/target feasibility checkers, never importing the constructor; +exhaustive n ≤ 5; `hypothesis` with ≥2 strategies; reproduces both examples; ≥5000 checks. Point +it at the reduction's risk: identity -> all-zero/all-one/alternating configs, n ≤ 6; algebraic -> +case boundaries (e.g. S = 2T, 2T ± 1); gadget -> widget invariants and traversal patterns. + +Cross-compare both `reduce` outputs on shared instances: targets structurally identical (up to +documented isomorphism) and feasibility in agreement. + +**VERIFIED** only when both scripts pass and cross-comparison agrees. Anything else is **FAILED**, +labelled by cause (constructor bug, adversary bug or ambiguous proof, proof bug, disagreement to +investigate). Never dismiss a disagreement. + +## 5. Certificate + +When a PR exists (or as soon as it is created), post the evidence: + +```bash +python3 scripts/pipeline_pr.py comment --repo "$REPO" --pr "$PR" --body-file /cert.md +``` + +`cert.md` starts with `## Verification certificate` and contains: verdict; resolved type chain; +constructor/adversary check counts per section; random seeds and `hypothesis` settings; the n +range and instance families covered; cross-comparison count and disagreements; then both scripts +(and the Typst proof), each collapsed in `

…` code blocks. Without a PR, +report the same content in the conversation. + +## Topology checks + +```bash +cargo run --example detect_isolated_problems # problems with no edges in or out +cargo run --example detect_unreachable_from_3sat # NP-hardness chains from KSatisfiability +pred show # incoming/outgoing edges per variant +pred path --limit all # every path with per-step and overall transforms +``` + +- **Orphans:** list isolated problems with variant counts, and search open issues that would + connect each (`gh issue list --state open --search ""`). Meta-issue #610 tracks orphans. +- **NP-hardness:** unreachable NP-hard problems need a new incoming reduction. Problems in P + (MaximumMatching, 2-SAT, 2-coloring, ...) and intermediate ones (Factoring) are correctly + unreachable — don't file them. The detectors report base names; confirm the specific variant + with `pred show` / `pred path`. +- **Redundancy of a proposed rule S -> T:** run `pred path` with variant-qualified endpoints. If no + path exists, the rule is novel. Otherwise compare each existing path's `Overall:` transform with + the proposed rule's, parameter by parameter. The rule is dominated if some path is no worse on + every target parameter (asymptotically, same parameters). `unavailable` overall transforms and + incomparable expressions (exp/log, different variables) count as not dominated. Report + **PASS** (not dominated or novel) or **WARN** (name the dominating path with variants), and + note practical merits a dominated rule may still have (simpler extraction, fewer steps, teaching). diff --git a/.claude/skills/how-to-write-manual/SKILL.md b/.claude/skills/how-to-write-manual/SKILL.md new file mode 100644 index 000000000..21c17dfa1 --- /dev/null +++ b/.claude/skills/how-to-write-manual/SKILL.md @@ -0,0 +1,126 @@ +--- +name: how-to-write-manual +description: Use when writing, improving, or auditing a problem-def or reduction-rule entry in docs/paper/reductions.typ, or user documentation under docs/src/ (mdBook). +--- + +# How to Write the Manual + +The paper `docs/paper/reductions.typ` is the mathematical manual: one `problem-def` per model, one +`reduction-rule` per directed edge. `docs/src/` is the mdBook user guide (`book.toml` points there). + +## Gold standards + +- Model: `problem-def("MaximumIndependentSet")` — definition, background with cited algorithms, + example from the fixture, figure, `pred-commands`. +- Rule: `reduction-rule("MinimumVertexCover", "MaximumIndependentSet"` and its reverse + `reduction-rule("MaximumIndependentSet", "MinimumVertexCover")` — minimal but complete proof + structure. For a richer worked example, read `reduction-rule("MaximumIndependentSet", "MaximumClique"`. + +## Paper mechanics + +- `display-name` dict (near the top) maps `ProblemName` to display text; every problem needs an entry. +- `problem-def(name, variant: none)[def][body]` auto-inserts complexity table, reduction links and + schema between `def` and `body`; label ``. +- `reduction-rule(source, target, example:, example-source-variant:, example-target-variant:, + example-caption:, extra:)[statement][proof]` auto-derives overhead from the graph export and + registers the edge for the completeness check; label ``. +- Every directed reduction needs its own entry, including reverses. Only `Decision

<-> P` pairs are + exempt. The completeness warning block at the end of the paper lists missing edges; your change + must never increase it. + +## Example data — never invent + +Examples come only from `docs/paper/data/examples.json`, produced by the canonical example specs: +`canonical_model_example_specs()` beside each model and `canonical_rule_example_specs()` beside each +rule (collected in `src/rules/mod.rs`). + +- Models: `#let x = load-model-example("Name")` → keys `problem`, `variant`, `instance`, + `optimal_config`, `optimal_value`. +- Rules: `#let r = load-example("Src", "Tgt")` → `r.source` / `r.target` (`problem`, `variant`, + `instance`) and `r.solutions` (each `source_config`, `target_config`). +- Pass `variant:` (models) or `source-variant:` / `target-variant:` (rules, plus the matching + `example-*-variant:` args on `reduction-rule`) whenever the pair has several fixtures — the loaders + panic on ambiguity or a missing fixture. +- Pull numbers from the loaded data with Typst expressions; do not hardcode values the fixture holds. +- If the fixture is unsuitable (too big, degenerate, doesn't show the structure), fix the canonical + spec in code and rebuild — never substitute a hand-made instance. +- Witness semantics: `solutions.at(0)` is the one canonical witness. Never derive "number of optimal + solutions" from fixture length; if multiplicity matters, argue it from the construction. + +Reproducibility block (after the `*Example.*` paragraph for models, at the top of `extra:` for +rules), using the shared helpers only: + +```typst +#pred-commands( + "pred create --example " + problem-spec(x) + " -o x.json", + "pred solve x.json", + "pred evaluate x.json --config " + cli-config(x.optimal_config), +) +``` + +- Models: `problem-spec(x)` renders `Problem/variant/...`; use it rather than a guessed alias — + canonical fixtures often live on non-default variants. Do not redefine it locally. +- Rules: `"pred create --example " + rule-spec(ex)` (with `ex` from `load-example`) renders + `Src/k=v/... --to Tgt/k=v/...`, which recreates the rule fixture's source instance, so the + `--config` from `ex.solutions` applies to it. `problem-spec(ex.source)` or a bare name creates the + source *model* example instead, a different instance. Keyed tokens avoid ambiguity such as `ILP/i64/bool`. +- `cli-config(...)` renders the JSON `--config`; never `map(str).join(",")`. +- Rules reduce with `pred reduce x.json --via route.json -o bundle.json` (`route.json` = one entry + selected from `pred path Src Tgt -o paths.json`), then `pred solve bundle.json`. `pred reduce` has no + `--to`; `target-spec()` and `load-results` do not exist. + +## Writing content + +**Definitions.** `def` is one self-contained statement: inputs with domains, then objective or +constraint. Every symbol defined before use, in both `def` and proofs. + +**Background.** History, applications, notable special cases, and the best known algorithm woven into +prose. Every complexity claim carries `@citation` naming the algorithm. Brute force with nothing +better known → footnote saying so. Unverified reference → `#footnote[Complexity not independently +verified from literature.]`. Keep prose compatible with the auto-generated variant complexity table. + +**Rule statement.** Construction summary plus overhead hint in source terms; cite the reduction's +source or add the unverified footnote. + +**Proof.** Italic sections `_Construction._`, `_Correctness._` (both directions, $arrow.r.double$ / +$arrow.l.double$), `_Variable mapping._` (only if non-trivial), `_Solution extraction._`. +Reproducible: enough to reimplement. Heavy reductions (roughly 300+ LOC) may sketch the approach and +cite the full construction instead. + +**Source priority for math:** the GitHub issue (`gh issue view N`) → derivation documents if +provided → the implementation (`src/rules/_.rs`, ground truth for the construction). +Never invent a proof; if sources disagree with the code, flag it rather than paper over it. + +**Worked example.** Show the source instance, walk the construction with concrete numbers and where +each target dimension comes from, verify the canonical witness end to end. Figures use the helpers in +`docs/paper/lib.typ` (`g-node`, `g-edge`, `graph-colors`) — follow the MIS figure. + +## Audit checklist + +Report each as PASS / WARN / FAIL with a specific reason; read the Rust source before judging math. + +Problem-def: +1. `display-name` entry; `def` non-empty and self-contained. +2. Background ≥ 2 informative sentences (applications, history, special cases). +3. Complexity claims cited or footnoted; consistent with `declare_variants!` complexity. +4. `*Example.*` loaded via `load-model-example`, matches `examples.json`, hand-checkable, exercises + the defining structure, shows the objective/verifier computation. +5. `#figure(` present where the structure is visual; `pred-commands` uses `problem-spec` (models) or `rule-spec` (rules) + `cli-config`. +6. Definition matches `evaluate()` in `src/models/`; no important special case or relation missing. + +Reduction-rule: +1. Statement matches what `reduce_to()` builds; prose overhead matches the `#[reduction]` transform. +2. Proof has construction, both correctness directions, extraction; more than a one-liner. +3. Example (`example: true`) loaded via `load-example`, verifies a witness end to end, correct + `pred-commands` (`--via`, `cli-config`). +4. Reverse edge, if in the graph, has its own entry. +5. Complexity citation or footnote present; no multiplicity claim drawn from fixture length. + +Report only issues verifiable from the source; "background is one sentence with no applications" +beats "background is thin". + +## Gates + +- `make paper` — runs the example/graph/schema exports itself, then compiles the PDF. Must compile + with no new completeness warnings. +- `make doc` — mdBook plus exports and CLI snippets, for changes under `docs/src/`. diff --git a/.claude/skills/issue-to-pr/SKILL.md b/.claude/skills/issue-to-pr/SKILL.md deleted file mode 100644 index 87734784e..000000000 --- a/.claude/skills/issue-to-pr/SKILL.md +++ /dev/null @@ -1,304 +0,0 @@ ---- -name: issue-to-pr -description: Use when you have a GitHub issue and want to create a PR with an implementation plan that triggers automated execution ---- - -# Issue to PR - -Convert a GitHub issue into a PR: write a plan, create the PR, then execute the plan using subagent-driven-development. - -## Invocation - -- `/issue-to-pr 42` — create PR with plan, then execute (for `[Rule]` issues, verification runs by default) -- `/issue-to-pr 42 --no-verify` — skip mathematical verification for `[Rule]` issues - -For Codex, open this `SKILL.md` directly and treat the slash-command forms above as aliases. The Makefile `run-issue` target already does this translation. - -## Workflow - -``` -Receive issue number - -> Fetch structured issue preflight report - -> Verify Good label and rule-model guards - -> If guards fail: STOP - -> If guards pass: research references, write plan, create or resume PR - -> Execute plan via subagent-driven-development -``` - -## Steps - -### 1. Parse Input - -Extract issue number, repo, and flags from arguments: -- `123` -> issue #123 -- `https://github.com/owner/repo/issues/123` -> issue #123 -- `owner/repo#123` -> issue #123 in owner/repo - -Normalize to: -- `ISSUE=` -- `REPO=` (default `CodingThrust/problem-reductions`) -- `EXECUTE=true|false` -- `NO_VERIFY=true|false` (default `false`; pass `--no-verify` to skip mathematical verification for `[Rule]` issues) - -### 2. Fetch Issue + Preflight Guards - -```bash -ISSUE_JSON=$(python3 scripts/pipeline_checks.py issue-context \ - --repo "$REPO" \ - --issue "$ISSUE" \ - --format json) -``` - -This `issue-context` packet is the expensive deterministic preflight call for `issue-to-pr`. It is allowed exactly once per top-level `issue-to-pr` invocation. After it succeeds, reuse `ISSUE_JSON` for all later guards, resume/create decisions, and summaries instead of calling `issue-context` again. - -Treat `ISSUE_JSON` as the source of truth for the deterministic preflight data: -- `title`, `body`, `labels`, and `comments` provide the issue summary and comment thread -- `kind`, `source_problem`, and `target_problem` provide parsed issue metadata -- `checks.good_label`, `checks.source_model`, and `checks.target_model` provide guard outcomes -- `existing_prs`, `resume_pr`, and `action` tell you whether to resume an open PR instead of creating a new one - -Present the issue summary to the user. **Also review all comments** — contributors and maintainers may have posted clarifications, corrections, additional context, or design decisions that refine or override parts of the original issue body. Incorporate relevant comment content when writing the plan. - -### 3. Verify Issue Has Passed check-issue - -The issue must have already passed the `check-issue` quality gate (Stage 1 validation). Do NOT re-validate the issue here. - -Use `ISSUE_JSON.checks.good_label`: -- If it is `fail` → **STOP**: "Issue #N has not passed check-issue. Please run `/check-issue ` first." -- If it is `pass` → continue. - -### 3.5. Model-Existence Guard (for `[Rule]` issues only) - -For `[Rule]` issues, `ISSUE_JSON` already includes `source_problem`, `target_problem`, and the deterministic model-existence checks. - -- If both `checks.source_model` and `checks.target_model` are `pass` → continue to step 4. -- If either is `fail` → **STOP**. Comment on the issue: "Blocked: model `` does not exist in main yet. Please implement it first (or file a `[Model]` issue)." - -**One item per PR, with one exception:** Do NOT implement a missing model as part of a `[Rule]` PR. `[Rule]` issues still require both models to exist on `main`. The only exception is a `[Model]` issue that explicitly claims direct ILP solvability: that PR should implement both the model and the direct ` -> ILP` rule together. - -### 4. Research References - -Use web search to look up the reference URL provided in the issue. This helps: -- Clarify the formal problem definition and notation -- Understand the reduction algorithm in detail (variable mapping, penalty terms, proof of correctness) -- Resolve any ambiguities in the issue description without bothering the contributor - -If the reference is a paper or textbook, search for accessible summaries, lecture notes, or Wikipedia articles on the same reduction. - -### 5. Write Plan - -Write implementation plan to `docs/plans/YYYY-MM-DD-.md` using `superpowers`. - -The plan MUST reference the appropriate implementation skill and follow its steps: - -- **For ordinary `[Model]` issues:** Follow [add-model](../add-model/SKILL.md) Steps 1-7 as the action pipeline -- **For `[Model]` issues that explicitly claim direct ILP solving:** Follow [add-model](../add-model/SKILL.md) Steps 1-7 **and** [add-rule](../add-rule/SKILL.md) Steps 1-7 for the direct ` -> ILP` rule in the same plan / PR -- **For `[Rule]` issues:** Follow [add-rule](../add-rule/SKILL.md) Steps 1-7 as the action pipeline. By default, `/add-rule` runs mathematical verification (Step 1) before implementation. If `--no-verify` was passed, include `--no-verify` when invoking `/add-rule` to skip verification. - -Include the concrete details from the issue (problem definition, reduction algorithm, example, etc.) mapped onto each step. - -**Plan batching:** The paper writing step (add-model Step 6 / add-rule Step 6) MUST be in a **separate batch** from the implementation steps, so it gets its own subagent with fresh context. It depends on the implementation being complete (needs exports). Example batch structure for a `[Model]` plan: -- Batch 1: Steps 1-5.5 (implement model, register, CLI, tests) -- Batch 2: Step 6 (write paper entry — depends on batch 1 for exports) - -For a `[Model]` issue with an explicit direct ILP claim, use: -- Batch 1: implement the model, register it, add the direct ` -> ILP` rule, and add model + rule tests -- Batch 2: write both the `problem-def(...)` and `reduction-rule(...)` paper entries, regenerate exports / fixtures, and run final ILP-enabled verification - -**Solver rules:** -- Ensure at least one solver is provided in the issue template. Check if the solving strategy is valid. If not, reply under issue to ask for clarification. -- If a `[Model]` issue explicitly claims direct ILP solving, implement the model and the direct ` -> ILP` reduction together in the same PR. Do not leave the ILP rule as a follow-up. -- The direct ILP rule must meet the same completeness bar as a standalone production ILP reduction: exact overhead metadata, feature-gated registration, strong closed-loop / extraction / weighted / infeasible / pathological tests when applicable, CLI/example-db/paper integration, and ILP-enabled workspace verification. -- Otherwise, ensure the information provided is enough to implement a solver. - -**Example rules:** -- Implement the user-provided example in `src/example_db/model_builders.rs` for a model, or in the rule-local `canonical_rule_example_specs()` for a rule. -- Run the relevant exports and verify the generated example data against the user-provided information. -- Present in `docs/paper/reductions.typ` in tutorial style with clear intuition (see KColoring->QUBO section for reference). - -### 6. Create PR (or Resume Existing) - -Use the `ISSUE_JSON.action` and `ISSUE_JSON.resume_pr` fields from Step 2. - -**Validate `resume_pr` before trusting it:** If `action == "resume-pr"`, verify that `resume_pr.head_ref_name` contains the current issue number (e.g., branch name includes `issue-{N}`). If it doesn't match, treat as `action = "create-pr"` instead — the script may have matched an unrelated PR. - -**If an open PR already exists** (`action == "resume-pr"`, validated): -- Switch to its branch: `git checkout ` -- Capture `PR=` -- Skip plan creation — jump directly to Step 7 (execute) - -**Worktree-aware branching:** If you are already inside a `run-pipeline` worktree (the CWD is under `.worktrees/`), the branch is already created — skip `prepare-issue-branch` and use the current branch directly. Only call `prepare-issue-branch` when running standalone (not inside a worktree). - -**If no open PR exists** (`action == "create-pr"`) — create one with only the plan file: - -```bash -# If NOT inside a worktree: prepare or reuse the issue branch -# (skip this if already in a run-pipeline worktree — branch already exists) -BRANCH_JSON=$(python3 scripts/pipeline_worktree.py prepare-issue-branch \ - --issue \ - --slug \ - --base main \ - --format json) -BRANCH=$(printf '%s\n' "$BRANCH_JSON" | python3 -c "import sys,json; print(json.load(sys.stdin)['branch'])") -# If inside a worktree, just use: BRANCH=$(git branch --show-current) - -# Stage the plan file -git add docs/plans/.md - -# Commit -git commit -m "Add plan for #: " - -# Push -git push -u origin "$BRANCH" - -# Create PR body -PR_BODY_FILE=$(mktemp) -cat > "$PR_BODY_FILE" <<'EOF' -## Summary -<Brief description> - -Fixes #<number> -EOF - -# Create PR and capture the created PR number -PR_JSON=$(python3 scripts/pipeline_pr.py create \ - --repo "$REPO" \ - --title "Fix #<number>: <title>" \ - --body-file "$PR_BODY_FILE" \ - --base main \ - --head "$BRANCH" \ - --format json) -PR=$(printf '%s\n' "$PR_JSON" | python3 -c "import sys,json; print(json.load(sys.stdin)['pr_number'])") -rm -f "$PR_BODY_FILE" -``` - -### 7. Execute Plan - -#### 7a. Implement - -Execute the plan using `superpowers:subagent-driven-development`: - -1. **Read the plan** from `docs/plans/<plan-file>.md` -2. **Clear context** — summarize only the plan content and essential file paths, then invoke subagent-driven-development with a clean prompt. Do not carry forward research notes, issue comments, or other accumulated context from prior steps. -3. **Invoke subagent-driven-development** with the plan as input — this dispatches parallel subagents for independent tasks in the plan - -If execution fails, leave the PR open with the plan commit only — the user can run `make run-plan` manually later. Skip remaining sub-steps. - -#### 7b. Commit - -Structural and quality review is handled by the `review-pipeline` stage, not here. The run stage just needs to produce working code. - -Ensure all implementation changes are committed before cleanup. A small coherent commit stack is acceptable, especially when resuming an existing PR or integrating subagent work; do not rewrite history just to collapse commits. If there are still uncommitted implementation changes, commit them now: -```bash -git add -A -git commit -m "Implement #<number>: <title>" -``` - -#### 7c. Clean Up Plan File - -Delete the plan file from the branch — it served its purpose during implementation and should not be merged into main: - -```bash -git rm docs/plans/<plan-file>.md -git commit -m "chore: remove plan file after implementation" -``` - -#### 7d. Push and Post Summary - -Post an implementation summary comment on the PR **before** pushing. This comment should: -- Summarize what was implemented (files added/changed) -- Highlight any **deviations from the plan** — design changes, unexpected issues, or workarounds discovered during implementation -- Note any open questions or trade-offs made - -```bash -COMMENT_FILE=$(mktemp) -cat > "$COMMENT_FILE" <<'EOF' -## Implementation Summary - -### Changes -- [list of files added/modified and what they do] - -### Deviations from Plan -- [any design changes, accidents, or workarounds — or "None"] - -### Open Questions -- [any trade-offs or items needing review — or "None"] -EOF -python3 scripts/pipeline_pr.py comment --repo "$REPO" --pr "$PR" --body-file "$COMMENT_FILE" -rm -f "$COMMENT_FILE" - -# Repo verification may regenerate ignored doc exports (notably after `make paper`). -# Inspect the tree once more before pushing. -git status --short - -# Generated doc exports under docs/src/reductions/ are ignored; do not stage them. - -# The issue plan file must be gone before push. -test ! -e docs/plans/<plan-file>.md - -git push -``` - -#### 7e. Done - -Report final status: -- PR URL and number -- Implementation summary - -The PR is **not merged** and review is **not** handled here. The separate `review-pipeline` skill picks up PRs from the `Review pool` board column to run agentic review (structural check, quality check, agentic feature tests). - -## Example - -``` -User: /issue-to-pr 42 - -Claude: Let me fetch issue #42... - -[Fetches issue: "[Rule] IndependentSet to QUBO"] -[Verifies Good label — passed] -[Researches references] -[Writes docs/plans/2026-02-09-independentset-to-qubo.md] -[Creates branch, commits, pushes] -[Creates PR] - -Created PR #45: Fix #42: Add IndependentSet -> QUBO reduction -``` - -``` -User: /issue-to-pr 42 - -Claude: Let me fetch issue #42... - -[Fetches issue: "[Rule] IndependentSet to QUBO"] -[Verifies Good label — passed] -[Researches references] -[Writes docs/plans/2026-02-09-independentset-to-qubo.md] -[Creates branch, commits, pushes] -[Creates PR] -[Continues to execute...] - -Executing plan via subagent-driven-development... -[Subagents implement the plan steps] -[Pushes] - -PR #45 created and pushed. -Run /review-pipeline to run agentic review (structural check, quality check, agentic tests). -``` - -## Common Mistakes - -| Mistake | Fix | -|---------|-----| -| Issue not checked | Run `/check-issue <N>` first — issue-to-pr requires it | -| Issue has failure labels | Fix the issue, re-run `/check-issue`, then retry | -| Including implementation code in initial PR | First PR: plan only | -| Generic plan | Use specifics from the issue, mapped to add-model/add-rule steps | -| Skipping CLI registration in plan | add-model still requires alias/create/example-db planning, but not manual CLI dispatch-table edits | -| Not verifying facts from issue | Use WebSearch/WebFetch to cross-check claims | -| Branch already exists on retry | Use `pipeline_worktree.py prepare-issue-branch` — it will reuse the existing branch instead of failing on `git checkout -b` | -| Dirty working tree | Use `pipeline_worktree.py prepare-issue-branch` — it stops before branching if the worktree is dirty | -| Resuming wrong PR | Always validate `resume_pr.head_ref_name` contains `issue-{N}` before trusting it — GitHub search can return false positives | -| `prepare-issue-branch` inside worktree | Skip it when inside a `run-pipeline` worktree (CWD under `.worktrees/`) — the branch already exists | -| Bundling unrelated model + rule in one PR | Keep the normal one-item-per-PR rule. The only exception is a `[Model]` issue that explicitly claims direct ILP solving, which should ship with its direct `<Model> -> ILP` rule | -| Plan files left in PR | Delete plan files before final push (Step 7c) | -| `make paper` or export steps changed tracked JSON after verification | Run `git status --short`, stage expected generated exports, and STOP if unexpected files remain before push | diff --git a/.claude/skills/propose/SKILL.md b/.claude/skills/propose/SKILL.md index 5024e0af5..782748d96 100644 --- a/.claude/skills/propose/SKILL.md +++ b/.claude/skills/propose/SKILL.md @@ -1,837 +1,139 @@ --- name: propose -description: Use when a user wants to propose a new problem model or reduction rule — guides them through brainstorming, clarifies the design, and files a GitHub issue +description: Use when someone wants to propose a new problem model or reduction rule for the library — interviews them in mathematical language, fills gaps from the reduction graph and literature, checks the draft, and files `[Model]` / `[Rule]` GitHub issues --- -# Propose a New Model or Rule +# Propose a Model or Rule -Interactive brainstorming skill that helps domain experts (who may not know the codebase) design a new problem model or reduction rule, then files well-formed GitHub issues. +Turn a domain expert's idea into GitHub issues that pass `how-to-triage-issue` on the +first try. The proposer may know no Rust; speak mathematics only. -**No programming knowledge required.** This skill works entirely in mathematical / domain language. - -## Invocation - -``` -/propose -/propose model -/propose rule -``` +Invocation: `/propose`, `/propose model`, `/propose rule`. <HARD-GATE> -Do NOT write any code, create any files, or invoke implementation skills (add-model, add-rule, issue-to-pr). -Exception: downloading main reference PDFs into `docs/research/raw/` is allowed and required by the literature-check step. -The ONLY project output of this skill is GitHub issues filed via `gh issue create`; reference PDFs are supporting evidence, not implementation artifacts. +No code, no repo edits, no implementation skills. The only repo write allowed is downloading +reference PDFs into `docs/research/raw/`. The output is GitHub issues filed with `gh issue create`, +and only after the user explicitly approves the final drafts. </HARD-GATE> -## Process - -```dot -digraph propose { - rankdir=TB; - "Start" [shape=doublecircle]; - "Detect type" [shape=diamond]; - "Brainstorm Model" [shape=box]; - "Topology analysis" [shape=box]; - "Propose rules?" [shape=diamond]; - "Brainstorm Rule(s)" [shape=box]; - "Select rule pair" [shape=box]; - "Study models" [shape=box]; - "Literature check" [shape=diamond]; - "Guided brainstorming" [shape=box]; - "Present draft(s)" [shape=box]; - "Run check-issue on draft" [shape=box]; - "Checks pass?" [shape=diamond]; - "User approves?" [shape=diamond]; - "File issue(s)" [shape=box]; - "Done" [shape=doublecircle]; - - "Study conventions" [shape=box]; - "Overlap check" [shape=diamond]; - - "Start" -> "Detect type"; - "Detect type" -> "Study conventions" [label="model or rule"]; - "Detect type" -> "Start" [label="ask user"]; - "Study conventions" -> "Overlap check" [label="model"]; - "Overlap check" -> "Brainstorm Model" [label="new or variant"]; - "Overlap check" -> "Select rule pair" [label="pivot to rules"]; - "Study conventions" -> "Select rule pair" [label="rule"]; - "Select rule pair" -> "Study models"; - "Study models" -> "Literature check"; - "Literature check" -> "Guided brainstorming" [label="pre-fill if found"]; - "Guided brainstorming" -> "Present draft(s)"; - "Brainstorm Model" -> "Topology analysis"; - "Topology analysis" -> "Propose rules?"; - "Propose rules?" -> "Brainstorm Rule(s)" [label="yes"]; - "Propose rules?" -> "Present draft(s)" [label="no, later"]; - "Brainstorm Rule(s)" -> "Present draft(s)"; - "Present draft(s)" -> "Run check-issue on draft"; - "Run check-issue on draft" -> "Checks pass?"; - "Checks pass?" -> "Present draft(s)" [label="fix issues"]; - "Checks pass?" -> "User approves?" [label="pass"]; - "User approves?" -> "Present draft(s)" [label="revise"]; - "User approves?" -> "File issue(s)" [label="yes"]; - "File issue(s)" -> "Done"; -} -``` - ---- - -## Step 1: Detect Type - -If the user didn't specify, use `AskUserQuestion`: - -``` -AskUserQuestion: - question: "What would you like to propose?" - header: "Type" - options: - - label: "New problem (model)" - description: "Define a new computational problem to add to the reduction graph" - - label: "New reduction rule" - description: "Add a reduction between two existing problems" -``` - ---- - -## Step 1b: Study Conventions - -Right after the user picks model or rule, **study at least one existing case** in the relevant category before asking any brainstorming questions. This grounds the conversation in the project's actual conventions and helps produce higher-quality drafts. - -### For Models - -1. Ask the user a brief orienting question (free text): - > "What problem are you thinking of? A name or rough description is enough." - -2. Based on the answer, identify the most similar existing problem in the graph. Use `pred list --json` to find candidates, then use `pred show <similar_problem>` to study one in detail: - - ```bash - pred show <similar_problem> --json - ``` - - Also find and read one closed `[Model]` issue in the same category: - - ```bash - gh issue list --label model --state closed --limit 20 --json number,title,body | jq '[.[] | select(.title | test("<keyword>"; "i"))] | .[0]' - ``` - - If no keyword match, just read the most recent closed model issue to see the template conventions. - -3. **Note internally** (do not dump raw output to the user): - - What fields / parameters the similar problem has - - How the issue defines variables, schema, complexity - - What level of mathematical detail is expected in examples - - How the "Reduction Rule Crossref" section is structured - - Use these conventions to guide the brainstorming questions and draft formatting in later steps. - -4. **Check for overlap with existing problems.** If the most similar problem is the *same problem* the user described (or a close generalization/restriction), surface this immediately via `AskUserQuestion` before continuing brainstorming: - - ``` - AskUserQuestion: - question: "I found that <ExistingProblem> already exists [status]. Your idea looks like a [variant/restriction/generalization]. How should we proceed?" - header: "Overlap" - options: - - label: "Propose a specialized variant" - description: "Define a new problem type for the restricted case (e.g., tournament restriction of a general digraph problem)" - - label: "Propose rules for the existing problem" - description: "Connect the existing problem to the graph — switches to the rule proposal flow" - - label: "Both — variant + rules" - description: "Propose the specialized variant AND rules connecting both versions" - ``` - - - If the user picks **"Propose rules"**, pivot to the **For Rules** flow (Step 3 for Rules). This is the only supported mid-flow type pivot. - - If the user picks **"Both"**, continue the model flow and flag that companion rules should also connect the existing problem. - - **Show the existing problem's schema** so the user can design for compatibility: - > "Here's how <ExistingProblem> is defined: [field summary from `pred show`]. Your variant can mirror this structure where applicable." - - If the existing problem is an orphan, mention this — it increases the value of connecting both problems. - -### For Rules - -1. **Run topology analysis first** to identify the most impactful missing reductions, then present the top candidates as recommendations. Only ask the user an open-ended "which two problems?" question if they don't have a specific pair in mind — otherwise, use the topology data to populate `AskUserQuestion` options in Step 3.1 directly. - - Run these commands silently before asking any questions: - ```bash - # Core data (fast — uses pre-built pred binary) - pred list --json - gh issue list --label rule --state open --limit 500 --json number,title - - # Topology analysis (slower — compiles example binaries, but gives orphan/NP-hardness data) - cargo run --example detect_isolated_problems 2>/dev/null - cargo run --example detect_unreachable_from_3sat 2>/dev/null - ``` - Run the first two commands in parallel. The example binaries take longer but provide essential orphan/NP-hardness gap data for ranking recommendations. - -2. Based on the topology results, study one existing reduction between similar problems. Use `pred to` and `pred from` to find existing reductions, then pick the most relevant one and examine it: - - ```bash - pred to <problem> --json # problems that reduce TO this one (incoming) - pred from <problem> --json # problems this reduces FROM (outgoing) - ``` - - Also find and read one closed `[Rule]` issue in a similar domain: - - ```bash - gh issue list --label rule --state closed --limit 20 --json number,title,body | jq '[.[] | select(.title | test("<keyword>"; "i"))] | .[0]' - ``` - -3. **Note internally**: - - How the reduction algorithm is structured (numbered steps, symbol definitions) - - How the parameter transform table is formatted (field names, formulas) - - How the example is worked through (source → construction → target → solution) - - What references and validation methods are used - - Use these conventions to guide the brainstorming questions and draft formatting in later steps. - -> **Key:** This step asks the user only one light question (to orient the search), then does silent research. Do not show the user raw JSON or code output — just absorb the conventions and let them shape your subsequent questions. - ---- - -## Step 2: Explore Context - -Before asking questions, check what already exists. Use `pred` if it's already installed; only build if the command is missing. - -```bash -# Only build if pred is not already installed — make cli takes >1 minute -command -v pred >/dev/null 2>&1 || make cli -pred list --json -``` - -Also **search for existing GitHub issues** to avoid duplicates and surface related work: - -```bash -# Search open rule issues for related reductions -gh issue list --label rule --state open --limit 500 --json number,title - -# Search open model issues for related problems -gh issue list --label model --state open --limit 500 --json number,title -``` - -Filter the results for keywords matching the user's area of interest (e.g., "knapsack", "traveling", "coloring"). When presenting suggestions in Step 3, **note any existing issues** that overlap — e.g., "Note: #138 SubsetSum→Knapsack already filed." - -This tells you what problems and reductions are already in the graph — essential for: -- Avoiding duplicate model proposals -- Avoiding duplicate rule proposals (check existing issues!) -- Identifying which problems a new rule could connect to -- Suggesting natural reduction targets - ---- - -## Step 3: Brainstorm (one question at a time) - -Ask questions **one at a time**. Prefer multiple-choice when possible. Use mathematical language, not programming language. - -### For Models - -Work through these topics in order, using `AskUserQuestion` where multiple-choice is natural. Adapt based on answers. (The orienting "What problem?" question was already asked in Step 1b.) - -**Auto-inference rule:** For questions 1 (motivation), 2 (problem type), 3 (variables), and 7 (data representation), the answer is often determinable from the user's orienting description. When this is the case, **state the inferred answer as a brief confirmation** instead of presenting an open-ended `AskUserQuestion`: -> "Based on your description, this is a minimization problem on a graph input with permutation variables — correct?" - -Only fall back to the full `AskUserQuestion` if the inference is genuinely ambiguous. Expert users find obvious multiple-choice questions annoying. - -1. **Why useful?** — If the user already explained the motivation in the orienting question, acknowledge it and move on. Only use `AskUserQuestion` if the motivation is unclear: - ``` - AskUserQuestion: - question: "What's the motivation for this problem? Where does it appear?" - header: "Motivation" - options: - - label: "Combinatorial optimization" - description: "Scheduling, routing, packing, allocation problems" - - label: "Physics / simulation" - description: "Spin systems, ground states, quantum computing" - - label: "Cryptography / number theory" - description: "Factoring, lattice problems, code-based crypto" - - label: "Something else" - description: "I'll describe the domain" - ``` - -2. **Definition** — Infer the problem type from the user's description (e.g., "find the largest..." → maximize, "find the smallest..." → minimize, "does there exist..." → satisfaction). If the inference is clear, confirm it inline: "This is a minimization problem — correct?" Only use the full `AskUserQuestion` if ambiguous: - ``` - AskUserQuestion: - question: "What kind of problem is this?" - header: "Problem type" - options: - - label: "Optimization (maximize)" - description: "Find a solution that maximizes an objective function" - - label: "Optimization (minimize)" - description: "Find a solution that minimizes an objective function" - - label: "Satisfaction (yes/no)" - description: "Find any solution that meets all constraints, or decide if one exists" - ``` - Then ask: "Can you state the problem formally? What's the input, constraints, and objective?" (Skip if the user already provided a formal definition in the orienting question.) - -3. **Variables** — Infer from the problem structure (vertex/edge selection → binary, coloring → k-valued, routing/ranking → permutation). If the inference is clear, confirm inline. Only use `AskUserQuestion` if ambiguous: - ``` - AskUserQuestion: - question: "How would you represent a solution? What are the decision variables?" - header: "Variables" - options: - - label: "Binary selection" - description: "Each variable is 0 or 1 (e.g., include/exclude)" - - label: "k-valued assignment" - description: "Each variable takes one of k values (e.g., coloring)" - - label: "Permutation" - description: "An ordering of all elements (e.g., tour)" - - label: "Other domain" - description: "I'll describe the variable structure" - ``` - -4. **Complexity & Reference** — Before asking, use WebSearch to research the best known exact algorithms and canonical references for this problem. Then present up to 3 candidates via `AskUserQuestion`, each combining the complexity bound with its source: - - ``` - AskUserQuestion: - question: "What is the best known exact algorithm for this problem?" - header: "Complexity" - options: - - label: "O(<expression>) — <algorithm/author>" - description: "<paper title, year> — <URL>" - - label: "O(<expression>) — <algorithm/author>" - description: "<paper title, year> — <URL>" - - label: "O(<expression>) — <algorithm/author>" - description: "<paper title, year> — <URL>" - - label: "I know a different bound" - description: "I'll provide the complexity, reference, and link" - ``` - - Requirements: - - Use concrete numeric exponents (e.g., `1.1996^n`, not `(2-ε)^n`) - - Every option must include a link to the paper or resource - - For the selected main algorithm/reference, download the PDF into `docs/research/raw/` and read the relevant theorem/algorithm/proof before drafting. Prefer existing paper tooling (`python3 scripts/fetch_papers.py lookup/download/scihub`) when the reference is or will be in `docs/paper/references.bib`; otherwise save the directly fetched PDF with a clear slug. If no full paper is obtainable, state that explicitly and do not present unverified abstract-level claims as checked facts. - - After the user picks one, fetch the BibTeX entry for the chosen reference (from the paper's page, DOI resolver, or Google Scholar) and record it — the BibTeX will be included in the filed issue - -5. **Solving strategy** — The library's brute-force solver works on every problem by enumerating the configuration space. **Auto-fill "Brute-force" as the baseline** — do not present it as a choice. - - If the problem also admits a natural **direct** ILP formulation, ask whether the issue should explicitly claim ILP solver support: - ``` - AskUserQuestion: - question: "This problem appears to admit a direct ILP formulation. Should the issue explicitly claim ILP solver support?" - header: "ILP solver" - options: - - label: "Yes — add direct <Problem> → ILP (Recommended)" - description: "File a direct companion rule issue and mark the model as solvable via ILP" - - label: "No — brute-force only" - description: "Keep the model issue's solver section limited to brute-force unless another concrete solver exists" - ``` - - If the user picks **Yes**: - - Set `requires_ilp_companion = true` - - The companion rule must be a direct `<Problem> → ILP` issue filed in the **same** `/propose` session - - State in the draft that later implementation is expected to ship the model and this direct ILP rule in the same PR - - If the user picks **No**, keep the model issue's "How to solve" section limited to brute-force unless a different specialized solver is available. - - **Do not mention ILP/QUBO in the model issue's "How to solve" section unless there is a concrete companion rule issue number.** - - **Only use `AskUserQuestion`** beyond that if the problem is polynomial-time solvable or has a specialized exact algorithm that should replace brute-force: - ``` - AskUserQuestion: - question: "This problem appears to be solvable in polynomial time. Which algorithm should be the primary solver?" - header: "Solver" - options: - - label: "<algorithm> (Recommended)" - description: "<why — e.g., runs in O(n^3) via Hungarian method>" - - label: "Brute-force anyway" - description: "Use generic brute-force even though faster algorithms exist" - ``` - -6. **Example** — Generate **at least 3** candidate examples yourself (varying in size and structure), then present via `AskUserQuestion`. **3 options is the minimum — never fewer.** Always include a "Generate new batch" escape hatch: - - ``` - AskUserQuestion: - question: "Which example instance should we use?" - header: "Example" - options: - - label: "<small instance summary>" - description: "<brief description — minimal but valid>" - - label: "<medium instance summary>" - description: "<brief description — exercises core structure>" - - label: "<larger instance summary>" - description: "<brief description — richer, more illustrative>" - - label: "Generate new batch" - description: "None of these work — generate a fresh set of examples" - ``` - - If the user picks "Generate new batch", create 3 new examples with different sizes/structures and re-present. - - After the user picks a concrete example, provide a complete instance with its expected outcome. - - For optimization problems: give at least one optimal solution and the optimal objective value - - For satisfaction problems: give at least one valid / satisfying solution and explain briefly why it is valid; also provide a NO instance with a clear infeasibility argument - - Must exercise the problem's core structure — pick instances where constraints interact nontrivially (e.g., multiple constraints are tight), not degenerate cases - - Must be small enough to verify by hand but large enough that a naive/incorrect implementation would produce a wrong answer - -7. **Data representation** — Infer from the problem definition (e.g., "vertices and edges" → graph, "rows and columns" → matrix, "universe and subsets" → set system). If the inference is clear from the user's description, confirm inline: "The input is a graph — correct?" Only use `AskUserQuestion` if ambiguous: - ``` - AskUserQuestion: - question: "What data defines an instance of this problem?" - header: "Input data" - options: - - label: "A graph" - description: "Vertices and edges, possibly weighted" - - label: "A matrix" - description: "Rows and columns of numbers" - - label: "A set system" - description: "A universe of elements and a collection of subsets" - - label: "Something else" - description: "I'll describe the input structure" - ``` - -8. **Variants** — Based on the data representation answer, ask about applicable variants using `AskUserQuestion`. Only show options that are viable for the problem's input structure. Skip this question entirely if no variants apply (e.g., the problem has a fixed unique input structure like Knapsack, Factoring, SubsetSum). - - **If the input is a graph** (from step 7), ask about graph topology: - ``` - AskUserQuestion: - question: "Which graph topologies should this problem support?" - header: "Graph topology" - multiSelect: true - options: - - label: "General graphs" - description: "No structural restriction (SimpleGraph) — default, almost always needed" - - label: "Planar graphs" - description: "Graphs embeddable in the plane without edge crossings" - - label: "Bipartite graphs" - description: "Graphs whose vertices split into two groups with edges only between groups" - - label: "Unit disk graphs" - description: "Intersection graphs of unit disks in the plane" - - label: "Kings subgraph" - description: "Subgraphs of the king's graph on a grid" - - label: "Triangular subgraph" - description: "Subgraphs of the triangular lattice" - ``` - Only include topology options that are meaningful for the problem (e.g., don't offer "Kings subgraph" for a problem that doesn't have special structure on grids). - - **If the problem can be weighted or unweighted**, ask: - ``` - AskUserQuestion: - question: "Should this problem support weighted instances?" - header: "Weights" - options: - - label: "Unweighted only" - description: "All elements have unit weight — simpler formulation" - - label: "Weighted (integers)" - description: "Elements have integer weights" - - label: "Weighted (real numbers)" - description: "Elements have real-valued weights" - - label: "Both weighted and unweighted" - description: "Support unit weight and integer weight variants" - ``` - Skip this if the problem inherently requires specific numeric values (e.g., QUBO always has a weight matrix, Knapsack always has item values). - - **If the problem has a parameter K** (e.g., K-coloring, K-satisfiability), ask: - ``` - AskUserQuestion: - question: "Should K be a fixed constant or a general parameter?" - header: "K parameter" - options: - - label: "General K" - description: "K is part of the input — problem is NP-hard for general K" - - label: "Fixed small K values" - description: "Define variants for specific K (e.g., K=2, K=3) with different complexities" - - label: "Both" - description: "General K plus specific fixed-K variants with known better algorithms" - ``` - Skip this if the problem has no natural K parameter. - - Record the chosen variants — they will appear in the Schema section of the issue draft (the "Variants" field). - -After model brainstorming is complete, proceed to **Step 3b: Topology Analysis**. - -### For Rules (standalone) - -Topology analysis was already run in Step 1b. Conventions were studied. - -#### Step 3.1: Recommend rules - -Use the topology data from Step 1b to present **data-driven recommendations** via `AskUserQuestion`. The options should be populated from the analysis — do not ask the user to name problems before you have analyzed the graph. - -``` -AskUserQuestion: - question: "Which reduction would you like to propose?" - header: "Reduction" - options: - - label: "<Source> → <Target> (Recommended)" - description: "<why most valuable — e.g., connects orphan X, proves NP-hardness, existing issue #N>" - - label: "<Source> → <Target>" - description: "<why valuable — note existing issues if any>" - - label: "<Source> → <Target>" - description: "<why valuable — note existing issues if any>" - - label: "I have a different pair" - description: "I'll describe the source and target problems" -``` - -**Selection criteria** (in priority order) — only suggest rules where **both source and target already exist** in the codebase: -- **Priority 1:** Rules that connect orphan problems to the main component (check `detect_isolated_problems` output) -- **Priority 2:** Rules that fill NP-hardness proof gaps (check `detect_unreachable_from_3sat` output) -- **Priority 3: ILP solver path for leaf problems** — If a problem has **no outgoing reductions** (check `pred from <problem>` returns empty), prioritize proposing `<Problem> → ILP` when the problem has a natural ILP formulation (linear constraints over integer variables). This is the most common way to make a new problem solvable. Run `pred from <problem> --json` for each orphan/leaf to detect missing outgoing edges. -- **Priority 4:** Other rules to large clusters (QUBO, ILP, SAT families) -- **Filter:** Exclude pairs that already have a reduction registered **AND** pairs that already have an open GitHub issue filed (even if not yet implemented). Do not recommend duplicates — if an issue exists, it should be implemented via `/issue-to-pr`, not re-proposed. - -After selection, verify both problems exist (or one is being proposed alongside). - -#### Step 3.2: Study source and target models - -**Mandatory before any brainstorming.** Inspect both models in the codebase: - -```bash -pred show <source> --json -pred show <target> --json -``` - -Note internally (do not dump to the user): -- Field names, types, and size getters for both problems -- Whether source/target are optimization or satisfaction problems -- Type mismatches (e.g., `BigUint` vs `i64`) that the reduction must handle -- Existing reductions to/from both problems (use `pred to` and `pred from`) +## What a finished proposal contains + +The issue must fill every section of `.github/ISSUE_TEMPLATE/problem.md` (model) or +`.github/ISSUE_TEMPLATE/rule.md` (rule). Read the template before drafting. Quality bars: + +- **Reference read, not just cited.** Download the main reference PDF to `docs/research/raw/` + and read the actual theorem, construction, and proof. Name the exact theorem/section in the + issue. A rule's algorithm section must be implementable (notation, gadgets, constraint + families, parameter choices, solution extraction); a citation-only algorithm fails. + Flag ambiguities or transcription risks in the source instead of smoothing them over. If no + full text is obtainable, say so and mark the proposal unverified. +- **Bug-catching example.** Small enough to verify by hand, but with constraints that interact + and are tight, so a wrong implementation gives a wrong answer. No triangles, all-zeros, or + single-element degenerate cases. Models: optimization gives an optimal witness and value; + feasibility gives a YES witness with justification plus a NO instance with an infeasibility + argument. Rules: fully worked source → construction → target, with optimal vs suboptimal (or + YES vs NO) behavior visible. Build the witnesses yourself; never ask the proposer for them. +- **No orphans.** A model needs at least one companion rule issue connecting it to the graph. + An orphan model gets rejected in review; if the proposer insists on skipping, put a visible + warning in the Reduction Rule Crossref section. +- **ILP claim needs its rule.** Check "solvable by reducing directly to ILP" only when a direct + `[Rule] <Model> to ILP` companion issue is filed in this same session. Checking it means the + model and that rule ship together in one PR. Otherwise the model is brute-force only; never + mention ILP/QUBO under "How to solve" without a concrete issue number. +- **Concrete complexity.** Best known exact algorithm with numeric bases/exponents + (`1.1996^n`, not `(2-ε)^n`, not `2^(ω n/3)`), author, year, link, and BibTeX. Polynomial + problems must not be given exponential bounds. +- **Math language, not implementation types.** Fields are "a graph", "nonnegative integer + weights", "a list of subsets of {0,…,n−1}". Do not ask the proposer to pick Rust types, + integer widths, or trait names; implementers derive those. Variables are described as a + configuration vector: count, per-variable domain, meaning. +- **Every symbol defined before use.** Size-overhead formulas use names from the target's + `parameters` (`pred show <target> --json`), and source parameters on the right-hand side. +- **BibTeX** for each reference at the end of the issue. + +## Infer silently, ask only for gaps + +Before asking anything, gather what the tools already know. Build `pred` if missing +(`command -v pred || make cli`; it takes minutes, so don't rebuild when present). -Also check for existing GitHub issues for this specific pair: ```bash -gh issue list --label rule --state open --json number,title | jq '.[] | select(.title | test("<source>.*<target>"; "i"))' -``` - -This information is essential for writing correct overhead tables and identifying implementation concerns. - -#### Step 3.3: Literature check - -After studying the models, check whether this is a **well-known textbook reduction**: -- Use WebSearch to check standard references (Garey & Johnson, Karp's 21, CLRS, Sipser, Arora & Barak) -- Check if an existing GitHub issue already describes this reduction - -Then identify the **main algorithm reference** for the reduction: the paper/book chapter whose construction will be implemented. Download its PDF into `docs/research/raw/` and read it carefully before drafting: +pred list --json # catalog; `pred list <query>` to search names/aliases +pred list --rules --all # every registered rule +pred show <Problem> --json # fields (schema/inputs), parameters, complexity, reduces_to/from +pred to <Problem> --hops 2 --json # incoming neighbors +pred from <Problem> --hops 2 --json # outgoing neighbors +pred path <Src> <Tgt> --limit 5 # set of existing paths (not a single "best" one) +gh issue list --label model --state all --limit 500 --json number,title,state +gh issue list --label rule --state all --limit 500 --json number,title,state +cargo run --example detect_isolated_problems # orphan problems +cargo run --example detect_unreachable_from_3sat # problems lacking an NP-hardness chain +``` + +Read one closed issue of the same kind and domain to match the expected level of detail. +Use WebSearch for literature; for references already in `docs/paper/references.bib`, +`python3 scripts/fetch_papers.py lookup|download|scihub` fetches PDFs, otherwise +`curl -L '<pdf-url>' -o docs/research/raw/<author-year-short-title>.pdf` and check with `file`. + +Then interview **one question at a time**, each with a recommended answer the user can accept. +State inferred facts as one-line confirmations ("This is a minimization over binary vertex +choices on a general graph — right?") instead of multiple-choice menus. Pre-fill textbook +reductions (Garey & Johnson, Karp) from the literature, but still let the user confirm the +algorithm, correctness argument, overhead, and example before drafting. + +## Model proposals + +1. Get a name or rough description. Check for overlap: if `pred show` finds the same problem + or a restriction/generalization, say so and let the user choose between a new variant, + rules for the existing problem instead, or both. Mention if the existing one is an orphan. +2. Fill definition, variables, input data, and variants (graph topologies, weighted or not, + fixed vs general K) only as far as they are meaningful for this problem. +3. Complexity and reference: offer up to three literature candidates with links, read the + chosen paper, fetch its BibTeX. +4. Solving: brute force is the baseline, never a question. Ask about a direct ILP claim only + when a natural linear formulation exists (see the ILP bar above). Mention a polynomial-time + algorithm if one is known. +5. Example and expected outcome: propose about three candidates of different sizes; the user + picks or asks for a new batch. +6. Companion rules: rank candidates as in the rule priorities below, with `<Model> → ILP` on + top when the model has no outgoing edge and admits a linear formulation, and mandatory when + the ILP claim is made. Brainstorm each chosen rule with the rule flow, lighter on context. + +## Rule proposals + +Rank candidate pairs from the topology data before asking which pair the user wants. Both +endpoints must exist (or be proposed in this session). Priorities: + +1. Connect an orphan from `detect_isolated_problems`. +2. Fill an NP-hardness gap from `detect_unreachable_from_3sat` (a reduction *from* a problem + already reachable from 3-SAT). +3. Give a leaf problem (empty `pred from`) a path to ILP via `<Leaf> → ILP`. +4. Connect to a large cluster (QUBO, ILP, SAT families). + +Exclude pairs that are already registered and pairs with an existing open or closed issue — +point the user to that issue instead. Also run `pred path <Src> <Tgt>`: if paths already exist, +the rule must justify itself (better overhead, NP-hardness direction, simpler witness mapping). + +Then, for the chosen pair, study both endpoints with `pred show --json` (fields, parameters, +optimization vs feasibility, value domains such as big integers vs bounded integers), read the +main reference, and settle with the user: motivation, algorithm, correctness argument +(feasibility or optimality preserved), size overhead, validation method, worked example. + +## Check, approve, file + +1. Draft every issue in full and run the `how-to-triage-issue` checks against the drafts: + usefulness (`pred show <NewModel>` fails; for rules, no redundant existing path), non- + triviality (not a renaming, subtype coercion, or variable substitution), correctness + (references exist in `docs/paper/references.bib` or verifiably online; claims match the + PDF), completeness (all template sections), writing (symbols consistent, overhead names + match target `parameters`, example fully worked), ILP-claim consistency. Fix what you can; + ask for the rest. +2. Show the drafts and get explicit approval ("file it"). Revisions loop back to step 1. +3. File the model first, then its rules, then patch the model's crossref with real numbers: ```bash -# Preferred when the reference is already in or being added to docs/paper/references.bib -python3 scripts/fetch_papers.py lookup -python3 scripts/fetch_papers.py download -python3 scripts/fetch_papers.py scihub - -# Fallback for an open PDF URL -curl -L '<pdf-url>' -o docs/research/raw/<author-year-short-title>.pdf +gh issue create --title "[Model] <Name>" --label model --body-file <draft> +gh issue create --title "[Rule] <Source> to <Target>" --label rule --body-file <draft> # mention #<model> +gh issue edit <model-number> --body-file <updated-draft> # real rule issue numbers in crossref / How to solve ``` -Requirements: -- Do not rely only on abstracts, snippets, or secondary summaries for the construction. -- Read the theorem statement, construction/gadget definitions, proof lemmas, and any assumptions/normalization steps. -- If the source has ambiguities, transcription risks, or missing details, call them out in the issue draft rather than smoothing them over. -- The issue must include an implementable algorithm, not just a citation. -- If no full reference can be obtained, tell the user and mark the proposal as unverified until the paper is available. - -If the reduction is well-known, use the literature to **pre-fill** answers in Step 3.4 — but still present each step to the user for confirmation. Do NOT skip the guided brainstorming. - -#### Step 3.4: Guided brainstorming - -**Always run this step**, whether the reduction is well-known or novel. For well-known reductions, pre-fill answers from literature and present them for confirmation. For novel reductions, ask the user to provide answers. Work through these topics in order, **one at a time**. - -1. **Why useful?** — State the motivation (e.g., connects orphan, fills NP-hardness gap) and present for confirmation via `AskUserQuestion`: - ``` - AskUserQuestion: - question: "What's the main motivation for this reduction?" - header: "Motivation" - options: - - label: "<inferred motivation> (Recommended)" - description: "<why — e.g., connects orphan PaintShop to QUBO hub>" - - label: "<alternative motivation>" - description: "<why>" - - label: "<alternative motivation>" - description: "<why>" - ``` - -2. **Algorithm** — Research the reduction algorithm (use WebSearch for well-known reductions, ask the user for novel ones). Present candidate approaches via `AskUserQuestion`: - ``` - AskUserQuestion: - question: "Which reduction approach should we use?" - header: "Algorithm" - options: - - label: "<approach 1> (Recommended)" - description: "<brief summary of how it works>" - - label: "<approach 2>" - description: "<brief summary>" - - label: "<approach 3>" - description: "<brief summary>" - ``` - After the user picks one, present the full algorithm write-up for confirmation. - - Must define all symbols before using them - - Must be detailed enough that someone could implement it - - Must reflect the downloaded main reference directly: include normalization assumptions, gadget/variable definitions, edge/constraint families, parameter settings, and solution extraction when the reference provides them - -3. **Explanation** — Present a correctness argument explaining why the reduction preserves feasibility (for satisfaction problems) or optimality (for optimization problems), then ask for feedback via `AskUserQuestion`: - ``` - AskUserQuestion: - question: "How does this explanation look?" - header: "Explanation" - options: - - label: "Looks good" - description: "The correctness argument is clear and complete" - - label: "More detail" - description: "Please expand the argument with more steps or formal reasoning" - - label: "Less detail" - description: "Too verbose — please shorten to the key insight" - ``` - If the user asks for more or less detail, revise and re-present. - -4. **Parameter transform** — Compute overhead from the algorithm using the target's parameters from `pred show <target> --json`. Present the overhead table and ask for confirmation: - > "Based on the algorithm, the parameter transform is: [table]. Does this look correct?" - -5. **Example** — Generate **at least 3** candidate examples yourself (varying in size and structure), then present via `AskUserQuestion`. **3 options is the minimum — never fewer.** Always include a "Generate new batch" escape hatch: - - ``` - AskUserQuestion: - question: "Which example instance should we use?" - header: "Example" - options: - - label: "<small instance summary>" - description: "<brief description — e.g., 3 items, capacity 5, optimal: items {1,2}>" - - label: "<medium instance summary>" - description: "<brief description — shows a non-obvious optimum>" - - label: "<larger instance summary>" - description: "<brief description — richer structure, more trade-offs>" - - label: "Generate new batch" - description: "None of these work — generate a fresh set of examples" - ``` - - If the user picks "Generate new batch", create 3 new examples with different sizes/structures and re-present. - - After the user picks a concrete example, fully work out the example: show source instance, each construction step, and the resulting target instance. - - **Validation-oriented examples are mandatory.** The example's primary purpose is to catch bugs in the reduction implementation during closed-loop testing, not just to illustrate the construction. Design examples that: - - **Exercise constraint interactions** — pick instances where multiple constraints are simultaneously tight or near-tight, so an incorrect reduction would produce a wrong answer (e.g., a flow instance where both lower bounds and capacity limits are active on different edges) - - **Distinguish correct from incorrect reductions** — a trivial instance (all zeros, single-element, identity-like) often passes even with a buggy reduction. Choose instances where a naive or partially wrong construction would yield the wrong feasibility/optimality answer - - **Include both YES and NO witnesses** (for satisfaction problems) or **optimal vs suboptimal configurations** (for optimization problems) — show that the reduction correctly maps solutions in both directions - - Must be hand-verifiable but **not** so small that it degenerates into a trivial case - - Do not ask the user to provide solved witnesses manually - -6. **Reference** — Use WebSearch to find references. Present candidate references via `AskUserQuestion`: - ``` - AskUserQuestion: - question: "Which reference should we cite?" - header: "Reference" - options: - - label: "<reference 1> (Recommended)" - description: "<paper title, year> — <URL>" - - label: "<reference 2>" - description: "<paper title, year> — <URL>" - - label: "<reference 3>" - description: "<paper title, year> — <URL>" - ``` - If no references are found, ask the user if this is a novel reduction. - After a reference is chosen, confirm the downloaded PDF path under `docs/research/raw/` and summarize what sections/theorems were read. - ---- - -## Step 3b: Topology Analysis (models only) - -After the model definition is clear, analyze the reduction graph to suggest which rules would be most valuable. Run: - -```bash -# Check orphan problems (to understand graph structure) -cargo run --example detect_isolated_problems 2>/dev/null - -# Check NP-hardness proof gaps (to find problems that need connections) -cargo run --example detect_unreachable_from_3sat 2>/dev/null - -# List existing problems and reductions -pred list --json - -# Check if paths exist between the new problem's likely reduction targets -pred path <similar_problem_A> <similar_problem_B> --json -``` - -Based on the topology analysis, present the user with **suggested reductions** via `AskUserQuestion` (use `multiSelect: true`): - -``` -AskUserQuestion: - question: "Which reductions would you like to propose to connect your problem to the graph? (select one or more)" - header: "Rules" - multiSelect: true - options: - - label: "<Source> → <Target> (Recommended)" - description: "<why most valuable — e.g., proves NP-hardness, connects to main cluster>" - - label: "<Source> → <Target>" - description: "<why valuable>" - - label: "<Source> → <Target>" - description: "<why valuable>" - - label: "I'll file rules separately" - description: "⚠ WARNING: A model with no reduction rules is an orphan node and WILL be rejected during review" -``` - -**Ranking criteria** (in order of priority): -- Connections that establish NP-hardness (from a problem reachable from 3-SAT) -- **ILP solver path** — if the new model has no outgoing edges and admits a natural ILP formulation, `<NewModel> → ILP` should be the top companion rule recommendation. This is the fastest way to make the problem solvable via the existing ILP solver infrastructure. -- If `requires_ilp_companion = true`, the direct `<NewModel> → ILP` rule is mandatory, not optional -- Connections to large clusters (QUBO, ILP, SAT families) -- Connections that reduce orphan count or bridge disconnected components -- Connections the user specifically mentioned during brainstorming - ---- - -## Step 3c: Brainstorm Companion Rules (models only) - -If the user picks one or more rules from Step 3b (or proposes their own): - -For **each** selected rule, run through the rule brainstorming flow (algorithm, correctness, overhead, example, reference) — but keep it lighter since the model context is already established. - -If `requires_ilp_companion = true`, one selected rule **must** be the direct `<NewModel> → ILP` companion. Do not keep the ILP solver claim in the model draft while deferring that rule to a later issue. - -If the user declines ("I'll file rules separately later"): -- **Strongly warn** via `AskUserQuestion`: - ``` - AskUserQuestion: - question: "A problem with no reduction rules is an orphan node — it will be isolated in the graph and REJECTED during review. Are you sure you want to skip?" - header: "⚠ Orphan Warning" - options: - - label: "Let me propose a rule now" - description: "I'll define at least one reduction rule to connect this problem to the graph" - - label: "Skip anyway — I'll file rule issues separately" - description: "I understand the risk. I will file companion rule issues before review." - ``` -- If the user chooses "Let me propose a rule now", go back to Step 3b and let them pick a rule, then brainstorm it. -- If `requires_ilp_companion = true`, the "Skip anyway" option is unavailable unless the model draft is revised to remove the ILP solver claim and revert to brute-force only. -- If the user still declines, include a placeholder in the model's "Reduction Rule Crossref" section noting which rules are planned, and add a visible warning in the draft: "⚠ No companion rule filed — this model will be an orphan node until a rule issue is created." - ---- - -## Step 4: Present Draft Issue(s) - -Once all information is collected, compose the full issue body following the GitHub issue template format. - -If proposing a model + rules, present all drafts together: - -> "Here are the draft issues. Please review — I can revise any section before filing." -> -> **Issue 1: [Model] ProblemName** -> (full draft) -> -> **Issue 2: [Rule] ProblemName to QUBO** -> (full draft) - -**For models**, the draft must include all template sections: -- Motivation -- Definition (Name, Reference, formal definition) -- Variables (Count, Per-variable domain, Meaning) -- Schema (Type name, Variants, Field table — use mathematical types, not programming types) -- Complexity (expression + citation + BibTeX) -- Extra Remark (if applicable) -- Reduction Rule Crossref (linking to companion rule issues or noting planned rules) -- How to solve (brute-force, direct ILP via companion rule if explicitly claimed, or other specialized solver — never claim ILP without a direct companion rule issue) -- Example Instance -- Expected Outcome - - Optimization problems: optimal solution + optimal objective value - - Satisfaction problems: valid / satisfying solution + brief justification -- BibTeX (include the BibTeX entry for the complexity/definition reference at the end of the issue) - -**For rules**, the draft must include: -- Source, Target, Motivation, Reference (with BibTeX) -- Reduction Algorithm (numbered steps, all symbols defined) -- Size Overhead (table with target metrics and formulas) -- Validation Method -- Example (fully worked: source instance, construction, target instance) -- BibTeX (include the BibTeX entry for the reference at the end of the issue) - ---- - -## Step 5: Run Check-Issue on Draft (BEFORE filing) - -**Critical: Run the check-issue logic on the draft BEFORE filing.** This catches problems early and avoids filing issues that will fail review. - -Apply the `/check-issue` quality checks against the draft content: - -### Rule draft checks -1. **Usefulness:** `pred path <source> <target>` — verify no existing path. If path exists, run redundancy analysis. -2. **Non-trivial:** Review the algorithm for genuine structural transformation (not just variable substitution or subtype coercion). -3. **Correctness:** Verify references exist (check `check-issue/references.md`, `docs/paper/references.bib`, then WebSearch). Cross-check claims. -4. **PDF/read-through:** Verify the main algorithm reference PDF was downloaded to `docs/research/raw/`, the relevant theorem/construction/proof was read, and the issue draft names the exact theorem/section used. If the issue only cites a paper without transcribing an implementable construction, this check fails. -5. **Well-written:** Verify all sections present, symbols consistent, overhead table field names match `pred show <target> --json` → `size_fields`, example is fully worked. - -### Model draft checks -1. **Usefulness:** `pred show <name>` must fail (problem doesn't exist). At least one reduction planned. -2. **Non-trivial:** Not isomorphic to existing problem. -3. **Correctness:** Complexity expression verified against literature. -4. **Well-written:** All template sections present, symbols consistent, example exercises core structure, and Expected Outcome matches the problem type (valid solution for satisfaction, optimal solution/value for optimization). -5. **ILP claim consistency:** If the draft claims direct ILP solvability, verify the Reduction Rule Crossref includes a direct `[Rule] <ProblemName> to ILP` companion issue. Claiming ILP without that concrete rule is a fail. - -**If any check fails:** Fix the draft automatically if possible. If user input is needed, ask. Loop back to Step 4 with the corrected draft. - -**If all checks pass:** Show the user a summary: "Draft passes all quality checks (Usefulness ✅, Non-trivial ✅, Correctness ✅, PDF/read-through ✅, Well-written ✅). Ready to file." - -Then present for approval via `AskUserQuestion`: - -``` -AskUserQuestion: - question: "The draft passes all quality checks. Ready to file?" - header: "Approval" - options: - - label: "File it" - description: "File the GitHub issue as-is" - - label: "Revise first" - description: "I have changes to suggest before filing" -``` - ---- - -## Step 6: File the Issue(s) - -Once the user approves, file all issues. For model + rule bundles, file the model issue first so rule issues can cross-reference it. - -```bash -# File model issue first -gh issue create \ - --title "[Model] ProblemName" \ - --label "model" \ - --body "$(cat <<'EOF' -<model issue body> -EOF -)" -``` - -Capture the model issue number, then file companion rule issues with cross-references: - -```bash -gh issue create \ - --title "[Rule] ProblemName to Target" \ - --label "rule" \ - --body "$(cat <<'EOF' -<rule issue body, referencing #model-issue-number> -EOF -)" -``` - -After filing all rule issues, update the model issue's "Reduction Rule Crossref" section with the actual issue numbers: - -```bash -# Update model issue body to replace placeholder with real issue numbers -gh issue edit <model-issue-number> --body "$(cat <<'EOF' -<updated body with real rule issue numbers> -EOF -)" -``` - -Print all issue URLs when done. - ---- - -## Key Principles - -- **Use `AskUserQuestion` only when genuine user input is needed** — use it for choices where the answer is NOT determinable from context (type detection, problem selection, example selection, approval). Do NOT use it when the answer is already clear from topology analysis, model inspection, or literature (e.g., don't ask "why is this useful?" when the topology analysis already shows it connects an orphan). -- **Auto-infer obvious answers** — When the user's orienting description clearly determines the answer (problem type, data representation, variable structure, motivation), confirm inline rather than presenting an open-ended `AskUserQuestion`. Expert users find obvious multiple-choice questions patronizing. -- **Study models before brainstorming** — always run `pred show <source> --json` and `pred show <target> --json` before asking questions. This reveals field types, size getters, and schema details that are essential for correct overhead tables. -- **Pre-fill well-known reductions** — if the reduction appears in standard textbooks, pre-fill answers from literature but still present each step to the user for confirmation. Never skip brainstorming steps. -- **Read the main reference, not just metadata** — for any rule whose algorithm comes from a paper, download the PDF to `docs/research/raw/`, read the actual construction/proof, and include implementation-level details in the issue. A citation-only algorithm section is not acceptable. -- **One question at a time** — don't overwhelm; each `AskUserQuestion` call has one focused question -- **Mathematical language only** — never mention Rust types, traits, macros, or code patterns to the user -- **Help find references** — use WebSearch to help locate papers, verify claims -- **Always provide a recommendation** — for every `AskUserQuestion` with multiple choices, analyze the problem context and mark one option as "(Recommended)" with a brief reason. Domain experts benefit from an informed default they can override. Base recommendations on the problem description, existing graph topology, and literature conventions. -- **Suggest, don't prescribe** — if the user is unsure about complexity or reductions, propose candidates and let them choose -- **Topology-driven suggestions** — run topology analysis first, then populate `AskUserQuestion` options with the most needed reductions ranked by value -- **Self-check before filing** — catch problems before they reach review -- **No implementation** — this skill produces issues, nothing else - -## Common Mistakes - -- **Don't ask questions with obvious answers.** If the topology analysis shows the rule connects an orphan, don't ask "What makes this reduction valuable?" — state it. If the user described "minimizing backward arcs," don't present a 3-option problem-type question — just confirm "This is a minimization problem — correct?" Only use full `AskUserQuestion` when the answer requires genuine user input or is ambiguous. -- **Don't skip model inspection.** Always run `pred show <source> --json` and `pred show <target> --json` before brainstorming. Missing this leads to wrong overhead tables and missed type mismatches (e.g., `BigUint` vs `i64`). -- **Don't skip confirmation for textbook reductions.** Even if SubsetSum → Knapsack is in Garey & Johnson, still present each brainstorming step with pre-filled answers for the user to confirm or revise. Never jump straight to the draft. -- **Don't rebuild `pred` unnecessarily.** Use `command -v pred` to check if it's installed before running `make cli` (which takes >1 minute). -- **Don't ask all questions at once.** One `AskUserQuestion` call per message. -- **Don't use programming jargon.** Say "list of weights" not "Vec<W>". Say "graph" not "SimpleGraph". Say "integer" not "i64". -- **Don't skip the reduction crossref.** An orphan model will be rejected. -- **Don't file without user approval.** Always show the draft first. -- **Don't implement anything.** The output is issues, not code. -- **Don't skip topology analysis for rules.** Always run topology analysis first, then populate `AskUserQuestion` options with the most needed reductions. +Keep draft files in the scratchpad, not the repo. Print all issue URLs at the end. diff --git a/.claude/skills/release/SKILL.md b/.claude/skills/release/SKILL.md index 8ae5ab65d..037ed5316 100644 --- a/.claude/skills/release/SKILL.md +++ b/.claude/skills/release/SKILL.md @@ -1,39 +1,39 @@ --- name: release -description: Use when preparing a new crate release, bumping versions, or tagging a release +description: Use when cutting a new problemreductions crate release — checks the repo is releasable, proposes the version bump with a changelog summary, and runs make release after explicit confirmation --- # Release -Guide for creating a new release of problemreductions. +`make release V=x.y.z` bumps the version in `Cargo.toml`, `problemreductions-macros/Cargo.toml`, +and `problemreductions-cli/Cargo.toml` (plus the inter-crate dependency versions), runs +`cargo check`, commits `release: vX.Y.Z`, tags `vX.Y.Z`, and pushes `main` and tags. CI then +publishes all three crates to crates.io. A pushed tag cannot be taken back cleanly, so every +gate below is hard. -## Step 1: Determine Version Bump - -Compare against the last release tag: +## Gates (refuse if any fails) ```bash -git tag -l 'v0.*' | sort -V # find latest tag -git log <last-tag>..HEAD --oneline # review commits -git diff <last-tag>..HEAD --stat # review scope +git branch --show-current # must be main +git status --porcelain # must be empty +git fetch origin main && git rev-parse HEAD origin/main # must be equal +make check # fmt-check + clippy + test must pass ``` -Apply semver for 0.x (pre-1.0): -- **Patch** (0.x.Y) — bug fixes, docs, CI only -- **Minor** (0.X.0) — new features, new reductions, new public API -- **Major** — reserved for post-1.0 +`make release` enforces the same branch/clean/up-to-date guard and runs `make check` itself. -## Step 2: Verify Clean State +## Version and changelog ```bash -make test clippy +last=$(git tag -l 'v0.*' | sort -V | tail -1) +git log "$last"..HEAD --oneline +git diff "$last"..HEAD --stat ``` -Both must pass with zero warnings before proceeding. - -## Step 3: Release - -```bash -make release V=x.y.z -``` +Pre-1.0 semver: **patch** (0.x.Y) for fixes, docs, CI; **minor** (0.X.0) for new models, rules, +CLI features, or any public API change, including breaking ones. -This target bumps versions in `Cargo.toml`, `problemreductions-macros/Cargo.toml`, and `problemreductions-cli/Cargo.toml`, runs `cargo check`, commits, tags, and pushes. CI publishes all three crates to crates.io. +Show the user the proposed version, the reason, and a grouped changelog summary (new models, +new rules, CLI, fixes, breaking changes). Run `make release V=x.y.z` only after the user +explicitly confirms that exact version. Afterwards report the tag and point to the release CI +run (`gh run list --workflow release.yml --limit 1`). diff --git a/.claude/skills/review-paper/SKILL.md b/.claude/skills/review-paper/SKILL.md deleted file mode 100644 index e5a8e2378..000000000 --- a/.claude/skills/review-paper/SKILL.md +++ /dev/null @@ -1,142 +0,0 @@ ---- -name: review-paper -description: Review the Typst paper (docs/paper/reductions.typ) for quality issues — evaluates 10 entries per session, reports mechanical and critical issues without fixing ---- - -# Review Paper - -Evaluate the quality of problem definitions and reduction rules in `docs/paper/reductions.typ`. Each session reviews **10 entries** (problems or rules), producing a structured report. **Read-only — do not modify any files.** - -## Usage - -``` -/review-paper # review next 10 unreviewed problem-defs -/review-paper rules # review next 10 unreviewed reduction-rules -/review-paper ProblemName # review a specific problem-def -/review-paper Source Target # review a specific reduction-rule -``` - -## Step 0: Determine Scope - -Parse the argument: -- No argument or `problems` → review problem-defs -- `rules` → review reduction-rules -- A specific name → review that single entry - -To pick which 10 to review, scan `docs/paper/reductions.typ` for all `problem-def(...)` or `reduction-rule(...)` entries. Start from the beginning of the file, skipping any that have been reviewed in a previous session (check memory for `paper-review-progress`). If all have been reviewed, report completion. - -## Step 1: Load Gold Standard - -Read the reference examples before reviewing: -- **Problem gold standard:** search for `problem-def("MaximumIndependentSet")` in `reductions.typ` — note its structure, depth, and components -- **Rule gold standard:** search for `reduction-rule("MaximumIndependentSet", "MinimumVertexCover"` — note its proof depth and example - -## Step 2: Review Each Entry - -For each of the 10 entries, read the full entry text and evaluate against the checklists below. - -### Problem-Def Checklist - -**Mechanical checks** (objective, can be verified by reading): - -| Check | Criterion | -|-------|-----------| -| M1. Display name | Entry exists in `display-name` dictionary | -| M2. Formal definition | `def` parameter is present and non-empty | -| M3. Self-contained notation | Every symbol in `def` is defined before first use | -| M4. Background text | Body contains at least 2 sentences of background/motivation | -| M5. Example present | Body contains `*Example.*` or `Example.` | -| M6. Example from fixture | Example data matches `docs/paper/data/examples.json` (not invented) — check by loading the JSON and comparing | -| M7. Figure present | Body contains `#figure(` | -| M8. Pred commands | Body contains `pred-commands(` or `pred create` | -| M9. Algorithm citation | Complexity claims have `@citation` or a footnote explaining absence | -| M10. Evaluation shown | Example shows how the objective/verifier computes the value | - -**Critical checks** (require judgment): - -| Check | Criterion | -|-------|-----------| -| C1. Definition correctness | Does the formal definition accurately describe the problem? Compare with the Rust implementation (`src/models/`) and literature | -| C2. Background quality | Is the background informative? Does it mention applications, history, special cases, or algorithmic context? | -| C3. Example pedagogy | Is the example small enough to verify by hand? Does it illustrate the key aspects of the problem? | -| C4. Completeness | Are there important aspects of the problem that are missing (e.g., well-known special cases, relationship to other problems)? | - -### Reduction-Rule Checklist - -**Mechanical checks:** - -| Check | Criterion | -|-------|-----------| -| M1. Theorem statement | Rule body describes the construction | -| M2. Proof present | Proof body is non-empty | -| M3. Proof length | Proof is at least 3 sentences (not just "trivial" or a one-liner) | -| M4. Overhead documented | Overhead is auto-generated from JSON (verify edge exists in `reduction_graph.json`) | -| M5. Example present | `example: true` and example renders correctly | -| M6. Example from fixture | Example data matches `docs/paper/data/examples.json` | -| M7. Pred commands | Example section contains `pred-commands(` with create/reduce/evaluate pipeline | -| M8. Both directions | If the reverse rule also exists in the graph, check it has its own entry | - -**Critical checks:** - -| Check | Criterion | -|-------|-----------| -| C1. Construction correctness | Does the theorem statement accurately describe what `reduce_to()` does? Read `src/rules/<source>_<target>.rs` to verify | -| C2. Proof correctness | Does the proof correctly argue that the reduction preserves solutions? | -| C3. Example clarity | Does the example clearly show source → target → solution extraction? | -| C4. Proof-only flag | If this is a proof-only reduction (not solver-executable), is that stated? | - -## Step 3: Generate Report - -Present results **one entry at a time** in this format: - -``` -### [N/10] ProblemName (or Source → Target) - -**Mechanical Issues:** -- [PASS] M1. Display name -- [FAIL] M5. Example present — no worked example in body -- [WARN] M9. Algorithm citation — complexity claim "O*(2^n)" has no @citation - -**Critical Issues:** -- [FAIL] C2. Background quality — body is only one sentence ("This is NP-hard.") - with no applications, history, or algorithmic context -- [OK] C1. Definition correctness — matches Rust implementation - -**Verdict:** 2 mechanical fails, 1 critical fail — needs improvement -``` - -After each entry, pause and ask: **"Continue to next entry, or discuss this one?"** - -Use these severity levels: -- **PASS** — meets criterion -- **WARN** — minor issue, could be improved but acceptable -- **FAIL** — does not meet criterion, should be fixed - -## Step 4: Session Summary - -After all 10 entries, print a summary table: - -``` -## Session Summary - -| Entry | Mechanical | Critical | Verdict | -|-------|-----------|----------|---------| -| ProblemA | 9/10 pass | 4/4 pass | Good | -| ProblemB | 7/10 pass | 3/4 pass | Needs work | -| ... | ... | ... | ... | - -Overall: X/10 entries pass all checks. -Top priorities for improvement: [list the 3 worst entries] -``` - -## Step 5: Save Progress - -Save progress to memory so the next session can continue where this one left off. Record which entries have been reviewed and their verdicts. - -## Important Rules - -1. **Do not modify any files.** This skill is read-only. -2. **Do not invent issues.** Only report problems you can verify by reading the source. -3. **Check the Rust source** for critical checks — don't guess whether the math is right. -4. **Be specific.** "Background is thin" is not useful. "Background is one sentence with no applications or algorithmic context" is useful. -5. **Compare to gold standard.** The MIS entry is the reference — entries don't need to be as long, but they should cover the same structural elements. diff --git a/.claude/skills/review-pipeline/SKILL.md b/.claude/skills/review-pipeline/SKILL.md deleted file mode 100644 index 51930ca79..000000000 --- a/.claude/skills/review-pipeline/SKILL.md +++ /dev/null @@ -1,266 +0,0 @@ ---- -name: review-pipeline -description: Agentic review for PRs in the Review pool — runs structural, quality, and agentic-test sub-reviews (no code changes), posts combined verdict, moves to Final review ---- - -# Review Pipeline - -Pick PRs from the `Review pool` column on the [GitHub Project board](https://github.com/orgs/CodingThrust/projects/8/views/1). For each PR: claim it into `Under review`, run three read-only sub-reviews in parallel (structural check, quality check, agentic feature tests), post a combined verdict as a PR comment, then move to `Final review`. - -**This skill does NOT modify the PR.** No commits, no pushes, no merging main. It only evaluates and reports. - -## Invocation - -- `/review-pipeline` -- pick the next Review pool item -- `/review-pipeline 570` -- process a specific PR number - -For Codex, open this `SKILL.md` directly and treat the slash-command forms above as aliases. The Makefile `run-review` target already does this translation. - -## Constants - -GitHub Project board IDs (for `gh project item-edit`): - -| Constant | Value | -|----------|-------| -| `PROJECT_ID` | `PVT_kwDOBrtarc4BRNVy` | -| `STATUS_FIELD_ID` | `PVTSSF_lADOBrtarc4BRNVyzg_GmQc` | -| `STATUS_REVIEW_POOL` | `7082ed60` | -| `STATUS_UNDER_REVIEW` | `f04790ca` | -| `STATUS_FINAL_REVIEW` | `51a3d8bb` | - -## Prerequisites - -- **agentic-tests** must be installed (`~/.claude/commands/agentic-tests:test-feature.md` must exist). If missing, STOP with: `agentic-tests not installed. Run: gh clone GiggleLiu/agentic-tests ~/.claude/agentic-tests && mkdir -p ~/.claude/commands && ln -s ~/.claude/agentic-tests/skills/test-feature/SKILL.md ~/.claude/commands/agentic-tests:test-feature.md` - -## Autonomous Mode - -This skill runs **fully autonomously** except for one case: if a Review pool card links multiple repo PRs and the intended target is unclear, STOP and ask the user which PR is the intended target. - -## Steps - -### 0a. Triage Review Pool - -Before spending the expensive full-context packet, do one lightweight Review-pool scan: - -```bash -REPO=$(gh repo view --json nameWithOwner --jq .nameWithOwner) -QUEUE=$(python3 scripts/pipeline_board.py list review --repo "$REPO" --format json) -``` - -If `PR` was explicitly supplied (for example `/review-pipeline 570`), do **not** pick a different item from the queue. Find that PR in the `QUEUE` JSON output to get its `ITEM_ID` and confirm it is in Review pool. - -`pipeline_board.py` only supports these subcommands: `next`, `claim-next`, `ack`, `list`, `move`, `backlog`. To look up a specific PR's board status, use `list` and filter the JSON output. - -Pick one candidate with a lightweight heuristic: -- prefer direct PR cards over issue cards -- any open PR in Review pool is eligible -- if multiple candidates are tied, pick one at random (e.g., use the current minute mod candidate count) to avoid always picking the same item on retries - -**Review-ready criteria** — a PR is ready for review if all of these hold: -- PR state is `OPEN` (not draft, not closed) -- The diff contains at least one model file (`src/models/`) or rule file (`src/rules/`) or other substantive code -- A test file exists for the new code -- The PR body does not say "WIP" or "DO NOT REVIEW" - -If the PR is not review-ready, post a diagnostic comment and move to Final review for human triage: - -```bash -gh pr comment <PR_NUMBER> --body "review-pipeline: PR not review-ready. <brief concrete reason>. Skipping full review, moving to Final review for human triage." -python3 scripts/pipeline_board.py move <ITEM_ID> final-review -``` - -For untargeted runs, then return to Step 0a to pick another item. -For explicit `PR` runs, STOP after reporting. - -If no candidate is both open and ready for review, STOP with `No Review pool PRs are currently ready for review-pipeline processing.` - -### 0b. Generate Review-Pipeline Report, Create Worktree, Generate Implementation Report - -Only after Step 0a has identified a review-ready PR should you spend the expensive context packets. - -**Generate review-pipeline context** (from the repo root, before entering the worktree — this queries GitHub APIs only): - -```bash -REPO_ROOT=$(pwd) - -# 1. Review-pipeline context (selection, comments, CI, linked issue) -set -- python3 scripts/pipeline_skill_context.py review-pipeline --repo "$REPO" --pr "$PR" --format text -REPORT=$("$@") -printf '%s\n' "$REPORT" -``` - -The review-pipeline report should already include: -- Selection: board item, PR number, linked issue, title, URL -- Recommendation Seed: suggested mode and deterministic blockers -- Comment Summary -- CI / Coverage -- PR head branch -- Linked Issue Context - -**Create worktree and check out the PR branch:** - -```bash -WORKTREE_JSON=$(python3 scripts/pipeline_worktree.py enter --name "review-pr-$PR" --format json) -WORKTREE_DIR=$(printf '%s\n' "$WORKTREE_JSON" | python3 -c "import sys,json; print(json.load(sys.stdin)['worktree_dir'])") -cd "$WORKTREE_DIR" -gh pr checkout "$PR" -``` - -**Generate review-implementation context** (inside the worktree — needs git diff against main): - -```bash -# 2. Review-implementation context (scope, checklists, diff) -IMPL_REPORT=$(python3 scripts/pipeline_skill_context.py review-implementation --repo-root . --format text) -printf '%s\n' "$IMPL_REPORT" -``` - -The review-implementation report should already include: -- Review Range: base SHA, head SHA -- Scope: review type (model/rule/generic), subject metadata -- Deterministic Checks: whitelist + completeness status -- Changed Files and Diff Stat - -The two expensive context calls are allowed exactly once each per top-level `review-pipeline` invocation. Both reports are reused for the rest of the skill — do not regenerate either. - -Branch from the review-pipeline report: -- `Bundle status: empty` => the selected PR is no longer eligible; run `cd "$REPO_ROOT" && python3 scripts/pipeline_worktree.py cleanup --worktree "$WORKTREE_DIR"`, then for untargeted runs return to Step 0a, for explicit `PR` runs STOP -- `Bundle status: needs-user-choice` => run `cd "$REPO_ROOT" && python3 scripts/pipeline_worktree.py cleanup --worktree "$WORKTREE_DIR"`, STOP and ask the user which PR is intended -- `Bundle status: ready` => claim the item and continue - -**Claim the item** (move to Under review) only after confirming `Bundle status: ready`: - -```bash -python3 scripts/pipeline_board.py move <ITEM_ID> under-review -``` - -Use the identifiers from the report for all subsequent operations. All subsequent steps run inside the worktree and should read facts from the reports instead of re-fetching them. - -### 1. Run Three Sub-Reviews (Parallel) - -Run three independent sub-reviews. All three are **read-only** — they evaluate the PR but do NOT commit or push anything. Dispatch them as parallel subagents where possible. - -**Pass `IMPL_REPORT` to both structural and quality subagents** so they skip their own context generation step. Include the full text of `IMPL_REPORT` in each subagent prompt with a prefix like: - -> The review-implementation context has already been generated. Use this report instead of running `pipeline_skill_context.py review-implementation` yourself: -> -> ``` -> <IMPL_REPORT content> -> ``` - -#### 1a. Structural Check (project-specific) - -Invoke `/review-structural` (file: `.claude/skills/review-structural/SKILL.md`) with the pre-generated `IMPL_REPORT`. This runs the model/rule checklists, build checks, semantic review, and issue compliance checks. - -**Mathematical correctness is critical.** In addition to the standard structural checks, verify: -- **For rules**: Is the reduction mathematically correct? Trace through the `reduce_to()` logic with a small example and confirm the target instance encodes the same problem. Check that `extract_solution` correctly inverts the mapping. Verify the paper proof sketch is sound — not just present, but logically valid. -- **For models**: Does `evaluate()` correctly compute the objective for the mathematical definition? Are edge cases handled (empty graph, zero weights, infeasible configs)? -- **Overhead expressions**: Manually count the sizes in `reduce_to()` output and verify they match the `overhead = { ... }` formulas. - -**Do NOT auto-fix anything.** Collect the output report for Step 2. - -#### 1b. Quality Check (generic) - -Invoke `/review-quality` (file: `.claude/skills/review-quality/SKILL.md`) with the pre-generated `IMPL_REPORT`. This runs DRY/KISS/HC-LC checks, test quality review, and HCI checks (if CLI changed). - -**Do NOT auto-fix anything.** Collect the output report for Step 2. - -#### 1c. Agentic Feature Tests - -**This step is mandatory — do NOT skip.** - -1. **Identify the feature** from the PR title and changed files: - - `[Model]` PRs: the new problem model name - - `[Rule]` PRs: the new reduction rule (source -> target) - -2. **Invoke `/agentic-tests:test-feature`** (file: `~/.claude/commands/agentic-tests:test-feature.md`) with the identified feature. This simulates a downstream user exercising the feature from docs and examples. - - **Minimum test checklist** for the agentic tester: - - For models, `pred list <Name>`; for rules, `pred list --rules <Source>` — verify the new catalog entry appears - - `pred show <Name>` — verify details display correctly - - `pred create --example <Name>` — verify example instance creation works - - `pred solve <instance>` — verify solving works on the example - - For rules: `pred reduce <source-instance>` — verify reduction produces valid target - -3. **Collect the test report.** For each issue found: - - Reproduce it from the current PR worktree to confirm it's real - - Classify as: `confirmed` / `not reproducible in current worktree` - - For confirmed issues, note severity and recommended fix - -**Do NOT fix any issues.** Only report them. When dispatching the agentic-test subagent, explicitly instruct it: "This is a read-only review run. Do NOT offer to fix issues, do NOT select option (a) 'Review together and fix', and do NOT modify any files. Report findings only and stop after generating the report." - -### 2. Compose Combined Review Comment - -Merge the results from all three sub-reviews into one structured PR comment. - -Paste the **structured report section** from each subagent — the formatted output (checklist tables, issue lists, test results), not raw transcripts or internal reasoning. The human in final-review reads these reports to make merge/hold decisions, so every finding matters. - -If the report's `Merge Prep` section indicates merge conflicts with main, include a note at the top: - -> **Note:** This PR has merge conflicts with `main`. These must be resolved before merging (handled in final-review Step 1). - -```bash -COMMENT_FILE=$(mktemp) -cat > "$COMMENT_FILE" <<'EOF' -## Agentic Review Report - -### Structural Check - -[Paste structured report from `/review-structural` here — full checklist table, build status, semantic review, issue compliance. Do not include internal reasoning.] - ---- - -### Quality Check - -[Paste structured report from `/review-quality` here — design principles review, HCI (if applicable), test quality, all issues with severity and file:line references.] - ---- - -### Agentic Feature Tests - -[Paste structured report from `/agentic-tests:test-feature` here — test results with all findings, reproduction results, and classifications.] - ---- - -Generated by review-pipeline -EOF -python3 scripts/pipeline_pr.py comment --repo "$REPO" --pr "$PR" --body-file "$COMMENT_FILE" -rm -f "$COMMENT_FILE" -``` - -The review stage does not judge pass/fail — it reports findings. The human in final-review decides. - -### 3. Move PR to Final Review - -Always move to Final review — the human decides what to do with the findings: - -```bash -python3 scripts/pipeline_board.py move <ITEM_ID> final-review -``` - -### 4. Clean Up Worktree - -```bash -cd "$REPO_ROOT" -python3 scripts/pipeline_worktree.py cleanup --worktree "$WORKTREE_DIR" -``` - -## Common Mistakes - -| Mistake | Fix | -|---------|-----| -| PR not in Review pool column | Verify status before processing; STOP if not Review pool | -| Processing a closed PR from a stale issue card | Require PR state `OPEN`; skip stale closed PRs | -| Guessing on an issue card with multiple linked repo PRs | Stop, show options to the user, and recommend the most likely correct OPEN PR | -| Committing or pushing changes | This skill is read-only — evaluate and report only, never modify the PR | -| Moving items backward to Ready | Never move backward — always forward to Final review | -| Missing project scopes | Run `gh auth refresh -s read:project,project` | -| Skipping structural check | Always run `/review-structural` — it catches gaps in paper entries, registrations, tests | -| Skipping agentic tests | Always run `/agentic-tests:test-feature` even if CI is green | -| Not checking out the right branch | Use `gh pr checkout <PR_NUMBER>` after `pipeline_worktree.py enter` | -| Worktree left behind on failure | Always run `pipeline_worktree.py cleanup` in Step 4 | -| Working in main checkout | All work happens in the worktree — never modify the main checkout | -| Fixing issues instead of reporting them | The review stage judges, it does not fix. Report findings for human/final-review to act on | -| Pasting raw agent transcripts | Paste the structured report sections only — checklist tables, issue lists, test results — not internal reasoning or scratch work | -| Regenerating context in subagents | Pass `IMPL_REPORT` to structural/quality subagents so they skip `pipeline_skill_context.py review-implementation` | -| Always picking the same PR on retry | Use randomized tie-breaking when multiple candidates are eligible | -| Inventing `pipeline_board.py` subcommands | Only `next`, `claim-next`, `ack`, `list`, `move`, `backlog` exist. Use `list` to look up a PR's board status | diff --git a/.claude/skills/review-quality/SKILL.md b/.claude/skills/review-quality/SKILL.md deleted file mode 100644 index 7c43c8241..000000000 --- a/.claude/skills/review-quality/SKILL.md +++ /dev/null @@ -1,112 +0,0 @@ ---- -name: review-quality -description: Generic code quality review — evaluates DRY, KISS, cohesion/coupling, test quality, and HCI. Read-only, no code changes. ---- - -# Quality Review - -Generic code quality review that applies to any code change. Evaluates design principles, test quality, and (if applicable) CLI/HCI quality. - -**This skill is read-only.** It evaluates and reports — it does NOT fix, commit, or push anything. - -## Invocation - -- `/review-quality` -- auto-detect from git diff - -Called by `review-pipeline` as one of three parallel sub-reviews. - -## Step 1: Get Context - -**If the caller (e.g., `review-pipeline`) already provided a pre-generated review-implementation report in the prompt, use that directly and skip the generation command below.** - -Otherwise, generate the context yourself: - -```bash -REPORT=$(python3 scripts/pipeline_skill_context.py review-implementation --repo-root . --format text) -printf '%s\n' "$REPORT" -``` - -Extract from the report: -- `Review Range`: base SHA and head SHA -- `Changed Files` and `Diff Stat` -- `Linked Issue Context` - -## Step 2: Read the Diff - -```bash -git diff --stat {BASE_SHA}..{HEAD_SHA} -git diff {BASE_SHA}..{HEAD_SHA} -``` - -Then read all changed files in full. - -## Step 3: Evaluate Design Principles - -### DRY (Don't Repeat Yourself) -Is there duplicated logic that should be extracted into a shared helper? Check for copy-pasted code blocks across files (similar graph construction, weight handling, or solution extraction patterns). - -### KISS (Keep It Simple, Stupid) -Is the implementation unnecessarily complex? Look for: over-engineered abstractions, convoluted control flow, premature generalization, layers of indirection that add no value. - -### High Cohesion, Low Coupling (HC/LC) -Does each module/function/struct have a single, well-defined responsibility? -- **Low cohesion**: Function doing unrelated things -- **High coupling**: Modules depending on each other's internals -- **Mixed concerns**: A single file containing both problem logic and CLI/serialization logic -- **God functions**: Functions longer than ~50 lines doing multiple conceptually distinct things - -## Step 4: Evaluate HCI (if CLI/MCP files changed) - -Only check these if the diff touches `problemreductions-cli/`: - -- **Error messages** — Are they actionable? Bad: `"invalid parameter"`. Good: `"KColoring requires --k <value> (e.g., --k 3)"`. -- **Discoverability** — Missing `--help` examples? Undocumented flags? Silent failures that should suggest alternatives? -- **Consistency** — Similar operations expressed similarly? Parameter names, output formats, error styles uniform? -- **Least surprise** — Output matches expectations? No contradictory output or silent data loss? -- **Feedback** — Tool confirms what it did? Echoes interpreted parameters for ambiguous operations? - -## Step 5: Evaluate Test Quality - -Flag tests that: -- **Only check types/shapes, not values**: e.g., `assert!(result.is_some())` without checking the solution is correct -- **Mirror the implementation**: Tests recomputing the same formula as the code prove nothing -- **Lack adversarial cases**: Only happy path. Tests must include infeasible configs and boundary cases -- **Use trivial instances only**: Single-edge or 2-node tests may pass with bugs. Need 5+ vertex instances -- **Closed-loop without verification**: Must verify extracted solution is **optimal** (compare brute-force on both source and target) -- **Assert count too low**: 1-2 asserts for non-trivial code is insufficient - -## Output Format - -``` -## Quality Review - -### Design Principles -- DRY: OK / ISSUE — [description with file:line] -- KISS: OK / ISSUE — [description with file:line] -- HC/LC: OK / ISSUE — [description with file:line] - -### HCI (if CLI/MCP changed) -- Error messages: OK / ISSUE — [description] -- Discoverability: OK / ISSUE — [description] -- Consistency: OK / ISSUE — [description] -- Least surprise: OK / ISSUE — [description] -- Feedback: OK / ISSUE — [description] - -### Test Quality -- Naive test detection: OK / ISSUE - - [specific tests flagged with reason and file:line] - -### Issues - -#### Critical (Must Fix) -[Bugs, correctness issues, data loss risks] - -#### Important (Should Fix) -[Architecture problems, missing tests, poor error handling] - -#### Minor (Nice to Have) -[Code style, optimization opportunities] - -### Summary -- [list of all ISSUE items as bullet points with severity] -``` diff --git a/.claude/skills/review-structural/SKILL.md b/.claude/skills/review-structural/SKILL.md deleted file mode 100644 index d291a7e02..000000000 --- a/.claude/skills/review-structural/SKILL.md +++ /dev/null @@ -1,179 +0,0 @@ ---- -name: review-structural -description: Project-specific structural completeness check for a PR — verifies model/rule checklists, build, semantic correctness, issue compliance. Read-only, no code changes. ---- - -# Structural Review - -Project-specific structural completeness check. Verifies that a model or rule implementation has all required components, passes build checks, and matches the linked issue specification. - -**This skill is read-only.** It evaluates and reports — it does NOT fix, commit, or push anything. - -## Invocation - -- `/review-structural` -- auto-detect from git diff -- `/review-structural model MaximumClique` -- review a specific model -- `/review-structural rule mis_qubo` -- review a specific rule - -Called by `review-pipeline` as one of three parallel sub-reviews. - -## Step 1: Get Context - -**If the caller (e.g., `review-pipeline`) already provided a pre-generated review-implementation report in the prompt, use that directly and skip the generation command below.** - -Otherwise, generate the context yourself: - -```bash -set -- python3 scripts/pipeline_skill_context.py review-implementation --repo-root . --format text - -# Explicit subject overrides: -# set -- "$@" --kind model --name MaximumClique -# set -- "$@" --kind rule --name mis_qubo --source MaximumIndependentSet --target QUBO - -REPORT=$("$@") -printf '%s\n' "$REPORT" -``` - -Extract from the report: -- `Scope`: review type (model/rule/generic), problem name, category, file stem -- `Deterministic Checks`: whitelist + completeness status -- `Linked Issue Context`: issue requirements to check against -- `Changed Files` and `Diff Stat` - -If review type is `generic` (no new model/rule detected), report "No structural review needed for generic changes" and stop. - -## Step 2: Run Structural Checklist - -### Model Checklist - -Only run if review type includes "model". Given: problem name `P`, category `C`, file stem `F`. - -| # | Check | How to verify | -|---|-------|--------------| -| 1 | Model file exists | `Glob("src/models/{C}/{F}.rs")` | -| 2 | `inventory::submit!` present | `Grep("inventory::submit", file)` | -| 3 | `#[derive(...Serialize, Deserialize)]` on struct | `Grep("Serialize.*Deserialize", file)` | -| 4 | `Problem` trait impl | `Grep("impl.*Problem for.*{P}", file)` | -| 5 | Aggregate value is present | `Grep("type Value =", file)` | -| 6 | `#[cfg(test)]` + `#[path = "..."]` test link | `Grep("#\\[path =", file)` | -| 7 | Test file exists | `Glob("src/unit_tests/models/{C}/{F}.rs")` | -| 8 | Test file has >= 3 test functions | `Grep("fn test_", test_file)` — count matches, FAIL if < 3 | -| 9 | Registered in `{C}/mod.rs` | `Grep("mod {F}", "src/models/{C}/mod.rs")` | -| 10 | Re-exported in `models/mod.rs` | `Grep("{P}", "src/models/mod.rs")` | -| 11 | Variant registration exists | `Grep("declare_variants!|VariantEntry", file)` | -| 12 | Alias registration | If aliases are claimed, verify problem aliases are in `ProblemSchemaEntry.aliases` and variant aliases are in `declare_variants!`; no frontend alias branch | -| 13 | CLI `create` support | Run `pred create <problem-spec> --help` for the concrete variant. Verify its flags and types come from the registered construction inputs (`ProblemSchemaEntry.fields` or the model-local `CreateSpec`), with a reusable codec for any unusual transport syntax. | -| 14 | Canonical model example registered | `Grep("{P}", "src/example_db/model_builders.rs")` | -| 15 | Paper `display-name` entry | `Grep('"{P}"', "docs/paper/reductions.typ")` | -| 16 | Paper `problem-def` block | `Grep('problem-def.*"{P}"', "docs/paper/reductions.typ")` | -| 17 | Numeric and error contracts | Derive the expected boundary representation from the mathematical definition, then compare schema types, Rust fields, aggregate/total type, constructor and serde validation, conversions, overflow behavior, and boundary tests against `docs/src/design.md#numeric-types-and-arithmetic`. Verify construction paths return `ConstructionError`, `evaluate()` returns `EvaluationError`, and no public model path returns `Result<_, String>`. | - -### Rule Checklist - -Only run if review type includes "rule". Given: source `S`, target `T`, rule file stem `R`, example stem `E`. - -| # | Check | How to verify | -|---|-------|--------------| -| 1 | Rule file exists | `Glob("src/rules/{R}.rs")` | -| 2 | `#[reduction(...)]` macro present | `Grep("#\\[reduction", file)` | -| 3 | `ReductionResult` impl present | `Grep("impl.*ReductionResult", file)` | -| 4 | `ReduceTo` impl present | `Grep("impl.*ReduceTo", file)` | -| 5 | `#[cfg(test)]` + `#[path = "..."]` test link | `Grep("#\\[path =", file)` | -| 6 | Test file exists | `Glob("src/unit_tests/rules/{R}.rs")` | -| 7 | Closed-loop test present | `Grep("fn test_.*closed_loop\|fn test_.*to_.*basic", test_file)` | -| 8 | Registered in `rules/mod.rs` | `Grep("mod {R}", "src/rules/mod.rs")` | -| 9 | Canonical rule example registered | `Grep("canonical_rule_example_specs", rule file)` and verify it is included by `src/rules/mod.rs` | -| 10 | Example-db lookup tests exist | `Grep("find_rule_example|build_rule_db", "src/unit_tests/example_db.rs")` | -| 11 | Paper `reduction-rule` entry | `Grep('reduction-rule.*"{S}".*"{T}"', "docs/paper/reductions.typ")` | -| 12 | Extraction contract | Direct decoders call `validate_target_solution()`, enforce rule-specific structure, and test malformed cases; the helper does not establish feasibility or optimality. Composed extractors may delegate. | -| 13 | Numeric and error contracts | Compare source/target boundary types, size arithmetic, coefficients, bounds, auxiliary IDs, conversions, overflow behavior, and boundary tests against `docs/src/design.md#numeric-types-and-arithmetic`. Verify public reduction paths return `ReductionError`, preserve target `ConstructionError` as its construction cause, and never stringify or silently handle either failure. | - -## Step 2b: Blacklisted File Check - -Scan the PR's changed files for auto-generated files that must never be committed: -- `docs/src/reductions/reduction_graph.json` -- `docs/src/reductions/problem_schemas.json` -- `docs/paper/data/examples.json` (current output path, gitignored) - -If any of these files appear in the diff, report **FAIL — blacklisted auto-generated file committed**. These files are rebuilt by CI/`make doc`/`make paper` and must not be in PRs. - -## Step 3: Build Check - -Run: -```bash -make test clippy -``` - -Report pass/fail. If tests fail, identify which tests. **Do NOT fix anything** — just report. - -## Step 4: Semantic Review - -### For Models: -1. **`evaluate()` correctness** — Does it check feasibility before computing the objective when the model has invalid configurations? Objective models should return `Max/Min/Extremum(None)` for infeasible configs, witness problems should return `false`, and aggregate-only models should return the per-configuration contribution that matches the intended fold semantics. -2. **`dims()` correctness** — Does it return the actual configuration space? (e.g., `vec![2; n]` for binary) -3. **Size getter consistency** — Do inherent getter methods (e.g., `num_vertices()`, `num_edges()`) match names used in overhead expressions? -4. **Weight handling** — Are weights managed via inherent methods, not traits? -5. **Numeric safety** — Are element and total types distinct where required, do serde and constructors enforce the same range, and are overflow and non-finite values rejected explicitly? - -### For Rules: -1. **`extract_solution` correctness** — Does it implement the mathematical inverse? Is every branch either a defined mathematical case or an `ExtractionError`, with no defaulting, truncation, clamping, panic, or recovery? -2. **Overhead accuracy** — Does `overhead = { field = "expr" }` reflect the actual size relationship? -3. **Example quality** — Is it tutorial-style? Does the JSON export include both source and target data? -4. **Paper quality** — Is the reduction-rule statement precise? Is the proof sketch sound? -5. **Numeric safety** — Are target sizes and auxiliary IDs checked before construction, with no unchecked narrowing or exact-to-`f64` shortcut? - -## Step 5: Issue Compliance Review - -Only if a linked issue was provided. - -### For Models (check against issue): -| # | Check | -|---|-------| -| 1 | Problem name matches issue | -| 2 | Mathematical definition matches | -| 3 | Problem framing (objective / witness / aggregate-only) matches | -| 4 | Type parameters match | -| 5 | Configuration space matches | -| 6 | Feasibility check matches | -| 7 | Objective function matches | -| 8 | Complexity matches | - -### For Rules (check against issue): -| # | Check | -|---|-------| -| 1 | Source/target match issue | -| 2 | Reduction algorithm matches | -| 3 | Solution extraction matches | -| 4 | Correctness preserved | -| 5 | Overhead expressions match | -| 6 | Example matches | - -Flag any deviation as ISSUE. - -## Output Format - -``` -## Structural Review: [model/rule] [Name] - -### Structural Completeness -| # | Check | Status | -|---|-------|--------| -| 1 | ... | PASS / FAIL — reason | - -### Build Status -- `make test`: PASS / FAIL -- `make clippy`: PASS / FAIL - -### Semantic Review -- [check]: OK / ISSUE — description - -### Issue Compliance (if linked issue found) -| # | Check | Status | -|---|-------|--------| -| 1 | ... | OK / ISSUE — deviation description | - -### Summary -- X/Y structural checks passed -- X/Y issue compliance checks passed (if applicable) -- [list of all FAIL/ISSUE items as bullet points] -``` diff --git a/.claude/skills/run-pipeline/SKILL.md b/.claude/skills/run-pipeline/SKILL.md deleted file mode 100644 index 731b7e550..000000000 --- a/.claude/skills/run-pipeline/SKILL.md +++ /dev/null @@ -1,218 +0,0 @@ ---- -name: run-pipeline -description: Pick a Ready issue from the GitHub Project board, move it from In Progress through issue-to-pr into Review pool ---- - -# Run Pipeline - -Pick a "Ready" issue from the [GitHub Project board](https://github.com/orgs/CodingThrust/projects/8/views/1), claim it into "In Progress", run `issue-to-pr`, then move it to "Review pool". The separate `review-pipeline` handles agentic review (structural check, quality check, agentic feature tests). - -## Invocation - -- `/run-pipeline` -- pick the highest-ranked Ready issue (ranked by importance, relatedness, pending rules) -- `/run-pipeline 97` -- process a specific issue number from the Ready column - -For Codex, open this `SKILL.md` directly and treat the slash-command forms above as aliases. The Makefile `run-pipeline` target already does this translation. - -## Constants - -GitHub Project board IDs (for `gh project item-edit`): - -| Constant | Value | -|----------|-------| -| `PROJECT_ID` | `PVT_kwDOBrtarc4BRNVy` | -| `STATUS_FIELD_ID` | `PVTSSF_lADOBrtarc4BRNVyzg_GmQc` | -| `STATUS_READY` | `f37d0d80` | -| `STATUS_IN_PROGRESS` | `a12cfc9c` | -| `STATUS_REVIEW_POOL` | `7082ed60` | -| `STATUS_UNDER_REVIEW` | `f04790ca` | -| `STATUS_FINAL_REVIEW` | `51a3d8bb` | -| `STATUS_DONE` | `6aca54fa` | - -## Autonomous Mode - -This skill runs **fully autonomously** — no confirmation prompts, no user questions. It picks the next issue and processes it end-to-end. All sub-skills (`issue-to-pr`, `check-issue`, `add-model`, `add-rule`, etc.) should also auto-approve any confirmation prompts. - -## Steps - -### 0. Generate the Project-Pipeline Report - -Step 0 should be a single report-generation step. Do not manually list Ready items, list In-progress items, grep model declarations, or re-derive blocked rules with separate shell commands. -The expensive full-context call here is `python3 scripts/pipeline_skill_context.py project-pipeline ...` (backed by `build_project_pipeline_context()`). For a single top-level `run-pipeline` invocation, call it once and reuse the packet for scoring, ranking, and choosing the issue. Do not rerun it in the single-issue path after the packet exists. - -```bash -set -- python3 scripts/pipeline_skill_context.py project-pipeline --repo CodingThrust/problem-reductions --repo-root . --format text - -# If a specific issue number was provided, validate it through the same bundle: -# set -- "$@" --issue <number> - -REPORT=$("$@") -printf '%s\n' "$REPORT" -``` - -The report is the Step 0 packet. It should already include: -- Queue Summary -- Eligible Ready Issues -- Blocked Ready Issues -- In Progress Issues -- Requested Issue validation when a specific issue was supplied - -Branch from the report: -- `Bundle status: empty` => STOP with `No Ready issues are currently available.` -- `Bundle status: no-eligible-issues` => STOP with `Ready issues exist, but all current rule candidates are blocked by missing models on main.` -- `Bundle status: requested-missing` => STOP with `Issue #N is not currently in the Ready column.` -- `Bundle status: requested-blocked` => STOP with the blocking reason from the report -- `Bundle status: ready` => continue - -The report already handled the deterministic setup: -- it loaded the Ready and In-progress issue sets -- it scanned existing problems on main -- it marked blocked `[Rule]` issues whose source or target model is still missing -- it computed the pending-rule unblock counts used for C3 - -#### 0a. Score Eligible Issues - -**Short-circuit:** If there is only 1 eligible issue, skip scoring and pick it directly. Print "Only 1 eligible issue, picking it." and jump to Step 0c. - -Score only **eligible** issues on three criteria. For `[Model]` issues, extract the problem name. For `[Rule]` issues, extract both source and target problem names. - -| Criterion | Weight | How to Assess | -|-----------|--------|---------------| -| **C1: Industrial/Theoretical Importance** | 3 | Read the report's issue summary for each eligible issue. Score 0-2: **2** = widely used in industry or foundational in complexity theory (e.g., ILP, SAT, MaxFlow, TSP, GraphColoring); **1** = moderately important or well-studied (e.g., SubsetSum, SetCover, Knapsack); **0** = niche or primarily academic | -| **C2: Related to Existing Problems** | 2 | Use the report's Ready/In-progress context plus `pred list <candidate>` or `pred list --json` if needed. Score 0-2: **2** = directly related (shares input structure or has known reductions to/from ≥2 existing problems, but is NOT a trivial variant of an existing one); **1** = loosely related (same domain, connects to 1 existing problem); **0** = isolated or is essentially a variant/renaming of an existing problem | -| **C3: Unblocks Pending Rules** | 2 | Read the `Pending rules unblocked` count already printed in the report for each eligible issue. Score 0-2: **2** = unblocks ≥2 pending rules; **1** = unblocks 1 pending rule; **0** = does not unblock any pending rule | - -**Final score** = C1 × 3 + C2 × 2 + C3 × 2 (max = 12) - -**Tie-breaking:** Models before Rules, then by lower issue number. - -**Important for C2:** A problem that is merely a weighted/unweighted variant or a graph-subtype specialization of an existing problem scores **0** on C2, not 2. The goal is to add genuinely new problem types that expand the graph's reach. - -#### 0b. Print Ranked List - -Print all Ready issues with their scores for visibility (no confirmation needed). Blocked rules appear at the bottom with their reason: - -``` -Ready issues (ranked): - Score Issue Title C1 C2 C3 - ───────────────────────────────────────────────────────────── - 10 #117 [Model] GraphPartitioning 2 2 2 - 8 #129 [Model] MultivariateQuadratic 2 1 1 - 7 #97 [Rule] BinPacking to ILP 1 2 1 - 6 #110 [Rule] LCS to ILP 1 1 1 - 4 #126 [Rule] KSatisfiability to SubsetSum 0 2 0 - - Blocked: - 3 #130 [Rule] MultivariateQuadratic to ILP -- model "MultivariateQuadratic" not yet implemented -``` - -#### 0c. Pick Issues - -**If a specific issue number was provided:** validate and claim it through the scripted bundle: - -```bash -STATE_FILE=/tmp/problemreductions-ready-selection.json -CLAIM=$(python3 scripts/pipeline_board.py claim-next ready "$STATE_FILE" --number <number> --format json) -``` - -The report should already have stopped you before this point if the requested issue was missing or blocked. - -After successful validation, extract `ITEM_ID`, `ISSUE`, and `TITLE` from `CLAIM` using the same commands shown below. - -**Otherwise (no args):** score the eligible issues from the report, pick the highest-scored one, and proceed immediately (no confirmation). After picking the issue number, claim it through the scripted bundle: - -```bash -STATE_FILE=/tmp/problemreductions-ready-selection.json -CLAIM=$(python3 scripts/pipeline_board.py claim-next ready "$STATE_FILE" --number <chosen-issue-number> --format json) -``` - -Extract the board item metadata from `CLAIM`: - -```bash -ITEM_ID=$(printf '%s\n' "$CLAIM" | python3 -c "import sys,json; print(json.load(sys.stdin)['item_id'])") -ISSUE=$(printf '%s\n' "$CLAIM" | python3 -c "import sys,json; data=json.load(sys.stdin); print(data['issue_number'] or data['number'])") -TITLE=$(printf '%s\n' "$CLAIM" | python3 -c "import sys,json; print(json.load(sys.stdin)['title'])") -``` - -### 1. Create Worktree - -Create an isolated worktree for this issue: - -```bash -REPO_ROOT=$(pwd) -WORKTREE_JSON=$(python3 scripts/pipeline_worktree.py enter --name "issue-$ISSUE" --format json) -WORKTREE_DIR=$(printf '%s\n' "$WORKTREE_JSON" | python3 -c "import sys,json; print(json.load(sys.stdin)['worktree_dir'])") -cd "$WORKTREE_DIR" -``` - -All subsequent steps run inside the worktree. This ensures the user's main checkout is never modified. - -`issue-to-pr` (Step 3) handles all PR detection and branch management — if an existing open PR exists, it checks out that branch and resumes; otherwise it creates a fresh branch from `origin/main`. - -### 2. Claim Result - -`claim-next ready` has already moved the selected issue from `Ready` to `In progress`. Keep using `ITEM_ID` from the `CLAIM` JSON payload for later board transitions. - -### 3. Run issue-to-pr - -Invoke the `issue-to-pr` skill (working directory is the worktree): - -``` -/issue-to-pr "$ISSUE" -``` - -This handles the full pipeline: fetch issue, verify Good label, research, write plan, create PR, implement. If an existing open PR is detected, `issue-to-pr` will resume it (skip plan creation, jump to execution). - -**If `issue-to-pr` fails:** move the issue to OnHold with a diagnostic comment (see Step 4). - -### 4. Move to "Review pool" - -After `issue-to-pr` fully succeeds, move the issue to the `Review pool` column. "Fully succeeds" means the implementation work is committed, the temporary plan file has been deleted, the PR implementation summary comment has been posted, the branch has been pushed, and the working tree is clean aside from ignored/generated files: - -```bash -python3 scripts/pipeline_board.py move <ITEM_ID> review-pool -``` - -**If `issue-to-pr` failed (whether or not a PR was created):** move the issue to `OnHold` with a diagnostic comment explaining what went wrong: - -```bash -gh issue comment <ISSUE> --body "run-pipeline: implementation failed. <brief reason>" -python3 scripts/pipeline_board.py move <ITEM_ID> on-hold -``` - -**Forward-only rule:** never move items backward (e.g., back to Ready). All failures go to OnHold for human triage. - -### 5. Clean Up Worktree - -After the issue is processed (success or failure), clean up the worktree: - -```bash -cd "$REPO_ROOT" -python3 scripts/pipeline_worktree.py cleanup --worktree "$WORKTREE_DIR" -``` - -### 6. Report - -Print a summary: - -``` -Pipeline complete: - Issue: #97 [Rule] BinPacking to ILP - PR: #200 - Status: Awaiting agentic review - Board: Moved Ready -> In Progress -> Review pool -``` - -## Common Mistakes - -| Mistake | Fix | -|---------|-----| -| Issue not in Ready column | Verify status before processing; STOP if not Ready | -| Picking a Rule whose model doesn't exist | Hard constraint: both source and target models must exist on `main` — pending Model issues do NOT count | -| Missing project scopes | Run `gh auth refresh -s read:project,project` | -| Moving items backward to Ready | Never move backward — all failures go to OnHold with diagnostic comment | -| Scoring a variant as "related" | Weighted/unweighted variants or graph-subtype specializations of existing problems score 0 on C2 | -| Worktree left behind on failure | Always run `pipeline_worktree.py cleanup` in Step 5 | -| Working in main checkout | All work happens in the worktree — never modify the main checkout | -| Missing items from project board | `gh project item-list` defaults to 30 items — always use `--limit 500` | -| Inventing `pipeline_board.py` subcommands | Only `next`, `claim-next`, `ack`, `list`, `move`, `backlog` exist | diff --git a/.claude/skills/topology-sanity-check/SKILL.md b/.claude/skills/topology-sanity-check/SKILL.md deleted file mode 100644 index 12c405ede..000000000 --- a/.claude/skills/topology-sanity-check/SKILL.md +++ /dev/null @@ -1,284 +0,0 @@ ---- -name: topology-sanity-check -description: Run sanity checks on the reduction graph topology — detect orphan (isolated) problems, NP-hardness proof gaps, and redundant reduction rules ---- - -# Topology Sanity Check - -Runs structural health checks on the reduction graph. Detects orphan problems, verifies NP-hardness proof chains from 3-SAT, and identifies redundant reduction rules dominated by composite paths. - -## Invocation - -``` -/topology-sanity-check # Run ALL checks -/topology-sanity-check orphans # Orphan detection only -/topology-sanity-check np-hardness # 3-SAT reachability only -/topology-sanity-check redundancy # Rule redundancy only -/topology-sanity-check redundancy <source> <target> # Check a specific rule -``` - -Examples: -``` -/topology-sanity-check -/topology-sanity-check orphans -/topology-sanity-check np-hardness -/topology-sanity-check redundancy -/topology-sanity-check redundancy MIS ILP -``` - ---- - -## Check 1: Orphan Detection (`orphans`) - -Finds problem types that have no reduction rules connecting them to the main graph — they are registered but unreachable. - -### Run - -```bash -cargo run --example detect_isolated_problems 2>&1 -``` - -### Report - -Parse the output and produce: - -```markdown -## Orphan Detection Report - -### Graph Summary -- Total problem types: N -- Connected components: N - -### Isolated Problems (N) - -| # | Problem | Variants | Category | -|---|---------|----------|----------| -| 1 | ProblemName | 2 (f64, i64) | algebraic | - -### Existing Issues - -For each isolated problem, search GitHub for open issues that would connect it: - - gh issue list --repo CodingThrust/problem-reductions --state open --search "<problem name>" --json number,title - -Report matches or "No issues filed". - -### Verdict - -- **All connected**: Every problem type is reachable from the main component. -- **N orphans found**: List them and reference issue #610 (meta-issue tracking connectivity). -``` - ---- - -## Check 2: NP-Hardness Proof Chains (`np-hardness`) - -Verifies that every NP-hard problem has a directed reduction path from 3-SAT, constituting a proof of NP-hardness. Problems without such a path are classified as: in P (correctly unreachable), intermediate complexity, orphans, or NP-hard with a missing proof chain. - -### Run - -```bash -cargo run --example detect_unreachable_from_3sat 2>&1 -``` - -### Report - -Parse the output and produce: - -```markdown -## NP-Hardness Proof Chain Report - -### Summary -- Total problem types: N -- Reachable from 3-SAT: N -- Not reachable: N - -### Reachable from 3-SAT (N) - -| # | Problem | Hops | -|---|---------|------| -| 1 | KSatisfiability | 0 | -| 2 | Satisfiability | 1 | - -### NP-hard but missing proof chain (N) — needs new reductions - -| # | Problem | Outgoing | Incoming | -|---|---------|----------|----------| -| 1 | MaximumClique | 1 | 0 | - -### Correctly unreachable - -**In P:** MaximumMatching, KSatisfiability(K2), KColoring(K2) -**Intermediate complexity:** Factoring - -### Orphans (no edges at all) -[list] - -### Verdict - -- **PASS**: All NP-hard problems have proof chains from 3-SAT -- **WARN**: N NP-hard problems missing proof chains (list them) -``` - -The script automatically classifies unreachable problems. Problems with 0 incoming AND 0 outgoing reductions are orphans. Known P-time problems (MaximumMatching, 2-SAT, 2-Coloring) and intermediate-complexity problems (Factoring) are flagged as correctly unreachable. Everything else is reported as a missing proof chain. - ---- - -## Check 3: Rule Redundancy (`redundancy`) - -Determines whether reduction rules are redundant (dominated by composite paths through the reduction graph). Can check all primitive rules or a single source-target pair. - -### Mode A: Check All Rules (no arguments after `redundancy`) - -Run the codebase's `find_dominated_rules` analysis test: - -```bash -cargo test test_find_dominated_rules_returns_known_set -- --nocapture 2>&1 -``` - -This runs the analysis from `src/rules/analysis.rs` which: -1. Enumerates every primitive reduction rule (direct edge) in the graph -2. For each, finds all alternative composite paths -3. Uses polynomial normalization and monomial-dominance to compare overheads -4. Reports dominated rules and unknown comparisons - -Always report rules with full variant-qualified endpoints, not just base names. -Use the same display style as `ReductionStep`, e.g. -`MaximumIndependentSet {graph: "SimpleGraph", weight: "One"} -> MaximumIndependentSet {graph: "KingsSubgraph", weight: "i64"}`. -Base-name-only summaries are ambiguous and can hide cast-only paths. - -Parse the test output and report: - -```markdown -## Rule Redundancy Report - -### Dominated Rules (N) - -| # | Rule | Dominating Path | -|---|------|-----------------| -| 1 | Source {variant...} -> Target {variant...} | A -> B -> C | - -### Unknown Comparisons (N) - -| # | Rule | Reason | -|---|------|--------| -| 1 | Source {variant...} -> Target {variant...} | expression comparison returned Unknown | - -### Allowed (acknowledged) dominated rules - -List the entries from the `allowed` set in `test_find_dominated_rules_returns_known_set` -(file: `src/unit_tests/rules/analysis.rs`), and note when that allow-list is keyed only by base names while the reported dominated rule is variant-specific. - -### Verdict - -- If test passes: all dominated rules are acknowledged in the allow-list. -- If test fails: report the unexpected dominated rule or stale allow-list entry. -``` - -### Mode B: Check Single Rule (source target arguments) - -#### Step 1: Resolve Problem Names - -Use MCP tools (`show_problem`) to validate and resolve aliases (MIS = MaximumIndependentSet, MVC = MinimumVertexCover, SAT = Satisfiability, etc.). - -#### Step 2: Check if Rule Already Exists - -Use `show_problem` on the source and check its `reduces_to` array for a direct edge to the target. - -- **Direct edge exists**: Report "Direct rule `<source> -> <target>` already exists" and proceed to redundancy analysis (Step 3). -- **No direct edge**: Report "No direct rule from `<source> -> <target>` exists yet." Then check if any path exists: - - Use `find_path` MCP tool. - - **Path exists**: Report the cheapest existing path and its overhead. This is the baseline the proposed new rule must beat to be non-redundant. - - **No path exists**: Report "No path exists — a new rule would be novel (not redundant)." Stop here. - -#### Step 3: Find All Paths - -Use `find_path` with `all: true` to get all paths between source and target. - -#### Step 4: Compare Overheads - -For each composite path (length > 1 step): - -1. Extract the **overall overhead** from the path result -2. Extract the **direct rule's overhead** from the single-step path -3. Compare field by field: - - For polynomial expressions: compare degree — lower degree means the composite is better - - For equal-degree polynomials: compare leading coefficients - - For non-polynomial (exp, log): report as "Unknown — manual review needed" - -**Dominance definition:** A composite path **dominates** the direct rule if, on every common overhead field, the composite's expression has equal or smaller asymptotic growth. - -#### Step 5: Report Results - -```markdown -## Redundancy Check: <Source> -> <Target> - -### Direct Rule -- Rule: `Source {variant...} -> Target {variant...}` -- Overhead: [field = expr, ...] - -### Composite Paths Found: N - -| # | Path | Steps | Overhead | Comparison | -|---|------|-------|----------|------------| -| 1 | A -> B -> C | 2 | field = expr | Dominates / Worse / Unknown | - -### Verdict - -- **Redundant**: At least one composite path dominates the direct rule -- **Not Redundant**: No composite path dominates the direct rule -- **Inconclusive**: Some paths have Unknown comparison (non-polynomial overhead) - -### Recommendation - -If redundant: -> The direct rule `Source {variant...} -> Target {variant...}` is dominated by the composite path `[path]`. -> Consider removing it unless it provides value for: -> - Simpler solution extraction (fewer intermediate steps) -> - Educational/documentation clarity -> - Better numerical behavior in practice - -If not redundant: -> The direct rule `Source {variant...} -> Target {variant...}` is not dominated by any composite path. -> It provides overhead that cannot be achieved through existing reductions. -``` - ---- - -## Combined Report (no arguments) - -When invoked with no arguments, run all three checks and produce a combined report. Run Check 1 and Check 2 in parallel (both are `cargo run --example`), then Check 3 sequentially: - -```markdown -# Topology Sanity Check - -## 1. Orphan Detection -[orphan report] - -## 2. NP-Hardness Proof Chains -[np-hardness report] - -## 3. Rule Redundancy -[redundancy report] - -## Summary -- Orphans: N isolated problems -- NP-hard without proof chain: N -- Dominated rules: N (M acknowledged) -- Unknown comparisons: N -- Overall: PASS / WARN / FAIL -``` - -**Overall verdict:** -- **PASS**: No orphans, all NP-hard problems have proof chains, no unexpected dominated rules -- **WARN**: Has orphans (tracked in #610), missing proof chains, or unknown comparisons, but no unexpected dominated rules -- **FAIL**: Unexpected dominated rules found (test failure) - -## Notes - -- "Equal overhead" does not necessarily mean a rule should be removed — direct rules have practical advantages (simpler extraction, fewer steps) -- The analysis uses asymptotic comparison (big-O), so constant factors are ignored -- This means the check can produce false alarms, especially when overhead metadata keeps only leading terms or when a long composite path is asymptotically comparable but practically much worse -- Treat "dominated" as "potentially redundant, requires manual review" unless the composite path is also clearly preferable structurally -- When overhead expressions involve variables from different problems (e.g., `num_vertices` vs `num_clauses`), comparison may not be meaningful — report as Unknown -- The ground truth for what the codebase considers dominated is `src/rules/analysis.rs` (`find_dominated_rules`) with the allow-list in `src/unit_tests/rules/analysis.rs` (`test_find_dominated_rules_returns_known_set`) diff --git a/.claude/skills/update-papers/SKILL.md b/.claude/skills/update-papers/SKILL.md index 82e5a539c..f2b5d7f0f 100644 --- a/.claude/skills/update-papers/SKILL.md +++ b/.claude/skills/update-papers/SKILL.md @@ -1,129 +1,43 @@ --- name: update-papers -description: Update the research paper collection — download new papers from references.bib, retry failed downloads, sync to Google Drive, and regenerate index.md +description: Use when the research paper collection in docs/research/ needs refreshing — fetches PDFs for new references.bib entries, retries missing ones, regenerates the index, and syncs PDFs to the shared rclone remote --- # Update Papers -Maintain the research paper collection in `docs/research/`. Downloads papers referenced in `docs/paper/references.bib`, manages the manifest, syncs to Google Drive, and keeps `docs/research/index.md` current. +PDFs for `docs/paper/references.bib` live in `docs/research/raw/`, tracked by +`docs/research/manifest.json`; `docs/research/index.md` cross-references them against +`docs/paper/reductions.typ`. Scripts: `scripts/fetch_papers.py`, `scripts/gen_paper_index.py`. -## Prerequisites +Requires `rclone` with a `gdrive` remote and `PAPERS_REMOTE` set (e.g. +`export PAPERS_REMOTE=gdrive:problemreductions-papers`) for the push step. -- `rclone` installed and configured with a `gdrive` remote -- `PAPERS_REMOTE` env var set (e.g., `gdrive:problemreductions-papers`) - -## Step 1: Check Current Status - -```bash -make papers-status -``` - -Note the counts: total entries, PDFs on disk, pending downloads, missing papers. - -## Step 2: Lookup New Papers - -Run the lookup to find arxiv/OA URLs for any new entries in `references.bib` since the last run. This is incremental — it skips entries already found in the manifest. - -```bash -make papers-lookup -``` - -Review the output: -- New arxiv papers found -- New OA (open access) papers found -- Papers with no free source (will need Sci-Hub in Step 4) - -## Step 3: Download Free Papers - -Download papers with known free URLs (arxiv + open access). Skips PDFs already on disk. +## Flow ```bash -make papers-download +make papers # lookup (arXiv/OA via Semantic Scholar) -> download -> scihub -> status +make papers-index # regenerate docs/research/index.md +make papers-push # upload new/changed PDFs + manifest to $PAPERS_REMOTE ``` -If some OA downloads fail with 403, that's expected — publisher paywalls. These will be picked up by Sci-Hub in the next step. +All steps are incremental and idempotent; OA 403s are normal and fall through to Sci-Hub. -## Step 4: Fetch Remaining via Sci-Hub +## Leftovers -For papers with DOIs that aren't on disk yet, try Sci-Hub mirrors. This is the slowest step (~5 seconds per paper). +After `make papers`, check `make papers-status` for missing entries. For each one, search +`"<title>" <first author> pdf`: author homepages, arXiv by title, and open-access fallbacks +(LIPIcs/Dagstuhl, HAL, ECCC, IACR ePrint). Download with +`curl -L -o docs/research/raw/<bibkey>.pdf '<url>'` and confirm with `file` that it is a PDF, +not an HTML paywall page. Rerun `make papers-index` and `make papers-push` afterwards. -```bash -make papers-scihub -``` - -The script tries multiple mirrors (`sci-hub.ru`, `sci-hub.do`, `sci-hub.it.nf`, `sci-hub.es.ht`). If all mirrors are down, retry later — the script is fully idempotent. - -## Step 4b: Manual Web Search for Remaining Failures - -After Sci-Hub, check `make papers-status` for papers still missing. For each one with a DOI that Sci-Hub couldn't find: - -1. **Web search** for `"<title>" <first-author> PDF` — try: - - Author homepages (Stanford, university pages) - - Open-access publishers: LIPIcs/Dagstuhl (all free), HAL archives, ECCC - - Preprint servers: arxiv (search by title), IACR ePrint -2. **Download manually** with `curl -L -o docs/research/raw/<key>.pdf "<url>"` -3. **Verify** the file is a real PDF: `file docs/research/raw/<key>.pdf` - -Skip textbooks (garey1979, sipser2012, cormen2022, conway1967) — these aren't available as single PDFs. - -## Step 5: Regenerate Index - -Update `docs/research/index.md` with the latest paper collection, cross-referenced against reduction rules and problem definitions in `reductions.typ`. - -```bash -make papers-index -``` - -Verify the index looks correct: -- Check the download count at the top -- Spot-check that new papers appear in the correct section (rules / problems / other) -- Confirm PDF links resolve for newly downloaded papers - -## Step 6: Sync to Google Drive - -Push updated PDFs and manifest to the shared Google Drive remote. Only uploads new/changed files. - -First verify the remote is configured: - -```bash -echo $PAPERS_REMOTE # should show e.g. gdrive:problemreductions-papers -# If empty, set it: -export PAPERS_REMOTE=gdrive:problemreductions-papers -``` - -Then push: - -```bash -make papers-push -``` - -## Step 7: Final Status - -```bash -make papers-status -``` - -Report to the user: -- How many new papers were downloaded -- How many remain missing (and why: no DOI, textbooks, Sci-Hub mirrors down) -- Whether the Google Drive sync succeeded - -## One-Liner - -For a full update in one command: - -```bash -make papers && make papers-index -``` - -This runs: lookup → download → scihub → status → index. +Skip textbooks; they have no single PDF: garey1979, sipser2012, cormen2022, conway1967. ## Troubleshooting -**Sci-Hub mirrors all fail**: Mirrors rotate frequently. Update `SCIHUB_DOMAINS` in `scripts/fetch_papers.py` or retry later. - -**rclone auth expired**: Run `rclone config reconnect gdrive:` to refresh the OAuth token. - -**Manifest is stale**: Delete `docs/research/manifest.json` and re-run `make papers-lookup` to rebuild from scratch. Existing PDFs on disk are preserved. +- All Sci-Hub mirrors fail: retry later or update `SCIHUB_DOMAINS` in `scripts/fetch_papers.py`. +- rclone auth expired: `rclone config reconnect gdrive:`. +- Stale manifest: delete `docs/research/manifest.json` and rerun `make papers`; PDFs on disk are + kept. +- New bib entry ignored: it must parse as `@type{key, ...}` with `title` and ideally `doi`. -**New bib entry not appearing**: Ensure the entry is in `docs/paper/references.bib` with proper formatting. The parser expects `@type{key, ... }` with fields like `title`, `doi`, `author`, `year`. +Report: newly downloaded, still missing (and why), and whether the push succeeded. diff --git a/.claude/skills/verify-reduction/SKILL.md b/.claude/skills/verify-reduction/SKILL.md deleted file mode 100644 index 91765ce3f..000000000 --- a/.claude/skills/verify-reduction/SKILL.md +++ /dev/null @@ -1,295 +0,0 @@ ---- -name: verify-reduction -description: Standalone mathematical verification of a reduction rule — generates a Typst proof plus constructor and independent adversary scripts with at least 5000 checks each. Reports a verdict without saving artifacts. ---- - -# Verify Reduction - -Mathematical verification of a reduction rule. Produces a Typst proof + dual Python verification scripts, iterating until all checks pass. Reports a VERIFIED/FAILED verdict. All artifacts are ephemeral — nothing is committed to the repository. - -Use standalone to check correctness before implementation, or as a subroutine of `/add-rule` (which calls this by default). - -## Invocation - -``` -/verify-reduction 868 -/verify-reduction SubsetSum Partition -``` - -## Step 0: Parse Input - -```bash -REPO=$(gh repo view --json nameWithOwner --jq .nameWithOwner) -ISSUE=<number> -ISSUE_JSON=$(gh issue view "$ISSUE" --json title,body,number) -``` - -If invoked with problem names instead of an issue number, use the names directly. - -## Step 1: Read Issue, Study Models, Type Check - -```bash -gh issue view "$ISSUE" --json title,body -pred show <Source> --json -pred show <Target> --json -``` - -### Type compatibility gate — MANDATORY - -Check source/target `Value` types before any work. The `grep` only locates the definitions; it does -not resolve generic parameters or associated types: - -```bash -grep "type Value = " src/models/*/<source_file>.rs src/models/*/<target_file>.rs -``` - -Resolve both concrete types completely before declaring compatibility: - -1. Substitute every concrete generic argument from the proposed rule. -2. Follow every type alias and associated type to its defining `impl`. -3. Record the substitution chain and the source file evidence in the verification report. -4. If any generic or associated type remains unresolved, run a compile-backed temporary Rust probe - using `std::any::type_name::<<ConcreteProblem as Problem>::Value>()`. Build the probe from `/tmp` - with a path dependency on this repository; do not modify the repository. - -Never infer a Rust value type from the mathematical problem name, from unit-weight terminology, or -from the Python verifier's integer representation. In particular, arbitrary-precision Python -integers do not establish that a Rust objective type is `usize` or that it is closed under all -legal source instances. - -Required report format: - -```text -TYPE RESOLUTION: - Source syntax: Min<W::Sum> - Substitutions: W = One; <One as WeightElement>::Sum = i64 - Source resolved: Min<i64> - Target syntax: Min<usize> - Target resolved: Min<usize> - Full-domain compatibility: FAILED -``` - -**Compatible pairs for `ReduceTo` (witness-capable):** -- `Or`->`Or` -- `Min<V>`->`Min<V>`, `Max<V>`->`Max<V>` (identical resolved inner type) -- `Or`->`Min`, `Or`->`Max` (feasibility embeds into optimization) - -`Min<S>`->`Min<T>` or `Max<S>`->`Max<T>` with `S != T` is not automatically compatible. Proceed -only if the rule or source model declares a bound covering every legal source instance and the -verification proves a total, order-preserving conversion over that full declared domain. Otherwise -STOP and report a value-domain mismatch. - -**Incompatible — STOP if any of these:** -- `Min`->`Or` or `Max`->`Or` — optimization source has no threshold K; needs a decision-variant source model -- `Max`->`Min` or `Min`->`Max` — opposite optimization directions; needs `ReduceToAggregate` or a decision-variant wrapper -- `Or`->`Sum` or `Min`->`Sum` — Sum is aggregate-only; needs `ReduceToAggregate` -- Any pair involving `And` or `Sum` on the target side - -**Regression case:** `MinimumDominatingSet<SimpleGraph, One>` resolves to `Min<i64>` because -`<One as WeightElement>::Sum = i64`; `MinimumHittingSet` resolves to `Min<usize>`. Report -`Min<i64> -> Min<usize>`, not `Min<usize> -> Min<usize>`. Without an explicit source-size bound, -the full-domain type gate fails even though the classical cardinality reduction is mathematically -correct and exhaustive small-instance checks pass. - -If incompatible, STOP and report the type mismatch and options. Do NOT proceed. - -### If compatible - -Extract: construction algorithm, correctness argument, overhead formulas, worked example, reference. Use WebSearch if the issue is incomplete. - -## Step 2: Write Typst Proof - -Write a standalone Typst proof (in a temp file, not committed). - -**Mandatory structure:** - -```typst -== Source $arrow.r$ Target <sec:source-target> -#theorem[...] <thm:source-target> -#proof[ - _Construction._ (numbered steps, every symbol defined before first use) - _Correctness._ - ($arrow.r.double$) ... (genuinely independent, NOT "the converse is similar") - ($arrow.l.double$) ... - _Solution extraction._ ... -] -*Overhead.* (table with target metric -> formula) -*Feasible example.* (YES instance, >=3 variables, fully worked with numbers) -*Infeasible example.* (NO instance, fully worked — show WHY no solution exists) -``` - -**Hard rules:** -- Zero instances of "clearly", "obviously", "it is easy to see", "straightforward" -- Zero scratch work ("Wait", "Hmm", "Actually", "Let me try") -- Two examples minimum, both with >=3 variables/vertices -- Every symbol defined before first use - -## Step 3: Write Constructor Python Script - -Write a Python verification script (temp file) with ALL 7 mandatory sections: - -| Section | What to verify | Notes | -|---------|---------------|-------| -| 1. Symbolic (sympy) | Overhead formulas symbolically for general n | "The overhead is trivial" is NOT an excuse to skip | -| 2. Exhaustive forward+backward | Source feasible <=> target feasible | n <= 5 minimum. ALL instances or >=300 sampled per (n,m) | -| 3. Solution extraction | Extract source solution from every feasible target witness | Most commonly skipped section. DO NOT SKIP | -| 4. Overhead formula | Build target, measure actual size, compare against formula | Catches off-by-one in construction | -| 5. Structural properties | Target well-formed, no degenerate cases | Gadget reductions: girth, connectivity, widget structure | -| 6. YES example | Reproduce exact Typst feasible example numbers | Every value must match | -| 7. NO example | Reproduce exact Typst infeasible example, verify both sides infeasible | Must verify WHY infeasible | - -### Minimum check counts — STRICTLY ENFORCED - -| Type | Minimum checks | Minimum n | -|------|---------------|-----------| -| Identity (same graph, different objective) | 10,000 | n <= 6 | -| Algebraic (padding, complement, case split) | 10,000 | n <= 5 | -| Gadget (widget, cycle construction) | 5,000 | n <= 5 | - -Every reduction gets at least 5,000 checks regardless of perceived simplicity. - -## Step 4: Run and Iterate - -```bash -python3 /tmp/verify_<source>_<target>.py -``` - -### Iteration 1: Fix failures - -Run the script. Fix any failures. Re-run until 0 failures. - -### Iteration 2: Check count audit - -Print and fill this table honestly: - -``` -CHECK COUNT AUDIT: - Total checks: ___ (minimum: 5,000) - Forward direction: ___ instances (minimum: all n <= 5) - Backward direction: ___ instances (minimum: all n <= 5) - Solution extraction: ___ feasible instances tested - Overhead formula: ___ instances compared - Symbolic (sympy): ___ identities verified - YES example: verified? [yes/no] - NO example: verified? [yes/no] - Structural properties: ___ checks -``` - -If ANY line is below minimum, enhance the script and re-run. Do NOT proceed. - -### Iteration 3: Gap analysis - -List EVERY claim in the Typst proof and whether it's tested: - -``` -CLAIM TESTED BY -"Universe has 2n elements" Section 4: overhead -"Complementarity forces consistency" Section 3: extraction -"Forward: NAE-sat -> valid splitting" Section 2: exhaustive -... -``` - -If any claim has no test, add one. If untestable, document WHY. - -## Step 5: Adversary Verification - -Dispatch a subagent that reads ONLY the Typst proof (not the constructor script) and independently implements + tests the reduction. - -**Adversary requirements:** -- Own `reduce()` function from scratch -- Own `extract_solution()` function -- Own `is_feasible_source()` and `is_feasible_target()` validators -- Exhaustive forward + backward for n <= 5 -- `hypothesis` property-based testing (>=2 strategies) -- Reproduce both Typst examples (YES and NO) -- >=5,000 total checks -- Must NOT import from the constructor script - -**Typed adversary focus** (include in prompt): -- **Identity reductions:** exhaustive enumeration n <= 6, edge-case configs (all-zero, all-one, alternating) -- **Algebraic reductions:** case boundary conditions (e.g., S = 2T exactly, S = 2T +/- 1), per-case extraction -- **Gadget reductions:** widget structure invariants, traversal patterns, interior vertex isolation - -### Cross-comparison - -After both scripts pass, compare `reduce()` outputs on shared instances. Both must produce structurally identical targets and agree on feasibility for all tested instances. - -### Verdict table - -| Constructor | Adversary | Cross-compare | Verdict | Action | -|-------------|-----------|---------------|---------|--------| -| Pass | Pass | Agree | **VERIFIED** | Done (or proceed to add-rule Step 2) | -| Pass | Pass | Disagree | **Suspect** | Investigate — may be isomorphic or latent bug | -| Pass | Fail | -- | **Adversary bug** | Fix adversary or clarify Typst spec | -| Fail | Pass | -- | **Constructor bug** | Fix constructor, re-run from Step 4 | -| Fail | Fail | -- | **Proof bug** | Re-examine Typst proof, return to Step 2 | - -## Step 6: Self-Review Checklist - -Every item must be YES. If any is NO, go back and fix. - -### Typst proof -- [ ] Construction with numbered steps, symbols defined before use -- [ ] Correctness with independent => and <= paragraphs -- [ ] Solution extraction section present -- [ ] Overhead table with formulas -- [ ] YES example (>=3 variables, fully worked) -- [ ] NO example (fully worked, explains WHY infeasible) -- [ ] Zero hand-waving language -- [ ] Zero scratch work - -### Type gate -- [ ] Concrete Rust `Value` types fully resolved with substitution evidence -- [ ] Different numeric domains either rejected or covered by an explicit full-domain range proof - -### Constructor Python -- [ ] 0 failures, >=5,000 total checks -- [ ] All 7 sections present and non-empty -- [ ] Exhaustive n <= 5 -- [ ] Extraction tested for every feasible instance -- [ ] Gap analysis: every Typst claim has a test - -### Adversary Python -- [ ] 0 failures, >=5,000 total checks -- [ ] Independent implementation (no imports from constructor) -- [ ] `hypothesis` PBT with >=2 strategies -- [ ] Reproduces both Typst examples - -### Cross-consistency -- [ ] Cross-comparison: 0 disagreements, 0 feasibility mismatches - -## Step 7: Report Verdict - -Report the final verdict to the user: - -``` -VERIFICATION RESULT: VERIFIED / FAILED - Source: <Source> - Target: <Target> - Constructor checks: <N> - Adversary checks: <N> - Cross-comparison: <N> instances, 0 disagreements - Issue: #<N> -``` - -If called as a subroutine of `/add-rule`, the verified Python `reduce()`, `extract_solution()`, and YES/NO instances remain in conversation context for use in the Rust implementation steps. No files are saved. - -If called standalone, the verdict is the final output. The user can inspect the proof and scripts interactively during the session. - -## Common Mistakes - -| Mistake | Consequence | -|---------|-------------| -| Proceeding past type gate with incompatible types | Wasted work — math may be correct but `ReduceTo` impl is impossible | -| Adversary imports from constructor script | Rejected — must be independent | -| No `hypothesis` PBT in adversary | Rejected | -| Section 1 (symbolic) empty | Rejected — "overhead is trivial" is not an excuse | -| Only YES example, no NO example | Rejected | -| n <= 3 or n <= 4 "because it's simple" | Rejected — minimum n <= 5 | -| No gap analysis | Rejected — perform before proceeding | -| Example has < 3 variables | Rejected — too degenerate | -| Either script has < 5,000 checks | Rejected — enhance testing | -| Extraction (Section 3) not tested | Rejected — most commonly skipped | -| Cross-comparison skipped | Rejected | -| Disagreements dismissed without investigation | Rejected | -| Saving artifacts to the repository | All files are ephemeral — use temp directory, nothing committed | diff --git a/.claude/skills/write-model-in-paper/SKILL.md b/.claude/skills/write-model-in-paper/SKILL.md deleted file mode 100644 index d8ca47ba2..000000000 --- a/.claude/skills/write-model-in-paper/SKILL.md +++ /dev/null @@ -1,212 +0,0 @@ ---- -name: write-model-in-paper -description: Use when writing or improving a problem-def entry in the Typst paper (docs/paper/reductions.typ) ---- - -# Write Problem Model in Paper - -Full authoring guide for writing a `problem-def` entry in `docs/paper/reductions.typ`. Covers formal definition, background, examples with visualization, and verification. - -> **Note:** This content is also inlined in `add-model` Step 6 (condensed form). This standalone version has more detail and is useful for improving existing entries. - -## Prerequisites - -Before using this skill, ensure: -- The problem model is implemented (`src/models/<category>/<name>.rs`) -- The problem is registered with schema and variant metadata -- A canonical example exists in `src/example_db/model_builders.rs` -- JSON exports are up to date (`cargo run --example export_graph && cargo run --example export_schemas`) - -## Reference Example - -**MaximumIndependentSet** in `docs/paper/reductions.typ` is the gold-standard model example. Search for `problem-def("MaximumIndependentSet")` to see the complete entry. Use it as a template for style, depth, and structure. - -## The `problem-def` Function - -```typst -#problem-def("ProblemName")[ - Formal definition... // parameter 1: def -][ - Background, example, figure... // parameter 2: body -] -``` - -**Three parameters:** -- `name` (string) — problem name matching `display-name` dictionary key -- `def` (content) — formal mathematical definition -- `body` (content) — background, examples, figures, algorithm list - -**Auto-generated between `def` and `body`:** -- Variant complexity table (from Rust `declare_variants!` metadata) -- Reduction links (from reduction graph JSON) -- Schema field table (from problem schema JSON) - -## Step 1: Register Display Name - -Add to the `display-name` dictionary near the top of `reductions.typ`: - -```typst -"ProblemName": [Display Name], -``` - -## Step 2: Write the Formal Definition (`def` parameter) - -One self-contained sentence or short paragraph. Requirements: - -1. **Introduce all inputs first** — graph, weights, sets, variables with their domains -2. **State the objective or constraint** — what is being optimized or satisfied -3. **Define all notation before use** — every symbol must be introduced before it appears - -### Pattern for optimization problems - -```typst -Given [inputs with domains], find [solution variable] [maximizing/minimizing] [objective] such that [constraints]. -``` - -### Pattern for satisfaction problems - -```typst -Given [inputs with domains], find [solution variable] such that [constraints]. -``` - -### Example (MIS) - -```typst -Given $G = (V, E)$ with vertex weights $w: V -> RR$, find $S subset.eq V$ -maximizing $sum_(v in S) w(v)$ such that no two vertices in $S$ are -adjacent: $forall u, v in S: (u, v) in.not E$. -``` - -## Step 3: Write the Body - -The body goes AFTER the auto-generated sections (complexity, reductions, schema). It contains four parts in order: - -### 3a. Background & Motivation - -1-3 sentences covering: -- Historical context (e.g., "One of Karp's 21 NP-complete problems") -- Applications (e.g., "appears in wireless network scheduling, register allocation") -- Notable structural properties (e.g., "Solvable in polynomial time on bipartite graphs, interval graphs, chordal graphs") - -If the user provides specific justification or motivation, incorporate it here. - -### 3b. Best Known Algorithms - -Must clearly state which algorithm gives the best complexity and cite reference. Add a warning as footnote if no reliable reference is found. - -Integrate algorithm complexity naturally into the background prose — do NOT append a terse "Best known: $O^*(...)$" at the end: - -```typst -% Good: names the algorithm, cites reference -The best known algorithm runs in $O^*(1.1996^n)$ time via measure-and-conquer -branching @xiao2017. - -% Good: brute-force with footnote when no better algorithm is known -The best known algorithm runs in $O^*(2^n)$ by brute-force -enumeration#footnote[No algorithm improving on brute-force is known for ...]. - -% Bad: terse appendage, no algorithm name, no reference -Best known: $O^*(2^n)$. -``` - -For problems with multiple notable algorithms or special cases, weave them into the text: -```typst -Solvable in $O(n+m)$ for $k=2$ via bipartiteness testing. For $k=3$, the best -known algorithm runs in $O^*(1.3289^n)$ @beigel2005; in general, inclusion-exclusion -achieves $O^*(2^n)$ @bjorklund2009. -``` - -**Citation rules:** -- Every complexity claim MUST have a citation (`@key`) identifying the algorithm -- If the best known is brute-force enumeration with no specialized algorithm, add footnote: `#footnote[No algorithm improving on brute-force enumeration is known for ...]` -- If a reference exists but has not been independently verified, add footnote: `#footnote[Complexity not independently verified from literature.]` -- Include approximation results where relevant (e.g., "0.878-approximation @goemans1995") - -**Consistency note:** The auto-generated complexity table (from `declare_variants!`) also shows complexity per variant. The written text and the auto-generated table may overlap. Keep both — the written text provides references and context; the auto-generated table provides per-variant detail. A future verification step will check consistency between them. - -### 3c. Example with Visualization - -A concrete small instance that illustrates the problem. **Use the generated canonical example data**, not an independently invented instance. - -#### Sourcing example data - -1. If you changed example builders/specs, run `cargo run --features "example-db" --example export_examples`. -2. Find the problem's entry in `docs/paper/data/examples.json` under `models` — it contains the canonical `instance`, `samples`, and `optimal` fields. -3. Use the values from `instance` in the paper example (translating 0-indexed code values to 1-indexed math notation where conventional, e.g., vertices {0,...,n-1} → {1,...,n}). -4. Use `optimal` configurations to show the solution. - -**Do not invent a different instance.** If the canonical example is unsuitable, fix it in `canonical_model_example_specs()`, re-run `export_examples`, then use the updated JSON. - -#### Requirements - -1. **Small enough to verify by hand** — readers should be able to check the solution -2. **Include a diagram/graph** using the paper's visualization helpers -3. **Show a valid/optimal solution** and explain why it is valid/optimal -4. **Walk through evaluation** — show how the objective/verifier computes the solution value - -#### Structure - -```typst -*Example.* Consider [instance description with concrete numbers from `examples.json`]. -[Describe the solution and why it's valid/optimal]. - -#figure({ - // visualization code — see MaximumIndependentSet for graph rendering pattern -}, -caption: [Caption describing the figure with key parameters], -) <fig:problem-example> -``` - -#### Reproducibility Commands - -Add a `pred-commands()` block after the `*Example.*` paragraph and before the `#figure`. The `pred create --example ...` spec must be derived from the loaded example data, not handwritten: - -```typst -#let problem-spec(data) = { - if data.variant.len() == 0 { data.problem } - else { data.problem + "/" + data.variant.values().join("/") } -} - -#pred-commands( - "pred create --example " + problem-spec(x) + " -o <name>.json", - "pred solve <name>.json", - "pred evaluate <name>.json --config " + x.optimal_config.map(str).join(","), -) -``` - -If `docs/paper/reductions.typ` already defines a shared `problem-spec()` helper, reuse it instead of reintroducing it locally. Do **not** guess whether the default variant matches the canonical example; canonical fixtures may live on non-default variants, and handwritten bare aliases can silently produce broken commands. - -**For graph problems**, use the paper's existing graph helpers: -- `petersen-graph()`, `house-graph()` or define custom vertex/edge lists -- `canvas(length: ..., { ... })` with `g-node()` and `g-edge()` -- Highlight solution elements with `graph-colors.at(0)` (blue) and use `white` fill for non-solution - -Refer to the **MaximumIndependentSet** entry for the complete graph rendering pattern. Adapt it to your problem. - -### 3d. Evaluation Explanation - -Explain how a configuration is evaluated — this maps to the Rust `evaluate()` method: -- For optimization: show the cost function computation on the example solution -- For satisfaction: show the verifier check on the example solution - -This can be woven into the example text (as MIS does: "$w(S) = sum_(v in S) w(v) = 4 = alpha(G)$"). - -## Step 4: Build and Verify - -```bash -# Build the paper (auto-runs export_graph + export_schemas) -make paper -``` - -### Verification Checklist - -- [ ] **Display name registered**: entry exists in `display-name` dictionary -- [ ] **Notation self-contained**: every symbol in `def` is defined before first use -- [ ] **Background present**: historical context, applications, or structural properties -- [ ] **Algorithms cited**: every complexity claim has `@citation` or footnote warning -- [ ] **Example from JSON**: instance data matches the canonical entry in `docs/paper/data/examples.json` -- [ ] **Evaluation shown**: objective/verifier computed on the example solution -- [ ] **Diagram included**: figure with caption and label for graph/matrix/set visualization -- [ ] **Paper compiles**: `make paper` succeeds without errors -- [ ] **Pred commands present**: `pred-commands()` block after example text, before figure, with create/solve/evaluate pipeline -- [ ] **Complexity consistency**: written complexity and auto-generated variant table are compatible (note any discrepancies for later review) diff --git a/.claude/skills/write-rule-in-paper/SKILL.md b/.claude/skills/write-rule-in-paper/SKILL.md deleted file mode 100644 index e7d3c884b..000000000 --- a/.claude/skills/write-rule-in-paper/SKILL.md +++ /dev/null @@ -1,277 +0,0 @@ ---- -name: write-rule-in-paper -description: Use when writing or improving a reduction-rule entry in the Typst paper (docs/paper/reductions.typ) ---- - -# Write Reduction Rule in Paper - -Full authoring guide for writing a `reduction-rule` entry in `docs/paper/reductions.typ`. Covers Typst mechanics, writing quality, and verification. - -> **Note:** This content is also inlined in `add-rule` Step 6 (condensed form). This standalone version has more detail and is useful for improving existing entries. - -## Reference Example - -**KColoring → QUBO** in `docs/paper/reductions.typ` is the gold-standard reduction example. Search for `reduction-rule("KColoring", "QUBO"` to see the complete entry. Use it as a template for style, depth, and structure. - -## Prerequisites - -Before using this skill, ensure: -- The reduction is implemented and tested (`src/rules/<source>_<target>.rs`) -- A rule-local `canonical_rule_example_specs()` exists and is included by `src/rules/mod.rs` -- If the canonical example changed, regenerate the paper data with `cargo run --features "example-db" --example export_examples` -- The reduction graph and schemas are up to date (`cargo run --example export_graph && cargo run --example export_schemas`) - -## Source Material - -For mathematical content (theorems, proofs, examples), consult these sources in priority order: -1. **GitHub issue** for the rule (`gh issue view <number>`): contains the verified reduction algorithm, correctness proof, parameter transform, and worked examples written during issue creation -2. **Derivation documents** (if available): e.g., `~/Downloads/reduction_derivations_*.typ` — these contain batch-verified proofs with explicit theorem/proof blocks -3. **The implementation** (`src/rules/<source>_<target>.rs`): the code is the ground truth for the construction - -Do NOT invent proofs — always cross-check against the issue and derivation sources. The issue body typically has: Reduction Algorithm, Correctness (forward/backward), Size Overhead, Example, and References sections that map directly to the paper entry structure. - -## Step 1: Load Example Data - -```typst -#let src_tgt = load-example("Source", "Target") -#let src_tgt_sol = src_tgt.solutions.at(0) -``` - -Where: -- `load-example(source, target, ...)` looks up the canonical rule entry from `docs/paper/data/examples.json` -- The returned record contains `source`, `target`, and `solutions` -- Access fields: `src_tgt.source.instance`, `src_tgt.target.instance`, `src_tgt_sol.source_config`, `src_tgt_sol.target_config` - -## Step 2: Write the Theorem Body (Rule Statement) - -The theorem body is a concise block with three parts: - -### 2a. Complexity with Reference - -State the reduction's time complexity with a citation. Examples: - -```typst -% With verified reference: -This $O(n + m)$ reduction @Author2023 constructs ... - -% Without verified reference — add footnote: -This $O(n^2)$ reduction#footnote[Complexity not independently verified from literature.] constructs ... -``` - -**Verification**: Identify the best known reference for this reduction's complexity. If you cannot find a peer-reviewed or textbook source, you MUST add the footnote. - -### 2b. Construction Summary - -One sentence describing what the reduction builds: - -```typst -... constructs an intersection graph $G' = (V', E')$ where ... -``` - -### 2c. Overhead Hint - -State target dimensions in terms of source. This complements the auto-derived overhead (which appears automatically from JSON edge data): - -```typst -... ($n k$ variables indexed by $v dot k + c$). -``` - -### Complete theorem body example - -```typst -][ - Given $G = (V, E)$ with $k$ colors, construct upper-triangular - $Q in RR^(n k times n k)$ using one-hot encoding $x_(v,c) in {0,1}$ - ($n k$ variables indexed by $v dot k + c$). -] -``` - -## Step 3: Write the Proof Body - -The proof must be **self-contained** (all notation defined before use) and **reproducible** (enough detail to reimplement the reduction from the proof alone). - -### Structure - -Use these subsections in order. Use italic labels exactly as shown: - -```typst -][ - _Construction._ ... - - _Correctness._ ... - - _Variable mapping._ ... // only if the reduction has a non-trivial variable mapping - - _Solution extraction._ ... -] -``` - -### 3a. Construction - -Full mathematical construction of the target instance. Define all symbols and notation here. - -**For standard reductions** (< 300 LOC): Write the complete construction with enough math to reimplement. - -**For heavy reductions** (300+ LOC): Briefly describe the approach and cite a reference: -```typst -_Construction._ The reduction follows the standard Cook–Levin construction @Cook1971, -encoding each gate as a set of clauses. See @Source for full details. -``` - -### 3b. Correctness - -Bidirectional (iff) argument showing solution correspondence. Use ($arrow.r.double$) and ($arrow.l.double$) for each direction: - -```typst -_Correctness._ ($arrow.r.double$) If $S$ is independent, then ... -($arrow.l.double$) If $C$ is a vertex cover, then ... -``` - -### 3c. Variable Mapping (if applicable) - -Explicitly state how source variables map to target variables. Include this section when the mapping is non-trivial (encoding, expansion, reindexing). Omit for identity mappings or trivial complement operations. - -```typst -_Variable mapping._ Vertices $= {S_1, ..., S_m}$, edges $= {(S_i, S_j) : S_i inter S_j != emptyset}$, $w(v_i) = w(S_i)$. -``` - -### 3d. Solution Extraction - -How to convert a target solution back to a source solution: - -```typst -_Solution extraction._ For each vertex $v$, find $c$ with $x_(v,c) = 1$. -``` - -## Step 4: Write the Worked Example (Extra Block) - -Detailed by default. Only use a brief example for trivially obvious reductions (complement, identity). - -### 4a. Typst Skeleton - -```typst -#reduction-rule("Source", "Target", - example: true, - example-caption: [Description ($n = ...$, $|E| = ...$)], - extra: [ - #pred-commands( - "pred create --example <ALIAS> -o <name>.json", - "pred reduce <name>.json --to " + target-spec(src_tgt) + " -o bundle.json", - "pred solve bundle.json", - "pred evaluate <name>.json --config " + src_tgt_sol.source_config.map(str).join(","), - ) - - // Optional: graph visualization - #{ - // canvas code for graph rendering - } - - *Step 1 -- [action].* [description with concrete numbers] - - *Step 2 -- [action].* [construction details] - - // ... more steps as needed - - *Step N -- Verify a solution.* [end-to-end verification] - - *Multiplicity:* The fixture stores one canonical witness. If total multiplicity matters, explain it from the construction. - ], -) -``` - -### 4a2. Reproducibility Commands - -Add a `pred-commands()` block at the top of the `extra:` content, before any data display or visualization. The source-side `pred create --example ...` spec must be derived from the loaded example data, not handwritten: - -```typst -#let problem-spec(data) = { - if data.variant.len() == 0 { data.problem } - else { data.problem + "/" + data.variant.values().join("/") } -} - -#pred-commands( - "pred create --example " + problem-spec(src_tgt.source) + " -o <source>.json", - "pred reduce <source>.json --to " + target-spec(src_tgt) + " -o bundle.json", - "pred solve bundle.json", - "pred evaluate <source>.json --config " + src_tgt_sol.source_config.map(str).join(","), -) -``` - -If `docs/paper/reductions.typ` already defines a shared `problem-spec()` helper, reuse it instead of duplicating it. Do **not** guess whether the default variant matches the canonical example; canonical rule fixtures can point at non-default source variants, and handwritten bare aliases can silently break the reproducibility block. - -The `target-spec()` helper handles empty variant dicts automatically. The `--config` is composed from `src_tgt_sol.source_config`. - -### 4b. Step-by-Step Content - -Each step should: -1. **Name the action** in bold: `*Step K -- [verb phrase].*` -2. **Show concrete numbers** from the example instance (use Typst expressions to extract from JSON, not hardcoded values) -3. **Explain where overhead comes from** — e.g., "5 vertices x 3 colors = 15 QUBO variables" - -### 4c. Required Steps - -| Step | Content | -|------|---------| -| First | Show the source instance (dimensions, structure). Include graph visualization if applicable. | -| Middle | Walk through the construction. Show intermediate values. Explicitly quantify overhead. | -| Second-to-last | Verify a concrete solution end-to-end (source config → target config, check validity). | -| Last | State that the fixture stores one canonical witness; if multiplicity matters, justify it mathematically from the construction. | - -### 4d. Graph Visualization (if applicable) - -```typst -#{ - let fills = src_tgt_sol.source_config.map(c => graph-colors.at(c)) - align(center, canvas(length: 0.8cm, { - for (u, v) in graph.edges { g-edge(graph.vertices.at(u), graph.vertices.at(v)) } - for (k, pos) in graph.vertices.enumerate() { - g-node(pos, name: str(k), fill: fills.at(k), label: str(k)) - } - })) -} -``` - -### 4e. Accessing Solution Data - -```typst -// Source configuration (e.g., color assignments) -#src_tgt_sol.source_config.map(str).join(", ") - -// Target configuration (e.g., binary encoding) -#src_tgt_sol.target_config.map(str).join(", ") - -// The canonical witness pair -#src_tgt.solutions.at(0) - -// Source instance fields -#src_tgt.source.instance.num_vertices -``` - -## Step 5: Register Display Name (if new problem) - -If this is a new problem not yet in the paper, add to the `display-name` dictionary near the top of `reductions.typ`: - -```typst -"ProblemName": [Display Name], -``` - -## Step 6: Build and Verify - -```bash -# Build the paper -make paper -``` - -### Verification Checklist - -- [ ] **Notation self-contained**: every symbol is defined before first use within the proof -- [ ] **Complexity cited**: reference exists, or footnote added for unverified claims -- [ ] **Overhead consistent**: prose dimensions match auto-derived overhead from JSON edge data -- [ ] **Example uses JSON data**: concrete values come from `load-example`/`load-results`, not hardcoded -- [ ] **Solution verified**: at least one solution checked end-to-end in the example -- [ ] **Witness semantics**: text treats `solutions.at(0)` as the canonical witness; any multiplicity claim is derived mathematically, not from fixture length -- [ ] **Pred commands present**: `pred-commands()` block at top of example with create/reduce/solve/evaluate pipeline -- [ ] **Paper compiles**: `make paper` succeeds without errors -- [ ] **Completeness check**: no new warnings about missing edges in the paper - -For simpler reductions, see MinimumVertexCover ↔ MaximumIndependentSet as a minimal example. diff --git a/.github/workflows/community-call.yml b/.github/workflows/community-call.yml index cd5fe64cc..2362f53f9 100644 --- a/.github/workflows/community-call.yml +++ b/.github/workflows/community-call.yml @@ -9,69 +9,48 @@ on: jobs: post-reminder: runs-on: ubuntu-latest + permissions: + contents: read + issues: read + pull-requests: read + checks: read + statuses: read steps: - - name: Collect agenda from project board + - name: Collect agenda from open PRs and ready issues id: agenda env: - GH_TOKEN: ${{ secrets.PROJECT_READ_TOKEN }} + GH_TOKEN: ${{ github.token }} run: | - # Fetch issues in "Review pool", "Final review", and "Ready" columns - QUERY='query { - node(id: "PVT_kwDOBrtarc4BRNVy") { - ... on ProjectV2 { - items(first: 50) { - nodes { - fieldValueByName(name: "Status") { - ... on ProjectV2ItemFieldSingleSelectValue { name } - } - content { - ... on Issue { - title number url - labels(first: 5) { nodes { name } } - author { login } - } - ... on PullRequest { - title number url - labels(first: 5) { nodes { name } } - author { login } - } - } - } - } - } - } - }' + fmt='"- [#\(.number)](\(.url)) \(.title) (by @\(.author.login))"' + PRS=$(gh pr list --repo "$GITHUB_REPOSITORY" --state open --limit 100 \ + --json number,title,url,author,statusCheckRollup) + # A PR is green when it has checks and every check concluded successfully (or was skipped). + green='(.statusCheckRollup | length > 0) and all(.statusCheckRollup[]; (.conclusion // .state) as $c | $c == "SUCCESS" or $c == "SKIPPED" or $c == "NEUTRAL")' + GREEN=$(echo "$PRS" | jq -r ".[] | select($green) | $fmt") + WORK=$(echo "$PRS" | jq -r ".[] | select(($green) | not) | $fmt") + READY=$(gh issue list --repo "$GITHUB_REPOSITORY" --state open --label Good \ + --search "-linked:pr" --limit 100 --json number,title,url,author --jq ".[] | $fmt") - RESULT=$(gh api graphql -f query="$QUERY") - - # Build agenda sections - build_section() { - local status="$1" - local header="$2" - local items - items=$(echo "$RESULT" | jq -r --arg s "$status" ' - .data.node.items.nodes[] - | select(.fieldValueByName.name == $s) - | select(.content != null) - | "- [#\(.content.number)](\(.content.url)) \(.content.title) (by @\(.content.author.login))" - ') - if [ -n "$items" ]; then - echo "**$header:**" - echo "$items" + section() { + if [ -n "$2" ]; then + echo "**$1:**" + echo "$2" echo "" fi } { echo 'AGENDA<<EOF' - build_section "OnHold" "On hold" + section "PRs ready for maintainer review (CI green)" "$GREEN" + section "PRs needing work" "$WORK" + section "Ready issues without a PR" "$READY" echo 'EOF' } >> "$GITHUB_OUTPUT" - name: Collect closed issues from last week id: closed env: - GH_TOKEN: ${{ secrets.PROJECT_READ_TOKEN }} + GH_TOKEN: ${{ github.token }} run: | SINCE=$(date -u -d '7 days ago' +%Y-%m-%dT%H:%M:%SZ) ITEMS=$(gh issue list --repo "$GITHUB_REPOSITORY" --state closed \ diff --git a/Makefile b/Makefile index 3238e53de..6b0a5db5c 100644 --- a/Makefile +++ b/Makefile @@ -1,10 +1,7 @@ # Makefile for problemreductions -.PHONY: help build test bench mcp-test fmt clippy doc mdbook website paper paper-data clean coverage rust-export compare qubo-testdata export-schemas release run-plan run-issue run-pipeline run-pipeline-forever run-review run-review-forever board-next board-claim board-ack board-move issue-context issue-guards pr-context pr-wait-ci worktree-issue worktree-pr diagrams jl-testdata cli cli-demo copilot-review papers papers-lookup papers-download papers-scihub papers-status papers-push papers-pull papers-index +.PHONY: help build test bench mcp-test fmt clippy doc mdbook website paper paper-data clean coverage rust-export compare qubo-testdata export-schemas release diagrams jl-testdata cli cli-demo copilot-review papers papers-lookup papers-download papers-scihub papers-status papers-push papers-pull papers-index -RUNNER ?= codex -CLAUDE_MODEL ?= opus -CODEX_MODEL ?= gpt-5.4 TEST_FEATURES := example-db # Cross-platform sed in-place: macOS needs -i '', Linux needs -i @@ -36,22 +33,6 @@ help: @echo " release V=x.y.z - Tag and push a new release (triggers CI publish)" @echo " cli - Build the pred CLI tool" @echo " cli-demo - Run closed-loop CLI demo (build + exercise all commands)" - @echo " run-plan - Execute a plan with Codex or Claude (latest plan in docs/plans/)" - @echo " run-issue N=<number> - Run issue-to-pr --execute for a GitHub issue" - @echo " run-pipeline [N=<number>] - Pick a Ready issue, implement, move to Review pool" - @echo " run-pipeline-forever - Loop: drain eligible Ready issues forever; poll only when the queue is empty" - @echo " run-review [N=<number>] - Pick PR from Review pool, fix comments/CI, run agentic tests" - @echo " run-review-forever - Loop: drain eligible Review pool PRs forever; poll only when the queue is empty" - @echo " board-next MODE=<ready|review|final-review> [NUMBER=<n>] [FORMAT=text|json] - Get the next eligible queued project item" - @echo " board-claim MODE=<ready|review> [NUMBER=<n>] [FORMAT=text|json] - Claim and move the next eligible queued project item" - @echo " board-ack MODE=<ready|review|final-review> ITEM=<id> - Acknowledge a queued project item" - @echo " board-move ITEM=<id> STATUS=<status> - Move a project item to a named status" - @echo " issue-context ISSUE=<number> [REPO=<owner/repo>] - Fetch structured issue preflight JSON" - @echo " issue-guards ISSUE=<number> [REPO=<owner/repo>] - Backward-compatible alias for issue-context" - @echo " pr-context PR=<number> [REPO=<owner/repo>] - Fetch structured PR snapshot JSON" - @echo " pr-wait-ci PR=<number> [REPO=<owner/repo>] - Poll CI until terminal state and print JSON" - @echo " worktree-issue ISSUE=<number> SLUG=<slug> - Create an issue worktree from origin/main" - @echo " worktree-pr PR=<number> [REPO=<owner/repo>] - Checkout a PR into an isolated worktree" @echo " copilot-review - Request Copilot code review on current PR" @echo "" @echo " papers - Full paper fetch: lookup + download + scihub" @@ -59,9 +40,6 @@ help: @echo " papers-download - Download available free PDFs" @echo " papers-scihub - Fetch remaining papers via Sci-Hub" @echo " papers-status - Show paper collection stats" - @echo "" - @echo " Set RUNNER=claude to use Claude instead of Codex (default: codex)" - @echo " Override CODEX_MODEL or CLAUDE_MODEL to pick a different model" # Build the project build: @@ -192,6 +170,10 @@ release: ifndef V $(error Usage: make release V=x.y.z) endif + @test "$$(git branch --show-current)" = main || { echo "release: must be on main"; exit 1; } + @test -z "$$(git status --porcelain)" || { echo "release: working tree not clean"; exit 1; } + @git fetch -q origin main && test "$$(git rev-parse HEAD)" = "$$(git rev-parse origin/main)" || { echo "release: main is not up to date with origin/main"; exit 1; } + $(MAKE) check @echo "Releasing v$(V)..." $(SED_I) 's/^version = ".*"/version = "$(V)"/' Cargo.toml $(SED_I) 's/^version = ".*"/version = "$(V)"/' problemreductions-macros/Cargo.toml @@ -246,44 +228,6 @@ compare: rust-export echo "Julia: $$julia"; echo "Rust: $$rust"; test "$$julia" = "$$rust" || exit 1; \ done -# Run a plan with Codex or Claude -# Usage: make run-plan [INSTRUCTIONS="..."] [OUTPUT=output.log] [AGENT_TYPE=<codex|claude>] -# PLAN_FILE defaults to the most recently modified file in docs/plans/ -INSTRUCTIONS ?= -OUTPUT ?= run-plan-output.log -AGENT_TYPE ?= $(RUNNER) -PLAN_FILE ?= $(shell ls -t docs/plans/*.md 2>/dev/null | head -1) - -run-plan: - @. scripts/make_helpers.sh; \ - NL=$$'\n'; \ - BRANCH=$$(git branch --show-current); \ - PLAN_FILE="$(PLAN_FILE)"; \ - if [ "$(AGENT_TYPE)" = "claude" ]; then \ - PROCESS="1. Read the plan file$${NL}2. Execute the plan — it specifies which skill(s) to use$${NL}3. Push: git push origin $$BRANCH$${NL}4. If a PR already exists for this branch, skip. Otherwise create one."; \ - else \ - PROCESS="1. Read the plan file$${NL}2. If the plan references repo-local workflow docs under .claude/skills/*/SKILL.md, open and follow them directly. Treat slash-command names as aliases for those files.$${NL}3. Execute the tasks step by step. For each task, implement and test before moving on.$${NL}4. Push: git push origin $$BRANCH$${NL}5. If a PR already exists for this branch, skip. Otherwise create one."; \ - fi; \ - PROMPT="Execute the plan in '$$PLAN_FILE'."; \ - if [ "$(AGENT_TYPE)" != "claude" ]; then \ - PROMPT="$${PROMPT}$${NL}$${NL}Repo-local skills live in .claude/skills/*/SKILL.md. Treat any slash-command references in the plan as aliases for those skill files."; \ - fi; \ - if [ -n "$(INSTRUCTIONS)" ]; then \ - PROMPT="$${PROMPT}$${NL}$${NL}## Additional Instructions$${NL}$(INSTRUCTIONS)"; \ - fi; \ - PROMPT="$${PROMPT}$${NL}$${NL}## Process$${NL}$${PROCESS}$${NL}$${NL}## Rules$${NL}- Tests should be strong enough to catch regressions.$${NL}- Do not modify tests to make them pass.$${NL}- Test failure must be reported."; \ - echo "=== Prompt ===" && echo "$$PROMPT" && echo "===" ; \ - RUNNER="$(AGENT_TYPE)" run_agent "$(OUTPUT)" "$$PROMPT" - -# Run issue-to-pr --execute for a GitHub issue -# Usage: make run-issue N=42 -N ?= -run-issue: - @. scripts/make_helpers.sh; \ - if [ -z "$(N)" ]; then echo "Usage: make run-issue N=<issue-number>"; exit 1; fi; \ - PROMPT=$$(skill_prompt issue-to-pr "/issue-to-pr $(N) --execute" "process GitHub issue $(N) with --execute behavior"); \ - run_agent "issue-$(N)-output.log" "$$PROMPT" - # Closed-loop CLI demo: exercises all commands end-to-end PRED := cargo run -p problemreductions-cli --release -- CLI_DEMO_DIR := /tmp/pred-cli-demo @@ -419,231 +363,6 @@ cli-demo: cli echo "=== Demo complete: $$(ls $(CLI_DEMO_DIR)/*.json | wc -l | tr -d ' ') JSON files in $(CLI_DEMO_DIR) ===" @echo "=== All 20 steps passed ✅ ===" -# Run project-pipeline: pick a Ready issue, implement, move to Review pool -# Usage: make run-pipeline (picks next Ready issue automatically) -# make run-pipeline N=97 (processes specific issue) -run-pipeline: - @. scripts/make_helpers.sh; \ - selection=""; \ - if [ -n "$(N)" ]; then \ - issue="$(N)"; \ - if [ -n "$$WATCH_MODE" ]; then \ - status=0; \ - tmp_state=$$(mktemp); \ - selection=$$(board_next_json ready "" "$(N)" "$$tmp_state") || status=$$?; \ - rm -f "$$tmp_state"; \ - if [ "$$status" -eq 1 ]; then \ - echo "Ready issue #$(N) is no longer eligible."; \ - watch_emit_outcome gone; \ - exit 0; \ - elif [ "$$status" -ne 0 ]; then \ - exit "$$status"; \ - fi; \ - issue=$$(printf '%s\n' "$$selection" | python3 -c "import sys,json; data=json.load(sys.stdin); print(data['issue_number'] or data['number'])"); \ - fi; \ - else \ - status=0; \ - tmp_state=$$(mktemp); \ - selection=$$(board_next_json ready "" "" "$$tmp_state") || status=$$?; \ - rm -f "$$tmp_state"; \ - if [ "$$status" -eq 1 ]; then \ - echo "No Ready issues are currently eligible."; \ - exit 1; \ - elif [ "$$status" -ne 0 ]; then \ - exit "$$status"; \ - fi; \ - issue=$$(printf '%s\n' "$$selection" | python3 -c "import sys,json; data=json.load(sys.stdin); print(data['issue_number'] or data['number'])"); \ - fi; \ - PROMPT=$$(skill_prompt_with_context project-pipeline "/project-pipeline $$issue" "process GitHub issue $$issue" "Selected queue item" "$$selection"); \ - run_agent_with_watch_outcome "pipeline-output.log" "$$PROMPT" - -# Drain Ready issues forever, polling only when the eligible queue is empty -# Checks every 30 minutes while idle; successful dispatches immediately re-check the queue -run-pipeline-forever: - @. scripts/make_helpers.sh; \ - MAKE=$(MAKE) watch_and_dispatch ready run-pipeline "Ready issues" - -# Get the next eligible board item from the scripted queue logic -# Usage: make board-next MODE=ready -# make board-next MODE=review REPO=CodingThrust/problem-reductions -# make board-next MODE=final-review REPO=CodingThrust/problem-reductions -# make board-next MODE=review REPO=CodingThrust/problem-reductions NUMBER=570 FORMAT=json -# STATE_FILE=/tmp/custom.json make board-next MODE=ready -board-next: - @if [ -z "$(MODE)" ]; then \ - echo "MODE=ready|review|final-review is required"; \ - exit 2; \ - fi - @. scripts/make_helpers.sh; \ - state_file=$${STATE_FILE:-/tmp/problemreductions-$(MODE)-state.json}; \ - case "$(MODE)" in \ - review|final-review) \ - repo=$${REPO:-$$(gh repo view --json nameWithOwner --jq .nameWithOwner)}; \ - poll_project_items "$(MODE)" "$$state_file" "$$repo" "$(NUMBER)" "$(if $(FORMAT),$(FORMAT),text)"; \ - ;; \ - *) \ - poll_project_items "$(MODE)" "$$state_file" "" "$(NUMBER)" "$(if $(FORMAT),$(FORMAT),text)"; \ - ;; \ - esac - -# Claim and move the next eligible board item through the scripted queue logic -# Usage: make board-claim MODE=ready -# make board-claim MODE=review REPO=CodingThrust/problem-reductions -# make board-claim MODE=review REPO=CodingThrust/problem-reductions NUMBER=570 FORMAT=json -# STATE_FILE=/tmp/custom.json make board-claim MODE=ready -board-claim: - @if [ -z "$(MODE)" ]; then \ - echo "MODE=ready|review is required"; \ - exit 2; \ - fi - @. scripts/make_helpers.sh; \ - state_file=$${STATE_FILE:-/tmp/problemreductions-$(MODE)-state.json}; \ - case "$(MODE)" in \ - review) \ - repo=$${REPO:-$$(gh repo view --json nameWithOwner --jq .nameWithOwner)}; \ - claim_project_items "$(MODE)" "$$state_file" "$$repo" "$(NUMBER)" "$(if $(FORMAT),$(FORMAT),json)"; \ - ;; \ - ready) \ - claim_project_items "$(MODE)" "$$state_file" "" "$(NUMBER)" "$(if $(FORMAT),$(FORMAT),json)"; \ - ;; \ - *) \ - echo "MODE=ready|review is required"; \ - exit 2; \ - ;; \ - esac - -# Advance a scripted board queue after an item is processed -# Usage: make board-ack MODE=ready ITEM=PVTI_xxx -# STATE_FILE=/tmp/custom.json make board-ack MODE=review ITEM=PVTI_xxx -# STATE_FILE=/tmp/custom.json make board-ack MODE=final-review ITEM=PVTI_xxx -board-ack: - @if [ -z "$(MODE)" ] || [ -z "$(ITEM)" ]; then \ - echo "MODE=ready|review|final-review and ITEM=<project-item-id> are required"; \ - exit 2; \ - fi - @. scripts/make_helpers.sh; \ - state_file=$${STATE_FILE:-/tmp/problemreductions-$(MODE)-state.json}; \ - ack_polled_item "$$state_file" "$(ITEM)" - -# Move a project board item to a named status through the shared board script -# Usage: make board-move ITEM=PVTI_xxx STATUS=under-review -board-move: - @if [ -z "$(ITEM)" ] || [ -z "$(STATUS)" ]; then \ - echo "ITEM=<project-item-id> and STATUS=<backlog|ready|in-progress|review-pool|under-review|final-review|on-hold|done> are required"; \ - exit 2; \ - fi - @. scripts/make_helpers.sh; \ - move_board_item "$(ITEM)" "$(STATUS)" - -# Fetch deterministic issue preflight JSON for issue-to-pr -# Usage: make issue-context ISSUE=117 -# make issue-context ISSUE=117 REPO=CodingThrust/problem-reductions -issue-context: - @if [ -z "$(ISSUE)" ]; then \ - echo "ISSUE=<number> is required"; \ - exit 2; \ - fi - @. scripts/make_helpers.sh; \ - repo=$${REPO:-CodingThrust/problem-reductions}; \ - issue_context "$$repo" "$(ISSUE)" - -# Fetch deterministic issue preflight JSON for issue-to-pr -# Usage: make issue-guards ISSUE=117 -# make issue-guards ISSUE=117 REPO=CodingThrust/problem-reductions -issue-guards: - @if [ -z "$(ISSUE)" ]; then \ - echo "ISSUE=<number> is required"; \ - exit 2; \ - fi - @. scripts/make_helpers.sh; \ - repo=$${REPO:-CodingThrust/problem-reductions}; \ - issue_guards "$$repo" "$(ISSUE)" - -# Fetch structured PR snapshot JSON from the shared helper -# Usage: make pr-context PR=570 -# make pr-context PR=570 REPO=CodingThrust/problem-reductions -pr-context: - @if [ -z "$(PR)" ]; then \ - echo "PR=<number> is required"; \ - exit 2; \ - fi - @. scripts/make_helpers.sh; \ - repo=$${REPO:-$$(gh repo view --json nameWithOwner --jq .nameWithOwner)}; \ - pr_snapshot "$$repo" "$(PR)" - -# Poll CI for a PR until it reaches a terminal state -# Usage: make pr-wait-ci PR=570 -# make pr-wait-ci PR=570 TIMEOUT=1200 INTERVAL=15 -pr-wait-ci: - @if [ -z "$(PR)" ]; then \ - echo "PR=<number> is required"; \ - exit 2; \ - fi - @. scripts/make_helpers.sh; \ - repo=$${REPO:-$$(gh repo view --json nameWithOwner --jq .nameWithOwner)}; \ - timeout=$${TIMEOUT:-900}; \ - interval=$${INTERVAL:-30}; \ - pr_wait_ci "$$repo" "$(PR)" "$$timeout" "$$interval" - -# Create an issue worktree from origin/main -# Usage: make worktree-issue ISSUE=117 SLUG=graph-partitioning -worktree-issue: - @if [ -z "$(ISSUE)" ] || [ -z "$(SLUG)" ]; then \ - echo "ISSUE=<number> and SLUG=<slug> are required"; \ - exit 2; \ - fi - @. scripts/make_helpers.sh; \ - base=$${BASE:-origin/main}; \ - create_issue_worktree "$(ISSUE)" "$(SLUG)" "$$base" - -# Checkout a PR into an isolated worktree -# Usage: make worktree-pr PR=570 -# make worktree-pr PR=570 REPO=CodingThrust/problem-reductions -worktree-pr: - @if [ -z "$(PR)" ]; then \ - echo "PR=<number> is required"; \ - exit 2; \ - fi - @. scripts/make_helpers.sh; \ - repo=$${REPO:-$$(gh repo view --json nameWithOwner --jq .nameWithOwner)}; \ - checkout_pr_worktree "$$repo" "$(PR)" - -# Usage: make run-review (picks next Review pool PR automatically) -# make run-review N=570 (processes specific PR) -# RUNNER=claude make run-review (use Claude instead of Codex) -run-review: - @. scripts/make_helpers.sh; \ - repo=$${REPO:-$$(gh repo view --json nameWithOwner --jq .nameWithOwner)}; \ - pr="$(N)"; \ - selection=$$(review_pipeline_context "$$repo" "$$pr"); \ - status_name=$$(printf '%s\n' "$$selection" | python3 -c "import sys,json; print(json.load(sys.stdin)['status'])"); \ - if [ "$$status_name" = "empty" ]; then \ - if [ -n "$$WATCH_MODE" ]; then \ - echo "Review item for PR #$$pr is no longer eligible."; \ - watch_emit_outcome gone; \ - exit 0; \ - fi; \ - echo "No Review pool PRs are currently eligible."; \ - exit 1; \ - fi; \ - if [ "$$status_name" = "ready" ]; then \ - pr=$$(printf '%s\n' "$$selection" | python3 -c "import sys,json; print(json.load(sys.stdin)['selection']['pr_number'])"); \ - slash_cmd="/review-pipeline $$pr"; \ - codex_desc="process PR #$$pr"; \ - else \ - slash_cmd="/review-pipeline"; \ - codex_desc="inspect the review pipeline bundle and resolve the next action"; \ - fi; \ - PROMPT=$$(skill_prompt_with_context review-pipeline "$$slash_cmd" "$$codex_desc" "Review pipeline context" "$$selection"); \ - run_agent_with_watch_outcome "review-output.log" "$$PROMPT" - -# Drain Review pool PRs forever, polling only when the eligible queue is empty -run-review-forever: - @. scripts/make_helpers.sh; \ - REPO=$$(gh repo view --json nameWithOwner --jq .nameWithOwner) || { echo "Failed to detect repo (gh repo view failed)"; exit 1; }; \ - if [ -z "$$REPO" ]; then echo "Failed to detect repo (empty result)"; exit 1; fi; \ - MAKE=$(MAKE) watch_and_dispatch review run-review "Review pool PRs" "$$REPO" - # Request Copilot code review on the current PR # Requires: gh extension install ChrisCarini/gh-copilot-review copilot-review: diff --git a/docs/agent-profiles/SKILLS.md b/docs/agent-profiles/SKILLS.md index df3ffde16..90a8e297e 100644 --- a/docs/agent-profiles/SKILLS.md +++ b/docs/agent-profiles/SKILLS.md @@ -14,20 +14,22 @@ Post-refactor extension points: - new model load/serialize/brute-force dispatch comes from `declare_variants!` in the model file, with an optional `default` - alias resolution lives in `problemreductions-cli/src/problem_name.rs` - `pred create` UX lives in `problemreductions-cli/src/commands/create.rs` -- model examples live in `src/example_db/model_builders.rs`; rule examples live beside their rules and are collected by `src/rules/mod.rs` - -- [issue-to-pr] — Convert a GitHub issue into a PR with an implementation plan -- [add-model] — Add a new problem model to the codebase -- [add-rule] — Add a new reduction rule to the codebase -- [review-implementation] — Review implementation completeness via parallel subagents -- [fix-pr] — Resolve PR review comments, CI failures, and coverage gaps -- [check-issue] — Quality gate for Rule and Model GitHub issues -- [topology-sanity-check] — Run sanity checks on the reduction graph: detect orphan problems and redundant rules -- [project-pipeline] — Pick the next ready issue, implement it, and move it through the project workflow -- [review-pipeline] — Process PRs in Review pool: fix comments, fix CI, run agentic review, move to Final review -- [propose] — Interactive brainstorming that turns a new model or rule idea into a GitHub issue -- [final-review] — Interactive maintainer review for PRs in the Final review column -- [dev-setup] — Install and configure the maintainer development environment -- [write-model-in-paper] — Write or improve a problem-def entry in the Typst paper -- [write-rule-in-paper] — Write or improve a reduction-rule entry in the Typst paper -- [release] — Create a new crate release with version bump +- model examples live beside each model in `canonical_model_example_specs()` (collected by `src/example_db/model_builders.rs`); rule examples live beside their rules and are collected by `src/rules/mod.rs` + +Guides (auto-invoked while working): + +- [how-to-code] — Implement or modify a problem model or reduction rule +- [how-to-verify] — Certify a reduction and check reduction-graph topology +- [how-to-write-manual] — Write or audit Typst manual entries and mdBook docs +- [how-to-review] — Fresh-context PR review with a `pred` feature test +- [how-to-triage-issue] — Quality-check and fix Model and Rule GitHub issues +- [how-to-ship] — Take an issue to a merge-ready pull request + +Tools (invoked on request): + +- [propose] — Turn a new model or rule idea into a GitHub issue +- [find-solver] — Match a real-world problem to a model, route, and solver +- [find-problem] — Find source problems a given solver handles +- [dev-setup] — Install and configure the development environment +- [release] — Guarded crate release with version bump +- [update-papers] — Refresh the research paper collection diff --git a/docs/paper/reductions.typ b/docs/paper/reductions.typ index 781280a3a..5c357942a 100644 --- a/docs/paper/reductions.typ +++ b/docs/paper/reductions.typ @@ -564,6 +564,14 @@ else { data.problem + "/" + data.variant.values().join("/") } } +// Keyed spec (Problem/key=value/...): positional values can be ambiguous (e.g. ILP/i64/bool) +#let keyed-spec(data) = { + data.problem + data.variant.pairs().map(((k, v)) => "/" + k + "=" + v).join() +} + +// `pred create --example` arguments reproducing a rule example's source instance +#let rule-spec(example) = keyed-spec(example.source) + " --to " + keyed-spec(example.target) + #let theorem = thmplain("theorem", [#h(-1.2em)Rule], base_level: 1) #let proof = thmproof("proof", "Proof") #let definition = thmbox( @@ -871,7 +879,7 @@ In all graph problems below, $G = (V, E)$ denotes an undirected graph with $|V| *Example.* Consider the Petersen graph $G$ with $n = #nv$ vertices, $|E| = #ne$ edges, and unit weights $w(v) = 1$ for all $v in V$. The graph is 3-regular (every vertex has degree 3). A maximum independent set is $S = {#S.map(i => $v_#i$).join(", ")}$ with $w(S) = sum_(v in S) w(v) = #alpha = #sym.alpha (G)$. No two vertices in $S$ share an edge, and no vertex can be added without violating independence. #pred-commands( - "pred create --example MIS -o mis.json", + "pred create --example " + problem-spec(x) + " -o mis.json", "pred solve mis.json", "pred evaluate mis.json --config " + cli-config(x.optimal_config), ) @@ -1405,7 +1413,7 @@ In all graph problems below, $G = (V, E)$ denotes an undirected graph with $|V| }).join("; "). The complement ${#complement.map(i => $v_#i$).join(", ")}$ is a maximum independent set ($#sym.alpha (G) = #alpha$, confirming $|"VC"| = n - alpha = #wS$). #pred-commands( - "pred create --example MVC -o mvc.json", + "pred create --example " + problem-spec(x) + " -o mvc.json", "pred solve mvc.json", "pred evaluate mvc.json --config " + cli-config(x.optimal_config), ) @@ -3663,7 +3671,7 @@ In all graph problems below, $G = (V, E)$ denotes an undirected graph with $|V| *Example.* Consider the house graph $G$ with $n = #nv$ vertices and $|E| = #ne$ edges. The triangle $K = {#K.map(i => $v_#i$).join(", ")}$ is a maximum clique of size $omega(G) = #omega$: all three pairs #clique-edges.map(((u, v)) => $(v_#u, v_#v)$).join(", ") are edges. No #(omega + 1)-clique exists because vertices $v_0$ and $v_1$ each have degree 2 and are not adjacent to all of ${#K.map(i => $v_#i$).join(", ")}$. #pred-commands( - "pred create --example MaximumClique -o maximum-clique.json", + "pred create --example " + problem-spec(x) + " -o maximum-clique.json", "pred solve maximum-clique.json", "pred evaluate maximum-clique.json --config " + cli-config(x.optimal_config), ) @@ -11315,7 +11323,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. example-caption: [$n = #max2sat_mc.source.instance.num_vars$ variables, $m = #max2sat_mc.source.instance.clauses.len()$ clauses, target has #max2sat_mc.target.instance.graph.num_vertices vertices and #max2sat_mc.target.instance.graph.edges.len() edges], extra: [ #pred-commands( - "pred create --example " + problem-spec(max2sat_mc.source) + " -o max2sat.json", + "pred create --example " + rule-spec(max2sat_mc) + " -o max2sat.json", "pred reduce max2sat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate max2sat.json --config " + cli-config(max2sat_mc_sol.source_config), @@ -11366,7 +11374,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. example-caption: [$n = #max2sat_ilp.source.instance.num_vars$ variables, $m = #max2sat_ilp.source.instance.clauses.len()$ clauses], extra: [ #pred-commands( - "pred create --example Maximum2Satisfiability -o max2sat.json", + "pred create --example " + rule-spec(max2sat_ilp) + " -o max2sat.json", "pred reduce max2sat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate max2sat.json --config " + cli-config(max2sat_ilp_sol.source_config), @@ -11528,7 +11536,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. example-caption: [Petersen graph ($n = 10$): VC $arrow.l.r$ IS], extra: [ #pred-commands( - "pred create --example MVC -o mvc.json", + "pred create --example " + rule-spec(mvc_mis) + " -o mvc.json", "pred reduce mvc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate mvc.json --config " + cli-config(mvc_mis_sol.source_config), @@ -11561,7 +11569,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. example-caption: [6-vertex source: two auxiliary centers encode the domination threshold], extra: [ #pred-commands( - "pred create --example " + problem-spec(dmds_mmmc.source) + " -o dmds.json", + "pred create --example " + rule-spec(dmds_mmmc) + " -o dmds.json", "pred reduce dmds.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate dmds.json --config " + cli-config(dmds_mmmc_sol.source_config), @@ -11596,7 +11604,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. example-caption: [6-vertex unit graph: dominating set of size 2 gives total distance 4], extra: [ #pred-commands( - "pred create --example " + problem-spec(dmds_msmc.source) + " -o dmds.json", + "pred create --example " + rule-spec(dmds_msmc) + " -o dmds.json", "pred reduce dmds.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate dmds.json --config " + cli-config(dmds_msmc_sol.source_config), @@ -11681,7 +11689,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. let fmt-seq(xs) = "(" + fmt-values(xs) + ")" [ #pred-commands( - "pred create --example " + problem-spec(mvc_lcs.source) + " -o mvc.json", + "pred create --example " + rule-spec(mvc_lcs) + " -o mvc.json", "pred reduce mvc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate mvc.json --config " + cli-config(mvc_lcs_sol.source_config), @@ -11722,7 +11730,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. example-caption: [7-vertex graph: each source edge becomes a directed 2-cycle], extra: [ #pred-commands( - "pred create --example MVC -o mvc.json", + "pred create --example " + rule-spec(mvc_fvs) + " -o mvc.json", "pred reduce mvc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate mvc.json --config " + cli-config(mvc_fvs_sol.source_config), @@ -11766,7 +11774,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. example-caption: [Path graph $P_5$: IS $arrow.r$ Clique via complement], extra: [ #pred-commands( - "pred create --example MIS -o mis.json", + "pred create --example " + rule-spec(mis_clique) + " -o mis.json", "pred reduce mis.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate mis.json --config " + cli-config(mis_clique_sol.source_config), @@ -11822,7 +11830,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. example-caption: [Path $P_4$: $n = 4$ vertices, $K = 2$ bound], extra: [ #pred-commands( - "pred create --example " + problem-spec(dmvc_cc.source) + " -o source.json", + "pred create --example " + rule-spec(dmvc_cc) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(dmvc_cc_sol.source_config), @@ -11848,7 +11856,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. example-caption: [Three-slot disjoint-union circuit on four atoms], extra: [ #pred-commands( - "pred create --example " + problem-spec(ec_ilp.source) + " -o ensemble.json", + "pred create --example " + rule-spec(ec_ilp) + " -o ensemble.json", "pred reduce ensemble.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ensemble.json --config " + cli-config(ec_ilp_sol.source_config), @@ -11903,7 +11911,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. example-caption: [Path $P_3$: vertex cover ${1}$ maps to an AND/OR graph of weight 5], extra: [ #pred-commands( - "pred create --example MVC -o mvc.json", + "pred create --example " + rule-spec(mvc_aog) + " -o mvc.json", "pred reduce mvc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate mvc.json --config " + cli-config(mvc_aog_sol.source_config), @@ -11964,7 +11972,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. example-caption: [10-spin Ising model on Petersen graph], extra: [ #pred-commands( - "pred create --example SpinGlass -o spinglass.json", + "pred create --example " + rule-spec(sg_qubo) + " -o spinglass.json", "pred reduce spinglass.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate spinglass.json --config " + cli-config(sg_qubo_sol.source_config), @@ -12014,7 +12022,7 @@ the displayed rule, extracted from the corresponding `pred path` entry. example-caption: [2D standard CVP with a coefficient box derived by the reduction], extra: [ #pred-commands( - "pred create --example CVP -o cvp.json", + "pred create --example " + rule-spec(cvp_qubo) + " -o cvp.json", "pred reduce cvp.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate cvp.json --config " + cli-config(cvp_qubo_sol.source_config), @@ -12060,7 +12068,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [House graph ($n = 5$, $|E| = 6$, $chi = 3$) with $k = 3$ colors], extra: [ #pred-commands( - "pred create --example " + problem-spec(kc_qubo.source) + " -o kcoloring.json", + "pred create --example " + rule-spec(kc_qubo) + " -o kcoloring.json", "pred reduce kcoloring.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate kcoloring.json --config " + cli-config(kc_qubo_sol.source_config), @@ -12145,7 +12153,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [3-SAT with $n = #ksat_qc.source.instance.num_vars$ variables and $m = #sat-num-clauses(ksat_qc.source.instance)$ clause mapped to a quadratic congruence with a $#ksat_qc.target.instance.b.len()$-digit modulus], extra: [ #pred-commands( - "pred create --example " + problem-spec(ksat_qc.source) + " -o ksat.json", + "pred create --example " + rule-spec(ksat_qc) + " -o ksat.json", "pred reduce ksat.json --via route.json -o bundle.json", "pred evaluate ksat.json --config " + cli-config(ksat_qc_sol.source_config), ) @@ -12217,7 +12225,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [3-SAT with 3 variables and 2 clauses], extra: [ #pred-commands( - "pred create --example " + problem-spec(ksat_ss.source) + " -o ksat.json", + "pred create --example " + rule-spec(ksat_ss) + " -o ksat.json", "pred reduce ksat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ksat.json --config " + cli-config(ksat_ss_sol.source_config), @@ -12258,7 +12266,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [#ss-cvp-n elements, target sum $B = #ss-cvp-target$], extra: [ #pred-commands( - "pred create --example SubsetSum -o subsetsum.json", + "pred create --example " + rule-spec(ss-cvp) + " -o subsetsum.json", "pred reduce subsetsum.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate subsetsum.json --config " + cli-config(ss-cvp-sol.source_config), @@ -12342,7 +12350,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [#part_ks_n elements, total sum $S = #part_ks_total$], extra: [ #pred-commands( - "pred create --example " + problem-spec(part_ks.source) + " -o partition.json", + "pred create --example " + rule-spec(part_ks) + " -o partition.json", "pred reduce partition.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate partition.json --config " + cli-config(part_ks_sol.source_config), @@ -12384,7 +12392,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [#part_ss_n elements, total sum $S = #part_ss_total$], extra: [ #pred-commands( - "pred create --example " + problem-spec(part_ss.source) + " -o partition.json", + "pred create --example " + rule-spec(part_ss) + " -o partition.json", "pred reduce partition.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate partition.json --config " + cli-config(part_ss_sol.source_config), @@ -12426,7 +12434,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [#part_ifwm_n elements, total sum $S = #part_ifwm_total$, bottleneck $R = #part_ifwm_half$], extra: [ #pred-commands( - "pred create --example " + problem-spec(part_ifwm.source) + " -o partition.json", + "pred create --example " + rule-spec(part_ifwm) + " -o partition.json", "pred reduce partition.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate partition.json --config " + cli-config(part_ifwm_sol.source_config), @@ -12469,7 +12477,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [$n = #ks_qubo_num_items$ items, capacity $C = #ks_qubo.source.instance.capacity$], extra: [ #pred-commands( - "pred create --example Knapsack -o knapsack.json", + "pred create --example " + rule-spec(ks_qubo) + " -o knapsack.json", "pred reduce knapsack.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate knapsack.json --config " + cli-config(ks_qubo_sol.source_config), @@ -12510,7 +12518,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [2-link worked example with 4 QUBO variables], extra: [ #pred-commands( - "pred create --example MinimumDiscretePlanarInverseKinematics -o ik.json", + "pred create --example " + rule-spec(mdpik_qubo) + " -o ik.json", "pred reduce ik.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ik.json --config " + cli-config(mdpik_qubo_sol.source_config), @@ -12565,7 +12573,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [$n = #mwc_qubo_n$ vertices, $k = #mwc_qubo_k$ terminals $T = {#fmt-values(mwc_qubo_terminals)}$, $|E| = #mwc_qubo_edges.len()$ edges], extra: [ #pred-commands( - "pred create --example MinimumMultiwayCut -o minimummultiwaycut.json", + "pred create --example " + rule-spec(mwc_qubo) + " -o minimummultiwaycut.json", "pred reduce minimummultiwaycut.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate minimummultiwaycut.json --config " + cli-config(mwc_qubo_sol.source_config), @@ -12631,7 +12639,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [4-variable QUBO with 3 quadratic terms], extra: [ #pred-commands( - "pred create --example QUBO/f64 -o qubo.json", + "pred create --example " + rule-spec(qubo_ilp) + " -o qubo.json", "pred reduce qubo.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate qubo.json --config " + cli-config(qubo_ilp_sol.source_config), @@ -12673,7 +12681,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [1-bit full adder to ILP], extra: [ #pred-commands( - "pred create --example CircuitSAT -o circuitsat.json", + "pred create --example " + rule-spec(cs_ilp) + " -o circuitsat.json", "pred reduce circuitsat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate circuitsat.json --config " + cli-config(cs_ilp_sol.source_config), @@ -12729,7 +12737,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [3-SAT with 5 variables and 7 clauses], extra: [ #pred-commands( - "pred create --example SAT -o sat.json", + "pred create --example " + rule-spec(sat_mis) + " -o sat.json", "pred reduce sat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate sat.json --config " + cli-config(sat_mis_sol.source_config), @@ -12761,7 +12769,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [5-variable SAT with 3 unit clauses to 3-coloring], extra: [ #pred-commands( - "pred create --example SAT -o sat.json", + "pred create --example " + rule-spec(sat_kc) + " -o sat.json", "pred reduce sat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate sat.json --config " + cli-config(sat_kc_sol.source_config), @@ -12789,7 +12797,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [5-variable 7-clause 3-SAT to dominating set], extra: [ #pred-commands( - "pred create --example SAT -o sat.json", + "pred create --example " + rule-spec(sat_ds) + " -o sat.json", "pred reduce sat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate sat.json --config " + cli-config(sat_ds_sol.source_config), @@ -12817,7 +12825,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [3-variable 4-clause SAT to equality-constrained integral flow], extra: [ #pred-commands( - "pred create --example " + problem-spec(sat_ifha.source) + " -o sat.json", + "pred create --example " + rule-spec(sat_ifha) + " -o sat.json", "pred reduce sat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate sat.json --config " + cli-config(sat_ifha_sol.source_config), @@ -12873,7 +12881,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [Mixed-size clauses (sizes 1 to 5) to 3-SAT], extra: [ #pred-commands( - "pred create --example SAT -o sat.json", + "pred create --example " + rule-spec(sat_ksat) + " -o sat.json", "pred reduce sat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate sat.json --config " + cli-config(sat_ksat_sol.source_config), @@ -12904,7 +12912,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [3-variable 2-clause SAT to MAX-2-SAT], extra: [ #pred-commands( - "pred create --example " + problem-spec(sat_max2sat.source) + " -o sat.json", + "pred create --example " + rule-spec(sat_max2sat) + " -o sat.json", "pred reduce sat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate sat.json --config " + cli-config(sat_max2sat_sol.source_config), @@ -12971,7 +12979,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [3-variable SAT formula to boolean circuit], extra: [ #pred-commands( - "pred create --example SAT -o sat.json", + "pred create --example " + rule-spec(sat_cs) + " -o sat.json", "pred reduce sat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate sat.json --config " + cli-config(sat_cs_sol.source_config), @@ -13004,7 +13012,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [Tseitin encoding of a circuit equation], extra: [ #pred-commands( - "pred create --example " + problem-spec(cs_sat.source) + " -o circuitsat.json", + "pred create --example " + rule-spec(cs_sat) + " -o circuitsat.json", "pred reduce circuitsat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate circuitsat.json --config " + cli-config(cs_sat_sol.source_config), @@ -13044,7 +13052,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [1-bit full adder to Ising model], extra: [ #pred-commands( - "pred create --example CircuitSAT -o circuitsat.json", + "pred create --example " + rule-spec(cs_sg) + " -o circuitsat.json", "pred reduce circuitsat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate circuitsat.json --config " + cli-config(cs_sg_sol.source_config), @@ -13097,7 +13105,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [Factor $N = #fact_cs.source.instance.target$], extra: [ #pred-commands( - "pred create --example Factoring -o factoring.json", + "pred create --example " + rule-spec(fact_cs) + " -o factoring.json", "pred reduce factoring.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate factoring.json --config " + cli-config(fact_cs_sol.source_config), @@ -13136,7 +13144,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [Petersen graph ($n = 10$, unit weights) to Ising], extra: [ #pred-commands( - "pred create --example MaxCut -o maxcut.json", + "pred create --example " + rule-spec(mc_sg) + " -o maxcut.json", "pred reduce maxcut.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate maxcut.json --config " + cli-config(mc_sg_sol.source_config), @@ -13162,7 +13170,7 @@ where $P$ is a penalty weight large enough that any constraint violation costs m example-caption: [10-spin Ising with alternating $J_(i j) in {plus.minus 1}$], extra: [ #pred-commands( - "pred create --example SpinGlass -o spinglass.json", + "pred create --example " + rule-spec(sg_mc) + " -o spinglass.json", "pred reduce spinglass.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate spinglass.json --config " + cli-config(sg_mc_sol.source_config), @@ -13342,7 +13350,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [DAG with #mfdts_ilp.source.instance.num_vertices vertices, #mfdts_ilp.source.instance.inputs.len() inputs, #mfdts_ilp.source.instance.outputs.len() outputs, and #(mfdts_ilp.source.instance.num_vertices - mfdts_ilp.source.instance.inputs.len() - mfdts_ilp.source.instance.outputs.len()) internal vertices], extra: [ #pred-commands( - "pred create --example " + problem-spec(mfdts_ilp.source) + " -o mfdts.json", + "pred create --example " + rule-spec(mfdts_ilp) + " -o mfdts.json", "pred reduce mfdts.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate mfdts.json --config " + cli-config(mfdts_ilp_sol.source_config), @@ -13418,7 +13426,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [3-cycle digraph: FVS of size 1 maps to an expression DAG needing 1 LOAD], extra: [ #pred-commands( - "pred create --example MinimumFeedbackVertexSet/One -o fvs.json", + "pred create --example " + rule-spec(fvs_cg) + " -o fvs.json", "pred reduce fvs.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate fvs.json --config " + cli-config(fvs_cg_sol.source_config), @@ -13466,7 +13474,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Weighted 5-cycle ($n = 5$), $k = 2$], extra: [ #pred-commands( - "pred create --example " + problem-spec(mckp_ilp.source) + " -o source.json", + "pred create --example " + rule-spec(mckp_ilp) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(mckp_ilp_sol.source_config), @@ -13500,7 +13508,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Two labelled 3-vertex digraphs with 2 arcs each], extra: [ #pred-commands( - "pred create --example " + problem-spec(mces_ilp.source) + " -o source.json", + "pred create --example " + rule-spec(mces_ilp) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(mces_ilp_sol.source_config), @@ -13538,7 +13546,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [$|V_1| = #cmo_ilp.source.instance.num_vertices_1$, $|E_1| = #cmo_ilp.source.instance.contacts_1.len()$, $|V_2| = #cmo_ilp.source.instance.num_vertices_2$, $|E_2| = #cmo_ilp.source.instance.contacts_2.len()$], extra: [ #pred-commands( - "pred create --example " + problem-spec(cmo_ilp.source) + " -o source.json", + "pred create --example " + rule-spec(cmo_ilp) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(cmo_ilp_sol.source_config), @@ -13578,7 +13586,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [$n = 4$ vertices, $m = 5$ edges, $k = 3$], extra: [ #pred-commands( - "pred create --example " + problem-spec(mewkc_ilp.source) + " -o source.json", + "pred create --example " + rule-spec(mewkc_ilp) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(mewkc_ilp_sol.source_config), @@ -13614,7 +13622,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [$n = #ks_ilp.source.instance.weights.len()$ items, capacity $C = #ks_ilp.source.instance.capacity$], extra: [ #pred-commands( - "pred create --example Knapsack -o knapsack.json", + "pred create --example " + rule-spec(ks_ilp) + " -o knapsack.json", "pred reduce knapsack.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate knapsack.json --config " + cli-config(ks_ilp_sol.source_config), @@ -13664,7 +13672,7 @@ The following reductions to Integer Linear Programming are straightforward formu let total_value = chosen.map(((i, c)) => c * values.at(i)).sum() [ #pred-commands( - "pred create --example " + problem-spec(ik_ilp.source) + " -o integer-knapsack.json", + "pred create --example " + rule-spec(ik_ilp) + " -o integer-knapsack.json", "pred reduce integer-knapsack.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate integer-knapsack.json --config " + cli-config(ik_ilp_sol.source_config), @@ -13719,7 +13727,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Path graph $P_4$: clique in $G$ maps to independent set in complement $overline(G)$.], extra: [ #pred-commands( - "pred create --example MaximumClique -o maximumclique.json", + "pred create --example " + rule-spec(clique_mis) + " -o maximumclique.json", "pred reduce maximumclique.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate maximumclique.json --config " + cli-config(clique_mis_sol.source_config), @@ -13796,7 +13804,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Path $P_4$: $n = 4$ vertices, $m = 3$ edges], extra: [ #pred-commands( - "pred create --example " + problem-spec(ola_seqmwct.source) + " -o source.json", + "pred create --example " + rule-spec(ola_seqmwct) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(ola_seqmwct_sol.source_config), @@ -13832,7 +13840,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [6-vertex, 7-edge graph: arrangement of length $11$ gives $4$ augmentations], extra: [ #pred-commands( - "pred create --example " + problem-spec(dola_c1ma.source) + " -o source.json", + "pred create --example " + rule-spec(dola_c1ma) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(dola_c1ma_sol.source_config), @@ -13898,7 +13906,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Cycle graph on $#hc_tsp_n$ vertices to weighted $K_#hc_tsp_n$], extra: [ #pred-commands( - "pred create --example " + problem-spec(hc_tsp.source) + " -o hc.json", + "pred create --example " + rule-spec(hc_tsp) + " -o hc.json", "pred reduce hc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate hc.json --config " + cli-config(hc_tsp_sol.source_config), @@ -13929,7 +13937,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Weighted $K_4$: the optimal tour $0 arrow 1 arrow 3 arrow 2 arrow 0$ with cost 80 is found by position-based ILP.], extra: [ #pred-commands( - "pred create --example TSP -o tsp.json", + "pred create --example " + rule-spec(tsp_ilp) + " -o tsp.json", "pred reduce tsp.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate tsp.json --config " + cli-config(tsp_ilp_sol.source_config), @@ -13975,7 +13983,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [The 3-vertex path $0 arrow 1 arrow 2$ encoded as a 7-variable ILP with optimum 5.], extra: [ #pred-commands( - "pred create --example LongestPath -o longest-path.json", + "pred create --example " + rule-spec(lp_ilp) + " -o longest-path.json", "pred reduce longest-path.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate longest-path.json --config " + cli-config(lp_ilp_sol.source_config), @@ -14020,7 +14028,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [TSP on $K_3$ with weights $w_(01) = 1$, $w_(02) = 2$, $w_(12) = 3$: the QUBO ground state encodes the optimal tour with cost $1 + 2 + 3 = 6$.], extra: [ #pred-commands( - "pred create --example TSP -o tsp.json", + "pred create --example " + rule-spec(tsp_qubo) + " -o tsp.json", "pred reduce tsp.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate tsp.json --config " + cli-config(tsp_qubo_sol.source_config), @@ -14059,7 +14067,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [LCS of two strings over a 3-symbol alphabet], extra: [ #pred-commands( - "pred create --example LCS -o lcs.json", + "pred create --example " + rule-spec(lcs_mis) + " -o lcs.json", "pred reduce lcs.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate lcs.json --config " + cli-config(lcs_mis_sol.source_config), @@ -14094,7 +14102,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Binary alphabet, 4 length-3 strings], extra: [ #pred-commands( - "pred create --example " + problem-spec(cs_ilp_str.source) + " -o source.json", + "pred create --example " + rule-spec(cs_ilp_str) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(cs_ilp_str_sol.source_config), @@ -14137,7 +14145,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Binary alphabet, 3 length-5 strings, length-3 windows], extra: [ #pred-commands( - "pred create --example " + problem-spec(css_ilp.source) + " -o source.json", + "pred create --example " + rule-spec(css_ilp) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(css_ilp_sol.source_config), @@ -14230,7 +14238,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Signed-weight exact tree formulation], extra: [ #pred-commands( - "pred create --example SteinerTree -o steinertree.json", + "pred create --example " + rule-spec(st_ilp) + " -o steinertree.json", "pred reduce steinertree.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate steinertree.json --config " + cli-config(st_ilp_sol.source_config), @@ -14257,7 +14265,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Unit-weight VC to Hitting Set ($n = #graph-num-vertices(mvc_hs.source.instance)$, $|E| = #graph-num-edges(mvc_hs.source.instance)$)], extra: [ #pred-commands( - "pred create --example 'MVC {weight: One}' -o mvc.json", + "pred create --example " + rule-spec(mvc_hs) + " -o mvc.json", "pred reduce mvc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate mvc.json --config " + cli-config(mvc_hs_sol.source_config), @@ -14352,7 +14360,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [$K_4$ with $n = #graph-num-vertices(mono_ilp.source.instance)$ vertices, $m = #graph-num-edges(mono_ilp.source.instance)$ edges, and $#mono_ilp.source.instance.triangles.len()$ triangles], extra: [ #pred-commands( - "pred create --example " + problem-spec(mono_ilp.source) + " -o monochromatic-triangle.json", + "pred create --example " + rule-spec(mono_ilp) + " -o monochromatic-triangle.json", "pred reduce monochromatic-triangle.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate monochromatic-triangle.json --config " + cli-config(mono_ilp_sol.source_config), @@ -14391,7 +14399,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [$|U| = #ss_bt.source.instance.universe_size$, $#ss_bt.source.instance.subsets.len()$ subsets, no normalization auxiliaries needed], extra: [ #pred-commands( - "pred create --example " + problem-spec(ss_bt.source) + " -o set-splitting.json", + "pred create --example " + rule-spec(ss_bt) + " -o set-splitting.json", "pred reduce set-splitting.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate set-splitting.json --config " + cli-config(ss_bt_sol.source_config), @@ -14466,7 +14474,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [#n\-vertex graph with $k = #k$: non-incidence gadget construction], extra: [ #pred-commands( - "pred create --example " + problem-spec(kc_bcbs.source) + " -o kclique.json", + "pred create --example " + rule-spec(kc_bcbs) + " -o kclique.json", "pred reduce kclique.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate kclique.json --config " + cli-config(kc_bcbs_sol.source_config), @@ -14531,7 +14539,7 @@ The following reductions to Integer Linear Programming are straightforward formu .map(((i, _)) => i) [ #pred-commands( - "pred create --example " + problem-spec(mmm_ach.source) + " -o mmm.json", + "pred create --example " + rule-spec(mmm_ach) + " -o mmm.json", "pred reduce mmm.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate mmm.json --config " + cli-config(mmm_ach_sol.source_config), @@ -14592,7 +14600,7 @@ The following reductions to Integer Linear Programming are straightforward formu .map(((i, _)) => i) [ #pred-commands( - "pred create --example " + problem-spec(mmm_mmd.source) + " -o mmm.json", + "pred create --example " + rule-spec(mmm_mmd) + " -o mmm.json", "pred reduce mmm.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate mmm.json --config " + cli-config(s-cfg), @@ -14875,7 +14883,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [6-vertex graph ($n = 6$, $q = 2$): two $P_3$ paths], extra: [ #pred-commands( - "pred create --example PartitionIntoPathsOfLength2 -o ppl2.json", + "pred create --example " + rule-spec(ppl2_bcsf) + " -o ppl2.json", "pred reduce ppl2.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ppl2.json --config " + cli-config(ppl2_bcsf_sol.source_config), @@ -15302,7 +15310,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [A bounded open-shop schedule], extra: [ #pred-commands( - "pred create --example " + problem-spec(doss_ilp.source) + " -o schedule.json", + "pred create --example " + rule-spec(doss_ilp) + " -o schedule.json", "pred reduce schedule.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate schedule.json --config " + cli-config(doss_ilp.solutions.at(0).source_config), @@ -15517,7 +15525,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Triangle plus pendant: $n = 4$ vertices, $m = 4$ edges], extra: [ #pred-commands( - "pred create --example " + problem-spec(hcd_ilp.source) + " -o source.json", + "pred create --example " + rule-spec(hcd_ilp) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(hcd_ilp_sol.source_config), @@ -15551,7 +15559,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [3-vertex digraph with 4 arcs (parallel edges)], extra: [ #pred-commands( - "pred create --example " + problem-spec(ep_ilp.source) + " -o source.json", + "pred create --example " + rule-spec(ep_ilp) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(ep_ilp_sol.source_config), @@ -15616,7 +15624,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Cycle graph on $#hc_lc_n$ vertices with unit edge lengths], extra: [ #pred-commands( - "pred create --example " + problem-spec(hc_lc.source) + " -o hc.json", + "pred create --example " + rule-spec(hc_lc) + " -o hc.json", "pred reduce hc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate hc.json --config " + cli-config(hc_lc_sol.source_config), @@ -15666,7 +15674,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [A circuit meeting a length bound], extra: [ #pred-commands( - "pred create --example " + problem-spec(dlc_ilp.source) + " -o circuit.json", + "pred create --example " + rule-spec(dlc_ilp) + " -o circuit.json", "pred reduce circuit.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate circuit.json --config " + cli-config(dlc_ilp.solutions.at(0).source_config), @@ -16206,7 +16214,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [4 cars, sequence length 8], extra: [ #pred-commands( - "pred create --example " + problem-spec(ps_qubo.source) + " -o paintshop.json", + "pred create --example " + rule-spec(ps_qubo) + " -o paintshop.json", "pred reduce paintshop.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate paintshop.json --config " + cli-config(ps_qubo_sol.source_config), @@ -16286,7 +16294,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Path graph $P_4$ ($n = 4$, $|E| = 3$, $K = 5$)], extra: [ #pred-commands( - "pred create --example " + problem-spec(rta_rtsa.source) + " -o rta.json", + "pred create --example " + rule-spec(rta_rtsa) + " -o rta.json", "pred reduce rta.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate rta.json --config " + cli-config(rta_rtsa_sol.source_config), @@ -16409,7 +16417,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [Diamond network: $n = 4$ vertices, $m = 5$ arcs, max flow $= 3$], extra: [ #pred-commands( - "pred create --example " + problem-spec(mcmf_mcc.source) + " -o source.json", + "pred create --example " + rule-spec(mcmf_mcc) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(mcmf_mcc_sol.source_config), @@ -16467,7 +16475,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [$n = #mlr_n$ items, $#mlr_nv$ pairwise variables, $#mlr_nc$ transitivity constraints], extra: [ #pred-commands( - "pred create --example MaximumLikelihoodRanking -o mlr.json", + "pred create --example " + rule-spec(mlr_ilp) + " -o mlr.json", "pred reduce mlr.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate mlr.json --config " + cli-config(mlr_ilp_sol.source_config), @@ -16511,7 +16519,7 @@ The following reductions to Integer Linear Programming are straightforward formu example-caption: [$K_#ocst_n$, #ocst_nv variables, #ocst_nc constraints], extra: [ #pred-commands( - "pred create --example OptimumCommunicationSpanningTree -o ocst.json", + "pred create --example " + rule-spec(ocst_ilp) + " -o ocst.json", "pred reduce ocst.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ocst.json --config " + cli-config(ocst_ilp_sol.source_config), @@ -16966,7 +16974,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Cycle $C_#hc_hp_n$ ($n = #hc_hp_n$): split $v_0$ into two copies with pendants], extra: [ #pred-commands( - "pred create --example " + problem-spec(hc_hp.source) + " -o hc.json", + "pred create --example " + rule-spec(hc_hp) + " -o hc.json", "pred reduce hc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate hc.json --config " + cli-config(hc_hp_sol.source_config), @@ -16997,7 +17005,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [5-vertex graph with $k = #kc_si.source.instance.k$: clique detection via subgraph isomorphism], extra: [ #pred-commands( - "pred create --example " + problem-spec(kc_si.source) + " -o kclique.json", + "pred create --example " + rule-spec(kc_si) + " -o kclique.json", "pred reduce kclique.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate kclique.json --config " + cli-config(kc_si_sol.source_config), @@ -17061,7 +17069,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [#part_mps_n elements, total sum $S = #part_mps_total$, deadline $D = #part_mps_deadline$], extra: [ #pred-commands( - "pred create --example " + problem-spec(part_mps.source) + " -o partition.json", + "pred create --example " + rule-spec(part_mps) + " -o partition.json", "pred reduce partition.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate partition.json --config " + cli-config(part_mps_sol.source_config), @@ -17101,7 +17109,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [#part_sosp_n elements, total sum $S = #part_sosp_total$, optimum $S^2 / 2 = #part_sosp_opt$], extra: [ #pred-commands( - "pred create --example " + problem-spec(part_sosp.source) + " -o partition.json", + "pred create --example " + rule-spec(part_sosp) + " -o partition.json", "pred reduce partition.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate partition.json --config " + cli-config(part_sosp_sol.source_config), @@ -17134,7 +17142,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #graph-num-vertices(hc_btsp.source.instance)$ vertices, $|E| = #graph-num-edges(hc_btsp.source.instance)$ edges: HC $arrow.r$ BTSP], extra: [ #pred-commands( - "pred create --example " + problem-spec(hc_btsp.source) + " -o hc.json", + "pred create --example " + rule-spec(hc_btsp) + " -o hc.json", "pred reduce hc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate hc.json --config " + cli-config(hc_btsp_sol.source_config), @@ -17165,7 +17173,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [5-vertex graph ($n = #kc_cbq.source.instance.graph.num_vertices$, $|E| = #kc_cbq.source.instance.graph.edges.len()$, $k = #kc_cbq.source.instance.k$)], extra: [ #pred-commands( - "pred create --example " + problem-spec(kc_cbq.source) + " -o kclique.json", + "pred create --example " + rule-spec(kc_cbq) + " -o kclique.json", "pred reduce kclique.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate kclique.json --config " + cli-config(kc_cbq_sol.source_config), @@ -17214,7 +17222,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$|X| = #x3c_ss.source.instance.universe_size$, $|cal(C)| = #x3c_ss.source.instance.subsets.len()$ subsets], extra: [ #pred-commands( - "pred create --example " + problem-spec(x3c_ss.source) + " -o x3c.json", + "pred create --example " + rule-spec(x3c_ss) + " -o x3c.json", "pred reduce x3c.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate x3c.json --config " + cli-config(x3c_ss_sol.source_config), @@ -17253,7 +17261,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [3-SAT with $n = #ksat_dmvc.source.instance.num_vars$ variables, $m = #sat-num-clauses(ksat_dmvc.source.instance)$ clauses reduced to Decision Minimum Vertex Cover with bound $k = #ksat_dmvc.target.instance.bound$], extra: [ #pred-commands( - "pred create --example " + problem-spec(ksat_dmvc.source) + " -o ksat.json", + "pred create --example " + rule-spec(ksat_dmvc) + " -o ksat.json", "pred reduce ksat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ksat.json --config " + cli-config(ksat_dmvc_sol.source_config), @@ -17289,7 +17297,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [3-vertex path with bound $k = #dmvc_hc.source.instance.bound$: one selector threads two cover-testing gadgets], extra: [ #pred-commands( - "pred create --example " + problem-spec(dmvc_hc.source) + " -o dmvc.json", + "pred create --example " + rule-spec(dmvc_hc) + " -o dmvc.json", "pred reduce dmvc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate dmvc.json --config " + cli-config(dmvc_hc_sol.source_config), @@ -17362,7 +17370,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Single clause ($n = #ksat_1in3.source.instance.num_vars$, $m = #sat-num-clauses(ksat_1in3.source.instance)$): 3-SAT $arrow.r$ 1-in-3 SAT], extra: [ #pred-commands( - "pred create --example " + problem-spec(ksat_1in3.source) + " -o ksat.json", + "pred create --example " + rule-spec(ksat_1in3) + " -o ksat.json", "pred reduce ksat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ksat.json --config " + cli-config(ksat_1in3_sol.source_config), @@ -17414,7 +17422,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Two-clause 3-SAT instance ($n = #ksat_d2cif.source.instance.num_vars$, $m = #sat-num-clauses(ksat_d2cif.source.instance)$) reduced to Directed Two-Commodity Integral Flow], extra: [ #pred-commands( - "pred create --example " + problem-spec(ksat_d2cif.source) + " -o ksat.json", + "pred create --example " + rule-spec(ksat_d2cif) + " -o ksat.json", "pred reduce ksat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ksat.json --config " + cli-config(ksat_d2cif_sol.source_config), @@ -17465,7 +17473,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Two-clause 3-SAT instance ($n = #ksat_rs.source.instance.num_vars$, $m = #sat-num-clauses(ksat_rs.source.instance)$) reduced to Register Sufficiency], extra: [ #pred-commands( - "pred create --example " + problem-spec(ksat_rs.source) + " -o ksat.json", + "pred create --example " + rule-spec(ksat_rs) + " -o ksat.json", "pred reduce ksat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ksat.json --config " + cli-config(ksat_rs_sol.source_config), @@ -17535,7 +17543,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Triangle graph ($n = #graph-num-vertices(mvc_mfas.source.instance)$, $|E| = #graph-num-edges(mvc_mfas.source.instance)$): VC $arrow.r$ FAS via vertex splitting], extra: [ #pred-commands( - "pred create --example " + problem-spec(mvc_mfas.source) + " -o mvc.json", + "pred create --example " + rule-spec(mvc_mfas) + " -o mvc.json", "pred reduce mvc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate mvc.json --config " + cli-config(mvc_mfas_sol.source_config), @@ -17579,7 +17587,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [3-SAT with $m = #ksat_kc.source.instance.clauses.len()$ clauses, $n = #ksat_kc.source.instance.num_vars$ variables $arrow.r$ $k$-clique on $#ksat_kc.target.instance.graph.num_vertices$ vertices], extra: [ #pred-commands( - "pred create --example " + problem-spec(ksat_kc.source) + " -o ksat.json", + "pred create --example " + rule-spec(ksat_kc) + " -o ksat.json", "pred reduce ksat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ksat.json --config " + cli-config(ksat_kc_sol.source_config), @@ -17616,7 +17624,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Single-clause 3-SAT reduced to #ksat_co.target.instance.num_elements cyclic-order elements], extra: [ #pred-commands( - "pred create --example " + problem-spec(ksat_co.source) + " -o ksat.json", + "pred create --example " + rule-spec(ksat_co) + " -o ksat.json", "pred reduce ksat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ksat.json --config " + cli-config(ksat_co_sol.source_config), @@ -17660,7 +17668,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [3-SAT with $n = #ksat_ps.source.instance.num_vars$ variables and $m = #ksat_ps.source.instance.clauses.len()$ clause $arrow.r$ unit-task preemptive schedule on #ksat_ps.target.instance.lengths.len() jobs and $#ksat_ps.target.instance.num_processors$ processors], extra: [ #pred-commands( - "pred create --example " + problem-spec(ksat_ps.source) + " -o ksat.json", + "pred create --example " + rule-spec(ksat_ps) + " -o ksat.json", "pred reduce ksat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ksat.json --config " + cli-config(ksat_ps_sol.source_config), @@ -17715,7 +17723,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Two-clause satisfiable formula reduced to a timetable gadget instance], extra: [ #pred-commands( - "pred create --example " + problem-spec(ksat_td.source) + " -o ksat.json", + "pred create --example " + rule-spec(ksat_td) + " -o ksat.json", "pred reduce ksat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ksat.json --config " + cli-config(ksat_td_sol.source_config), @@ -17786,7 +17794,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [4-cycle graph ($n = #graph-num-vertices(hc_bicon.source.instance)$, $|E| = #graph-num-edges(hc_bicon.source.instance)$): HC $arrow.r$ biconnectivity augmentation], extra: [ #pred-commands( - "pred create --example " + problem-spec(hc_bicon.source) + " -o hc.json", + "pred create --example " + rule-spec(hc_bicon) + " -o hc.json", "pred reduce hc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate hc.json --config " + cli-config(hc_bicon_sol.source_config), @@ -17828,7 +17836,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [4-cycle ($n = #hc_sca_n$): HC to budget-#hc_sca_n SCA], extra: [ #pred-commands( - "pred create --example " + problem-spec(hc_sca.source) + " -o hc.json", + "pred create --example " + rule-spec(hc_sca) + " -o hc.json", "pred reduce hc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate hc.json --config " + cli-config(hc_sca_sol.source_config), @@ -17863,7 +17871,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Cycle $C_#hc_sc_n$ ($n = #hc_sc_n$): vertex splitting to Stacker Crane], extra: [ #pred-commands( - "pred create --example " + problem-spec(hc_sc.source) + " -o hc.json", + "pred create --example " + rule-spec(hc_sc) + " -o hc.json", "pred reduce hc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate hc.json --config " + cli-config(hc_sc_sol.source_config), @@ -17899,7 +17907,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Cycle $C_#hc_rp_n$ ($n = #hc_rp_n$): vertex splitting to Rural Postman], extra: [ #pred-commands( - "pred create --example " + problem-spec(hc_rp.source) + " -o hc.json", + "pred create --example " + rule-spec(hc_rp) + " -o hc.json", "pred reduce hc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate hc.json --config " + cli-config(hc_rp_sol.source_config), @@ -17931,7 +17939,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [The three-vertex path has an independent set of size at least two], extra: [ #pred-commands( - "pred create --example DecisionMaximumIndependentSet/One -o independent-set.json", + "pred create --example " + rule-spec(mis_ifb) + " -o independent-set.json", "pred reduce independent-set.json --via route.json -o bundle.json", "pred solve bundle.json", "pred extract bundle.json --config " + cli-config(mis_ifb_sol.target_config), @@ -17959,7 +17967,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Cycle graph $C_#hc_qa.source.instance.graph.num_vertices$ ($n = #hc_qa.source.instance.graph.num_vertices$, $|E| = #hc_qa.source.instance.graph.edges.len()$)], extra: [ #pred-commands( - "pred create --example " + problem-spec(hc_qa.source) + " -o hc.json", + "pred create --example " + rule-spec(hc_qa) + " -o hc.json", "pred reduce hc.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate hc.json --config " + cli-config(hc_qa_sol.source_config), @@ -18008,7 +18016,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [#part_bp_n elements, total sum $S = #part_bp_total$], extra: [ #pred-commands( - "pred create --example " + problem-spec(part_bp.source) + " -o partition.json", + "pred create --example " + rule-spec(part_bp) + " -o partition.json", "pred reduce partition.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate partition.json --config " + cli-config(part_bp_sol.source_config), @@ -18039,7 +18047,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [#x3c_msp.source.instance.subsets.len() subsets over $3q = #x3c_msp.source.instance.universe_size$ elements], extra: [ #pred-commands( - "pred create --example " + problem-spec(x3c_msp.source) + " -o x3c.json", + "pred create --example " + rule-spec(x3c_msp) + " -o x3c.json", "pred reduce x3c.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate x3c.json --config " + cli-config(x3c_msp_sol.source_config), @@ -18079,7 +18087,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [#x3c_mfdts.source.instance.subsets.len() triples over $3q = #x3c_mfdts.source.instance.universe_size$ elements, with one shared output], extra: [ #pred-commands( - "pred create --example " + problem-spec(x3c_mfdts.source) + " -o x3c.json", + "pred create --example " + rule-spec(x3c_mfdts) + " -o x3c.json", "pred reduce x3c.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate x3c.json --config " + cli-config(x3c_mfdts_sol.source_config), @@ -18114,7 +18122,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [#x3c_mas.source.instance.subsets.len() triples over $3q = #x3c_mas.source.instance.universe_size$ elements, with decision bound $q = #(x3c_mas.source.instance.universe_size / 3)$], extra: [ #pred-commands( - "pred create --example " + problem-spec(x3c_mas.source) + " -o x3c.json", + "pred create --example " + rule-spec(x3c_mas) + " -o x3c.json", "pred reduce x3c.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate x3c.json --config " + cli-config(x3c_mas_sol.source_config), @@ -18163,7 +18171,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [#subsetsum-num-elements(ss_part.source.instance) elements, target $T = #ss_part.source.instance.target$], extra: [ #pred-commands( - "pred create --example " + problem-spec(ss_part.source) + " -o subsetsum.json", + "pred create --example " + rule-spec(ss_part) + " -o subsetsum.json", "pred reduce subsetsum.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate subsetsum.json --config " + cli-config(ss_part_sol.source_config), @@ -18216,7 +18224,7 @@ The following table shows concrete target-variable counts for example instances, let chosen_sum = chosen.map(i => sizes.at(i)).sum() [ #pred-commands( - "pred create --example " + problem-spec(ss_ik.source) + " -o subsetsum.json", + "pred create --example " + rule-spec(ss_ik) + " -o subsetsum.json", "pred solve subsetsum.json", "pred create --example " + problem-spec(ss_ik.target) + " -o integer-knapsack.json", "pred solve integer-knapsack.json", @@ -18264,7 +18272,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #sat_nt.source.instance.num_vars$ variables, $m = #sat-num-clauses(sat_nt.source.instance)$ clauses], extra: [ #pred-commands( - "pred create --example " + problem-spec(sat_nt.source) + " -o sat.json", + "pred create --example " + rule-spec(sat_nt) + " -o sat.json", "pred reduce sat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate sat.json --config " + cli-config(sat_nt_sol.source_config), @@ -18296,7 +18304,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #graph-num-vertices(kc_pic.source.instance)$ vertices, $k = #kc_pic.source.instance.num_colors$ colors], extra: [ #pred-commands( - "pred create --example " + problem-spec(kc_pic.source) + " -o kcoloring.json", + "pred create --example " + rule-spec(kc_pic) + " -o kcoloring.json", "pred reduce kcoloring.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate kcoloring.json --config " + cli-config(kc_pic_sol.source_config), @@ -18392,7 +18400,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [4 elements, $K = 2$, $B = 1$ $arrow.r$ ILP with #clustering_ilp.target.instance.variables.len() variables and #clustering_ilp.target.instance.constraints.len() constraints], extra: [ #pred-commands( - "pred create --example " + problem-spec(clustering_ilp.source) + " -o clustering.json", + "pred create --example " + rule-spec(clustering_ilp) + " -o clustering.json", "pred reduce clustering.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate clustering.json --config " + cli-config(clustering_ilp_sol.source_config), @@ -18437,7 +18445,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #graph-num-vertices(pic_mcbc.source.instance)$ vertices, $m = #graph-num-edges(pic_mcbc.source.instance)$ edges, $K = #pic_mcbc.source.instance.num_cliques$], extra: [ #pred-commands( - "pred create --example " + problem-spec(pic_mcbc.source) + " -o partition-into-cliques.json", + "pred create --example " + rule-spec(pic_mcbc) + " -o partition-into-cliques.json", "pred reduce partition-into-cliques.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate partition-into-cliques.json --config " + cli-config(pic_mcbc_sol.source_config), @@ -18472,7 +18480,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Triangle plus pendant: $n = 4$ vertices, $m = 4$ edges], extra: [ #pred-commands( - "pred create --example " + problem-spec(mcbc_migb.source) + " -o source.json", + "pred create --example " + rule-spec(mcbc_migb) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(mcbc_migb_sol.source_config), @@ -18499,7 +18507,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #ksat_ker.source.instance.num_vars$ variables, $m = #sat-num-clauses(ksat_ker.source.instance)$ clauses], extra: [ #pred-commands( - "pred create --example " + problem-spec(ksat_ker.source) + " -o ksat.json", + "pred create --example " + rule-spec(ksat_ker) + " -o ksat.json", "pred reduce ksat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ksat.json --config " + cli-config(ksat_ker_sol.source_config), @@ -18547,7 +18555,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #graph-num-vertices(hp_dcst.source.instance)$ vertices, $K = 2$], extra: [ #pred-commands( - "pred create --example " + problem-spec(hp_dcst.source) + " -o hampath.json", + "pred create --example " + rule-spec(hp_dcst) + " -o hampath.json", "pred reduce hampath.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate hampath.json --config " + cli-config(hp_dcst_sol.source_config), @@ -18579,7 +18587,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #nae_ss.source.instance.num_vars$ variables, $m = #sat-num-clauses(nae_ss.source.instance)$ clauses], extra: [ #pred-commands( - "pred create --example " + problem-spec(nae_ss.source) + " -o naesat.json", + "pred create --example " + rule-spec(nae_ss) + " -o naesat.json", "pred reduce naesat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate naesat.json --config " + cli-config(nae_ss_sol.source_config), @@ -18618,7 +18626,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #nae_ppm.source.instance.num_vars$ variables, $m = #sat-num-clauses(nae_ppm.source.instance)$ clauses, target $K = #nae_ppm.target.instance.num_matchings$], extra: [ #pred-commands( - "pred create --example " + problem-spec(nae_ppm.source) + " -o naesat.json", + "pred create --example " + rule-spec(nae_ppm) + " -o naesat.json", "pred reduce naesat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate naesat.json --config " + cli-config(nae_ppm_sol.source_config), @@ -18664,7 +18672,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$|U| = #x3c_sp.source.instance.universe_size$, $|cal(C)| = #x3c_sp.source.instance.subsets.len()$ subsets], extra: [ #pred-commands( - "pred create --example " + problem-spec(x3c_sp.source) + " -o x3c.json", + "pred create --example " + rule-spec(x3c_sp) + " -o x3c.json", "pred reduce x3c.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate x3c.json --config " + cli-config(x3c_sp_sol.source_config), @@ -18705,7 +18713,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$|U| = #x3c_bdst.source.instance.universe_size$, $|cal(C)| = #x3c_bdst.source.instance.subsets.len()$ subsets; target $D = #x3c_bdst.target.instance.diameter_bound$, $B = #x3c_bdst.target.instance.weight_bound$], extra: [ #pred-commands( - "pred create --example " + problem-spec(x3c_bdst.source) + " -o x3c.json", + "pred create --example " + rule-spec(x3c_bdst) + " -o x3c.json", "pred reduce x3c.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate x3c.json --config " + cli-config(x3c_bdst_sol.source_config), @@ -18754,7 +18762,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [#subsetsum-num-elements(ss_iem.source.instance) elements, target $B = #ss_iem.source.instance.target$], extra: [ #pred-commands( - "pred create --example " + problem-spec(ss_iem.source) + " -o subsetsum.json", + "pred create --example " + rule-spec(ss_iem) + " -o subsetsum.json", "pred reduce subsetsum.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate subsetsum.json --config " + cli-config(ss_iem_sol.source_config), @@ -18795,7 +18803,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #ksat_si.source.instance.num_vars$ variables, $m = #sat-num-clauses(ksat_si.source.instance)$ clauses], extra: [ #pred-commands( - "pred create --example " + problem-spec(ksat_si.source) + " -o ksat.json", + "pred create --example " + rule-spec(ksat_si) + " -o ksat.json", "pred reduce ksat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate ksat.json --config " + cli-config(ksat_si_sol.source_config), @@ -18834,7 +18842,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$m = 2$ triples, target sum $B = 15$], extra: [ #pred-commands( - "pred create --example " + problem-spec(n3dm_nmts.source) + " -o source.json", + "pred create --example " + rule-spec(n3dm_nmts) + " -o source.json", "pred reduce source.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate source.json --config " + cli-config(n3dm_nmts_sol.source_config), @@ -18859,7 +18867,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [#part_stw.source.instance.sizes.len() elements, total $= #part_stw.source.instance.sizes.sum()$], extra: [ #pred-commands( - "pred create --example " + problem-spec(part_stw.source) + " -o partition.json", + "pred create --example " + rule-spec(part_stw) + " -o partition.json", "pred reduce partition.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate partition.json --config " + cli-config(part_stw_sol.source_config), @@ -18916,7 +18924,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [#part_oss.source.instance.sizes.len() elements, $m = #part_oss.target.instance.inner.num_machines$ machines], extra: [ #pred-commands( - "pred create --example " + problem-spec(part_oss.source) + " -o partition.json", + "pred create --example " + rule-spec(part_oss) + " -o partition.json", "pred reduce partition.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate partition.json --config " + cli-config(part_oss_sol.source_config), @@ -18970,7 +18978,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #nae_mc.source.instance.num_vars$ variables, $m = #sat-num-clauses(nae_mc.source.instance)$ clauses, $M = #(sat-num-clauses(nae_mc.source.instance) + 1)$], extra: [ #pred-commands( - "pred create --example " + problem-spec(nae_mc.source) + " -o naesat.json", + "pred create --example " + rule-spec(nae_mc) + " -o naesat.json", "pred reduce naesat.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate naesat.json --config " + cli-config(nae_mc_sol.source_config), @@ -19022,7 +19030,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$q = #tdm_tp.source.instance.universe_size$, $t = #tdm_tp.source.instance.triples.len()$, target has #tdm_tp.target.instance.sizes.len() numbers], extra: [ #pred-commands( - "pred create --example " + problem-spec(tdm_tp.source) + " -o three-dimensional-matching.json", + "pred create --example " + rule-spec(tdm_tp) + " -o three-dimensional-matching.json", "pred reduce three-dimensional-matching.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate three-dimensional-matching.json --config " + cli-config(tdm_tp_sol.source_config), @@ -19092,7 +19100,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$q = #tdm_ilp.source.instance.universe_size$, $t = #tdm_ilp.source.instance.triples.len()$ triples $arrow.r$ ILP with #tdm_ilp.target.instance.variables.len() variables and #tdm_ilp.target.instance.constraints.len() constraints], extra: [ #pred-commands( - "pred create --example " + problem-spec(tdm_ilp.source) + " -o three-dimensional-matching.json", + "pred create --example " + rule-spec(tdm_ilp) + " -o three-dimensional-matching.json", "pred reduce three-dimensional-matching.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate three-dimensional-matching.json --config " + cli-config(tdm_ilp_sol.source_config), @@ -19137,7 +19145,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$q = #tdm_mwd.source.instance.universe_size$, $m = #tdm_mwd.source.instance.triples.len()$ triples $arrow.r$ #tdm_mwd.target.instance.matrix.len() $times$ #tdm_mwd.target.instance.matrix.at(0).len() parity-check matrix], extra: [ #pred-commands( - "pred create --example " + problem-spec(tdm_mwd.source) + " -o three-dimensional-matching.json", + "pred create --example " + rule-spec(tdm_mwd) + " -o three-dimensional-matching.json", "pred reduce three-dimensional-matching.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate three-dimensional-matching.json --config " + cli-config(tdm_mwd_sol.source_config), @@ -19184,7 +19192,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [#tp_rcs.source.instance.sizes.len() elements, $B = #tp_rcs.source.instance.bound$], extra: [ #pred-commands( - "pred create --example " + problem-spec(tp_rcs.source) + " -o threepartition.json", + "pred create --example " + rule-spec(tp_rcs) + " -o threepartition.json", "pred reduce threepartition.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate threepartition.json --config " + cli-config(tp_rcs_sol.source_config), @@ -19219,7 +19227,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [3-Partition with $3m = #tp_srd.source.instance.sizes.len()$ elements and $B = #tp_srd.source.instance.bound$ mapped to #tp_srd.target.instance.lengths.len() sequencing tasks], extra: [ #pred-commands( - "pred create --example " + problem-spec(tp_srd.source) + " -o tp.json", + "pred create --example " + rule-spec(tp_srd) + " -o tp.json", "pred reduce tp.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate tp.json --config " + cli-config(tp_srd_sol.source_config), @@ -19258,7 +19266,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Triangle graph ($n = #mc_mcbs.source.instance.graph.num_vertices$, $|E| = #mc_mcbs.source.instance.graph.edges.len()$, unit weights) mapped to $K_#mc_mcbs.target.instance.graph.num_vertices$], extra: [ #pred-commands( - "pred create --example " + problem-spec(mc_mcbs.source) + " -o maxcut.json", + "pred create --example " + rule-spec(mc_mcbs) + " -o maxcut.json", "pred reduce maxcut.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate maxcut.json --config " + cli-config(mc_mcbs_sol.source_config), @@ -19297,7 +19305,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Cycle $C_#mc_mmc_n$ (unit weights, $W = #mc_mmc_W$): adjacency matrix as quadratic form], extra: [ #pred-commands( - "pred create --example " + problem-spec(mc_mmc.source) + " -o maxcut.json", + "pred create --example " + rule-spec(mc_mmc) + " -o maxcut.json", "pred reduce maxcut.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate maxcut.json --config " + cli-config(mc_mmc_sol.source_config), @@ -19335,7 +19343,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #graph-num-vertices(hp_ist.source.instance)$ vertices, target tree $P_n$], extra: [ #pred-commands( - "pred create --example " + problem-spec(hp_ist.source) + " -o hampath.json", + "pred create --example " + rule-spec(hp_ist) + " -o hampath.json", "pred reduce hampath.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate hampath.json --config " + cli-config(hp_ist_sol.source_config), @@ -19367,7 +19375,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$|U| = #x3c_gf2.source.instance.universe_size$, $|cal(C)| = #x3c_gf2.source.instance.subsets.len()$ subsets], extra: [ #pred-commands( - "pred create --example " + problem-spec(x3c_gf2.source) + " -o x3c.json", + "pred create --example " + rule-spec(x3c_gf2) + " -o x3c.json", "pred reduce x3c.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate x3c.json --config " + cli-config(x3c_gf2_sol.source_config), @@ -19410,7 +19418,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [#part_pp.source.instance.sizes.len() elements, total $= #part_pp.source.instance.sizes.sum()$], extra: [ #pred-commands( - "pred create --example " + problem-spec(part_pp.source) + " -o partition.json", + "pred create --example " + rule-spec(part_pp) + " -o partition.json", "pred reduce partition.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate partition.json --config " + cli-config(part_pp_sol.source_config), @@ -19462,7 +19470,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #graph-num-vertices(hpbtv_lp.source.instance)$ vertices, $s = #hpbtv_lp.source.instance.source_vertex$, $t = #hpbtv_lp.source.instance.target_vertex$], extra: [ #pred-commands( - "pred create --example " + problem-spec(hpbtv_lp.source) + " -o hampath2v.json", + "pred create --example " + rule-spec(hpbtv_lp) + " -o hampath2v.json", "pred reduce hampath2v.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate hampath2v.json --config " + cli-config(hpbtv_lp_sol.source_config), @@ -19494,7 +19502,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [$n = #graph-num-vertices(gp_mc.source.instance)$ vertices, $|E| = #graph-num-edges(gp_mc.source.instance)$], extra: [ #pred-commands( - "pred create --example " + problem-spec(gp_mc.source) + " -o graphpart.json", + "pred create --example " + rule-spec(gp_mc) + " -o graphpart.json", "pred reduce graphpart.json --via route.json -o bundle.json", "pred solve bundle.json", "pred evaluate graphpart.json --config " + cli-config(gp_mc_sol.source_config), @@ -19552,7 +19560,7 @@ The following table shows concrete target-variable counts for example instances, example-caption: [Canonical PCSF $arrow$ Steiner Tree instance (path $0 - 1 - 2$, $n = #pcsf_st_n$, $m = #pcsf_st_m$, $k = #pcsf_st_k$ prized vertices)], extra: [ #pred-commands( - "pred create --example PrizeCollectingSteinerForest -o pcsf.json", + "pred create --example " + rule-spec(pcsf_st) + " -o pcsf.json", "pred reduce pcsf.json --via route.json -o bundle.json", "pred solve bundle.json", ) diff --git a/docs/src/skills.md b/docs/src/skills.md index c60b6665d..a6a10dcc4 100644 --- a/docs/src/skills.md +++ b/docs/src/skills.md @@ -41,42 +41,30 @@ Ask the agent to construct a small instance, solve it, and evaluate the recovere | Skill | What it produces | |---|---| | `propose` | Turns a definition or a candidate source-to-target connection into a precise proposal and files a GitHub issue. | -| `check-issue` | Quality gate for `[Model]` and `[Rule]` issues: usefulness, non-triviality, literature, and writing. Posts a report. | -| `fix-issue` | Fixes problems found by `check-issue`, then re-checks and moves the issue to Ready. | -| `issue-to-pr` | Converts an approved issue into a pull request with an implementation plan. | -| `add-model` | Adds a problem model: source, variants, tests, canonical example, and paper entry. | -| `add-rule` | Adds a reduction rule with the same artifacts, verified mathematically by default. | -| `verify-reduction` | Standalone verification of a rule: Typst proof, a constructor script, and an adversary script with thousands of independent checks. | -| `fix-pr` | Resolves review comments, CI failures, and coverage gaps on a pull request. | -| `write-model-in-paper`, `write-rule-in-paper` | Write or improve an entry in the Typst paper. | + +The remaining workflow has no queue or project board. Point an agent at an issue ("implement #42") and it follows these guides on its own: + +| Guide | Covers | +|---|---| +| `how-to-triage-issue` | Quality check of `[Model]` and `[Rule]` issues: usefulness, non-triviality, literature, completeness, writing. Fixes mechanical problems and discusses substantive ones. | +| `how-to-code` | Adding a problem model or reduction rule: source, variants, tests, canonical example. | +| `how-to-verify` | Certifying a rule: type check, Typst proof, and a constructor and an adversary script with thousands of independent checks, posted as a verification certificate on the pull request. Also reduction-graph topology checks. | +| `how-to-write-manual` | Writing or auditing entries in the Typst manual and these docs. | +| `how-to-review` | Fresh-context review of a pull request: structural check, quality check, and a feature test through `pred`. | +| `how-to-ship` | From issue to a merge-ready pull request: gates, review, CI, review comments, and coverage. A maintainer merges. | For a rule, distinguish a construction supported by literature from a new conjecture, and record proof gaps and counterexamples explicitly. +A passing test suite is evidence for the tested instances, not a proof for all inputs. Keep the mathematical argument, the constructor and adversary checks, closed-loop tests, and review findings with the work. + ## Maintain | Skill | What it produces | |---|---| -| `run-pipeline` | Takes one Ready issue from the project board through implementation to the Review pool. | -| `review-pipeline` | Agentic review of a pull request: structural check, quality check, and feature tests. Moves it to Final review. | -| `review-structural`, `review-quality` | The two read-only sub-reviews, usable on their own. | -| `final-review` | Interactive maintainer review, then merge or hold. | -| `auto-pipeline` | Chains the steps above from a Backlog issue to Final review. | -| `topology-sanity-check` | Detects isolated problems, missing NP-hardness chains from 3-SAT, and dominated rules. | -| `review-paper` | Reviews ten paper entries for mechanical and critical issues. | -| `release` | Determines the version bump, verifies tests, and tags a release. | +| `release` | Checks the tree, proposes the version bump, and tags a release after confirmation. | | `dev-setup` | Installs and configures the development tools. | | `update-papers` | Downloads referenced papers and regenerates the collection index. | -The corresponding make targets run in a configured maintainer checkout: - -```bash -make run-issue N=42 # implement one issue -make run-pipeline # pick the next Ready issue -make run-review N=570 # review one pull request -``` - -A passing test suite is evidence for the tested instances, not a proof for all inputs. Keep the mathematical argument, the constructor and adversary checks, closed-loop tests, and review findings with the work. - ## Reading without a browser The [Markdown index](markdown/index.md) lists every page of this guide with code includes expanded. [reduction_graph.json](reductions/reduction_graph.json) and [problem_schemas.json](reductions/problem_schemas.json) provide the registry as structured data. diff --git a/scripts/make_helpers.sh b/scripts/make_helpers.sh deleted file mode 100644 index 93eb77b50..000000000 --- a/scripts/make_helpers.sh +++ /dev/null @@ -1,322 +0,0 @@ -#!/usr/bin/env sh - -# --- Runner --- - -# Build a prompt for the given skill, adapting for claude vs codex. -# skill_prompt <skill-name> <claude-slash-command> [codex-description] -# Examples: -# skill_prompt project-pipeline "/project-pipeline" "pick and process the next Ready issue" -# skill_prompt project-pipeline "/project-pipeline 97" "process GitHub issue 97" -# skill_prompt review-pipeline "/review-pipeline" "pick and process the next Review pool PR" -skill_prompt() { - skill=$1 - slash_cmd=$2 - codex_desc=${3-} - - if [ "${RUNNER:-codex}" = "claude" ]; then - echo "$slash_cmd" - else - echo "Use the repo-local skill at '.claude/skills/${skill}/SKILL.md'. Follow it to ${codex_desc}. Read the skill file directly instead of assuming Claude slash-command support." - fi -} - -# Build a prompt and optionally append structured context for Codex. -# skill_prompt_with_context <skill> <slash-cmd> <codex-desc> <context-label> <context-json> -skill_prompt_with_context() { - skill=$1 - slash_cmd=$2 - codex_desc=${3-} - context_label=${4-} - context_json=${5-} - - base_prompt=$(skill_prompt "$skill" "$slash_cmd" "$codex_desc") - if [ "${RUNNER:-codex}" = "claude" ] || [ -z "$context_json" ]; then - echo "$base_prompt" - else - printf '%s\n\n## %s\n%s\n' "$base_prompt" "$context_label" "$context_json" - fi -} - -# Run an agent with the configured runner (claude or codex). -# run_agent <log-file> <prompt> -run_agent() { - output_file=$1 - prompt=$2 - - if [ "${RUNNER:-codex}" = "claude" ]; then - claude --dangerously-skip-permissions \ - --model "${CLAUDE_MODEL:-opus}" \ - --verbose \ - --output-format text \ - --max-turns 500 \ - -p "$prompt" 2>&1 | tee "$output_file" - else - codex exec \ - --enable multi_agent \ - -m "${CODEX_MODEL:-gpt-5.4}" \ - -s danger-full-access \ - "$prompt" 2>&1 | tee "$output_file" - fi -} - -watch_emit_outcome() { - outcome=$1 - if [ -n "${WATCH_MODE:-}" ]; then - printf '__WATCH_OUTCOME__=%s\n' "$outcome" - fi -} - -run_agent_with_watch_outcome() { - output_file=$1 - prompt=$2 - - if run_agent "$output_file" "$prompt"; then - watch_emit_outcome processed - return 0 - fi - - rc=$? - watch_emit_outcome agent-failed - return "$rc" -} - -# --- Project board --- - -# Detect the next eligible item from the current board snapshot. -# poll_project_items <mode> <state-file> [repo] [number] [format] -poll_project_items() { - mode=$1 - state_file=$2 - repo=${3-} - number=${4-} - fmt=${5-text} - - set -- scripts/pipeline_board.py next "$mode" "$state_file" --format "$fmt" - if [ -n "$repo" ]; then - set -- "$@" --repo "$repo" - fi - if [ -n "$number" ]; then - set -- "$@" --number "$number" - fi - # Filter blocked [Rule] issues whose model dependency is missing on main - if [ "$mode" = "ready" ]; then - set -- "$@" --repo-root . - fi - python3 "$@" -} - -ack_polled_item() { - state_file=$1 - item_id=$2 - python3 scripts/pipeline_board.py ack "$state_file" "$item_id" -} - -board_next_json() { - mode=$1 - repo=${2-} - number=${3-} - state_file=${4-} - - if [ -z "$state_file" ]; then - state_file="/tmp/problemreductions-${mode}-state.json" - fi - - poll_project_items "$mode" "$state_file" "$repo" "$number" json -} - -claim_project_items() { - mode=$1 - state_file=$2 - repo=${3-} - number=${4-} - fmt=${5-json} - - set -- scripts/pipeline_board.py claim-next "$mode" "$state_file" --format "$fmt" - if [ -n "$repo" ]; then - set -- "$@" --repo "$repo" - fi - if [ -n "$number" ]; then - set -- "$@" --number "$number" - fi - # Filter blocked [Rule] issues whose model dependency is missing on main - if [ "$mode" = "ready" ]; then - set -- "$@" --repo-root . - fi - python3 "$@" -} - -board_claim_json() { - mode=$1 - repo=${2-} - number=${3-} - state_file=${4-} - - if [ -z "$state_file" ]; then - state_file="/tmp/problemreductions-${mode}-state.json" - fi - - claim_project_items "$mode" "$state_file" "$repo" "$number" json -} - -move_board_item() { - item_id=$1 - status=$2 - python3 scripts/pipeline_board.py move "$item_id" "$status" -} - -# --- PR helpers --- - -pr_snapshot() { - repo=$1 - pr=$2 - python3 scripts/pipeline_pr.py snapshot --repo "$repo" --pr "$pr" --format json -} - -pr_wait_ci() { - repo=$1 - pr=$2 - timeout=${3:-900} - interval=${4:-30} - python3 scripts/pipeline_pr.py wait-ci --repo "$repo" --pr "$pr" --timeout "$timeout" --interval "$interval" --format json -} - -review_pipeline_context() { - repo=$1 - pr=${2-} - fmt=${3:-json} - - set -- scripts/pipeline_skill_context.py review-pipeline --repo "$repo" --format "$fmt" - if [ -n "$pr" ]; then - set -- "$@" --pr "$pr" - fi - python3 "$@" -} - -# --- Issue helpers --- - -issue_guards() { - repo=$1 - issue=$2 - repo_root=${3:-.} - python3 scripts/pipeline_checks.py issue-guards --repo "$repo" --issue "$issue" --repo-root "$repo_root" --format json -} - -issue_context() { - repo=$1 - issue=$2 - repo_root=${3:-.} - python3 scripts/pipeline_checks.py issue-context --repo "$repo" --issue "$issue" --repo-root "$repo_root" --format json -} - -# --- Worktree helpers --- - -create_issue_worktree() { - issue=$1 - slug=$2 - base=${3:-origin/main} - python3 scripts/pipeline_worktree.py create-issue --issue "$issue" --slug "$slug" --base "$base" --format json -} - -checkout_pr_worktree() { - repo=$1 - pr=$2 - python3 scripts/pipeline_worktree.py checkout-pr --repo "$repo" --pr "$pr" --format json -} - -merge_main_worktree() { - worktree=$1 - python3 scripts/pipeline_worktree.py merge-main --worktree "$worktree" --format json -} - -cleanup_pipeline_worktree() { - worktree=$1 - python3 scripts/pipeline_worktree.py cleanup --worktree "$worktree" --format json -} - -# Run a make target from watch_and_dispatch and capture any explicit outcome marker. -dispatch_watch_target() { - make_target=$1 - number=$2 - output_file=$(mktemp) - - if WATCH_MODE=1 ${MAKE:-make} "$make_target" N="$number" >"$output_file" 2>&1; then - rc=0 - else - rc=$? - fi - - outcome= - while IFS= read -r line || [ -n "$line" ]; do - case "$line" in - __WATCH_OUTCOME__=*) - outcome=${line#__WATCH_OUTCOME__=} - ;; - *) - printf '%s\n' "$line" - ;; - esac - done <"$output_file" - rm -f "$output_file" - - if [ -z "$outcome" ]; then - if [ "$rc" -eq 0 ]; then - outcome=processed - else - outcome=agent-failed - fi - fi - - WATCH_DISPATCH_OUTCOME=$outcome - WATCH_DISPATCH_RC=$rc -} - -# Poll a board column and dispatch a make target until the eligible queue is drained. -# watch_and_dispatch <mode> <make-target> <label> [repo] -# Example: -# watch_and_dispatch ready run-pipeline "Ready issues" -# watch_and_dispatch review run-review "Review pool PRs" "$REPO" -watch_and_dispatch() { - mode=$1 - make_target=$2 - label=$3 - repo=${4-} - interval=${POLL_INTERVAL:-1800} - - state_file=${STATE_FILE:-/tmp/problemreductions-${mode}-forever-state.json} - - trap 'exit 130' INT TERM - echo "Watching ${label} (polling every $((interval / 60))m when idle)..." - while true; do - next_item=$(poll_project_items "$mode" "$state_file" "$repo" "" text) - status=$? - if [ "$status" -eq 0 ]; then - item_id=$(printf '%s\n' "$next_item" | cut -f1) - number=$(printf '%s\n' "$next_item" | cut -f2) - echo "$(date '+%Y-%m-%d %H:%M:%S') Dispatching ${label} item $number ($item_id)" - dispatch_watch_target "$make_target" "$number" - dispatch_outcome=$WATCH_DISPATCH_OUTCOME - dispatch_rc=$WATCH_DISPATCH_RC - case "$dispatch_outcome" in - processed) - ack_polled_item "$state_file" "$item_id" || exit $? - echo "$(date '+%Y-%m-%d %H:%M:%S') Processed ${label} item $number; rechecking queue immediately..." - continue - ;; - gone) - echo "$(date '+%Y-%m-%d %H:%M:%S') ${label} item $number is no longer eligible; rechecking queue immediately..." - continue - ;; - *) - echo "$(date '+%Y-%m-%d %H:%M:%S') Dispatch for ${label} item $number ended with outcome '$dispatch_outcome' (rc=${dispatch_rc}); sleeping $((interval / 60))m before retrying..." >&2 - sleep "$interval" - continue - ;; - esac - elif [ "$status" -eq 1 ]; then - echo "$(date '+%Y-%m-%d %H:%M:%S') No eligible ${label}, sleeping $((interval / 60))m..." - sleep "$interval" - else - exit "$status" - fi - done -} diff --git a/scripts/pipeline_board.py b/scripts/pipeline_board.py deleted file mode 100644 index 70130179c..000000000 --- a/scripts/pipeline_board.py +++ /dev/null @@ -1,1716 +0,0 @@ -#!/usr/bin/env python3 -"""Shared project-board logic for polling, recovery, and board CLI helpers.""" - -from __future__ import annotations - -import argparse -import json -import subprocess -import sys -import time -from collections import Counter -from datetime import datetime, timezone -from pathlib import Path -from typing import Callable - -PROJECT_ID = "PVT_kwDOBrtarc4BRNVy" -STATUS_FIELD_ID = "PVTSSF_lADOBrtarc4BRNVyzg_GmQc" - -STATUS_BACKLOG = "Backlog" -STATUS_READY = "Ready" -STATUS_IN_PROGRESS = "In progress" -STATUS_REVIEW_POOL = "Review pool" -STATUS_UNDER_REVIEW = "Under review" -STATUS_FINAL_REVIEW = "Final review" -STATUS_ON_HOLD = "OnHold" -STATUS_DONE = "Done" - -STATUS_OPTION_IDS = { - STATUS_BACKLOG: "ab337660", - STATUS_READY: "f37d0d80", - STATUS_IN_PROGRESS: "a12cfc9c", - STATUS_REVIEW_POOL: "7082ed60", - STATUS_UNDER_REVIEW: "f04790ca", - STATUS_FINAL_REVIEW: "51a3d8bb", - STATUS_ON_HOLD: "48dfe446", - STATUS_DONE: "6aca54fa", -} - -STATUS_ALIASES = { - "backlog": STATUS_BACKLOG, - "ready": STATUS_READY, - "in-progress": STATUS_IN_PROGRESS, - "in_progress": STATUS_IN_PROGRESS, - "in progress": STATUS_IN_PROGRESS, - "review-pool": STATUS_REVIEW_POOL, - "review_pool": STATUS_REVIEW_POOL, - "review pool": STATUS_REVIEW_POOL, - "under-review": STATUS_UNDER_REVIEW, - "under_review": STATUS_UNDER_REVIEW, - "under review": STATUS_UNDER_REVIEW, - "final-review": STATUS_FINAL_REVIEW, - "final_review": STATUS_FINAL_REVIEW, - "final review": STATUS_FINAL_REVIEW, - "on-hold": STATUS_ON_HOLD, - "on_hold": STATUS_ON_HOLD, - "on hold": STATUS_ON_HOLD, - "onhold": STATUS_ON_HOLD, - "done": STATUS_DONE, -} - -FAILURE_LABELS = {"PoorWritten", "Wrong", "Trivial", "Useless"} - - -def run_gh(*args: str, retries: int = 3, retry_delay: float = 5.0) -> str: - """Run a ``gh`` CLI command, retrying on transient failures. - - The ``gh project`` subcommands occasionally fail with cryptic errors - like "unknown owner type" due to transient API issues or token - refresh races (see cli/cli#7985, cli/cli#8885). Retrying after a - short delay resolves these reliably. - """ - last_exc: subprocess.CalledProcessError | None = None - for attempt in range(retries): - try: - return subprocess.check_output( - ["gh", *args], text=True, stderr=subprocess.PIPE, - ) - except subprocess.CalledProcessError as exc: - last_exc = exc - stderr = (exc.stderr or "").strip() - if attempt < retries - 1: - print( - f"[run_gh] attempt {attempt + 1}/{retries} failed " - f"(rc={exc.returncode}, stderr={stderr!r}), " - f"retrying in {retry_delay}s…", - file=sys.stderr, - ) - time.sleep(retry_delay) - else: - print( - f"[run_gh] all {retries} attempts failed " - f"(stderr={stderr!r})", - file=sys.stderr, - ) - raise last_exc # type: ignore[misc] - - -def _graphql_board_query(project_id: str, page_size: int, cursor: str | None) -> str: - """Build a lightweight GraphQL query for project board items. - - Fetches only id, status, and content (type/number/title) — no other field - values. This costs ~1 GraphQL point per page vs ~20-50 for - ``gh project item-list`` which pulls all field values. - """ - after = f', after: "{cursor}"' if cursor else "" - return ( - "query {" - f' node(id: "{project_id}") {{' - " ... on ProjectV2 {" - f" items(first: {page_size}{after}) {{" - " pageInfo { hasNextPage endCursor }" - " nodes {" - " id" - ' fieldValueByName(name: "Status") {' - " ... on ProjectV2ItemFieldSingleSelectValue { name }" - " }" - " content {" - " __typename" - " ... on Issue { number title }" - " ... on PullRequest { number title }" - " }" - " }" - " }" - " }" - " }" - "}" - ) - - -def _parse_graphql_board_items(raw: dict) -> tuple[list[dict], bool, str | None]: - """Parse a GraphQL board response into (items, has_next_page, cursor).""" - node = raw.get("data", {}).get("node") or {} - project = node # already a ProjectV2 via inline fragment - connection = project.get("items") or {} - page_info = connection.get("pageInfo") or {} - has_next = page_info.get("hasNextPage", False) - cursor = page_info.get("endCursor") - - items: list[dict] = [] - for gql_node in connection.get("nodes") or []: - status_field = gql_node.get("fieldValueByName") or {} - status = status_field.get("name") - - content_raw = gql_node.get("content") or {} - typename = content_raw.get("__typename") - if typename not in ("Issue", "PullRequest"): - continue - - content = { - "type": typename, - "number": content_raw.get("number"), - "title": content_raw.get("title"), - } - items.append({ - "id": gql_node.get("id"), - "status": status, - "content": content, - "title": content_raw.get("title"), - }) - - return items, has_next, cursor - - -def fetch_board_items_graphql( - project_id: str, - limit: int, - *, - page_size: int = 100, - cache_file: Path | None = None, - cache_max_age: float = 120, -) -> dict: - """Fetch board items using a lightweight custom GraphQL query. - - Returns ``{"items": [...]}`` in the same shape as ``gh project item-list`` - (minus ``linked pull requests`` which callers resolve via ``pr_resolver``). - Costs ~1 GraphQL point per 100-item page vs ~20-50 for the full CLI fetch. - """ - if cache_file is not None: - try: - age = time.time() - cache_file.stat().st_mtime - if age < cache_max_age: - return json.loads(cache_file.read_text()) - except (FileNotFoundError, json.JSONDecodeError): - pass - - all_items: list[dict] = [] - cursor: str | None = None - while len(all_items) < limit: - batch = min(page_size, limit - len(all_items)) - query = _graphql_board_query(project_id, batch, cursor) - raw = json.loads(run_gh("api", "graphql", "-f", f"query={query}")) - items, has_next, cursor = _parse_graphql_board_items(raw) - all_items.extend(items) - if not has_next: - break - - data = {"items": all_items} - if cache_file is not None: - cache_file.parent.mkdir(parents=True, exist_ok=True) - cache_file.write_text(json.dumps(data)) - - return data - - -DEFAULT_OWNER = "CodingThrust" -DEFAULT_PROJECT_NUMBER = 8 - - -def fetch_board_items( - owner: str, - project_number: int, - limit: int, - *, - cache_file: Path | None = None, - cache_max_age: float = 120, - lite: bool = False, -) -> dict: - """Fetch project board items, optionally using a file cache. - - When *lite* is True **and** the target is the default project, uses a - lightweight custom GraphQL query (~1 point per 100 items) that omits - ``linked pull requests``. Callers that need linked PRs (review, - final-review) should pass ``lite=False`` (the default). - - For non-default owner/project_number, always falls back to the full - ``gh project item-list`` fetch regardless of *lite*. - """ - if cache_file is not None: - try: - age = time.time() - cache_file.stat().st_mtime - if age < cache_max_age: - return json.loads(cache_file.read_text()) - except (FileNotFoundError, json.JSONDecodeError): - pass - - if lite and owner == DEFAULT_OWNER and project_number == DEFAULT_PROJECT_NUMBER: - data = fetch_board_items_graphql( - PROJECT_ID, - limit, - ) - else: - data = json.loads( - run_gh( - "project", - "item-list", - str(project_number), - "--owner", - owner, - "--format", - "json", - "--limit", - str(limit), - ) - ) - - if cache_file is not None: - cache_file.parent.mkdir(parents=True, exist_ok=True) - cache_file.write_text(json.dumps(data)) - - return data - - -def fetch_pr_reviews(repo: str, pr_number: int) -> list[dict]: - data = json.loads(run_gh("api", f"repos/{repo}/pulls/{pr_number}/reviews")) - if not isinstance(data, list): - raise ValueError(f"Unexpected PR review payload for #{pr_number}: {data!r}") - return data - - -def fetch_pr_state(repo: str, pr_number: int) -> str: - return run_gh( - "pr", - "view", - str(pr_number), - "--repo", - repo, - "--json", - "state", - "--jq", - ".state", - ).strip() - - -def fetch_pr_info(repo: str, pr_number: int) -> dict: - data = json.loads( - run_gh( - "pr", - "view", - str(pr_number), - "--repo", - repo, - "--json", - "number,state,title,url", - ) - ) - if not isinstance(data, dict): - raise ValueError(f"Unexpected PR payload for #{pr_number}: {data!r}") - return data - - -def batch_fetch_prs_with_reviews(repo: str, pr_numbers: list[int]) -> dict[int, dict]: - """Fetch PR info + reviews for multiple PRs in a single GraphQL call. - - Returns dict keyed by PR number with keys: number, state, title, url, reviews. - """ - if not pr_numbers: - return {} - - owner, name = repo.split("/", 1) - - fragments = [] - for n in pr_numbers: - fragments.append( - f"pr_{n}: pullRequest(number: {n}) {{" - f" number state title url" - f" reviews(last: 50) {{ nodes {{ author {{ login }} state }} }}" - f"}}" - ) - - query = ( - f'query {{ repository(owner: "{owner}", name: "{name}") {{' - + " ".join(fragments) - + "}}" - ) - - raw = json.loads(run_gh("api", "graphql", "-f", f"query={query}")) - repo_data = raw.get("data", {}).get("repository", {}) - - result: dict[int, dict] = {} - for n in pr_numbers: - pr_data = repo_data.get(f"pr_{n}") - if pr_data is None: - continue - reviews_nodes = pr_data.get("reviews", {}).get("nodes", []) - result[n] = { - "number": pr_data["number"], - "state": pr_data["state"], - "title": pr_data.get("title", ""), - "url": pr_data.get("url", ""), - "reviews": reviews_nodes, - } - return result - - -def batch_fetch_issues(repo: str, issue_numbers: list[int]) -> dict[int, dict]: - """Fetch issue metadata for multiple issues in a single GraphQL call. - - Returns dict keyed by issue number with keys: number, title, body, state, url, labels, comments. - """ - if not issue_numbers: - return {} - - owner, name = repo.split("/", 1) - - fragments = [] - for n in issue_numbers: - fragments.append( - f"issue_{n}: issue(number: {n}) {{" - f" number title body state url" - f" labels(first: 20) {{ nodes {{ name }} }}" - f" comments(first: 50) {{ nodes {{ body }} }}" - f"}}" - ) - - query = ( - f'query {{ repository(owner: "{owner}", name: "{name}") {{' - + " ".join(fragments) - + "}}" - ) - - raw = json.loads(run_gh("api", "graphql", "-f", f"query={query}")) - repo_data = raw.get("data", {}).get("repository", {}) - - result: dict[int, dict] = {} - for n in issue_numbers: - issue_data = repo_data.get(f"issue_{n}") - if issue_data is None: - continue - result[n] = { - "number": issue_data["number"], - "title": issue_data.get("title", ""), - "body": issue_data.get("body", ""), - "state": issue_data.get("state", ""), - "url": issue_data.get("url", ""), - "labels": issue_data.get("labels", {}).get("nodes", []), - "comments": issue_data.get("comments", {}).get("nodes", []), - } - return result - - -def resolve_issue_pr(repo: str, issue_number: int) -> int | None: - data = json.loads( - run_gh( - "pr", - "list", - "-R", - repo, - "--search", - f"Fix #{issue_number} in:title state:open", - "--json", - "number", - "--limit", - "1", - ) - ) - if not data: - return None - return int(data[0]["number"]) - - -def item_identity(item: dict) -> str: - item_id = item.get("id") - if item_id is not None: - return str(item_id) - - content = item.get("content") or {} - number = content.get("number") - item_type = content.get("type", "item") - if number is not None: - return f"{item_type}:{number}" - - title = item.get("title") - if title: - return str(title) - - raise ValueError(f"Board item has no stable identity: {item!r}") - - -def load_state(state_file: Path) -> dict: - if not state_file.exists(): - return {"retries": {}} - - raw = state_file.read_text().strip() - if not raw: - return {"retries": {}} - - data = json.loads(raw) - if not isinstance(data, dict): - raise ValueError(f"State file must contain a JSON object: {state_file}") - - retries = data.get("retries", {}) - if not isinstance(retries, dict): - raise ValueError(f"Invalid poll state format: {state_file}") - - normalized_retries: dict[str, int] = {} - for item_id, count in retries.items(): - normalized_retries[str(item_id)] = int(count) - return {"retries": normalized_retries} - - -def save_state(state_file: Path, state: dict) -> None: - state_file.parent.mkdir(parents=True, exist_ok=True) - state_file.write_text(json.dumps(state, indent=2, sort_keys=True) + "\n") - - -def linked_pr_numbers(item: dict, repo: str | None = None) -> list[int]: - urls = item.get("linked pull requests") or [] - numbers: list[int] = [] - - if repo is not None: - prefix = f"https://github.com/{repo}/pull/" - for url in urls: - if not isinstance(url, str) or not url.startswith(prefix): - continue - suffix = url.removeprefix(prefix) - if suffix.isdigit(): - numbers.append(int(suffix)) - return numbers - - for url in urls: - if not isinstance(url, str): - continue - try: - numbers.append(int(url.rstrip("/").split("/")[-1])) - except ValueError: - continue - return numbers - - -def linked_repo_pr_numbers(item: dict, repo: str) -> list[int]: - return linked_pr_numbers(item, repo) - - -def entry_title(item: dict) -> str | None: - content = item.get("content") or {} - return content.get("title") or item.get("title") - - -def build_entry( - item: dict, - *, - number: int, - issue_number: int | None = None, - pr_number: int | None = None, -) -> dict: - return { - "number": number, - "issue_number": issue_number, - "pr_number": pr_number, - "status": item.get("status"), - "title": entry_title(item), - } - - -def _scan_existing_problems(repo_root: str | Path) -> set[str]: - """Return public struct/enum names under src/models/ (lightweight model scan).""" - import re as _re - - models_root = Path(repo_root) / "src" / "models" - if not models_root.exists(): - return set() - names: set[str] = set() - for path in models_root.rglob("*.rs"): - for m in _re.finditer(r"\bpub\s+(?:struct|enum)\s+([A-Z][A-Za-z0-9_]*)\b", path.read_text()): - names.add(m.group(1)) - return names - - -def _is_rule_blocked(title: str | None, existing_problems: set[str]) -> bool: - """Return True if *title* is a [Rule] whose source or target model is missing.""" - import re as _re - - if not title: - return False - m = _re.match(r"^\[Rule\]\s+(?P<source>.+?)\s+to\s+(?P<target>.+?)\s*$", title) - if not m: - return False - return m.group("source") not in existing_problems or m.group("target") not in existing_problems - - -def ready_entries(board_data: dict, *, repo_root: str | Path | None = None) -> dict[str, dict]: - existing: set[str] | None = None - if repo_root is not None: - existing = _scan_existing_problems(repo_root) - - entries = {} - for item in board_data.get("items", []): - if item.get("status") != STATUS_READY: - continue - - content = item.get("content") or {} - number = content.get("number") - if number is None: - continue - - if existing is not None: - title = content.get("title") or item.get("title") - if _is_rule_blocked(title, existing): - continue - - issue_number = int(number) - entries[item_identity(item)] = build_entry( - item, - number=issue_number, - issue_number=issue_number, - ) - return entries - - -def status_items( - board_data: dict, - status_name: str, - *, - content_types: set[str] | None = None, -) -> list[dict]: - if content_types is None: - content_types = {"Issue"} - items = [] - for item in board_data.get("items", []): - if item.get("status") != status_name: - continue - - content = item.get("content") or {} - item_type = content.get("type") - if item_type not in content_types: - continue - - number = content.get("number") - if number is None: - continue - - issue_number = int(number) if item_type == "Issue" else None - pr_number = int(number) if item_type == "PullRequest" else None - entry = build_entry( - item, - number=int(number), - issue_number=issue_number, - pr_number=pr_number, - ) - entry["item_id"] = item_identity(item) - items.append(entry) - - return sorted(items, key=lambda entry: (entry["number"], entry["item_id"])) - - -def review_entries( - board_data: dict, - repo: str, - pr_resolver: Callable[[str, int], int | None] | None, - pr_state_fetcher: Callable[[str, int], str], - *, - batch_pr_fetcher: Callable[[str, list[int]], dict[int, dict]] | None = None, -) -> dict[str, dict]: - pr_cache: dict[int, dict] = {} - if batch_pr_fetcher is not None: - all_pr_numbers: list[int] = [] - for item in board_data.get("items", []): - if item.get("status") != STATUS_REVIEW_POOL: - continue - content = item.get("content") or {} - number = content.get("number") - if number is None: - continue - if content.get("type") == "PullRequest": - all_pr_numbers.append(int(number)) - else: - all_pr_numbers.extend(linked_pr_numbers(item, repo)) - if all_pr_numbers: - pr_cache = batch_pr_fetcher(repo, all_pr_numbers) - - def _get_pr_state(pr_num: int) -> str: - if pr_num in pr_cache: - return str(pr_cache[pr_num].get("state", "")) - return pr_state_fetcher(repo, pr_num) - - entries = {} - for item in board_data.get("items", []): - if item.get("status") != STATUS_REVIEW_POOL: - continue - - content = item.get("content") or {} - item_type = content.get("type") - number = content.get("number") - if number is None: - continue - - pr_number: int | None - if item_type == "PullRequest": - pr_number = int(number) - if _get_pr_state(pr_number) != "OPEN": - continue - elif item_type == "Issue": - linked_numbers = linked_pr_numbers(item, repo) - if len(linked_numbers) > 1: - continue - if len(linked_numbers) == 1: - pr_number = linked_numbers[0] - if _get_pr_state(pr_number) != "OPEN": - continue - else: - if pr_resolver is None: - raise ValueError("review mode requires pr_resolver for issue cards without linked PRs") - pr_number = pr_resolver(repo, int(number)) - if pr_number is None: - continue - if _get_pr_state(pr_number) != "OPEN": - continue - else: - pr_number = None - - if pr_number is None: - continue - - issue_number = int(number) if item_type == "Issue" else None - entries[item_identity(item)] = build_entry( - item, - number=pr_number, - issue_number=issue_number, - pr_number=pr_number, - ) - return entries - - -def review_candidates( - board_data: dict, - repo: str, - pr_resolver: Callable[[str, int], int | None] | None, - pr_info_fetcher: Callable[[str, int], dict], - *, - batch_pr_fetcher: Callable[[str, list[int]], dict[int, dict]] | None = None, -) -> list[dict]: - # Pre-fetch all known PR numbers in one batch if batch fetcher is available - pr_cache: dict[int, dict] = {} - if batch_pr_fetcher is not None: - all_pr_numbers: list[int] = [] - for item in board_data.get("items", []): - if item.get("status") != STATUS_REVIEW_POOL: - continue - content = item.get("content") or {} - number = content.get("number") - if number is None: - continue - if content.get("type") == "PullRequest": - all_pr_numbers.append(int(number)) - else: - all_pr_numbers.extend(linked_pr_numbers(item, repo)) - if all_pr_numbers: - pr_cache = batch_pr_fetcher(repo, all_pr_numbers) - - def _get_pr_info(pr_num: int) -> dict: - if pr_num in pr_cache: - return pr_cache[pr_num] - return pr_info_fetcher(repo, pr_num) - - candidates = [] - for item in board_data.get("items", []): - if item.get("status") != STATUS_REVIEW_POOL: - continue - - content = item.get("content") or {} - item_type = content.get("type") - number = content.get("number") - if number is None: - continue - - base_entry = build_entry(item, number=0) - base_entry["item_id"] = item_identity(item) - issue_number = int(number) if item_type == "Issue" else None - - if item_type == "PullRequest": - pr_number = int(number) - pr_info = _get_pr_info(pr_number) - state = pr_info.get("state") - base_entry.update({"number": pr_number, "pr_number": pr_number}) - if state != "OPEN": - base_entry.update( - { - "eligibility": "stale-closed-pr", - "reason": f"linked PR #{pr_number} is {state}", - } - ) - candidates.append(base_entry) - continue - - base_entry.update({"eligibility": "eligible", "reason": "open PR"}) - candidates.append(base_entry) - continue - - if item_type != "Issue": - continue - - base_entry["issue_number"] = issue_number - linked_numbers = linked_pr_numbers(item, repo) - if len(linked_numbers) > 1: - linked_infos = [_get_pr_info(pr_number) for pr_number in linked_numbers] - open_numbers = [ - int(info["number"]) - for info in linked_infos - if str(info.get("state")).upper() == "OPEN" - ] - recommendation = open_numbers[0] if len(open_numbers) == 1 else None - base_entry.update( - { - "number": recommendation or int(linked_infos[0]["number"]), - "pr_number": recommendation, - "eligibility": "ambiguous-linked-prs", - "reason": "multiple linked repo PRs require confirmation", - "recommendation": recommendation, - "linked_repo_prs": [ - { - "number": int(info["number"]), - "state": str(info.get("state")), - "title": info.get("title"), - } - for info in linked_infos - ], - } - ) - candidates.append(base_entry) - continue - - if len(linked_numbers) == 1: - pr_number = linked_numbers[0] - pr_info = _get_pr_info(pr_number) - state = pr_info.get("state") - base_entry.update({"number": pr_number, "pr_number": pr_number}) - if state != "OPEN": - base_entry.update( - { - "eligibility": "stale-closed-pr", - "reason": f"linked PR #{pr_number} is {state}", - } - ) - candidates.append(base_entry) - continue - else: - if pr_resolver is None: - raise ValueError("review candidate listing requires pr_resolver for issue cards without linked PRs") - pr_number = pr_resolver(repo, issue_number) - if pr_number is None: - base_entry.update( - { - "number": issue_number, - "pr_number": None, - "eligibility": "no-open-pr", - "reason": f"issue #{issue_number} has no open PR", - } - ) - candidates.append(base_entry) - continue - base_entry.update({"number": pr_number, "pr_number": pr_number}) - - base_entry.update({"eligibility": "eligible", "reason": "open PR"}) - candidates.append(base_entry) - - return sorted( - candidates, - key=lambda entry: ( - entry["pr_number"] is None, - entry["number"], - entry["item_id"], - ), - ) - - -def final_review_entries( - board_data: dict, - repo: str, - pr_resolver: Callable[[str, int], int | None] | None, - pr_state_fetcher: Callable[[str, int], str], - *, - batch_pr_fetcher: Callable[[str, list[int]], dict[int, dict]] | None = None, -) -> dict[str, dict]: - pr_cache: dict[int, dict] = {} - if batch_pr_fetcher is not None: - all_pr_numbers: list[int] = [] - for item in board_data.get("items", []): - if item.get("status") != STATUS_FINAL_REVIEW: - continue - content = item.get("content") or {} - number = content.get("number") - if number is None: - continue - if content.get("type") == "PullRequest": - all_pr_numbers.append(int(number)) - else: - all_pr_numbers.extend(linked_pr_numbers(item, repo)) - if all_pr_numbers: - pr_cache = batch_pr_fetcher(repo, all_pr_numbers) - - def _get_pr_state(pr_num: int) -> str: - if pr_num in pr_cache: - return str(pr_cache[pr_num].get("state", "")) - return pr_state_fetcher(repo, pr_num) - - entries = {} - for item in board_data.get("items", []): - if item.get("status") != STATUS_FINAL_REVIEW: - continue - - content = item.get("content") or {} - item_type = content.get("type") - number = content.get("number") - if number is None: - continue - - pr_number: int | None - if item_type == "PullRequest": - pr_number = int(number) - if _get_pr_state(pr_number) != "OPEN": - continue - elif item_type == "Issue": - linked_numbers = linked_pr_numbers(item, repo) - if len(linked_numbers) > 1: - continue - if len(linked_numbers) == 1: - pr_number = linked_numbers[0] - if _get_pr_state(pr_number) != "OPEN": - continue - else: - if pr_resolver is None: - raise ValueError( - "final-review mode requires pr_resolver for issue cards without linked PRs" - ) - pr_number = pr_resolver(repo, int(number)) - if pr_number is None: - continue - if _get_pr_state(pr_number) != "OPEN": - continue - else: - pr_number = None - - if pr_number is None: - continue - - issue_number = int(number) if item_type == "Issue" else None - entries[item_identity(item)] = build_entry( - item, - number=pr_number, - issue_number=issue_number, - pr_number=pr_number, - ) - return entries - - -def current_entries( - mode: str, - board_data: dict, - repo: str | None = None, - pr_resolver: Callable[[str, int], int | None] | None = None, - pr_state_fetcher: Callable[[str, int], str] | None = None, - *, - batch_pr_fetcher: Callable[[str, list[int]], dict[int, dict]] | None = None, - repo_root: str | Path | None = None, -) -> dict[str, dict]: - if mode == "ready": - return ready_entries(board_data, repo_root=repo_root) - if mode == "review": - if repo is None: - raise ValueError("repo is required in review mode") - if pr_state_fetcher is None: - raise ValueError("review mode requires pr_state_fetcher") - return review_entries( - board_data, - repo, - pr_resolver, - pr_state_fetcher, - batch_pr_fetcher=batch_pr_fetcher, - ) - if mode == "final-review": - if repo is None: - raise ValueError("repo is required in final-review mode") - if pr_state_fetcher is None: - raise ValueError("final-review mode requires pr_state_fetcher") - return final_review_entries( - board_data, - repo, - pr_resolver, - pr_state_fetcher, - batch_pr_fetcher=batch_pr_fetcher, - ) - raise ValueError(f"Unsupported mode: {mode}") - - -def process_snapshot( - mode: str, - board_data: dict, - state_file: Path, - repo: str | None = None, - pr_resolver: Callable[[str, int], int | None] | None = None, - pr_state_fetcher: Callable[[str, int], str] | None = None, - target_number: int | None = None, - *, - batch_pr_fetcher: Callable[[str, list[int]], dict[int, dict]] | None = None, - repo_root: str | Path | None = None, -) -> tuple[str, int] | None: - next_entry = select_next_entry( - mode, - board_data, - state_file, - repo, - pr_resolver, - pr_state_fetcher, - target_number, - batch_pr_fetcher=batch_pr_fetcher, - repo_root=repo_root, - ) - if next_entry is None: - return None - return str(next_entry["item_id"]), int(next_entry["number"]) - - -def select_entry_from_entries( - current_visible: dict[str, dict], - state_file: Path, - target_number: int | None = None, -) -> dict | None: - del state_file - return select_current_entry_from_entries( - "review", - current_visible, - target_number=target_number, - ) - - -def current_entry_sort_key( - mode: str, - entry: dict, - item_id: str, -) -> tuple[int, int, str] | tuple[int, str]: - if mode == "ready": - title = entry.get("title") or "" - if title.startswith("[Model]"): - kind_priority = 0 - elif title.startswith("[Rule]"): - kind_priority = 1 - else: - kind_priority = 2 - return (kind_priority, int(entry["number"]), item_id) - return (int(entry["number"]), item_id) - - -def select_current_entry_from_entries( - mode: str, - current_visible: dict[str, dict], - target_number: int | None = None, -) -> dict | None: - if target_number is not None: - matching_item_id = next( - ( - item_id - for item_id, entry in current_visible.items() - if int(entry["number"]) == target_number - ), - None, - ) - if matching_item_id is None: - return None - return { - **current_visible[matching_item_id], - "item_id": matching_item_id, - } - - if not current_visible: - return None - - item_id = min( - current_visible, - key=lambda candidate_id: current_entry_sort_key( - mode, - current_visible[candidate_id], - candidate_id, - ), - ) - return { - **current_visible[item_id], - "item_id": item_id, - } - - -def select_polled_entry_from_entries( - mode: str, - current_visible: dict[str, dict], - state_file: Path, - target_number: int | None = None, -) -> dict | None: - del state_file - return select_current_entry_from_entries( - mode, - current_visible, - target_number=target_number, - ) - - -def select_next_entry( - mode: str, - board_data: dict, - state_file: Path, - repo: str | None = None, - pr_resolver: Callable[[str, int], int | None] | None = None, - pr_state_fetcher: Callable[[str, int], str] | None = None, - target_number: int | None = None, - *, - batch_pr_fetcher: Callable[[str, list[int]], dict[int, dict]] | None = None, - repo_root: str | Path | None = None, -) -> dict | None: - current_visible = current_entries( - mode, - board_data, - repo, - pr_resolver, - pr_state_fetcher, - batch_pr_fetcher=batch_pr_fetcher, - repo_root=repo_root, - ) - return select_polled_entry_from_entries( - mode, - current_visible, - state_file, - target_number=target_number, - ) - - -def ack_item(state_file: Path, item_id: str) -> None: - state = load_state(state_file) - retries = state.get("retries", {}) - retries.pop(str(item_id), None) - state["retries"] = retries - save_state(state_file, state) - - -def label_names(issue: dict) -> set[str]: - return {label["name"] for label in issue.get("labels", [])} - - -def is_tracked_issue_title(title: str | None) -> bool: - if not title: - return False - return title.startswith("[Model]") or title.startswith("[Rule]") - - -def all_checks_green(pr: dict) -> bool: - statuses = pr.get("statusCheckRollup") or [] - if not statuses: - return False - - for status in statuses: - typename = status.get("__typename") - if typename == "CheckRun": - if status.get("status") != "COMPLETED": - return False - if status.get("conclusion") not in {"SUCCESS", "SKIPPED", "NEUTRAL"}: - return False - elif typename == "StatusContext": - if status.get("state") != "SUCCESS": - return False - return True - - -def infer_issue_status( - issue: dict, - linked_prs: list[dict], -) -> tuple[str, str]: - labels = label_names(issue) - merged_prs = [pr for pr in linked_prs if pr.get("mergedAt")] - open_prs = [pr for pr in linked_prs if pr.get("state") == "OPEN"] - - if merged_prs: - pr_numbers = ", ".join(f"#{pr['number']}" for pr in merged_prs) - return STATUS_DONE, f"linked merged PR {pr_numbers}" - - if issue.get("state") == "CLOSED": - return STATUS_DONE, "issue itself is closed" - - if open_prs: - green_prs = [pr for pr in open_prs if all_checks_green(pr)] - if len(green_prs) == len(open_prs): - pr_numbers = ", ".join(f"#{pr['number']}" for pr in open_prs) - return STATUS_FINAL_REVIEW, f"green open PR {pr_numbers}" - - pr_numbers = ", ".join(f"#{pr['number']}" for pr in open_prs) - return STATUS_REVIEW_POOL, f"open PR {pr_numbers} still implementing or fixing review" - - if "Good" in labels: - return STATUS_READY, 'label "Good" present and no linked PR' - - if labels & FAILURE_LABELS: - bad = ", ".join(sorted(labels & FAILURE_LABELS)) - return STATUS_BACKLOG, f"failure labels present: {bad}" - - return STATUS_BACKLOG, "default backlog: no linked PR and no Ready signal" - - -def build_recovery_plan( - board_data: dict, - issues: list[dict], - prs: list[dict], -) -> list[dict]: - issues_by_number = {issue["number"]: issue for issue in issues} - prs_by_number = {pr["number"]: pr for pr in prs} - - plan = [] - for item in board_data.get("items", []): - content = item.get("content") or {} - issue_number = content.get("number") - if issue_number is None: - continue - - issue = issues_by_number.get(issue_number) - if issue is None: - continue - - title = content.get("title") or issue.get("title") - if not is_tracked_issue_title(title): - continue - - linked_prs = [ - prs_by_number[pr_number] - for pr_number in linked_pr_numbers(item) - if pr_number in prs_by_number - ] - status_name, reason = infer_issue_status(issue, linked_prs) - plan.append( - { - "item_id": item["id"], - "issue_number": issue_number, - "title": title, - "current_status": item.get("status"), - "proposed_status": status_name, - "option_id": STATUS_OPTION_IDS[status_name], - "reason": reason, - } - ) - - return sorted(plan, key=lambda entry: entry["issue_number"]) - - -def save_backup( - backup_file: Path, - *, - board_data: dict, - issues: list[dict], - prs: list[dict], - plan: list[dict], -) -> None: - backup_file.parent.mkdir(parents=True, exist_ok=True) - payload = { - "generated_at": datetime.now(timezone.utc).isoformat(), - "board_data": board_data, - "issues": issues, - "prs": prs, - "plan": plan, - } - backup_file.write_text(json.dumps(payload, indent=2, sort_keys=True) + "\n") - - -def default_backup_path(project_number: int) -> Path: - stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ") - return Path("/tmp") / f"project-{project_number}-status-recovery-{stamp}.json" - - -def print_summary(plan: list[dict]) -> None: - counts = Counter(entry["proposed_status"] for entry in plan) - print("Proposed status counts:") - for status_name in [ - STATUS_BACKLOG, - STATUS_READY, - STATUS_REVIEW_POOL, - STATUS_FINAL_REVIEW, - STATUS_DONE, - ]: - print(f" {status_name}: {counts.get(status_name, 0)}") - - -def print_examples(plan: list[dict], limit: int = 20) -> None: - print("") - print(f"First {min(limit, len(plan))} assignments:") - for entry in plan[:limit]: - print( - f" #{entry['issue_number']:<4} {entry['proposed_status']:<13} " - f"{entry['reason']} | {entry['title']}" - ) - - -def normalize_status_name(status: str) -> str: - normalized = status.strip() - if normalized in STATUS_OPTION_IDS: - return normalized - - alias = STATUS_ALIASES.get(normalized.lower()) - if alias is None: - choices = ", ".join(sorted(STATUS_OPTION_IDS)) - raise ValueError(f"Unsupported status {status!r}. Expected one of: {choices}") - return alias - - -def claimed_status_for_mode(mode: str) -> str: - if mode == "ready": - return STATUS_IN_PROGRESS - if mode == "review": - return STATUS_UNDER_REVIEW - raise ValueError(f"Unsupported claim-next mode: {mode}") - - -def claim_entry_from_entries( - mode: str, - current_visible: dict[str, dict], - state_file: Path, - target_number: int | None = None, - mover: Callable[[str, str], None] | None = None, -) -> dict | None: - next_entry = select_polled_entry_from_entries( - mode, - current_visible, - state_file, - target_number=target_number, - ) - if next_entry is None: - return None - - claimed_status = claimed_status_for_mode(mode) - move = mover or move_item - move(str(next_entry["item_id"]), claimed_status) - return { - **next_entry, - "claimed": True, - "claimed_status": claimed_status, - } - - -def claim_next_entry( - mode: str, - board_data: dict, - state_file: Path, - repo: str | None = None, - pr_resolver: Callable[[str, int], int | None] | None = None, - pr_state_fetcher: Callable[[str, int], str] | None = None, - target_number: int | None = None, - mover: Callable[[str, str], None] | None = None, - *, - batch_pr_fetcher: Callable[[str, list[int]], dict[int, dict]] | None = None, - repo_root: str | Path | None = None, -) -> dict | None: - current_visible = current_entries( - mode, - board_data, - repo=repo, - pr_resolver=pr_resolver, - pr_state_fetcher=pr_state_fetcher, - batch_pr_fetcher=batch_pr_fetcher, - repo_root=repo_root, - ) - return claim_entry_from_entries( - mode, - current_visible, - state_file, - target_number=target_number, - mover=mover, - ) - - -def eligible_review_candidate_entries(candidates: list[dict]) -> dict[str, dict]: - entries: dict[str, dict] = {} - for candidate in candidates: - if candidate.get("eligibility") != "eligible": - continue - - item_id = candidate.get("item_id") - number = candidate.get("pr_number") or candidate.get("number") - if item_id is None or number is None: - continue - - issue_number = candidate.get("issue_number") - pr_number = candidate.get("pr_number") - entries[str(item_id)] = { - "number": int(number), - "issue_number": int(issue_number) if issue_number is not None else None, - "pr_number": int(pr_number) if pr_number is not None else int(number), - "status": candidate.get("status"), - "title": candidate.get("title"), - } - return entries - - -def move_item( - item_id: str, - status: str, - *, - project_id: str = PROJECT_ID, - field_id: str = STATUS_FIELD_ID, -) -> None: - status_name = normalize_status_name(status) - subprocess.check_call( - [ - "gh", - "project", - "item-edit", - "--project-id", - project_id, - "--id", - item_id, - "--field-id", - field_id, - "--single-select-option-id", - STATUS_OPTION_IDS[status_name], - ] - ) - - -def apply_plan( - plan: list[dict], - *, - project_id: str = PROJECT_ID, - field_id: str = STATUS_FIELD_ID, -) -> int: - changed = 0 - for entry in plan: - if entry["current_status"] == entry["proposed_status"]: - continue - move_item( - entry["item_id"], - entry["proposed_status"], - project_id=project_id, - field_id=field_id, - ) - changed += 1 - return changed - - -def print_next_item( - next_item: dict | None, - *, - mode: str, - fmt: str = "text", -) -> int: - if next_item is None: - return 1 - - if fmt == "json": - payload = {"mode": mode, **next_item} - print(json.dumps(payload)) - else: - print(f"{next_item['item_id']}\t{next_item['number']}") - return 0 - - -def print_claim_result( - claim_result: dict | None, - *, - mode: str, - fmt: str = "json", -) -> int: - if claim_result is None: - return 1 - - if fmt == "json": - payload = {"mode": mode, **claim_result} - print(json.dumps(payload)) - else: - print( - f"{claim_result['item_id']}\t" - f"{claim_result['number']}\t" - f"{claim_result['claimed_status']}" - ) - return 0 - - -def print_candidate_list( - mode: str, - items: list[dict], - *, - fmt: str = "text", -) -> int: - if fmt == "json": - print(json.dumps({"mode": mode, "items": items})) - return 0 - - for item in items: - number = item.get("pr_number") or item.get("issue_number") or item["number"] - title = item.get("title") or "" - eligibility = item.get("eligibility") or "" - print(f"{item['item_id']}\t{number}\t{eligibility}\t{title}") - return 0 - - -def backlog_issues( - board_data: dict, - issue_type: str, -) -> list[dict]: - """List Backlog issues of the given type, sorted by Good label first then by number. - - Returns a list of dicts with: number, title, item_id, labels, has_good. - """ - prefix = "[Model]" if issue_type == "model" else "[Rule]" - - check_labels = FAILURE_LABELS | {"Good"} - - results = [] - for item in board_data.get("items", []): - if item.get("status") != STATUS_BACKLOG: - continue - content = item.get("content") or {} - if content.get("type") != "Issue": - continue - title = content.get("title") or "" - if not title.startswith(prefix): - continue - number = content.get("number") - if number is None: - continue - item_labels = set(item.get("labels") or []) - # Only include issues that have been through check-issue - if not (item_labels & check_labels): - continue - has_good = "Good" in item_labels - results.append({ - "number": int(number), - "title": title, - "item_id": item_identity(item), - "labels": sorted(item_labels), - "has_good": has_good, - }) - - # Good-labeled first, then by issue number - results.sort(key=lambda r: (not r["has_good"], r["number"])) - return results - - -def parse_args(argv: list[str]) -> argparse.Namespace: - parser = argparse.ArgumentParser(description="Project board automation helpers.") - subparsers = parser.add_subparsers(dest="command", required=True) - - next_parser = subparsers.add_parser("next") - next_parser.add_argument("mode", choices=["ready", "review", "final-review"]) - next_parser.add_argument("state_file", type=Path) - next_parser.add_argument("--repo") - next_parser.add_argument("--owner", default="CodingThrust") - next_parser.add_argument("--project-number", type=int, default=8) - next_parser.add_argument("--limit", type=int, default=500) - next_parser.add_argument("--number", type=int) - next_parser.add_argument("--format", choices=["text", "json"], default="text") - next_parser.add_argument("--board-cache", type=Path, default=None) - next_parser.add_argument("--board-cache-max-age", type=float, default=120) - next_parser.add_argument("--repo-root", type=Path, default=None, - help="Repo root for filtering blocked [Rule] issues in ready mode") - - claim_parser = subparsers.add_parser("claim-next") - claim_parser.add_argument("mode", choices=["ready", "review"]) - claim_parser.add_argument("state_file", type=Path) - claim_parser.add_argument("--repo") - claim_parser.add_argument("--owner", default="CodingThrust") - claim_parser.add_argument("--project-number", type=int, default=8) - claim_parser.add_argument("--limit", type=int, default=500) - claim_parser.add_argument("--number", type=int) - claim_parser.add_argument("--format", choices=["text", "json"], default="json") - claim_parser.add_argument("--project-id", default=PROJECT_ID) - claim_parser.add_argument("--field-id", default=STATUS_FIELD_ID) - claim_parser.add_argument("--board-cache", type=Path, default=None) - claim_parser.add_argument("--board-cache-max-age", type=float, default=120) - claim_parser.add_argument("--repo-root", type=Path, default=None, - help="Repo root for filtering blocked [Rule] issues in ready mode") - - ack_parser = subparsers.add_parser("ack") - ack_parser.add_argument("state_file", type=Path) - ack_parser.add_argument("item_id") - - list_parser = subparsers.add_parser("list") - list_parser.add_argument( - "mode", - choices=["ready", "in-progress", "review-pool", "review", "final-review"], - ) - list_parser.add_argument("--repo") - list_parser.add_argument("--owner", default="CodingThrust") - list_parser.add_argument("--project-number", type=int, default=8) - list_parser.add_argument("--limit", type=int, default=500) - list_parser.add_argument("--format", choices=["text", "json"], default="text") - list_parser.add_argument("--board-cache", type=Path, default=None) - list_parser.add_argument("--board-cache-max-age", type=float, default=120) - - move_parser = subparsers.add_parser("move") - move_parser.add_argument("item_id") - move_parser.add_argument("status") - move_parser.add_argument("--project-id", default=PROJECT_ID) - move_parser.add_argument("--field-id", default=STATUS_FIELD_ID) - - fix_parser = subparsers.add_parser("backlog") - fix_parser.add_argument("issue_type", choices=["model", "rule"]) - fix_parser.add_argument("--owner", default="CodingThrust") - fix_parser.add_argument("--project-number", type=int, default=8) - fix_parser.add_argument("--limit", type=int, default=500) - fix_parser.add_argument("--format", choices=["text", "json"], default="json") - - find_parser = subparsers.add_parser("find") - find_parser.add_argument("number", type=int, help="Issue or PR number to look up") - find_parser.add_argument("--owner", default="CodingThrust") - find_parser.add_argument("--project-number", type=int, default=8) - find_parser.add_argument("--limit", type=int, default=500) - - return parser.parse_args(argv) - - -def main(argv: list[str] | None = None) -> int: - args = parse_args(argv or sys.argv[1:]) - - if args.command == "ack": - ack_item(args.state_file, args.item_id) - return 0 - - if args.command == "move": - move_item( - args.item_id, - args.status, - project_id=args.project_id, - field_id=args.field_id, - ) - return 0 - - if args.command == "backlog": - board_data = fetch_board_items(args.owner, args.project_number, args.limit) - results = backlog_issues(board_data, args.issue_type) - if args.format == "json": - print(json.dumps({"issue_type": args.issue_type, "items": results})) - else: - for r in results: - good = "Good" if r["has_good"] else "" - print(f"#{r['number']:<5} {good:5s} {r['title']}") - return 0 if results else 1 - - if args.command == "find": - board_data = fetch_board_items( - args.owner, args.project_number, args.limit, lite=True, - ) - for item in board_data.get("items", []): - content = item.get("content") or {} - if content.get("number") == args.number: - result = { - "item_id": item["id"], - "status": item.get("status"), - "number": content["number"], - "title": content.get("title"), - } - print(json.dumps(result)) - return 0 - print(json.dumps({"error": f"Issue #{args.number} not found on project board"})) - return 1 - - if args.command == "claim-next": - if args.mode == "review" and not args.repo: - raise SystemExit("--repo is required in claim-next review mode") - lite = args.mode not in {"review", "final-review"} - board_data = fetch_board_items(args.owner, args.project_number, args.limit, cache_file=args.board_cache, cache_max_age=args.board_cache_max_age, lite=lite) - if args.mode == "review": - review_entries_map = eligible_review_candidate_entries( - review_candidates( - board_data, - args.repo, - resolve_issue_pr, - fetch_pr_info, - batch_pr_fetcher=batch_fetch_prs_with_reviews, - ) - ) - claim_result = claim_entry_from_entries( - args.mode, - review_entries_map, - args.state_file, - target_number=args.number, - mover=lambda item_id, status: move_item( - item_id, - status, - project_id=args.project_id, - field_id=args.field_id, - ), - ) - else: - claim_result = claim_next_entry( - args.mode, - board_data, - args.state_file, - repo=args.repo, - pr_resolver=resolve_issue_pr, - pr_state_fetcher=fetch_pr_state, - target_number=args.number, - mover=lambda item_id, status: move_item( - item_id, - status, - project_id=args.project_id, - field_id=args.field_id, - ), - repo_root=getattr(args, "repo_root", None), - ) - return print_claim_result(claim_result, mode=args.mode, fmt=args.format) - - if args.command == "list": - if args.mode == "review" and not args.repo: - raise SystemExit("--repo is required in list review mode") - lite = args.mode not in {"review", "final-review"} - board_data = fetch_board_items(args.owner, args.project_number, args.limit, cache_file=args.board_cache, cache_max_age=args.board_cache_max_age, lite=lite) - if args.mode == "ready": - items = status_items(board_data, STATUS_READY) - return print_candidate_list(args.mode, items, fmt=args.format) - if args.mode == "in-progress": - items = status_items(board_data, STATUS_IN_PROGRESS) - return print_candidate_list(args.mode, items, fmt=args.format) - if args.mode == "review-pool": - items = status_items( - board_data, - STATUS_REVIEW_POOL, - content_types={"Issue", "PullRequest"}, - ) - return print_candidate_list(args.mode, items, fmt=args.format) - if args.mode == "review": - items = review_candidates( - board_data, - args.repo, - resolve_issue_pr, - fetch_pr_info, - batch_pr_fetcher=batch_fetch_prs_with_reviews, - ) - return print_candidate_list(args.mode, items, fmt=args.format) - if args.mode == "final-review": - items = status_items( - board_data, - STATUS_FINAL_REVIEW, - content_types={"Issue", "PullRequest"}, - ) - return print_candidate_list(args.mode, items, fmt=args.format) - raise SystemExit(f"Unsupported list mode: {args.mode}") - - if args.mode in {"review", "final-review"} and not args.repo: - raise SystemExit(f"--repo is required in {args.mode} mode") - - lite = args.mode not in {"review", "final-review"} - board_data = fetch_board_items(args.owner, args.project_number, args.limit, cache_file=args.board_cache, cache_max_age=args.board_cache_max_age, lite=lite) - if args.mode == "review": - review_entries_map = eligible_review_candidate_entries( - review_candidates( - board_data, - args.repo, - resolve_issue_pr, - fetch_pr_info, - batch_pr_fetcher=batch_fetch_prs_with_reviews, - ) - ) - next_item = select_polled_entry_from_entries( - args.mode, - review_entries_map, - args.state_file, - target_number=args.number, - ) - else: - next_item = select_next_entry( - args.mode, - board_data, - args.state_file, - repo=args.repo, - pr_resolver=resolve_issue_pr, - pr_state_fetcher=fetch_pr_state, - target_number=args.number, - batch_pr_fetcher=batch_fetch_prs_with_reviews, - repo_root=getattr(args, "repo_root", None), - ) - return print_next_item(next_item, mode=args.mode, fmt=args.format) - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/scripts/pipeline_checks.py b/scripts/pipeline_checks.py index f65e8e76d..ef10c1f85 100644 --- a/scripts/pipeline_checks.py +++ b/scripts/pipeline_checks.py @@ -1,5 +1,5 @@ #!/usr/bin/env python3 -"""Deterministic review checks for scope detection and file whitelists.""" +"""Deterministic review checks (scope detection, whitelists, completeness) and issue-context CLI.""" from __future__ import annotations @@ -645,62 +645,14 @@ def fetch_existing_prs(repo: str, issue_number: int) -> list[dict]: return [pr for pr in data if _pr_references_issue(pr, issue_number)] -def git_output(*args: str) -> list[str]: - output = subprocess.check_output(["git", *args], text=True) - return [line for line in output.splitlines() if line] - - -def git_text(*args: str) -> str: - return subprocess.check_output(["git", *args], text=True) - - -def load_file_list(path: str | Path) -> list[str]: - lines = Path(path).read_text().splitlines() - return [line.strip() for line in lines if line.strip()] - - def emit_result(result: dict, fmt: str) -> None: print(json.dumps(result, indent=2, sort_keys=True)) def parse_args(argv: list[str]) -> argparse.Namespace: - parser = argparse.ArgumentParser(description="Pipeline review checks.") + parser = argparse.ArgumentParser(description="Issue context checks.") subparsers = parser.add_subparsers(dest="command", required=True) - detect = subparsers.add_parser("detect-scope") - detect.add_argument("--base", required=True) - detect.add_argument("--head", required=True) - detect.add_argument("--format", choices=["json", "text"], default="json") - - whitelist = subparsers.add_parser("file-whitelist") - whitelist.add_argument("--kind", choices=["model", "rule"], required=True) - whitelist.add_argument("--files-file", required=True) - whitelist.add_argument("--format", choices=["json", "text"], default="json") - - completeness = subparsers.add_parser("completeness") - completeness.add_argument("--kind", choices=["model", "rule"], required=True) - completeness.add_argument("--name", required=True) - completeness.add_argument("--source") - completeness.add_argument("--target") - completeness.add_argument("--repo-root", default=".") - completeness.add_argument("--format", choices=["json", "text"], default="json") - - review_context = subparsers.add_parser("review-context") - review_context.add_argument("--repo-root", default=".") - review_context.add_argument("--base", required=True) - review_context.add_argument("--head", required=True) - review_context.add_argument("--kind", choices=["model", "rule", "generic"]) - review_context.add_argument("--name") - review_context.add_argument("--source") - review_context.add_argument("--target") - review_context.add_argument("--format", choices=["json", "text"], default="json") - - issue_guards = subparsers.add_parser("issue-guards") - issue_guards.add_argument("--repo", required=True) - issue_guards.add_argument("--issue", required=True, type=int) - issue_guards.add_argument("--repo-root", default=".") - issue_guards.add_argument("--format", choices=["json", "text"], default="json") - issue_context = subparsers.add_parser("issue-context") issue_context.add_argument("--repo", required=True) issue_context.add_argument("--issue", required=True, type=int) @@ -713,74 +665,7 @@ def parse_args(argv: list[str]) -> argparse.Namespace: def main(argv: list[str] | None = None) -> int: args = parse_args(argv or sys.argv[1:]) - if args.command == "detect-scope": - changed_files = git_output("diff", "--name-only", f"{args.base}..{args.head}") - added_files = git_output( - "diff", - "--name-only", - "--diff-filter=A", - f"{args.base}..{args.head}", - ) - emit_result( - detect_scope_from_paths( - added_files=added_files, - changed_files=changed_files, - ), - args.format, - ) - return 0 - - if args.command == "file-whitelist": - emit_result( - file_whitelist_check(args.kind, load_file_list(args.files_file)), - args.format, - ) - return 0 - - if args.command == "completeness": - emit_result( - completeness_check( - args.kind, - args.repo_root, - name=args.name, - source=args.source, - target=args.target, - ), - args.format, - ) - return 0 - - if args.command == "review-context": - changed_files = git_output("diff", "--name-only", f"{args.base}..{args.head}") - added_files = git_output( - "diff", - "--name-only", - "--diff-filter=A", - f"{args.base}..{args.head}", - ) - scope = detect_scope_from_paths( - added_files=added_files, - changed_files=changed_files, - ) - subject = infer_review_subject( - scope, - kind=args.kind, - name=args.name, - source=args.source, - target=args.target, - ) - emit_result( - build_review_context( - args.repo_root, - diff_stat=git_text("diff", "--stat", f"{args.base}..{args.head}"), - scope=scope, - subject=subject, - ), - args.format, - ) - return 0 - - if args.command in {"issue-guards", "issue-context"}: + if args.command == "issue-context": emit_result( issue_context_check( args.repo_root, diff --git a/scripts/pipeline_pr.py b/scripts/pipeline_pr.py index e56e0045d..db9543b38 100644 --- a/scripts/pipeline_pr.py +++ b/scripts/pipeline_pr.py @@ -721,18 +721,6 @@ def post_pr_comment(repo: str, pr_number: int, body_file: str) -> None: ) -def edit_pr_body(repo: str, pr_number: int, body_file: str) -> None: - run_gh_checked( - "pr", - "edit", - str(pr_number), - "--repo", - repo, - "--body-file", - body_file, - ) - - def render_context_text(result: dict) -> str: comments = result.get("comments") or {} counts = comments.get("counts") or {} @@ -829,25 +817,11 @@ def parse_args(argv: list[str]) -> argparse.Namespace: context.add_argument("--current", action="store_true") context.add_argument("--format", choices=["json", "text"], default="json") - for name in [ - "current", - "snapshot", - "comments", - "ci", - "wait-ci", - "codecov", - "linked-issue", - "create", - "comment", - "edit-body", - ]: + for name in ["comments", "ci", "wait-ci", "codecov", "create", "comment"]: command = subparsers.add_parser(name) - if name == "current": - command.add_argument("--format", choices=["json", "text"], default="json") - else: - command.add_argument("--repo", required=True) - if name != "create": - command.add_argument("--pr", required=True, type=int) + command.add_argument("--repo", required=True) + if name != "create": + command.add_argument("--pr", required=True, type=int) if name == "wait-ci": command.add_argument("--timeout", type=float, default=900) command.add_argument("--interval", type=float, default=30) @@ -857,9 +831,9 @@ def parse_args(argv: list[str]) -> argparse.Namespace: command.add_argument("--base") command.add_argument("--head") command.add_argument("--format", choices=["json", "text"], default="json") - elif name in {"comment", "edit-body"}: + elif name == "comment": command.add_argument("--body-file", required=True) - elif name != "current": + else: command.add_argument("--format", choices=["json", "text"], default="json") return parser.parse_args(argv) @@ -880,17 +854,6 @@ def main(argv: list[str] | None = None) -> int: emit_result(build_pr_context(repo, pr_number), args.format) return 0 - if args.command == "current": - emit_result( - build_current_pr_context(fetch_current_repo(), fetch_current_pr_data()), - args.format, - ) - return 0 - - if args.command == "snapshot": - emit_result(build_pr_snapshot(args.repo, args.pr), args.format) - return 0 - if args.command == "comments": emit_result(build_comments_summary(args.repo, args.pr), args.format) return 0 @@ -912,25 +875,6 @@ def main(argv: list[str] | None = None) -> int: emit_result(build_codecov_summary(args.repo, args.pr), args.format) return 0 - if args.command == "linked-issue": - pr_data = fetch_pr_data(args.repo, args.pr) - issue_number, issue = fetch_linked_issue_bundle(args.repo, pr_data) - issue_comments = ( - fetch_issue_comments(args.repo, issue_number) - if issue_number is not None - else [] - ) - emit_result( - build_linked_issue_result( - pr_number=args.pr, - linked_issue_number=issue_number, - linked_issue=issue, - linked_issue_comments=issue_comments, - ), - args.format, - ) - return 0 - if args.command == "create": emit_result( create_pr( @@ -948,10 +892,6 @@ def main(argv: list[str] | None = None) -> int: post_pr_comment(args.repo, args.pr, args.body_file) return 0 - if args.command == "edit-body": - edit_pr_body(args.repo, args.pr, args.body_file) - return 0 - raise AssertionError(f"Unhandled command: {args.command}") diff --git a/scripts/pipeline_skill_context.py b/scripts/pipeline_skill_context.py index 7f6e70a80..88c708c68 100644 --- a/scripts/pipeline_skill_context.py +++ b/scripts/pipeline_skill_context.py @@ -1,5 +1,5 @@ #!/usr/bin/env python3 -"""Skill-scoped context bundle CLI skeleton.""" +"""Skill-scoped context bundle CLI (review-implementation packet).""" from __future__ import annotations @@ -15,26 +15,7 @@ # ``from scripts.pipeline_skill_context import ...``). sys.path.insert(0, str(Path(__file__).resolve().parent)) -import pipeline_board # noqa: E402 import pipeline_checks # noqa: E402 -import pipeline_pr # noqa: E402 -import pipeline_worktree # noqa: E402 - - -PROJECT_BOARD_NUMBER = 8 -PROJECT_BOARD_LIMIT = 500 -DEFAULT_REPO = "CodingThrust/problem-reductions" - - -def build_status_result(skill: str, *, status: str, **fields: object) -> dict: - result = { - "skill": skill, - "status": status, - } - for key, value in fields.items(): - if value is not None: - result[key] = value - return result def report_check_status(check: dict | None) -> str: @@ -45,387 +26,6 @@ def report_check_status(check: dict | None) -> str: return "pass" if check.get("ok") else "fail" -def first_paragraph(text: str | None) -> str: - if not text: - return "" - paragraphs = [chunk.strip() for chunk in text.split("\n\n") if chunk.strip()] - if not paragraphs: - return "" - return " ".join(paragraphs[0].split()) - - -def scan_existing_problems(repo_root: str | Path) -> set[str]: - problem_names: set[str] = set() - models_root = Path(repo_root) / "src/models" - if not models_root.exists(): - return problem_names - - for path in sorted(models_root.rglob("*.rs")): - text = path.read_text() - for match in pipeline_checks.re.finditer( - r"\bpub\s+(?:struct|enum)\s+([A-Z][A-Za-z0-9_]*)\b", - text, - ): - problem_names.add(match.group(1)) - return problem_names - - -def review_pipeline_suggested_mode(result: dict) -> str: - status = result.get("status") - if status == "empty": - return "empty" - if status == "needs-user-choice": - return "needs-user-choice" - - merge_status = ((result.get("prep") or {}).get("merge") or {}).get("status") - if merge_status == "conflicted": - return "conflicted-fix" - if merge_status == "aborted": - return "manual-followup" - - ci_state = ((result.get("pr") or {}).get("ci") or {}).get("state") - if ci_state == "failure": - return "fix-ci" - return "normal-fix" - - -def review_pipeline_seed_items(result: dict) -> list[str]: - blockers: list[str] = [] - prep = result.get("prep") or {} - merge_status = (prep.get("merge") or {}).get("status") - if merge_status == "conflicted": - blockers.append("merge conflicts with main") - elif merge_status == "aborted": - blockers.append("merge prep aborted") - - pr = result.get("pr") or {} - ci_state = (pr.get("ci") or {}).get("state") - if ci_state == "failure": - blockers.append("CI is failing") - - comment_counts = (pr.get("comments") or {}).get("counts") or {} - human_count = sum( - int(comment_counts.get(key, 0)) - for key in [ - "human_inline_comments", - "human_issue_comments", - "human_linked_issue_comments", - "human_reviews", - ] - ) - if human_count: - blockers.append(f"{human_count} human review items to audit") - - deduped: list[str] = [] - for blocker in blockers: - if blocker not in deduped: - deduped.append(blocker) - return deduped - - -def render_review_pipeline_text(result: dict) -> str: - lines = [ - "# Review Pipeline Packet", - "", - "## Selection", - f"- Bundle status: {result.get('status')}", - ] - - if result.get("status") == "empty": - lines.append("- No eligible review-pipeline item is currently available.") - return "\n".join(lines) + "\n" - - if result.get("status") == "needs-user-choice": - lines.extend( - [ - "", - "## Ambiguous PR Options", - ] - ) - for option in result.get("options") or []: - lines.append( - f"- PR #{option.get('number')} [{option.get('state', 'UNKNOWN')}] {option.get('title') or ''}".rstrip() - ) - if result.get("recommendation") is not None: - lines.append(f"- Recommended PR: #{result['recommendation']}") - return "\n".join(lines) + "\n" - - selection = result.get("selection") or {} - pr = result.get("pr") or {} - prep = result.get("prep") or {} - comments = pr.get("comments") or {} - counts = comments.get("counts") or {} - ci = pr.get("ci") or {} - codecov = pr.get("codecov") or {} - checkout = prep.get("checkout") or {} - merge = prep.get("merge") or {} - - if selection.get("pr_number") is not None: - lines.append(f"- PR: #{selection['pr_number']}") - if selection.get("item_id"): - lines.append(f"- Board item: `{selection['item_id']}`") - if selection.get("issue_number") is not None: - lines.append(f"- Linked issue: #{selection['issue_number']}") - if pr.get("title") or selection.get("title"): - lines.append(f"- Title: {pr.get('title') or selection.get('title')}") - if pr.get("url"): - lines.append(f"- URL: {pr['url']}") - - lines.extend( - [ - "", - "## Recommendation Seed", - f"- Suggested mode: {review_pipeline_suggested_mode(result)}", - ] - ) - seed_items = review_pipeline_seed_items(result) - if seed_items: - lines.append("- Attention points:") - lines.extend(f" - {item}" for item in seed_items) - else: - lines.append("- Attention points: none from deterministic checks") - - lines.extend( - [ - "", - "## Comment Summary", - f"- Human inline comments: {counts.get('human_inline_comments', 0)}", - f"- Human PR issue comments: {counts.get('human_issue_comments', 0)}", - f"- Human linked-issue comments: {counts.get('human_linked_issue_comments', 0)}", - f"- Human review bodies: {counts.get('human_reviews', 0)}", - ] - ) - - lines.extend( - [ - "", - "## CI / Coverage", - f"- CI state: {ci.get('state', 'unknown')}", - ] - ) - if ci: - lines.append(f"- Failing checks: {ci.get('failing', 0)}") - lines.append(f"- Pending checks: {ci.get('pending', 0)}") - if codecov.get("found"): - lines.append(f"- Patch coverage: {codecov.get('patch_coverage')}%") - if codecov.get("project_coverage") is not None: - lines.append(f"- Project coverage: {codecov.get('project_coverage')}%") - - lines.extend( - [ - "", - "## Merge Prep", - f"- Ready: {str(prep.get('ready')).lower()}", - f"- Merge status: {merge.get('status', 'unknown')}", - ] - ) - if checkout.get("worktree_dir"): - lines.append(f"- Worktree: `{checkout['worktree_dir']}`") - if checkout.get("head_ref_name"): - lines.append(f"- PR head branch: `{checkout['head_ref_name']}`") - conflicts = merge.get("conflicts") or [] - if conflicts: - lines.append("- Conflicts:") - lines.extend(f" - `{conflict}`" for conflict in conflicts) - - if pr.get("issue_context_text"): - lines.extend( - [ - "", - "## Linked Issue Context", - pr["issue_context_text"], - ] - ) - - return "\n".join(lines) + "\n" - - -def final_review_suggested_mode(result: dict) -> str: - status = result.get("status") - if status == "empty": - return "empty" - if status == "ready-with-warnings": - return "warning-fallback" - - merge_status = ((result.get("prep") or {}).get("merge") or {}).get("status") - if merge_status == "conflicted": - return "conflicted-review" - if merge_status == "aborted": - return "warning-fallback" - return "normal-review" - - -def final_review_seed_items(result: dict) -> list[str]: - review_context = result.get("review_context") or {} - prep = result.get("prep") or {} - warnings = list(result.get("warnings") or []) - blockers = list(warnings) - - merge_status = (prep.get("merge") or {}).get("status") - if merge_status == "conflicted": - blockers.append("merge conflicts with main") - elif merge_status == "aborted": - blockers.append("merge prep aborted") - - whitelist = review_context.get("whitelist") or {} - if whitelist and not whitelist.get("ok"): - blockers.append("files outside expected whitelist") - - completeness = review_context.get("completeness") or {} - for missing in completeness.get("missing", []): - blockers.append(f"missing completeness item: {missing}") - - comment_counts = ((result.get("pr") or {}).get("comments") or {}).get("counts") or {} - manual_comment_count = sum( - int(comment_counts.get(key, 0)) - for key in [ - "human_inline_comments", - "human_issue_comments", - "human_linked_issue_comments", - "human_reviews", - ] - ) - if manual_comment_count: - blockers.append( - f"manual comment audit required for {manual_comment_count} human review items" - ) - - deduped: list[str] = [] - for blocker in blockers: - if blocker not in deduped: - deduped.append(blocker) - return deduped - - -def render_final_review_text(result: dict) -> str: - selection = result.get("selection") or {} - pr = result.get("pr") or {} - prep = result.get("prep") or {} - review_context = result.get("review_context") or {} - subject = review_context.get("subject") or {} - comments = pr.get("comments") or {} - counts = comments.get("counts") or {} - checkout = prep.get("checkout") or {} - merge = prep.get("merge") or {} - - lines = [ - "# Final Review Packet", - "", - "## Selection", - f"- Bundle status: {result.get('status')}", - ] - if selection.get("pr_number") is not None: - lines.append(f"- PR: #{selection['pr_number']}") - if selection.get("item_id"): - lines.append(f"- Board item: `{selection['item_id']}`") - if selection.get("issue_number") is not None: - lines.append(f"- Linked issue: #{selection['issue_number']}") - if pr.get("title") or selection.get("title"): - lines.append(f"- Title: {pr.get('title') or selection.get('title')}") - if pr.get("url"): - lines.append(f"- URL: {pr['url']}") - - lines.extend( - [ - "", - "## Recommendation Seed", - f"- Suggested mode: {final_review_suggested_mode(result)}", - ] - ) - seed_items = final_review_seed_items(result) - if seed_items: - lines.append("- Review blockers / attention points:") - lines.extend(f" - {item}" for item in seed_items) - else: - lines.append("- Review blockers / attention points: none from deterministic checks") - - lines.extend( - [ - "", - "## Subject", - f"- Kind: {subject.get('kind', 'unknown')}", - ] - ) - if subject.get("name"): - lines.append(f"- Name: {subject['name']}") - if subject.get("source"): - lines.append(f"- Source: {subject['source']}") - if subject.get("target"): - lines.append(f"- Target: {subject['target']}") - - lines.extend( - [ - "", - "## Comment Summary", - f"- Human reviews: {counts.get('human_reviews', 0)}", - f"- Human inline comments: {counts.get('human_inline_comments', 0)}", - f"- Human PR issue comments: {counts.get('human_issue_comments', 0)}", - f"- Human linked-issue comments: {counts.get('human_linked_issue_comments', 0)}", - ] - ) - if pr.get("issue_context_text"): - lines.extend( - [ - "", - "### Linked Issue Context", - pr["issue_context_text"], - ] - ) - - lines.extend( - [ - "", - "## Merge Prep", - f"- Ready: {str(prep.get('ready')).lower()}", - f"- Merge status: {merge.get('status', 'unknown')}", - ] - ) - if checkout.get("worktree_dir"): - lines.append(f"- Worktree: `{checkout['worktree_dir']}`") - conflicts = merge.get("conflicts") or [] - if conflicts: - lines.append("- Conflicts:") - lines.extend(f" - `{conflict}`" for conflict in conflicts) - warnings = result.get("warnings") or [] - if warnings: - lines.append("- Warnings:") - lines.extend(f" - {warning}" for warning in warnings) - - lines.extend( - [ - "", - "## Deterministic Checks", - f"- Whitelist: {report_check_status(review_context.get('whitelist'))}", - f"- Completeness: {report_check_status(review_context.get('completeness'))}", - ] - ) - missing = (review_context.get("completeness") or {}).get("missing") or [] - if missing: - lines.append("- Missing items:") - lines.extend(f" - `{item}`" for item in missing) - - changed_files = review_context.get("changed_files") or [] - lines.extend(["", "## Changed Files"]) - if changed_files: - lines.extend(f"- `{path}`" for path in changed_files) - else: - lines.append("- None captured") - - diff_stat = review_context.get("diff_stat") - if diff_stat: - lines.extend(["", "## Diff Stat", "```text", diff_stat, "```"]) - - full_diff = review_context.get("full_diff") - if full_diff: - lines.extend(["", "## Full Diff", "```diff", full_diff, "```"]) - - pred_list = review_context.get("pred_list") - if pred_list: - lines.extend(["", "## Problem Catalog (`pred list`)", "```text", pred_list, "```"]) - - return "\n".join(lines) + "\n" - - def render_review_implementation_text(result: dict) -> str: git = result.get("git") or {} review_context = result.get("review_context") or {} @@ -492,78 +92,9 @@ def render_review_implementation_text(result: dict) -> str: return "\n".join(lines) + "\n" -def render_project_pipeline_text(result: dict) -> str: - ready_issues = result.get("ready_issues") or [] - eligible = [issue for issue in ready_issues if issue.get("eligible")] - blocked = [issue for issue in ready_issues if not issue.get("eligible")] - in_progress = result.get("in_progress_issues") or [] - requested = result.get("requested_issue") - - lines = [ - "# Project Pipeline Packet", - "", - "## Queue Summary", - f"- Bundle status: {result.get('status')}", - f"- Ready issues: {len(ready_issues)}", - f"- Eligible ready issues: {len(eligible)}", - f"- Blocked ready issues: {len(blocked)}", - f"- In progress issues: {len(in_progress)}", - f"- Existing problems on main: {len(result.get('existing_problems') or [])}", - ] - - if requested is not None: - lines.extend( - [ - "", - "## Requested Issue", - f"- Issue: #{requested.get('issue_number')}", - f"- Title: {requested.get('title') or 'unknown'}", - f"- Eligible: {str(bool(requested.get('eligible'))).lower()}", - ] - ) - if requested.get("blocking_reason"): - lines.append(f"- Blocking reason: {requested['blocking_reason']}") - - lines.extend(["", "## Eligible Ready Issues"]) - if eligible: - for issue in eligible: - lines.append(f"- #{issue.get('issue_number')} {issue.get('title')}") - lines.append(f" - Kind: {issue.get('kind', 'unknown')}") - lines.append( - f" - Pending rules unblocked: {issue.get('pending_rule_count', 0)}" - ) - if issue.get("summary"): - lines.append(f" - Summary: {issue['summary']}") - else: - lines.append("- None") - - lines.extend(["", "## Blocked Ready Issues"]) - if blocked: - for issue in blocked: - lines.append(f"- #{issue.get('issue_number')} {issue.get('title')}") - lines.append(f" - Blocking reason: {issue.get('blocking_reason')}") - if issue.get("summary"): - lines.append(f" - Summary: {issue['summary']}") - else: - lines.append("- None") - - if in_progress: - lines.extend(["", "## In Progress Issues"]) - for issue in in_progress: - lines.append(f"- #{issue.get('issue_number')} {issue.get('title')}") - - return "\n".join(lines) + "\n" - - def render_text(result: dict) -> str: - if result.get("skill") == "review-pipeline": - return render_review_pipeline_text(result) - if result.get("skill") == "final-review": - return render_final_review_text(result) if result.get("skill") == "review-implementation": return render_review_implementation_text(result) - if result.get("skill") == "project-pipeline": - return render_project_pipeline_text(result) return json.dumps(result, indent=2, sort_keys=True) + "\n" @@ -574,106 +105,6 @@ def emit_result(result: dict, fmt: str) -> None: print(json.dumps(result, indent=2, sort_keys=True)) -def fetch_review_candidates(repo: str) -> list[dict]: - owner = repo.split("/", 1)[0] - board_data = pipeline_board.fetch_board_items( - owner, - PROJECT_BOARD_NUMBER, - PROJECT_BOARD_LIMIT, - ) - return pipeline_board.review_candidates( - board_data, - repo, - pipeline_board.resolve_issue_pr, - pipeline_board.fetch_pr_info, - batch_pr_fetcher=pipeline_board.batch_fetch_prs_with_reviews, - ) - - -def build_ready_result(*, skill: str, selection: dict, pr: dict, prep: dict) -> dict: - return build_status_result( - skill, - status="ready", - selection=selection, - pr=pr, - prep=prep, - ) - - -def build_ambiguous_selection(candidate: dict, *, pr_number: int) -> dict: - return { - "item_id": candidate["item_id"], - "number": pr_number, - "issue_number": candidate.get("issue_number"), - "pr_number": pr_number, - "status": candidate.get("status"), - "title": candidate.get("title"), - } - - -def select_final_review_entry( - *, - repo: str, - pr_number: int | None, -) -> dict | None: - """Find a Final-review board entry for the given PR (or the first available one).""" - owner = repo.split("/", 1)[0] - board_data = pipeline_board.fetch_board_items( - owner, - PROJECT_BOARD_NUMBER, - PROJECT_BOARD_LIMIT, - ) - items = [ - item - for item in board_data.get("items", []) - if item.get("status") == pipeline_board.STATUS_FINAL_REVIEW - ] - for item in items: - content = item.get("content") or {} - number = content.get("number") - if number is None: - continue - item_type = content.get("type", "") - if item_type == "PullRequest": - item_pr = int(number) - elif item_type == "Issue": - item_pr = pipeline_board.resolve_issue_pr(repo, int(number)) - if item_pr is None: - continue - else: - continue - entry = { - "item_id": pipeline_board.item_identity(item), - "number": item_pr, - "pr_number": item_pr, - "status": item.get("status"), - "title": content.get("title"), - } - if item_type == "Issue": - entry["issue_number"] = int(number) - if pr_number is not None: - if item_pr == pr_number: - return entry - else: - state = pipeline_board.fetch_pr_state(repo, item_pr) - if state == "OPEN": - return entry - return None - - -def _get_current_gh_user() -> str: - """Return the GitHub login of the currently authenticated user.""" - try: - output = subprocess.check_output( - ["gh", "api", "user", "--jq", ".login"], - text=True, - stderr=subprocess.DEVNULL, - ) - return output.strip() - except Exception: - return "" - - def git_output_in(repo_root: str | Path, *args: str) -> list[str]: output = subprocess.check_output( ["git", "-C", str(repo_root), *args], @@ -689,83 +120,6 @@ def git_text_in(repo_root: str | Path, *args: str) -> str: ) -def infer_final_review_subject(scope: dict, pr_context: dict) -> dict: - subject = pipeline_checks.infer_review_subject(scope) - linked_issue = pr_context.get("linked_issue") or {} - linked_title = (linked_issue.get("title") or "").strip() - - rule_match = pipeline_checks.RULE_TITLE_RE.match(linked_title) - if rule_match: - subject["kind"] = "rule" - subject["source"] = rule_match.group("source") - subject["target"] = rule_match.group("target") - return subject - - model_match = pipeline_checks.MODEL_TITLE_RE.match(linked_title) - if model_match: - subject["kind"] = "model" - subject["name"] = subject.get("name") or model_match.group("name") - subject["source"] = None - subject["target"] = None - - return subject - - -def build_final_review_checks(*, prep: dict, pr_context: dict) -> dict: - checkout = prep.get("checkout") or {} - worktree_dir = checkout.get("worktree_dir") - base_sha = checkout.get("base_sha") - head_sha = checkout.get("head_sha") - if not worktree_dir or not base_sha or not head_sha: - raise ValueError("prepare-review output missing checkout diff range") - - diff_range = f"{base_sha}..{head_sha}" - changed_files = git_output_in(worktree_dir, "diff", "--name-only", diff_range) - added_files = git_output_in( - worktree_dir, - "diff", - "--name-only", - "--diff-filter=A", - diff_range, - ) - scope = pipeline_checks.detect_scope_from_paths( - added_files=added_files, - changed_files=changed_files, - ) - subject = infer_final_review_subject(scope, pr_context) - review_context = pipeline_checks.build_review_context( - worktree_dir, - diff_stat=git_text_in(worktree_dir, "diff", "--stat", diff_range), - scope=scope, - subject=subject, - ) - review_context["full_diff"] = git_text_in(worktree_dir, "diff", diff_range) - review_context["pred_list"] = _run_pred_list(worktree_dir) - return review_context - - -def _run_pred_list(worktree_dir: str | Path) -> str | None: - """Run ``pred list`` in *worktree_dir*, building the CLI first if needed.""" - pred_cmd = ["cargo", "run", "-p", "problemreductions-cli", "--bin", "pred", "--", "list"] - try: - return subprocess.check_output( - pred_cmd, cwd=str(worktree_dir), text=True, stderr=subprocess.DEVNULL, - ) - except (subprocess.CalledProcessError, FileNotFoundError): - pass - # Binary may not exist yet — build it and retry. - try: - subprocess.check_call( - ["make", "cli"], cwd=str(worktree_dir), - stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, - ) - return subprocess.check_output( - pred_cmd, cwd=str(worktree_dir), text=True, stderr=subprocess.DEVNULL, - ) - except Exception: - return None - - def default_review_implementation_context_builder( repo_root: str | Path, *, @@ -872,375 +226,6 @@ def build_review_implementation_context( } -def classify_project_issue( - entry: dict, - *, - issue: dict, - existing_problems: set[str], - pending_rule_counts: dict[str, int], -) -> dict: - kind, source_problem, target_problem = pipeline_checks.issue_kind_from_title( - entry.get("title") - ) - blocking_reason = None - eligible = True - if kind == "rule": - missing = [ - problem - for problem in [source_problem, target_problem] - if problem and problem not in existing_problems - ] - if missing: - eligible = False - blocking_reason = f'model "{missing[0]}" not yet implemented on main' - - issue_number = int(entry["issue_number"]) - return { - "item_id": entry.get("item_id"), - "issue_number": issue_number, - "title": entry.get("title"), - "kind": kind, - "source_problem": source_problem, - "target_problem": target_problem, - "eligible": eligible, - "blocking_reason": blocking_reason, - "pending_rule_count": pending_rule_counts.get(entry.get("title", ""), 0) - if kind == "rule" - else pending_rule_counts.get( - pipeline_checks.MODEL_TITLE_RE.match(entry.get("title", "")).group("name") - if pipeline_checks.MODEL_TITLE_RE.match(entry.get("title", "")) - else "", - 0, - ), - "summary": first_paragraph(issue.get("body")), - "issue": issue, - } - - -def build_pending_rule_counts( - ready_entries: list[dict], - in_progress_entries: list[dict], -) -> dict[str, int]: - counts: dict[str, int] = {} - for entry in [*ready_entries, *in_progress_entries]: - kind, source_problem, target_problem = pipeline_checks.issue_kind_from_title( - entry.get("title") - ) - if kind != "rule": - continue - for problem in [source_problem, target_problem]: - if not problem: - continue - counts[problem] = counts.get(problem, 0) + 1 - return counts - - -def fetch_project_board_data(repo: str) -> dict: - owner = repo.split("/", 1)[0] - return pipeline_board.fetch_board_items( - owner, - PROJECT_BOARD_NUMBER, - PROJECT_BOARD_LIMIT, - ) - - -def build_project_pipeline_context( - *, - repo: str, - issue_number: int | None, - repo_root: Path, - board_fetcher: Callable[[str], dict] | None = None, - issue_fetcher: Callable[[str, int], dict] | None = None, - batch_issue_fetcher: Callable[[str, list[int]], dict[int, dict]] | None = None, - existing_problem_finder: Callable[[Path], set[str]] | None = None, -) -> dict: - board_fetcher = board_fetcher or fetch_project_board_data - _custom_issue_fetcher = issue_fetcher is not None - issue_fetcher = issue_fetcher or pipeline_checks.fetch_issue - # Only use batch fetcher when no custom per-item fetcher was injected (e.g. tests) - if batch_issue_fetcher is None and not _custom_issue_fetcher: - batch_issue_fetcher = pipeline_board.batch_fetch_issues - existing_problem_finder = existing_problem_finder or scan_existing_problems - - board_data = board_fetcher(repo) - ready_entries = sorted( - pipeline_board.ready_entries(board_data).values(), - key=lambda entry: entry["issue_number"], - ) - in_progress_entries = pipeline_board.status_items( - board_data, - pipeline_board.STATUS_IN_PROGRESS, - ) - existing_problems = existing_problem_finder(repo_root) - pending_rule_counts = build_pending_rule_counts(ready_entries, in_progress_entries) - - ready_entries_items = sorted( - pipeline_board.ready_entries(board_data).items(), - key=lambda pair: pair[1]["issue_number"], - ) - - # Batch-fetch all issue data in one API call when batch fetcher is available - if batch_issue_fetcher is not None: - all_issue_numbers = [int(entry["issue_number"]) for _, entry in ready_entries_items] - issues_cache = batch_issue_fetcher(repo, all_issue_numbers) - - def _fetch_one(repo: str, n: int) -> dict: - if n in issues_cache: - return issues_cache[n] - return issue_fetcher(repo, n) - else: - _fetch_one = issue_fetcher - - ready_issues = [ - classify_project_issue( - dict(entry, item_id=item_id), - issue=_fetch_one(repo, int(entry["issue_number"])), - existing_problems=existing_problems, - pending_rule_counts=pending_rule_counts, - ) - for item_id, entry in ready_entries_items - ] - - requested_issue = None - if issue_number is not None: - requested_issue = next( - ( - issue - for issue in ready_issues - if int(issue["issue_number"]) == issue_number - ), - None, - ) - - eligible_ready_issues = [issue for issue in ready_issues if issue.get("eligible")] - - if not ready_issues: - status = "empty" - elif issue_number is not None and requested_issue is None: - status = "requested-missing" - elif requested_issue is not None and not requested_issue.get("eligible"): - status = "requested-blocked" - elif not eligible_ready_issues: - status = "no-eligible-issues" - else: - status = "ready" - - return build_status_result( - "project-pipeline", - status=status, - repo=repo, - existing_problems=sorted(existing_problems), - ready_issues=ready_issues, - in_progress_issues=in_progress_entries, - requested_issue=requested_issue, - ) - - -def _select_candidate( - candidates: list[dict], - pr_number: int | None, -) -> dict | None: - """Pick a candidate from the review-candidates list (no state file, no claiming).""" - if pr_number is not None: - return next( - ( - c - for c in candidates - if int(c.get("pr_number") or c.get("number") or -1) == pr_number - ), - None, - ) - eligible = [c for c in candidates if c.get("eligibility") == "eligible"] - return eligible[0] if eligible else None - - -def _selection_from_candidate(candidate: dict) -> dict: - """Build a selection dict from a candidate (read-only, no board move).""" - item_id = str(candidate["item_id"]) - return { - "item_id": item_id, - "number": int(candidate.get("pr_number") or candidate["number"]), - "issue_number": candidate.get("issue_number"), - "pr_number": int(candidate.get("pr_number") or candidate["number"]), - "status": candidate.get("status"), - "title": candidate.get("title"), - } - - -def build_review_pipeline_context( - *, - repo: str, - pr_number: int | None, - review_candidate_fetcher: Callable[[str], list[dict]] | None = None, - pr_context_builder: Callable[[str, int], dict] | None = None, - review_preparer: Callable[[str, int], dict] | None = None, -) -> dict: - """Build review-pipeline context (read-only, no board move). - - The agent is responsible for claiming the item (moving to Under review) - after it has verified the PR is review-ready and is about to start work. - """ - review_candidate_fetcher = review_candidate_fetcher or fetch_review_candidates - pr_context_builder = pr_context_builder or pipeline_pr.build_pr_context - review_preparer = review_preparer or ( - lambda repo, pr_number: pipeline_worktree.prepare_review_from_cwd( - repo=repo, - pr_number=pr_number, - ) - ) - - candidates = review_candidate_fetcher(repo) - if not candidates: - return build_status_result("review-pipeline", status="empty") - - # Check for ambiguous cards first - if pr_number is None: - ambiguous = next( - (c for c in candidates if c.get("eligibility") == "ambiguous-linked-prs"), - None, - ) - if ambiguous is not None: - return build_status_result( - "review-pipeline", - status="needs-user-choice", - options=ambiguous.get("linked_repo_prs", []), - recommendation=ambiguous.get("recommendation"), - ) - else: - matching_ambiguous = next( - ( - c - for c in candidates - if c.get("eligibility") == "ambiguous-linked-prs" - and any( - int(option["number"]) == pr_number - for option in c.get("linked_repo_prs", []) - ) - ), - None, - ) - if matching_ambiguous is not None: - selection = build_ambiguous_selection(matching_ambiguous, pr_number=pr_number) - return build_ready_result( - skill="review-pipeline", - selection=selection, - pr=pr_context_builder(repo, pr_number), - prep=review_preparer(repo, pr_number), - ) - - candidate = _select_candidate(candidates, pr_number) - if candidate is None: - return build_status_result("review-pipeline", status="empty") - - if candidate.get("eligibility") != "eligible": - return build_status_result("review-pipeline", status="empty") - - selection = _selection_from_candidate(candidate) - selected_pr_number = int(selection["pr_number"]) - return build_ready_result( - skill="review-pipeline", - selection=selection, - pr=pr_context_builder(repo, selected_pr_number), - prep=review_preparer(repo, selected_pr_number), - ) - - -def build_final_review_context( - *, - repo: str, - pr_number: int | None, - selection_fetcher: Callable[..., dict | None] | None = None, - pr_context_builder: Callable[[str, int], dict] | None = None, - review_preparer: Callable[[str, int], dict] | None = None, - review_context_builder: Callable[..., dict] | None = None, -) -> dict: - selection_fetcher = selection_fetcher or select_final_review_entry - pr_context_builder = pr_context_builder or pipeline_pr.build_pr_context - review_preparer = review_preparer or ( - lambda repo, pr_number: pipeline_worktree.prepare_review_from_cwd( - repo=repo, - pr_number=pr_number, - ) - ) - review_context_builder = review_context_builder or build_final_review_checks - - selection = selection_fetcher( - repo=repo, - pr_number=pr_number, - ) - if selection is None: - return build_status_result("final-review", status="empty") - - selected_pr_number = int(selection.get("pr_number") or selection["number"]) - pr_context = pr_context_builder(repo, selected_pr_number) - - # Self-review warning: flag if reviewer is the PR author (unless repo owner). - pr_author = (pr_context.get("author") or "").lower() - current_user = _get_current_gh_user().lower() - repo_owner = repo.split("/", 1)[0].lower() if "/" in repo else "" - self_review_warning = None - if pr_author and current_user and pr_author == current_user and current_user != repo_owner: - self_review_warning = f"Self-review: PR author '{pr_author}' is the current reviewer" - - prep: dict - try: - prep = review_preparer(repo, selected_pr_number) - except Exception as exc: - return { - "skill": "final-review", - "status": "ready-with-warnings", - "selection": selection, - "pr": pr_context, - "prep": { - "ready": False, - "error": str(exc), - }, - "review_context": None, - "warnings": [ - f"failed to prepare final-review worktree: {exc}", - ], - } - - try: - review_context = review_context_builder( - prep=prep, - pr_context=pr_context, - ) - except Exception as exc: - return { - "skill": "final-review", - "status": "ready-with-warnings", - "selection": selection, - "pr": pr_context, - "prep": prep, - "review_context": None, - "warnings": [ - f"failed to derive final-review review context: {exc}", - ], - } - - warnings = [self_review_warning] if self_review_warning else [] - return build_status_result( - "final-review", - status="ready", - selection=selection, - pr=pr_context, - prep=prep, - review_context=review_context, - warnings=warnings or None, - ) - - -def add_bundle_parser( - subparsers, - command: str, -) -> None: - parser = subparsers.add_parser(command) - parser.add_argument("--repo", required=True) - parser.add_argument("--pr", type=int) - parser.add_argument("--format", choices=["json", "text"], default="json") - - def add_review_implementation_parser(subparsers) -> None: parser = subparsers.add_parser("review-implementation") parser.add_argument("--repo-root", type=Path, default=Path(".")) @@ -1251,22 +236,11 @@ def add_review_implementation_parser(subparsers) -> None: parser.add_argument("--format", choices=["json", "text"], default="json") -def add_project_pipeline_parser(subparsers) -> None: - parser = subparsers.add_parser("project-pipeline") - parser.add_argument("--repo", default=DEFAULT_REPO) - parser.add_argument("--issue", type=int) - parser.add_argument("--repo-root", type=Path, default=Path(".")) - parser.add_argument("--format", choices=["json", "text"], default="json") - - def parse_args(argv: list[str]) -> argparse.Namespace: - parser = argparse.ArgumentParser(description="Skill-scoped pipeline context bundles.") + parser = argparse.ArgumentParser(description="Skill-scoped context bundles.") subparsers = parser.add_subparsers(dest="command", required=True) - add_bundle_parser(subparsers, "review-pipeline") - add_bundle_parser(subparsers, "final-review") add_review_implementation_parser(subparsers) - add_project_pipeline_parser(subparsers) return parser.parse_args(argv) @@ -1274,26 +248,6 @@ def parse_args(argv: list[str]) -> argparse.Namespace: def main(argv: list[str] | None = None) -> int: args = parse_args(argv or sys.argv[1:]) - if args.command == "review-pipeline": - emit_result( - build_review_pipeline_context( - repo=args.repo, - pr_number=args.pr, - ), - args.format, - ) - return 0 - - if args.command == "final-review": - emit_result( - build_final_review_context( - repo=args.repo, - pr_number=args.pr, - ), - args.format, - ) - return 0 - if args.command == "review-implementation": emit_result( build_review_implementation_context( @@ -1307,17 +261,6 @@ def main(argv: list[str] | None = None) -> int: ) return 0 - if args.command == "project-pipeline": - emit_result( - build_project_pipeline_context( - repo=args.repo, - issue_number=args.issue, - repo_root=args.repo_root, - ), - args.format, - ) - return 0 - raise AssertionError(f"Unhandled command: {args.command}") diff --git a/scripts/pipeline_worktree.py b/scripts/pipeline_worktree.py deleted file mode 100644 index a32005f4b..000000000 --- a/scripts/pipeline_worktree.py +++ /dev/null @@ -1,519 +0,0 @@ -#!/usr/bin/env python3 -"""Shared worktree helpers for issue and PR pipeline flows.""" - -from __future__ import annotations - -import argparse -import json -import re -import subprocess -import sys -from pathlib import Path - - -def sanitize_component(text: str) -> str: - normalized = re.sub(r"[^A-Za-z0-9]+", "-", text.strip().lower()).strip("-") - return normalized or "work" - - -def plan_issue_worktree( - repo_root: str | Path, - *, - issue_number: int, - slug: str, - base_ref: str = "origin/main", -) -> dict: - repo_root = str(Path(repo_root)) - branch = f"issue-{issue_number}-{sanitize_component(slug)}" - worktree_dir = str(Path(repo_root) / ".worktrees" / branch) - return { - "issue_number": issue_number, - "slug": slug, - "branch": branch, - "worktree_dir": worktree_dir, - "base_ref": base_ref, - } - - -def plan_pr_worktree( - repo_root: str | Path, - *, - pr_number: int, - head_ref_name: str, - base_sha: str, - head_sha: str, -) -> dict: - repo_root = str(Path(repo_root)) - local_branch = f"review-pr-{pr_number}-{sanitize_component(head_ref_name)}" - worktree_dir = str(Path(repo_root) / ".worktrees" / local_branch) - return { - "pr_number": pr_number, - "head_ref_name": head_ref_name, - "local_branch": local_branch, - "worktree_dir": worktree_dir, - "fetch_ref": f"pull/{pr_number}/head:{local_branch}", - "base_sha": base_sha, - "head_sha": head_sha, - } - - -def summarize_merge( - *, - worktree: str | Path, - exit_code: int, - conflicts: list[str], -) -> dict: - conflicts = sorted(conflicts) - if exit_code == 0: - status = "clean" - elif conflicts: - status = "conflicted" - else: - status = "aborted" - - likely_complex = len(conflicts) > 1 or any( - path.startswith(".claude/skills/add-model/") - or path.startswith(".claude/skills/add-rule/") - for path in conflicts - ) - - return { - "worktree": str(worktree), - "status": status, - "conflicts": conflicts, - "likely_complex": likely_complex, - } - - -def run_git(repo_root: str | Path, *args: str) -> str: - return subprocess.check_output(["git", "-C", str(repo_root), *args], text=True) - - -def run_git_checked(repo_root: str | Path, *args: str) -> None: - subprocess.check_output(["git", "-C", str(repo_root), *args], stderr=subprocess.STDOUT) - - -def run_gh_json(*args: str): - return json.loads(subprocess.check_output(["gh", *args], text=True)) - - -def repo_root_from(path: str | Path) -> Path: - return Path(run_git(path, "rev-parse", "--show-toplevel").strip()) - - -def branch_exists(repo_root: str | Path, branch: str) -> bool: - proc = subprocess.run( - ["git", "-C", str(repo_root), "rev-parse", "--verify", branch], - capture_output=True, - text=True, - ) - return proc.returncode == 0 - - -def prepare_issue_branch( - *, - issue_number: int, - slug: str, - base_ref: str = "main", - repo_root: str | Path | None = None, -) -> dict: - repo_root = Path(repo_root or repo_root_from(Path.cwd())).resolve() - plan = plan_issue_worktree( - repo_root, - issue_number=issue_number, - slug=slug, - base_ref=base_ref, - ) - - status_output = run_git(repo_root, "status", "--porcelain").strip() - if status_output: - raise RuntimeError("working tree is dirty; stash or commit changes before branching") - - run_git_checked(repo_root, "checkout", base_ref) - existing_branch = branch_exists(repo_root, plan["branch"]) - if existing_branch: - run_git_checked(repo_root, "checkout", plan["branch"]) - action = "checkout-existing" - else: - run_git_checked(repo_root, "checkout", "-b", plan["branch"]) - action = "create-branch" - - base_sha = run_git(repo_root, "rev-parse", base_ref).strip() - head_sha = run_git(repo_root, "rev-parse", "HEAD").strip() - return { - **plan, - "existing_branch": existing_branch, - "action": action, - "base_sha": base_sha, - "head_sha": head_sha, - } - - -def create_issue_worktree( - *, - issue_number: int, - slug: str, - base_ref: str = "origin/main", - repo_root: str | Path | None = None, -) -> dict: - repo_root = Path(repo_root or repo_root_from(Path.cwd())).resolve() - plan = plan_issue_worktree( - repo_root, - issue_number=issue_number, - slug=slug, - base_ref=base_ref, - ) - - Path(plan["worktree_dir"]).parent.mkdir(parents=True, exist_ok=True) - remote, _, branch_name = base_ref.partition("/") - if remote and branch_name: - run_git_checked(repo_root, "fetch", remote, branch_name) - run_git_checked( - repo_root, - "worktree", - "add", - plan["worktree_dir"], - "-b", - plan["branch"], - base_ref, - ) - - base_sha = run_git(repo_root, "rev-parse", base_ref).strip() - head_sha = run_git(plan["worktree_dir"], "rev-parse", "HEAD").strip() - return { - **plan, - "base_sha": base_sha, - "head_sha": head_sha, - } - - -def checkout_pr_worktree( - *, - repo: str, - pr_number: int, - repo_root: str | Path | None = None, -) -> dict: - repo_root = Path(repo_root or repo_root_from(Path.cwd())).resolve() - pr_data = run_gh_json( - "pr", - "view", - str(pr_number), - "--repo", - repo, - "--json", - "headRefName,headRefOid,baseRefName", - ) - - # baseRefOid is not available via gh pr view --json; resolve it locally - run_git_checked(repo_root, "fetch", "origin", pr_data["baseRefName"]) - base_sha = run_git(repo_root, "rev-parse", f"origin/{pr_data['baseRefName']}").strip() - - plan = plan_pr_worktree( - repo_root, - pr_number=pr_number, - head_ref_name=pr_data["headRefName"], - base_sha=base_sha, - head_sha=pr_data["headRefOid"], - ) - - worktree_dir = Path(plan["worktree_dir"]) - - # If the worktree already exists from a previous run, remove it first - if worktree_dir.exists(): - run_git_checked(repo_root, "worktree", "remove", "--force", str(worktree_dir)) - # Also clean up the local branch if it exists (may be left over after worktree removal) - branch_check = subprocess.run( - ["git", "-C", str(repo_root), "rev-parse", "--verify", plan["local_branch"]], - capture_output=True, - ) - if branch_check.returncode == 0: - subprocess.run( - ["git", "-C", str(repo_root), "branch", "-D", plan["local_branch"]], - capture_output=True, - ) - - worktree_dir.parent.mkdir(parents=True, exist_ok=True) - run_git_checked(repo_root, "fetch", "origin", plan["fetch_ref"]) - run_git_checked(repo_root, "worktree", "add", plan["worktree_dir"], plan["local_branch"]) - return plan - - -def merge_main( - *, - worktree: str | Path, -) -> dict: - worktree = Path(worktree).resolve() - run_git_checked(worktree, "fetch", "origin", "main") - proc = subprocess.run( - ["git", "-C", str(worktree), "merge", "origin/main", "--no-edit"], - text=True, - capture_output=True, - ) - - conflict_output = run_git(worktree, "diff", "--name-only", "--diff-filter=U").strip() - conflicts = [line for line in conflict_output.splitlines() if line] - summary = summarize_merge(worktree=worktree, exit_code=proc.returncode, conflicts=conflicts) - summary["stdout"] = proc.stdout - summary["stderr"] = proc.stderr - return summary - - -def prepare_review( - *, - repo: str, - pr_number: int, - repo_root: str | Path | None = None, -) -> dict: - checkout = checkout_pr_worktree( - repo=repo, - pr_number=pr_number, - repo_root=repo_root, - ) - merge = merge_main(worktree=checkout["worktree_dir"]) - return { - "repo": repo, - "pr_number": pr_number, - "ready": merge["status"] == "clean", - "checkout": checkout, - "merge": merge, - } - - -def enter( - *, - name: str, - base_ref: str = "origin/main", - repo_root: str | Path | None = None, -) -> dict: - """Create a named worktree from base_ref. Idempotent — removes stale worktree/branch if they exist.""" - repo_root = Path(repo_root or repo_root_from(Path.cwd())).resolve() - branch = sanitize_component(name) - worktree_dir = str(repo_root / ".worktrees" / branch) - - # Remove stale worktree if it exists from a previous run - if Path(worktree_dir).exists(): - run_git_checked(repo_root, "worktree", "remove", "--force", worktree_dir) - # Clean up stale branch if it exists - if branch_exists(repo_root, branch): - result = subprocess.run( - ["git", "-C", str(repo_root), "branch", "-D", branch], - capture_output=True, - text=True, - ) - if result.returncode != 0: - raise RuntimeError( - f"Failed to delete stale branch '{branch}': {result.stderr.strip()}" - ) - - # Fetch the base ref - remote, _, branch_name = base_ref.partition("/") - if remote and branch_name: - run_git_checked(repo_root, "fetch", remote, branch_name) - - Path(worktree_dir).parent.mkdir(parents=True, exist_ok=True) - run_git_checked(repo_root, "worktree", "add", worktree_dir, "-b", branch, base_ref) - - return { - "worktree_dir": worktree_dir, - "branch": branch, - "base_ref": base_ref, - } - - -def prepare_review_from_cwd(*, repo: str, pr_number: int) -> dict: - """Build prep data from CWD (assumes PR branch already checked out via EnterWorktree + gh pr checkout).""" - cwd = str(Path.cwd()) - run_git_checked(cwd, "fetch", "origin", "main") - base_sha = run_git(cwd, "merge-base", "origin/main", "HEAD").strip() - head_sha = run_git(cwd, "rev-parse", "HEAD").strip() - return { - "repo": repo, - "pr_number": pr_number, - "ready": True, - "checkout": { - "worktree_dir": cwd, - "base_sha": base_sha, - "head_sha": head_sha, - }, - "merge": {"status": "skipped"}, - } - - -def cleanup_worktree(*, worktree: str | Path) -> dict: - worktree = Path(worktree).resolve() - # Determine repo root before cleanup — if the worktree is already gone, - # fall back to the parent .worktrees directory's repo. - if worktree.is_dir(): - try: - repo_root = repo_root_from(worktree) - except Exception: - repo_root = worktree.parent.parent - else: - # Worktree dir already deleted; derive repo root from parent. - repo_root = worktree.parent.parent # .worktrees/<branch> -> repo root - repo_root = Path(repo_root).resolve() - - # Validate that repo_root is actually a git repository. - if not (repo_root / ".git").exists(): - return { - "worktree": str(worktree), - "removed": not worktree.exists(), - "branch_still_exists": False, - "error": f"derived repo root {repo_root} is not a git repository", - } - - # Always prune first — this cleans up stale worktree entries for - # directories that were already deleted outside of git. - subprocess.run( - ["git", "-C", str(repo_root), "worktree", "prune"], - capture_output=True, - ) - - # Remove worktree if it still exists in git's tracking. - subprocess.run( - ["git", "-C", str(repo_root), "worktree", "remove", str(worktree), "--force"], - capture_output=True, - text=True, - ) - - branch_name = worktree.name - try: - branch_list = run_git( - repo_root, "branch", "--list", "--format=%(refname:short)" - ).splitlines() - except subprocess.CalledProcessError: - branch_list = [] - return { - "worktree": str(worktree), - "removed": not worktree.exists(), - "branch_still_exists": branch_name in branch_list, - } - - -def emit_result(result: dict, fmt: str) -> None: - if fmt == "json": - print(json.dumps(result, indent=2, sort_keys=True)) - else: - print(json.dumps(result, indent=2, sort_keys=True)) - - -def parse_args(argv: list[str]) -> argparse.Namespace: - parser = argparse.ArgumentParser(description="Pipeline worktree helpers.") - subparsers = parser.add_subparsers(dest="command", required=True) - - enter_parser = subparsers.add_parser("enter") - enter_parser.add_argument("--name", required=True) - enter_parser.add_argument("--base", default="origin/main") - enter_parser.add_argument("--repo-root") - enter_parser.add_argument("--format", choices=["json", "text"], default="json") - - create_issue = subparsers.add_parser("create-issue") - create_issue.add_argument("--issue", required=True, type=int) - create_issue.add_argument("--slug", required=True) - create_issue.add_argument("--base", default="origin/main") - create_issue.add_argument("--repo-root") - create_issue.add_argument("--format", choices=["json", "text"], default="json") - - prepare_issue = subparsers.add_parser("prepare-issue-branch") - prepare_issue.add_argument("--issue", required=True, type=int) - prepare_issue.add_argument("--slug", required=True) - prepare_issue.add_argument("--base", default="main") - prepare_issue.add_argument("--repo-root") - prepare_issue.add_argument("--format", choices=["json", "text"], default="json") - - checkout_pr = subparsers.add_parser("checkout-pr") - checkout_pr.add_argument("--repo", required=True) - checkout_pr.add_argument("--pr", required=True, type=int) - checkout_pr.add_argument("--repo-root") - checkout_pr.add_argument("--format", choices=["json", "text"], default="json") - - prepare_review = subparsers.add_parser("prepare-review") - prepare_review.add_argument("--repo", required=True) - prepare_review.add_argument("--pr", required=True, type=int) - prepare_review.add_argument("--repo-root") - prepare_review.add_argument("--format", choices=["json", "text"], default="json") - - merge_parser = subparsers.add_parser("merge-main") - merge_parser.add_argument("--worktree", required=True) - merge_parser.add_argument("--format", choices=["json", "text"], default="json") - - cleanup_parser = subparsers.add_parser("cleanup") - cleanup_parser.add_argument("--worktree", required=True) - cleanup_parser.add_argument("--format", choices=["json", "text"], default="json") - - return parser.parse_args(argv) - - -def main(argv: list[str] | None = None) -> int: - args = parse_args(argv or sys.argv[1:]) - - if args.command == "enter": - emit_result( - enter( - name=args.name, - base_ref=args.base, - repo_root=args.repo_root, - ), - args.format, - ) - return 0 - - if args.command == "create-issue": - emit_result( - create_issue_worktree( - issue_number=args.issue, - slug=args.slug, - base_ref=args.base, - repo_root=args.repo_root, - ), - args.format, - ) - return 0 - - if args.command == "prepare-issue-branch": - emit_result( - prepare_issue_branch( - issue_number=args.issue, - slug=args.slug, - base_ref=args.base, - repo_root=args.repo_root, - ), - args.format, - ) - return 0 - - if args.command == "checkout-pr": - emit_result( - checkout_pr_worktree( - repo=args.repo, - pr_number=args.pr, - repo_root=args.repo_root, - ), - args.format, - ) - return 0 - - if args.command == "prepare-review": - emit_result( - prepare_review( - repo=args.repo, - pr_number=args.pr, - repo_root=args.repo_root, - ), - args.format, - ) - return 0 - - if args.command == "merge-main": - emit_result(merge_main(worktree=args.worktree), args.format) - return 0 - - if args.command == "cleanup": - emit_result(cleanup_worktree(worktree=args.worktree), args.format) - return 0 - - raise AssertionError(f"Unhandled command: {args.command}") - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/scripts/project_board_poll.py b/scripts/project_board_poll.py deleted file mode 100644 index b25a04972..000000000 --- a/scripts/project_board_poll.py +++ /dev/null @@ -1,168 +0,0 @@ -#!/usr/bin/env python3 -"""Compatibility wrapper for the board poller CLI.""" - -from __future__ import annotations - -import argparse -import subprocess -import sys -from pathlib import Path - -import pipeline_board - -item_identity = pipeline_board.item_identity -load_state = pipeline_board.load_state -save_state = pipeline_board.save_state -ready_entries = pipeline_board.ready_entries -ack_item = pipeline_board.ack_item - - -def fetch_pr_reviews(repo: str, pr_number: int) -> list[dict]: - output = subprocess.check_output( - ["gh", "api", f"repos/{repo}/pulls/{pr_number}/reviews"], - text=True, - ) - data = pipeline_board.json.loads(output) - if not isinstance(data, list): - raise ValueError(f"Unexpected PR review payload for #{pr_number}: {data!r}") - return data - - -def fetch_pr_state(repo: str, pr_number: int) -> str: - return subprocess.check_output( - [ - "gh", - "pr", - "view", - str(pr_number), - "--repo", - repo, - "--json", - "state", - "--jq", - ".state", - ], - text=True, - ).strip() - - -def resolve_issue_pr(repo: str, issue_number: int) -> int | None: - output = subprocess.check_output( - [ - "gh", - "pr", - "list", - "-R", - repo, - "--search", - f"Fix #{issue_number} in:title state:open", - "--json", - "number", - "--limit", - "1", - ], - text=True, - ) - data = pipeline_board.json.loads(output) - if not data: - return None - return int(data[0]["number"]) - - -def linked_repo_pr_numbers(item: dict, repo: str) -> list[int]: - return pipeline_board.linked_repo_pr_numbers(item, repo) - - -def review_entries( - board_data: dict, - repo: str, - pr_resolver=resolve_issue_pr, - pr_state_fetcher=fetch_pr_state, -) -> dict[str, dict]: - return pipeline_board.review_entries( - board_data, - repo, - pr_resolver, - pr_state_fetcher, - ) - - -def current_entries( - mode: str, - board_data: dict, - repo: str | None = None, - pr_resolver=resolve_issue_pr, - pr_state_fetcher=fetch_pr_state, -) -> dict[str, dict]: - return pipeline_board.current_entries( - mode, - board_data, - repo, - pr_resolver, - pr_state_fetcher, - ) - - -def process_snapshot( - mode: str, - board_data: dict, - state_file: Path, - repo: str | None = None, - pr_resolver=resolve_issue_pr, - pr_state_fetcher=fetch_pr_state, -) -> tuple[str, int] | None: - return pipeline_board.process_snapshot( - mode, - board_data, - state_file, - repo, - pr_resolver, - pr_state_fetcher, - ) - - -def parse_args(argv: list[str]) -> argparse.Namespace: - parser = argparse.ArgumentParser( - description="Select eligible board items from the current project-board snapshot." - ) - subparsers = parser.add_subparsers(dest="command", required=True) - - poll = subparsers.add_parser("poll") - poll.add_argument("mode", choices=["ready", "review"]) - poll.add_argument("state_file", type=Path) - poll.add_argument("--repo") - - ack = subparsers.add_parser("ack") - ack.add_argument("state_file", type=Path) - ack.add_argument("item_id") - - return parser.parse_args(argv) - - -def main(argv: list[str] | None = None) -> int: - args = parse_args(argv or sys.argv[1:]) - - if args.command == "ack": - ack_item(args.state_file, args.item_id) - return 0 - - if args.mode == "review" and not args.repo: - raise SystemExit("--repo is required in review mode") - - board_data = pipeline_board.json.load(sys.stdin) - next_item = process_snapshot( - args.mode, - board_data, - args.state_file, - repo=args.repo, - ) - if next_item is None: - return 1 - - item_id, number = next_item - print(f"{item_id}\t{number}") - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/scripts/project_board_recover.py b/scripts/project_board_recover.py deleted file mode 100644 index b821006e6..000000000 --- a/scripts/project_board_recover.py +++ /dev/null @@ -1,147 +0,0 @@ -#!/usr/bin/env python3 -"""Compatibility wrapper for project-board status recovery.""" - -from __future__ import annotations - -import argparse -import subprocess -import sys -from pathlib import Path - -import pipeline_board - -PROJECT_ID = "PVT_kwDOBrtarc4BRNVy" -STATUS_FIELD_ID = "PVTSSF_lADOBrtarc4BRNVyzg_GmQc" - -STATUS_BACKLOG = pipeline_board.STATUS_BACKLOG -STATUS_READY = pipeline_board.STATUS_READY -STATUS_REVIEW_POOL = pipeline_board.STATUS_REVIEW_POOL -STATUS_FINAL_REVIEW = pipeline_board.STATUS_FINAL_REVIEW -STATUS_DONE = pipeline_board.STATUS_DONE -STATUS_OPTION_IDS = pipeline_board.STATUS_OPTION_IDS -FAILURE_LABELS = pipeline_board.FAILURE_LABELS -label_names = pipeline_board.label_names -linked_pr_numbers = pipeline_board.linked_pr_numbers -is_tracked_issue_title = pipeline_board.is_tracked_issue_title -all_checks_green = pipeline_board.all_checks_green -infer_issue_status = pipeline_board.infer_issue_status -build_recovery_plan = pipeline_board.build_recovery_plan -apply_plan = pipeline_board.apply_plan -save_backup = pipeline_board.save_backup -print_summary = pipeline_board.print_summary -print_examples = pipeline_board.print_examples - - -def run_gh(*args: str) -> str: - return subprocess.check_output(["gh", *args], text=True) - - -def fetch_board_items(owner: str, project_number: int, limit: int) -> dict: - return pipeline_board.json.loads( - run_gh( - "project", - "item-list", - str(project_number), - "--owner", - owner, - "--format", - "json", - "--limit", - str(limit), - ) - ) - - -def fetch_issues(repo: str, limit: int) -> list[dict]: - return pipeline_board.json.loads( - run_gh( - "issue", - "list", - "-R", - repo, - "--state", - "all", - "--limit", - str(limit), - "--json", - "number,state,closedAt,title,labels", - ) - ) - - -def fetch_prs(repo: str, limit: int) -> list[dict]: - return pipeline_board.json.loads( - run_gh( - "pr", - "list", - "-R", - repo, - "--state", - "all", - "--limit", - str(limit), - "--json", - "number,state,isDraft,mergedAt,title,url,reviewDecision,statusCheckRollup,closingIssuesReferences", - ) - ) - - -def default_backup_path(project_number: int) -> Path: - return pipeline_board.default_backup_path(project_number) - - -def parse_args(argv: list[str]) -> argparse.Namespace: - parser = argparse.ArgumentParser( - description="Recover project board item statuses after Status-field recreation." - ) - parser.add_argument("--owner", default="CodingThrust") - parser.add_argument("--repo", default="CodingThrust/problem-reductions") - parser.add_argument("--project-number", type=int, default=8) - parser.add_argument("--project-id", default=PROJECT_ID) - parser.add_argument("--field-id", default=STATUS_FIELD_ID) - parser.add_argument("--limit", type=int, default=500) - parser.add_argument("--apply", action="store_true") - parser.add_argument("--backup-file", type=Path) - parser.add_argument("--plan-file", type=Path) - parser.add_argument("--no-examples", action="store_true") - return parser.parse_args(argv) - - -def main(argv: list[str] | None = None) -> int: - args = parse_args(argv or sys.argv[1:]) - - board_data = fetch_board_items(args.owner, args.project_number, args.limit) - issues = fetch_issues(args.repo, args.limit) - prs = fetch_prs(args.repo, args.limit) - - plan = build_recovery_plan(board_data, issues, prs) - if args.plan_file is not None: - args.plan_file.parent.mkdir(parents=True, exist_ok=True) - args.plan_file.write_text( - pipeline_board.json.dumps(plan, indent=2, sort_keys=True) + "\n" - ) - - print_summary(plan) - if not args.no_examples: - print_examples(plan) - - if not args.apply: - return 0 - - backup_file = args.backup_file or default_backup_path(args.project_number) - save_backup( - backup_file, - board_data=board_data, - issues=issues, - prs=prs, - plan=plan, - ) - changed = apply_plan(plan, project_id=args.project_id, field_id=args.field_id) - print("") - print(f"Applied {changed} status updates.") - print(f"Backup written to {backup_file}") - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/scripts/test_make_helpers.py b/scripts/test_make_helpers.py deleted file mode 100644 index 916eb26cc..000000000 --- a/scripts/test_make_helpers.py +++ /dev/null @@ -1,825 +0,0 @@ -#!/usr/bin/env python3 -import shutil -import subprocess -import tempfile -import unittest -from pathlib import Path - - -REPO_ROOT = Path(__file__).resolve().parents[1] - - -class MakeHelpersTests(unittest.TestCase): - def test_helper_sources_under_dash(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - ["dash", "-c", ". scripts/make_helpers.sh"], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - - def test_run_agent_enables_multi_agent_for_codex(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "codex() { printf '%s\\n' \"$@\"; }; " - "RUNNER=codex CODEX_MODEL=test-model " - "run_agent /dev/null 'test prompt'" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "exec", - "--enable", - "multi_agent", - "-m", - "test-model", - "-s", - "danger-full-access", - "test prompt", - ], - ) - - def test_skill_prompt_with_context_appends_json_for_codex(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "RUNNER=codex " - "skill_prompt_with_context review-pipeline '/review-pipeline 570' " - "'process PR #570' 'Selected queue item' '{\"pr_number\":570}'" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertIn("Use the repo-local skill", proc.stdout) - self.assertIn("Selected queue item", proc.stdout) - self.assertIn('{"pr_number":570}', proc.stdout) - - def test_skill_prompt_with_context_keeps_claude_slash_command_clean(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "RUNNER=claude " - "skill_prompt_with_context review-pipeline '/review-pipeline 570' " - "'process PR #570' 'Selected queue item' '{\"pr_number\":570}'" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual(proc.stdout.strip(), "/review-pipeline 570") - - def test_poll_project_items_uses_pipeline_board_cli(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "python3() { printf '%s\\n' \"$@\"; }; " - "poll_project_items ready /tmp/state.json" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "scripts/pipeline_board.py", - "next", - "ready", - "/tmp/state.json", - "--format", - "text", - "--repo-root", - ".", - ], - ) - - def test_move_board_item_uses_pipeline_board_cli(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "python3() { printf '%s\\n' \"$@\"; }; " - "move_board_item PVTI_demo final-review" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "scripts/pipeline_board.py", - "move", - "PVTI_demo", - "final-review", - ], - ) - - def test_claim_project_items_uses_pipeline_board_cli(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "python3() { printf '%s\\n' \"$@\"; }; " - "claim_project_items ready /tmp/state.json" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "scripts/pipeline_board.py", - "claim-next", - "ready", - "/tmp/state.json", - "--format", - "json", - "--repo-root", - ".", - ], - ) - - def test_make_board_next_final_review_passes_repo(self) -> None: - proc = subprocess.run( - [ - "make", - "-n", - "board-next", - "MODE=final-review", - "REPO=CodingThrust/problem-reductions", - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertIn( - 'poll_project_items "final-review" "$state_file" "$repo"', - proc.stdout, - ) - - def test_make_board_next_review_forwards_number_and_format(self) -> None: - proc = subprocess.run( - [ - "make", - "-n", - "board-next", - "MODE=review", - "REPO=CodingThrust/problem-reductions", - "NUMBER=570", - "FORMAT=json", - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertIn( - 'poll_project_items "review" "$state_file" "$repo" "570" "json"', - proc.stdout, - ) - - def test_make_board_claim_review_forwards_repo_number_and_format(self) -> None: - proc = subprocess.run( - [ - "make", - "-n", - "board-claim", - "MODE=review", - "REPO=CodingThrust/problem-reductions", - "NUMBER=570", - "FORMAT=json", - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertIn( - 'claim_project_items "review" "$state_file" "$repo" "570" "json"', - proc.stdout, - ) - - def test_board_next_json_uses_scripted_json_poll(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "python3() { printf '%s\\n' \"$@\"; }; " - "board_next_json review CodingThrust/problem-reductions 570 /tmp/review.json" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "scripts/pipeline_board.py", - "next", - "review", - "/tmp/review.json", - "--format", - "json", - "--repo", - "CodingThrust/problem-reductions", - "--number", - "570", - ], - ) - - def test_board_claim_json_uses_scripted_json_claim(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "python3() { printf '%s\\n' \"$@\"; }; " - "board_claim_json review CodingThrust/problem-reductions 570 /tmp/review.json" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "scripts/pipeline_board.py", - "claim-next", - "review", - "/tmp/review.json", - "--format", - "json", - "--repo", - "CodingThrust/problem-reductions", - "--number", - "570", - ], - ) - - def test_review_pipeline_context_uses_skill_bundle_cli(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "python3() { printf '%s\\n' \"$@\"; }; " - "review_pipeline_context CodingThrust/problem-reductions 570" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "scripts/pipeline_skill_context.py", - "review-pipeline", - "--repo", - "CodingThrust/problem-reductions", - "--format", - "json", - "--pr", - "570", - ], - ) - - def test_make_run_review_uses_skill_bundle_context(self) -> None: - proc = subprocess.run( - ["make", "-n", "run-review"], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertIn('review_pipeline_context "$repo"', proc.stdout) - self.assertIn('skill_prompt_with_context review-pipeline', proc.stdout) - - def test_make_run_review_watch_mode_emits_outcome_markers(self) -> None: - proc = subprocess.run( - ["make", "-n", "run-review", "WATCH_MODE=1", "N=570"], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertIn("watch_emit_outcome gone", proc.stdout) - self.assertIn('run_agent_with_watch_outcome "review-output.log" "$PROMPT"', proc.stdout) - - def test_make_run_pipeline_uses_scripted_board_selection(self) -> None: - proc = subprocess.run( - ["make", "-n", "run-pipeline"], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertIn('board_next_json ready "" "" "$tmp_state"', proc.stdout) - self.assertIn('skill_prompt_with_context project-pipeline', proc.stdout) - - def test_make_run_pipeline_watch_mode_emits_outcome_markers(self) -> None: - proc = subprocess.run( - ["make", "-n", "run-pipeline", "WATCH_MODE=1", "N=42"], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertIn("watch_emit_outcome gone", proc.stdout) - self.assertIn('run_agent_with_watch_outcome "pipeline-output.log" "$PROMPT"', proc.stdout) - - def test_watch_and_dispatch_uses_persistent_default_state_file(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "poll_project_items() { printf 'state:%s\\n' \"$2\" >&2; return 2; }; " - "watch_and_dispatch ready run-pipeline 'Ready issues'" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 2, proc.stderr) - self.assertIn( - "state:/tmp/problemreductions-ready-forever-state.json", - proc.stderr, - ) - - def test_watch_and_dispatch_rechecks_immediately_after_successful_dispatch(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "flag=/tmp/test-watch-and-dispatch-$$; " - "rm -f \"$flag\"; " - "date() { printf '2026-03-16 00:00:00'; }; " - "poll_project_items() { " - " if [ ! -f \"$flag\" ]; then : > \"$flag\"; printf 'PVTI_1\\t42\\n'; return 0; fi; " - " return 2; " - "}; " - "make() { printf 'make:%s %s\\n' \"$1\" \"$2\"; return 0; }; " - "ack_polled_item() { printf 'ack:%s\\n' \"$2\"; }; " - "sleep() { printf 'sleep:%s\\n' \"$1\"; return 0; }; " - "MAKE=make POLL_INTERVAL=600 " - "watch_and_dispatch ready run-pipeline 'Ready issues'" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 2, proc.stderr) - self.assertIn("make:run-pipeline N=42", proc.stdout) - self.assertIn("ack:PVTI_1", proc.stdout) - self.assertNotIn("sleep:600", proc.stdout) - - def test_watch_and_dispatch_drains_ready_items_before_sleeping(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "counter_file=/tmp/test-watch-drain-$$.count; " - "rm -f \"$counter_file\"; " - "date() { printf '2026-03-16 00:00:00'; }; " - "poll_project_items() { " - " count=$(cat \"$counter_file\" 2>/dev/null || printf '0'); " - " count=$((count + 1)); " - " printf '%s' \"$count\" > \"$counter_file\"; " - " case \"$count\" in " - " 1) printf 'PVTI_1\\t42\\n'; return 0 ;; " - " 2) printf 'PVTI_2\\t43\\n'; return 0 ;; " - " 3) return 1 ;; " - " *) return 2 ;; " - " esac; " - "}; " - "make() { printf 'make:%s %s\\n' \"$1\" \"$2\"; return 0; }; " - "ack_polled_item() { printf 'ack:%s\\n' \"$2\"; }; " - "sleep() { printf 'sleep:%s\\n' \"$1\"; return 0; }; " - "MAKE=make POLL_INTERVAL=600 " - "watch_and_dispatch ready run-pipeline 'Ready issues'" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 2, proc.stderr) - self.assertIn("make:run-pipeline N=42", proc.stdout) - self.assertIn("ack:PVTI_1", proc.stdout) - self.assertIn("make:run-pipeline N=43", proc.stdout) - self.assertIn("ack:PVTI_2", proc.stdout) - self.assertEqual(proc.stdout.count("sleep:600"), 1) - self.assertLess( - proc.stdout.index("make:run-pipeline N=43"), - proc.stdout.index("sleep:600"), - ) - - def test_watch_and_dispatch_skips_retry_state_for_explicit_gone_outcome(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "gone-state.json" - counter_file = Path(tmpdir) / "gone-counter.txt" - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - f"state_file={state_file}; " - f"counter_file={counter_file}; " - "rm -f \"$state_file\" \"$counter_file\"; " - "poll_project_items() { " - " count=$(cat \"$counter_file\" 2>/dev/null || printf '0'); " - " count=$((count + 1)); " - " printf '%s' \"$count\" > \"$counter_file\"; " - " if [ \"$count\" -eq 1 ]; then printf 'PVTI_1\\t42\\n'; return 0; fi; " - " return 2; " - "}; " - "make() { printf '__WATCH_OUTCOME__=gone\\n'; return 3; }; " - "move_board_item() { printf 'move:%s %s\\n' \"$1\" \"$2\"; }; " - "sleep() { printf 'sleep:%s\\n' \"$1\"; return 0; }; " - "STATE_FILE=\"$state_file\" MAKE=make POLL_INTERVAL=600 " - "watch_and_dispatch ready run-pipeline 'Ready issues'" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - state_file_exists = state_file.exists() - self.assertEqual(proc.returncode, 2, proc.stderr) - self.assertNotIn("move:PVTI_1 on-hold", proc.stdout) - self.assertFalse(state_file_exists) - - def test_watch_and_dispatch_calls_poll_with_correct_args(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "state_tmp=$(mktemp); " - 'printf \'{"pending":["a","b"],"visible":{}}\' > "$state_tmp"; ' - "poll_project_items() { printf 'mode:%s state:%s\\n' \"$1\" \"$2\" >&2; return 2; }; " - "STATE_FILE=\"$state_tmp\" " - "watch_and_dispatch ready run-pipeline 'Ready issues'" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 2, proc.stderr) - self.assertIn("mode:ready state:", proc.stderr) - - def test_watch_and_dispatch_handles_missing_state_file(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "poll_project_items() { printf 'mode:%s\\n' \"$1\" >&2; return 2; }; " - "STATE_FILE=/tmp/nonexistent-state-$$.json " - "watch_and_dispatch ready run-pipeline 'Ready issues'" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 2, proc.stderr) - self.assertIn("mode:ready", proc.stderr) - - def test_pr_snapshot_uses_pipeline_pr_cli(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "python3() { printf '%s\\n' \"$@\"; }; " - "pr_snapshot CodingThrust/problem-reductions 570" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "scripts/pipeline_pr.py", - "snapshot", - "--repo", - "CodingThrust/problem-reductions", - "--pr", - "570", - "--format", - "json", - ], - ) - - def test_issue_guards_uses_pipeline_checks_cli(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "python3() { printf '%s\\n' \"$@\"; }; " - "issue_guards CodingThrust/problem-reductions 117 /tmp/repo" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "scripts/pipeline_checks.py", - "issue-guards", - "--repo", - "CodingThrust/problem-reductions", - "--issue", - "117", - "--repo-root", - "/tmp/repo", - "--format", - "json", - ], - ) - - def test_pr_wait_ci_uses_pipeline_pr_cli(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "python3() { printf '%s\\n' \"$@\"; }; " - "pr_wait_ci CodingThrust/problem-reductions 570 1200 15" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "scripts/pipeline_pr.py", - "wait-ci", - "--repo", - "CodingThrust/problem-reductions", - "--pr", - "570", - "--timeout", - "1200", - "--interval", - "15", - "--format", - "json", - ], - ) - - def test_issue_context_uses_pipeline_checks_cli(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "python3() { printf '%s\\n' \"$@\"; }; " - "issue_context CodingThrust/problem-reductions 117 /tmp/repo" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "scripts/pipeline_checks.py", - "issue-context", - "--repo", - "CodingThrust/problem-reductions", - "--issue", - "117", - "--repo-root", - "/tmp/repo", - "--format", - "json", - ], - ) - - def test_make_issue_context_uses_shared_helper(self) -> None: - proc = subprocess.run( - [ - "make", - "-n", - "issue-context", - "ISSUE=117", - "REPO=CodingThrust/problem-reductions", - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertIn( - 'issue_context "$repo" "117"', - proc.stdout, - ) - - def test_create_issue_worktree_uses_pipeline_worktree_cli(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "python3() { printf '%s\\n' \"$@\"; }; " - "create_issue_worktree 117 graph-partitioning origin/main" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "scripts/pipeline_worktree.py", - "create-issue", - "--issue", - "117", - "--slug", - "graph-partitioning", - "--base", - "origin/main", - "--format", - "json", - ], - ) - - def test_checkout_pr_worktree_uses_pipeline_worktree_cli(self) -> None: - if shutil.which("dash") is None: - self.skipTest("dash is not installed") - - proc = subprocess.run( - [ - "dash", - "-c", - ( - ". scripts/make_helpers.sh; " - "python3() { printf '%s\\n' \"$@\"; }; " - "checkout_pr_worktree CodingThrust/problem-reductions 570" - ), - ], - cwd=REPO_ROOT, - capture_output=True, - text=True, - ) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertEqual( - proc.stdout.splitlines(), - [ - "scripts/pipeline_worktree.py", - "checkout-pr", - "--repo", - "CodingThrust/problem-reductions", - "--pr", - "570", - "--format", - "json", - ], - ) - - -if __name__ == "__main__": - unittest.main() diff --git a/scripts/test_pipeline_board.py b/scripts/test_pipeline_board.py deleted file mode 100644 index 58c0ac6c9..000000000 --- a/scripts/test_pipeline_board.py +++ /dev/null @@ -1,1340 +0,0 @@ -#!/usr/bin/env python3 -import io -import json -import tempfile -import unittest -from contextlib import redirect_stdout -from pathlib import Path - -from unittest.mock import patch - -from pipeline_board import ( - STATUS_DONE, - STATUS_FINAL_REVIEW, - STATUS_IN_PROGRESS, - STATUS_ON_HOLD, - STATUS_READY, - STATUS_REVIEW_POOL, - STATUS_UNDER_REVIEW, - _is_rule_blocked, - _scan_existing_problems, - ack_item, - batch_fetch_issues, - batch_fetch_prs_with_reviews, - claim_next_entry, - build_recovery_plan, - claim_entry_from_entries, - eligible_review_candidate_entries, - final_review_entries, - load_state, - normalize_status_name, - print_next_item, - process_snapshot, - ready_entries, - review_candidates, - review_entries, - select_next_entry, - status_items, -) - - -def make_issue_item( - item_id: str, - number: int, - *, - status: str = "Ready", - title: str | None = None, - linked_prs: list[int] | None = None, -) -> dict: - item = { - "id": item_id, - "status": status, - "content": { - "type": "Issue", - "number": number, - "title": title or f"[Model] Issue {number}", - }, - "title": title or f"[Model] Issue {number}", - } - if linked_prs is not None: - item["linked pull requests"] = [ - f"https://github.com/CodingThrust/problem-reductions/pull/{pr_number}" - for pr_number in linked_prs - ] - return item - - -def make_pr_item(item_id: str, number: int, status: str = "Review pool") -> dict: - return { - "id": item_id, - "status": status, - "content": {"type": "PullRequest", "number": number}, - } - - -def make_issue(number: int, *, state: str = "OPEN", labels: list[str] | None = None) -> dict: - return { - "number": number, - "state": state, - "title": f"[Model] Issue {number}", - "labels": [{"name": label} for label in (labels or [])], - } - - -def make_pr( - number: int, - *, - state: str = "OPEN", - merged: bool = False, - checks: list[dict] | None = None, -) -> dict: - return { - "number": number, - "state": state, - "mergedAt": "2026-03-15T00:00:00Z" if merged else None, - "statusCheckRollup": checks or [], - } - - -def success_check(name: str = "ci") -> dict: - return { - "__typename": "CheckRun", - "name": name, - "status": "COMPLETED", - "conclusion": "SUCCESS", - } - - -class PipelineBoardPollTests(unittest.TestCase): - def test_ready_selection_ignores_ack_until_board_changes(self) -> None: - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "ready-state.json" - snapshot = { - "items": [ - make_issue_item("PVTI_1", 101), - make_issue_item("PVTI_2", 102), - ] - } - - item_id, number = process_snapshot("ready", snapshot, state_file) - self.assertEqual((item_id, number), ("PVTI_1", 101)) - - item_id, number = process_snapshot("ready", snapshot, state_file) - self.assertEqual((item_id, number), ("PVTI_1", 101)) - - ack_item(state_file, "PVTI_1") - item_id, number = process_snapshot("ready", snapshot, state_file) - self.assertEqual((item_id, number), ("PVTI_1", 101)) - - def test_ready_entries_filters_blocked_rules(self) -> None: - """[Rule] issues whose source/target model is missing are excluded when repo_root is set.""" - with tempfile.TemporaryDirectory() as tmpdir: - # Create a fake src/models/ with ILP and MaxCut - models_dir = Path(tmpdir) / "src" / "models" - models_dir.mkdir(parents=True) - (models_dir / "ilp.rs").write_text("pub struct ILP {}") - (models_dir / "max_cut.rs").write_text("pub struct MaxCut {}") - - snapshot = { - "items": [ - make_issue_item("PVTI_1", 184, title="[Model] MinimumMultiwayCut"), - make_issue_item("PVTI_2", 185, title="[Rule] MinimumMultiwayCut to ILP"), - make_issue_item("PVTI_3", 186, title="[Rule] MinimumMultiwayCut to QUBO"), - make_issue_item("PVTI_4", 100, title="[Rule] MaxCut to ILP"), - ] - } - - # Without repo_root: all 4 items returned - entries = ready_entries(snapshot) - self.assertEqual(len(entries), 4) - - # With repo_root: blocked rules filtered out - # 185: source MinimumMultiwayCut missing - # 186: source MinimumMultiwayCut missing, target QUBO missing - # 184: [Model] — never filtered - # 100: both MaxCut and ILP exist — not filtered - entries = ready_entries(snapshot, repo_root=tmpdir) - numbers = {e["number"] for e in entries.values()} - self.assertEqual(numbers, {184, 100}) - - def test_select_next_entry_ready_ignores_stale_state_file(self) -> None: - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "ready-state.json" - state_file.write_text( - json.dumps( - { - "pending": [], - "visible": { - "PVTI_1": { - "number": 101, - "issue_number": 101, - "pr_number": None, - "status": STATUS_READY, - "title": "[Model] ExactCoverBy3Sets", - } - }, - } - ) - ) - - entry = select_next_entry( - "ready", - { - "items": [ - make_issue_item( - "PVTI_1", - 101, - title="[Model] ExactCoverBy3Sets", - ) - ] - }, - state_file, - ) - - self.assertEqual( - entry, - { - "item_id": "PVTI_1", - "number": 101, - "issue_number": 101, - "pr_number": None, - "status": STATUS_READY, - "title": "[Model] ExactCoverBy3Sets", - }, - ) - - def test_select_next_entry_ready_prefers_models_before_rules(self) -> None: - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "ready-state.json" - entry = select_next_entry( - "ready", - { - "items": [ - make_issue_item( - "PVTI_2", - 100, - title="[Rule] MaxCut to ILP", - ), - make_issue_item( - "PVTI_1", - 200, - title="[Model] ExactCoverBy3Sets", - ), - ] - }, - state_file, - ) - - self.assertEqual( - entry, - { - "item_id": "PVTI_1", - "number": 200, - "issue_number": 200, - "pr_number": None, - "status": STATUS_READY, - "title": "[Model] ExactCoverBy3Sets", - }, - ) - - def test_is_rule_blocked_helper(self) -> None: - existing = {"ILP", "MaximumIndependentSet"} - # Not a rule - self.assertFalse(_is_rule_blocked("[Model] Foo", existing)) - # Both exist - self.assertFalse(_is_rule_blocked("[Rule] ILP to MaximumIndependentSet", existing)) - # Source missing - self.assertTrue(_is_rule_blocked("[Rule] Foo to ILP", existing)) - # Target missing - self.assertTrue(_is_rule_blocked("[Rule] ILP to Bar", existing)) - - def test_review_queue_resolves_issue_cards_to_prs(self) -> None: - def fake_pr_resolver(repo: str, issue_number: int) -> int | None: - self.assertEqual(repo, "CodingThrust/problem-reductions") - return 570 if issue_number == 117 else None - - def fake_pr_state_fetcher(repo: str, pr_number: int) -> str: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(pr_number, 570) - return "OPEN" - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "review-state.json" - item_id, number = process_snapshot( - "review", - {"items": [make_issue_item("PVTI_10", 117, status="Review pool")]}, - state_file, - repo="CodingThrust/problem-reductions", - pr_resolver=fake_pr_resolver, - pr_state_fetcher=fake_pr_state_fetcher, - ) - self.assertEqual((item_id, number), ("PVTI_10", 570)) - - def test_review_queue_skips_closed_pr_cards(self) -> None: - def fake_pr_state_fetcher(repo: str, pr_number: int) -> str: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(pr_number, 570) - return "CLOSED" - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "review-state.json" - no_item = process_snapshot( - "review", - {"items": [make_pr_item("PVTI_10", 570)]}, - state_file, - repo="CodingThrust/problem-reductions", - pr_state_fetcher=fake_pr_state_fetcher, - ) - self.assertIsNone(no_item) - - def test_final_review_queue_resolves_issue_cards_to_open_prs(self) -> None: - def fake_pr_resolver(repo: str, issue_number: int) -> int | None: - self.assertEqual(repo, "CodingThrust/problem-reductions") - return 615 if issue_number == 101 else None - - def fake_pr_state_fetcher(repo: str, pr_number: int) -> str: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(pr_number, 615) - return "OPEN" - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "final-review-state.json" - item_id, number = process_snapshot( - "final-review", - {"items": [make_issue_item("PVTI_20", 101, status="Final review")]}, - state_file, - repo="CodingThrust/problem-reductions", - pr_resolver=fake_pr_resolver, - pr_state_fetcher=fake_pr_state_fetcher, - ) - self.assertEqual((item_id, number), ("PVTI_20", 615)) - - def test_final_review_queue_skips_closed_pr_cards(self) -> None: - def fake_pr_state_fetcher(repo: str, pr_number: int) -> str: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(pr_number, 621) - return "CLOSED" - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "final-review-state.json" - no_item = process_snapshot( - "final-review", - {"items": [make_pr_item("PVTI_21", 621, status="Final review")]}, - state_file, - repo="CodingThrust/problem-reductions", - pr_state_fetcher=fake_pr_state_fetcher, - ) - self.assertIsNone(no_item) - - def test_parse_args_accepts_final_review_list_mode(self) -> None: - import pipeline_board - - args = pipeline_board.parse_args(["list", "final-review", "--repo", "CodingThrust/problem-reductions"]) - - self.assertEqual(args.command, "list") - self.assertEqual(args.mode, "final-review") - - def test_main_lists_final_review_items(self) -> None: - import pipeline_board - - board_data = { - "items": [make_issue_item("PVTI_30", 239, status="Final review", title="[Model] BalancedCompleteBipartiteSubgraph")] - } - - with ( - patch.object(pipeline_board, "fetch_board_items", return_value=board_data) as fetch_board_items, - patch.object(pipeline_board, "print_candidate_list", return_value=0) as print_candidate_list, - ): - rc = pipeline_board.main( - ["list", "final-review", "--repo", "CodingThrust/problem-reductions", "--format", "json"] - ) - - self.assertEqual(rc, 0) - fetch_board_items.assert_called_once() - print_candidate_list.assert_called_once_with( - "final-review", - [ - { - "number": 239, - "issue_number": 239, - "pr_number": None, - "status": "Final review", - "title": "[Model] BalancedCompleteBipartiteSubgraph", - "item_id": "PVTI_30", - } - ], - fmt="json", - ) - - -class ReviewCandidateQueueTests(unittest.TestCase): - def test_select_next_entry_review_ignores_stale_state_file(self) -> None: - def fake_pr_resolver(repo: str, issue_number: int) -> int | None: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(issue_number, 117) - return 570 - - def fake_pr_state_fetcher(repo: str, pr_number: int) -> str: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(pr_number, 570) - return "OPEN" - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "review-state.json" - state_file.write_text( - json.dumps( - { - "pending": [], - "visible": { - "PVTI_1": { - "issue_number": 117, - "number": 570, - "pr_number": 570, - "status": "Review pool", - "title": "[Model] GraphPartitioning", - } - }, - } - ) - ) - - entry = select_next_entry( - "review", - { - "items": [ - make_issue_item( - "PVTI_1", - 117, - status="Review pool", - title="[Model] GraphPartitioning", - ) - ] - }, - state_file, - repo="CodingThrust/problem-reductions", - pr_resolver=fake_pr_resolver, - pr_state_fetcher=fake_pr_state_fetcher, - ) - - self.assertEqual( - entry, - { - "item_id": "PVTI_1", - "number": 570, - "issue_number": 117, - "pr_number": 570, - "status": STATUS_REVIEW_POOL, - "title": "[Model] GraphPartitioning", - }, - ) - - def test_ack_item_clears_retry_state_only(self) -> None: - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "review-candidates.json" - state_file.write_text( - json.dumps( - { - "retries": { - "PVTI_1": 2, - "PVTI_2": 1, - } - } - ) - ) - - ack_item(state_file, "PVTI_1") - self.assertEqual( - load_state(state_file), - {"retries": {"PVTI_2": 1}}, - ) - - def test_claim_entry_from_entries_moves_selected_review_item(self) -> None: - entries = eligible_review_candidate_entries( - [ - { - "item_id": "PVTI_11", - "issue_number": 117, - "pr_number": 570, - "status": "Review pool", - "title": "[Model] GraphPartitioning", - "eligibility": "eligible", - } - ] - ) - moves: list[tuple[str, str]] = [] - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "review-candidates.json" - claimed = claim_entry_from_entries( - "review", - entries, - state_file, - mover=lambda item_id, status: moves.append((item_id, status)), - ) - - self.assertEqual(claimed["item_id"], "PVTI_11") - self.assertEqual(claimed["claimed_status"], STATUS_UNDER_REVIEW) - self.assertEqual(moves, [("PVTI_11", STATUS_UNDER_REVIEW)]) - - -class PipelineBoardRecoveryTests(unittest.TestCase): - def test_recovery_plan_marks_merged_pr_items_done(self) -> None: - board_data = { - "items": [ - make_issue_item( - "PVTI_1", - 101, - status="Review pool", - title="[Model] MinimumFeedbackVertexSet", - linked_prs=[615], - ) - ] - } - issues = [make_issue(101, labels=["Good"])] - prs = [make_pr(615, state="MERGED", merged=True)] - - plan = build_recovery_plan(board_data, issues, prs) - - self.assertEqual(len(plan), 1) - self.assertEqual(plan[0]["proposed_status"], STATUS_DONE) - - def test_recovery_plan_marks_green_open_prs_final_review(self) -> None: - board_data = { - "items": [ - make_issue_item( - "PVTI_1", - 101, - status="Review pool", - title="[Model] HamiltonianPath", - linked_prs=[621], - ) - ] - } - issues = [make_issue(101, labels=["Good"])] - prs = [make_pr(621, checks=[success_check()])] - - plan = build_recovery_plan(board_data, issues, prs) - - self.assertEqual(plan[0]["proposed_status"], STATUS_FINAL_REVIEW) - self.assertIn("green open PR", plan[0]["reason"]) - - def test_recovery_plan_marks_open_pr_with_failing_checks_review_pool(self) -> None: - board_data = { - "items": [ - make_issue_item( - "PVTI_1", - 101, - status="In progress", - title="[Model] SteinerTree", - linked_prs=[192], - ) - ] - } - issues = [make_issue(101, labels=["Good"])] - prs = [make_pr(192, checks=[])] # no checks → not green - - plan = build_recovery_plan(board_data, issues, prs) - - self.assertEqual(plan[0]["proposed_status"], STATUS_REVIEW_POOL) - - def test_recovery_plan_marks_good_issue_without_pr_ready(self) -> None: - board_data = { - "items": [ - make_issue_item( - "PVTI_1", - 101, - status="Backlog", - title="[Model] ExactCoverBy3Sets", - ) - ] - } - issues = [make_issue(101, labels=["Good"])] - - plan = build_recovery_plan(board_data, issues, prs=[]) - - self.assertEqual(plan[0]["proposed_status"], STATUS_READY) - - -class PipelineBoardStatusTests(unittest.TestCase): - def test_normalize_status_name_accepts_pipeline_aliases(self) -> None: - self.assertEqual(normalize_status_name("ready"), STATUS_READY) - self.assertEqual(normalize_status_name("review-pool"), STATUS_REVIEW_POOL) - self.assertEqual(normalize_status_name("in-progress"), STATUS_IN_PROGRESS) - self.assertEqual(normalize_status_name("under review"), STATUS_UNDER_REVIEW) - self.assertEqual(normalize_status_name("on-hold"), STATUS_ON_HOLD) - self.assertEqual(normalize_status_name("done"), STATUS_DONE) - - -class PipelineBoardOutputTests(unittest.TestCase): - def test_claim_next_ready_moves_selected_item_to_in_progress(self) -> None: - moves: list[tuple[str, str]] = [] - - def fake_mover(item_id: str, status: str) -> None: - moves.append((item_id, status)) - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "ready-state.json" - result = claim_next_entry( - "ready", - { - "items": [ - make_issue_item( - "PVTI_1", - 101, - title="[Model] ExactCoverBy3Sets", - ) - ] - }, - state_file, - mover=fake_mover, - ) - - self.assertEqual(moves, [("PVTI_1", STATUS_IN_PROGRESS)]) - self.assertEqual( - result, - { - "item_id": "PVTI_1", - "number": 101, - "issue_number": 101, - "pr_number": None, - "status": STATUS_READY, - "title": "[Model] ExactCoverBy3Sets", - "claimed": True, - "claimed_status": STATUS_IN_PROGRESS, - }, - ) - - def test_claim_next_review_moves_selected_item_to_under_review(self) -> None: - moves: list[tuple[str, str]] = [] - - def fake_pr_resolver(repo: str, issue_number: int) -> int | None: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(issue_number, 117) - return 570 - - def fake_pr_state_fetcher(repo: str, pr_number: int) -> str: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(pr_number, 570) - return "OPEN" - - def fake_mover(item_id: str, status: str) -> None: - moves.append((item_id, status)) - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "review-state.json" - result = claim_next_entry( - "review", - { - "items": [ - make_issue_item( - "PVTI_10", - 117, - status="Review pool", - title="[Model] GraphPartitioning", - ) - ] - }, - state_file, - repo="CodingThrust/problem-reductions", - pr_resolver=fake_pr_resolver, - pr_state_fetcher=fake_pr_state_fetcher, - mover=fake_mover, - ) - - self.assertEqual(moves, [("PVTI_10", STATUS_UNDER_REVIEW)]) - self.assertEqual( - result, - { - "item_id": "PVTI_10", - "number": 570, - "issue_number": 117, - "pr_number": 570, - "status": STATUS_REVIEW_POOL, - "title": "[Model] GraphPartitioning", - "claimed": True, - "claimed_status": STATUS_UNDER_REVIEW, - }, - ) - - def test_select_next_entry_honors_requested_number(self) -> None: - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "ready-state.json" - entry = select_next_entry( - "ready", - { - "items": [ - make_issue_item("PVTI_1", 101, title="[Model] A"), - make_issue_item("PVTI_2", 102, title="[Model] B"), - ] - }, - state_file, - target_number=102, - ) - self.assertEqual( - entry, - { - "item_id": "PVTI_2", - "number": 102, - "issue_number": 102, - "pr_number": None, - "status": STATUS_READY, - "title": "[Model] B", - }, - ) - - def test_select_next_entry_includes_ready_issue_metadata(self) -> None: - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "ready-state.json" - entry = select_next_entry( - "ready", - { - "items": [ - make_issue_item( - "PVTI_1", - 101, - title="[Model] ExactCoverBy3Sets", - ) - ] - }, - state_file, - ) - self.assertEqual( - entry, - { - "item_id": "PVTI_1", - "number": 101, - "issue_number": 101, - "pr_number": None, - "status": STATUS_READY, - "title": "[Model] ExactCoverBy3Sets", - }, - ) - - def test_select_next_entry_includes_review_metadata(self) -> None: - def fake_pr_resolver(repo: str, issue_number: int) -> int | None: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(issue_number, 117) - return 570 - - def fake_pr_state_fetcher(repo: str, pr_number: int) -> str: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(pr_number, 570) - return "OPEN" - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "review-state.json" - entry = select_next_entry( - "review", - { - "items": [ - make_issue_item( - "PVTI_10", - 117, - status="Review pool", - title="[Model] GraphPartitioning", - ) - ] - }, - state_file, - repo="CodingThrust/problem-reductions", - pr_resolver=fake_pr_resolver, - pr_state_fetcher=fake_pr_state_fetcher, - ) - self.assertEqual( - entry, - { - "item_id": "PVTI_10", - "number": 570, - "issue_number": 117, - "pr_number": 570, - "status": STATUS_REVIEW_POOL, - "title": "[Model] GraphPartitioning", - }, - ) - - def test_print_next_item_json_emits_rich_payload(self) -> None: - buffer = io.StringIO() - with redirect_stdout(buffer): - rc = print_next_item( - { - "item_id": "PVTI_20", - "number": 615, - "issue_number": 101, - "pr_number": 615, - "status": STATUS_FINAL_REVIEW, - "title": "[Model] MinimumFeedbackVertexSet", - }, - mode="final-review", - fmt="json", - ) - - self.assertEqual(rc, 0) - self.assertEqual( - json.loads(buffer.getvalue()), - { - "mode": "final-review", - "item_id": "PVTI_20", - "number": 615, - "issue_number": 101, - "pr_number": 615, - "status": STATUS_FINAL_REVIEW, - "title": "[Model] MinimumFeedbackVertexSet", - }, - ) - - -class PipelineBoardReviewCandidateTests(unittest.TestCase): - def test_review_candidates_report_ambiguous_issue_cards(self) -> None: - def fake_pr_resolver(repo: str, issue_number: int) -> int | None: - raise AssertionError("ambiguous cards should not resolve by issue search") - - def fake_pr_info_fetcher(repo: str, pr_number: int) -> dict: - self.assertEqual(repo, "CodingThrust/problem-reductions") - return { - 170: {"number": 170, "state": "CLOSED", "title": "Superseded LCS model"}, - 173: { - "number": 173, - "state": "OPEN", - "title": "Fix #109: Add LCS reduction", - }, - }[pr_number] - - candidates = review_candidates( - { - "items": [ - make_issue_item( - "PVTI_10", - 108, - status="Review pool", - title="[Model] LongestCommonSubsequence", - linked_prs=[170, 173], - ) - ] - }, - "CodingThrust/problem-reductions", - fake_pr_resolver, - fake_pr_info_fetcher, - ) - - self.assertEqual(len(candidates), 1) - self.assertEqual( - candidates[0], - { - "item_id": "PVTI_10", - "number": 173, - "issue_number": 108, - "pr_number": 173, - "status": STATUS_REVIEW_POOL, - "title": "[Model] LongestCommonSubsequence", - "eligibility": "ambiguous-linked-prs", - "reason": "multiple linked repo PRs require confirmation", - "recommendation": 173, - "linked_repo_prs": [ - { - "number": 170, - "state": "CLOSED", - "title": "Superseded LCS model", - }, - { - "number": 173, - "state": "OPEN", - "title": "Fix #109: Add LCS reduction", - }, - ], - }, - ) - - def test_review_candidates_report_eligible_for_open_pr(self) -> None: - def fake_pr_resolver(repo: str, issue_number: int) -> int | None: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(issue_number, 117) - return 570 - - def fake_pr_info_fetcher(repo: str, pr_number: int) -> dict: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(pr_number, 570) - return { - "number": 570, - "state": "OPEN", - "title": "Fix #117: [Model] GraphPartitioning", - } - - candidates = review_candidates( - { - "items": [ - make_issue_item( - "PVTI_11", - 117, - status="Review pool", - title="[Model] GraphPartitioning", - ) - ] - }, - "CodingThrust/problem-reductions", - fake_pr_resolver, - fake_pr_info_fetcher, - ) - - self.assertEqual(candidates[0]["eligibility"], "eligible") - self.assertEqual(candidates[0]["reason"], "open PR") - - -class PipelineBoardStatusListTests(unittest.TestCase): - def test_status_items_list_ready_issues(self) -> None: - items = status_items( - { - "items": [ - make_issue_item( - "PVTI_1", - 101, - status="Ready", - title="[Model] ExactCoverBy3Sets", - ), - make_issue_item( - "PVTI_2", - 102, - status="In progress", - title="[Rule] A to B", - ), - ] - }, - STATUS_READY, - ) - self.assertEqual( - items, - [ - { - "item_id": "PVTI_1", - "number": 101, - "issue_number": 101, - "pr_number": None, - "status": STATUS_READY, - "title": "[Model] ExactCoverBy3Sets", - } - ], - ) - - def test_status_items_list_in_progress_issues(self) -> None: - items = status_items( - { - "items": [ - make_issue_item( - "PVTI_1", - 101, - status="Ready", - title="[Model] ExactCoverBy3Sets", - ), - make_issue_item( - "PVTI_2", - 102, - status="In progress", - title="[Rule] A to B", - ), - ] - }, - STATUS_IN_PROGRESS, - ) - self.assertEqual( - items, - [ - { - "item_id": "PVTI_2", - "number": 102, - "issue_number": 102, - "pr_number": None, - "status": STATUS_IN_PROGRESS, - "title": "[Rule] A to B", - } - ], - ) - - -class ReviewCandidatesBatchTests(unittest.TestCase): - def test_review_candidates_uses_batch_fetcher(self) -> None: - """When batch_pr_fetcher is provided, individual fetchers are NOT called.""" - - def fail_pr_info_fetcher(repo: str, pr_number: int) -> dict: - raise AssertionError("should not be called when batch is available") - - def fake_batch_pr_fetcher( - repo: str, pr_numbers: list[int] - ) -> dict[int, dict]: - return { - 570: { - "number": 570, - "state": "OPEN", - "title": "Fix #117", - "url": "https://github.com/o/r/pull/570", - "reviews": [], - } - } - - candidates = review_candidates( - { - "items": [ - make_pr_item("PVTI_1", 570, status="Review pool"), - ] - }, - "CodingThrust/problem-reductions", - None, - fail_pr_info_fetcher, - batch_pr_fetcher=fake_batch_pr_fetcher, - ) - - self.assertEqual(len(candidates), 1) - self.assertEqual(candidates[0]["eligibility"], "eligible") - - def test_review_candidates_batch_falls_back_for_resolved_prs(self) -> None: - """pr_resolver results are not in the batch cache, so individual fetchers are used.""" - resolve_called = [] - - def fake_pr_resolver(repo: str, issue_number: int) -> int | None: - resolve_called.append(issue_number) - return 580 - - def fake_pr_info_fetcher(repo: str, pr_number: int) -> dict: - return {"number": 580, "state": "OPEN", "title": "Fix #120"} - - def fake_batch_pr_fetcher( - repo: str, pr_numbers: list[int] - ) -> dict[int, dict]: - # No linked PRs known ahead of time for this issue - return {} - - candidates = review_candidates( - { - "items": [ - make_issue_item( - "PVTI_2", 120, status="Review pool", title="[Model] Foo" - ), - ] - }, - "CodingThrust/problem-reductions", - fake_pr_resolver, - fake_pr_info_fetcher, - batch_pr_fetcher=fake_batch_pr_fetcher, - ) - - self.assertEqual(len(candidates), 1) - self.assertEqual(candidates[0]["eligibility"], "eligible") - self.assertEqual(resolve_called, [120]) - - -class ReviewEntriesBatchTests(unittest.TestCase): - def test_review_entries_uses_batch_pr_fetcher(self) -> None: - """review_entries should use batch_pr_fetcher and skip individual calls.""" - individual_called: list[int] = [] - - def fail_pr_state_fetcher(repo: str, pr_number: int) -> str: - individual_called.append(pr_number) - raise AssertionError("should not be called when batch cache has the PR") - - def fake_batch_pr_fetcher( - repo: str, pr_numbers: list[int] - ) -> dict[int, dict]: - return { - 570: { - "number": 570, - "state": "OPEN", - "title": "Fix #117", - "reviews": [], - } - } - - entries = review_entries( - {"items": [make_pr_item("PVTI_1", 570, status="Review pool")]}, - "CodingThrust/problem-reductions", - None, - fail_pr_state_fetcher, - batch_pr_fetcher=fake_batch_pr_fetcher, - ) - - self.assertEqual(len(entries), 1) - entry = list(entries.values())[0] - self.assertEqual(entry["pr_number"], 570) - self.assertEqual(individual_called, []) - - def test_review_entries_falls_back_without_batch(self) -> None: - """Without batch_pr_fetcher, individual pr_state_fetcher is used.""" - def fake_pr_state_fetcher(repo: str, pr_number: int) -> str: - return "OPEN" - - entries = review_entries( - {"items": [make_pr_item("PVTI_1", 570, status="Review pool")]}, - "CodingThrust/problem-reductions", - None, - fake_pr_state_fetcher, - ) - - self.assertEqual(len(entries), 1) - - def test_final_review_entries_uses_batch_pr_fetcher(self) -> None: - """final_review_entries should use batch_pr_fetcher and skip individual calls.""" - individual_called: list[int] = [] - - def fail_pr_state_fetcher(repo: str, pr_number: int) -> str: - individual_called.append(pr_number) - raise AssertionError("should not be called when batch cache has the PR") - - def fake_batch_pr_fetcher( - repo: str, pr_numbers: list[int] - ) -> dict[int, dict]: - return { - 570: { - "number": 570, - "state": "OPEN", - "title": "Fix #117", - } - } - - entries = final_review_entries( - {"items": [make_pr_item("PVTI_1", 570, status="Final review")]}, - "CodingThrust/problem-reductions", - None, - fail_pr_state_fetcher, - batch_pr_fetcher=fake_batch_pr_fetcher, - ) - - self.assertEqual(len(entries), 1) - entry = list(entries.values())[0] - self.assertEqual(entry["pr_number"], 570) - self.assertEqual(individual_called, []) - - def test_final_review_entries_falls_back_without_batch(self) -> None: - """Without batch_pr_fetcher, individual pr_state_fetcher is used.""" - def fake_pr_state_fetcher(repo: str, pr_number: int) -> str: - return "OPEN" - - entries = final_review_entries( - {"items": [make_pr_item("PVTI_1", 570, status="Final review")]}, - "CodingThrust/problem-reductions", - None, - fake_pr_state_fetcher, - ) - - self.assertEqual(len(entries), 1) - - def test_final_review_entries_batch_linked_issue(self) -> None: - """final_review_entries batch path works for issue items with linked PRs.""" - def fail_pr_state_fetcher(repo: str, pr_number: int) -> str: - raise AssertionError("should not be called") - - def fake_pr_resolver(repo: str, issue_number: int) -> int | None: - return None # should not be called — linked PR exists - - def fake_batch_pr_fetcher( - repo: str, pr_numbers: list[int] - ) -> dict[int, dict]: - return { - 580: { - "number": 580, - "state": "OPEN", - "title": "Fix #120", - } - } - - entries = final_review_entries( - { - "items": [ - make_issue_item( - "PVTI_2", 120, - status="Final review", - linked_prs=[580], - ), - ] - }, - "CodingThrust/problem-reductions", - fake_pr_resolver, - fail_pr_state_fetcher, - batch_pr_fetcher=fake_batch_pr_fetcher, - ) - - self.assertEqual(len(entries), 1) - entry = list(entries.values())[0] - self.assertEqual(entry["pr_number"], 580) - - -class BatchFetchTests(unittest.TestCase): - def test_batch_fetch_prs_with_reviews_builds_correct_query(self) -> None: - fake_response = { - "data": { - "repository": { - "pr_42": { - "number": 42, - "state": "OPEN", - "title": "Fix foo", - "url": "https://github.com/o/r/pull/42", - "reviews": { - "nodes": [ - { - "author": {"login": "copilot-pull-request-reviewer"}, - "state": "COMMENTED", - }, - ] - }, - }, - "pr_99": { - "number": 99, - "state": "CLOSED", - "title": "Old PR", - "url": "https://github.com/o/r/pull/99", - "reviews": {"nodes": []}, - }, - } - } - } - - with patch( - "pipeline_board.run_gh", return_value=json.dumps(fake_response) - ) as mock_gh: - result = batch_fetch_prs_with_reviews( - "CodingThrust/problem-reductions", [42, 99] - ) - - self.assertEqual(set(result.keys()), {42, 99}) - self.assertEqual(result[42]["state"], "OPEN") - self.assertEqual( - result[42]["reviews"][0]["author"]["login"], - "copilot-pull-request-reviewer", - ) - self.assertEqual(result[99]["state"], "CLOSED") - self.assertEqual(result[99]["reviews"], []) - - mock_gh.assert_called_once() - call_args = mock_gh.call_args[0] - self.assertEqual(call_args[0], "api") - self.assertEqual(call_args[1], "graphql") - - def test_batch_fetch_prs_with_reviews_empty_list(self) -> None: - result = batch_fetch_prs_with_reviews( - "CodingThrust/problem-reductions", [] - ) - self.assertEqual(result, {}) - - def test_batch_fetch_issues_builds_correct_query(self) -> None: - fake_response = { - "data": { - "repository": { - "issue_42": { - "number": 42, - "title": "[Model] Foo", - "body": "## Definition\n...", - "state": "OPEN", - "url": "https://github.com/o/r/issues/42", - "labels": {"nodes": [{"name": "Model"}]}, - "comments": {"nodes": [{"body": "looks good"}]}, - }, - } - } - } - - with patch( - "pipeline_board.run_gh", return_value=json.dumps(fake_response) - ) as mock_gh: - result = batch_fetch_issues("CodingThrust/problem-reductions", [42]) - - self.assertEqual(set(result.keys()), {42}) - self.assertEqual(result[42]["title"], "[Model] Foo") - self.assertEqual(result[42]["state"], "OPEN") - self.assertEqual(result[42]["labels"], [{"name": "Model"}]) - self.assertEqual(result[42]["comments"], [{"body": "looks good"}]) - mock_gh.assert_called_once() - - def test_batch_fetch_issues_empty_list(self) -> None: - result = batch_fetch_issues("CodingThrust/problem-reductions", []) - self.assertEqual(result, {}) - - -class GraphQLBoardFetchTests(unittest.TestCase): - def test_parse_graphql_board_items_extracts_issue_and_pr(self) -> None: - from pipeline_board import _parse_graphql_board_items - - raw = { - "data": { - "node": { - "items": { - "pageInfo": {"hasNextPage": True, "endCursor": "abc123"}, - "nodes": [ - { - "id": "PVTI_1", - "fieldValueByName": {"name": "Ready"}, - "content": {"__typename": "Issue", "number": 101, "title": "[Model] Foo"}, - }, - { - "id": "PVTI_2", - "fieldValueByName": {"name": "Review pool"}, - "content": {"__typename": "PullRequest", "number": 570, "title": "Fix #117"}, - }, - { - "id": "PVTI_3", - "fieldValueByName": {"name": "Backlog"}, - "content": {"__typename": "DraftIssue", "number": None, "title": "Draft"}, - }, - ], - } - } - } - } - - items, has_next, cursor = _parse_graphql_board_items(raw) - - self.assertTrue(has_next) - self.assertEqual(cursor, "abc123") - self.assertEqual(len(items), 2) - self.assertEqual(items[0], {"id": "PVTI_1", "status": "Ready", "content": {"type": "Issue", "number": 101, "title": "[Model] Foo"}, "title": "[Model] Foo"}) - self.assertEqual(items[1], {"id": "PVTI_2", "status": "Review pool", "content": {"type": "PullRequest", "number": 570, "title": "Fix #117"}, "title": "Fix #117"}) - - def test_fetch_board_items_graphql_paginates(self) -> None: - from pipeline_board import fetch_board_items_graphql - - page1 = {"data": {"node": {"items": {"pageInfo": {"hasNextPage": True, "endCursor": "cur1"}, "nodes": [{"id": "PVTI_1", "fieldValueByName": {"name": "Ready"}, "content": {"__typename": "Issue", "number": 1, "title": "A"}}]}}}} - page2 = {"data": {"node": {"items": {"pageInfo": {"hasNextPage": False, "endCursor": None}, "nodes": [{"id": "PVTI_2", "fieldValueByName": {"name": "Ready"}, "content": {"__typename": "Issue", "number": 2, "title": "B"}}]}}}} - - with patch("pipeline_board.run_gh", side_effect=[json.dumps(page1), json.dumps(page2)]) as mock_gh: - result = fetch_board_items_graphql("PVT_test", 200, page_size=1) - - self.assertEqual(len(result["items"]), 2) - self.assertEqual(result["items"][0]["id"], "PVTI_1") - self.assertEqual(result["items"][1]["id"], "PVTI_2") - self.assertEqual(mock_gh.call_count, 2) - - def test_fetch_board_items_lite_uses_graphql_for_default_project(self) -> None: - import pipeline_board - - graphql_response = {"data": {"node": {"items": {"pageInfo": {"hasNextPage": False, "endCursor": None}, "nodes": [{"id": "PVTI_1", "fieldValueByName": {"name": "Ready"}, "content": {"__typename": "Issue", "number": 42, "title": "Test"}}]}}}} - - with patch("pipeline_board.run_gh", return_value=json.dumps(graphql_response)) as mock_gh: - result = pipeline_board.fetch_board_items("CodingThrust", 8, 100, lite=True) - - self.assertEqual(len(result["items"]), 1) - call_args = mock_gh.call_args[0] - self.assertEqual(call_args[0], "api") - self.assertEqual(call_args[1], "graphql") - - def test_fetch_board_items_lite_false_uses_cli(self) -> None: - import pipeline_board - - cli_response = {"items": [{"id": "PVTI_1", "status": "Ready", "content": {"type": "Issue", "number": 42, "title": "Test"}, "linked pull requests": ["https://github.com/CodingThrust/problem-reductions/pull/100"]}]} - - with patch("pipeline_board.run_gh", return_value=json.dumps(cli_response)) as mock_gh: - result = pipeline_board.fetch_board_items("CodingThrust", 8, 100, lite=False) - - self.assertEqual(len(result["items"]), 1) - self.assertIn("linked pull requests", result["items"][0]) - call_args = mock_gh.call_args[0] - self.assertEqual(call_args[0], "project") - - def test_fetch_board_items_lite_non_default_project_uses_cli(self) -> None: - import pipeline_board - - cli_response = {"items": []} - - with patch("pipeline_board.run_gh", return_value=json.dumps(cli_response)) as mock_gh: - result = pipeline_board.fetch_board_items("OtherOrg", 99, 100, lite=True) - - call_args = mock_gh.call_args[0] - self.assertEqual(call_args[0], "project") - - -if __name__ == "__main__": - unittest.main() diff --git a/scripts/test_pipeline_checks.py b/scripts/test_pipeline_checks.py index 41256bbc3..cc9341120 100644 --- a/scripts/test_pipeline_checks.py +++ b/scripts/test_pipeline_checks.py @@ -491,25 +491,6 @@ def test_build_review_context_skips_checks_for_generic_scope(self) -> None: self.assertTrue(context["whitelist"]["skipped"]) self.assertTrue(context["completeness"]["skipped"]) - def test_parse_args_accepts_review_context(self) -> None: - args = parse_args( - [ - "review-context", - "--repo-root", - ".", - "--base", - "abc123", - "--head", - "def456", - "--format", - "json", - ] - ) - - self.assertEqual(args.command, "review-context") - self.assertEqual(args.base, "abc123") - self.assertEqual(args.head, "def456") - def test_parse_args_accepts_issue_context(self) -> None: args = parse_args( [ diff --git a/scripts/test_pipeline_pr.py b/scripts/test_pipeline_pr.py index 0f4a151bf..279c876f2 100644 --- a/scripts/test_pipeline_pr.py +++ b/scripts/test_pipeline_pr.py @@ -12,7 +12,6 @@ build_snapshot, create_pr, emit_result, - edit_pr_body, extract_codecov_summary, extract_linked_issue_number, fetch_linked_issue_bundle, @@ -605,27 +604,6 @@ def test_post_pr_comment_uses_gh_pr_comment_with_body_file(self, check_call: moc ] ) - @mock.patch("pipeline_pr.subprocess.check_call") - def test_edit_pr_body_uses_gh_pr_edit_with_body_file(self, check_call: mock.Mock) -> None: - edit_pr_body( - "CodingThrust/problem-reductions", - 570, - "/tmp/body.md", - ) - - check_call.assert_called_once_with( - [ - "gh", - "pr", - "edit", - "570", - "--repo", - "CodingThrust/problem-reductions", - "--body-file", - "/tmp/body.md", - ] - ) - @mock.patch("pipeline_pr.fetch_current_pr_data_for_repo") @mock.patch("pipeline_pr.run_gh_checked") def test_create_pr_uses_gh_pr_create_and_returns_current_context( @@ -666,7 +644,7 @@ def test_create_pr_uses_gh_pr_create_and_returns_current_context( self.assertEqual(result["pr_number"], 570) self.assertEqual(result["repo"], "CodingThrust/problem-reductions") - def test_parse_args_accepts_comment_and_edit_body_commands(self) -> None: + def test_parse_args_accepts_comment_and_create_commands(self) -> None: comment_args = parse_args( [ "comment", @@ -681,36 +659,6 @@ def test_parse_args_accepts_comment_and_edit_body_commands(self) -> None: self.assertEqual(comment_args.command, "comment") self.assertEqual(comment_args.body_file, "/tmp/comment.md") - edit_args = parse_args( - [ - "edit-body", - "--repo", - "CodingThrust/problem-reductions", - "--pr", - "570", - "--body-file", - "/tmp/body.md", - ] - ) - self.assertEqual(edit_args.command, "edit-body") - self.assertEqual(edit_args.body_file, "/tmp/body.md") - - current_args = parse_args(["current", "--format", "json"]) - self.assertEqual(current_args.command, "current") - - linked_issue_args = parse_args( - [ - "linked-issue", - "--repo", - "CodingThrust/problem-reductions", - "--pr", - "570", - "--format", - "json", - ] - ) - self.assertEqual(linked_issue_args.command, "linked-issue") - create_args = parse_args( [ "create", diff --git a/scripts/test_pipeline_skill_context.py b/scripts/test_pipeline_skill_context.py index c68aac934..f4887e6b1 100644 --- a/scripts/test_pipeline_skill_context.py +++ b/scripts/test_pipeline_skill_context.py @@ -10,40 +10,6 @@ class PipelineSkillContextTests(unittest.TestCase): - def test_parse_args_review_pipeline_defaults(self) -> None: - args = pipeline_skill_context.parse_args( - [ - "review-pipeline", - "--repo", - "CodingThrust/problem-reductions", - ] - ) - - self.assertEqual(args.command, "review-pipeline") - self.assertEqual(args.repo, "CodingThrust/problem-reductions") - self.assertIsNone(args.pr) - self.assertFalse(hasattr(args, "state_file")) - self.assertEqual(args.format, "json") - - def test_parse_args_final_review_with_explicit_values(self) -> None: - args = pipeline_skill_context.parse_args( - [ - "final-review", - "--repo", - "CodingThrust/problem-reductions", - "--pr", - "615", - "--format", - "text", - ] - ) - - self.assertEqual(args.command, "final-review") - self.assertEqual(args.repo, "CodingThrust/problem-reductions") - self.assertEqual(args.pr, 615) - self.assertFalse(hasattr(args, "state_file")) - self.assertEqual(args.format, "text") - def test_parse_args_review_implementation_defaults(self) -> None: args = pipeline_skill_context.parse_args( [ @@ -58,27 +24,6 @@ def test_parse_args_review_implementation_defaults(self) -> None: self.assertIsNone(args.kind) self.assertEqual(args.format, "text") - def test_parse_args_project_pipeline_with_explicit_values(self) -> None: - args = pipeline_skill_context.parse_args( - [ - "project-pipeline", - "--repo", - "CodingThrust/problem-reductions", - "--issue", - "117", - "--repo-root", - "/tmp/repo", - "--format", - "text", - ] - ) - - self.assertEqual(args.command, "project-pipeline") - self.assertEqual(args.repo, "CodingThrust/problem-reductions") - self.assertEqual(args.issue, 117) - self.assertEqual(args.repo_root, Path("/tmp/repo")) - self.assertEqual(args.format, "text") - def test_emit_result_prints_sorted_json_for_all_formats(self) -> None: expected_output = '{\n "a": 2,\n "b": 1\n}\n' @@ -89,110 +34,6 @@ def test_emit_result_prints_sorted_json_for_all_formats(self) -> None: pipeline_skill_context.emit_result({"b": 1, "a": 2}, fmt) self.assertEqual(stdout.getvalue(), expected_output) - def test_emit_result_prints_final_review_text_report(self) -> None: - result = { - "skill": "final-review", - "status": "ready", - "selection": { - "item_id": "PVTI_22", - "pr_number": 615, - "issue_number": 117, - "title": "[Model] GraphPartitioning", - }, - "pr": { - "number": 615, - "title": "Fix #117: [Model] GraphPartitioning", - "url": "https://github.com/CodingThrust/problem-reductions/pull/615", - "comments": { - "counts": { - "human_inline_comments": 1, - "human_issue_comments": 2, - "human_linked_issue_comments": 1, - "human_reviews": 1, - } - }, - "issue_context_text": "Issue #117: Add GraphPartitioning model", - }, - "prep": { - "ready": False, - "checkout": {"worktree_dir": "/tmp/final-pr-615"}, - "merge": {"status": "conflicted", "conflicts": ["src/models/graph_partitioning.rs"]}, - }, - "review_context": { - "subject": {"kind": "model", "name": "GraphPartitioning"}, - "whitelist": {"ok": True, "violations": []}, - "completeness": {"ok": False, "missing": ["paper_display_name"]}, - "changed_files": ["src/models/graph_partitioning.rs", "docs/paper/reductions.typ"], - "diff_stat": "2 files changed, 30 insertions(+), 2 deletions(-)", - }, - } - - stdout = io.StringIO() - with redirect_stdout(stdout): - pipeline_skill_context.emit_result(result, "text") - - rendered = stdout.getvalue() - self.assertIn("# Final Review Packet", rendered) - self.assertIn("- PR: #615", rendered) - self.assertIn("- Board item: `PVTI_22`", rendered) - self.assertIn("## Recommendation Seed", rendered) - self.assertIn("- Suggested mode: conflicted-review", rendered) - self.assertIn("## Deterministic Checks", rendered) - self.assertIn("- Completeness: fail", rendered) - self.assertIn("- `paper_display_name`", rendered) - self.assertIn("## Changed Files", rendered) - - def test_emit_result_prints_review_pipeline_text_report(self) -> None: - result = { - "skill": "review-pipeline", - "status": "ready", - "selection": { - "item_id": "PVTI_11", - "pr_number": 570, - "issue_number": 117, - "title": "[Model] GraphPartitioning", - }, - "pr": { - "number": 570, - "title": "Fix #117: [Model] GraphPartitioning", - "url": "https://github.com/CodingThrust/problem-reductions/pull/570", - "comments": { - "counts": { - "human_inline_comments": 1, - "human_issue_comments": 1, - "human_linked_issue_comments": 1, - } - }, - "issue_context_text": "Issue #117: Add GraphPartitioning model", - "ci": {"state": "failure", "failing": 1, "pending": 0}, - "codecov": {"found": True, "patch_coverage": 84.21}, - }, - "prep": { - "ready": False, - "checkout": { - "worktree_dir": "/tmp/review-pr-570", - "head_ref_name": "issue-117-graph-partitioning", - }, - "merge": {"status": "conflicted", "conflicts": ["src/models/graph_partitioning.rs"]}, - }, - } - - stdout = io.StringIO() - with redirect_stdout(stdout): - pipeline_skill_context.emit_result(result, "text") - - rendered = stdout.getvalue() - self.assertIn("# Review Pipeline Packet", rendered) - self.assertIn("- PR: #570", rendered) - self.assertIn("- Board item: `PVTI_11`", rendered) - self.assertIn("## Recommendation Seed", rendered) - self.assertIn("- Suggested mode: conflicted-fix", rendered) - self.assertIn("- CI state: failure", rendered) - self.assertIn("## Merge Prep", rendered) - self.assertIn("- Worktree: `/tmp/review-pr-570`", rendered) - self.assertIn("- PR head branch: `issue-117-graph-partitioning`", rendered) - self.assertIn("## Linked Issue Context", rendered) - def test_emit_result_prints_review_implementation_text_report(self) -> None: result = { "skill": "review-implementation", @@ -241,155 +82,6 @@ def test_emit_result_prints_review_implementation_text_report(self) -> None: self.assertIn("## Deterministic Checks", rendered) self.assertIn("- Completeness: fail", rendered) - def test_emit_result_prints_project_pipeline_text_report(self) -> None: - result = { - "skill": "project-pipeline", - "status": "ready", - "repo": "CodingThrust/problem-reductions", - "existing_problems": ["BinPacking", "ILP", "GraphColoring"], - "requested_issue": None, - "ready_issues": [ - { - "item_id": "PVTI_1", - "issue_number": 117, - "title": "[Model] GraphPartitioning", - "kind": "model", - "eligible": True, - "blocking_reason": None, - "pending_rule_count": 2, - "summary": "Partition graph vertices into balanced groups.", - }, - { - "item_id": "PVTI_2", - "issue_number": 130, - "title": "[Rule] MultivariateQuadratic to ILP", - "kind": "rule", - "eligible": False, - "blocking_reason": 'model "MultivariateQuadratic" not yet implemented on main', - "pending_rule_count": 0, - "source_problem": "MultivariateQuadratic", - "target_problem": "ILP", - "summary": "Linearize quadratic constraints.", - }, - ], - "in_progress_issues": [ - { - "issue_number": 129, - "title": "[Model] MultivariateQuadratic", - } - ], - } - - stdout = io.StringIO() - with redirect_stdout(stdout): - pipeline_skill_context.emit_result(result, "text") - - rendered = stdout.getvalue() - self.assertIn("# Project Pipeline Packet", rendered) - self.assertIn("- Bundle status: ready", rendered) - self.assertIn("- Ready issues: 2", rendered) - self.assertIn("- In progress issues: 1", rendered) - self.assertIn("## Eligible Ready Issues", rendered) - self.assertIn("- #117 [Model] GraphPartitioning", rendered) - self.assertIn("- Pending rules unblocked: 2", rendered) - self.assertIn("## Blocked Ready Issues", rendered) - self.assertIn('model "MultivariateQuadratic" not yet implemented on main', rendered) - - def test_build_status_result_normalizes_empty_state(self) -> None: - self.assertEqual( - pipeline_skill_context.build_status_result( - "review-pipeline", - status="empty", - ), - { - "skill": "review-pipeline", - "status": "empty", - }, - ) - - def test_build_status_result_normalizes_manual_choice_state(self) -> None: - options = [{"item_id": "PVTI_1", "pr_number": 173}] - - self.assertEqual( - pipeline_skill_context.build_status_result( - "review-pipeline", - status="needs-user-choice", - options=options, - recommendation=173, - ), - { - "skill": "review-pipeline", - "status": "needs-user-choice", - "options": options, - "recommendation": 173, - }, - ) - - def test_main_review_pipeline_emits_ready_bundle_shape(self) -> None: - result = { - "skill": "review-pipeline", - "status": "ready", - "selection": {"item_id": "PVTI_1", "pr_number": 173}, - "prep": {"ready": True}, - "pr": {"number": 173}, - } - - with mock.patch.object( - pipeline_skill_context, - "build_review_pipeline_context", - return_value=result, - ) as builder: - stdout = io.StringIO() - with redirect_stdout(stdout): - exit_code = pipeline_skill_context.main( - [ - "review-pipeline", - "--repo", - "CodingThrust/problem-reductions", - ] - ) - - builder.assert_called_once_with( - repo="CodingThrust/problem-reductions", - pr_number=None, - ) - self.assertEqual(exit_code, 0) - self.assertEqual(json.loads(stdout.getvalue()), result) - - def test_main_final_review_emits_ready_bundle_shape(self) -> None: - result = { - "skill": "final-review", - "status": "ready", - "selection": {"item_id": "PVTI_2", "pr_number": 615}, - "prep": {"ready": True}, - "pr": {"number": 615}, - "review_context": {"files": ["src/lib.rs"]}, - } - - with mock.patch.object( - pipeline_skill_context, - "build_final_review_context", - return_value=result, - ) as builder: - stdout = io.StringIO() - with redirect_stdout(stdout): - exit_code = pipeline_skill_context.main( - [ - "final-review", - "--repo", - "CodingThrust/problem-reductions", - "--pr", - "615", - ] - ) - - builder.assert_called_once_with( - repo="CodingThrust/problem-reductions", - pr_number=615, - ) - self.assertEqual(exit_code, 0) - self.assertEqual(json.loads(stdout.getvalue()), result) - def test_main_review_implementation_emits_ready_bundle_shape(self) -> None: result = { "skill": "review-implementation", @@ -423,376 +115,6 @@ def test_main_review_implementation_emits_ready_bundle_shape(self) -> None: self.assertEqual(exit_code, 0) self.assertEqual(json.loads(stdout.getvalue()), result) - def test_main_project_pipeline_emits_ready_bundle_shape(self) -> None: - result = { - "skill": "project-pipeline", - "status": "ready", - "ready_issues": [{"issue_number": 117}], - } - - with mock.patch.object( - pipeline_skill_context, - "build_project_pipeline_context", - return_value=result, - ) as builder: - stdout = io.StringIO() - with redirect_stdout(stdout): - exit_code = pipeline_skill_context.main( - [ - "project-pipeline", - "--repo", - "CodingThrust/problem-reductions", - "--issue", - "117", - ] - ) - - builder.assert_called_once_with( - repo="CodingThrust/problem-reductions", - issue_number=117, - repo_root=Path("."), - ) - self.assertEqual(exit_code, 0) - self.assertEqual(json.loads(stdout.getvalue()), result) - - def test_build_review_pipeline_context_reports_empty_queue(self) -> None: - result = pipeline_skill_context.build_review_pipeline_context( - repo="CodingThrust/problem-reductions", - pr_number=None, - - review_candidate_fetcher=lambda repo: [], - ) - - self.assertEqual( - result, - { - "skill": "review-pipeline", - "status": "empty", - }, - ) - - def test_build_review_pipeline_context_reports_manual_choice_for_ambiguous_card(self) -> None: - result = pipeline_skill_context.build_review_pipeline_context( - repo="CodingThrust/problem-reductions", - pr_number=None, - - review_candidate_fetcher=lambda repo: [ - { - "item_id": "PVTI_10", - "issue_number": 108, - "pr_number": 173, - "status": "Review pool", - "title": "[Model] LongestCommonSubsequence", - "eligibility": "ambiguous-linked-prs", - "recommendation": 173, - "linked_repo_prs": [ - {"number": 170, "state": "CLOSED", "title": "Superseded LCS model"}, - {"number": 173, "state": "OPEN", "title": "Fix #109: Add LCS reduction"}, - ], - } - ], - ) - - self.assertEqual( - result, - { - "skill": "review-pipeline", - "status": "needs-user-choice", - "options": [ - {"number": 170, "state": "CLOSED", "title": "Superseded LCS model"}, - {"number": 173, "state": "OPEN", "title": "Fix #109: Add LCS reduction"}, - ], - "recommendation": 173, - }, - ) - - def test_build_review_pipeline_context_disambiguates_explicit_pr_choice(self) -> None: - result = pipeline_skill_context.build_review_pipeline_context( - repo="CodingThrust/problem-reductions", - pr_number=173, - - review_candidate_fetcher=lambda repo: [ - { - "item_id": "PVTI_10", - "issue_number": 108, - "pr_number": 173, - "status": "Review pool", - "title": "[Model] LongestCommonSubsequence", - "eligibility": "ambiguous-linked-prs", - "recommendation": 173, - "linked_repo_prs": [ - {"number": 170, "state": "CLOSED", "title": "Superseded LCS model"}, - {"number": 173, "state": "OPEN", "title": "Fix #109: Add LCS reduction"}, - ], - } - ], - pr_context_builder=lambda repo, pr_number: { - "number": pr_number, - "title": "Fix #109: Add LCS reduction", - }, - review_preparer=lambda repo, pr_number: { - "ready": True, - "checkout": {"worktree_dir": "/tmp/review-pr-173"}, - }, - ) - - self.assertEqual( - result, - { - "skill": "review-pipeline", - "status": "ready", - "selection": { - "item_id": "PVTI_10", - "number": 173, - "issue_number": 108, - "pr_number": 173, - "status": "Review pool", - "title": "[Model] LongestCommonSubsequence", - }, - "pr": { - "number": 173, - "title": "Fix #109: Add LCS reduction", - }, - "prep": { - "ready": True, - "checkout": {"worktree_dir": "/tmp/review-pr-173"}, - }, - }, - ) - - def test_build_review_pipeline_context_returns_ready_bundle_for_eligible_pr(self) -> None: - result = pipeline_skill_context.build_review_pipeline_context( - repo="CodingThrust/problem-reductions", - pr_number=None, - review_candidate_fetcher=lambda repo: [ - { - "item_id": "PVTI_11", - "issue_number": 117, - "pr_number": 570, - "number": 570, - "status": "Review pool", - "title": "[Model] GraphPartitioning", - "eligibility": "eligible", - "reason": "open PR", - } - ], - pr_context_builder=lambda repo, pr_number: { - "number": pr_number, - "comments": {"counts": {"human_inline_comments": 0}}, - }, - review_preparer=lambda repo, pr_number: { - "ready": True, - "checkout": {"worktree_dir": "/tmp/review-pr-570"}, - }, - ) - - self.assertEqual(result["status"], "ready") - self.assertEqual(result["selection"]["pr_number"], 570) - self.assertNotIn("claimed", result["selection"]) - self.assertEqual( - result["pr"], - {"number": 570, "comments": {"counts": {"human_inline_comments": 0}}}, - ) - self.assertEqual( - result["prep"], - {"ready": True, "checkout": {"worktree_dir": "/tmp/review-pr-570"}}, - ) - - def test_build_review_pipeline_context_does_not_claim(self) -> None: - """Context generation is read-only — the agent claims after verifying.""" - result = pipeline_skill_context.build_review_pipeline_context( - repo="CodingThrust/problem-reductions", - pr_number=None, - review_candidate_fetcher=lambda repo: [ - { - "item_id": "PVTI_11", - "issue_number": 117, - "pr_number": 570, - "number": 570, - "status": "Review pool", - "title": "[Model] GraphPartitioning", - "eligibility": "eligible", - "reason": "open PR", - } - ], - pr_context_builder=lambda repo, pr_number: {"number": pr_number}, - review_preparer=lambda repo, pr_number: {"ready": True}, - ) - - self.assertEqual(result["status"], "ready") - self.assertEqual(result["selection"]["pr_number"], 570) - self.assertNotIn("claimed", result["selection"]) - - def test_build_review_pipeline_context_explicit_pr_does_not_claim(self) -> None: - """Explicit PR selection is also read-only.""" - result = pipeline_skill_context.build_review_pipeline_context( - repo="CodingThrust/problem-reductions", - pr_number=570, - review_candidate_fetcher=lambda repo: [ - { - "item_id": "PVTI_11", - "issue_number": 117, - "pr_number": 570, - "number": 570, - "status": "Review pool", - "title": "[Model] GraphPartitioning", - "eligibility": "eligible", - "reason": "open PR", - } - ], - pr_context_builder=lambda repo, pr_number: {"number": pr_number}, - review_preparer=lambda repo, pr_number: {"ready": True}, - ) - - self.assertEqual(result["status"], "ready") - self.assertEqual(result["selection"]["pr_number"], 570) - self.assertNotIn("claimed", result["selection"]) - - def test_build_final_review_context_reports_empty_queue(self) -> None: - result = pipeline_skill_context.build_final_review_context( - repo="CodingThrust/problem-reductions", - pr_number=None, - - selection_fetcher=lambda **kwargs: None, - ) - - self.assertEqual( - result, - { - "skill": "final-review", - "status": "empty", - }, - ) - - def test_build_final_review_context_returns_ready_bundle_for_clean_prep(self) -> None: - selection = { - "item_id": "PVTI_22", - "number": 615, - "issue_number": 117, - "pr_number": 615, - "status": "Final review", - "title": "[Model] GraphPartitioning", - } - prep = { - "ready": True, - "checkout": { - "worktree_dir": "/tmp/final-pr-615", - "base_sha": "abc123", - "head_sha": "def456", - }, - "merge": {"status": "clean", "conflicts": [], "likely_complex": False}, - } - pr_context = { - "number": 615, - "title": "Fix #117: [Model] GraphPartitioning", - } - review_context = { - "subject": {"kind": "model", "name": "GraphPartitioning"}, - "whitelist": {"ok": True}, - "completeness": {"ok": True}, - } - - result = pipeline_skill_context.build_final_review_context( - repo="CodingThrust/problem-reductions", - pr_number=None, - - selection_fetcher=lambda **kwargs: selection, - pr_context_builder=lambda repo, pr_number: pr_context, - review_preparer=lambda repo, pr_number: prep, - review_context_builder=lambda *, prep, pr_context: review_context, - ) - - self.assertEqual( - result, - { - "skill": "final-review", - "status": "ready", - "selection": selection, - "pr": pr_context, - "prep": prep, - "review_context": review_context, - }, - ) - - def test_build_final_review_context_keeps_review_context_on_conflicted_prep(self) -> None: - selection = { - "item_id": "PVTI_23", - "number": 620, - "issue_number": 118, - "pr_number": 620, - "status": "Final review", - "title": "[Rule] BinPacking to ILP", - } - prep = { - "ready": False, - "checkout": { - "worktree_dir": "/tmp/final-pr-620", - "base_sha": "abc123", - "head_sha": "def456", - }, - "merge": { - "status": "conflicted", - "conflicts": ["src/rules/binpacking_ilp.rs"], - "likely_complex": False, - }, - } - review_context = { - "subject": {"kind": "rule", "name": "binpacking_ilp"}, - "whitelist": {"ok": True}, - "completeness": {"ok": True}, - } - - result = pipeline_skill_context.build_final_review_context( - repo="CodingThrust/problem-reductions", - pr_number=620, - - selection_fetcher=lambda **kwargs: selection, - pr_context_builder=lambda repo, pr_number: {"number": pr_number}, - review_preparer=lambda repo, pr_number: prep, - review_context_builder=lambda *, prep, pr_context: review_context, - ) - - self.assertEqual(result["status"], "ready") - self.assertEqual(result["prep"]["merge"]["status"], "conflicted") - self.assertEqual(result["review_context"], review_context) - - def test_build_final_review_context_returns_warning_state_on_prep_failure(self) -> None: - selection = { - "item_id": "PVTI_24", - "number": 621, - "issue_number": 119, - "pr_number": 621, - "status": "Final review", - "title": "[Model] FlowShopScheduling", - } - - def fail_prepare(repo: str, pr_number: int) -> dict: - raise RuntimeError("checkout failed") - - result = pipeline_skill_context.build_final_review_context( - repo="CodingThrust/problem-reductions", - pr_number=621, - - selection_fetcher=lambda **kwargs: selection, - pr_context_builder=lambda repo, pr_number: {"number": pr_number}, - review_preparer=fail_prepare, - ) - - self.assertEqual( - result, - { - "skill": "final-review", - "status": "ready-with-warnings", - "selection": selection, - "pr": {"number": 621}, - "prep": {"ready": False, "error": "checkout failed"}, - "review_context": None, - "warnings": [ - "failed to prepare final-review worktree: checkout failed", - ], - }, - ) - def test_build_review_implementation_context_without_current_pr(self) -> None: result = pipeline_skill_context.build_review_implementation_context( repo_root=Path("/tmp/repo"), @@ -824,48 +146,5 @@ def test_build_review_implementation_context_without_current_pr(self) -> None: self.assertNotIn("current_pr", result) self.assertEqual(result["review_context"]["subject"]["kind"], "generic") - def test_build_project_pipeline_context_reports_requested_blocked_issue(self) -> None: - board_data = { - "items": [ - { - "id": "PVTI_2", - "status": "Ready", - "content": { - "type": "Issue", - "number": 130, - "title": "[Rule] MultivariateQuadratic to ILP", - }, - } - ] - } - issue_data = { - 130: { - "number": 130, - "title": "[Rule] MultivariateQuadratic to ILP", - "body": "Linearize quadratic constraints.", - "comments": [], - "labels": [], - "url": "https://github.com/CodingThrust/problem-reductions/issues/130", - } - } - - result = pipeline_skill_context.build_project_pipeline_context( - repo="CodingThrust/problem-reductions", - issue_number=130, - repo_root=Path("/tmp/repo"), - board_fetcher=lambda repo: board_data, - issue_fetcher=lambda repo, issue_number: issue_data[issue_number], - existing_problem_finder=lambda repo_root: {"ILP"}, - ) - - self.assertEqual(result["skill"], "project-pipeline") - self.assertEqual(result["status"], "requested-blocked") - self.assertEqual(result["requested_issue"]["issue_number"], 130) - self.assertEqual( - result["requested_issue"]["blocking_reason"], - 'model "MultivariateQuadratic" not yet implemented on main', - ) - - if __name__ == "__main__": unittest.main() diff --git a/scripts/test_pipeline_worktree.py b/scripts/test_pipeline_worktree.py deleted file mode 100644 index 89a57da19..000000000 --- a/scripts/test_pipeline_worktree.py +++ /dev/null @@ -1,296 +0,0 @@ -#!/usr/bin/env python3 -import unittest -from unittest import mock - -import pipeline_worktree -from pipeline_worktree import ( - prepare_issue_branch, - plan_issue_worktree, - plan_pr_worktree, - summarize_merge, -) - - -class PipelineWorktreeTests(unittest.TestCase): - def test_plan_issue_worktree_sanitizes_slug_and_uses_worktrees_dir(self) -> None: - plan = plan_issue_worktree( - "/tmp/problemreductions", - issue_number=117, - slug="Graph Partitioning / Exact", - base_ref="origin/main", - ) - - self.assertEqual(plan["branch"], "issue-117-graph-partitioning-exact") - self.assertEqual( - plan["worktree_dir"], - "/tmp/problemreductions/.worktrees/issue-117-graph-partitioning-exact", - ) - self.assertEqual(plan["base_ref"], "origin/main") - - def test_plan_pr_worktree_uses_pull_ref_and_sanitized_local_branch(self) -> None: - plan = plan_pr_worktree( - "/tmp/problemreductions", - pr_number=570, - head_ref_name="feature/lcs cleanup", - base_sha="base123", - head_sha="head456", - ) - - self.assertEqual(plan["local_branch"], "review-pr-570-feature-lcs-cleanup") - self.assertEqual( - plan["worktree_dir"], - "/tmp/problemreductions/.worktrees/review-pr-570-feature-lcs-cleanup", - ) - self.assertEqual( - plan["fetch_ref"], - "pull/570/head:review-pr-570-feature-lcs-cleanup", - ) - self.assertEqual(plan["base_sha"], "base123") - self.assertEqual(plan["head_sha"], "head456") - - def test_summarize_merge_clean_result(self) -> None: - summary = summarize_merge( - worktree="/tmp/problemreductions/.worktrees/review-pr-570", - exit_code=0, - conflicts=[], - ) - - self.assertEqual(summary["status"], "clean") - self.assertFalse(summary["likely_complex"]) - self.assertEqual(summary["conflicts"], []) - - def test_summarize_merge_conflicted_result_marks_complex_skill_conflicts(self) -> None: - summary = summarize_merge( - worktree="/tmp/problemreductions/.worktrees/review-pr-570", - exit_code=1, - conflicts=[ - ".claude/skills/add-model/SKILL.md", - "src/models/graph/graph_partitioning.rs", - ], - ) - - self.assertEqual(summary["status"], "conflicted") - self.assertTrue(summary["likely_complex"]) - self.assertEqual( - summary["conflicts"], - [ - ".claude/skills/add-model/SKILL.md", - "src/models/graph/graph_partitioning.rs", - ], - ) - - def test_summarize_merge_without_conflicts_is_aborted(self) -> None: - summary = summarize_merge( - worktree="/tmp/problemreductions/.worktrees/review-pr-570", - exit_code=128, - conflicts=[], - ) - - self.assertEqual(summary["status"], "aborted") - self.assertFalse(summary["likely_complex"]) - - @mock.patch("pipeline_worktree.merge_main") - @mock.patch("pipeline_worktree.checkout_pr_worktree") - def test_prepare_review_bundles_checkout_and_clean_merge( - self, - checkout_pr_worktree: mock.Mock, - merge_main: mock.Mock, - ) -> None: - prepare_review = getattr(pipeline_worktree, "prepare_review", None) - self.assertIsNotNone(prepare_review) - - checkout_payload = { - "pr_number": 570, - "head_ref_name": "feature/lcs cleanup", - "local_branch": "review-pr-570-feature-lcs-cleanup", - "worktree_dir": "/tmp/problemreductions/.worktrees/review-pr-570-feature-lcs-cleanup", - "fetch_ref": "pull/570/head:review-pr-570-feature-lcs-cleanup", - "base_sha": "base123", - "head_sha": "head456", - } - merge_payload = { - "worktree": "/tmp/problemreductions/.worktrees/review-pr-570-feature-lcs-cleanup", - "status": "clean", - "conflicts": [], - "likely_complex": False, - "stdout": "Already up to date.\n", - "stderr": "", - } - checkout_pr_worktree.return_value = checkout_payload - merge_main.return_value = merge_payload - - result = prepare_review( - repo="CodingThrust/problem-reductions", - pr_number=570, - repo_root="/tmp/problemreductions", - ) - - self.assertEqual(result["checkout"], checkout_payload) - self.assertEqual(result["merge"], merge_payload) - self.assertTrue(result["ready"]) - checkout_pr_worktree.assert_called_once_with( - repo="CodingThrust/problem-reductions", - pr_number=570, - repo_root="/tmp/problemreductions", - ) - merge_main.assert_called_once_with( - worktree="/tmp/problemreductions/.worktrees/review-pr-570-feature-lcs-cleanup" - ) - - @mock.patch("pipeline_worktree.merge_main") - @mock.patch("pipeline_worktree.checkout_pr_worktree") - def test_prepare_review_marks_conflicted_merge_not_ready( - self, - checkout_pr_worktree: mock.Mock, - merge_main: mock.Mock, - ) -> None: - prepare_review = getattr(pipeline_worktree, "prepare_review", None) - self.assertIsNotNone(prepare_review) - - checkout_pr_worktree.return_value = { - "pr_number": 570, - "head_ref_name": "feature/lcs cleanup", - "local_branch": "review-pr-570-feature-lcs-cleanup", - "worktree_dir": "/tmp/problemreductions/.worktrees/review-pr-570-feature-lcs-cleanup", - "fetch_ref": "pull/570/head:review-pr-570-feature-lcs-cleanup", - "base_sha": "base123", - "head_sha": "head456", - } - merge_main.return_value = { - "worktree": "/tmp/problemreductions/.worktrees/review-pr-570-feature-lcs-cleanup", - "status": "conflicted", - "conflicts": [ - ".claude/skills/add-model/SKILL.md", - "src/models/graph/graph_partitioning.rs", - ], - "likely_complex": True, - "stdout": "Auto-merging ...\n", - "stderr": "CONFLICT (content): Merge conflict in .claude/skills/add-model/SKILL.md\n", - } - - result = prepare_review( - repo="CodingThrust/problem-reductions", - pr_number=570, - repo_root="/tmp/problemreductions", - ) - - self.assertFalse(result["ready"]) - self.assertEqual( - result["merge"]["conflicts"], - [ - ".claude/skills/add-model/SKILL.md", - "src/models/graph/graph_partitioning.rs", - ], - ) - self.assertTrue(result["merge"]["likely_complex"]) - - @mock.patch("pipeline_worktree.run_git_checked") - @mock.patch("pipeline_worktree.run_git") - def test_prepare_issue_branch_creates_new_branch_when_missing( - self, - run_git: mock.Mock, - run_git_checked: mock.Mock, - ) -> None: - run_git.side_effect = [ - "", # git status --porcelain - "abc123\n", # rev-parse main - "def456\n", # rev-parse HEAD - ] - - with mock.patch("pipeline_worktree.branch_exists", return_value=False): - result = prepare_issue_branch( - issue_number=117, - slug="Graph Partitioning", - base_ref="main", - repo_root="/tmp/problemreductions", - ) - - self.assertEqual(result["branch"], "issue-117-graph-partitioning") - self.assertEqual(result["action"], "create-branch") - self.assertFalse(result["existing_branch"]) - tails = [call.args[1:] for call in run_git_checked.call_args_list] - self.assertIn(("checkout", "main"), tails) - self.assertIn(("checkout", "-b", "issue-117-graph-partitioning"), tails) - - @mock.patch("pipeline_worktree.run_git_checked") - @mock.patch("pipeline_worktree.run_git") - def test_prepare_issue_branch_reuses_existing_branch( - self, - run_git: mock.Mock, - run_git_checked: mock.Mock, - ) -> None: - run_git.side_effect = [ - "", # git status --porcelain - "abc123\n", # rev-parse main - "def456\n", # rev-parse HEAD - ] - - with mock.patch("pipeline_worktree.branch_exists", return_value=True): - result = prepare_issue_branch( - issue_number=117, - slug="Graph Partitioning", - base_ref="main", - repo_root="/tmp/problemreductions", - ) - - self.assertEqual(result["branch"], "issue-117-graph-partitioning") - self.assertEqual(result["action"], "checkout-existing") - self.assertTrue(result["existing_branch"]) - tails = [call.args[1:] for call in run_git_checked.call_args_list] - self.assertIn(("checkout", "main"), tails) - self.assertIn(("checkout", "issue-117-graph-partitioning"), tails) - - - @mock.patch("pipeline_worktree.run_git_checked") - @mock.patch("pipeline_worktree.branch_exists", return_value=False) - def test_enter_creates_worktree_from_base_ref( - self, - branch_exists: mock.Mock, - run_git_checked: mock.Mock, - ) -> None: - import tempfile - with tempfile.TemporaryDirectory() as tmpdir: - result = pipeline_worktree.enter( - name="issue-42", - base_ref="origin/main", - repo_root=tmpdir, - ) - - self.assertEqual(result["branch"], "issue-42") - self.assertIn("issue-42", result["worktree_dir"]) - self.assertIn(".worktrees", result["worktree_dir"]) - self.assertEqual(result["base_ref"], "origin/main") - # Should have called fetch + worktree add - calls = [c.args[1:] for c in run_git_checked.call_args_list] - self.assertIn(("fetch", "origin", "main"), calls) - - @mock.patch("pipeline_worktree.run_git") - @mock.patch("pipeline_worktree.run_git_checked") - @mock.patch("pipeline_worktree.Path") - def test_prepare_review_from_cwd_builds_prep_from_current_dir( - self, - mock_path: mock.Mock, - run_git_checked: mock.Mock, - run_git: mock.Mock, - ) -> None: - mock_path.cwd.return_value = "/tmp/worktree" - run_git.side_effect = [ - "base_abc\n", # merge-base - "head_def\n", # rev-parse HEAD - ] - - result = pipeline_worktree.prepare_review_from_cwd( - repo="CodingThrust/problem-reductions", - pr_number=42, - ) - - self.assertTrue(result["ready"]) - self.assertEqual(result["checkout"]["worktree_dir"], "/tmp/worktree") - self.assertEqual(result["checkout"]["base_sha"], "base_abc") - self.assertEqual(result["checkout"]["head_sha"], "head_def") - self.assertEqual(result["merge"]["status"], "skipped") - run_git_checked.assert_called_once_with("/tmp/worktree", "fetch", "origin", "main") - - -if __name__ == "__main__": - unittest.main() diff --git a/scripts/test_project_board_poll.py b/scripts/test_project_board_poll.py deleted file mode 100644 index 064416059..000000000 --- a/scripts/test_project_board_poll.py +++ /dev/null @@ -1,163 +0,0 @@ -#!/usr/bin/env python3 -import sys -import tempfile -import unittest -from pathlib import Path - -sys.path.insert(0, str(Path(__file__).parent)) - -from project_board_poll import ack_item, process_snapshot - - -def make_issue_item(item_id: str, number: int, status: str = "Ready") -> dict: - return { - "id": item_id, - "status": status, - "content": {"type": "Issue", "number": number}, - } - - -def make_pr_item(item_id: str, number: int, status: str = "Review pool") -> dict: - return { - "id": item_id, - "status": status, - "content": {"type": "PullRequest", "number": number}, - } - - -def with_linked_prs(item: dict, *pr_numbers: int) -> dict: - updated = dict(item) - updated["linked pull requests"] = [ - f"https://github.com/CodingThrust/problem-reductions/pull/{number}" - for number in pr_numbers - ] - return updated - - -class ProjectBoardPollTests(unittest.TestCase): - def test_ready_selection_ignores_ack_until_board_changes(self) -> None: - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "ready-state.json" - snapshot = { - "items": [ - make_issue_item("PVTI_1", 101), - make_issue_item("PVTI_2", 102), - ] - } - - item_id, number = process_snapshot("ready", snapshot, state_file) - self.assertEqual((item_id, number), ("PVTI_1", 101)) - - item_id, number = process_snapshot("ready", snapshot, state_file) - self.assertEqual((item_id, number), ("PVTI_1", 101)) - - ack_item(state_file, "PVTI_1") - item_id, number = process_snapshot("ready", snapshot, state_file) - self.assertEqual((item_id, number), ("PVTI_1", 101)) - - def test_ready_queue_detects_new_item_after_queue_drops_to_zero(self) -> None: - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "ready-state.json" - - item_id, number = process_snapshot( - "ready", - {"items": [make_issue_item("PVTI_1", 101)]}, - state_file, - ) - self.assertEqual((item_id, number), ("PVTI_1", 101)) - - ack_item(state_file, "PVTI_1") - no_item = process_snapshot("ready", {"items": []}, state_file) - self.assertIsNone(no_item) - - item_id, number = process_snapshot( - "ready", - {"items": [make_issue_item("PVTI_2", 102)]}, - state_file, - ) - self.assertEqual((item_id, number), ("PVTI_2", 102)) - - def test_empty_state_file_is_treated_as_no_previous_items(self) -> None: - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "ready-state.json" - state_file.write_text("") - - item_id, number = process_snapshot( - "ready", - {"items": [make_issue_item("PVTI_1", 101)]}, - state_file, - ) - self.assertEqual((item_id, number), ("PVTI_1", 101)) - - def test_review_queue_resolves_issue_cards_to_prs(self) -> None: - def fake_pr_resolver(repo: str, issue_number: int) -> int | None: - self.assertEqual(repo, "CodingThrust/problem-reductions") - return 570 if issue_number == 117 else None - - def fake_pr_state_fetcher(repo: str, pr_number: int) -> str: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(pr_number, 570) - return "OPEN" - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "review-state.json" - item_id, number = process_snapshot( - "review", - {"items": [make_issue_item("PVTI_10", 117, status="Review pool")]}, - state_file, - repo="CodingThrust/problem-reductions", - pr_resolver=fake_pr_resolver, - pr_state_fetcher=fake_pr_state_fetcher, - ) - self.assertEqual((item_id, number), ("PVTI_10", 570)) - - def test_review_queue_skips_closed_pr_cards(self) -> None: - def fake_pr_state_fetcher(repo: str, pr_number: int) -> str: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(pr_number, 570) - return "CLOSED" - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "review-state.json" - no_item = process_snapshot( - "review", - {"items": [make_pr_item("PVTI_10", 570)]}, - state_file, - repo="CodingThrust/problem-reductions", - pr_state_fetcher=fake_pr_state_fetcher, - ) - self.assertIsNone(no_item) - - def test_review_queue_skips_issue_cards_with_mixed_linked_pr_states(self) -> None: - def fake_pr_resolver(repo: str, issue_number: int) -> int | None: - self.assertEqual(repo, "CodingThrust/problem-reductions") - self.assertEqual(issue_number, 108) - return 173 - - def fake_pr_state_fetcher(repo: str, pr_number: int) -> str: - self.assertEqual(repo, "CodingThrust/problem-reductions") - return {170: "CLOSED", 173: "OPEN"}[pr_number] - - with tempfile.TemporaryDirectory() as tmpdir: - state_file = Path(tmpdir) / "review-state.json" - no_item = process_snapshot( - "review", - { - "items": [ - with_linked_prs( - make_issue_item("PVTI_10", 108, status="Review pool"), - 170, - 173, - ) - ] - }, - state_file, - repo="CodingThrust/problem-reductions", - pr_resolver=fake_pr_resolver, - pr_state_fetcher=fake_pr_state_fetcher, - ) - self.assertIsNone(no_item) - - -if __name__ == "__main__": - unittest.main() diff --git a/scripts/test_project_board_recover.py b/scripts/test_project_board_recover.py deleted file mode 100644 index 004b1f14a..000000000 --- a/scripts/test_project_board_recover.py +++ /dev/null @@ -1,189 +0,0 @@ -#!/usr/bin/env python3 -import sys -import unittest -from pathlib import Path - -sys.path.insert(0, str(Path(__file__).parent)) - -from project_board_recover import ( - STATUS_BACKLOG, - STATUS_DONE, - STATUS_FINAL_REVIEW, - STATUS_READY, - STATUS_REVIEW_POOL, - all_checks_green, - build_recovery_plan, - infer_issue_status, -) - - -def make_issue( - number: int, - *, - state: str = "OPEN", - labels: list[str] | None = None, -) -> dict: - return { - "number": number, - "state": state, - "labels": [{"name": name} for name in (labels or [])], - } - - -def make_pr( - number: int, - *, - state: str = "OPEN", - merged: bool = False, - checks: list[dict] | None = None, -) -> dict: - return { - "number": number, - "state": state, - "mergedAt": "2026-03-15T00:00:00Z" if merged else None, - "statusCheckRollup": checks or [], - } - - -class ProjectBoardRecoverTests(unittest.TestCase): - def test_all_checks_green_accepts_successful_runs_and_statuses(self) -> None: - pr = make_pr( - 101, - checks=[ - { - "__typename": "CheckRun", - "status": "COMPLETED", - "conclusion": "SUCCESS", - }, - { - "__typename": "StatusContext", - "state": "SUCCESS", - }, - ], - ) - self.assertTrue(all_checks_green(pr)) - - def test_all_checks_green_rejects_pending_or_failing_checks(self) -> None: - pending = make_pr( - 101, - checks=[ - { - "__typename": "CheckRun", - "status": "IN_PROGRESS", - "conclusion": None, - } - ], - ) - failing = make_pr( - 102, - checks=[ - { - "__typename": "StatusContext", - "state": "FAILURE", - } - ], - ) - self.assertFalse(all_checks_green(pending)) - self.assertFalse(all_checks_green(failing)) - - def test_closed_issue_recovers_to_done(self) -> None: - status, reason = infer_issue_status(make_issue(10, state="CLOSED"), []) - self.assertEqual(status, STATUS_DONE) - self.assertIn("closed", reason) - - def test_good_issue_without_pr_recovers_to_ready(self) -> None: - status, reason = infer_issue_status(make_issue(10, labels=["Good"]), []) - self.assertEqual(status, STATUS_READY) - self.assertIn("Good", reason) - - def test_unchecked_issue_without_pr_recovers_to_backlog(self) -> None: - status, reason = infer_issue_status(make_issue(10), []) - self.assertEqual(status, STATUS_BACKLOG) - self.assertIn("no linked PR", reason) - - def test_issue_with_failure_labels_recovers_to_backlog(self) -> None: - status, reason = infer_issue_status( - make_issue(10, labels=["PoorWritten", "Wrong"]), - [], - ) - self.assertEqual(status, STATUS_BACKLOG) - self.assertIn("failure", reason) - - def test_issue_with_open_pr_no_green_checks_recovers_to_review_pool(self) -> None: - pr = make_pr(200) - status, reason = infer_issue_status( - make_issue(10, labels=["Good"]), - [pr], - ) - self.assertEqual(status, STATUS_REVIEW_POOL) - self.assertIn("still implementing", reason) - - def test_issue_with_green_open_pr_recovers_to_final_review(self) -> None: - pr = make_pr( - 200, - checks=[ - { - "__typename": "CheckRun", - "status": "COMPLETED", - "conclusion": "SUCCESS", - } - ], - ) - status, reason = infer_issue_status( - make_issue(10, labels=["Good"]), - [pr], - ) - self.assertEqual(status, STATUS_FINAL_REVIEW) - self.assertIn("green open PR", reason) - - def test_issue_with_non_green_open_pr_stays_in_review_pool(self) -> None: - pr = make_pr( - 200, - checks=[ - { - "__typename": "CheckRun", - "status": "IN_PROGRESS", - "conclusion": None, - } - ], - ) - status, reason = infer_issue_status( - make_issue(10, labels=["Good"]), - [pr], - ) - self.assertEqual(status, STATUS_REVIEW_POOL) - self.assertIn("still implementing", reason) - - def test_issue_with_closed_unmerged_pr_falls_back_to_backlog(self) -> None: - pr = make_pr(200, state="CLOSED") - status, reason = infer_issue_status(make_issue(10), [pr]) - self.assertEqual(status, STATUS_BACKLOG) - self.assertIn("default", reason) - - def test_build_recovery_plan_ignores_non_model_and_non_rule_titles(self) -> None: - plan = build_recovery_plan( - { - "items": [ - { - "id": "PVTI_1", - "status": None, - "content": {"number": 10, "title": "[Model] Foo"}, - }, - { - "id": "PVTI_2", - "status": None, - "content": {"number": 11, "title": "Meta task"}, - }, - ] - }, - [ - make_issue(10, labels=["Good"]) | {"title": "[Model] Foo"}, - make_issue(11, labels=["Good"]) | {"title": "Meta task"}, - ], - [], - ) - self.assertEqual([entry["issue_number"] for entry in plan], [10]) - - -if __name__ == "__main__": - unittest.main()