diff --git a/CHANGELOG.md b/CHANGELOG.md index 5ae82b0..4d52a1e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -15,6 +15,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 [Datasets](docs/DATASETS.md), [FAQ](docs/FAQ.md), [Roadmap](docs/ROADMAP.md). - `scripts/download_datasets.py` — fetch the six benchmark datasets from scPerturb with checksum verification. +- `examples/` — ready-to-run task descriptions for each benchmark. - GitHub Actions CI: test matrix on Python 3.9–3.12, lint, package build, `CITATION.cff` validation, and a secret scan. - `requirements-ci.txt` — the minimal dependency set needed to run `pytest tests`, diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 6698da3..24adea3 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -107,6 +107,29 @@ Use the [bad research plan](.github/ISSUE_TEMPLATE/bad_plan.yml) template. A run completed successfully but produced wrong science is more informative than a crash, and we have nowhere else to get that data. +### 6. Contribute a real example bundle + +[`examples/outputs/`](examples/outputs/) currently holds a *reference* bundle: the +documents are hand-written to match the schemas the code emits, because whoever +assembled it had no provider key and no downloaded dataset. A bundle from a genuine run +is strictly better, and replacing it is one of the most useful things you can do for +new users — it is the page people read while deciding whether to try this at all. + +```bash +python scripts/export_example_run.py --dataset --out examples/outputs/ +``` + +The script collects the artifacts, strips API keys, absolute paths, usernames and +hostnames, and leaves a `PROVENANCE.md` stub. Two rules: + +- **Complete the provenance stub.** State what produced each file and what data the + metrics came from. An unedited stub will be rejected. +- **Re-read every file yourself.** Scrubbing is best-effort, not a guarantee. Agent + traces can echo secrets in shapes no regex anticipates. + +If your metrics came from a subsample or from synthetic data, say so plainly rather than +letting them read as benchmark results. + --- ## Changes to agent behaviour diff --git a/README.md b/README.md index 3f0899e..1b240c0 100644 --- a/README.md +++ b/README.md @@ -26,7 +26,7 @@ Building a virtual cell model is a months-long loop: read the perturbation-model CellForge collapses that loop into a single command. ```bash -cellforge --dataset-path "data/datasets/.h5ad" \ +cellforge --dataset-path data/datasets/adamson.h5ad \ --task "Predict single-cell gene expression after CRISPRi knockdown in K562." ``` @@ -81,13 +81,11 @@ cellforge --doctor # verify config, dataset dir, literature dir, LLM key, Pyt python scripts/download_datasets.py adamson --out data/datasets/ ``` -The downloader preserves the original scPerturb filenames and downloads three Adamson files; it does not create `adamson.h5ad` or merge them. Replace `.h5ad` below with the input you have selected and prepared for your task. See [dataset preparation](docs/DATASETS.md). - **5. Run it** ```bash cellforge \ - --dataset-path "data/datasets/.h5ad" \ + --dataset-path data/datasets/adamson.h5ad \ --task "Predict post-perturbation gene expression in K562 cells after CRISPRi knockdown. \ Report MSE, PCC, R2 overall and restricted to differentially expressed genes." ``` @@ -104,7 +102,7 @@ data/ Want to also *train* the model it wrote? Add the opt-in execution stage: ```bash -cellforge --phase autorun --dataset-path "data/datasets/.h5ad" --executor local --workers 2 +cellforge --phase autorun --dataset-path data/datasets/adamson.h5ad --executor local --workers 2 ``` 📖 **Full walkthrough:** [docs/QUICKSTART.md](docs/QUICKSTART.md) · **Running on a cluster:** [docs/QUICKSTART.md#running-on-slurm](docs/QUICKSTART.md#running-on-slurm) @@ -166,6 +164,34 @@ Debate runs until every agent clears $c \ge 0.8$ **and** the widest pairwise gap --- +## 👀 See the output before you spend a token + +A complex run costs roughly 80k prompt / 400k completion tokens, plus 4–8 GPU-hours if you train what it writes. You should be able to look at the deliverables first. + +**[`examples/outputs/`](examples/outputs/)** is a complete worked bundle for the Adamson CRISPRi task — the task analysis, the research plan, and the training script, laid out exactly as a run leaves them: + +| Stage | Artifact | | +|---|---|---| +| Task Analysis | [`task_analysis_report.md`](examples/outputs/adamson_crispri/analyses/task_analysis_report.md) | dataset, problem, baselines, agent refinement round | +| Method Design | [`research_plan.md`](examples/outputs/adamson_crispri/plans/research_plan.md) | architecture + protocol — 👤 **your approval gate** | +| Code Generation | [`result.py`](examples/outputs/adamson_crispri/workspace/result.py) | runnable training script — 👤 **approval before any GPU job** | +| Verification | [`verification.json`](examples/outputs/adamson_crispri/workspace/verification.json) | the deterministic, non-LLM acceptance check | + +```bash +cd examples/outputs/adamson_crispri/workspace +python result.py --help # no third-party dependencies needed +python result.py --selftest +``` + +> [!IMPORTANT] +> **The two documents are reference artifacts, not transcripts of a live run.** They are written to match the schemas the code serialises, section for section. `result.py` is real, working, verified code, and its `metrics.json` is the genuine output of really running it — on **synthetic** data, so those numbers are a smoke test and nothing more. +> +> Every file's provenance is stated individually in [`PROVENANCE.md`](examples/outputs/adamson_crispri/PROVENANCE.md). Real benchmark numbers are in [docs/RESULTS.md](docs/RESULTS.md). +> +> Have you run the pipeline for real? [`scripts/export_example_run.py`](scripts/export_example_run.py) packages and scrubs a run into a bundle — a real one should replace this, and we would take that PR gladly. + +--- + ## 🔬 What CellForge designed Six datasets in, six architectures out. None were templates — each was specified by the agents, then implemented and trained end to end. Full model cards: [docs/MODELS.md](docs/MODELS.md). @@ -266,6 +292,7 @@ For a cheap smoke test: `MODEL_NAME=gpt-4o-mini` with `METHOD_DESIGN_MAX_ROUNDS= | | | |---|---| | [**Quickstart**](docs/QUICKSTART.md) | Install, configure, first run, Slurm | +| [**Example outputs**](examples/outputs/) | What the pipeline actually hands you, stage by stage | | [**Architecture**](docs/ARCHITECTURE.md) | The stages, agent roster, coordination score, ablations | | [**Results**](docs/RESULTS.md) | Full benchmark tables, all metrics, all baselines, judge study | | [**Model cards**](docs/MODELS.md) | The six generated architectures in detail | diff --git a/docs/FAQ.md b/docs/FAQ.md index 86a3709..183f7bb 100644 --- a/docs/FAQ.md +++ b/docs/FAQ.md @@ -81,7 +81,7 @@ METHOD_DESIGN_MAX_ROUNDS=2 \ METHOD_DESIGN_MAX_EXPERTS=2 \ METHOD_DESIGN_MAX_TOKENS_PER_CALL=200 \ CELLFORGE_ONLINE_RETRIEVAL=false \ -cellforge --phase task_analysis --dataset-path "data/datasets/.h5ad" --task "..." +cellforge --phase task_analysis --dataset-path data/datasets/adamson.h5ad --task "..." ``` Also: run stage by stage rather than end to end. Reading the analysis report before diff --git a/docs/QUICKSTART.md b/docs/QUICKSTART.md index b2d448e..5a208e0 100644 --- a/docs/QUICKSTART.md +++ b/docs/QUICKSTART.md @@ -140,8 +140,6 @@ python scripts/download_datasets.py --list python scripts/download_datasets.py adamson --out data/datasets/ ``` -The downloader preserves the three original Adamson filenames; it does not create `adamson.h5ad` or merge them. Replace `.h5ad` in the commands below with the input selected and prepared for your task. - Or bring your own `.h5ad`. CellForge expects an AnnData object with a perturbation label in `.obs`; see [DATASETS.md](DATASETS.md#bring-your-own-data) for the expected fields. @@ -154,12 +152,18 @@ fields. ```bash cellforge \ - --dataset-path "data/datasets/.h5ad" \ + --dataset-path data/datasets/adamson.h5ad \ --task "Predict post-perturbation gene expression in K562 cells after CRISPRi \ knockdown. Evaluate on unseen perturbations and unseen cell contexts. Report MSE, \ PCC, and R2, both overall and restricted to differentially expressed genes." ``` +Long task descriptions are easier to keep in a file: + +```bash +cellforge --dataset-path data/datasets/adamson.h5ad --task-file examples/adamson_crispri.txt +``` + ### Stage by stage Recommended the first time, so you can read each artifact before paying for the next @@ -168,8 +172,8 @@ stage. ```bash # ① understand the task and the data, ground it in literature cellforge --phase task_analysis \ - --dataset-path "data/datasets/.h5ad" \ - --task "Predict post-perturbation gene expression in K562 cells after CRISPRi knockdown." + --dataset-path data/datasets/adamson.h5ad \ + --task-file examples/adamson_crispri.txt # ② let the expert agents debate an architecture cellforge --phase method_design @@ -184,7 +188,7 @@ The execution stage is opt-in and never runs implicitly. ```bash cellforge --phase autorun \ - --dataset-path "data/datasets/.h5ad" \ + --dataset-path data/datasets/adamson.h5ad \ --executor local \ --workers 2 ``` @@ -220,7 +224,7 @@ model-generated commands and output — review before sharing. ```bash cellforge --phase autorun \ - --dataset-path "data/datasets/.h5ad" \ + --dataset-path data/datasets/adamson.h5ad \ --codegen-backend codex \ --executor slurm \ --partition gpu \ @@ -263,7 +267,7 @@ METHOD_DESIGN_MAX_ROUNDS=2 \ METHOD_DESIGN_MAX_EXPERTS=2 \ METHOD_DESIGN_MAX_TOKENS_PER_CALL=200 \ CELLFORGE_ONLINE_RETRIEVAL=false \ -cellforge --phase task_analysis --dataset-path "data/datasets/.h5ad" \ +cellforge --phase task_analysis --dataset-path data/datasets/adamson.h5ad \ --task "Smoke test." ``` diff --git a/docs/README.md b/docs/README.md index 07b26e2..a2309f6 100644 --- a/docs/README.md +++ b/docs/README.md @@ -3,6 +3,7 @@ | Document | Read it when | |---|---| | [Quickstart](QUICKSTART.md) | You want a first run working, locally or on Slurm | +| [Example outputs](../examples/outputs/) | You want to see what a run produces before paying for one | | [Architecture](ARCHITECTURE.md) | You want to know how the agents actually reach a decision | | [Results](RESULTS.md) | You want the benchmark numbers and how they were produced | | [Model cards](MODELS.md) | You want details of the six architectures CellForge designed | diff --git a/docs/RESULTS.md b/docs/RESULTS.md index 55f0777..fca38aa 100644 --- a/docs/RESULTS.md +++ b/docs/RESULTS.md @@ -204,21 +204,19 @@ loop itself gives up. --- -## Running on your prepared benchmark input - -The downloader preserves original filenames and does not merge the three Adamson files. Replace `.h5ad` with your prepared benchmark input, using the preprocessing and split definitions in [DATASETS.md](DATASETS.md). +## Reproducing this ```bash # 1. get the data python scripts/download_datasets.py --all --out data/datasets/ # 2. run the pipeline for one benchmark -cellforge --dataset-path "data/datasets/.h5ad" \ - --task "Predict post-perturbation gene expression in K562 cells after CRISPRi knockdown." +cellforge --dataset-path data/datasets/adamson.h5ad \ + --task-file examples/adamson_crispri.txt # 3. train the model it designed, with the paper's split settings cellforge --phase autorun \ - --dataset-path "data/datasets/.h5ad" \ + --dataset-path data/datasets/adamson.h5ad \ --executor slurm --partition \ --gres gpu:1 --mem 32G --slurm-time 08:00:00 \ --split-ood-ratio 0.2 --split-val-ratio 0.1 --split-seed 42 diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index 5d71156..2f8c601 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -15,6 +15,7 @@ The gap between "this looks interesting" and "I ran it" is where most people lea | | Item | Why | |---|---|---| | 🔴 | **Colab / notebook demo** | A hosted notebook that runs Task Analysis on a small dataset with one API key. No conda, no npm, no cluster. This is the single highest-leverage thing on this list. | +| 🔴 | **Pre-computed example outputs** | Commit one real analysis report, research plan, and `result.py` under `examples/outputs/` so people can see what CellForge produces before spending a token. | | 🔴 | **`--dry-run`** | Walk the pipeline with a stub LLM, print what *would* be called and the estimated token cost. Catches configuration mistakes for free. | | 🔴 | **Cost estimate before the expensive stage** | Print projected tokens and dollars before Method Design starts, and ask. | | 🔴 | **Docker image** | `docker run cellforge` with the scientific stack baked in. The full pip install is several GB and a common first-run failure. | diff --git a/package-lock.json b/package-lock.json index 68f8ba3..03bf46c 100644 --- a/package-lock.json +++ b/package-lock.json @@ -8,13 +8,13 @@ "name": "cellforge-codex-autorun", "version": "0.1.0", "dependencies": { - "@openai/codex-sdk": "^0.146.0" + "@openai/codex-sdk": "^0.158.0" } }, "node_modules/@openai/codex": { - "version": "0.146.0", - "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.146.0.tgz", - "integrity": "sha512-yG3sPWNda/2YAIQIDq9MrrjoCTIQ7rxYM5IasrG3VBcuhCLTkgeg/JzqmJq1V98RE4MJ5jCxDXXQlOjrditFRw==", + "version": "0.158.0", + "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.158.0.tgz", + "integrity": "sha512-GBhcKpQmVLsCtEP5mUf6WFye6QQgTKotwsXrYPM0GFmEsoOSahik6hkf2FHabX9kZBl0Y0/PQowLu7uYMKT8dg==", "license": "Apache-2.0", "bin": { "codex": "bin/codex.js" @@ -23,19 +23,19 @@ "node": ">=16" }, "optionalDependencies": { - "@openai/codex-darwin-arm64": "npm:@openai/codex@0.146.0-darwin-arm64", - "@openai/codex-darwin-x64": "npm:@openai/codex@0.146.0-darwin-x64", - "@openai/codex-linux-arm64": "npm:@openai/codex@0.146.0-linux-arm64", - "@openai/codex-linux-x64": "npm:@openai/codex@0.146.0-linux-x64", - "@openai/codex-win32-arm64": "npm:@openai/codex@0.146.0-win32-arm64", - "@openai/codex-win32-x64": "npm:@openai/codex@0.146.0-win32-x64" + "@openai/codex-darwin-arm64": "npm:@openai/codex@0.158.0-darwin-arm64", + "@openai/codex-darwin-x64": "npm:@openai/codex@0.158.0-darwin-x64", + "@openai/codex-linux-arm64": "npm:@openai/codex@0.158.0-linux-arm64", + "@openai/codex-linux-x64": "npm:@openai/codex@0.158.0-linux-x64", + "@openai/codex-win32-arm64": "npm:@openai/codex@0.158.0-win32-arm64", + "@openai/codex-win32-x64": "npm:@openai/codex@0.158.0-win32-x64" } }, "node_modules/@openai/codex-darwin-arm64": { "name": "@openai/codex", - "version": "0.146.0-darwin-arm64", - "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.146.0-darwin-arm64.tgz", - "integrity": "sha512-nb61yX4r5L6Z0dlC4o3u0GAK1YCd4TUvjaB382bajDoh84V+uv2hTBIVZ++fgXWV9yoeuNrNnNcn7GoTGOe2Tg==", + "version": "0.158.0-darwin-arm64", + "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.158.0-darwin-arm64.tgz", + "integrity": "sha512-0OKSjlWY1j4Ld1fT87QttNw3Y2SthcXi4GcrWSHjleZg1n86eG3+shJl4Pv+siUmsBJhWlNmm6rRO4Yv8ZyQLg==", "cpu": [ "arm64" ], @@ -50,9 +50,9 @@ }, "node_modules/@openai/codex-darwin-x64": { "name": "@openai/codex", - "version": "0.146.0-darwin-x64", - "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.146.0-darwin-x64.tgz", - "integrity": "sha512-hTQR5jy/ObfTf1MDnuJCZJAe+SljKE8DDwQWN6lDFgjsPhMQz852U2tILt8Ei+G5GkQSzemHYKl2AYPwW0Y5xw==", + "version": "0.158.0-darwin-x64", + "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.158.0-darwin-x64.tgz", + "integrity": "sha512-FrX1o3APrL7F6QkO8z08Rq8lJitH2sNI7pkebA02eYA103hDs3fyy9d32zeLLQJUeAhmZA4pV9xJR4m7cVyNpQ==", "cpu": [ "x64" ], @@ -67,9 +67,9 @@ }, "node_modules/@openai/codex-linux-arm64": { "name": "@openai/codex", - "version": "0.146.0-linux-arm64", - "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.146.0-linux-arm64.tgz", - "integrity": "sha512-qiYDxkkEFnXG7joadJW6Q+XcgyDXCpGdpa9nk/c+i0gEomur1j7bHvx12NfWWCF/y8Tqri6ay+FLuC2MjdehtA==", + "version": "0.158.0-linux-arm64", + "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.158.0-linux-arm64.tgz", + "integrity": "sha512-T9AgGcoU7HNsxJ4prT/R9YvztyEmz9kWP/VlT2cbCop4PVwiqfG9YMJtLo4j2ScewvSaTl4sQFwla+fwGpoRZg==", "cpu": [ "arm64" ], @@ -84,9 +84,9 @@ }, "node_modules/@openai/codex-linux-x64": { "name": "@openai/codex", - "version": "0.146.0-linux-x64", - "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.146.0-linux-x64.tgz", - "integrity": "sha512-fswvyGprAPCMiOEue/7MKMk7pCjh9kZIJfJX5i9atmfnmGYbYCcUhZsEH9LEP0+0t5xyPqDbfNXY7NSxIVuXxA==", + "version": "0.158.0-linux-x64", + "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.158.0-linux-x64.tgz", + "integrity": "sha512-mY12GZPM8TuOWVGxCyNV2NSA8t1uIupm9Nhu9VeQVEvvxBKN08U8bLH+vn7oISUHZPC8bwhROIbo/ZUxqV6XMg==", "cpu": [ "x64" ], @@ -100,12 +100,12 @@ } }, "node_modules/@openai/codex-sdk": { - "version": "0.146.0", - "resolved": "https://registry.npmjs.org/@openai/codex-sdk/-/codex-sdk-0.146.0.tgz", - "integrity": "sha512-lhlcfmufd4EvjquERH3HNG0/fuxPQJBjFsAjtP2LJjAsuyMZk245ci8xfvT4QoZ7Vnbe15y7rj/NtMNcUJIJOg==", + "version": "0.158.0", + "resolved": "https://registry.npmjs.org/@openai/codex-sdk/-/codex-sdk-0.158.0.tgz", + "integrity": "sha512-lZIsVxPjTkgaz5nBlbUy5qd/cB7kjqXjtKMqvQJ+gaWkXevp9bWcyk3Meyix07GFE8+P8cYH3JxdgBikzaPvdg==", "license": "Apache-2.0", "dependencies": { - "@openai/codex": "0.146.0" + "@openai/codex": "0.158.0" }, "engines": { "node": ">=18" @@ -113,9 +113,9 @@ }, "node_modules/@openai/codex-win32-arm64": { "name": "@openai/codex", - "version": "0.146.0-win32-arm64", - "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.146.0-win32-arm64.tgz", - "integrity": "sha512-EW6zdjDe+SLX2Iw+xymJ5+Pz2+DGexdstfFHXh4Ub+TfJsQPiMjGfZfNaoWgdJ2FsqSIzVKu2+G0KCMGYz2W8g==", + "version": "0.158.0-win32-arm64", + "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.158.0-win32-arm64.tgz", + "integrity": "sha512-Jw1u0q0+5PG97jPkINxE3UCFtsYBN8Af+IjjM0zlCO675Sv5lNsXU2K9aIaXwVeQIKw8q2lbPAR4SEkq2rSOoA==", "cpu": [ "arm64" ], @@ -130,9 +130,9 @@ }, "node_modules/@openai/codex-win32-x64": { "name": "@openai/codex", - "version": "0.146.0-win32-x64", - "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.146.0-win32-x64.tgz", - "integrity": "sha512-b3lxMYeR0+IhstNo4JjX1P9cPc1xwVcCVkPd1lD1wpWPJ0SBhpIkPczwbu3ZRkJcdyl342+rgyf4DUrbZLdrGA==", + "version": "0.158.0-win32-x64", + "resolved": "https://registry.npmjs.org/@openai/codex/-/codex-0.158.0-win32-x64.tgz", + "integrity": "sha512-IaUmY11Zdqa/Zok6kE0Z5375pXtClKRbS8P1Gh2C73Aq65CaGn9N8gT6BXGC3+XD6yamAIiEQW5xM842UCCGow==", "cpu": [ "x64" ], diff --git a/package.json b/package.json index 52294de..0aba156 100644 --- a/package.json +++ b/package.json @@ -4,6 +4,6 @@ "private": true, "type": "module", "dependencies": { - "@openai/codex-sdk": "^0.146.0" + "@openai/codex-sdk": "^0.158.0" } } diff --git a/scripts/download_datasets.py b/scripts/download_datasets.py index 6e3e679..f152bb7 100755 --- a/scripts/download_datasets.py +++ b/scripts/download_datasets.py @@ -368,7 +368,7 @@ def main() -> int: print(f" ✅ All files downloaded and verified in {out}\n") print(" Next:") - print(f" cellforge --dataset-path {out}/.h5ad --task \"\"\n") + print(f" cellforge --dataset-path {out}/.h5ad --task-file examples/.txt\n") return 0 diff --git a/scripts/export_example_run.py b/scripts/export_example_run.py new file mode 100644 index 0000000..046b7eb --- /dev/null +++ b/scripts/export_example_run.py @@ -0,0 +1,331 @@ +#!/usr/bin/env python3 +"""Package a real CellForge run into a committable example bundle. + +A run leaves artifacts scattered across ``data/analyses/``, ``data/plans/`` and a +generation workspace, and those artifacts are not safe to commit as they stand: +agent traces echo API keys, and every path in them is absolute and specific to +the machine that produced them. This script collects the artifacts, scrubs them, +and writes the directory layout that ``examples/outputs/`` expects. + + python scripts/export_example_run.py --dataset adamson_crispri \\ + --out examples/outputs/adamson_crispri + +It refuses to write anything it could not scrub. If a file still looks like it +contains a credential after redaction, the export fails loudly rather than +quietly committing a secret. + +Nothing here needs an API key or a GPU — it only moves and rewrites text. +""" + +from __future__ import annotations + +import argparse +import json +import os +import re +import shutil +import socket +import sys +from pathlib import Path +from typing import Dict, Iterable, List, Optional, Sequence, Tuple + +REPO_ROOT = Path(__file__).resolve().parent.parent + +# Text extensions get scrubbed; anything else is copied verbatim or skipped. +TEXT_SUFFIXES = {".md", ".json", ".jsonl", ".mmd", ".py", ".txt", ".yaml", ".yml", ".cfg", ".toml"} + +# Files that must never end up in a bundle regardless of location. +NEVER_COPY = { + ".env", + "config.local.json", + "id_rsa", + "credentials.json", + "token.json", +} + +SKIP_SUFFIXES = {".h5ad", ".h5", ".hdf5", ".loom", ".mtx", ".npz", ".npy", + ".pkl", ".pickle", ".pt", ".pth", ".ckpt", ".onnx"} + +# Patterns replaced during scrubbing. Order matters: longer/more specific first. +SECRET_PATTERNS: Sequence[Tuple[str, str]] = ( + (r"sk-ant-[A-Za-z0-9_\-]{16,}", "$ANTHROPIC_API_KEY"), + (r"sk-proj-[A-Za-z0-9_\-]{16,}", "$OPENAI_API_KEY"), + (r"sk-[A-Za-z0-9]{32,}", "$API_KEY"), + (r"\bgh[pousr]_[A-Za-z0-9]{20,}", "$GITHUB_TOKEN"), + (r"\bAKIA[0-9A-Z]{16}\b", "$AWS_ACCESS_KEY_ID"), + (r"\bAIza[0-9A-Za-z_\-]{35}\b", "$GOOGLE_API_KEY"), + (r"\bey[A-Za-z0-9_\-]{10,}\.[A-Za-z0-9_\-]{10,}\.[A-Za-z0-9_\-]{10,}", "$JWT"), + (r"(?i)\b(api[_\-]?key|secret|token|password)\b(\s*[:=]\s*)[\"']?[A-Za-z0-9_\-]{16,}[\"']?", + r"\1\2$REDACTED"), +) + +# Anything still matching these after scrubbing aborts the export. +TRIPWIRES: Sequence[str] = ( + r"sk-ant-[A-Za-z0-9_\-]{16,}", + r"sk-[A-Za-z0-9]{32,}", + r"\bgh[pousr]_[A-Za-z0-9]{20,}", + r"\bAKIA[0-9A-Z]{16}\b", + r"\bAIza[0-9A-Za-z_\-]{35}\b", +) + + +class Scrubber: + """Rewrites machine-specific and secret-looking text out of run artifacts.""" + + def __init__(self, workspace: Path, extra: Optional[Dict[str, str]] = None) -> None: + self.replacements: List[Tuple[str, str]] = [] + + # Longest paths first, so /home/u/CellForge/data resolves before /home/u. + raw: Dict[str, str] = { + str(workspace.resolve()): "$WORKSPACE", + str(REPO_ROOT): "$REPO", + str(Path.home()): "$HOME", + sys.executable: "python3", + } + try: + raw[socket.gethostname()] = "$HOSTNAME" + except OSError: + pass + for name in ("USER", "LOGNAME"): + value = os.environ.get(name) + if value and len(value) > 2: + raw[value] = "$USER" + if extra: + raw.update(extra) + + for needle, token in sorted(raw.items(), key=lambda kv: -len(kv[0])): + if needle and needle not in {"/", "."}: + self.replacements.append((needle, token)) + + def scrub(self, text: str) -> str: + for needle, token in self.replacements: + text = text.replace(needle, token) + for pattern, token in SECRET_PATTERNS: + text = re.sub(pattern, token, text) + return text + + @staticmethod + def tripwire(text: str) -> List[str]: + return [p for p in TRIPWIRES if re.search(p, text)] + + +def newest(directory: Path, pattern: str) -> Optional[Path]: + """Most recently modified file matching a glob, or None.""" + matches = sorted(directory.glob(pattern), key=lambda p: p.stat().st_mtime, reverse=True) + return matches[0] if matches else None + + +def copy_file(src: Path, dst: Path, scrubber: Scrubber, dry_run: bool) -> Tuple[bool, str]: + """Copy one file, scrubbing it if it is text. Returns (copied, note).""" + if src.name in NEVER_COPY: + return False, "refused (never-copy list)" + if src.suffix.lower() in SKIP_SUFFIXES: + return False, "skipped (data/model artifact)" + + if src.suffix.lower() not in TEXT_SUFFIXES: + if dry_run: + return True, "would copy verbatim" + dst.parent.mkdir(parents=True, exist_ok=True) + shutil.copy2(src, dst) + return True, "copied verbatim" + + try: + text = src.read_text(encoding="utf-8") + except (UnicodeDecodeError, OSError) as exc: + return False, f"skipped (unreadable: {exc})" + + cleaned = scrubber.scrub(text) + hits = scrubber.tripwire(cleaned) + if hits: + raise SystemExit( + f"\nABORT: {src} still matches credential patterns after scrubbing:\n" + + "\n".join(f" - {h}" for h in hits) + + "\n\nNothing was written. Redact the file by hand, or add a pattern to " + "SECRET_PATTERNS in this script, then re-run.\n" + ) + + scrubbed = " (scrubbed)" if cleaned != text else "" + if dry_run: + return True, "would copy" + scrubbed + note = "copied" + scrubbed + dst.parent.mkdir(parents=True, exist_ok=True) + dst.write_text(cleaned, encoding="utf-8") + return True, note + + +def collect(dataset: str, workspace: Path, code_dir: Optional[Path]) -> List[Tuple[Path, str]]: + """Find the artifacts of a run. Returns (source, destination-relative) pairs.""" + found: List[Tuple[Path, str]] = [] + + analyses = workspace / "data" / "analyses" / dataset + if analyses.is_dir(): + for name in ("task_analysis_report.md", "task_analysis.json"): + candidate = analyses / name + if candidate.is_file(): + found.append((candidate, f"analyses/{name}")) + for extra in sorted(analyses.glob("*.md")): + if extra.name != "task_analysis_report.md": + found.append((extra, f"analyses/{extra.name}")) + + plans = workspace / "data" / "plans" / dataset + if not plans.is_dir(): + plans = workspace / "data" / "plans" + if plans.is_dir(): + # A run stamps the filename; normalise it so the bundle is stable. + for suffix in ("md", "json", "mmd"): + latest = newest(plans, f"research_plan_*.{suffix}") + if latest is not None: + found.append((latest, f"plans/research_plan.{suffix}")) + + if code_dir is not None and code_dir.is_dir(): + for name in ("result.py", "metrics.json", "verification.json", "requirements.txt"): + candidate = code_dir / name + if candidate.is_file(): + found.append((candidate, f"workspace/{name}")) + + return found + + +PROVENANCE_STUB = """# Provenance + + + +## Summary + +This bundle was exported from a real CellForge run with +`scripts/export_example_run.py`. + +## Run details + +| Field | Value | +| --- | --- | +| Dataset | {dataset} | +| Date of run | | +| CellForge commit | {commit} | +| LLM provider / model | | +| Code generation backend | | +| Hardware | | +| Wall-clock time | | +| Approximate token cost | | + +## Per-file provenance + +| File | What it is | Real run output? | +| --- | --- | --- | +{rows} + +## Metrics + + + +## Known caveats + + +""" + + +def write_provenance(out: Path, dataset: str, copied: List[str], dry_run: bool) -> None: + target = out / "PROVENANCE.md" + if target.exists(): + print(" PROVENANCE.md exists, leaving it alone") + return + + commit = "unknown" + head = REPO_ROOT / ".git" / "HEAD" + try: + ref = head.read_text(encoding="utf-8").strip() + if ref.startswith("ref: "): + commit = (REPO_ROOT / ".git" / ref[5:]).read_text(encoding="utf-8").strip()[:12] + elif ref: + commit = ref[:12] + except OSError: + pass + + rows = "\n".join(f"| `{path}` | | Yes |" for path in copied) + body = PROVENANCE_STUB.format(dataset=dataset, commit=commit, rows=rows or "| | | |") + if dry_run: + print(" would write PROVENANCE.md stub") + return + target.write_text(body, encoding="utf-8") + print(" wrote PROVENANCE.md stub — complete it before opening a PR") + + +def main(argv: Optional[Sequence[str]] = None) -> int: + parser = argparse.ArgumentParser( + description="Package a real CellForge run into a committable example bundle.", + formatter_class=argparse.ArgumentDefaultsHelpFormatter, + ) + parser.add_argument("--dataset", required=True, + help="Dataset name as used under data/analyses/ and data/plans/.") + parser.add_argument("--out", type=Path, required=True, + help="Destination bundle directory.") + parser.add_argument("--workspace", type=Path, default=Path.cwd(), + help="Run workspace containing data/analyses and data/plans.") + parser.add_argument("--code-dir", type=Path, default=None, + help="Directory holding the generated result.py. Defaults to " + "the newest .cellforge_workspaces/* under --workspace.") + parser.add_argument("--dry-run", action="store_true", + help="Report what would be written without writing it.") + parser.add_argument("--force", action="store_true", + help="Overwrite a non-empty --out directory.") + args = parser.parse_args(argv) + + workspace = args.workspace.expanduser().resolve() + if not workspace.is_dir(): + parser.error(f"--workspace is not a directory: {workspace}") + + code_dir = args.code_dir + if code_dir is None: + pool = workspace / ".cellforge_workspaces" + if pool.is_dir(): + candidates = [p for p in pool.iterdir() if p.is_dir()] + if candidates: + code_dir = max(candidates, key=lambda p: p.stat().st_mtime) + if code_dir is not None: + code_dir = code_dir.expanduser().resolve() + + out = args.out.expanduser().resolve() + if out.exists() and any(out.iterdir()) and not args.force and not args.dry_run: + parser.error(f"--out is not empty: {out}. Pass --force to overwrite.") + + print(f"workspace : {workspace}") + print(f"code dir : {code_dir or '(none found)'}") + print(f"dataset : {args.dataset}") + print(f"out : {out}\n") + + artifacts = collect(args.dataset, workspace, code_dir) + if not artifacts: + print("No run artifacts found. Expected at least one of:") + print(f" {workspace}/data/analyses/{args.dataset}/task_analysis_report.md") + print(f" {workspace}/data/plans/{args.dataset}/research_plan_*.md") + print(f" {code_dir or ''}/result.py") + print("\nPass --workspace / --code-dir if the run lives elsewhere.") + return 1 + + scrubber = Scrubber(workspace) + copied: List[str] = [] + for src, relative in artifacts: + ok, note = copy_file(src, out / relative, scrubber, args.dry_run) + print(f" [{'ok' if ok else '--'}] {relative:44s} {note}") + if ok: + copied.append(relative) + + if not copied: + print("\nNothing was eligible for copying.") + return 1 + + write_provenance(out, args.dataset, copied, args.dry_run) + + print(f"\n{len(copied)} file(s) {'would be ' if args.dry_run else ''}exported.") + if not args.dry_run: + print("\nBefore committing:") + print(" 1. Complete PROVENANCE.md — an unedited stub will be rejected.") + print(" 2. Re-read every file. Scrubbing is best-effort, not a guarantee.") + print(" 3. Confirm no dataset files or model weights were included.") + return 0 + + +if __name__ == "__main__": + sys.exit(main())