Skip to content

Fix scoring provenance, input validation and report integrity - #34

Merged
DimaMolod merged 8 commits into
mainfrom
fix/scoring-and-report-integrity
Oct 1, 2026
Merged

DimaMolod merged 8 commits into
mainfrom
fix/scoring-and-report-integrity

Conversation

@DimaMolod

@DimaMolod DimaMolod commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

AlphaJudge could reuse stale scores, reject native AlphaFold-Multimer output, choose structures nondeterministically, and omit PDF table rows. This change makes input handling and cache reuse explicit and preserves all report rows.

  • Validate cached CSVs against scoring options, input fingerprints, software/calibration and the CSV checksum. Default input checks use size, mtime, ctime and file identity; --cache_validation content enables full hashes. Metadata can miss same-size edits with coarse or preserved timestamps. Old manifests and incomplete scoring runs are recomputed; files are published atomically.
  • Separate PAE rendering from scoring. PNG options and rendering failures do not invalidate CSVs. Missing or stale PNGs can be regenerated on a cache hit without constructing interfaces or rerunning biophysical scores.
  • Accept native AF2 Multimer rankings and recover individual confidences/full PAE from result pickles, including gzip/xz. Preserve missing individual scores when only combined confidence exists.
  • Record the parser backend and selected structure. Explicit backend provenance prevents Boltz filenames from selecting an AF2 calibration. AF2 selection is deterministic; --af2_structure unrelaxed overrides the default relaxed-first preference. Historical calibration predates per-row structure provenance, so its relaxation policy cannot be reconstructed reliably; the frozen distributions are retained.
  • Reject corrupt or unalignable PAE with file context and per-model tracebacks; discard partial rows from failed models. Missing AF3 full-confidence files raise FileNotFoundError even when the summary contains chain-pair minima; combined exports carrying actual full PAE remain supported.
  • Preserve Bio.PDB residue hierarchy, paginate interface and aggregate ranking tables, preserve the caller's Matplotlib backend/settings, and reject invalid aggregate-report CLI options before scoring.
  • Record cutoffs, calibration identity and the actual contributing features/count. Reports explain recalculation and qualify custom/unknown cutoffs. Remove the unsupported interface-area percentile registry entry while retaining raw area.

pDockQ2 coefficients and its existing pairwise formula are unchanged. A regression preserves agreement with the larger directional score from IPSAE at 6174cf9. Frozen quantile ladders and percentile tie policy are unchanged. README is shorter than on main, with a plain cache note and changelog link. New CSV field definitions are in the changelog.

Validation:

  • Regressions reproduced the failures before the fixes. The focused suite now passes 128 tests after the pre-release cleanup, including PNG recovery without rescoring, stat/content cache policies, explicit coordinate preference, AF3 error distinctions and pDockQ2 compatibility.
  • Cache identity benchmark on 32 fixture files (59,719,098 input bytes), three warm measurements: old median 149.5 ms, default metadata mode 34.5 ms with zero input payload bytes read, explicit content mode 128.9 ms. Software and CSV checksums remain enabled.
  • Rendered PDF review: all 36 interface rows and all 42 aggregate ranking rows remain visible across repeated-header pages.
  • Final broader integration suite after cleanup: 182 passed, 6 skipped on Python 3.12 in 11m15s. CI runs all six slow CCP4/SC references across Python 3.10–3.13. The 1.4.4 wheel and source archive also pass strict metadata checks; an isolated wheel installation passes CLI scoring and warm-cache reuse.

Pre-release cleanup: shared parser options, typed cache settings, separate scoring/publication steps, shared PAE and manifest helpers, unified report labels/page sizes, and shared test fixtures. Version is 1.4.4. Input discovery remains conservative so newly available preferred structures invalidate the cache. The final broader suite and all pre-merge CI checks passed, including Python 3.10–3.13 and Docker. Released as 1.4.4.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@DimaMolod
DimaMolod merged commit cab1794 into main Oct 1, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant