Skip to content

Fix solvation energy, Boltz-2 token alignment, report and AF3 parser issues - #32

Merged
DimaMolod merged 5 commits into
30-alphajudge-pae-handling-for-af3x-outputs-cross-linker-tokensfrom
30-fixes
Sep 28, 2026
Merged

DimaMolod merged 5 commits into
30-alphajudge-pae-handling-for-af3x-outputs-cross-linker-tokensfrom
30-fixes

Conversation

@DimaMolod

Copy link
Copy Markdown
Collaborator

These are fixes for the problems found while reviewing #31, plus a check of parsers/af3.py. There is one commit per fix, and each behaviour change has a test that fails on the previous code.

This PR is stacked on #31, so it targets 30-simplify. GitHub will retarget it to the #30 branch once #31 is merged.

Fixes

  1. Solvation energy now matches PISA (a9774d6).
    • Cause: ProSurf skipped the per-atom SAS calculation for atoms with no partner atom nearby. Charged groups (Arg NE/NH1/NH2, His ND1/NE2, carboxylate O1/O2) are typed by comparing SAS between sibling atoms of the same residue, so a residue that straddles the cutoff could be typed charged in one state and neutral in the other.
    • Fix: SAS is now computed for every atom of any residue near the partner.
    • Effect: on the six CCP4 references, int_solv_en was off by +0.02 to +1.09 kcal/mol and now agrees to within 1e-4. The test tolerance drops from abs=1.2 to 1e-3. Area, H-bonds and salt bridges are unchanged.
  2. Boltz-2 PAE, pLDDT and pair ipTM are aligned by token and chain (69c9f1b).
    • Cause: Boltz gives one token per ligand atom. The parser kept the leading N×N block and indexed pair_chains_iptm by position among the scored chains. That is only right when every ligand chain comes after the proteins, which is the same class of bug as AlphaJudge PAE handling for AF3x outputs (cross-linker tokens) #30.
    • Fix: the token layout is rebuilt from the structure, and pair ipTM is looked up by chain ID in structure order. A mis-sized matrix is dropped with a warning.
    • Effect: 6OGE (ligands last) is unchanged. test_boltz.py covers the ligand-first and ligand-between cases.
  3. Report fixes (15cf721).
    • Empty model_used: it crashed the aggregate report through the invalid ** glob on Python 3.10–3.12. Model names are now matched literally in PAE image lookup.
    • Feature count: the table caption said 10 metascore features; it now gives the real count, 11.
    • Median: the aggregate cover used the upper-middle value for an even number of interfaces; it now uses the true median.
    • Appendix numbering: headings were A.<model> cut to eight characters; they are now numbered A.1, A.2, ….
  4. ProSurf interface cache is keyed on the atoms it reads (92ed184).
    • Cause: the cache key was each residue's full_id plus its first-atom coordinate, and every structure is loaded as "complex". A model whose side chains moved 50 Å still returned the cached 2717.28 Ų.
    • Fix: the key is now all atom identities and coordinates, and the cache keeps four results instead of up to 33 structures.
  5. AF3 parser (a996347).
    • Ranking order: a missing or unparsable ranking_score gave a NaN sort key, which scrambles the order, so "best" could be the wrong model. Unscored rows now rank last.
    • Flat-layout lookup: seed-1_sample-1 matched and preferred seed-1_sample-10's structure.
    • Empty ranking file: in best mode this raised IndexError; it now writes the explained empty CSV.
    • Cleanup: removed the unused model_dir and the duplicate _find_existing helper.

Verification

  • pytest test/ with ALPHAJUDGE_RUN_SLOW_SC_REFERENCE=1: 103/103 pass (92 existing + 11 new regression tests).
  • Golden run over all bundled AF2/AF3/AF3x/Boltz-2 runs, compared to Simplify scoring, report and biophysics code without changing outputs #31: across 28 CSVs only interface_solv_en (99 values, max |Δ| 1.085 kcal/mol) and interface_meta_score (99 values, max |Δ| 0.0012) change; every other column is identical. All 32 PAE PNGs are byte-identical; 10 of 11 PDFs change, as expected from the metascore and report-text fixes.

Worth knowing

  • Calibration: the frozen interface_solv_en percentile ladder was built with the previous solvation values. The correction is small next to its decile spacing (roughly 3–5 kcal/mol in the lower half). A recalibration run on the benchmark would make the ladder exact.
  • Benchmark numbers: numbers in the benchmark paper that come from interface_solv_en or the metascore were computed with the old solvation code.

🤖 Generated with Claude Code

DimaMolod and others added 5 commits September 26, 2026 11:35
ProSurf skipped the per-atom accessible-area calculation for atoms with
no partner-chain atom in reach, leaving their SAS at 0 even when a
sibling atom of the same residue was at the interface. The solvation
energy types charged groups (Arg NE/NH1/NH2, His ND1/NE2, carboxylate
O1/O2) by comparing SAS between those siblings, so a residue straddling
the cutoff could be typed charged in one state and neutral in the other.

SAS is now computed for every atom of any residue with an atom near the
other side, as PISA does for the atoms that matter here. Residues wholly
away from the interface still keep SAS 0; their terms are identical in
the free and bound states and cancel.

On the six CCP4 references int_solv_en moves from +0.02..+1.09 kcal/mol
off PISA to within 1e-4, so the test tolerance drops from 1.2 to 1e-3.
Buried area, interface residues, H-bonds and salt bridges are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Boltz-2 gives each ligand atom its own token, and the parser kept the
leading N x N block of the PAE (and the first N pLDDT values) and read
pair_chains_iptm by position among the scored chains. That is only right
when every ligand chain follows the proteins; a ligand chain first or in
between shifted every PAE entry and could return another pair's ipTM,
the same class of error #30 fixed for AF3x.

The token layout is rebuilt from the structure (one token per standard
residue, one per atom for ligands and modified residues, whose scored
token is the CA/C1' atom's as for AF3), and PAE/pLDDT are selected by it
when the sizes match. pair_chains_iptm is recorded against the
structure's chain order and dropped with a warning when its size does
not match. The centre-atom token rule moves to BaseParser so AF3 and
Boltz-2 share it.

The bundled 6OGE run (crosslinker chains last) is unchanged; the new
test_boltz.py covers ligand-first, ligand-between and mis-sized matrices
and fails on the previous code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- _find_pae_png built "*<model>*" globs from the raw model name: an empty
  name made the invalid pattern "**PAE*plot*.png" (ValueError on Python
  3.10-3.12, crashing generate_aggregate_report), and characters such as
  "[" were read as glob syntax. The name is now escaped, and without one
  only the ranked_0 PAE plot is considered.
- The per-interface caption said "10 metascore features"; it now counts
  META_SCORE_FEATURES (11).
- The aggregate cover's median took the upper middle value for an even
  number of interfaces; it now uses statistics.median.
- Appendix pages were numbered "A.<model name>", which the heading cuts
  to eight characters ("A.model…"); they are now A.1, A.2, ... with the
  model name kept in the title.

Each has a test that fails on the previous code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cache shared by interface area, H-bonds, salt bridges and solvation
energy was keyed on each residue's full_id and first-atom coordinate.
Every structure is parsed under the id "complex", so another model of
the same chain pair whose residues start at the same N atoms (e.g. a
side-chain repack) hit the stale result: a model with the second chain's
side chains moved 50 A away reported the original 2717.28 A^2.

The key now covers every residue id and name, atom name and element, and
all coordinates. The cache keeps the four most recent results rather
than up to 33 (which also held their structures in memory); one
interface's four scores are computed back to back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- _read_csv_order sorted rows on float(ranking_score) with NaN for a
  missing or unparsable score, and NaN keys leave Python's sort order
  arbitrary, so the "best" model could be wrong. Unscored rows now rank
  last and the rest keep their order.
- The flat-layout structure lookup globbed "*seed-1_sample-1*", which
  also matches "..._sample-10_model.cif" and sorts it first. Glob hits
  must now carry the model name not followed by another digit.
- process() indexed run.order[0] in best mode and raised IndexError when
  the ranking file listed no models; it now takes order[:1] and writes
  the explained empty CSV, as "all" mode already did.
- Drop the unused model_dir in load_model and the private _find_existing
  in favour of BaseParser._first_existing. Tidy the mid-file imports in
  test_parsers_and_runner.py.

Each behaviour change has a test that fails on the previous code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Base automatically changed from 30-simplify to 30-alphajudge-pae-handling-for-af3x-outputs-cross-linker-tokens September 28, 2026 12:18
@DimaMolod
DimaMolod marked this pull request as ready for review September 28, 2026 12:18
@DimaMolod
DimaMolod merged commit 1d07261 into 30-alphajudge-pae-handling-for-af3x-outputs-cross-linker-tokens Sep 28, 2026
12 checks passed
@DimaMolod
DimaMolod deleted the 30-fixes branch September 28, 2026 12:20

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a9963477e3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

def rank(row: dict) -> float:
# A missing score ranks last; a NaN sort key would scramble the order.
value = score(row)
return -np.inf if np.isnan(value) else value

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Demote infinite ranking scores with other invalid values

When ranking_scores.csv contains inf (for example from an overflowed upstream score), float() accepts it and this key ranks it above every finite model, so --models_to_analyse best selects an invalid model. This also conflicts with the subsequent np.isfinite check, which deliberately omits that model from scores; treat all non-finite values as unscored when constructing the sort key.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant