Conversation
AF3x models each crosslink as a ligand, and AF3 tokenizes ligands and ions per atom, so confidences.json carries more PAE tokens than AlphaJudge has scored residues. When the shapes differed, _normalize_pae_af3 filled every residue pair of a chain pair with that block's minimum PAE (#30). On a 4G3Y B-C prediction with one DSSO crosslink the whole B->C block read 5.8 A against a true mean of 21.1 A, inflating ipSAE, LIS and pDockQ2. PAE now goes through the same token-to-residue alignment as the contact probabilities: tokens are matched by chain and residue identifiers, and both directions of each residue pair are kept. The alignment no longer falls back to positional assignment. Supplied residue identifiers are authoritative, and without them each retained chain must match by ID with one token per residue. A PAE that cannot be mapped now raises instead of being substituted, so summary-only runs (chain_pair_pae_min), wrong-shape matrices without token identifiers (previously filled with 100 A) and ambiguous tokens skip the model with an explicit error. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Interface.iptm_chainpair indexed chain_pair_iptm by each chain's position among the structure's scored chains, but AF3 orders that matrix over every chain, ligands included. AF3x names each crosslinker ligand with the first unused chain letter, so a crosslinker can precede or sit between the proteins: on the 4G3Y B-C prediction (DSSO as chain A) the B-C interface reported the DSSO-B ipTM of 0.78 instead of 0.25. The parser now records the matrix's chain order from token_chain_ids, and the interface looks both chains up by ID. A chain_pair_iptm whose size does not match that order is dropped with a warning rather than indexed, and a non-finite pair value is treated as missing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
AF3's pTM, ipTM and ranking score are computed over every token, including the ligands, ions and AF3x crosslinkers that AlphaJudge excludes from interface scoring, so they do not describe the scored residues alone. The AF3 parser now labels its global confidences. global_confidence_scope is includes_excluded_tokens when the PAE had more tokens than scored residues, scored_residues otherwise, and unknown for the other parsers; the runner writes it to the score table together with iptm_scope (chain_pair or global). For includes_excluded_tokens rows, feature_is_comparable leaves confidence_score, and ipTM unless it is the chain-pair value, out of interface_meta_score and the report percentiles. The raw values stay in the table, the report page says why they are excluded, and a stale precomputed metascore is not reused for such rows. This also applies to plain AF3 predictions containing ligands or ions, whose metascore therefore changes relative to 1.4.2. Cached per-run CSVs without the two new columns are recomputed once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The other AF3x tests build their tokens by hand. This adds one real AF3x output so the whole path is checked against the model's own matrices: a 4G3Y B-C prediction with one DSSO crosslink (seed 30, one recycle, one sample) from the AF3x commit recorded in provenance.json. confidences.json is xz-compressed, and the AF3 output terms of use are included. The test extracts the protein block from the raw token matrices independently of the parser and compares PAE, contact probabilities, pair ipTM and the per-interface scores. It runs on the output as generated, with the crosslinker chain first, and on a copy whose tokens and chain summary are permuted consistently. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
AF3 gives a standard residue one token centred on CA (C1' for nucleic acids) but splits a modified residue such as SEP into one token per atom, all with the residue's chain and residue ID. AlphaJudge scores modified amino acids, since is_pae_token_residue keeps any residue with a CA, so rejecting duplicate IDs skipped every AF3 model that contained a PTM. A scored residue with several tokens now takes the token of its own CA or C1' atom; AF3 emits those tokens in the residue's atom order (tokenizer in alphafold3/model/features.py). The mapping is used only when the tokens and the residue's atoms correspond one to one, and anything else is still rejected as ambiguous. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PAE and contact probabilities each ran their own token-to-residue alignment on the same identifiers, through two near-identical functions that copied the matrices block by block, and load_model re-read the raw PAE to tell whether any tokens had been excluded. _residue_tokens_af3 now works out once which token carries each scored residue, and _ResidueTokens.select takes the residue block of any token matrix with one fancy index. The PAE step builds the map where it first needed alignment, so errors keep their order and messages, and returns it for the contact probabilities and the global-confidence scope. The rules are unchanged: supplied residue identifiers are authoritative, chain-only matching needs one token per residue, and matrices without token metadata must already be in residue order. Outputs are unchanged. On 1,000 generated AF3/AF3x runs (ligands, per-atom modified residues, nucleic acids, shuffled tokens, valid and broken metadata and matrices) the aligned matrices, errors, log messages and per-interface CSVs are identical to the previous commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Interface.iptm_chainpair carried all of the chain_pair_iptm handling: choosing the chain order (the parser's recorded order, else the scored chains), finding both chains, and reading an entry that may be missing or sit in a ragged matrix. That belongs with the matrix it reads. Confidence.pair_iptm(chain_a, chain_b, default_chain_ids) now does it by the same rules, and Interface passes the scored chains as the default order for sources that do not record one (Boltz-2, or a Confidence built directly). A unit test covers the recorded order, the default order, an unknown chain, a ragged matrix and a missing value. Outputs are unchanged: the 1,000 generated AF3/AF3x runs give the same per-interface ipTM and CSVs as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The scope strings were spelled out wherever they were produced or read: "includes_excluded_tokens" and "scored_residues" in the AF3 parser, metascore and report, "chain_pair" and "global" in the runner and metascore. A typo in any reader would quietly turn the metascore exclusion off, because an unrecognised scope counts as comparable. confidence.py now defines them next to the field they describe, and the parser, runner, metascore and report import them. The CSV values are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A pass over the package for dead code, duplicated logic and needless indirection. Scores, CSVs, report PDFs and PAE PNGs are byte-identical to the previous commit on every bundled AF2, AF3, AF3x and Boltz-2 run (all models, per-run and aggregate reports, cache reuse, non-default thresholds), the raw Connolly dot surfaces are bitwise equal, and the full test suite passes including the slow CCP4 SC reference checks. report.py (2031 -> ~1370 lines): drop the page-total/footer plumbing that nothing read, the never-drawn logo and info icon, an unused row grouper and slider-marker branch. Repeated axes setup and text boilerplate go through small _text_axes/_label_axes/_kv_rows helpers. runner.py: one worker partial serves the serial and process-pool paths, the cache/recompute branches share the report and aggregation tail, and each CSV row is built by _interface_row. Unused matplotlib/numpy imports and a deferred import of an already-imported module are gone. Interface/Complex: one pair-index generator replaces four copies of the residue-to-PAE lookup; mpDockQ's global contact count is the sum of the per-interface counts (the same <= r^2 test Biopython's KD-tree applies) instead of a second neighbour search; a per-pair atom cache that could never hit is removed. Interface.label and Complex.chain_boundaries replace private-attribute access from the runner. Parsers: magic-byte decompression is shared by JSON and pickle loading, path lookups share _first_existing, and ParserManager keeps only register/pick. Biophysics: bonds.py uses one KD-tree contact generator for H-bonds, salt bridges and disulfides; prosurf.py computes each side's area through one coverage helper; connolly.py loses an unused trim(), _disptl and a never-read neighbour list, and shares the probe-torus geometry between its two users; sc.py and solvation.py drop redundant bookkeeping. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ProSurf skipped the per-atom accessible-area calculation for atoms with no partner-chain atom in reach, leaving their SAS at 0 even when a sibling atom of the same residue was at the interface. The solvation energy types charged groups (Arg NE/NH1/NH2, His ND1/NE2, carboxylate O1/O2) by comparing SAS between those siblings, so a residue straddling the cutoff could be typed charged in one state and neutral in the other. SAS is now computed for every atom of any residue with an atom near the other side, as PISA does for the atoms that matter here. Residues wholly away from the interface still keep SAS 0; their terms are identical in the free and bound states and cancel. On the six CCP4 references int_solv_en moves from +0.02..+1.09 kcal/mol off PISA to within 1e-4, so the test tolerance drops from 1.2 to 1e-3. Buried area, interface residues, H-bonds and salt bridges are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Boltz-2 gives each ligand atom its own token, and the parser kept the leading N x N block of the PAE (and the first N pLDDT values) and read pair_chains_iptm by position among the scored chains. That is only right when every ligand chain follows the proteins; a ligand chain first or in between shifted every PAE entry and could return another pair's ipTM, the same class of error #30 fixed for AF3x. The token layout is rebuilt from the structure (one token per standard residue, one per atom for ligands and modified residues, whose scored token is the CA/C1' atom's as for AF3), and PAE/pLDDT are selected by it when the sizes match. pair_chains_iptm is recorded against the structure's chain order and dropped with a warning when its size does not match. The centre-atom token rule moves to BaseParser so AF3 and Boltz-2 share it. The bundled 6OGE run (crosslinker chains last) is unchanged; the new test_boltz.py covers ligand-first, ligand-between and mis-sized matrices and fails on the previous code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- _find_pae_png built "*<model>*" globs from the raw model name: an empty
name made the invalid pattern "**PAE*plot*.png" (ValueError on Python
3.10-3.12, crashing generate_aggregate_report), and characters such as
"[" were read as glob syntax. The name is now escaped, and without one
only the ranked_0 PAE plot is considered.
- The per-interface caption said "10 metascore features"; it now counts
META_SCORE_FEATURES (11).
- The aggregate cover's median took the upper middle value for an even
number of interfaces; it now uses statistics.median.
- Appendix pages were numbered "A.<model name>", which the heading cuts
to eight characters ("A.model…"); they are now A.1, A.2, ... with the
model name kept in the title.
Each has a test that fails on the previous code.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cache shared by interface area, H-bonds, salt bridges and solvation energy was keyed on each residue's full_id and first-atom coordinate. Every structure is parsed under the id "complex", so another model of the same chain pair whose residues start at the same N atoms (e.g. a side-chain repack) hit the stale result: a model with the second chain's side chains moved 50 A away reported the original 2717.28 A^2. The key now covers every residue id and name, atom name and element, and all coordinates. The cache keeps the four most recent results rather than up to 33 (which also held their structures in memory); one interface's four scores are computed back to back. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- _read_csv_order sorted rows on float(ranking_score) with NaN for a missing or unparsable score, and NaN keys leave Python's sort order arbitrary, so the "best" model could be wrong. Unscored rows now rank last and the rest keep their order. - The flat-layout structure lookup globbed "*seed-1_sample-1*", which also matches "..._sample-10_model.cif" and sorts it first. Glob hits must now carry the model name not followed by another digit. - process() indexed run.order[0] in best mode and raised IndexError when the ranking file listed no models; it now takes order[:1] and writes the explained empty CSV, as "all" mode already did. - Drop the unused model_dir in load_model and the private _find_existing in favour of BaseParser._first_existing. Tidy the mid-file imports in test_parsers_and_runner.py. Each behaviour change has a test that fails on the previous code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Simplify scoring, report and biophysics code without changing outputs
Fix #30 more accurate solvation energy, safier Boltz-2 token alignment
DimaMolod
deleted the
30-alphajudge-pae-handling-for-af3x-outputs-cross-linker-tokens
branch
September 28, 2026 12:29
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.