Skip to content

feat: schedule local MMseqs2 features for AlphaFold 2 runs - #55

Merged
DimaMolod merged 10 commits into
mainfrom
feat/af2-local-mmseqs
Sep 14, 2026
Merged

DimaMolod merged 10 commits into
mainfrom
feat/af2-local-mmseqs

Conversation

@DimaMolod

@DimaMolod DimaMolod commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator

Local MMseqs2 features for AlphaFold 2 runs

mmseqs2_features refused every data pipeline except AlphaFold 3. AlphaPulldown (KosinskiLab/AlphaPulldown#642) now finalizes the same MSA bundles into AlphaFold 2 pickles, so this PR schedules the stage for --data_pipeline: alphafold2 too. It uses the same bounded GPU/CPU search shards and separate CPU finalization jobs, and writes .pkl output through the existing backend-aware FEATURE_OUTPUT_PATTERN. The finalize rule gets its backend from --data_pipeline rather than a hard-coded alphafold3. For AlphaFold 2 it forwards the template-stack arguments from create_feature_arguments (--use_hhsearch, the PDB seqres/PDB70/mmCIF paths, the search binaries), but never the native MSA arguments this stage replaces.

Running it needs a prediction image built from an AlphaPulldown that includes KosinskiLab/AlphaPulldown#642.

Fixes that affect AlphaFold 3 runs too

  • Every local MMseqs2 shard failed with a KeyError on a run that starts from UniProt accessions, on every retry and before any search ran. The first real AlphaFold 2 end-to-end run found it. Snakemake re-parses the workflow in every SLURM job, and the shard plan depended on things that change between parses: lengths learned from FASTAs the run itself downloads, a shared feature store other runs write to, and (for repairs) the MSA cache directory's mtime, which every running shard changes. Planning is now one module, schedule_msa_shards, over an append-only shard registry beside the MSA cache. A protein once planned keeps its shard, only uncovered proteins are planned and appended, and every registered job id keeps resolving. A grown sample sheet re-plans nothing. A real-Snakemake regression test reproduces the original failure.
  • Switching use_gpu silently reused the previous alignments. The workflow's MSA cache key decides whether the core stage runs at all, and it did not include search mode. It does now. num_iterations was unreachable and is now configurable and keyed. db_load_mode was documented but read by nothing; it is now passed through, and deliberately left out of the key because it is memory-only.
  • AlphaFold 3 template overrides were dropped. AlphaFold 3's template search honours --pdb_seqres_database_path, --template_mmcif_dir and the hmmsearch/hmmbuild binaries, exactly as the native run does. The rule forwarded none of them, so an override silently searched the defaults under --data_dir. A run that overrides nothing keeps its feature cache namespace; a run that does override gets features built from the templates it asked for.

Template cache identity

For AF2 with --use_hhsearch: true, mmseqs2_features.template_database_ids.pdb70 is required and forwarded as --template_pdb70_database_id. Other searches require pdb_seqres; all require mmcif. The command and feature key share one database-selection helper. Update the selected immutable ID whenever its database is rebuilt, even at the same path: only finalized features are invalidated, and the MSA bundles stay reusable. Unused database identities do not move the cache. This needs the companion core update in KosinskiLab/AlphaPulldown#642.

Tests

137 tests pass at 96e92bc (2026-09-14, SLURM 61831784). Unit coverage reaches all eight new executable functions in the workflow module. PDB70 regressions verify command forwarding, missing-ID rejection, feature invalidation, and MSA reuse. They include the real-Snakemake regression test for the shard KeyError, an AlphaFold 2 dry-run DAG test, and schedule tests for inputs arriving between parses, growth, a protein leaving mid-run, repairs surviving other shards publishing, successive repairs, and an unusable registry.

End to end on the cluster: the historical 2026-09-11 run used workflow dd5b1e7 and core 71eef1c4, completing all 35 steps. It started from UniProt accessions and local FASTAs, put five proteins through one search shard (54 min on a MIG slice, first attempt), five finalizations (29–51 s each), nine AF2-multimer folds and their analysis. The folds cover a heterodimer, a homodimer, a chopped chain and a monomer, plus three edge cases on the real databases, alone and paired: thioredoxin (a deep MSA), a random 90-residue orphan with no homologs or templates, and a 26-residue peptide. A second invocation found nothing to do. This run used hmmsearch and predates the final resource-default commit and current PDB70 identity fix; the current unit and core integration/real-binary gates validate those revisions. The earlier run with the old shard planning is what surfaced the KeyError fixed here.

Benchmark

Numbers are in KosinskiLab/AlphaPulldown#642. For AlphaFold 2, top-ranked DockQ over 12 heterodimers averaged 0.56 with local MMseqs2 features against 0.59 with native reduced_dbs ones, and 9 of 12 interfaces were acceptable either way.

Also from the measurements

  • AlphaFold 2 finalization has its own resource defaults. It had borrowed AlphaFold 3's (2 GB + 1 MB per residue, 30 min, ×1.1 per retry). AlphaFold 2's template featurization peaked at 18.8 GB and 87 min over 56 chains (median ~1 GB and 2 min), and it follows which structures the templates come from, not length. The defaults are now per backend: AlphaFold 2 gets a flat 16 GB base and 60 min, while AlphaFold 3 is unchanged and explicit settings still win.
  • Cold GPU searches on network storage are documented. The search reads ~375 GB of padded databases in full, so from NFS a cold first attempt is I/O-bound: ~150 MB/s with the GPU idle, about an hour per shard, against minutes from page cache. The README says what to do when first attempts time out.

Follow-up

🤖 Generated with Claude Code

DimaMolod and others added 5 commits September 11, 2026 15:17
The workflow's msa_cache_key decides whether the core MSA stage runs at all: a
valid shard summary short-circuits the core's own provenance check. So a search
setting missing from the key is worse than a slow cache miss -- the job that
would notice never starts, and old alignments are reused under new settings.

That was the case for use_gpu. The core records cpu/gpu in bundle provenance,
the key did not, so flipping it kept the previous alignments silently. Now keyed.

num_iterations (iterative profile search) was unreachable: msa_cli_arguments
never emitted --mmseqs_num_iterations and config had no key. Now configurable,
passed through, and keyed.

db_load_mode was documented in config.yaml and read by nothing. Now passed
through as --mmseqs_db_load_mode -- and deliberately NOT keyed, because it
changes memory behaviour and never the alignment; keying it would re-search
hours of work to reproduce byte-identical output.

Each new setting is recorded in the key and on the command line only when it
differs from its default, so an unchanged configuration keeps its cache
namespace and issues the same command it did before. db_load_mode needs an
image carrying the core flag of the same name; left unset, nothing changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The local MMseqs2 stage refused any data pipeline but AlphaFold 3. Its search
is backend-neutral, and AlphaPulldown now finalizes MSA bundles into AlphaFold 2
pickles as well, so an AlphaFold 2 run can use it: same bounded search shards,
same separate CPU finalization jobs, output named by the existing backend-aware
FEATURE_OUTPUT_PATTERN.

The adapter records which backend it serves, and the finalize rule is told it
through --data_pipeline instead of hard-coding alphafold3. For AlphaFold 2 the
rule forwards the template-stack arguments from create_feature_arguments --
--use_hhsearch, the PDB seqres/PDB70/mmCIF paths, the search binaries -- so the
finalizer builds the template stack native features would have used. The native
MSA arguments (jackhmmer, HHblits, their databases) are deliberately not
forwarded: they describe a search this stage replaces.

The feature cache key gains the backend and the forwarded template arguments,
because hmmsearch and hhsearch find different templates -- but only for
AlphaFold 2, so every AlphaFold 3 feature cache keeps the namespace it had. A
test pins that equality.

A real Snakemake dry run of an AlphaFold 2 configuration asserts the DAG
schedules the shared search shard and a finalization carrying
--data_pipeline=alphafold2 into a .pkl output.

AlphaFold 2 finalization resources are not measured yet; the rule borrows the
AlphaFold 3 constants and the OOM retry escalation until they are. Running it
needs a prediction image built from an AlphaPulldown carrying the AlphaFold 2
finalizer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Say how to enable it for AlphaFold 2, that its recipe is reduced_dbs, which
create_feature_arguments reach finalization and which deliberately do not, and
what is still missing: measured AlphaFold 2 finalization resources, a prediction
image carrying the finalizer, and a benchmark against native features.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A run starting from UniProt accessions failed every local MMseqs2 shard with a
KeyError before any search ran, on every retry. Found by the first real
end-to-end run of an AlphaFold 2 configuration; AlphaFold 3 is affected alike.

Snakemake parses the workflow again inside every SLURM job, and a job has to
find the shard it was submitted for. Three things a parse read could change
that answer between the head node's parse and a job's:

- lengths learned from FASTAs the run itself downloads: the head node plans
  before download_uniprot runs, so each protein is planned alone; the job's
  parse sees the lengths, groups them, and its shard id is gone
- the requested protein set, which excludes proteins with precomputed features
  in a shared store other runs write to
- for repairs, the MSA cache directory's mtime, which repair_shard_identifier
  hashed as a nonce -- and every running shard changes. Reproduced: one other
  shard publishing moved the repair id from ...e99f94eaca90 to ...2345589f69b4.

Planning is now one module, schedule_msa_shards, over an append-only shard
registry beside the MSA cache. A protein once planned keeps its shard; only
proteins no usable registered shard covers are planned, and appended. Every
registered shard's job id keeps resolving on every later parse, so a running
job always finds itself. A grown sample sheet re-plans nothing -- no GPU job
per existing shard -- and a protein leaving mid-run orphans nothing. A partly
requested shard that never completed is left alone and its requested proteins
planned afresh; one that completed keeps serving them.

The repair id drops the nonce: each completed repair leaves a summary that is
part of the next repair's state, so successive repairs still get distinct ids.
Candidate summaries are now ordered totally; ties within one timestamp tick had
been broken by hash-seeded set order, so two parses could pick different valid
summaries -- a flake the new tests caught.

Lengths for the first plan now come from the length filter's resolver whether
or not filtering is on, so a run from accessions groups its proteins instead of
scanning the databases once per protein. The Snakefile's planning globals
collapse to one schedule the rules ask for a shard and a completion target.

Tests move to the schedule's interface: inputs arriving between parses, growth,
a protein leaving mid-run, a repair surviving another shard publishing,
successive repairs, partly requested shards, an unusable registry. The real
Snakemake regression test still reproduces the original failure end to end.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…tion

The finalization stage runs AlphaFold 3's template search, which honours
--pdb_seqres_database_path, --template_mmcif_dir and the hmmsearch/hmmbuild
binaries exactly as the native run does. The rule forwarded none of them for
AlphaFold 3, so an override silently searched the default under --data_dir.

Each backend now has its own template argument set. The feature cache key
records them only when set, so an AlphaFold 3 run overriding nothing keeps its
namespace while one that does override gets features built from the templates
it asked for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The GPU search reads every padded database in full, ~375 GB for the four
protein ones. On NFS a cold first attempt is bound by that read -- measured
~150 MB/s with the GPU idle, about an hour per shard -- against minutes once the
node's page cache holds them. Say so, and what to do when first attempts time
out. Replace "not measured" and "not benchmarked" for AlphaFold 2 with the
measured finalization cost and the DockQ comparison against native features.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…zation

AlphaFold 2 finalization borrowed AlphaFold 3's constants (2 GB + 1 MB per
residue, 30 min), measured on AlphaFold 3's light template search. AlphaFold 2's
template featurization is far heavier: over 56 benchmark chains the median was
~1 GB / 2 min, but the peak 18.8 GB and 87 min, and it follows which structures
the template hits come from, not length or MSA depth -- an 89-residue chain was
the slowest. The first attempt ran out of memory, and a 1.1x escalation per retry
never caught up.

The finalization defaults are now per backend: AlphaFold 2 gets a flat 16 GB
base, which covers that peak at the default safety factor for the shortest
chains, and 60 min, which finishes ~95% of chains first time and the rest on the
retry. AlphaFold 3 keeps its defaults, and an explicit setting overrides either.
The config and README described these defaults one commit early.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@DimaMolod
DimaMolod marked this pull request as ready for review September 11, 2026 19:25
Share selected template database identities between finalization arguments and the feature cache key. Require PDB70 only for HHsearch, retain seqres for other searches, and preserve reusable MSA bundles when only templates change.
@DimaMolod
DimaMolod merged commit dae7e5e into main Sep 14, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant