feat: schedule local MMseqs2 features for AlphaFold 2 runs - #55
Merged
Merged
Conversation
The workflow's msa_cache_key decides whether the core MSA stage runs at all: a valid shard summary short-circuits the core's own provenance check. So a search setting missing from the key is worse than a slow cache miss -- the job that would notice never starts, and old alignments are reused under new settings. That was the case for use_gpu. The core records cpu/gpu in bundle provenance, the key did not, so flipping it kept the previous alignments silently. Now keyed. num_iterations (iterative profile search) was unreachable: msa_cli_arguments never emitted --mmseqs_num_iterations and config had no key. Now configurable, passed through, and keyed. db_load_mode was documented in config.yaml and read by nothing. Now passed through as --mmseqs_db_load_mode -- and deliberately NOT keyed, because it changes memory behaviour and never the alignment; keying it would re-search hours of work to reproduce byte-identical output. Each new setting is recorded in the key and on the command line only when it differs from its default, so an unchanged configuration keeps its cache namespace and issues the same command it did before. db_load_mode needs an image carrying the core flag of the same name; left unset, nothing changes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The local MMseqs2 stage refused any data pipeline but AlphaFold 3. Its search is backend-neutral, and AlphaPulldown now finalizes MSA bundles into AlphaFold 2 pickles as well, so an AlphaFold 2 run can use it: same bounded search shards, same separate CPU finalization jobs, output named by the existing backend-aware FEATURE_OUTPUT_PATTERN. The adapter records which backend it serves, and the finalize rule is told it through --data_pipeline instead of hard-coding alphafold3. For AlphaFold 2 the rule forwards the template-stack arguments from create_feature_arguments -- --use_hhsearch, the PDB seqres/PDB70/mmCIF paths, the search binaries -- so the finalizer builds the template stack native features would have used. The native MSA arguments (jackhmmer, HHblits, their databases) are deliberately not forwarded: they describe a search this stage replaces. The feature cache key gains the backend and the forwarded template arguments, because hmmsearch and hhsearch find different templates -- but only for AlphaFold 2, so every AlphaFold 3 feature cache keeps the namespace it had. A test pins that equality. A real Snakemake dry run of an AlphaFold 2 configuration asserts the DAG schedules the shared search shard and a finalization carrying --data_pipeline=alphafold2 into a .pkl output. AlphaFold 2 finalization resources are not measured yet; the rule borrows the AlphaFold 3 constants and the OOM retry escalation until they are. Running it needs a prediction image built from an AlphaPulldown carrying the AlphaFold 2 finalizer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Say how to enable it for AlphaFold 2, that its recipe is reduced_dbs, which create_feature_arguments reach finalization and which deliberately do not, and what is still missing: measured AlphaFold 2 finalization resources, a prediction image carrying the finalizer, and a benchmark against native features. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A run starting from UniProt accessions failed every local MMseqs2 shard with a KeyError before any search ran, on every retry. Found by the first real end-to-end run of an AlphaFold 2 configuration; AlphaFold 3 is affected alike. Snakemake parses the workflow again inside every SLURM job, and a job has to find the shard it was submitted for. Three things a parse read could change that answer between the head node's parse and a job's: - lengths learned from FASTAs the run itself downloads: the head node plans before download_uniprot runs, so each protein is planned alone; the job's parse sees the lengths, groups them, and its shard id is gone - the requested protein set, which excludes proteins with precomputed features in a shared store other runs write to - for repairs, the MSA cache directory's mtime, which repair_shard_identifier hashed as a nonce -- and every running shard changes. Reproduced: one other shard publishing moved the repair id from ...e99f94eaca90 to ...2345589f69b4. Planning is now one module, schedule_msa_shards, over an append-only shard registry beside the MSA cache. A protein once planned keeps its shard; only proteins no usable registered shard covers are planned, and appended. Every registered shard's job id keeps resolving on every later parse, so a running job always finds itself. A grown sample sheet re-plans nothing -- no GPU job per existing shard -- and a protein leaving mid-run orphans nothing. A partly requested shard that never completed is left alone and its requested proteins planned afresh; one that completed keeps serving them. The repair id drops the nonce: each completed repair leaves a summary that is part of the next repair's state, so successive repairs still get distinct ids. Candidate summaries are now ordered totally; ties within one timestamp tick had been broken by hash-seeded set order, so two parses could pick different valid summaries -- a flake the new tests caught. Lengths for the first plan now come from the length filter's resolver whether or not filtering is on, so a run from accessions groups its proteins instead of scanning the databases once per protein. The Snakefile's planning globals collapse to one schedule the rules ask for a shard and a completion target. Tests move to the schedule's interface: inputs arriving between parses, growth, a protein leaving mid-run, a repair surviving another shard publishing, successive repairs, partly requested shards, an unusable registry. The real Snakemake regression test still reproduces the original failure end to end. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…tion The finalization stage runs AlphaFold 3's template search, which honours --pdb_seqres_database_path, --template_mmcif_dir and the hmmsearch/hmmbuild binaries exactly as the native run does. The rule forwarded none of them for AlphaFold 3, so an override silently searched the default under --data_dir. Each backend now has its own template argument set. The feature cache key records them only when set, so an AlphaFold 3 run overriding nothing keeps its namespace while one that does override gets features built from the templates it asked for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The GPU search reads every padded database in full, ~375 GB for the four protein ones. On NFS a cold first attempt is bound by that read -- measured ~150 MB/s with the GPU idle, about an hour per shard -- against minutes once the node's page cache holds them. Say so, and what to do when first attempts time out. Replace "not measured" and "not benchmarked" for AlphaFold 2 with the measured finalization cost and the DockQ comparison against native features. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…zation AlphaFold 2 finalization borrowed AlphaFold 3's constants (2 GB + 1 MB per residue, 30 min), measured on AlphaFold 3's light template search. AlphaFold 2's template featurization is far heavier: over 56 benchmark chains the median was ~1 GB / 2 min, but the peak 18.8 GB and 87 min, and it follows which structures the template hits come from, not length or MSA depth -- an 89-residue chain was the slowest. The first attempt ran out of memory, and a 1.1x escalation per retry never caught up. The finalization defaults are now per backend: AlphaFold 2 gets a flat 16 GB base, which covers that peak at the default safety factor for the shortest chains, and 60 min, which finishes ~95% of chains first time and the rest on the retry. AlphaFold 3 keeps its defaults, and an explicit setting overrides either. The config and README described these defaults one commit early. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DimaMolod
marked this pull request as ready for review
September 11, 2026 19:25
Share selected template database identities between finalization arguments and the feature cache key. Require PDB70 only for HHsearch, retain seqres for other searches, and preserve reusable MSA bundles when only templates change.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Local MMseqs2 features for AlphaFold 2 runs
mmseqs2_featuresrefused every data pipeline except AlphaFold 3. AlphaPulldown (KosinskiLab/AlphaPulldown#642) now finalizes the same MSA bundles into AlphaFold 2 pickles, so this PR schedules the stage for--data_pipeline: alphafold2too. It uses the same bounded GPU/CPU search shards and separate CPU finalization jobs, and writes.pkloutput through the existing backend-awareFEATURE_OUTPUT_PATTERN. The finalize rule gets its backend from--data_pipelinerather than a hard-codedalphafold3. For AlphaFold 2 it forwards the template-stack arguments fromcreate_feature_arguments(--use_hhsearch, the PDB seqres/PDB70/mmCIF paths, the search binaries), but never the native MSA arguments this stage replaces.Running it needs a prediction image built from an AlphaPulldown that includes KosinskiLab/AlphaPulldown#642.
Fixes that affect AlphaFold 3 runs too
KeyErroron a run that starts from UniProt accessions, on every retry and before any search ran. The first real AlphaFold 2 end-to-end run found it. Snakemake re-parses the workflow in every SLURM job, and the shard plan depended on things that change between parses: lengths learned from FASTAs the run itself downloads, a shared feature store other runs write to, and (for repairs) the MSA cache directory's mtime, which every running shard changes. Planning is now one module,schedule_msa_shards, over an append-only shard registry beside the MSA cache. A protein once planned keeps its shard, only uncovered proteins are planned and appended, and every registered job id keeps resolving. A grown sample sheet re-plans nothing. A real-Snakemake regression test reproduces the original failure.use_gpusilently reused the previous alignments. The workflow's MSA cache key decides whether the core stage runs at all, and it did not include search mode. It does now.num_iterationswas unreachable and is now configurable and keyed.db_load_modewas documented but read by nothing; it is now passed through, and deliberately left out of the key because it is memory-only.--pdb_seqres_database_path,--template_mmcif_dirand the hmmsearch/hmmbuild binaries, exactly as the native run does. The rule forwarded none of them, so an override silently searched the defaults under--data_dir. A run that overrides nothing keeps its feature cache namespace; a run that does override gets features built from the templates it asked for.Template cache identity
For AF2 with
--use_hhsearch: true,mmseqs2_features.template_database_ids.pdb70is required and forwarded as--template_pdb70_database_id. Other searches requirepdb_seqres; all requiremmcif. The command and feature key share one database-selection helper. Update the selected immutable ID whenever its database is rebuilt, even at the same path: only finalized features are invalidated, and the MSA bundles stay reusable. Unused database identities do not move the cache. This needs the companion core update in KosinskiLab/AlphaPulldown#642.Tests
137 tests pass at
96e92bc(2026-09-14, SLURM61831784). Unit coverage reaches all eight new executable functions in the workflow module. PDB70 regressions verify command forwarding, missing-ID rejection, feature invalidation, and MSA reuse. They include the real-Snakemake regression test for the shardKeyError, an AlphaFold 2 dry-run DAG test, and schedule tests for inputs arriving between parses, growth, a protein leaving mid-run, repairs surviving other shards publishing, successive repairs, and an unusable registry.End to end on the cluster: the historical 2026-09-11 run used workflow
dd5b1e7and core71eef1c4, completing all 35 steps. It started from UniProt accessions and local FASTAs, put five proteins through one search shard (54 min on a MIG slice, first attempt), five finalizations (29–51 s each), nine AF2-multimer folds and their analysis. The folds cover a heterodimer, a homodimer, a chopped chain and a monomer, plus three edge cases on the real databases, alone and paired: thioredoxin (a deep MSA), a random 90-residue orphan with no homologs or templates, and a 26-residue peptide. A second invocation found nothing to do. This run used hmmsearch and predates the final resource-default commit and current PDB70 identity fix; the current unit and core integration/real-binary gates validate those revisions. The earlier run with the old shard planning is what surfaced theKeyErrorfixed here.Benchmark
Numbers are in KosinskiLab/AlphaPulldown#642. For AlphaFold 2, top-ranked DockQ over 12 heterodimers averaged 0.56 with local MMseqs2 features against 0.59 with native
reduced_dbsones, and 9 of 12 interfaces were acceptable either way.Also from the measurements
Follow-up
identifieris never re-searched. Keying it at parse time would break the every-parse-agrees invariant for databases the submit host cannot see, so it needs a design decision.🤖 Generated with Claude Code