A local MMseqs2 database rebuilt or re-copied in place, under the same identifier, is never searched again by the workflow. Completed shards keep serving MSAs from the old build, while proteins added later are searched against the new one, so the cache mixes the two.
Why
AlphaPulldown stamps each database's .index file size into bundle provenance (feature_batch.py, _cache_signature → index_size). This is a cheap content witness, because identifier is operator-supplied and cannot detect a rebuild, truncation or half-finished copy. The core would re-search on a mismatch, but it never runs:
LocalMmseqsFeatureConfig.msa_cache_key records each database's identifier and max_sequences, not its content
validate_shard_summary checks the bundles' digests and sizes, not the databases they came from
So a completed shard stays complete, and the job whose cache check would have caught the change is never scheduled. This is the same failure mode use_gpu had before it was keyed.
Why it isn't a one-line fix
The key names the MSA cache directory, a path that every parse of a run must agree on. Snakemake re-parses inside each SLURM job, and the new shard registry exists to keep those parses in agreement. Statting {path}.index at parse time gives different answers whenever the submit host and the compute nodes see the databases differently. The obvious case is databases staged on node-local SSD, where the submit host reads 0 and a compute node reads the real size. That is the same class of cross-parse failure as the shard KeyError fixed in #55.
Options
- Stat at parse time and fail loudly when a configured database is unreadable. Simple, but it rules out databases the submit host cannot see.
- Record the witness in the shard registry when a shard is first planned, and have the search job (which can see the databases) compare it. A mismatch then becomes a repair rather than a silent reuse.
- Document it instead: a rebuilt database must get a new
identifier.
Until then
Give a rebuilt or re-copied database a new identifier. That moves the MSA cache key, and everything is re-searched.
A local MMseqs2 database rebuilt or re-copied in place, under the same
identifier, is never searched again by the workflow. Completed shards keep serving MSAs from the old build, while proteins added later are searched against the new one, so the cache mixes the two.Why
AlphaPulldown stamps each database's
.indexfile size into bundle provenance (feature_batch.py,_cache_signature→index_size). This is a cheap content witness, becauseidentifieris operator-supplied and cannot detect a rebuild, truncation or half-finished copy. The core would re-search on a mismatch, but it never runs:LocalMmseqsFeatureConfig.msa_cache_keyrecords each database'sidentifierandmax_sequences, not its contentvalidate_shard_summarychecks the bundles' digests and sizes, not the databases they came fromSo a completed shard stays complete, and the job whose cache check would have caught the change is never scheduled. This is the same failure mode
use_gpuhad before it was keyed.Why it isn't a one-line fix
The key names the MSA cache directory, a path that every parse of a run must agree on. Snakemake re-parses inside each SLURM job, and the new shard registry exists to keep those parses in agreement. Statting
{path}.indexat parse time gives different answers whenever the submit host and the compute nodes see the databases differently. The obvious case is databases staged on node-local SSD, where the submit host reads 0 and a compute node reads the real size. That is the same class of cross-parse failure as the shardKeyErrorfixed in #55.Options
identifier.Until then
Give a rebuilt or re-copied database a new
identifier. That moves the MSA cache key, and everything is re-searched.