Skip to content

Local MMseqs2: a database rebuilt in place under the same identifier is never re-searched #56

Description

@DimaMolod

A local MMseqs2 database rebuilt or re-copied in place, under the same identifier, is never searched again by the workflow. Completed shards keep serving MSAs from the old build, while proteins added later are searched against the new one, so the cache mixes the two.

Why

AlphaPulldown stamps each database's .index file size into bundle provenance (feature_batch.py, _cache_signature → index_size). This is a cheap content witness, because identifier is operator-supplied and cannot detect a rebuild, truncation or half-finished copy. The core would re-search on a mismatch, but it never runs:

  • LocalMmseqsFeatureConfig.msa_cache_key records each database's identifier and max_sequences, not its content
  • validate_shard_summary checks the bundles' digests and sizes, not the databases they came from

So a completed shard stays complete, and the job whose cache check would have caught the change is never scheduled. This is the same failure mode use_gpu had before it was keyed.

Why it isn't a one-line fix

The key names the MSA cache directory, a path that every parse of a run must agree on. Snakemake re-parses inside each SLURM job, and the new shard registry exists to keep those parses in agreement. Statting {path}.index at parse time gives different answers whenever the submit host and the compute nodes see the databases differently. The obvious case is databases staged on node-local SSD, where the submit host reads 0 and a compute node reads the real size. That is the same class of cross-parse failure as the shard KeyError fixed in #55.

Options

  1. Stat at parse time and fail loudly when a configured database is unreadable. Simple, but it rules out databases the submit host cannot see.
  2. Record the witness in the shard registry when a shard is first planned, and have the search job (which can see the databases) compare it. A mismatch then becomes a repair rather than a silent reuse.
  3. Document it instead: a rebuilt database must get a new identifier.

Until then

Give a rebuilt or re-copied database a new identifier. That moves the MSA cache key, and everything is re-searched.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions