Skip to content

perf(lindblad): store the expm generator cache as blocked flat CSC - #229

Merged
david-pl merged 2 commits into
split/5-kossakowskifrom
perf/lindblad-mf-cols
Sep 28, 2026
Merged

david-pl merged 2 commits into
split/5-kossakowskifrom
perf/lindblad-mf-cols

Conversation

@AlexSchuckert

Copy link
Copy Markdown
Collaborator

Stacked on #222 (split/5-kossakowski).

Summary

build_mf_cols / build_orbit_rep_cols returned one Vec per column, each reserved at the full L* output count although only in-basis entries are kept. The expm generator cache is now a BlockCsc: blocks of up to 4096 columns, each an exactly sized offsets/rows/vals triple (structure of arrays, 12 B per real entry instead of 16 B). CscOp::dot keeps the same per-thread column partition, so results are bit-identical.

Single commit, crates/ppvm-lindblad/src/mf_expm.rs only.

Measurements

pc_step on the two-leg XY ladder with a local probe (not part of this PR): L = 41 rungs (N = 82), max_basis = 2^18, admit_basis = 3·2^18, dt = 0.1, 25 steps, 4 threads, drop_tol = 0. Apple M4 MacBook Air (fanless, so wall times are only comparable within a rep). Max RSS from /usr/bin/time -l; live-heap peak from a counting global allocator.

mimalloc, 3 interleaved reps wall (s) max RSS (MiB) live-heap peak (MiB)
base (#222) 53.7 / 53.5 / 62.5 484 / 484 / 483 441.7
this PR 46.2 / 48.9 / 60.1 501 / 491 / 470 275.2
  • The live-heap peak drops by 38% and wall time by 4–14%, with each expm call about 20–30% faster.
  • On macOS + mimalloc the heap saving shows up only weakly in the process footprint: max RSS moves by −3% to +4%, and peak footprint (/usr/bin/time -l) drops by about 3–13%. With the macOS system allocator, RSS is not reproducible (768–1237 MiB for the same binary), so the effect cannot be resolved there.
  • This has not been measured on Linux/glibc, where fragmentation from the per-column allocations of the old layout may matter more.
  • The final coefficients are bit-identical to the base in every run.

Tests: cargo test -p ppvm-lindblad passes (8/8) with this change stacked together with the skip-empty-hop2 PR; not yet run on this branch alone (CI will).

🤖 Generated with Claude Code

build_mf_cols / build_orbit_rep_cols returned one Vec per column, each
reserved at the full L* output count although only in-basis entries are
kept. On large bases this over-reserves the cache severalfold and leaves
one small allocation per basis string per expm call for the allocator to
retain.

The cache is now a BlockCsc: blocks of up to 4096 columns, each an exactly
sized offsets/rows/vals triple (structure of arrays, 12 B per real entry
instead of 16 B). CscOp::dot keeps the same per-thread column partition,
so results are bit-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-09-28 13:44 UTC

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@david-pl
david-pl merged commit 78a4b6d into split/5-kossakowski Sep 28, 2026
13 checks passed
@david-pl
david-pl deleted the perf/lindblad-mf-cols branch September 28, 2026 13:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants