Skip to content

Fix CUDA OOM reclaim - #18

Merged
GiggleLiu merged 13 commits into
mainfrom
fix/cuda-oom-and-complex-casts
Aug 31, 2026
Merged

GiggleLiu merged 13 commits into
mainfrom
fix/cuda-oom-and-complex-casts

Conversation

@GiggleLiu

@GiggleLiu GiggleLiu commented Aug 31, 2026 •

Copy link
Copy Markdown

Fix CUDA OOM by reclaiming and retrying, and stop stranding memory in oversized slabs.

  • IoError::OutOfMemory (new variant): distinguishes a transient device OOM from BufferTooBig.
  • CUDA Command::reserve: on OOM, fence → memory_cleanup → fence → retry once.
  • SubSlices preset: top pool is now exclusive with exact-size pages instead of a
    whole-device slab (an 800 MB tensor no longer reserves a 6 GB slab).
  • Sliced pools: slices > page/4 get an exact page; on OOM, fall back to an exact page.
  • Exclusive pool: pages are exact-size and the tightest free page is reused.

Risk: the exclusive-pool changes are backend-agnostic and also affect wgpu's staging/
uniform pools (less cross-size page reuse, bounded by the ~1.4× bucket ratio).
IoError gains a variant (non-#[non_exhaustive]; downstream exhaustive matches break).

Verified: cargo test -p t4a-cubecl-runtime --lib memory_management (43 pass);
cargo check -p t4a-cubecl-cuda. The CUDA retry path itself needs a device to exercise.

Note: the branch history also carries the CUDA complex-cast commits, but those are
already on main via #17 and contribute no net diff here; squash-merge.

Maintained by tensor4all to support tenferro-rs until merged upstream.

@GiggleLiu GiggleLiu changed the title Fix CUDA OOM reclaim + CUDA complex casts Fix CUDA OOM reclaim Aug 31, 2026
@GiggleLiu

Copy link
Copy Markdown
Author

Review report (fresh-context code review + per-PR policy audit)

Gate: PASS — no secrets, dangerous files, generated artifacts, caches, or debug leftovers in the diff. Net diff vs main is the OOM-reclaim work only (7 files); the complex-cast commits in the branch history are already on main via #17 and contribute no diff — squash-merge.

Findings and how they were addressed

Severity Finding Resolution
Important ExclusiveMemoryPool::get_free_page picked a free page by free_count, not size. With this PR's exact-size pages, a small request could occupy a much larger free page and strand the difference — the stranding the PR targets. Fixed in a4fabe0: prefer the tightest fit, tie-break on free_count. Regression test exclusive_pool_reuses_the_tightest_free_page fails without the fix (2560 vs 1536 B reserved).
Important No tests for OOM error propagation, the single-retry guard in SlicedPool::alloc, or the SubSlices top-pool change. Added sliced_pool_oom_fallback_tries_an_exact_page_at_most_once and subslices_reserves_large_allocations_exactly (a4fabe0).
Important Dropping cur_avg_size is backend-agnostic: wgpu's ExclusivePages staging/uniform pools lose some cross-size page reuse. Accepted and documented as a risk in the PR body (bounded by the ~1.4× bucket ratio; tightest-fit reuse keeps exact pages efficient).
Important PR body described work already merged (#17) and lacked verification/risk notes. Body rewritten.
Minor (left as-is) Duplicated fence block in command.rs:107-115; reclaim is per-stream (sibling streams' free pages are not released); Fence::new unwraps cuEventCreate on the OOM path; IoError gains a variant on a non-#[non_exhaustive] enum → needs a version bump before the next crates.io publish (t4a-cubecl-runtime 0.10.0 is already published). —

Verification

  • Local: cargo test -p t4a-cubecl-runtime --lib memory_management → 43 passed; cargo fmt --check and cargo clippy --tests clean.
  • GPU host (A800 80 GB, CUDA 12.1), commit a4fabe0: cargo test -p t4a-cubecl-cuda → 711 passed, 24 ignored, 1 failed: reinterpret_slice_f16::global::read_from_i8x4 (left: [1.0, 0.0], right: [1.0, -8.5]). The same test fails identically on origin/main (5756141), so it is pre-existing and unrelated to this PR (CI does not run the CUDA suite).
  • End-to-end OOM reclaim on the A800 (scratch test, not committed): reserve 5…16 GiB exact pages back-to-back, dropping each — 126 GiB cumulative on an 80 GiB device, with the top pool never deallocating on its own:
    after 12 GiB: reserved 73014444032 B
    WARN cubecl_cuda::compute::command: device allocation of 13958643712 B failed; reclaiming and retrying
    after 13 GiB: reserved 13958643712 B
    after 14 GiB: reserved 28991029248 B
    after 15 GiB: reserved 45097156608 B
    after 16 GiB: reserved 62277025792 B
    test result: ok. 1 passed
    
    The eight dropped pages were reclaimed and the retry succeeded.
  • CI on a4fabe0: code-quality and documentation green; linux-std-tests (wgpu via lavapipe, including the exclusive-memory-only lib tests — the check that exercises the cross-backend pool change) still queued when this was posted.

@GiggleLiu
GiggleLiu marked this pull request as ready for review August 31, 2026 12:33
@GiggleLiu
GiggleLiu merged commit 1c9be42 into main Aug 31, 2026
3 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant