Wake idle device services with ready-batch notifications - #19
Merged
Merged
Conversation
This was referenced Sep 19, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Replace the custom device-service channel's final 150 us polling sleep with
park/unpark. Keep the existing batching, spin/yield budgets, task ownership, atomic publication, completion waits and client backoff. No CPU affinity is imposed.cpyin typos: the locally installed typos 1.39 flags four unchanged occurrences.Performance decision
Prior A100 experiments on CubeCL 5939d8e showed large f64 BMM around 2.12–2.14 ms without increased process CPU-seconds; the GPU GEMM itself remains about 2.06 ms. Busy polling was rejected. All 60 numerical records passed in the notification experiment.
The original unrestricted nonregression gate failed on small eager BMM (+40.8%) and chain1k (+20.2%). Both binaries show CPU/L3-placement-dependent latency; matched same-L3 diagnostics narrow it. The user explicitly accepted these small GPU cases as documented limitations rather than adoption blockers. This does not establish unrestricted nonregression or erase the failed gate. Final integrated performance will be remeasured after the tenferro dependency update merges. These earlier measurements do not cover this exact PR revision's added client-side startup registration.
Validation
On main 1c9be42 plus this change, CUDA benchmark container, explicit Cargo jobs=16:
-D warningsand workspace format check pass.xtask validateaudit, format and workspace lint pass; its typos failure was corrected by the identifier exception and that check passes.xtask build --ci, documentation build and doctests pass. Documentation retains an existing private intra-doc link warning.xtask test --ci --test-threads 1reaches WGPU but fails because this CUDA container has no Vulkan adapter (5 pass / 371 fail / 15 ignored in WGPU). Full local validation is therefore not claimed successful.tests::test_cubecl_std::reinterpret_slice_f16::global::read_from_i8x4, returning[1.0, 0.0]instead of[1.0, -8.5]. A predeclared 3-run baseline/3-run candidate reproduction fails identically in all six runs when the channel is restored to unmodified main 1c9be42 or to this PR, with all other sources/dependencies held fixed. This is an existing failure, not a passing full CUDA suite; no unrelated reinterpretation repair is bundled.Fork CI runner correction
On initial head
966a6d307f84a6bca4ac4f2588f41c90a0f5c62d, prepare-checks, code-quality and documentation passed. Linux stable/previous and nightly Miri remained queued for upstream-specific GCP labels, with zero available fork runners. Previous PRs #17/#18 also had these jobs cancelled rather than completed.The user explicitly requested fixing the configuration and passing the checks. Commit
62ff7e2aa1eaadac65aa25fc6f64e1bd957da594changes only the tworuns-ondeclarations toubuntu-24.04, retaining all three matrix jobs, setup steps and test commands. Workflow concurrency automatically superseded the old run. No test or check has been removed or allowed to fail.Runner-corrected CI https://github.com/tensor4all/cubecl/actions/runs/35442618407 completed both Linux versions: 19 test summaries / 1,521 passing test executions each, including all five notification tests. However, its Miri job silently skipped every crate: tracel-xtask 4.16 derives package names from directory names and the fork's
t4a-*package filter matches none. That job's green status is not Miri evidence.Commit
7697cf4e070217369a2a10ec1deca30fd2d37deainvokes Cargo Miri with the actualt4a-cubecl-commonpackage, retaining original UB-only-Zmiri-ignore-leaksmode and unit/integration targets. Local nightly Miri now actually passes 43 tests, including all five wakeup tests (three native tests are excluded by existing cfgs).Final hosted run: https://github.com/tensor4all/cubecl/actions/runs/35443446621 — all six jobs passed on head
7697cf4e070217369a2a10ec1deca30fd2d37dea. Logs confirm 43 actual Miri tests passing in each unit/integration invocation, including all five notification tests. PR merged asa2adda17affd40494393a1f40d90980e1235617con 2026-09-19.Benchmarks, protocols, prior failed gates and source hashes are retained in the benchmark worktree, with durable per-item notes under
notes/nvidia-gpu/gpu/dense.md. No tenferro pin or published benchmark timing is updated by this PR.