Skip to content

compiler: Schedule the readers of a waited-on Function after the wait - #3037

Open
mloubout wants to merge 3 commits into
mainfrom
fix-stream-wait-first-reader
Open

mloubout wants to merge 3 commits into
mainfrom
fix-stream-wait-first-reader

Conversation

@mloubout

Copy link
Copy Markdown
Contributor

No description provided.

memcpy_prefetch attached the WaitLock to the first Cluster reading the
buffer in the order it sees them. That Cluster need not run first: a
HaloTouch is only a placeholder, later scheduled beside the stencil that
needs the halo, and later passes reorder and fuse Clusters. Under the
CUDA backend a reader of the buffer was then launched before the wait
and read the previous snapshot, e.g. the mass term of a reflection-FWI
adjoint while its halo-reading neighbours waited correctly.

Every reader now carries the wait; one on a released lock is free, and
Clusters with the same syncs fuse again, which also removes a kernel
launch in that case. The prefetch writes the buffer, so data dependences
keep it after all readers.
@mloubout mloubout added GPU compiler bug-C bug in the generated code labels Sep 29, 2026
@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 83.93%. Comparing base (f976f24) to head (751250e).

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #3037      +/-   ##
==========================================
- Coverage   83.95%   83.93%   -0.02%     
==========================================
  Files         258      258              
  Lines       55679    55682       +3     
  Branches     4769     4770       +1     
==========================================
- Hits        46743    46738       -5     
- Misses       8123     8128       +5     
- Partials      813      816       +3     
Flag Coverage Δ
pytest-gpu-aomp-amdgpuX 68.68% <100.00%> (+<0.01%) ⬆️
pytest-gpu-gcc- 78.56% <33.33%> (-0.01%) ⬇️
pytest-gpu-icx- 78.50% <33.33%> (-0.02%) ⬇️
pytest-gpu-nvc-nvidiaX 69.21% <100.00%> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Waiting on the reader alone split an interpolation's point loop: the
Clusters defining the positions posx/posy, right before the reader in
the same IterationSpace, had no sync, landed in a separate nest and left
the reader with undeclared scalars. The Clusters that only define scalars
right before a reader, with no syncs of their own, now share its wait.
Fusion's toposort only ordered ClusterGroups by data hazards and fences.
Two readers of a prefetched buffer carry no hazard between them, so a
reader could be scheduled ahead of the Cluster holding the WaitLock on
it, and read the previous snapshot: under the CUDA backend the mass term
of a reflection-FWI adjoint ran before the wait while its halo readers
ran after. The DAG now orders every later reader of a waited-on Function
after the wait, and memcpy_prefetch is back to a single WaitLock on the
first reader.
@mloubout mloubout changed the title compiler: Make every reader of a prefetched buffer wait for it compiler: Schedule the readers of a waited-on Function after the wait Sep 30, 2026

# Functions `cg0` waits on, e.g. a prefetched buffer: reading them
# is no data hazard, but no reader may be scheduled before the wait
waited = {s.target for s in flatten(cg0.syncs.values())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

so what's s.target, and why does it not appear among cg0's equations?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug-C bug in the generated code compiler GPU

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants