Skip to content

Ada (RTX 4070) needs ~17 SM clocks per warp step where Ampere (A16) needs 11, same SASS #9

Description

@parda21

Measured while adding the fp8 element type (docs/results.md, section 8). The v9 decoder runs at 11.0–12.6 SM clocks per warp step (64 weights) on the A16 (GA107, 1.76 GHz nominal); on the RTX 4070 (AD104) the same SASS (identical instruction counts for sm_86 and sm_89) takes about 17 clocks per step at a verified 2.7 GHz.

What was ruled out on the 4070 (gate_up 19456 × 2560, fp8, 0.146 ms baseline):

  • memory: L2-resident data is as fast as DRAM; knockout without raw-plane loads 0.112 ms, without stream refills 0.144 ms, without either 0.140 ms
  • prefetch: this block's rows into L1 (-DPF_ROWS_L1=1) neutral, next block into L2 (-DPF_NEXT_L2=1) −5 to −10 %
  • clock: 2.7 GHz sustained (nvidia-smi -lms 250), power cap flag set but no clock drop

Candidates: shared-memory bank conflicts of the 12-bit LUT gather (32 random words per warp), dependent-chain latency in the SM (each step is a serial chain: funnel shift → peek → LDS → advance), Ada's scheduler. Nsight Compute (stall reasons, smsp__warp_issue_stalled_*) is the tool; it is not installed on the 4070 machine. If you have ncu on an Ada card: NWC_DLL=build/nwc_ops.dll ncu --section WarpStateStats --section SpeedOfLight -k regex:kern_block python scripts/kernbench.py --ncu --only gate_up.

Why it matters: closing the gap makes fp8 on Ada bandwidth-bound (1.14× a native fp8 matvec instead of parity) and gives BF16 headroom on the 4080 Super / 4090.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    help wantedwe cannot do this one aloneroadmapplanned work, shown in the README roadmap

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions