Measured while adding the fp8 element type (docs/results.md, section 8). The v9 decoder runs at 11.0–12.6 SM clocks per warp step (64 weights) on the A16 (GA107, 1.76 GHz nominal); on the RTX 4070 (AD104) the same SASS (identical instruction counts for sm_86 and sm_89) takes about 17 clocks per step at a verified 2.7 GHz.
What was ruled out on the 4070 (gate_up 19456 × 2560, fp8, 0.146 ms baseline):
- memory: L2-resident data is as fast as DRAM; knockout without raw-plane loads 0.112 ms, without stream refills 0.144 ms, without either 0.140 ms
- prefetch: this block's rows into L1 (
-DPF_ROWS_L1=1) neutral, next block into L2 (-DPF_NEXT_L2=1) −5 to −10 %
- clock: 2.7 GHz sustained (
nvidia-smi -lms 250), power cap flag set but no clock drop
Candidates: shared-memory bank conflicts of the 12-bit LUT gather (32 random words per warp), dependent-chain latency in the SM (each step is a serial chain: funnel shift → peek → LDS → advance), Ada's scheduler. Nsight Compute (stall reasons, smsp__warp_issue_stalled_*) is the tool; it is not installed on the 4070 machine. If you have ncu on an Ada card: NWC_DLL=build/nwc_ops.dll ncu --section WarpStateStats --section SpeedOfLight -k regex:kern_block python scripts/kernbench.py --ncu --only gate_up.
Why it matters: closing the gap makes fp8 on Ada bandwidth-bound (1.14× a native fp8 matvec instead of parity) and gives BF16 headroom on the 4080 Super / 4090.
Measured while adding the fp8 element type (docs/results.md, section 8). The v9 decoder runs at 11.0–12.6 SM clocks per warp step (64 weights) on the A16 (GA107, 1.76 GHz nominal); on the RTX 4070 (AD104) the same SASS (identical instruction counts for sm_86 and sm_89) takes about 17 clocks per step at a verified 2.7 GHz.
What was ruled out on the 4070 (gate_up 19456 × 2560, fp8, 0.146 ms baseline):
-DPF_ROWS_L1=1) neutral, next block into L2 (-DPF_NEXT_L2=1) −5 to −10 %nvidia-smi -lms 250), power cap flag set but no clock dropCandidates: shared-memory bank conflicts of the 12-bit LUT gather (32 random words per warp), dependent-chain latency in the SM (each step is a serial chain: funnel shift → peek → LDS → advance), Ada's scheduler. Nsight Compute (stall reasons,
smsp__warp_issue_stalled_*) is the tool; it is not installed on the 4070 machine. If you have ncu on an Ada card:NWC_DLL=build/nwc_ops.dll ncu --section WarpStateStats --section SpeedOfLight -k regex:kern_block python scripts/kernbench.py --ncu --only gate_up.Why it matters: closing the gap makes fp8 on Ada bandwidth-bound (1.14× a native fp8 matvec instead of parity) and gives BF16 headroom on the 4080 Super / 4090.