Run scripts/kernbench.py --runs 30, scripts/graph_decode.py --mode nwc --fusion and the native comparison on the RTX 4080 Super (80 SMs, 736 GB/s). Projection from the cost model in docs/results.md section 3: bandwidth-bound, ~1.4x over cuBLAS. Result goes into the hardware table in the README.
Run
scripts/kernbench.py --runs 30,scripts/graph_decode.py --mode nwc --fusionand the native comparison on the RTX 4080 Super (80 SMs, 736 GB/s). Projection from the cost model in docs/results.md section 3: bandwidth-bound, ~1.4x over cuBLAS. Result goes into the hardware table in the README.