NWC-fp8 is compared against the library's own reference fp8 matvec (nwc_ref_fp8, one byte per weight at the memory bandwidth) and, at model level, against python -m nwc.demo MODEL --native-fp8 (same quantization, uncompressed, reference matvec). Both are ours. An external number would make the claim independent: vLLM with --quantization fp8 (Marlin fp8 weight-only on Ampere/Ada) serving Qwen/Qwen3-4B at batch 1, tokens/s on the RTX 4070 and the A16, next to python -m nwc.demo Parda21/Qwen3-4B-NWC-fp8 --load --graph. Measured so far (docs/results.md, section 8): 4070 native fp8 75.0 vs NWC-fp8 73.6 tokens/s; A16 in progress.
NWC-fp8 is compared against the library's own reference fp8 matvec (
nwc_ref_fp8, one byte per weight at the memory bandwidth) and, at model level, againstpython -m nwc.demo MODEL --native-fp8(same quantization, uncompressed, reference matvec). Both are ours. An external number would make the claim independent: vLLM with--quantization fp8(Marlin fp8 weight-only on Ampere/Ada) serving Qwen/Qwen3-4B at batch 1, tokens/s on the RTX 4070 and the A16, next topython -m nwc.demo Parda21/Qwen3-4B-NWC-fp8 --load --graph. Measured so far (docs/results.md, section 8): 4070 native fp8 75.0 vs NWC-fp8 73.6 tokens/s; A16 in progress.