Skip to content

External fp8 baseline: vLLM (fp8 weight-only) tokens/s on the same GPUs #10

Description

@parda21

NWC-fp8 is compared against the library's own reference fp8 matvec (nwc_ref_fp8, one byte per weight at the memory bandwidth) and, at model level, against python -m nwc.demo MODEL --native-fp8 (same quantization, uncompressed, reference matvec). Both are ours. An external number would make the claim independent: vLLM with --quantization fp8 (Marlin fp8 weight-only on Ampere/Ada) serving Qwen/Qwen3-4B at batch 1, tokens/s on the RTX 4070 and the A16, next to python -m nwc.demo Parda21/Qwen3-4B-NWC-fp8 --load --graph. Measured so far (docs/results.md, section 8): 4070 native fp8 75.0 vs NWC-fp8 73.6 tokens/s; A16 in progress.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    benchmarkmeasurements from a GPUhelp wantedwe cannot do this one aloneroadmapplanned work, shown in the README roadmap

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions