Every GPU we have not measured ourselves goes into the hardware table in the README, credited to you. Especially wanted: H100 (PCIe / SXM), A100, RTX 4090 / 3090, T4 / RTX 20xx, Blackwell.
pip install neural-weight-compression transformers accelerate
python -m nwc.doctor
python -m nwc.demo Parda21/Qwen3-4B-NWC --load --graph
python -m nwc.demo Qwen/Qwen3-4B --native --graph
Post the output of all three commands here or in a benchmark issue. From a clone, python scripts/kernbench.py --runs 30 adds the per-layer kernel comparison. Thank you.
Every GPU we have not measured ourselves goes into the hardware table in the README, credited to you. Especially wanted: H100 (PCIe / SXM), A100, RTX 4090 / 3090, T4 / RTX 20xx, Blackwell.
Post the output of all three commands here or in a benchmark issue. From a clone,
python scripts/kernbench.py --runs 30adds the per-layer kernel comparison. Thank you.