Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing, Siddharth Joshi
University of Notre Dame · EMNLP 2026 (Main)
- 2026-08-21: QTEA is accepted to the EMNLP 2026 Main Conference.
QTEA is a post-training quantization method that quantizes the linear layers of a decoder-only LLM into effectively 1.7 bits per weight using just a single pass over 256 calibration sequences.
We treat salient weights as error compensators. QTEA first quantizes every
weight into a compact ternary base {-1, 0, +1}, then spends a small residual
budget on the columns where ternarization causes the largest drop in accuracy.
This allows for keeping one FP8 value per four rows inside those columns,
facilitating a column-semi sparse layout that recovers most of the accuracy
of unstructured salient weights while staying GPU-friendly. We also implement
two further changes to the GPTQ-style column-by-column sweep: we use a
per-column rescale factor that is jointly optimized with the ternary assignments,
and we introduce an error decay term that attenuates error propagation so late
columns are not over-compensated. A lookup-table CUDA kernel then evaluates the
result ensuring that the hardware can take advantage of the ternarization.
Quality. QTEA achieves the best average zero-shot accuracy among the methods we evaluated, with the advantage increasing with model size.
- Qwen3-14B: average zero-shot accuracy rises from 45.11% to 52.65%, a 16.7% relative gain, with 1.40× lower WikiText-2 perplexity (16.48 → 11.78) and 2.61× lower C4 perplexity (68.13 → 26.14).
- Llama3-8B: accuracy improves from 37.79% to 40.29% (6.6% relative), with 1.34× and 1.95× lower WikiText-2 and C4 perplexity.
- The 1:4 semi-sparse residual costs only 0.9 accuracy points against an unstructured salient-weight upper bound, at 4× lower residual storage.
Efficiency. The lookup-table kernel delivers practical speedups.
- 7.2× faster per-token generation than FP16 with CUDA Graphs on Llama2-70B (41.13 → 5.70 ms/token), and 13.3× against FP16 without CUDA Graphs; 3.62× on Qwen3-14B.
- Latency matches a ternary-only kernel to within 0.01–0.03 ms/token.
- Speedup is stable at 3.07–3.20× for batch size 1 across 512–4096 tokens of context, and still 2.54× at batch size 4.
- On a commercially available TSMC 22nm-based implementation, a co-designed accelerator delivers 3.83× lower latency and 69.4% lower energy than dense FP16 matrix multiplication.
| Method | Qwen3-14B Wiki2 PPL ↓ | Qwen3-14B C4 PPL ↓ | Qwen3-14B 0-shot avg ↑ | Llama3-8B Wiki2 PPL ↓ | Llama3-8B C4 PPL ↓ | Llama3-8B 0-shot avg ↑ |
|---|---|---|---|---|---|---|
| FP16 | 6.38 | 9.68 | 68.23 | 6.14 | 9.45 | 65.59 |
| GPTQ | 37.90 | 74.50 | 37.31 | 1480.43 | 394.74 | 33.30 |
| Slim-LLM | 22.85 | 68.38 | 44.13 | 38.21 | 390.02 | 34.52 |
| PB-LLM | 2.89e4 | 2.44e4 | 32.50 | 73.08 | 104.15 | 36.25 |
| PT²-LLM | 16.48 | 68.13 | 45.11 | 32.19 | 129.83 | 37.79 |
| QTEA (1.7 bit) | 11.78 | 26.14 | 52.65 | 24.09 | 66.45 | 40.29 |
conda create -n qtea python=3.12 -y
conda activate qtea
pip install -r requirements.txtIf you already maintain a PyTorch CUDA environment, install the packages from
requirements.txt there instead of creating a new one. lm_eval is only needed
for zero-shot accuracy, and a CUDA toolchain (nvcc) is only needed for the
lookup-table inference kernel.
To quantize a model yourself:
python quantize.py --model Qwen/Qwen3-8B-Base --output packed/qwen3-8b.ptThis writes a packed checkpoint holding the ternary codes, with the residuals and the
scales compressed. Only one decoder block is resident on the GPU at a time, so the
footprint is set by the calibration activations and the largest block rather than by
the model. Lower --nsamples if OOM;
A QTEA-quantized Qwen3-14B-Base checkpoint is also published on the Hugging Face
Hub as ims-lab/Qwen3-14B-base-QTEA,
so you can skip this step and go straight to Evaluation:
huggingface-cli download ims-lab/Qwen3-14B-base-QTEA packed_qwen3_14b.pt --local-dir packed/python eval/evaluate.py --model Qwen/Qwen3-8B-Base --checkpoint packed/qwen3-8b.ptReports WikiText-2 and C4 perplexity plus zero-shot accuracy on the seven tasks
above. Drop --checkpoint for the FP16 baseline, and add --backend lut to run
the packed weights through the lookup-table CUDA kernel instead of rebuilding
FP16 weights. See eval/README.md for the full set of
options and for how the kernel works.
bash scripts/quantize_qwen3.sh # Qwen3-Base, 0.6B to 14B
bash scripts/quantize_llama3.sh # Llama3-8BBoth scripts quantize and then evaluate, writing JSON results to results/. The
defaults in quantize.py are the values used for paper. Note that We apply an
additional ternary-fitting round for Llama3, this is captured in the script.
Reproduced here on one H200 with torch 2.8 / transformers 4.56:
| model | WikiText-2 | C4 | zero-shot avg |
|---|---|---|---|
| Qwen3-0.6B-Base | 63.58 | 172.66 | 34.37 |
| Qwen3-14B-Base | 11.78 | 26.14 | 52.67 |
| Llama3-8B | 24.09 | 66.39 | 40.13 |
Quantization is deterministic for a given GPU and library stack, but numerics are impacted by differences between environments: the ternary assignment is a hard threshold inside a sequential error-feedback loop, which can be impacted by rounding differences between hardware.
Llama-3 and Qwen3 are supported out of the box. Any decoder stack that exposes
model.model.layers with the usual seven projections works unchanged; other
families need their projection names added to qtea/sequential.py.
What is the best way to quantize an LLM below 2 bits without retraining? QTEA is a post-training method: one calibration pass, no gradient training and no QAT. It keeps a ternary base for every weight and adds a small, hardware-friendly sparse residual only where ternarization hurts most. Among the sub-2-bit PTQ methods we evaluated (PT²-LLM, PB-LLM, Slim-LLM, GPTQ), it gives the best accuracy.
How is QTEA different from BitNet b1.58? BitNet trains ternary models from scratch. QTEA ternarizes an existing pretrained checkpoint after training.
How is it different from GPTQ? QTEA builds on the GPTQ column-by-column sweep and adds three things: a ternary quantizer with per-column rescale factors optimized jointly with the ternary assignments, a salient-column sparse residual, and an error-decay term that stops late columns from being over-compensated.
Does it actually run faster? Yes. Five ternary values are packed per byte and evaluated with a lookup-table CUDA GEMV kernel: 7.2× faster per-token generation than FP16 on Llama2-70B with CUDA Graphs, and 3.62× on Qwen3-14B.
Which models are supported?
Llama-3 and Qwen3 (0.6B–14B) out of the box, plus any decoder that exposes
model.model.layers with the standard seven projections. A pre-quantized
Qwen3-14B-Base checkpoint is on
Hugging Face.
Quantization
quantize.py: command-line entry point — quantizes one model and saves the packed checkpoint.qtea/qtea.py: the QTEA algorithm for a single linear layer. Sweeps the columns left to right, picking the salient columns, refining their rescale factors, fitting the sparse residual and propagating the quantization error.qtea/quantizer.py: fits the ternary scale and centre that every group of 128 columns shares.qtea/sequential.py: applies the quantizer to the whole decoder stack, one block at a time, so only one block ever sits on the GPU.qtea/pack.py: the checkpoint format — packs and unpacks the ternary codes (five per byte), the FP8 residuals and the scales.qtea/data.py: WikiText-2 and C4 loaders for calibration and perplexity.qtea/model.py: HuggingFace model loading.
Evaluation
eval/evaluate.py: measures perplexity and zero-shot accuracy for a packed checkpoint, or for the FP16 baseline.eval/perplexity.py: perplexity, computed one decoder block at a time to keep memory low.eval/packed_linear.py: loads a packed checkpoint into a HuggingFace model.eval/ternary_kernel.pyandeval/csrc/ternary_gemv.cu: the lookup-table GEMV kernel — the Python wrapper and the CUDA source behind the paper's latency numbers.eval/check_kernel.py: checks the kernel against a dense FP16 matmul.
Other
scripts/: one reproduction script per model family.tests/: round-trip tests for the checkpoint format (python -m pytest tests).figs/: figures used by this README and byeval/README.md.
QTEA appears at the EMNLP 2026 Main Conference. If you find it useful in your research, please cite:
@misc{guo2026qteaternaryllmssparse,
title={QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization},
author={Yipin Guo and Arun M George and Jie Fu and Tareq Mahmoud and Sixue Xing and Siddharth Joshi},
year={2026},
eprint={2609.00224},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2609.00224},
}This project is released under the MIT License. See LICENSE for details. The
evaluation datasets and the models retain their own licences.

