Model quantization
Quantization stores a model's numbers with fewer bits, such as 8-bit integers instead of 32-bit decimals, so it needs less memory.
A trained model keeps what it learned in numbers called weights. Quantization saves each weight with fewer bits, so it is less precise. A common switch is from 32-bit floating-point numbers, which can hold decimals, to 8-bit integers, which are whole numbers. An 8-bit integer can take only 256 different values. So each weight gets rounded to the nearest one.
To do the rounding, you first pick a range and a scale. The scale is the step size between one allowed value and the next. You divide each weight by the scale, round it, and store the small whole number. When the model runs, it multiplies that number by the scale. That gives back something close to the original. Google’s LiteRT tools report full integer quantization as 4x smaller with a 3x or greater speedup.
Researchers have pushed further for large language models. LLM.int8 keeps a few rare, unusually large values (outliers) in 16-bit. Most other values use 8-bit. This halved the memory a model uses when running. GPTQ shrinks each weight to just 3 or 4 bits. QLoRA fine-tunes, or further trains, small add-on parts called adapters. They sit on top of a frozen 4-bit base model that is not changed.
Big models hold huge numbers of weights, and wide number formats make them heavy.
Follow six weights as they are squeezed from 32 bits to 8 bits.
- 1 · calibrateFind the range of values to cover, then work out a scale from it.
- 2 · roundDivide each value by the scale and round it to the nearest whole number.
- 3 · clipValues outside the range are pushed to the nearest allowed code.
- 4 · checkTest the smaller model to see if its accuracy is still good enough.
Quantization trades a little precision for a lot less memory, so the result always needs testing.
| Who | What they ask | What it works with |
|---|---|---|
| App developer | “Can this model fit on a small embedded device?” | Model size before and after quantization |
| Researcher | “Can I fine-tune a big model with less memory?” | A frozen 4-bit base model plus small trainable adapters |
| Evaluator | “Did accuracy drop after rounding?” | Accuracy of the quantized model |
- Fewer bits per weight means the model needs less memory.
- Integer maths can make operations like matrix multiplication faster.
- Post-training methods can shrink an already-trained model without training it again.
- Integer-only models can run more efficiently on common integer-only hardware.
- Weights are stored in lower precision, so accuracy has to be checked.
- Some methods need an extra calibration step for extreme compression.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsQuantization, Hugging Face Optimum · read 28 Sept 2026
- docsQuantization overview, Hugging Face Transformers · read 28 Sept 2026
- docsPost-training quantization, Google AI Edge · read 28 Sept 2026
- paperQuantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference, arXiv (Jacob et al.) · read 28 Sept 2026
- paperLLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, arXiv (Dettmers et al.) · read 28 Sept 2026
- paperGPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, arXiv (Frantar et al.) · read 28 Sept 2026
- paperQLoRA: Efficient Finetuning of Quantized LLMs, arXiv (Dettmers et al.) · read 28 Sept 2026