QuantizationConcepts

Model quantization

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Quantization stores a model's numbers with fewer bits, such as 8-bit integers instead of 32-bit decimals, so it needs less memory.

1 · What it is

A trained model keeps what it learned in numbers called weights. Quantization saves each weight with fewer bits, so it is less precise. A common switch is from 32-bit floating-point numbers, which can hold decimals, to 8-bit integers, which are whole numbers. An 8-bit integer can take only 256 different values. So each weight gets rounded to the nearest one.

To do the rounding, you first pick a range and a scale. The scale is the step size between one allowed value and the next. You divide each weight by the scale, round it, and store the small whole number. When the model runs, it multiplies that number by the scale. That gives back something close to the original. Google’s LiteRT tools report full integer quantization as 4x smaller with a 3x or greater speedup.

Researchers have pushed further for large language models. LLM.int8 keeps a few rare, unusually large values (outliers) in 16-bit. Most other values use 8-bit. This halved the memory a model uses when running. GPTQ shrinks each weight to just 3 or 4 bits. QLoRA fine-tunes, or further trains, small add-on parts called adapters. They sit on top of a frozen 4-bit base model that is not changed.

2 · Why it exists

Big models hold huge numbers of weights, and wide number formats make them heavy.

MemoryLarge language models need a lot of GPU memory just to run.
Size on diskWeights are often stored as 32-bit floating-point numbers.
Small devicesSome embedded devices only work with whole-number data types.
3 · How it works

Follow six weights as they are squeezed from 32 bits to 8 bits.

Quantization turns six 32-bit weights into 8-bit codes Six 32-bit weights go into a calibrate step that picks the range -2 to 2 and a scale S of 2 divided by 127. The key step divides each weight by S and rounds it, giving the codes -114, -44, 6, 38, 83 and 127. The codes are stored with 8 bits each, and multiplying a code by S gives back about -1.795 for -1.8. 1 · INPUT 2 · CALIBRATE 3 · KEY STEP 4 · STORE AND USE 32-bit weights Pick range Scale and round 8-bit codes -1.8 -0.7 0.1 0.6 1.3 2.0 -2 to 2 S = 2 / 127 q = round(w / S) -114 -44 6 38 83 127 8 bits each use: w ≈ q × S -114 × S ≈ -1.795 decimals, 32 bits each whole numbers -127 to 127 close to -1.8, not exact Each code is a quarter of the size, but every weight is now rounded.
Each code times the scale gives back a close, but not exact, copy of the original weight.
  1. 1 · calibrateFind the range of values to cover, then work out a scale from it.
  2. 2 · roundDivide each value by the scale and round it to the nearest whole number.
  3. 3 · clipValues outside the range are pushed to the nearest allowed code.
  4. 4 · checkTest the smaller model to see if its accuracy is still good enough.

Quantization trades a little precision for a lot less memory, so the result always needs testing.

4 · Where it's used
WhoWhat they askWhat it works with
App developer“Can this model fit on a small embedded device?”Model size before and after quantization
Researcher“Can I fine-tune a big model with less memory?”A frozen 4-bit base model plus small trainable adapters
Evaluator“Did accuracy drop after rounding?”Accuracy of the quantized model
5 · What it solves, and what it doesn't
solves
  • Fewer bits per weight means the model needs less memory.
  • Integer maths can make operations like matrix multiplication faster.
  • Post-training methods can shrink an already-trained model without training it again.
  • Integer-only models can run more efficiently on common integer-only hardware.
doesn't solve
  • Weights are stored in lower precision, so accuracy has to be checked.
  • Some methods need an extra calibration step for extreme compression.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsQuantization, Hugging Face Optimum · read 28 Sept 2026
  2. docsQuantization overview, Hugging Face Transformers · read 28 Sept 2026
  3. docsPost-training quantization, Google AI Edge · read 28 Sept 2026
  4. paperQuantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference, arXiv (Jacob et al.) · read 28 Sept 2026
  5. paperLLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, arXiv (Dettmers et al.) · read 28 Sept 2026
  6. paperGPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, arXiv (Frantar et al.) · read 28 Sept 2026
  7. paperQLoRA: Efficient Finetuning of Quantized LLMs, arXiv (Dettmers et al.) · read 28 Sept 2026