Mixed precisionConcepts

Mixed precision training

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Mixed precision trains with both 16-bit and 32-bit numbers so the run is faster and uses less memory.

1 · What it is

Ordinary training keeps each number in 32-bit single precision, called FP32. Mixed precision keeps the weights, activations, and gradients in 16-bit storage called FP16. That format is IEEE half precision. A 32-bit master copy of the weights is kept, and the optimizer adds the gradient to that copy. Each forward and backward pass starts from a 16-bit rounding of that master copy.

This nearly halves memory requirements. Activations are also stored in half precision, so overall training memory for deep networks is roughly halved. Half-precision math throughput on recent GPUs is 2 to 8 times higher than single precision. Volta Tensor Cores multiply FP16 matrices and can accumulate those products in FP32. Without FP32 accumulation, some FP16 models did not match the baseline accuracy.

FP16 writes a zero when a number’s magnitude is smaller than 2−24. Raising the loss before the backward pass shifts those small gradients into values FP16 can hold. The chain rule then applies that same scale to every gradient. Those gradients are divided back down before the optimizer step. Too large a scale overflows, filling the weight gradients with infinities and NaNs and irreversibly damaging the weights.

BFLOAT16 can represent the same range of values as FP32, and conversion either way is simple. The BFLOAT16 experiments were run without hyperparameter changes. Keeping some parts in 32-bit types for numeric stability lets the model train equally well on metrics such as accuracy. Most models use the float32 dtype, which takes 32 bits of memory. Float16 and bfloat16 each take 16 bits of memory. NVIDIA GPUs can run float16 faster than float32, and TPUs can run bfloat16 faster than float32. TensorFlow reports that its mixed precision API can improve performance by more than 3 times on modern GPUs. NVIDIA reports up to a 3 times overall speedup on the most arithmetic-heavy architectures. That automated setup keeps float32 master weights so each update can accumulate there. PyTorch automatic mixed precision runs some operations in float32 and others in float16. Linear layers and convolutions are much faster in float16 or bfloat16. Reductions often need the dynamic range of float32. The PyTorch recipe expects a significant 2 to 3 times speedup on Volta, Turing and Ampere Tensor Core GPUs.

2 · Why it exists

Larger networks need more memory and compute to train.

Memory costReduced-precision formats cut the memory required to store the same values during training.
Narrow rangeFP16 has a narrower range than single precision. When the magnitude falls below 2−24, FP16 stores that number as zero.
Weak gradientsScaling the loss shifts small gradients into the range FP16 can represent.
3 · How it works

Follow one training step from the master weights to the next update.

  1. 1 · copyKeep an FP32 master copy of the weights and update it with the weight gradient during the optimizer step.
  2. 2 · computeRound that copy to half precision for the forward and backward pass.
  3. 3 · scaleMultiply the loss before back-propagation so small gradients stay inside the FP16 range.
  4. 4 · updateUnscale the gradients before the weight update so the step size matches FP32 training.

A loss scale that overflows writes infinities into the gradients. Volta Tensor Cores multiply FP16 matrices and can accumulate the products in FP16 or FP32.

4 · Where it's used
WhoWhat they askWhat it works with
Training engineer“Can this model train in FP16 without losing accuracy?”The FP32 master weights, the loss scale, and the full-precision baseline
GPU operator“Which layers should stay in float32?”Reductions and other wide-range operations
Model owner“Will bfloat16 avoid a hand-tuned loss scale?”The format range and the training hyperparameters
Capacity planner“How much activation memory does half precision free?”Peak training memory against the FP32 run
5 · What it solves, and what it doesn't
solves
  • Half precision nearly halves memory requirements.
  • Updating an FP32 master copy of the weights matched FP32 training results.
  • A loss scale of 8 let the Multibox SSD detector match FP32 accuracy.
  • Volta Tensor Cores multiply FP16 matrices and can accumulate the products in FP32.
doesn't solve
  • The SSD detector failed to train in FP16 without loss scaling.
  • Too large a scale overflows and writes infinities and NaNs that irreversibly damage the weights.
  • Updating FP16 weights directly caused an 80 percent relative accuracy loss on a Mandarin speech model.
  • In the BFLOAT16 study, the FP16 results needed an extra loss-scaling hyperparameter.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperMixed Precision Training, Micikevicius et al., ICLR 2018 · read 28 Sept 2026
  2. paperA Study of BFLOAT16 for Deep Learning Training, Kalamkar et al. · read 28 Sept 2026
  3. docsMixed precision, TensorFlow · read 28 Sept 2026
  4. docsTrain With Mixed Precision, NVIDIA · read 28 Sept 2026
  5. docsAutomatic Mixed Precision, PyTorch · read 28 Sept 2026