FLOPsConcepts

Floating-point operations

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

FLOPs count the small arithmetic steps, like one multiply or one add on decimal numbers, that a computer does to train an AI model.

1 · What it is

A floating-point number is how a computer stores a number with a decimal point, such as 0.25 or 3.7. Inside the chip it is kept in binary, the ones-and-zeros system chips use. Most decimal fractions can only be stored approximately. A floating-point operation is one piece of arithmetic on such numbers: a single multiply, or a single add. FLOPs, with a small s at the end, is a count of those operations.

Why count them? Researchers who study scaling watch how a language model’s mistakes change as it grows, reads more data and gets more computing power. Their papers measure that computing in floating-point operations.

Watch the capital letters. FLOPs is an amount of work, like distance. Chip makers give speed as floating-point operations per second, written FLOPS, like speed.

So how do you count the operations in training a model? A shortcut comes from a paper on scaling laws for language models. Call the number of parameters, the adjustable numbers inside the model, N. When the model reads one token, a word or piece of a word, each parameter costs about two operations, because every multiply comes paired with an add. Reading it is called the forward pass. Then the model learns from the token by checking its mistake and nudging its numbers. That step is the backward pass, and it costs roughly twice as much. Add them up and one training token costs about 6N operations.

Repeat for all D training tokens and you get about 6ND. A later paper, known for a model called Chinchilla, calls this a common approximation. It skips small extra steps, such as layer normalization. Take an imaginary model with one billion parameters trained on 20 billion tokens. Six times 10^9 times 2 × 10^10 gives 1.2 × 10^20 FLOPs.

Papers also use the petaflop-day: a quadrillion operations a second for a whole day, about 8.64 × 10^19 operations. Our model needs roughly 1.4 of them.

A team often knows how many chips it has and for how long. The Chinchilla researchers built more than 400 test models, ranging in size from 70 million parameters up to more than 16 billion, each reading 5 to 500 billion tokens. Their finding was that model size and training tokens should grow together: double one, double the other.

Their model Chinchilla used the same compute budget as a model called Gopher, but had 70 billion parameters and four times more data. Chinchilla beat the 280-billion-parameter Gopher. The same count of operations can produce very different models.

Speed figures need care too. A chip’s spec sheet does not list one FLOPS number. It gives separate figures for different number formats, called precisions.

Precision is like a ruler. One marked in millimetres gives finer answers than one marked in centimetres. On almost every machine, an ordinary decimal number uses a 64-bit format called double precision, or FP64. It keeps 53 binary digits, so it is very exact but not perfect: the stored 0.1 is a hair above one tenth. Smaller formats, such as the 8-bit FP8, keep far fewer digits. In return, NVIDIA says its H100 chip uses less memory and runs faster with them.

In NVIDIA’s table for the H100 SXM, the FP64 row says 34 teraFLOPS, meaning 34 trillion operations a second. The FP8 row says 3,958 teraFLOPS, over a hundred times more. But that FP8 figure has an asterisk, and the footnote reads “with sparsity”. Read the fine print.

Back to our imaginary model and its 1.2 × 10^20 FLOPs. At 34 teraFLOPS, one H100 would need about 41 days. At the FP8 headline figure, it would need a little over eight hours. Same work, very different clock time. Both are sums on printed top speeds, not promises about a real run.

FLOPs have also reached the law. The European Union’s AI Act draws a line based on the floating-point operations used in training, to catch the most advanced general-purpose models. The European Commission says crossing it currently costs an estimated tens of millions of euros. Models can also be flagged on other grounds, such as capabilities or number of users.

In short, a FLOP count tells you how much arithmetic went into a model, not how good the model is.

2 · Why it exists

People need a way to say how much computing an AI model took.

Budgets are fixedTeams often know their chips and time in advance.
Size versus dataShould fixed compute go to a bigger model or more data?
Rules need a lineEU rules use training FLOPs to spot the most advanced general-purpose models.
3 · How it works

Estimate the training compute of a language model in four steps.

Estimating training compute: about 2N operations forward, about 4N more backward, so about 6N per token, times D tokens COUNTING THE WORK OF ONE TRAINING RUN Model size Forward pass Total training Add backward pass N parameters one token in ≈ 2N FLOPs multiply + add × D tokens C ≈ 6ND FLOPs + about 4N ≈ 6N per token FLOPs = how much work · FLOPS = operations per second
A common approximation: about 6 operations per parameter for every training token.
  1. 1 · countCount the model's parameters, the adjustable numbers inside it, and call that N.
  2. 2 · forwardReading one token costs roughly two operations per parameter, because each multiply comes paired with an add.
  3. 3 · backwardLearning from that token costs about twice as much again, so one training token costs about 6N operations.
  4. 4 · totalRepeat for all D tokens in the training set to get about 6ND FLOPs of training compute.

FLOPs measure how much work; FLOPS (per second) measure how fast a chip does it.

4 · Where it's used
WhoWhat they askWhat it works with
Research team“Should this compute buy a bigger model or more text?”Training FLOPs budget, parameter count and token count
Infrastructure planner“How many chips, for how long, does this budget need?”Number of accelerators and target training duration
Policy officer“Does this model cross the compute threshold in the rules?”Training compute in floating-point operations
Student“How fast is this chip at each number precision?”FLOPS figures on each chip's spec sheet
5 · What it solves, and what it doesn't
solves
  • It gives one number for a fixed training budget that researchers can plan around.
  • It lets researchers study how to split that budget between model size and training tokens.
  • It gives regulators a measurable line for very large training runs.
doesn't solve
  • Two models trained with the same compute budget can perform very differently.
  • The 6N-per-token estimate leaves out smaller costs, so it is an approximation.
  • EU rules can also flag models by capabilities or number of users, not compute alone.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsFloating-Point Arithmetic: Issues and Limitations, Python Software Foundation · read 28 Sept 2026
  2. paperScaling Laws for Neural Language Models, arXiv (Kaplan et al., OpenAI) · read 28 Sept 2026
  3. paperTraining Compute-Optimal Large Language Models, arXiv (Hoffmann et al., DeepMind) · read 28 Sept 2026
  4. officialH100 GPU, NVIDIA · read 28 Sept 2026
  5. officialGeneral-Purpose AI Models in the AI Act – Questions & Answers, European Commission · read 28 Sept 2026