Concepts

F1 score

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

The F1 score rolls a classifier's precision and recall into one number from 0 to 1, and it stays low if either of them is low.

1 · What it is

A classifier sorts items into yes and no, and two numbers describe its yes calls. Precision asks what share of the items the model called positive really were positive. Recall asks what share of the truly positive items the model found. F1 is the harmonic mean of the two, which works out to 2 × precision × recall ÷ (precision + recall). Written with counts it becomes 2TP ÷ (2TP + FP + FN), for true positives, false positives and false negatives. The score runs from 0, the worst, to 1, the best. Libraries also call it the balanced F-score or F-measure.

Picture a model with precision 0.9 and recall 0.1: it is careful, but it finds almost nothing. A plain average gives it 0.5, while F1 gives 2 × 0.9 × 0.1 ÷ 1.0 = 0.18. F1 is pulled toward the lower of its two inputs. It can never come out higher than the plain average. Search engines show the same trap at a bigger scale. A system that hands back every document finds all the relevant ones, so its plain average can’t drop below 50%. Yet if just one document in 10,000 matters, its F1 comes out at about 0.02%. True negatives, the items correctly left out, never enter the formula.

The general form, F-beta, is a weighted harmonic mean of precision and recall. A beta above 1 gives recall more weight, and a beta below 1 favours precision. F1 is the case where beta is 1 and both count equally. With more than two classes, F1 is computed for each class and then combined. Macro averaging takes the plain mean of the per-class scores. Weighted averaging weights each class by how many true examples it has. Micro averaging pools the counts from all classes first. In a multiclass problem with every class included, micro-averaged F1 equals accuracy.

The same harmonic mean first appeared in ecology in 1948. In the 1970s, C. J. van Rijsbergen brought it into the study of search systems. His book scores a search system with one number built from precision and recall. Its beta setting says how many times more a user values recall than precision. The notation F1 was adopted in 1992. Chicco and Jurman argue that F1 can look over-optimistic on imbalanced data. Their preferred alternative, the Matthews correlation coefficient, only gives a high score when true and false positives and negatives all come out well.

2 · Why it exists

Precision and recall are two numbers, and recall alone is easy to push to a perfect score.

Two scores pull apartPrecision and recall usually trade off against each other, so one model can win on one and lose on the other.
Recall is easy to gameA search system that returns every document reaches perfect recall, but its precision is very low.
Accuracy hides rare casesWhen almost every item is negative, a system that says no to everything can still look accurate.
3 · How it works

Follow two made-up models through the F1 formula.

Made-up counts. Both models average 0.50, but F1 drops Model A to 0.18.
  1. 1 · countTest the model and count its true positives, false positives and false negatives.
  2. 2 · rateTurn the counts into precision, the share of positive calls that were right, and recall, the share of real positives that were found.
  3. 3 · combineTake the harmonic mean, 2 × precision × recall ÷ (precision + recall).
  4. 4 · readRead the result from 0 to 1, where a big gap between precision and recall drags the score toward the lower one.
4 · Where it's used
WhoWhat they askWhat it works with
Search team“Does the new ranking find relevant pages without the junk?”Retrieved documents checked against relevance judgments
Question-answering benchmark“How close are the model's answers to the reference answers?”Exact match together with F1 on SQuAD
Summarisation team“How much of the reference summary does the generated one cover?”ROUGE precision and recall rolled into one ROUGE F1
Screening team“Can we count missed cases as worse than false alarms?”F-beta with beta 2, which makes recall twice as important
5 · What it solves, and what it doesn't
solves
  • It gives one number that is only high when precision and recall are both high.
  • It is preferable to accuracy when one class is much rarer than the other.
  • It stops a return-everything strategy from scoring well.
  • Its F-beta form can tilt the score toward recall or precision when one kind of mistake costs more.
doesn't solve
  • It ignores true negatives, so it never credits the negatives a model correctly leaves out.
  • Renaming which class counts as positive changes the score.
  • Plain F1 weights precision and recall equally, even when the real costs of the two mistakes differ.
  • Critics argue it can still flatter a model on imbalanced data.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  2. docsClassification: Accuracy, recall, precision, and related metrics (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  3. docsf1_score, scikit-learn · read 27 Sept 2026
  4. docsfbeta_score, scikit-learn · read 27 Sept 2026
  5. docs3.4. Metrics and scoring: quantifying the quality of predictions, scikit-learn · read 27 Sept 2026
  6. paperEvaluation of unranked retrieval sets (Introduction to Information Retrieval, chapter 8), Manning, Raghavan and Schütze, Cambridge University Press · read 27 Sept 2026
  7. paperInformation Retrieval, chapter 7, Evaluation, C. J. van Rijsbergen, University of Glasgow · read 27 Sept 2026
  8. paperThe advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation, Chicco and Jurman, BMC Genomics 2020 · read 27 Sept 2026