Naive Bayes
Naive Bayes sorts things into classes by multiplying how common each class is by how likely each feature is in that class, then picking the top score.
Naive Bayes is a family of classifiers built on Bayes’ theorem. A classifier looks at one example, such as an email, picks out useful features, and puts it into one of a fixed set of classes, such as spam or not spam. Bayes’ theorem lets you turn a probability around: from how likely the words are in spam to how likely the email is spam given its words. The method goes back to 1961, when it was invented to sort documents into subject categories.
The naive part is one bold shortcut: once you know the class, each feature tells you nothing about the others. For text that is plainly false, since words such as “hong” and “kong” tend to appear together. The usual text version also ignores word order and treats each document as a bag of words. In return, each word’s numbers can be learned on their own, which keeps training simple when there are many words.
Training is counting. The prior is the share of training examples in each class. A word’s likelihood for a class is how often it appears among all the words in that class’s training documents. To score a new email, the model multiplies the prior by the likelihood of each word and picks the class with the biggest product. Real systems add logarithms instead of multiplying, because multiplying many tiny numbers can underflow. One trap remains: a word never seen with a class gives a likelihood of zero, and that zero cancels all the other evidence. Add-one, or Laplace, smoothing fixes this by adding 1 to every count before dividing. In scikit-learn, the amount added is a setting called alpha, and MultinomialNB sets it to 1 by default.
Variants differ in how they model each feature. Multinomial naive Bayes uses word counts and is a classic choice for text. Bernoulli naive Bayes only records whether each word is present, and counts a missing word as evidence too. Gaussian naive Bayes handles measurements, assuming each follows a bell curve within a class. A 1998 comparison found Bernoulli sometimes better with small vocabularies and multinomial usually better with large ones.
Its probability estimates are often poor, yet it still classifies well, because picking a class only needs the right class to score highest. One study argues that dependencies between features can spread evenly across the classes or cancel each other out. The scikit-learn guide calls it extremely fast next to more complex methods. It also warns that its probability outputs are not to be taken too seriously.
Spam filtering was one of the first uses of naive Bayes for text. A 1998 junk-mail study built its filter with a naive Bayesian classifier. It marked a message as junk only when the model gave more than a 99.9% chance. The authors did not trust the model’s probabilities to be accurate, but found that threshold still reasonable.
Scoring every combination of words is impossible, so a simpler model is needed.
Score one short email against two classes.
- 1 · countCount how many training emails are in each class, and how often each word appears in each class.
- 2 · smoothTurn the counts into class priors and word likelihoods, adding 1 to every word count first.
- 3 · multiplyFor a new email, multiply each class's prior by the likelihood of every word in the email.
- 4 · normaliseDivide each product by their sum so the class probabilities add up to 1.
- 5 · decidePick the class with the highest score.
The model is naive because it multiplies word likelihoods as if the words had nothing to do with each other.
| Who | What they ask | What it works with |
|---|---|---|
| Email provider | “Is this message junk or a real email?” | The words in one incoming email, scored against a spam class and a not-spam class |
| Review site team | “Is this movie review positive or negative?” | The bag of words in one review |
| Library or research database | “Which subject heading fits this paper?” | The terms in the paper, compared with each subject's word counts |
- Trains in a single pass of counting over the data.
- Needs only one number per word and class, not one for every combination of words.
- Works with a small amount of training data.
- Often picks the right class even though its independence assumption is false.
- Its probabilities are poor estimates, so a score of 0.99 should not be read as a 99% chance.
- It ignores word order, so it treats China sues France and France sues China alike.
- Discriminative models such as logistic regression are often more accurate.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docs1.9. Naive Bayes (scikit-learn user guide), scikit-learn · read 27 Sept 2026
- docsMultinomialNB (scikit-learn API reference), scikit-learn · read 27 Sept 2026
- paperIntroduction to Information Retrieval, 13.2: Naive Bayes text classification, Manning, Raghavan and Schütze, Cambridge University Press · read 27 Sept 2026
- paperIntroduction to Information Retrieval, 13.4: Properties of Naive Bayes, Manning, Raghavan and Schütze, Cambridge University Press · read 27 Sept 2026
- paperSpeech and Language Processing, Appendix B: Naive Bayes, Text Classification, and Sentiment (draft of August 19, 2026), Jurafsky and Martin, Stanford University · read 27 Sept 2026
- paperA Bayesian Approach to Filtering Junk E-Mail, Sahami, Dumais, Heckerman and Horvitz, AAAI Workshop on Learning for Text Categorization 1998 · read 27 Sept 2026
- paperA Comparison of Event Models for Naive Bayes Text Classification, McCallum and Nigam, AAAI Workshop on Learning for Text Categorization 1998 · read 27 Sept 2026
- paperThe Optimality of Naive Bayes, Zhang, FLAIRS 2004 (AAAI) · read 27 Sept 2026