Concepts

Logistic regression

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Logistic regression weighs and adds up an example's features, squeezes the total into a probability between 0 and 1, and uses a threshold to pick a class.

1 · What it is

Logistic regression is a model that predicts a probability. Despite its name, it is used to classify things, not to predict a number such as a price. It also goes by logit regression and the log-linear classifier. Google’s crash course calls it an extremely efficient way to calculate probabilities.

It works in two moves. First it multiplies each feature by a weight, adds the results and adds a bias, giving one score, z. Then the sigmoid function squeezes z into a probability between 0 and 1. That S-shaped curve, plotted in the activation function entry, is also called the logistic function, which gives the model its name.

The score z has a meaning of its own: it is the log-odds, the natural log of the chance of yes divided by the chance of no. A positive weight is evidence for the positive class, and a negative weight is evidence against it. The model can also be paired with significance tests, such as the Wald test, to check whether a feature matters. In the statsmodels library it is called Logit and is fitted by maximum likelihood.

Training picks the weights and bias that make the true labels in the training data most likely. The loss it minimises is log loss, not the squared loss used for linear regression. Textbooks also call it cross-entropy loss. It stays small when the model is confident and right, and grows when it is confident and wrong. That loss is convex, so gradient descent cannot get stuck in a poor local minimum.

Regularisation adds a penalty on large weights. Scikit-learn switches it on unless told otherwise, a machine learning habit that statistics tools do not usually share. Its setting C is the inverse of the penalty’s strength, so a smaller C means a stronger penalty. L2 prefers many small weights, while L1 sets many weights to exactly zero. For more than two classes, the softmax function takes the place of the sigmoid. Another route, one-vs-rest, trains one yes-or-no model per class. Once each probability becomes a class, a confusion matrix counts the right and wrong calls.

2 · Why it exists

Many yes-or-no questions need a probability, and a plain weighted sum cannot give one.

Sums are unboundedA weighted sum of features can be any number from minus infinity to plus infinity, but a probability has to stay between 0 and 1.
Squared error misfitsSquared loss suits a model whose output changes at a constant rate. The output of logistic regression does not.
Answers need a callA spam filter must finally say spam or not spam. The threshold that makes that call is chosen by a person, not learned in training.
3 · How it works

Follow one email from its features to a spam decision.

Illustrative numbers. The same probability of 0.67 costs little if the email is spam and much more if it is not.
  1. 1 · scoreMultiply each feature by its weight, add the results and add the bias, giving one score called z.
  2. 2 · squashPass z through the sigmoid, which turns any number into a probability between 0 and 1.
  3. 3 · decideCompare the probability with a threshold, 0.5 by default in scikit-learn, to pick the class.
  4. 4 · learnIn training, log loss scores each probability against the true label, and the weights and bias are chosen to make the true labels most likely.

The score z is the log-odds. The sigmoid turns log-odds back into a probability.

4 · Where it's used
WhoWhat they askWhat it works with
Email provider“How likely is this message to be spam?”Features of one email, such as its links and trigger words
Review site“Is this review positive or negative?”Counts of telling words in the review text
Health research team“Which factors go with hospital readmission once age is accounted for?”Patient records, reading each weight as evidence for or against readmission
Data team“What does a simple, fast baseline score before we try a neural network?”The same features, fitted with scikit-learn's LogisticRegression
5 · What it solves, and what it doesn't
solves
  • Gives a probability that can be used as it is, or turned into a class.
  • Its loss is convex, so gradient descent cannot get trapped in a local minimum.
  • Each weight is readable. Its sign says which class that feature argues for.
  • Its probabilities are more likely to be well calibrated out of the box, because its loss matches the logit link.
doesn't solve
  • Its decision boundary is flat, a line or plane. With several classes, each boundary sits where one class's probability overtakes the others.
  • It has a linear architecture. A deep network can pick up tangled interactions between features that it misses.
  • With many features and no penalty, training can push the loss ever closer to zero. Some regularisation is needed to stop it overfitting.
  • It does not pick the threshold. That choice moves the balance of false positives and false negatives.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsLogistic regression: Calculating a probability with the sigmoid function (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  2. docsLogistic regression: Loss and regularization (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  3. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  4. paperSpeech and Language Processing, chapter 4: Logistic Regression and Text Classification (draft of August 19, 2026), Jurafsky and Martin, Stanford University · read 27 Sept 2026
  5. docsLinear Models: 1.1.11 Logistic regression (scikit-learn user guide), scikit-learn · read 27 Sept 2026
  6. docsLogisticRegression (scikit-learn API reference), scikit-learn · read 27 Sept 2026
  7. docsProbability calibration, scikit-learn · read 27 Sept 2026
  8. docsDecision Boundaries of Multinomial and One-vs-Rest Logistic Regression, scikit-learn · read 27 Sept 2026
  9. docsstatsmodels.discrete.discrete_model.Logit, statsmodels · read 27 Sept 2026