Linear regression
Linear regression predicts a number by multiplying each feature by a learned weight, adding the results and adding a bias.
Linear regression is a model that predicts a number. It learns how a label, the answer you want, relates to one or more features, the inputs. With one feature, the model is y′ = b + w1x1. The weight w1 is the slope of a line, and the bias b is where the line crosses the vertical axis. Training calculates both. With more features, each gets its own weight, as in y′ = b + w1x1 + w2x2. A fuel-economy model could add engine displacement, acceleration, cylinders and horsepower to car weight.
Google’s Machine Learning Crash Course works through seven cars, with weight in thousands of pounds and fuel economy in miles per gallon. Heavier cars generally manage fewer miles per gallon. Real measurements are noisy, so no line goes through every point. The gap between an actual value and the line’s prediction is called a residual. Least squares picks the weight and bias that make the sum of the squared residuals as small as possible. Squaring makes a big miss count far more than a small one, and the average of the squares is the mean squared error, or MSE. Gauss and Legendre developed the method independently in the late 1700s and early 1800s.
For the seven cars, the least-squares line is y′ = 33.59 − 4.57x. Its squared residuals add up to 10.29, so the MSE is 10.29 ÷ 7 = 1.47. The course’s own hand-drawn line, y′ = 34 − 4.6x, scores 1.56 on the same cars. The fitted line predicts 33.59 − 4.57 × 4.0 = 15.3 miles per gallon for a 4,000-pound car, where the hand-drawn line gives 15.6. The weight is a rate: each extra 1,000 pounds lowers the prediction by 4.57 miles per gallon. A weight of zero would mean the feature has no effect at all.
There are two routes to the best line. For linear regression a closed-form solution exists, so the best weights can be computed directly instead of by repeated steps. scikit-learn’s LinearRegression and NumPy’s lstsq both solve the least-squares problem this way, and scikit-learn uses a matrix method called singular value decomposition. Gradient descent, the loop covered in the training explainer, starts with the weight and bias near zero and repeatedly nudges them downhill. The squared-error loss of a linear model forms a single bowl, so when gradient descent converges it has found the lowest loss.
The word linear refers to the weights, not to the shape of the line. Adding squared features lets the same method fit a curve, and the model stays linear in its weights. The limits are real. A handful of extreme points can pull a least-squares fit well away from the rest of the data, because squared error pulls the line toward them. Far outside the data, its predictions only get more extreme. Still, no modelling method is used more often than linear least squares, and it can work well without a huge dataset. That makes it the usual baseline: the reference model a more complex model has to beat.
Predicting a number from examples raises three problems.
Follow seven cars from scattered dots to a fitted line.
- 1 · guessThe model predicts miles per gallon as the bias plus the weight times the car's weight, y′ = b + w1x1.
- 2 · measureFor each car, the residual is the actual value minus the line's predicted value.
- 3 · minimiseLeast squares picks the w and b that make the squared residuals, and so their average, the mean squared error, as small as possible.
- 4 · solveA direct closed-form calculation finds that pair at once, or gradient descent approaches it step by step.
- 5 · predictA new car's weight goes into the fitted line to give its predicted miles per gallon.
Least squares decides which line is best. The closed-form maths and gradient descent are two routes to that same line.
| Who | What they ask | What it works with |
|---|---|---|
| Energy analyst | “How much does fuel economy drop for every extra 1,000 pounds?” | Car weights and measured miles per gallon |
| Lab engineer | “What pressure should this gas show at 60 degrees?” | Paired temperature and pressure readings |
| Data scientist | “Does my deep model actually beat a simple baseline?” | The same features fed to a linear regression |
| Shop planner | “How many umbrellas will we sell next week?” | Past weekly sales, prices and rainfall |
- Gives a prediction you can check by hand, with one weight per feature plus a bias.
- Each weight reads as a rate, and a weight of zero means that feature is ignored.
- For squared error there is one lowest point, and a closed-form formula can compute it directly.
- A modest amount of data is often enough, and it makes a sensible baseline for bigger models.
- Curved relationships need extra features, such as squared inputs, added by hand.
- A single extreme point, or a couple of them, can drag the fitted line far off course.
- Far outside the training data its predictions keep getting more extreme, so extrapolation is risky.
- When features are strongly correlated, the weights become highly sensitive to noise.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsLinear regression (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsLinear regression: Loss (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsLinear regression: Gradient descent (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docs1.1. Linear Models (scikit-learn user guide), scikit-learn · read 27 Sept 2026
- official4.1.4.1. Linear Least Squares Regression (NIST/SEMATECH e-Handbook of Statistical Methods), NIST · read 27 Sept 2026
- official4.4.4. How can I tell if a model fits my data? (NIST/SEMATECH e-Handbook of Statistical Methods), NIST · read 27 Sept 2026
- docsnumpy.linalg.lstsq, NumPy · read 27 Sept 2026
- paperMathematics for Machine Learning, chapter 9: Linear Regression, Cambridge University Press (Deisenroth, Faisal and Ong) · read 27 Sept 2026