Support vector machines
A support vector machine separates two classes with the boundary that leaves the widest possible gap to the nearest training points of each class.
A support vector machine, or SVM, is a supervised learning method best known for sorting examples into classes, though it can also do regression and outlier detection. The approach came out of computer science research in the 1990s. Its standard form is usually traced to papers by Boser and colleagues in 1992 and by Cortes and Vapnik in 1995.
Picture each training example as a point, with one axis per feature. A new example gets the class of whichever side of the boundary it lands on. With two features that boundary is a straight line, and with more features it is a flat surface called a hyperplane. Many lines may separate the classes. The SVM keeps the one with the widest margin: the gap between the line and the closest training points from either class. As a rule, the wider that gap, the better the model tends to do on examples it has never seen. The points sitting on the edges of that margin are the support vectors. Moving any other point leaves the line where it is, as long as the point does not cross into the margin.
Real data rarely splits perfectly, so practical SVMs use a soft margin. Points may sit inside the margin or on the wrong side of the line, and each one adds a penalty. That penalty is called the hinge loss, and its aim is a boundary that stays well clear of the training points on both sides. Points safely on the right side of the margin add nothing, so they do not shape the result. In scikit-learn, C is the dial between two goals: few training mistakes and a simple boundary. Set C low and the model puts up with some mistakes in return for a smooth boundary. Set it high and it pushes to get every training point right. Some textbooks define C the other way round, as a budget for how much the margin may be violated, so check which convention a source uses.
For curved borders, SVMs turn to a kernel, a function that scores how similar two examples are. Using one lets the SVM work as if the data had been mapped into a much larger feature space, without ever building that space. The boundary is straight in that larger space, which shows up as a curve in the original features. Google’s glossary gives an example of a hundred input features mapped internally into a million dimensions. The popular RBF kernel, also called the radial kernel, works by distance, so only nearby training points sway a prediction. Its gamma setting sets that reach: raise gamma and each example only sways points very near it. The LIBSVM guide suggests starting with RBF. It recommends picking C and gamma with a grid search scored by cross-validation. Scikit-learn’s SVC class is built on LIBSVM. That library covers classification, regression and one-class SVMs.
Support vector regression, or SVR, applies the same idea to predicting numbers. Small misses are free: a training point whose predicted value lands near the true number adds no cost. When a straight boundary is enough, scikit-learn’s LinearSVC copes with millions of examples, and its training time grows roughly in step with the data. Logistic regression is often the better pick when the classes blur into each other; an SVM usually has the edge when a clear gap splits them.
When two classes can be split by a straight line, endless such lines exist, and a classifier needs a rule to pick one.
Follow one training set to the line an SVM draws.
- 1 · scalePut every feature on a similar range, such as 0 to 1, because SVMs are not scale invariant.
- 2 · widenAmong the lines that separate the classes, keep the one whose distance to the nearest point of either class is largest.
- 3 · softenAllow some points inside the margin or across the line, at a penalty whose strength is set by C.
- 4 · bendFor curved borders, use a kernel so the maximum-margin boundary is flat in a larger feature space.
The points on the margin, or on its wrong side, are the support vectors. Only they decide where the line goes.
| Who | What they ask | What it works with |
|---|---|---|
| Support team | “Is this ticket about billing or about a bug?” | Word counts from each ticket, scaled to a common range |
| Biology lab | “Which of three cell types does this sample look like?” | Measured features of each sample, with an RBF kernel |
| Data scientist | “Does a curved boundary beat a straight one on this data?” | The same training set fitted with linear and RBF kernels, compared by cross-validation |
| Energy analyst | “How much electricity will the grid need tomorrow?” | Past load readings, fitted with support vector regression |
- It still works when the features outnumber the training examples.
- It is memory efficient, because its decision uses only the support vectors.
- Swapping the kernel lets the same method draw straight or curved boundaries.
- Points far from the boundary cannot drag the line around.
- Kernel SVMs get slow on big data. Scikit-learn says SVC fit time grows at least with the square of the number of examples.
- It gives scores, not probabilities. Getting probabilities needs an extra, costly cross-validation.
- It does not choose C or gamma for you, and those choices strongly affect results.
- It does not scale features for you. A feature with a large numeric range can swamp the others.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docs1.4. Support Vector Machines (scikit-learn user guide), scikit-learn · read 27 Sept 2026
- docsSVC (scikit-learn API reference), scikit-learn · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- paperAn Introduction to Statistical Learning, chapter 9: Support Vector Machines (seventh printing), James, Witten, Hastie and Tibshirani, Springer · read 27 Sept 2026
- docsA Practical Guide to Support Vector Classification, Hsu, Chang and Lin, National Taiwan University · read 27 Sept 2026
- repoLIBSVM: A Library for Support Vector Machines, Chang and Lin, National Taiwan University · read 27 Sept 2026