Features
Features are the input facts a machine learning model reads about each example, such as a car's mileage or colour, turned into a list of numbers.
Features are the facts a machine learning model gets to look at for each example. A house-price model might see a home’s size, its number of bedrooms and its age. A model studying how weather affects test scores might see temperature, humidity and air pressure. The thing the model is trying to predict, such as the sale price, is not a feature. That is the label, and keeping the two apart is the whole point.
Models only work with numbers, so every example is turned into a feature vector: one list of decimal numbers in a fixed order. Numeric features, like mileage, are usually rescaled to a common range such as 0 to 1, so a column measured in hundreds of thousands does not swamp one measured in single digits. Categorical features, like colour, have a fixed set of values and are usually one-hot encoded: a row of zeros with a single 1 marking the category. Numbers that are really names, like postcodes, get treated as categories too. Choosing and shaping these inputs is called feature engineering.
Simpler models, such as logistic regression, can only use the features people hand them. A program that spots cars in photos might need someone to decide that “has a wheel” is worth measuring. Deep learning changed that: a neural network learns its own features in its hidden layers, and an embedding layer can turn a word or product into a short list of learned numbers. Those learned features often work well, but they are hard for people to read.
Features can also go wrong. Irrelevant, repeated or noisy columns add clutter, and feature selection tools exist to prune them. Worse is leakage, where a feature secretly carries the answer, such as a flag that is set only after a sale has happened. The model then scores brilliantly in testing and falls apart on real data.
Raw data is not something a model can read directly.
Follow one car listing into a model.
- 1 · chooseDecide which facts about each example might help predict the answer, and leave the answer itself out.
- 2 · encodeTurn categories, such as a car's colour, into numbers, usually with one-hot encoding.
- 3 · scalePut numeric features on a similar range, such as 0 to 1, so no feature dominates by size alone.
- 4 · assembleJoin every value into one list of decimal numbers, the feature vector, in the same order for every example.
- 5 · learnThe model learns a separate weight for each position in the vector.
The model never sees the raw row. It sees the feature vector, so how you build it shapes what the model can learn.
| Who | What they ask | What it works with |
|---|---|---|
| House-price team | “Should the postcode go in as a number or as a category?” | Size, bedrooms and postcode of past sales |
| Spam filter team | “Which words in an email should count as features?” | Word counts from each message |
| Online car marketplace | “Do rare paint colours each need their own column?” | The colour field on every listing |
| Recommendation team | “Can the model learn its own features for each dish instead of us writing them?” | Learned embeddings of menu items |
- Turns numbers, categories and text into one format a model can read.
- Scaling helps training settle faster and stops wide-ranging features from getting too much attention.
- One-hot encoding lets the model learn a separate weight for each category.
- Feature crosses let a simple linear model pick up combinations, such as two traits that matter only together.
- One-hot vectors get very long when a category has many values, such as tens of thousands of postcodes.
- A feature that secretly stands in for the answer makes test scores too good to trust.
- Transforms set up during training must be reused unchanged at prediction time, or the model's inputs stop matching.
- Features a deep network learns for itself are usually hard for people to interpret.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsNumerical data: How a model ingests data using feature vectors (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsNumerical data: Normalization (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsNumerical data: Binning (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsCategorical data: Vocabulary and one-hot encoding (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsCategorical data: Feature crosses (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsNeural networks (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsEmbeddings: Obtaining embeddings (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsPreprocessing data, scikit-learn · read 27 Sept 2026
- docsFeature extraction, scikit-learn · read 27 Sept 2026
- docsCommon pitfalls and recommended practices, scikit-learn · read 27 Sept 2026
- docsWorking with preprocessing layers, TensorFlow · read 27 Sept 2026
- paperDeep learning, Nature (LeCun, Bengio and Hinton, 2015) · read 27 Sept 2026
- paperDeep Learning, chapter 1: Introduction, MIT Press (Goodfellow, Bengio and Courville) · read 27 Sept 2026