Concepts

Features

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Features are the input facts a machine learning model reads about each example, such as a car's mileage or colour, turned into a list of numbers.

1 · What it is

Features are the facts a machine learning model gets to look at for each example. A house-price model might see a home’s size, its number of bedrooms and its age. A model studying how weather affects test scores might see temperature, humidity and air pressure. The thing the model is trying to predict, such as the sale price, is not a feature. That is the label, and keeping the two apart is the whole point.

Models only work with numbers, so every example is turned into a feature vector: one list of decimal numbers in a fixed order. Numeric features, like mileage, are usually rescaled to a common range such as 0 to 1, so a column measured in hundreds of thousands does not swamp one measured in single digits. Categorical features, like colour, have a fixed set of values and are usually one-hot encoded: a row of zeros with a single 1 marking the category. Numbers that are really names, like postcodes, get treated as categories too. Choosing and shaping these inputs is called feature engineering.

Simpler models, such as logistic regression, can only use the features people hand them. A program that spots cars in photos might need someone to decide that “has a wheel” is worth measuring. Deep learning changed that: a neural network learns its own features in its hidden layers, and an embedding layer can turn a word or product into a short list of learned numbers. Those learned features often work well, but they are hard for people to read.

Features can also go wrong. Irrelevant, repeated or noisy columns add clutter, and feature selection tools exist to prune them. Worse is leakage, where a feature secretly carries the answer, such as a flag that is set only after a sale has happened. The model then scores brilliantly in testing and falls apart on real data.

2 · Why it exists

Raw data is not something a model can read directly.

Models only read numbersEvery value a model takes in must be a number, but much real data is words, categories or free text.
Raw numbers can misleadA feature measured in millions can drown out one measured in single digits, and some numbers, like postcodes, are really names.
The wrong inputs teach wrong lessonsA column that secretly gives away the answer makes a model look brilliant in testing and fail in real use.
3 · How it works

Follow one car listing into a model.

An illustrative listing with assumed ranges. The price is the label, so it must never sneak into the feature vector.
  1. 1 · chooseDecide which facts about each example might help predict the answer, and leave the answer itself out.
  2. 2 · encodeTurn categories, such as a car's colour, into numbers, usually with one-hot encoding.
  3. 3 · scalePut numeric features on a similar range, such as 0 to 1, so no feature dominates by size alone.
  4. 4 · assembleJoin every value into one list of decimal numbers, the feature vector, in the same order for every example.
  5. 5 · learnThe model learns a separate weight for each position in the vector.

The model never sees the raw row. It sees the feature vector, so how you build it shapes what the model can learn.

4 · Where it's used
WhoWhat they askWhat it works with
House-price team“Should the postcode go in as a number or as a category?”Size, bedrooms and postcode of past sales
Spam filter team“Which words in an email should count as features?”Word counts from each message
Online car marketplace“Do rare paint colours each need their own column?”The colour field on every listing
Recommendation team“Can the model learn its own features for each dish instead of us writing them?”Learned embeddings of menu items
5 · What it solves, and what it doesn't
solves
  • Turns numbers, categories and text into one format a model can read.
  • Scaling helps training settle faster and stops wide-ranging features from getting too much attention.
  • One-hot encoding lets the model learn a separate weight for each category.
  • Feature crosses let a simple linear model pick up combinations, such as two traits that matter only together.
doesn't solve
  • One-hot vectors get very long when a category has many values, such as tens of thousands of postcodes.
  • A feature that secretly stands in for the answer makes test scores too good to trust.
  • Transforms set up during training must be reused unchanged at prediction time, or the model's inputs stop matching.
  • Features a deep network learns for itself are usually hard for people to interpret.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  2. docsNumerical data: How a model ingests data using feature vectors (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  3. docsNumerical data: Normalization (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  4. docsNumerical data: Binning (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  5. docsCategorical data: Vocabulary and one-hot encoding (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  6. docsCategorical data: Feature crosses (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  7. docsNeural networks (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  8. docsEmbeddings: Obtaining embeddings (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  9. docsPreprocessing data, scikit-learn · read 27 Sept 2026
  10. docsFeature extraction, scikit-learn · read 27 Sept 2026
  11. docsCommon pitfalls and recommended practices, scikit-learn · read 27 Sept 2026
  12. docsWorking with preprocessing layers, TensorFlow · read 27 Sept 2026
  13. paperDeep learning, Nature (LeCun, Bengio and Hinton, 2015) · read 27 Sept 2026
  14. paperDeep Learning, chapter 1: Introduction, MIT Press (Goodfellow, Bengio and Courville) · read 27 Sept 2026