Feature engineering
Feature engineering turns raw data, like prices and colour names, into the lists of numbers a model can learn from.
A model does not read a table the way you do. It reads one list of numbers per example, called a feature vector. Raw data rarely arrives like that. It has prices in the hundreds of thousands, ages in years and names like “north side”. Feature engineering is the work of turning each of those raw values into numbers the model can learn from well.
Numbers usually get reshaped. Scaling squeezes a column into a small, standard range, such as 0 to 1. Clipping caps extreme values at a chosen limit. Bucketing groups values into ranges, so the model learns one setting (a weight) for “0 to 10 years old” and another for “11 to 30”. Words need more work. One-hot encoding gives each possible category its own slot in the list: a 1 in its slot, 0s everywhere else. Rare categories can share one catch-all slot.
You can also combine columns. A feature cross joins two categories or buckets, such as neighbourhood and age range, into one new feature. That lets a simple model notice that a pairing matters, not just each part on its own.
These rules are part of the model. Work them out from the training data only. Then apply them the same way to new data. Tools such as scikit-learn’s pipelines chain the steps to the model, so no step gets forgotten.
A model reads a list of numbers, not a spreadsheet row the way you do.
Follow one house listing as it becomes a model-ready list of numbers.
- 1 · scaleSqueeze a wide-ranging number, such as price, into a small standard range such as 0 to 1.
- 2 · bucketGroup a number, such as age, into ranges so the model can learn one weight per range.
- 3 · encodeTurn a category, such as a neighbourhood name, into a list with a single 1 and the rest 0s.
- 4 · assembleJoin every result into one list of numbers, the feature vector the model reads.
- 5 · reuseApply exactly the same rules again whenever the model makes a prediction.
The rules are part of the model: use the same ones for training and for predictions.
| Who | What they ask | What it works with |
|---|---|---|
| House-price app | “How should price, age and neighbourhood go into the model?” | Listing fields |
| Car-pricing team | “How do we feed paint colour to the model?” | A colour column |
| Shop planner | “Does temperature affect how many shoppers come in?” | Daily temperatures, grouped into ranges |
- It turns words and other non-number values into numbers a model can use.
- Scaling puts columns with very different ranges on a similar scale.
- Grouping into ranges and combining columns let simple models pick up patterns that aren't a straight line.
- Swapping each category for a plain index number is not enough, because the model treats those numbers as amounts.
- Combining two sparse columns makes an even sparser one, with many more positions.
- If the rules are worked out using test data, the model's scores will look better than they really are.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsNumerical data: How a model ingests data using feature vectors, Google for Developers · read 29 Sept 2026
- docsNumerical data: Normalization, Google for Developers · read 29 Sept 2026
- docsNumerical data: Binning, Google for Developers · read 29 Sept 2026
- docsCategorical data: Vocabulary and one-hot encoding, Google for Developers · read 29 Sept 2026
- docsCategorical data: Feature crosses, Google for Developers · read 29 Sept 2026
- docsCommon pitfalls and recommended practices, scikit-learn developers · read 29 Sept 2026