Concepts

Feature engineering

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Feature engineering turns raw data, like prices and colour names, into the lists of numbers a model can learn from.

1 · What it is

A model does not read a table the way you do. It reads one list of numbers per example, called a feature vector. Raw data rarely arrives like that. It has prices in the hundreds of thousands, ages in years and names like “north side”. Feature engineering is the work of turning each of those raw values into numbers the model can learn from well.

Numbers usually get reshaped. Scaling squeezes a column into a small, standard range, such as 0 to 1. Clipping caps extreme values at a chosen limit. Bucketing groups values into ranges, so the model learns one setting (a weight) for “0 to 10 years old” and another for “11 to 30”. Words need more work. One-hot encoding gives each possible category its own slot in the list: a 1 in its slot, 0s everywhere else. Rare categories can share one catch-all slot.

You can also combine columns. A feature cross joins two categories or buckets, such as neighbourhood and age range, into one new feature. That lets a simple model notice that a pairing matters, not just each part on its own.

These rules are part of the model. Work them out from the training data only. Then apply them the same way to new data. Tools such as scikit-learn’s pipelines chain the steps to the model, so no step gets forgotten.

2 · Why it exists

A model reads a list of numbers, not a spreadsheet row the way you do.

Words aren't numbersA model can only work with numbers, so a word like "red" has to be turned into numbers first.
Big numbers shoutIf one column runs into the millions and another stays small, the model tends to lean far too hard on the big one.
Patterns hideSome patterns only show up after you group values into ranges or combine two columns.
3 · How it works

Follow one house listing as it becomes a model-ready list of numbers.

Each raw column gets its own rule, then the results are joined into one list of numbers.
  1. 1 · scaleSqueeze a wide-ranging number, such as price, into a small standard range such as 0 to 1.
  2. 2 · bucketGroup a number, such as age, into ranges so the model can learn one weight per range.
  3. 3 · encodeTurn a category, such as a neighbourhood name, into a list with a single 1 and the rest 0s.
  4. 4 · assembleJoin every result into one list of numbers, the feature vector the model reads.
  5. 5 · reuseApply exactly the same rules again whenever the model makes a prediction.

The rules are part of the model: use the same ones for training and for predictions.

4 · Where it's used
WhoWhat they askWhat it works with
House-price app“How should price, age and neighbourhood go into the model?”Listing fields
Car-pricing team“How do we feed paint colour to the model?”A colour column
Shop planner“Does temperature affect how many shoppers come in?”Daily temperatures, grouped into ranges
5 · What it solves, and what it doesn't
solves
  • It turns words and other non-number values into numbers a model can use.
  • Scaling puts columns with very different ranges on a similar scale.
  • Grouping into ranges and combining columns let simple models pick up patterns that aren't a straight line.
doesn't solve
  • Swapping each category for a plain index number is not enough, because the model treats those numbers as amounts.
  • Combining two sparse columns makes an even sparser one, with many more positions.
  • If the rules are worked out using test data, the model's scores will look better than they really are.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsNumerical data: How a model ingests data using feature vectors, Google for Developers · read 29 Sept 2026
  2. docsNumerical data: Normalization, Google for Developers · read 29 Sept 2026
  3. docsNumerical data: Binning, Google for Developers · read 29 Sept 2026
  4. docsCategorical data: Vocabulary and one-hot encoding, Google for Developers · read 29 Sept 2026
  5. docsCategorical data: Feature crosses, Google for Developers · read 29 Sept 2026
  6. docsCommon pitfalls and recommended practices, scikit-learn developers · read 29 Sept 2026