Concepts

One-hot encoding

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

One-hot encoding turns a category, such as a colour or a type of transport, into a list of 0s with a single 1 marking which category it is.

1 · What it is

One-hot encoding is how a machine learning model reads a category. Models can only do arithmetic on floating-point numbers, so a column holding words such as bus, car and train has to be converted first. The trick is to give every possible category its own position in a list. A row then gets a 1 in the position for its category and a 0 everywhere else.

Why not just number the categories 0, 1, 2 and 3? Because a model treats numbers as amounts, and categories have no natural order. It would read car as 2 and bus as 1, as though a car were twice a bus. Google’s course gives the same warning with colours: left as plain indexes, purple would count as six times orange. The model then learns a separate weight for every position, so it can discover that one category matters more than another.

Most of each vector is zeros, so tools often store only where the 1 sits. This sparse form takes far less memory, although the model still trains on the full vector. In scikit-learn, OneHotEncoder builds the vocabulary from the training data. In pandas, get_dummies does the same job, and PyTorch has a one_hot function for index numbers.

The catch is size. A one-hot vector is as long as the vocabulary, and Google’s course lists about 42,000 US postal codes and roughly 500,000 English words as examples of such features. At that scale one-hot is usually a poor choice: a neural network needs a huge number of weights, and training costs more. Embeddings are the usual fix, turning each category into a short list of learned numbers. They also fix a second flaw: one-hot treats every pair of categories as equally different, and the word2vec paper notes that words stored as vocabulary indexes carry no notion of similarity. Hashing is a less common shortcut that folds categories into a fixed number of buckets, at the cost of unrelated values sometimes sharing one.

New values cause trouble too. A category that never appeared in training has no position. By default scikit-learn stops with an error; told to ignore it, the encoder writes all zeros. Keras’s StringLookup layer instead reserves an out-of-vocabulary slot at the front, and rare values can be lumped into that same bucket. Some tools can also drop one column per feature, which helps an unregularized linear regression but can bias other models.

2 · Why it exists

Models need numbers, but categories are names.

Models only read numbersA model can only work with floating-point numbers, so words like "car" or "Blue" must be converted before training.
Plain codes invent an orderIf categories are simply numbered 0, 1, 2, a model treats those numbers as amounts, as if one colour were six times another.
Each category needs its own weightA car-price model should be free to learn that red cars sell for more than green ones, which needs a separate weight per colour.
3 · How it works

Follow one column of transport modes into numbers.

Illustrative data. Each column is one category from the vocabulary; the model reads the rows of 0s and 1s, never the words.
  1. 1 · collectList the categories found in the training data and give each one a fixed index number, which together form the vocabulary.
  2. 2 · sizeMake a vector with one position per category, so four categories give four positions.
  3. 3 · setFor each row, put a 1 at that category's position and a 0 in every other position.
  4. 4 · learnEach position becomes its own feature, and the model learns a separate weight for it.

The model never sees the word car. It sees [0, 0, 1, 0] and learns one weight per position.

4 · Where it's used
WhoWhat they askWhat it works with
Car marketplace“Should paint colour go in as one number or as one column per colour?”The colour of each listed car
Delivery app“Does the day of the week change how many orders we get?”The weekday each order was placed
Payments team“What happens when a country we have never seen shows up after launch?”The country on each payment
Language team“Can we give every English word its own column?”A vocabulary of words from the training text
5 · What it solves, and what it doesn't
solves
  • Lets a model read categories without inventing an order or size between them.
  • Gives the model a separate weight for each category.
  • Stores cheaply, since a sparse representation keeps only the position of the 1.
  • Comes built into common tools, such as scikit-learn's OneHotEncoder.
doesn't solve
  • Vectors grow with the number of categories, so tens of thousands of postcodes make one-hot a poor choice.
  • Every pair of categories looks equally different, so hot dog is no closer to shawarma than to salad.
  • A category never seen in training gets no column of its own; it becomes an error, all zeros or a shared unknown slot.
  • Hashing can shrink the vector, but unrelated categories can then collide in the same column.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsCategorical data: Vocabulary and one-hot encoding (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  2. docsCategorical data: Common issues (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  3. docsEmbeddings (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  4. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  5. docsOneHotEncoder, scikit-learn · read 27 Sept 2026
  6. docsPreprocessing data: Encoding categorical features, scikit-learn · read 27 Sept 2026
  7. docsFeature extraction: Feature hashing, scikit-learn · read 27 Sept 2026
  8. docspandas.get_dummies, pandas · read 27 Sept 2026
  9. docsStringLookup layer, Keras · read 27 Sept 2026
  10. docstorch.nn.functional.one_hot, PyTorch · read 27 Sept 2026
  11. paperEfficient Estimation of Word Representations in Vector Space, Mikolov et al., arXiv 2013 · read 27 Sept 2026