One-hot encoding
One-hot encoding turns a category, such as a colour or a type of transport, into a list of 0s with a single 1 marking which category it is.
One-hot encoding is how a machine learning model reads a category. Models can only do arithmetic on floating-point numbers, so a column holding words such as bus, car and train has to be converted first. The trick is to give every possible category its own position in a list. A row then gets a 1 in the position for its category and a 0 everywhere else.
Why not just number the categories 0, 1, 2 and 3? Because a model treats numbers as amounts, and categories have no natural order. It would read car as 2 and bus as 1, as though a car were twice a bus. Google’s course gives the same warning with colours: left as plain indexes, purple would count as six times orange. The model then learns a separate weight for every position, so it can discover that one category matters more than another.
Most of each vector is zeros, so tools often store only where the 1 sits. This sparse form takes far less memory, although the model still trains on the full vector. In scikit-learn, OneHotEncoder builds the vocabulary from the training data. In pandas, get_dummies does the same job, and PyTorch has a one_hot function for index numbers.
The catch is size. A one-hot vector is as long as the vocabulary, and Google’s course lists about 42,000 US postal codes and roughly 500,000 English words as examples of such features. At that scale one-hot is usually a poor choice: a neural network needs a huge number of weights, and training costs more. Embeddings are the usual fix, turning each category into a short list of learned numbers. They also fix a second flaw: one-hot treats every pair of categories as equally different, and the word2vec paper notes that words stored as vocabulary indexes carry no notion of similarity. Hashing is a less common shortcut that folds categories into a fixed number of buckets, at the cost of unrelated values sometimes sharing one.
New values cause trouble too. A category that never appeared in training has no position. By default scikit-learn stops with an error; told to ignore it, the encoder writes all zeros. Keras’s StringLookup layer instead reserves an out-of-vocabulary slot at the front, and rare values can be lumped into that same bucket. Some tools can also drop one column per feature, which helps an unregularized linear regression but can bias other models.
Models need numbers, but categories are names.
Follow one column of transport modes into numbers.
- 1 · collectList the categories found in the training data and give each one a fixed index number, which together form the vocabulary.
- 2 · sizeMake a vector with one position per category, so four categories give four positions.
- 3 · setFor each row, put a 1 at that category's position and a 0 in every other position.
- 4 · learnEach position becomes its own feature, and the model learns a separate weight for it.
The model never sees the word car. It sees [0, 0, 1, 0] and learns one weight per position.
| Who | What they ask | What it works with |
|---|---|---|
| Car marketplace | “Should paint colour go in as one number or as one column per colour?” | The colour of each listed car |
| Delivery app | “Does the day of the week change how many orders we get?” | The weekday each order was placed |
| Payments team | “What happens when a country we have never seen shows up after launch?” | The country on each payment |
| Language team | “Can we give every English word its own column?” | A vocabulary of words from the training text |
- Lets a model read categories without inventing an order or size between them.
- Gives the model a separate weight for each category.
- Stores cheaply, since a sparse representation keeps only the position of the 1.
- Comes built into common tools, such as scikit-learn's OneHotEncoder.
- Vectors grow with the number of categories, so tens of thousands of postcodes make one-hot a poor choice.
- Every pair of categories looks equally different, so hot dog is no closer to shawarma than to salad.
- A category never seen in training gets no column of its own; it becomes an error, all zeros or a shared unknown slot.
- Hashing can shrink the vector, but unrelated categories can then collide in the same column.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsCategorical data: Vocabulary and one-hot encoding (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsCategorical data: Common issues (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsEmbeddings (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsOneHotEncoder, scikit-learn · read 27 Sept 2026
- docsPreprocessing data: Encoding categorical features, scikit-learn · read 27 Sept 2026
- docsFeature extraction: Feature hashing, scikit-learn · read 27 Sept 2026
- docspandas.get_dummies, pandas · read 27 Sept 2026
- docsStringLookup layer, Keras · read 27 Sept 2026
- docstorch.nn.functional.one_hot, PyTorch · read 27 Sept 2026
- paperEfficient Estimation of Word Representations in Vector Space, Mikolov et al., arXiv 2013 · read 27 Sept 2026