Concepts

Collaborative filtering

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Collaborative filtering recommends items by learning from what many people rated, clicked or watched, then predicting the gaps in your own history.

1 · What it is

Collaborative filtering predicts what one person will like from what other people did. All it needs is a record of past behaviour, like purchases or star ratings. That behaviour goes into a grid called the feedback matrix, with a row for each user and a column for each item. Feedback is explicit when people give a numerical rating. It is implicit when the system infers interest from behaviour, such as someone watching a film. The MovieLens benchmark from GroupLens Research holds 32 million ratings of 87,585 movies by 200,948 users. Even so, most people never rate most films, so nearly all cells stay empty. Content-based filtering is the opposite, matching item traits to what one user already liked.

The earliest methods look for neighbours. User-based methods fill a blank by looking at how people with similar taste scored that item. Item-based methods instead look at how you yourself scored items that resemble it. At Amazon, two products count as related when shoppers who buy the first are unusually likely to pick up the second. Simply counting co-purchases would push a few bestsellers, such as trash bags, to the top for everyone. Item-based picks are also easier to justify to the person getting them.

Matrix factorisation takes a different route and gives every user and every item a short vector of numbers. A predicted rating is the inner product, or dot product, of the user’s vector and the item’s vector. These vectors are embeddings on one shared map, where a user lands near the items they tend to like. Nobody labels the numbers by hand, because training learns them from the matrix. For star ratings, a common recipe fits only the cells that have ratings and adds a regularisation penalty so the model does not overfit. Two common ways to fit the vectors are stochastic gradient descent and weighted alternating least squares. Alternating least squares starts from random vectors, fixes the user vectors while solving for the item vectors, then swaps. Apache Spark’s MLlib library learns these latent factors with alternating least squares. By default it uses 10 latent factors.

Clicks, views and purchases say a lot about what someone likes but little about what they dislike. One answer came from researchers building a recommender for TV shows. Their paper read each observation as a like-or-not signal paired with its own confidence level, not as a rating. It also fitted every user and item pair, not only the observed ones. Spark uses that approach when its implicit feedback option is switched on.

2 · Why it exists

Recommending from item descriptions alone runs into three problems.

Descriptions are costlyTo describe every item you need extra data about it, and that data can be missing or hard to gather.
Taste is hard to labelPart of why people enjoy something is subtle, and it rarely fits neatly into a list of item traits.
No help from othersContent-based filtering judges each person alone, so it never borrows signals from anyone else.
3 · How it works

Follow one empty cell from the grid to a recommendation.

Illustrative numbers. Each user and each film gets a short learned vector, and their dot product fills an empty cell.
  1. 1 · collectPut the feedback in a grid, with one row per user and one column per item.
  2. 2 · learnTraining learns a short vector for every user and every item, so their dot products match the known cells.
  3. 3 · predictFor an empty cell, the prediction is the dot product of that user's vector and that item's vector.
  4. 4 · rankRecommend each user the handful of items with the top scores.

Similar people end up with nearby vectors, so a film your look-alikes loved scores high for you.

4 · Where it's used
WhoWhat they askWhat it works with
Streaming service team“Which show should this viewer see next?”Play counts from all viewers
Online shop“What else do people who bought this kettle tend to buy?”Purchase histories, compared item by item
Music app“Which songs will suit this new listener's taste?”Plays and skips from listeners with similar habits
Student learning recommenders“How close can my model get to ratings it has not seen?”A public MovieLens rating set
5 · What it solves, and what it doesn't
solves
  • The feedback grid is enough on its own; nobody has to write item features by hand.
  • It can suggest something new to you because people with similar tastes liked it.
  • Two tables of short vectors take less room than the full grid.
  • In alternating least squares, each half-step has an exact answer and can be split across many machines.
doesn't solve
  • It cannot score an item or user that has no history yet, which is called the cold start problem.
  • Hit items drift into nearly every list, and scoring by dot product makes that worse.
  • Adding side features, such as a user's country or a film's genre, takes a more complex model.
  • With clicks and plays, a missing entry may mean the person never saw the item, not that they disliked it.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsCollaborative filtering, Google for Developers, Recommendation Systems course · read 27 Sept 2026
  2. docsMatrix factorization, Google for Developers, Recommendation Systems course · read 27 Sept 2026
  3. docsCollaborative filtering advantages & disadvantages, Google for Developers, Recommendation Systems course · read 27 Sept 2026
  4. docsContent-based filtering, Google for Developers, Recommendation Systems course · read 27 Sept 2026
  5. docsDeep neural network models, Google for Developers, Recommendation Systems course · read 27 Sept 2026
  6. paperCollaborative Filtering for Implicit Feedback Datasets, Hu, Koren and Volinsky (author copy) · read 27 Sept 2026
  7. docsCollaborative Filtering (Spark 4.2.0 documentation), Apache Spark · read 27 Sept 2026
  8. officialThe history of Amazon's recommendation algorithm, Amazon Science · read 27 Sept 2026
  9. officialMovieLens, GroupLens Research · read 27 Sept 2026
  10. docsRetrieval, Google for Developers, Recommendation Systems course · read 27 Sept 2026