All capabilities · Machine learning & data science

Build a recommendation or ranking model

Collaborative filtering, embeddings-based retrieval, offline ranking metrics.

~15 focused hoursintermediate
Explore 3 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

No job posting we have read names this yet. It stays on the map so the picture is complete; it will not appear on a path until a role asks for it.
What employers mean

You should be able to…

  1. Build a collaborative filtering model (user-item matrix factorization) from interaction data
  2. Build a content-based / embeddings-based retrieval component for cold-start items
  3. Combine retrieval (candidate generation) with ranking, not just one flat model
  4. Evaluate with offline ranking metrics (precision@k, recall@k, NDCG, MAP) not just RMSE
  5. Handle the cold-start problem for new users or new items
  6. Explain the exploration/exploitation and popularity-bias tradeoffs in a recommender
  7. Design an offline evaluation that reasonably approximates online (A/B) performance

Needs first: Train and evaluate classical ML models

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

3 tools to explore

TensorFlow Recommenders

Data · Test

Build a recommendation model and evaluate candidate retrieval and ranking.

Practice

Two-stage recommender with cold-start fallback on MovieLens

Using MovieLens ml-latest-small — 100,836 ratings from 610 users across 9,742 titles — build the two halves of a real recommender rather than one flat model: a collaborative-filtering candidate generator, then a ranking step over its output. Add a content-based fallback so items with no interaction history can still surface. Split by time, not at random, and evaluate with precision@k and NDCG@k against a most-popular baseline, which is a stronger opponent than most people expect.

Start from

MovieLens ml-latest-small — 100,836 ratings from 610 users across 9,742 titles, with tags and genres

Milestones
  1. Build the time-ordered split and the most-popular baseline with precision@k / NDCG@k · ~2.5h
  2. Train the collaborative-filtering candidate generator (ALS or matrix factorisation) · ~2.5h
  3. Add the content-based cold-start fallback and a ranking pass over the candidates · ~2h
  4. Sweep k, compare against the baseline, and write the offline-vs-online caveats · ~2h
Done when
  • Collaborative filtering model trained on user-item interactions with a documented train/test time split (not random split, to avoid leakage)
  • Cold-start fallback implemented for items with no interaction history
  • Precision@k and NDCG@k reported and compared against a most-popular baseline
  • README explains why the chosen offline metrics were used and what an online A/B test would additionally need to check
Prove it

Evidence a recruiter can check

  • The evaluation table: precision@k and NDCG@k for the popularity baseline, the CF model and the reranked two-stage output
  • The split code with the cutoff timestamp visible — the thing that separates an honest recsys eval from a leaky one
  • Five sampled users' top-5 recommendations next to their real watch history, including one obviously bad case you left in
  • A cold-start demonstration: a title with zero interactions still getting recommended, and the content features that got it there
Signal it

Built a two-stage recommender on MovieLens — collaborative-filtering retrieval with a content-based cold-start fallback, evaluated with precision@k and NDCG@k on a time-based split and beaten against a popularity baseline.

Interview

Questions you'll get asked

  1. Walk me through how you'd design a 'customers who bought this also bought' feature end to end
  2. How is a two-tower retrieval model different from a simple matrix factorization model?
  3. How do you handle the cold-start problem for a brand-new user with no history?
  4. Why is RMSE often a poor metric for evaluating a recommender, and what would you use instead?
  5. What's the difference between candidate generation and ranking in a recommendation pipeline?
  6. How would you measure and reduce popularity bias in your recommendations?
  7. How do you evaluate a recsys model offline before an A/B test, and what are the risks of trusting offline metrics alone?