scikit-learn
Data · Test
Build baseline models and evaluate them with consistent train/test splits.
Collaborative filtering, embeddings-based retrieval, offline ranking metrics.
Explore 3 tools for this projectShare of job postings in India, per role, that name this capability.
Needs first: Train and evaluate classical ML models
Start with one tool for each part of your project. You don’t need to learn them all.
3 tools to explore
Data · Test
Build baseline models and evaluate them with consistent train/test splits.
Data · Test
Build a recommendation model and evaluate candidate retrieval and ranking.
Data
Index dense vectors and compare nearest-neighbour retrieval methods.
Using MovieLens ml-latest-small — 100,836 ratings from 610 users across 9,742 titles — build the two halves of a real recommender rather than one flat model: a collaborative-filtering candidate generator, then a ranking step over its output. Add a content-based fallback so items with no interaction history can still surface. Split by time, not at random, and evaluate with precision@k and NDCG@k against a most-popular baseline, which is a stronger opponent than most people expect.
MovieLens ml-latest-small — 100,836 ratings from 610 users across 9,742 titles, with tags and genres
Built a two-stage recommender on MovieLens — collaborative-filtering retrieval with a content-based cold-start fallback, evaluated with precision@k and NDCG@k on a time-based split and beaten against a popularity baseline.