All capabilities · Retrieval & knowledge systems

Generate embeddings and run vector search

Choose embedding models, store vectors (pgvector/Chroma/Pinecone), run similarity and hybrid search.

~10 focused hoursintermediate
Explore 4 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

4 tools to explore

pgvector

Data · Build

Store embeddings in PostgreSQL and compare similarity-search results.

OpenAI API

Build · Test

Connect model calls, tool use and structured responses to your own application.

Practices & references

  • Hybrid search
  • Retrieval relevance
Practice

Hybrid search over Indian consumer-tech product reviews

Embed a few thousand Flipkart product reviews into pgvector or Chroma, build a BM25 keyword index over the same corpus, and fuse the two with a re-ranking step. Test it on queries that need both halves at once - 'battery drains fast Redmi phone' needs the model name matched literally and the complaint matched semantically. Benchmark hybrid against vector-only and keyword-only on a query set you write, and measure top-k retrieval latency at the indexed corpus size.

Start from

Flipkart Products Review Dataset on Kaggle - product reviews with ratings and product metadata

Milestones
  1. Load and clean the reviews, then embed 500+ into the vector store · ~2.5h
  2. Add a BM25 index and a fusion step over both result sets · ~2.5h
  3. Write 10 realistic queries and score all three retrieval modes · ~1.5h
  4. Add re-ranking, measure top-k latency, justify the embedding model · ~1.5h
Done when
  • At least 500 reviews embedded and indexed in a vector store (pgvector or Chroma)
  • A BM25 keyword index runs alongside vector search, with a combined re-ranked result set
  • A test set of 10 realistic queries shows hybrid search outperforming vector-only or keyword-only on relevance
  • Query latency for top-k retrieval is measured and reported (target: sub-second on the test dataset)
Prove it

Evidence a recruiter can check

  • A relevance table scoring hybrid against vector-only and keyword-only on the same 10 queries
  • The query 'battery drains fast Redmi phone' with all three result sets pasted, showing exactly what each mode misses
  • Top-k retrieval latency at the indexed corpus size, with the corpus size stated
  • The embedding-model choice argued against at least one alternative you actually ran on the same queries
Signal it

Built hybrid retrieval over a few thousand product reviews - vector embeddings fused with a BM25 index and re-ranked - and showed it beating vector-only and keyword-only search on a hand-scored query set at sub-second latency.

Interview

Questions you'll get asked

  1. How do you choose between pgvector, Chroma, and a managed service like Pinecone for a given project?
  2. What's the difference between exact and approximate nearest-neighbor search, and when does it matter?
  3. How would you implement hybrid search combining BM25 keyword matching with vector similarity?
  4. How do you decide on embedding dimensionality and its cost/latency/accuracy tradeoffs?
  5. How would you keep a vector index in sync as underlying documents get updated or deleted?
  6. How do you evaluate whether your embedding model is actually retrieving relevant results?