All capabilities · Retrieval & knowledge systems

Build a grounded RAG application with citations

Retrieve → rerank → generate grounded answers with citations over a private corpus.

~20 focused hoursintermediate
Explore 5 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Build a retrieve → rerank → generate pipeline that answers questions grounded in a private corpus
  2. Return citations/source references alongside generated answers
  3. Handle 'I don't know' gracefully when the corpus doesn't contain an answer, instead of hallucinating
  4. Tune retrieval (top-k, chunk size, reranking) to improve answer quality
  5. Support multi-turn Q&A where follow-up questions still retrieve correctly
  6. Add a reranking step to improve precision after initial vector retrieval
  7. Deploy the RAG system behind an API or simple UI for others to query

Needs first: Generate embeddings and run vector search, Design and version prompts systematically

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

5 tools to explore

LlamaIndex

Build · Data

Connect documents to retrieval and model workflows with explicit data handling.

pgvector

Data · Build

Store embeddings in PostgreSQL and compare similarity-search results.

Claude API

Build · Test

Build model-backed features with messages, tool use and responses you can evaluate.

Practices & references

  • Reranking
  • Citations
  • Failure analysis
Practice

Grounded RAG assistant over Indian labour-law documents

Build a RAG assistant over India's labour codes and related public labour-law documents, answering questions like 'how many days of casual leave am I entitled to?' with citations down to the section. Implement retrieve, rerank, then generate, and make the assistant say the corpus does not cover it rather than invent an answer. Expose it as a small API or UI that handles multi-turn follow-ups, and score it on a question set you label by hand.

Start from

The four Indian labour codes and related public labour-law PDFs published by the Ministry of Labour & Employment

Milestones
  1. Ingest the labour-law PDFs with section metadata and get vector retrieval working · ~4h
  2. Wire retrieve into generate with inline citations back to section and page · ~3h
  3. Add the reranker and compare retrieved chunks before and after · ~3h
  4. Add out-of-corpus refusal and multi-turn follow-up handling · ~3h
  5. Hand-label 10 questions, score the answers, deploy or record a demo · ~2h
Done when
  • Answers include inline citations pointing to the specific source document and section/page
  • A reranking step measurably improves top-result relevance over vector-only retrieval, shown with before/after examples
  • At least 5 out-of-corpus questions are correctly answered with 'not covered' instead of a hallucinated answer
  • A simple API or UI lets a user ask multi-turn follow-up questions that retrieve correctly
Prove it

Evidence a recruiter can check

  • Ten answers with inline citations, each traceable to the section of the code it quotes
  • The same query's top chunks before and after reranking, placed side by side
  • Five out-of-corpus questions answered with an explicit 'not covered', pasted verbatim
  • A hand-labelled score (e.g. 8 of 10 correct-with-citation) with the labelled question set committed
  • A hosted link or demo recording where someone else can ask a question of their own
Signal it

Built a grounded RAG assistant over India's labour codes - retrieve, rerank and generate with section-level citations - scored on a hand-labelled question set and refusing out-of-corpus questions instead of hallucinating.

Interview

Questions you'll get asked

  1. Walk me through the architecture of a RAG pipeline you've built, retrieval to generation.
  2. How do you get an LLM to say 'I don't know' instead of hallucinating when the answer isn't in the retrieved context?
  3. What's the role of a reranker in a RAG pipeline, and when is it worth the added latency?
  4. How do you generate accurate citations that point back to the exact source chunk?
  5. How would you handle a follow-up question in a multi-turn RAG conversation where the retrieval query needs the prior turn's context?
  6. How do you decide top-k for retrieval, and what happens if it's too low or too high?
  7. How would you extend a RAG system to work well over a mix of English and Hindi documents?