Ragas
Test
Evaluate a retrieval workflow with questions, reference answers and quality metrics.
Build eval sets, measure recall/faithfulness (RAGAS-style), tune chunking, reranking, hybrid search.
Explore 3 tools for this projectShare of job postings in India, per role, that name this capability.
Needs first: Build a grounded RAG application with citations
Start with one tool for each part of your project. You don’t need to learn them all.
3 tools to explore
Test
Evaluate a retrieval workflow with questions, reference answers and quality metrics.
Test · Monitor
Inspect model traces and compare outputs against an evaluation dataset.
Test
Add model-output checks and evaluation metrics to a repeatable test suite.
Take the corpus and pipeline from your RAG project and hand-label 25 question / answer / source triples over it. Build a RAGAS-based (or equivalent) harness that computes faithfulness, answer relevance and context recall automatically instead of by eye. Run the same corpus through three different chunking strategies, report every metric per strategy, and prove the harness actually works by deliberately worsening a config and watching the scores fall.
Your own RAG corpus and pipeline from a previous project, plus 25 question/answer/source triples you hand-label over it
Built a RAGAS-based eval harness for a RAG pipeline over a hand-labelled question set - compared three chunking strategies on faithfulness, relevance and context recall, and verified the harness catches a deliberately introduced regression.