LangSmith
Test · Monitor
Inspect model traces and compare outputs against an evaluation dataset.
Choose offline/online metrics, design human-rating rubrics, set launch bars for AI features.
Explore 4 tools for this projectShare of job postings in India, per role, that name this capability.
Needs first: Explain how LLMs work and where they fail
Start with one tool for each part of your project. You don’t need to learn them all.
4 tools to explore
Test · Monitor
Inspect model traces and compare outputs against an evaluation dataset.
Monitor · Test
Trace model calls and review prompts, latency, costs and evaluation results.
Data · Test
Label examples, compare annotations and export a dataset for review or evaluation.
Data · Plan & explain
Build a scoring sheet, clean a small dataset or make assumptions visible in a simple model.
Pull Hindi and English utterances from the MASSIVE dataset — real user requests already labelled with intent — and cut a 100-example eval set across the categories a support desk cares about. Build an LLM classifier over it and score it two ways: an automated metric (accuracy and per-class F1) and a human rubric for the edge cases the metric cannot see. Run 2-3 prompt variants through the same harness, then write the recommendation and the launch bar you would hold it to.
Built a 100-example Hindi/English intent eval set and harness, compared three prompt variants on per-class F1 plus a human rubric, and set the launch bar that decided which one shipped.