LangSmith
Test · Monitor
Inspect model traces and compare outputs against an evaluation dataset.
Trajectory evals, cost/latency budgets, failure taxonomies, sandboxing and permissioning.
Explore 3 tools for this projectShare of job postings in India, per role, that name this capability.
Needs first: Build a multi-step agent workflow, Build an LLM evaluation harness
Start with one tool for each part of your project. You don’t need to learn them all.
3 tools to explore
Test · Monitor
Inspect model traces and compare outputs against an evaluation dataset.
Test
Add model-output checks and evaluation metrics to a repeatable test suite.
Test
Turn expected behaviour and failure cases into repeatable Python tests.
Take the document-triage agent from the agent-workflow project and build a harness that replays recorded runs and scores the whole trajectory — which tools were called, in what order, and whether any unsafe action slipped through — not just the final answer. The inputs are runs you record yourself: freeze 20 traces as golden references, four of them deliberately bad. Score each replay against a written failure taxonomy and enforce a per-run token, cost and latency budget. Fail the CI run when trajectory accuracy drops below your threshold or any run blows its budget.
20 runs of your own triage agent, recorded and frozen as JSON trajectories (tool calls, arguments, order, outcome) — four of them deliberately broken
Built a trajectory-level eval harness for an agent — 20 golden runs scored against a four-category failure taxonomy with per-run cost and latency budgets, wired into CI so a prompt change that breaks tool selection fails the build.