All capabilities · Agents & workflows

Evaluate and harden agents

Trajectory evals, cost/latency budgets, failure taxonomies, sandboxing and permissioning.

~12 focused hoursadvanced
Explore 3 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Score an agent's full trajectory (sequence of tool calls and reasoning steps), not just the final answer
  2. Build a failure taxonomy (wrong tool chosen, hallucinated argument, infinite loop, unsafe action) and tag failures against it
  3. Set and enforce cost and latency budgets per agent run, failing or degrading gracefully when exceeded
  4. Sandbox agent tool execution (filesystem, network, shell) so a misbehaving agent can't cause real damage
  5. Add permissioning/approval gates for high-risk tool calls and test that they actually block unauthorized actions
  6. Write regression tests that catch when a prompt or model change breaks a previously-working agent trajectory
  7. Run agents against adversarial or edge-case inputs to find where the loop breaks

Needs first: Build a multi-step agent workflow, Build an LLM evaluation harness

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

3 tools to explore

Practices & references

  • Trajectory and step-level scoring
  • Cost and latency budgets
Practice

Trajectory eval harness and CI gate for a triage agent

Take the document-triage agent from the agent-workflow project and build a harness that replays recorded runs and scores the whole trajectory — which tools were called, in what order, and whether any unsafe action slipped through — not just the final answer. The inputs are runs you record yourself: freeze 20 traces as golden references, four of them deliberately bad. Score each replay against a written failure taxonomy and enforce a per-run token, cost and latency budget. Fail the CI run when trajectory accuracy drops below your threshold or any run blows its budget.

Start from

20 runs of your own triage agent, recorded and frozen as JSON trajectories (tool calls, arguments, order, outcome) — four of them deliberately broken

Milestones
  1. Record and freeze 20 agent runs as golden trajectories, including four you know go wrong · ~2.5h
  2. Write the failure taxonomy and a scorer that tags wrong-tool, wrong-order and unsafe-action separately · ~3.5h
  3. Add per-run token/cost and latency budgets that fail a run when exceeded · ~2h
  4. Wire the harness into CI with a per-category pass/fail summary · ~2h
  5. Break the agent on purpose and confirm the harness catches it · ~1h
Done when
  • 20+ recorded trajectories with expected tool-call sequences as golden references
  • A scoring function that flags wrong-tool, wrong-order, and unsafe-action failures separately (a failure taxonomy, not just pass/fail)
  • A hard cost/latency budget enforced per run, with runs that exceed it marked as failed
  • The harness runs in CI (or a script) and produces a pass/fail summary with per-category failure counts
Prove it

Evidence a recruiter can check

  • The per-category report from one harness run — counts for wrong-tool, wrong-order, unsafe-action and budget-exceeded — pasted in the README
  • The failure taxonomy document, with a real trajectory from your runs illustrating each category
  • A before/after CI run showing the harness catching a regression you introduced on purpose, then going green after the fix
  • The 20 golden trajectories committed to the repo, so anyone can reproduce your scores without your API key
Signal it

Built a trajectory-level eval harness for an agent — 20 golden runs scored against a four-category failure taxonomy with per-run cost and latency budgets, wired into CI so a prompt change that breaks tool selection fails the build.

Interview

Questions you'll get asked

  1. How do you evaluate an agent when there's no single 'correct' final answer, only a good-enough trajectory?
  2. Describe a failure taxonomy you'd build for a customer-support agent, and how you'd tag failures automatically.
  3. How do you sandbox an agent that can run shell commands or hit the filesystem?
  4. How would you set and enforce a per-request cost budget for an agentic workflow that can call tools recursively?
  5. What's the difference between step-level and outcome-level evaluation for agents?
  6. How do you catch an agent that gets stuck in a loop before it burns through your token budget?
  7. Tell me about a time an agent misbehaved in a way your evals didn't catch. What did you change?