All capabilities · Product, business & communication

Define quality metrics and eval plans for AI features

Choose offline/online metrics, design human-rating rubrics, set launch bars for AI features.

~10 focused hoursintermediate
Explore 4 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Choose offline metrics (accuracy, groundedness, exact/fuzzy match) appropriate to the task
  2. Design a human-rating rubric that two raters apply consistently
  3. Set an online metric and a launch bar tied to business outcome, not just model score
  4. Build a small labeled eval set from real or representative data before launch
  5. Track eval scores across prompt/model versions to justify a change
  6. Explain the gap between an offline eval score and real user satisfaction
  7. Design a human-in-the-loop review process for low-confidence outputs

Needs first: Explain how LLMs work and where they fail

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

4 tools to explore

Label Studio

Data · Test

Label examples, compare annotations and export a dataset for review or evaluation.

Google Sheets

Data · Plan & explain

Build a scoring sheet, clean a small dataset or make assumptions visible in a simple model.

Practice

Eval harness for a Hindi/English support-intent classifier

Pull Hindi and English utterances from the MASSIVE dataset — real user requests already labelled with intent — and cut a 100-example eval set across the categories a support desk cares about. Build an LLM classifier over it and score it two ways: an automated metric (accuracy and per-class F1) and a human rubric for the edge cases the metric cannot see. Run 2-3 prompt variants through the same harness, then write the recommendation and the launch bar you would hold it to.

Start from

MASSIVE (AmazonScience) — 60-intent labelled utterances across 51 languages, including hi-IN and en-US

Milestones
  1. Cut a 100-example set from MASSIVE hi-IN and en-US and write the labelling guideline · ~1.5h
  2. Build the classifier and the scoring script (accuracy + per-class F1) · ~2h
  3. Write the human rubric and apply it to 20 cases the metric and you disagree on · ~1.5h
  4. Run three prompt variants, tabulate, and write the launch-bar recommendation · ~1.5h
Done when
  • Eval set has at least 50 labeled examples with a documented labeling guideline
  • Results are reported for at least 2 prompt/model variants with a comparison table
  • A human-rating rubric (not just automated accuracy) is defined and applied to a sample of outputs
  • A written recommendation states the launch bar and whether the current best variant meets it
Prove it

Evidence a recruiter can check

  • A results table comparing three prompt variants on the same 100 examples with per-class F1, and the variant you would ship marked
  • The labelling guideline and the eval set itself, so a stranger can re-run your numbers and get them
  • A human rubric applied to 20 disagreements, with a note on what the automated metric was blind to
  • A written launch bar and an honest statement of whether your best variant clears it
Signal it

Built a 100-example Hindi/English intent eval set and harness, compared three prompt variants on per-class F1 plus a human rubric, and set the launch bar that decided which one shipped.

Interview

Questions you'll get asked

  1. How would you evaluate a customer-support chatbot before launch — walk me through your plan.
  2. What's the difference between an offline eval and an online (in-production) metric, and when do you need both?
  3. How do you build a rubric that two different human raters will score consistently?
  4. Tell me about a launch bar you set for an AI feature — how did you pick the threshold?
  5. How would you detect that a model update silently made quality worse?
  6. What metrics would you track for a RAG system beyond 'did it answer'?