All capabilities · Annotation, quality & human feedback

Write labeling guidelines and evaluation rubrics

Turn a vague quality goal into a rubric with examples and edge cases that others can apply.

~8 focused hoursintermediate
Explore 3 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Turn a vague quality bar ('good customer support response') into scored, checkable criteria
  2. Provide 2-3 worked examples per score level so raters apply the rubric consistently
  3. Cover edge cases explicitly instead of leaving them to individual judgment
  4. Pilot a rubric on a small sample and revise it before rolling out to a full annotation team
  5. Balance rubric strictness with practical rater throughput
  6. Version and document rubric changes so historical ratings stay interpretable
  7. Write rubrics usable by both human raters and automated LLM-judge pipelines

Needs first: Evaluate and rank model responses (RLHF / preference data)

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

3 tools to explore

Google Docs

Plan & explain

Write a rubric, project story or decision brief that others can review and comment on.

Google Sheets

Data · Plan & explain

Build a scoring sheet, clean a small dataset or make assumptions visible in a simple model.

Label Studio

Data · Test

Label examples, compare annotations and export a dataset for review or evaluation.

Practices & references

  • Golden examples
  • Likert and pass/fail criteria
Practice

Rubric and golden set for grading AI customer-support replies

Take real customer-support questions from the Bitext support dataset and generate AI replies to them, deliberately including some weak ones. Write a four-criterion rubric — accuracy, helpfulness, tone, policy compliance — on a 1-5 scale with two worked examples per level, including a Hindi-language example you write. Score a 20-item golden set by hand with rationales, then point an LLM-as-judge at the same rubric and compare. The rubric is not done until piloting has forced you to change it at least once.

Start from

Bitext customer-support dataset on Hugging Face — ~27k real support queries across 27 intents; you generate the AI replies to grade

Milestones
  1. Draft the four criteria with anchors and two worked examples per score level · ~1h
  2. Generate the 20 replies (some deliberately weak) and score them by hand with rationales · ~2h
  3. Run the LLM-as-judge on the same rubric and compare score by score · ~1h
  4. Revise the rubric on what the pilot exposed and log the change · ~0.5h
Done when
  • Rubric has 4+ criteria, each with a 1-5 scale and at least 2 worked examples per level
  • A 20-item golden set is scored manually with rationales referencing specific rubric criteria
  • An LLM-as-judge is run against the same rubric on the golden set and results are compared to manual scores
  • A short revision log documents at least one rubric change made after piloting, and why
Prove it

Evidence a recruiter can check

  • The rubric with two worked examples at every score level of every criterion — the anchors are the artefact, not the criteria names
  • Manual-vs-LLM-judge agreement on the golden set, with the items where the judge and you diverged shown side by side
  • The 20-item golden set with a rationale per score naming the criterion it turned on
  • A revision log showing at least one rubric change the pilot forced, with the before and after wording
  • The Hindi-language worked example, showing the tone criterion holds up outside English
Signal it

Wrote a four-criterion 1-5 rubric with worked anchors for grading AI support replies, built a 20-item hand-scored golden set, and benchmarked an LLM-as-judge against it — revising the rubric where the pilot exposed ambiguity.

Interview

Questions you'll get asked

  1. Turn 'the response should be helpful' into a concrete, scoreable rubric with 3+ criteria.
  2. How do you pilot-test a new rubric before rolling it out to 20 annotators?
  3. What do you include in a rubric so two different raters reach the same score on a borderline case?
  4. How would you version a rubric so old ratings remain interpretable after you update it?
  5. How is a rubric for a human rater different from one designed for an LLM-as-judge pipeline?
  6. Give an example of a rubric criterion that sounded clear but caused disagreement in practice — how did you fix it?
  7. How do you balance a rubric being thorough vs. annotators being able to apply it quickly?