All capabilities · Annotation, quality & human feedback

Audit label quality and compute agreement

Sampling, inter-annotator agreement, error taxonomies, feedback to annotators.

~8 focused hoursintermediate
Explore 3 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Design a sampling plan to audit label quality without re-reviewing every item
  2. Compute inter-annotator agreement (Cohen's kappa or similar) and interpret what the number means
  3. Build an error taxonomy that groups mistakes into actionable categories, not just 'wrong'
  4. Give specific, example-backed feedback to annotators that improves future accuracy
  5. Track quality metrics over time and flag a drifting annotator or a broken guideline early
  6. Distinguish a genuine annotator error from a guideline ambiguity that needs fixing upstream
  7. Report QA results to stakeholders in a way that supports a go/no-go decision on a dataset

Needs first: Label data accurately against guidelines

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

3 tools to explore

Google Sheets

Data · Plan & explain

Build a scoring sheet, clean a small dataset or make assumptions visible in a simple model.

Practices & references

  • Sampling plans
  • Inter-annotator agreement
  • Error taxonomies
Practice

Label-quality audit and inter-annotator agreement report

Use the GoEmotions raw release, where every Reddit comment carries labels from several named raters — real disagreement between real people, not simulated. Pick two raters with a decent overlap, take 150 shared items, compute Cohen's kappa and interpret it in plain language. Then do the part that matters: sort the disagreements into a named error taxonomy, write feedback an annotator could actually act on, and end with a go/no-go call on whether the dataset is ready to ship.

Start from

GoEmotions raw split on Hugging Face — 58k Reddit comments with per-rater labels, so two real annotators' overlapping items can be compared

Milestones
  1. Find two raters with 150+ overlapping items and build the paired label table · ~1h
  2. Compute Cohen's kappa per label and overall, and sanity-check it against raw agreement · ~2h
  3. Sort every disagreement into named error-taxonomy buckets with example items · ~2h
  4. Write the QA report with actionable feedback and the go/no-go call · ~2h
Done when
  • Cohen's kappa (or equivalent) is computed correctly and interpreted in plain language in the report
  • Disagreements are categorized into at least 4 named error-taxonomy buckets, not left as an undifferentiated list
  • Report includes at least 3 specific, example-backed feedback points an annotator could act on
  • Report ends with a clear go/no-go recommendation on whether the dataset is ready to ship
Prove it

Evidence a recruiter can check

  • Cohen's kappa computed per label and overall, with the labels the two raters diverge on most named and explained in plain language
  • The error taxonomy with named buckets and real example items in each — showing the disagreements were diagnosed, not just counted
  • Three example-backed feedback points written as instructions an annotator could follow tomorrow
  • A go/no-go recommendation on dataset readiness that commits to an answer and says what would change it
  • The paired-label extraction code, so a reader can check which raters and items the kappa was computed over
Signal it

Audited a multi-rater labelled dataset for quality — computed per-label Cohen's kappa on 150 overlapping items, diagnosed the disagreements into a named error taxonomy, and issued a go/no-go readiness call with actionable annotator feedback.

Interview

Questions you'll get asked

  1. How do you calculate Cohen's kappa and what does a kappa of 0.4 actually tell you?
  2. Design a sampling plan to audit 2% of a 100,000-item labeled dataset — how do you choose the sample?
  3. How do you tell whether a labeling error is the annotator's fault or the guideline's fault?
  4. Walk me through an error taxonomy you'd build for a text-classification annotation project.
  5. How would you give feedback to an annotator whose accuracy dropped this week?
  6. What's the difference between accuracy against gold labels and inter-annotator agreement, and when do you need each?
  7. How do you decide a dataset is 'good enough' to ship to the ML team?