All capabilities · LLM application development

Design and version prompts systematically

Write system prompts, few-shot examples, output constraints; manage prompt versions and regressions.

~10 focused hoursbeginner
Explore 4 tools for this project
Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

4 tools to explore

Claude API

Build · Test

Build model-backed features with messages, tool use and responses you can evaluate.

OpenAI API

Build · Test

Connect model calls, tool use and structured responses to your own application.

Practices & references

  • System prompts
  • Few-shot examples
  • Regression evaluations
Practice

Prompt regression harness for a support-ticket classifier

Sample 300 rows from a public customer-support dataset, hand-label 30 of them for category, priority and sentiment, and build a prompt plus an eval script that classifies each ticket. Version at least three prompt iterations in git and score every version against the same labelled set, so the improvement is a number rather than a feeling. Add a handful of code-mixed Hindi-English tickets you write yourself so the prompt has to survive how Indian customers actually type.

Start from

Bitext customer-support intent dataset on Hugging Face - tens of thousands of labelled support utterances; sample 300 and hand-label 30 as your gold set

Milestones
  1. Sample 300 tickets and hand-label a 30-ticket gold set · ~1.5h
  2. Write prompt v1 and get parseable output on all 30 tickets · ~1.5h
  3. Build the eval script that scores accuracy and F1 per prompt version · ~2.5h
  4. Iterate to v2 and v3 on the failures, recording what changed and why · ~2h
Done when
  • At least 3 prompt versions committed to git with a changelog explaining what changed and why
  • An eval script reports accuracy/F1 per prompt version on the same 30-ticket labeled set
  • The final prompt handles at least 3 edge cases (code-mixed Hindi-English, ambiguous category, empty ticket) without crashing
  • README shows a before/after accuracy table across prompt versions
Prove it

Evidence a recruiter can check

  • An accuracy and F1 table per prompt version, all scored on the same 30-ticket gold set
  • The labelled gold set committed alongside the harness, so anyone can re-run the scores themselves
  • One ticket that v1 got wrong and v3 gets right, with both raw outputs pasted side by side
  • A note naming which technique - few-shot, chain-of-thought, format constraint - moved the score most, with the delta it produced
Signal it

Built a prompt regression harness for a support-ticket classifier - hand-labelled a gold set, versioned three prompt iterations in git, and lifted classification accuracy with every change scored rather than eyeballed.

Interview

Questions you'll get asked

  1. How do you structure a system prompt for a customer-support bot that must never discuss competitors?
  2. A prompt worked last week and started failing after a model update — how do you debug it?
  3. How do you decide between zero-shot, few-shot, and chain-of-thought prompting for a task?
  4. How do you version and test prompts the way you'd version and test code?
  5. Show me how you'd reduce hallucination in a prompt that summarizes financial documents.
  6. How would you prompt a model to always reply in Hindi-English code-mixed text for an Indian audience, without breaking JSON output?