All capabilities · Product, business & communication

Explain how LLMs work and where they fail

Tokens, context, hallucination, RAG vs fine-tuning, cost drivers — well enough to make decisions with engineers.

~8 focused hoursbeginner
Explore 4 tools for this project
What employers mean

You should be able to…

  1. Explain tokens, context windows, and latency/cost trade-offs to a non-technical stakeholder
  2. Explain why an LLM hallucinates and how grounding/RAG reduces (not eliminates) it
  3. Compare RAG vs fine-tuning vs plain prompting for a given business problem and justify the choice
  4. Read a model provider's pricing page and estimate monthly API cost for a proposed feature
  5. Explain embeddings and vector search well enough for an engineer to trust your product calls
  6. Spot when a demo answer is confidently wrong and explain why in plain language
  7. Translate a vague 'add AI to the product' ask into a scoped, buildable capability
Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

4 tools to explore

Google AI Studio

Build

Test prompts and model inputs before turning a feasibility experiment into code.

OpenAI API

Build · Test

Connect model calls, tool use and structured responses to your own application.

Claude API

Build · Test

Build model-backed features with messages, tool use and responses you can evaluate.

tiktoken

Build

Count tokens and compare how prompt changes affect the input budget.

Practice

Explainer deck and cost model for a Hindi/Hinglish support bot

Pick a real Indian scenario — a lender's helpline fielding Hindi and Hinglish loan-repayment questions — and write a one-page, jargon-free explainer of how an LLM-based bot would answer it, where it goes wrong, and how grounding it in the lender's published policy documents changes the answer. Pair that with a spreadsheet estimating monthly token cost at three volume tiers (1k / 10k / 100k conversations) using a provider's actual published per-token prices. Present both in a 10-minute walkthrough recorded as if to a non-technical founder.

Start from

Anthropic's and OpenAI's published API pricing pages — the per-token prices and context limits your cost model is built from

Milestones
  1. Tokenise 20 real Hindi/Hinglish questions and measure what one conversation actually costs · ~1.5h
  2. Build the three-tier cost spreadsheet from the published per-token prices · ~1.5h
  3. Write the one-page explainer with two named hallucination failure modes · ~1.5h
  4. Record the 10-minute founder walkthrough · ~1.5h
Done when
  • Explainer avoids all ML jargon (no 'transformer', 'attention', 'logits') while staying technically accurate
  • Cost model cites the actual provider pricing page and shows the token-count assumptions used
  • At least 2 concrete hallucination failure modes are described with a mitigation for each
  • A recorded walkthrough (5-10 min) exists and is linked from the README
Prove it

Evidence a recruiter can check

  • A three-tier cost spreadsheet showing the token-count assumptions per conversation and linking the pricing page every number came from
  • A one-page explainer that survives a jargon check — no 'transformer', 'attention' or 'logits' — and still gets grounding right
  • Two hallucination failure modes you actually triggered in the playground, each with the mitigation you would ship
  • A 5-10 minute recorded walkthrough pitched at someone who has never used an LLM API
Signal it

Built a token-level cost model for a Hindi/Hinglish support bot across three volume tiers and an explainer that took a non-technical audience to a build/no-build decision without ML jargon.

Interview

Questions you'll get asked

  1. Explain to a non-technical VP why our chatbot sometimes makes things up.
  2. When would you use RAG instead of fine-tuning, and why?
  3. What is a token, and why does context length matter for both quality and cost?
  4. Walk me through what happens end to end when a user sends a prompt to an LLM API.
  5. What's the difference between embeddings-based search and a knowledge graph?
  6. How would you estimate the monthly API cost for a support-chat feature handling 10,000 conversations/month?
  7. Give me three distinct ways an LLM-powered feature can fail in production.