All capabilities · Data engineering & analytics

Use LLMs to accelerate analysis (text, SQL, summaries)

Text-to-SQL, classify free text at scale, summarize reviews/tickets with LLM APIs — with validation.

~10 focused hoursintermediate
Explore 4 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Build a text-to-SQL assistant that turns a natural-language question into a validated query
  2. Classify free-text at scale (support tickets, reviews, survey responses) using an LLM with a fixed taxonomy
  3. Summarize large volumes of unstructured feedback (reviews, call transcripts) into themes with counts
  4. Validate LLM output against ground truth before trusting it in a report (never ship un-spot-checked numbers)
  5. Design prompts that return structured, parseable output (JSON) instead of free-form text
  6. Know when NOT to use an LLM — e.g. exact aggregation is a SQL job, not a prompt job
  7. Estimate and control API cost/latency for a batch job over thousands of rows

Needs first: Clean and transform data with Pandas, Integrate LLM APIs into an application

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

4 tools to explore

OpenAI API

Build · Test

Connect model calls, tool use and structured responses to your own application.

Claude API

Build · Test

Build model-backed features with messages, tool use and responses you can evaluate.

Practice

LLM triage pipeline for a support-ticket backlog (English + Hinglish)

Sample 200 real support conversations from the Customer Support on Twitter dataset and write 10-15 Hinglish tickets yourself for the code-mixed cases. Build a pipeline that calls an LLM API to classify each ticket into a fixed taxonomy (billing, delivery, product defect, other), extract sentiment, and roll the week up into a top-issues summary, emitting structured JSON per ticket. Then do the unglamorous half: hand-label a held-out sample, score the model against it, and iterate the prompt until the number moves. Report the cost per 1,000 tickets.

Start from

Customer Support on Twitter dataset on Kaggle (~3M real support tweets — sample 200), plus 10-15 Hinglish tickets you write yourself for the code-mixed cases

Milestones
  1. Sample the tickets, write the Hinglish ones, and pin the taxonomy and JSON output schema · ~2h
  2. Build the batched classification pipeline with schema validation and retry on malformed output · ~2h
  3. Hand-label the held-out sample and score the LLM against it per category · ~1.5h
  4. Iterate the prompt on the worst category, re-score, and log token cost per 1,000 tickets · ~1.5h
Done when
  • Pipeline processes at least 200 tickets end-to-end and outputs structured JSON per ticket
  • A held-out sample of 30-50 tickets is manually labeled and compared to LLM output with a reported accuracy/agreement number
  • Handles at least one mixed Hindi/English (Hinglish) input correctly
  • README documents prompt design choices and estimated per-1000-ticket API cost
Prove it

Evidence a recruiter can check

  • A per-category accuracy table against your hand-labelled held-out sample, with the categories the model confuses called out by name
  • The prompt-iteration log: each version, what you changed, and the accuracy before and after
  • Structured JSON output for the full 200-ticket run, including the Hinglish tickets, so a reader can spot-check the labels
  • Measured token cost per 1,000 tickets, with the arithmetic shown rather than estimated
Signal it

Built an LLM triage pipeline that classifies support tickets into a fixed taxonomy with validated JSON output — measured against a hand-labelled sample, improved through prompt iteration, and costed per 1,000 tickets.

Interview

Questions you'll get asked

  1. How would you build a 'chat with your data' feature safely, so it can't run destructive SQL?
  2. You ask an LLM to classify 10,000 support tickets into 8 categories — how do you validate accuracy without reading all 10,000?
  3. What's your approach to getting reliable structured JSON output from an LLM?
  4. When would you use an LLM vs. a traditional NLP/regex approach for a text task?
  5. How do you handle hallucinated numbers when an LLM is asked to summarize a dataset?
  6. Describe a pipeline for summarizing 5,000 Hindi customer support chats into top 10 issues, with counts.
  7. How would you keep API costs under control when classifying a million rows?