All capabilities · LLM application development

Integrate LLM APIs into an application

Call OpenAI/Anthropic/Gemini/Bedrock APIs with streaming, retries, token accounting and cost awareness.

~12 focused hoursbeginner
Explore 4 tools for this project
What employers mean

You should be able to…

  1. Call OpenAI/Claude/Gemini APIs from a backend service with proper error handling
  2. Implement streaming responses so the UI shows tokens as they arrive
  3. Add retry/backoff logic for rate limits (429) and transient 5xx errors
  4. Track token usage and cost per request/user for billing or budgeting
  5. Swap between model providers (OpenAI, Anthropic, Bedrock, Vertex) behind a common interface
  6. Handle API keys/secrets securely (env vars, secret managers) not hardcoded
  7. Set sane timeouts and max_tokens to avoid hung requests in production

Needs first: Write production-quality Python for AI work

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

4 tools to explore

Claude API

Build · Test

Build model-backed features with messages, tool use and responses you can evaluate.

OpenAI API

Build · Test

Connect model calls, tool use and structured responses to your own application.

Amazon Bedrock

Build · Deploy

Connect managed models to an application and explore the surrounding AWS controls.

tiktoken

Build

Count tokens and compare how prompt changes affect the input budget.

Practice

Multi-provider LLM chat CLI with cost tracking

Build a command-line chat tool that talks to two different LLM providers behind one shared interface, streams the response token-by-token, and logs the cost of every call in INR (fixed USD-INR rate) to a local SQLite file. Any two providers work: if you don't already hold OpenAI or Anthropic credit, Google AI Studio and Groq both issue a key with no card, so the whole build runs on zero spend. Add retry-with-backoff for rate limits and a --budget flag that ends the session once cumulative spend crosses a limit.

Start from

Free API keys from two providers that issue one without a card (Google AI Studio, Groq) - or your own OpenAI/Anthropic keys - plus a 20-prompt benchmark set you write yourself

Milestones
  1. Get a streaming call working against one provider from the CLI · ~2.5h
  2. Put both providers behind one interface selected by --provider · ~2.5h
  3. Log tokens and INR cost per call to SQLite and add --report · ~2.5h
  4. Add backoff on 429/5xx and wire up the --budget cutoff · ~2.5h
Done when
  • Switching --provider anthropic|openai changes the backend with no other code changes
  • Responses stream to stdout token-by-token, not all at once
  • A 429/5xx from the API triggers exponential backoff and retries, visible in logs
  • Every call's token usage and estimated INR cost is persisted and summable via a --report command
Prove it

Evidence a recruiter can check

  • A --report table pasted in the README showing token counts and INR cost across 20+ real calls
  • A latency and spend comparison of both providers on the identical 20-prompt set, with the numbers rather than claims
  • A log excerpt from a forced 429 showing the backoff intervals and the retry that succeeded
  • Unit tests that mock the provider client and assert retry behaviour on 429/500 without spending tokens
Signal it

Built a multi-provider LLM chat CLI with streaming, retry-with-backoff and per-call INR cost tracking - benchmarked latency and spend across two providers on an identical prompt set, with a budget cap that ends a session before it overruns.

Interview

Questions you'll get asked

  1. How do you handle a 429 rate-limit error from the OpenAI/Claude API in production?
  2. Walk me through how you'd stream a response from an LLM API to a React frontend.
  3. How do you estimate and control the cost of an LLM feature before shipping it?
  4. What's the difference between synchronous and streaming API calls, and when do you use each?
  5. How would you design a wrapper that lets you switch from GPT-4o to Claude without rewriting business logic?
  6. How do you test code that calls an external LLM API without burning tokens on every CI run?