All capabilities · Programming foundations

Design scalable AI-backed systems

Reason about latency, queues, caching, storage and failure modes for systems that call models.

~30 focused hoursadvanced
Explore 4 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Design a system that serves an LLM API to thousands of concurrent users without falling over
  2. Choose between synchronous and queue-based (async) processing for a slow AI task
  3. Add caching to avoid re-calling an expensive model for repeated inputs
  4. Design for a model provider outage (fallback model, retries with backoff, circuit breaker)
  5. Estimate capacity: how many GPUs/requests-per-second does a given traffic pattern need
  6. Design rate limiting and quota enforcement per tenant/customer
  7. Reason about data flow for a RAG pipeline (ingestion, embedding, retrieval, generation) at scale

Needs first: Build and consume REST APIs, Deploy an AI service to the cloud

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

4 tools to explore

draw.io

Plan & explain

Draw a process or architecture diagram that can be reviewed alongside your project.

Redis

Build

Add a cache and measure the effect of expiry, hit rate and invalidation on an application.

Apache Kafka

Build · Data

Pass events between services and practise consumer groups and message handling.

Practices & references

  • Load balancing
  • Queues
  • Rate limiting
Practice

Resilient RAG API design for a growing document corpus

Design, and prototype the risky parts of, a RAG API answering questions over a corpus that keeps growing: an ingestion queue, an embedding and vector-store step, a cache for repeated questions, and a fallback path when the primary LLM provider errors. Use RBI Master Directions as the corpus — hundreds of public regulatory PDFs that grow over time, so the capacity maths is real. Write the capacity estimate and the failure-mode doc with numbers and stated assumptions, and build the two components most likely to bite you.

Start from

RBI Master Directions — several hundred public regulatory PDFs, downloadable without signup and updated over time

Milestones
  1. Size the corpus and write the capacity estimate — QPS, storage growth, monthly cost · ~4.5h
  2. Draw the architecture: ingestion queue, embedding, vector store, cache, fallback · ~5h
  3. Prototype the cache and measure its hit rate over a replayed question log · ~7.5h
  4. Prototype the provider fallback and test it by killing the primary mid-request · ~6h
  5. Write the failure-mode doc: queue backlog, provider outage, cold cache · ~2.5h
Done when
  • Architecture diagram covering ingestion, retrieval, generation, caching and failure fallback
  • Written capacity estimate (requests/sec, storage growth, cost) with assumptions stated
  • A working prototype of at least the caching layer or the fallback/retry logic, with tests
  • A documented failure scenario (provider outage or queue backlog) and how the design handles it
Prove it

Evidence a recruiter can check

  • A capacity estimate with the arithmetic shown — documents ingested per day, embedding spend, and vector-store growth over 12 months
  • A measured cache hit rate over a replayed question log, with the latency and cost it saved
  • A test that fails the primary provider mid-request and the log lines showing the fallback answering
  • An architecture diagram annotated with what breaks first under load and what the design does about it
Signal it

Designed and prototyped a resilient RAG API over a few hundred regulatory PDFs — a measured cache hit rate cut per-query cost, and a tested provider-fallback path kept answers flowing through a simulated outage.

Interview

Questions you'll get asked

  1. Design a chatbot backend that needs to serve 10,000 concurrent users with sub-2s latency.
  2. How would you cache LLM responses without serving stale or unsafe answers?
  3. Design a system that ingests 1M documents a day into a RAG pipeline.
  4. Your model provider's API is down for 10 minutes — how does your system degrade gracefully?
  5. How would you rate-limit an API per customer while keeping p99 latency low?
  6. Walk me through the trade-offs of a synchronous vs async (queue + webhook) API for a slow AI task.
  7. How would you scale a vector search index from 1M to 100M embeddings?