All capabilities · Evaluation, safety & observability

Optimize inference cost and latency

Model routing, caching, batching, smaller models, quantization/vLLM where self-hosting.

~10 focused hoursadvanced
Explore 3 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Route easy queries to a cheaper/smaller model and only escalate hard queries to a frontier model
  2. Use prompt caching to avoid re-paying for repeated long system prompts or context across requests
  3. Add a semantic cache so repeated/similar queries are served from cache instead of hitting the model again
  4. Batch requests where possible to improve throughput and reduce per-request overhead
  5. Self-host an open model with vLLM (continuous batching, quantization) when API costs exceed self-hosting costs at scale
  6. Set and monitor a per-feature cost budget, alerting before it's exceeded rather than after the bill arrives
  7. Quantize a model (int8/int4) to cut memory and latency when self-hosting, and measure the quality trade-off
  8. Measure and report cost-per-request and p95 latency before and after each optimization to justify it

Needs first: Trace, monitor and debug LLM apps in production

Learn — free, link-checked

The few resources that matter

Read · intermediate · 20 min · langfuse.com

Token & Cost Tracking

Shows exactly how per-call token/cost is captured and rolled up into dashboards, the basis for any cost-optimization work. — Langfuse
Read · intermediate · 20 min · platform.openai.com

Prompt Caching

OpenAI's automatic prompt-caching behavior and pricing, know the platform differences when discussing cost optimization. — OpenAI
Read · intermediate · 25 min · docs.anthropic.com

Prompt caching

Cuts repeated-context cost/latency up to 90%, a concrete, demoable optimization for agent systems with long system prompts. — Anthropic
Read · intermediate · 30 min · docs.litellm.ai

LiteLLM

Unified proxy/SDK for calling 100+ LLM providers with one interface, the standard tool for model routing and fallback in JDs. — LiteLLM (BerriAI)
Read · intermediate · 30 min · docs.litellm.ai

Routing, Fallbacks & Load Balancing

Shows exactly how to route cheap/fast models for easy queries and fall back to stronger models, the core cost/latency lever. — LiteLLM (BerriAI)
Course · intermediate · 90 min · deeplearning.ai

Quantization Fundamentals with Hugging Face

Hands-on quantization (int8/int4) of open models to cut memory/latency when self-hosting instead of calling paid APIs. — DeepLearning.AI (with Hugging Face)
Read · advanced · 45 min · docs.vllm.ai

vLLM

The standard high-throughput inference server (PagedAttention, continuous batching, quantization) for self-hosting open models cheaply. — vLLM Project
Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

3 tools to explore

tiktoken

Build

Count tokens and compare how prompt changes affect the input budget.

Practices & references

  • Prompt caching
  • Semantic caching
  • Model routing and quantisation
Practice

Cost and latency optimization pass on a support RAG chatbot

Take a working RAG support chatbot and cut its cost per request and p95 latency without losing answer quality. Build a 200-query replay set by sampling the Bitext customer-support dataset on Hugging Face, mixing easy FAQ-style questions with multi-part ones so a router has something real to route. Add prompt caching for the repeated system prompt and retrieved context, a LiteLLM routing layer sending easy queries to a small model and hard ones to a frontier model, and a semantic cache for repeated questions with a similarity safeguard. Measure cost per request and p95 latency after every single change, not just at the end.

Start from

A 200-query replay set you sample from the Bitext customer-support dataset on Hugging Face — mixed easy FAQ and multi-part questions across 27 intents

Milestones
  1. Build the replay harness and record baseline cost-per-request and p95 latency over the 200 queries · ~1.5h
  2. Turn on prompt caching for the system prompt and retrieved context, and measure the hit rate · ~1.5h
  3. Add the LiteLLM router across a cheap and a frontier model, and score routing accuracy · ~3h
  4. Add the semantic cache with a similarity threshold and a staleness safeguard · ~1.5h
  5. Write the before/after report with the numbers attributed per change · ~1h
Done when
  • Prompt caching enabled for the repeated system prompt/context, with a measured cache-hit rate
  • A LiteLLM-based router sending at least two query classes to two different-cost models, with routing accuracy reported
  • A semantic cache serving repeated/similar queries, with a safeguard against serving stale/wrong answers
  • A before/after report showing cost-per-request and p95 latency reduced by a measurable, stated percentage
Prove it

Evidence a recruiter can check

  • The before/after table — cost per request and p95 latency at baseline, then after caching, after routing, after the semantic cache, one row per change
  • Routing accuracy over the 200-query replay set, with the misrouted queries listed and what each one cost you
  • Cache-hit rate for the run, the similarity threshold you settled on, and the wrong answer that made you raise it
  • Traces showing which model each request actually reached, so the routing and caching claims are checkable rather than asserted
Signal it

Cut per-request cost and p95 latency on a support RAG chatbot using prompt caching, LiteLLM model routing and a semantic cache — measured on a 200-query replay set with before/after numbers attributed to each optimization.

Interview

Questions you'll get asked

  1. How would you cut a chatbot's per-request cost by 50% without noticeably hurting quality?
  2. Explain how prompt caching works and where it gives the biggest win.
  3. When does it make sense to self-host with vLLM instead of calling a hosted API, and what does that trade off?
  4. How would you design a model-routing layer that sends easy queries to a small model and hard ones to GPT-4/Claude?
  5. What's the risk of an aggressive semantic cache, and how do you avoid serving a stale or wrong cached answer?
  6. How do you quantify the latency/cost impact of quantizing a self-hosted model?
  7. Walk me through a real cost-optimization you shipped — before/after numbers?