All capabilities · LLM application development

Manage context windows and memory

Token budgeting, summarization, conversation memory, caching prompts across turns.

~8 focused hoursintermediate
Explore 4 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Budget tokens so a conversation doesn't silently exceed the model's context window
  2. Summarize or truncate older turns while preserving the facts that matter
  3. Use prompt caching to cut latency/cost on repeated system prompts or long documents
  4. Design conversation memory that persists key facts across sessions, not just within one context window
  5. Decide what to keep verbatim vs. summarize vs. drop from a long chat history
  6. Measure and report token/cost usage per conversation for budgeting

Needs first: Design and version prompts systematically

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

4 tools to explore

tiktoken

Build

Count tokens and compare how prompt changes affect the input budget.

Claude API

Build · Test

Build model-backed features with messages, tool use and responses you can evaluate.

OpenAI API

Build · Test

Connect model calls, tool use and structured responses to your own application.

Practices & references

  • Prompt caching
  • Conversation summarisation
  • Memory retrieval
Practice

Rolling-memory assistant for 100-turn conversations

Build a chat assistant that holds a 100+ turn conversation without ever hitting the context limit. Drive it from a transcript you script yourself - a borrower asking follow-up questions about a public bank loan FAQ page gives you the repetitive, back-referencing turns real support produces. Combine a sliding window of recent turns with a running summary of everything older, cache the frozen system prompt, and log token usage per turn so you can chart how cost grows across a long session.

Start from

A 100-turn conversation you script yourself - a simulated borrower asking follow-ups about a public bank loan FAQ page, replayed turn by turn

Milestones
  1. Script the 100-turn conversation and a runner that replays it · ~1h
  2. Add the sliding window and per-turn token logging · ~1.5h
  3. Summarise turns as they fall out of the window and refresh on an interval · ~1.5h
  4. Turn on prompt caching and chart cost per turn across the whole session · ~1.5h
Done when
  • A simulated 100-turn conversation completes without hitting a context-length error
  • Token usage per turn is logged and a chart/table shows caching reduces repeated-prefix cost after turn 1
  • A running summary is regenerated periodically and a spot-check shows it retains key facts from more than 20 turns back
  • Configurable window size and summary-refresh interval, documented in README
Prove it

Evidence a recruiter can check

  • A cost-per-turn chart across the full 100-turn run, showing where summarisation and caching bend the curve
  • cache_read_input_tokens (or the provider equivalent) non-zero in the logs from turn 2 onward
  • A recall test: a fact stated at turn 5 answered correctly at turn 90, with both turns quoted
  • The summarise-versus-truncate tradeoff written up with the window size and refresh interval you settled on
Signal it

Built a rolling-memory chat assistant that holds 100+ turn conversations inside a fixed context budget - sliding window plus periodic summarisation and prompt caching, with per-turn token cost charted across the session.

Interview

Questions you'll get asked

  1. How do you keep a multi-turn chatbot from exceeding the model's context window on a long conversation?
  2. What is prompt caching and when does it actually save cost — walk through prefix stability.
  3. How would you design memory for an assistant that needs to remember a user's preferences across sessions, weeks apart?
  4. What's the tradeoff between summarizing old turns vs. just truncating them?
  5. How do you count tokens before sending a request, and why does that matter for cost and reliability?
  6. A user's conversation is degrading in quality after 40 turns — how do you debug whether it's a context problem?