All capabilities · Cloud, deployment & production

Troubleshoot and support AI systems in production

Own L1-L3 tickets for an AI product or platform: read logs and traces, reproduce model and API failures, follow the runbook, escalate with a clean root-cause note and close within SLA.

~24 focused hoursbeginner
Explore 5 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Triage an incoming ticket about an AI product: reproduce the failure, assign a severity and acknowledge it within the SLA clock
  2. Read application and container logs to find the request that failed and the line that says why
  3. Tell an upstream model/API failure (429 rate limit, timeout, 5xx) apart from a bug in the customer's request (400, bad schema, missing key)
  4. Follow a runbook step by step to restore service, and note where the runbook was wrong or missing
  5. Reproduce a reported model or API failure with a minimal curl or Python call and attach it to the ticket
  6. Write a bug report Engineering can act on: description, reproduction steps, expected vs actual, log excerpt, timestamps
  7. Escalate L1 to L2/L3 with a clean hand-off note instead of forwarding the raw complaint
  8. Write a one-page blameless RCA: timeline, root cause, contributing factors, action items with owners
  9. Spot a memory/GPU limit or misconfigured environment variable as the cause of a crashing inference container

Needs first: Integrate LLM APIs into an application

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

5 tools to explore

Docker

Deploy · Build

Package a service with its dependencies and run a repeatable local environment.

Jira

Plan & explain

Record a reproducible incident, track investigation steps and document the resolution.

Practice

AI support lab: break a local LLM API five ways and support it back to health

Run a small model on Ollama behind a thin FastAPI wrapper in Docker Compose, with request ids and structured JSON logs on every call. Inject five realistic failures you can switch on at will: an upstream rate limit, a model timeout, a malformed request, a bad environment variable, and a container memory limit that kills the model. For each one write the ticket as a customer would file it, the log excerpt that diagnoses it, the runbook step that restores service, and a one-page blameless RCA. Finish with a small log dashboard that shows error rate by failure class.

Start from

A small open model pulled with Ollama, wrapped in your own FastAPI service; the five failures and their tickets are scripted by you

Milestones
  1. Stand up the lab: Ollama + FastAPI wrapper in Docker Compose, request ids, JSON logs and a /health endpoint · ~4h
  2. Inject the five failures behind toggles (rate limit, timeout, malformed request, bad env var, memory limit) and reproduce each on demand · ~5h
  3. For each failure: file the ticket, pull the diagnosing log excerpt, write the runbook step that fixes it · ~6h
  4. Write five one-page blameless RCAs plus a severity/SLA table and an L1-to-L3 escalation note · ~4.5h
  5. Build a log dashboard (Grafana Loki or a small Python viewer) showing error rate per failure class · ~3.25h
Done when
  • All five failures can be triggered and cleared with a documented command, and each produces a distinct, searchable log signature
  • Five tickets, each with severity, reproduction steps, the log excerpt that diagnoses it and the runbook step that resolved it
  • Five one-page RCAs with timeline, root cause, contributing factors and at least one action item each
  • The dashboard shows error rate by failure class and a request id can be traced from ticket to log line
Prove it

Evidence a recruiter can check

  • Five tickets side by side with the log excerpt that diagnoses each, showing a 429, a timeout, a 400, a bad env var and an OOM kill are told apart from the logs alone
  • The runbook, with the exact commands that clear each failure and a note on the one step you got wrong the first time
  • Five one-page blameless RCAs, each with a timestamped timeline and an action item
  • A dashboard screenshot showing error rate per failure class during a run where you triggered all five
  • The Docker Compose file with the failure toggles and memory limit visible as configuration
Signal it

Built a Dockerised LLM API support lab, injected five production-style failures (rate limit, timeout, malformed request, bad config, OOM) and diagnosed each from logs alone — with tickets, a runbook, blameless RCAs and an error-rate dashboard.

Interview

Questions you'll get asked

  1. A customer says the AI assistant 'stopped working'. Walk me through your first 15 minutes.
  2. How do you tell a rate-limit error from a timeout from a malformed request when all the customer sees is 'something went wrong'?
  3. A model container keeps restarting. Which logs do you read first and what are you looking for?
  4. What goes into a good escalation note to L3? What do you leave out?
  5. How do you write an RCA that doesn't blame a person?
  6. A P1 hits at 2 a.m. with a 30-minute response SLA. What do you do in what order?
  7. Describe a runbook you have followed that turned out to be wrong. What did you do?
  8. How would you reproduce an intermittent 500 from an LLM API with only the customer's request id?