All capabilities · Evaluation, safety & observability

Trace, monitor and debug LLM apps in production

Langfuse/LangSmith/OpenTelemetry traces, token & cost dashboards, drift and failure alerts.

~8 focused hoursintermediate
Explore 4 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Instrument an LLM app with traces/spans so every prompt, tool call, and model response is captured with timing
  2. Track token usage and cost per request, per user, and per feature, rolled up into a dashboard
  3. Set up alerts for latency spikes, error-rate increases, or cost anomalies in production
  4. Correlate a user-reported bad response back to its exact trace (prompt, retrieved context, tool calls) for debugging
  5. Detect drift: a model or prompt update silently changing output quality or latency distribution
  6. Self-host or configure a tracing backend (Langfuse) vs use a managed one (LangSmith) based on data-residency needs
  7. Export traces via OpenTelemetry so LLM observability integrates with existing infra monitoring

Needs first: Deploy an AI service to the cloud

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

4 tools to explore

OpenTelemetry

Monitor

Instrument an application with traces and metrics to follow a request across services.

Grafana

Monitor · Data

Build dashboards for service health and investigate changes in operational metrics.

Practices & references

  • Structured logging
  • Trace correlation
Practice

Per-tenant cost and latency dashboard for a RAG support bot

Stand up a small RAG support bot serving two tenants over two different sets of public policy/FAQ PDFs, then instrument every request with Langfuse tracing — nested spans for retrieval, prompt construction, the model call and any tool call. Use the Langfuse Cloud free tier (free signup, no card) or self-host it with Docker. Build a view showing token cost and p95 latency per tenant per day, and add an alert rule for a latency or error-rate threshold. Then generate a deliberately bad answer and trace it from the user report back to the retrieved chunk that caused it.

Start from

Langfuse Cloud free tier (free signup, no card) plus a two-tenant RAG bot you build over two sets of public policy/FAQ PDFs

Milestones
  1. Stand up the two-tenant RAG bot and get one end-to-end trace landing in Langfuse · ~1.5h
  2. Add nested spans for retrieval, prompt build and model call, tagged by tenant · ~1.5h
  3. Build the per-tenant cost and p95 latency view and configure the alert rule · ~1.5h
  4. Simulate a latency spike and a bad answer, then write the root-cause walkthrough · ~1.5h
Done when
  • Every request produces a trace with nested spans for retrieval, prompt construction, and model call
  • A dashboard (Langfuse UI or exported to Grafana) showing token cost and p95 latency broken down by tenant
  • An alert rule configured and demonstrated firing on a simulated latency spike or error burst
  • A written walkthrough of tracing one bad response from user report to root cause using the trace
Prove it

Evidence a recruiter can check

  • The per-tenant cost and p95 latency view, screenshotted with real numbers from a day of traffic you generated
  • A root-cause walkthrough that starts at one bad answer and ends at the retrieved chunk that caused it, with the trace and span ids shown at each hop
  • The alert firing on a simulated latency spike — the notification alongside the span that explains it
  • The instrumentation diff, showing exactly which spans wrap retrieval, prompt construction and the model call
Signal it

Instrumented a multi-tenant RAG support bot with Langfuse — nested spans for retrieval, prompt and model call, per-tenant cost and p95 latency dashboards, and alerting that caught a simulated latency spike before users did.

Interview

Questions you'll get asked

  1. How would you debug a single bad response from a production RAG chatbot down to its root cause?
  2. What would you put on a cost/latency dashboard for an LLM feature, and who looks at it?
  3. How do you detect that a vendor's silent model update degraded your app's quality?
  4. Explain the difference between a trace, a span, and a generation in LLM observability tooling.
  5. How would you set up alerting for a spike in token cost without false-paging on normal traffic growth?
  6. Why would you choose Langfuse (self-hosted) over LangSmith (managed), or vice versa, for an Indian fintech client?
  7. How does OpenTelemetry fit into an LLM observability stack that already uses Langfuse or LangSmith?