All capabilities · Agents & workflows

Implement tool / function calling

Define tools, let the model call them, execute safely, return results into the loop.

~8 focused hoursintermediate
Explore 4 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Define a tool/function schema (name, description, JSON-schema parameters) that an LLM can call reliably
  2. Parse a model's tool_call / function_call response and route it to the right Python function
  3. Validate and coerce tool arguments with Pydantic before executing side-effecting code
  4. Return tool results back into the conversation so the model can use them in its next turn
  5. Handle parallel/multiple tool calls in a single model turn
  6. Design clear tool names, descriptions and error messages so the model picks the right tool and recovers from bad input
  7. Add retries and timeouts around flaky external tool calls (APIs, DB queries, web search)
  8. Force or restrict tool choice (e.g. require a specific tool, or disable tools for a turn)

Needs first: Get reliable structured outputs from LLMs

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

4 tools to explore

OpenAI API

Build · Test

Connect model calls, tool use and structured responses to your own application.

Claude API

Build · Test

Build model-backed features with messages, tool use and responses you can evaluate.

Pydantic

Build · Test

Define typed schemas and validate the structured data entering your application.

Practices & references

  • JSON Schema
  • Tool argument validation
Practice

Transaction-support bot that answers from tools, not memory

Build a support chatbot that answers questions about payment transactions (status, refund eligibility, dispute filing) by calling four tools against a local transactions database instead of guessing. You create the data yourself: a short Python generator seeds a SQLite table with ~200 synthetic transactions covering settled, failed, pending and already-refunded cases, so no real payment system or merchant account is involved. Validate every tool's arguments with Pydantic before execution and feed structured results back into the conversation. Put a confirmation turn in front of the one tool that mutates data.

Start from

A SQLite transactions table you seed yourself — ~200 synthetic records (txn id, amount, status, timestamp, refund window) from a 30-line Python generator

Milestones
  1. Seed the synthetic transactions DB and write the three read-only tools · ~1.5h
  2. Wire the tool-call loop: parse tool_calls, dispatch, feed results back into the turn · ~2.5h
  3. Add Pydantic argument validation and a confirmation turn before the dispute tool · ~1.5h
  4. Log 10 sample conversations and write the README with the tool schemas · ~1h
Done when
  • At least 3 distinct tools defined with JSON-schema parameters and docstrings
  • Pydantic models validate and reject malformed tool arguments before execution
  • A destructive tool (raise dispute) requires an explicit user confirmation turn before executing
  • 10 sample conversations logged showing correct tool selection, including one where the model asks a clarifying question instead of guessing
Prove it

Evidence a recruiter can check

  • The four tool schemas beside a transcript log of 10 conversations, annotated with which tool fired for each intent
  • A transcript where the model asks a clarifying question instead of inventing a transaction id
  • Unit tests that feed malformed arguments to each tool and assert they are rejected before anything executes
  • A recorded run of the dispute tool refusing to mutate a row until the user confirms in a separate turn
Signal it

Built a transaction-support bot on LLM tool calling — four Pydantic-validated tools over a transactions database, with a confirmation gate on the one destructive action, so every answer traces back to a row instead of the model's guess.

Interview

Questions you'll get asked

  1. Walk me through the full request/response loop when a model decides to call a tool.
  2. How do you make a tool description good enough that the model calls it correctly on the first try?
  3. What happens if the model hallucinates arguments that don't match your Pydantic schema? How do you handle that?
  4. How would you let a model call two independent tools in parallel and merge the results?
  5. Design a 'get order status' tool for an e-commerce chatbot — what's in the schema, and what do you return on failure?
  6. How do you prevent a tool-calling agent from calling a destructive tool (e.g. refund, delete) without confirmation?
  7. What's the difference between forcing a specific tool call and letting the model choose freely?