LLM Tool Calling

Tool calling (also called function calling or tool use) lets a language model do more than generate text: it can request that your application run a function, such as searching a database, calling an API, reading a file, or sending a message, and then use the result to continue. The model doesn't execute anything itself. It emits a structured request (tool name plus JSON arguments), your code runs the tool, and you send the result back. That loop is the foundation of AI agents.

Every major model API supports it, and protocols like the Model Context Protocol standardize how tools are exposed. The quality of an agent depends heavily on tool design: clear names and descriptions, well-typed inputs, concise and informative outputs, and safe handling of anything with side effects.

TL;DR

Quick Example

A tool-use loop with the Anthropic Python SDK:

SDK helpers (tool runners) and agent frameworks implement this loop for you, but it's worth understanding the raw version.

Core Concepts

Tool Definitions

A tool definition is a contract the model reads:

The Tool-Use Loop

  1. Send the conversation plus the tool definitions.
  2. The model responds with text and/or one or more tool calls (the stop reason indicates tool use).
  3. Your application validates and executes each call.
  4. Append the assistant turn and a message containing tool results (matched by tool call ID).
  5. Repeat until the model returns a normal answer, or a step limit is reached.

Tools can also be server-side (executed by the provider, such as web search, code execution, or computer use) or exposed via MCP servers that your client connects to.

Parallel Tool Calls

Models can emit several independent calls in one turn ("get weather in Paris and Tokyo"). Execute them concurrently and return all results together, in one user message. That cuts latency substantially in agent workflows.

Tool Choice

Results and Errors

Tool results are ordinary content the model reads, so design them for the model:

Designing Good Tools

Test tools the way you'd test an API for a new colleague: run realistic tasks, read the transcripts, and refine descriptions where the model hesitates or misuses them. See context engineering.

Security

Tool calling connects model output to real systems, so treat the model as an untrusted caller:

Best Practices

Write Descriptions Like Documentation

Spend more effort on tool descriptions than on system prompts. Include when to use the tool, parameter semantics, units and formats, and examples of good inputs.

Keep the Toolset Focused

Dozens of overlapping tools confuse selection and bloat every request. Offer the minimal set for the task, load tools dynamically (tool search), or split work across specialized agents. See multi-agent systems.

Bound the Loop

Set maximum iterations, timeouts, and token budgets, and detect repeated identical calls, so a confused model can't loop forever.

Log Every Call

Record tool name, arguments, results, latency, and errors for debugging, evaluation, and auditing. Traces of agent runs are the primary way to improve tool design. See LLM evaluation.

Common Mistakes

Dumping Raw API Responses

Map responses to the handful of fields the model needs, and offer a detail tool or pagination for more.

Vague Tool Descriptions

"description": "Gets data" leads to wrong tool selection and invented arguments. Be specific about purpose, inputs, and outputs.

Executing Side Effects Without Confirmation

An agent that can send_email or issue_refund on its own interpretation of an ambiguous request will eventually do something unintended. Gate high-impact tools behind explicit approval.

FAQ

Does the model execute the tools itself?

No. For client-side tools, the model only produces a structured request; your application decides whether and how to execute it and returns the result. Provider-hosted tools (web search, code execution) run on the provider's infrastructure, but the model still only requests them.

What's the difference between tool calling and structured outputs?

Structured outputs constrain the model's final response to a schema. Tool calling lets the model request actions mid-conversation and continue with the results. Forcing a single tool call with a schema is also a common way to get structured data.

How many tools can I give a model?

Models handle dozens of well-described tools, but accuracy and cost degrade as the toolset grows and overlaps. For large toolsets, use dynamic tool loading or search, group tools by domain, or route tasks to specialized agents with smaller toolsets.

How is MCP related to tool calling?

The Model Context Protocol is a standard for exposing tools (and resources and prompts) from servers to AI applications. An MCP client discovers a server's tools and presents them to the model as ordinary tool definitions, so any MCP-compatible app can use any MCP server's tools.

Related Topics

References