A language model cannot call an API, read a database or send an email. What it can do is emit a structured request that says "I would like this function run with these arguments". Everything else, running the code, handling failures, returning the result, is your application. This guide explains how tool calling works in the Claude, OpenAI and Gemini APIs, using each vendor's own documentation, and where the three differ. For the larger picture of what wraps this loop, see how AI agents work and how to secure them, and for the protocol layer that standardises tool access, what MCP is.
The loop, as each vendor describes it
Anthropic describes tool use as letting Claude "call functions that you define or that Anthropic provides". Claude decides when to call a tool "based on the user's request and the tool's description", and returns "a structured call that your application executes (client tools) or that Anthropic executes (server tools)."[1]
Google's guide breaks the same loop into four steps: define a function declaration, call the model with the prompt and declarations, execute the function in your own code, then send the result back so the model can produce a final response. Its wording is explicit: "The model doesn't execute the function itself."[3]
OpenAI's guide lists five high-level steps: make a request with tools, receive a tool call, run your code, send the output back, then receive a final response or further tool calls.[4]
| Step | Claude | OpenAI | Gemini |
|---|---|---|---|
| Model asks for a call | Response with stop_reason: "tool_use" and a tool_use block[1] | A function call carrying call_id, name and JSON-encoded arguments[4] | A function_call step with a name and arguments[3] |
| You return the result | A tool_result block that references the tool_use id[1] | A function_call_output that references the call_id[4] | A function result sent back in a new request[3] |
| Loop repeats | Described on a separate "How tool use works" page not read for this article[1] | Yes, "for as many calls as a task needs"[4] | Yes, across multiple turns[3] |
The practical consequence is that you own the loop. You decide the maximum number of iterations, what happens when a tool throws, and what text the model sees in the result. Anthropic's SDK offers a Tool Runner that executes your tools and returns results automatically, but the underlying round trip is the same.[1]
Client tools and server tools
Anthropic draws a line that the other two guides handle differently. Client tools (your own, plus Anthropic-schema tools such as bash and text_editor) run in your application. Server tools such as web_search, web_fetch and code_execution run on Anthropic's infrastructure, so you see results without handling execution.[1] That matters for security review: a server tool moves execution out of your process, and a client tool leaves every side effect in your hands.
How a tool is defined
All three describe a tool with a name, a natural-language description and a JSON Schema for its inputs. Anthropic's example uses an input_schema object with properties and required.[1] OpenAI lists type, name, description, parameters and a strict flag; Google lists type (which must be "function"), name, description and parameters.[4][3]
The description is the model's only guide for when to call the tool. Anthropic notes that the boundary between calling a tool and answering directly "is steerable through your system prompt", and gives examples of lighter and stronger instructions to call tools.[1] Treat descriptions as prompts, and test them like prompts.
Controlling when the model calls a tool
Each API has a setting for this. The names differ, so read the one for your vendor:
| Behaviour | Claude | OpenAI | Gemini |
|---|---|---|---|
| Model decides | auto (default)[1] | auto (default)[4] | auto (default)[3] |
| Must call some tool | any[1] | required[4] | any[3] |
| Must call a specific tool | tool[1] | Name the function[4] | Not covered on the part of the page read[3] |
| No tool calls | none[1] | none[4] | none[3] |
| Restrict to a subset | Not covered on the pages read | allowed_tools[4] | allowed_tools[3] |
Two cautions. First, the Claude any and tool names appear in the documentation's pricing table; the exact request syntax for forcing a tool is on a separate "Define tools" page that was not read for this article, so check it before copying code.[1] Second, Gemini lists a fourth mode, validated, described as the model ensuring schema adherence; this article does not cover how it differs from the strict modes below.[3]
Missing arguments
Anthropic's documentation includes a caveat worth remembering. If the prompt lacks a required parameter, Claude Opus "is much more likely to recognize that a parameter is missing and ask for it", while Sonnet "might also infer a reasonable value", and the page's example shows Sonnet inventing a location. It adds: "This behavior is not guaranteed."[1] The design rule that follows is our interpretation: never let a guessed argument reach an irreversible action without validation or confirmation.
Making arguments match your schema
Without enforcement, a model can return the wrong type. Anthropic's example is "2" instead of 2, or an omitted required field.[2] Two vendors offer a switch that prevents this.
- Claude: set
strict: trueon the tool. Anthropic says this "guarantees Claude's tool inputs match your JSON Schema by constraining the model's token sampling to schema-valid outputs", a technique it calls grammar-constrained sampling. The computer-use and browser-use toolset entries do not accept it.[2] - OpenAI: strict mode requires
additionalProperties: falseon each object and every property listed inrequired, with optional fields expressed by addingnullto the type. The guide says "We recommend always enabling strict mode", and notes that Chat Completions requests remain non-strict by default.[4]
Schema-valid is not the same as correct. A strict schema guarantees passengers is an integer, not that it is the right integer. Business rules, authorisation and idempotency still belong in your code.
Anthropic also documents a data-handling detail: strict tool schemas are compiled into grammars and cached for up to 24 hours since last use, so protected health information must not appear in property names, enum values, const values or patterns.[2]
Parallel and chained calls
A model can request several independent calls in one turn. OpenAI says setting parallel_tool_calls to false "ensures exactly zero or one tool is called", and lists snapshot-specific caveats, including that strict mode is disabled for fine-tuned models when they call multiple functions in one turn.[4] Anthropic's example disables parallel use with disable_parallel_tool_use: true inside the auto setting.[1] Google describes parallel calling for independent functions and compositional calling for sequences where one result informs the next, such as checking a forecast and then setting a thermostat.[3]
If your tools have side effects that depend on order, turn parallel calls off or make the tools safe to run in any order.
What it costs
Tool use is billed through ordinary tokens. Anthropic lists three sources: the tools parameter (names, descriptions and schemas), tool_use blocks, and tool_result blocks. It also adds an automatic tool-use system prompt whose size depends on the model; in its table, Claude Opus 5.5 and Sonnet 5.5 show 286 tokens for auto and none.[1] OpenAI says function definitions count toward the context limit and are "billed as input tokens", and recommends "fewer than 20 functions available at the start of a turn".[4] Prices change often; see the cost article for how to turn token counts into spend.
A long tool list therefore costs tokens on every request and can crowd the context. Both Anthropic and OpenAI document tool search features for loading rarely used tools on demand; OpenAI notes its tool search needs gpt-5.4 or later.[1][4]
State, history and reasoning models
The loop depends on history. Google notes that with store=false you must resend the full history on each request, including the original input, every model step exactly as received and the function result.[3] OpenAI states that reasoning models must have their reasoning items passed back with tool outputs, and that GPT-6 Astra and GPT-6.1 Sol require the Responses API for tool calling.[4] Both are reminders that dropping or editing a returned step can break the next call.
A checklist before production
This is editorial guidance drawn from the documented behaviour above, not a vendor recommendation:
- Cap loop iterations and set a timeout on every tool.
- Enable strict schemas where the vendor offers them, then validate values yourself.
- Return errors to the model as readable results instead of crashing the loop.
- Keep the tool list short; measure its token cost.
- Disable parallel calls for order-dependent tools.
- Require confirmation for tools that spend money, send messages or delete data.
- Log every call and result; they are your audit trail.
What this article did not test
We read documentation; we did not run the three APIs side by side, so we make no claim about which model chooses tools more accurately. Field names and defaults change, so treat the tables as a reading guide, retrieved on 2026-10-08, and confirm against the linked pages before relying on them.




