ANTM
AgentsGuide

How Tool Calling Works: Claude, OpenAI and Gemini Compared

A model never runs your function. It returns a structured request, your code runs it, and you send the result back. Here is the loop, the control knobs and the pitfalls, as documented by three vendors.

Editorial desk

Published 7 min read

A glowing teal node on the left connected by a thin teal line to a teal square and then to a set of concentric coral rings, with two small coral dots, on a dark navy background
Illustration generated with AI (FLUX.1 [schnell] (Black Forest Labs) via Cloudflare Workers AI, Apache 2.0). Prompt and direction by ANTM.

A language model cannot call an API, read a database or send an email. What it can do is emit a structured request that says "I would like this function run with these arguments". Everything else, running the code, handling failures, returning the result, is your application. This guide explains how tool calling works in the Claude, OpenAI and Gemini APIs, using each vendor's own documentation, and where the three differ. For the larger picture of what wraps this loop, see how AI agents work and how to secure them, and for the protocol layer that standardises tool access, what MCP is.

The loop, as each vendor describes it

Anthropic describes tool use as letting Claude "call functions that you define or that Anthropic provides". Claude decides when to call a tool "based on the user's request and the tool's description", and returns "a structured call that your application executes (client tools) or that Anthropic executes (server tools)."[1]

Google's guide breaks the same loop into four steps: define a function declaration, call the model with the prompt and declarations, execute the function in your own code, then send the result back so the model can produce a final response. Its wording is explicit: "The model doesn't execute the function itself."[3]

OpenAI's guide lists five high-level steps: make a request with tools, receive a tool call, run your code, send the output back, then receive a final response or further tool calls.[4]

StepClaudeOpenAIGemini
Model asks for a callResponse with stop_reason: "tool_use" and a tool_use block[1]A function call carrying call_id, name and JSON-encoded arguments[4]A function_call step with a name and arguments[3]
You return the resultA tool_result block that references the tool_use id[1]A function_call_output that references the call_id[4]A function result sent back in a new request[3]
Loop repeatsDescribed on a separate "How tool use works" page not read for this article[1]Yes, "for as many calls as a task needs"[4]Yes, across multiple turns[3]

The practical consequence is that you own the loop. You decide the maximum number of iterations, what happens when a tool throws, and what text the model sees in the result. Anthropic's SDK offers a Tool Runner that executes your tools and returns results automatically, but the underlying round trip is the same.[1]

Client tools and server tools

Anthropic draws a line that the other two guides handle differently. Client tools (your own, plus Anthropic-schema tools such as bash and text_editor) run in your application. Server tools such as web_search, web_fetch and code_execution run on Anthropic's infrastructure, so you see results without handling execution.[1] That matters for security review: a server tool moves execution out of your process, and a client tool leaves every side effect in your hands.

How a tool is defined

All three describe a tool with a name, a natural-language description and a JSON Schema for its inputs. Anthropic's example uses an input_schema object with properties and required.[1] OpenAI lists type, name, description, parameters and a strict flag; Google lists type (which must be "function"), name, description and parameters.[4][3]

The description is the model's only guide for when to call the tool. Anthropic notes that the boundary between calling a tool and answering directly "is steerable through your system prompt", and gives examples of lighter and stronger instructions to call tools.[1] Treat descriptions as prompts, and test them like prompts.

Controlling when the model calls a tool

Each API has a setting for this. The names differ, so read the one for your vendor:

BehaviourClaudeOpenAIGemini
Model decidesauto (default)[1]auto (default)[4]auto (default)[3]
Must call some toolany[1]required[4]any[3]
Must call a specific tooltool[1]Name the function[4]Not covered on the part of the page read[3]
No tool callsnone[1]none[4]none[3]
Restrict to a subsetNot covered on the pages readallowed_tools[4]allowed_tools[3]

Two cautions. First, the Claude any and tool names appear in the documentation's pricing table; the exact request syntax for forcing a tool is on a separate "Define tools" page that was not read for this article, so check it before copying code.[1] Second, Gemini lists a fourth mode, validated, described as the model ensuring schema adherence; this article does not cover how it differs from the strict modes below.[3]

Missing arguments

Anthropic's documentation includes a caveat worth remembering. If the prompt lacks a required parameter, Claude Opus "is much more likely to recognize that a parameter is missing and ask for it", while Sonnet "might also infer a reasonable value", and the page's example shows Sonnet inventing a location. It adds: "This behavior is not guaranteed."[1] The design rule that follows is our interpretation: never let a guessed argument reach an irreversible action without validation or confirmation.

Making arguments match your schema

Without enforcement, a model can return the wrong type. Anthropic's example is "2" instead of 2, or an omitted required field.[2] Two vendors offer a switch that prevents this.

  • Claude: set strict: true on the tool. Anthropic says this "guarantees Claude's tool inputs match your JSON Schema by constraining the model's token sampling to schema-valid outputs", a technique it calls grammar-constrained sampling. The computer-use and browser-use toolset entries do not accept it.[2]
  • OpenAI: strict mode requires additionalProperties: false on each object and every property listed in required, with optional fields expressed by adding null to the type. The guide says "We recommend always enabling strict mode", and notes that Chat Completions requests remain non-strict by default.[4]

Schema-valid is not the same as correct. A strict schema guarantees passengers is an integer, not that it is the right integer. Business rules, authorisation and idempotency still belong in your code.

Anthropic also documents a data-handling detail: strict tool schemas are compiled into grammars and cached for up to 24 hours since last use, so protected health information must not appear in property names, enum values, const values or patterns.[2]

Parallel and chained calls

A model can request several independent calls in one turn. OpenAI says setting parallel_tool_calls to false "ensures exactly zero or one tool is called", and lists snapshot-specific caveats, including that strict mode is disabled for fine-tuned models when they call multiple functions in one turn.[4] Anthropic's example disables parallel use with disable_parallel_tool_use: true inside the auto setting.[1] Google describes parallel calling for independent functions and compositional calling for sequences where one result informs the next, such as checking a forecast and then setting a thermostat.[3]

If your tools have side effects that depend on order, turn parallel calls off or make the tools safe to run in any order.

What it costs

Tool use is billed through ordinary tokens. Anthropic lists three sources: the tools parameter (names, descriptions and schemas), tool_use blocks, and tool_result blocks. It also adds an automatic tool-use system prompt whose size depends on the model; in its table, Claude Opus 5.5 and Sonnet 5.5 show 286 tokens for auto and none.[1] OpenAI says function definitions count toward the context limit and are "billed as input tokens", and recommends "fewer than 20 functions available at the start of a turn".[4] Prices change often; see the cost article for how to turn token counts into spend.

A long tool list therefore costs tokens on every request and can crowd the context. Both Anthropic and OpenAI document tool search features for loading rarely used tools on demand; OpenAI notes its tool search needs gpt-5.4 or later.[1][4]

State, history and reasoning models

The loop depends on history. Google notes that with store=false you must resend the full history on each request, including the original input, every model step exactly as received and the function result.[3] OpenAI states that reasoning models must have their reasoning items passed back with tool outputs, and that GPT-6 Astra and GPT-6.1 Sol require the Responses API for tool calling.[4] Both are reminders that dropping or editing a returned step can break the next call.

A checklist before production

This is editorial guidance drawn from the documented behaviour above, not a vendor recommendation:

  1. Cap loop iterations and set a timeout on every tool.
  2. Enable strict schemas where the vendor offers them, then validate values yourself.
  3. Return errors to the model as readable results instead of crashing the loop.
  4. Keep the tool list short; measure its token cost.
  5. Disable parallel calls for order-dependent tools.
  6. Require confirmation for tools that spend money, send messages or delete data.
  7. Log every call and result; they are your audit trail.

What this article did not test

We read documentation; we did not run the three APIs side by side, so we make no claim about which model chooses tools more accurately. Field names and defaults change, so treat the tables as a reading guide, retrieved on 2026-10-08, and confirm against the linked pages before relying on them.

Frequently asked questions

Does the model run the function itself?
No for client or custom tools. Google's documentation states "The model doesn't execute the function itself", and Anthropic's says your application executes client tools. Anthropic also offers server tools that run on its infrastructure.[1][3]
What is the difference between function calling and tool use?
Anthropic's documentation uses the terms interchangeably: "Tool use (also called function calling)".[1]
How do I force a model to call a tool?
Each API has a setting for it: OpenAI's tool_choice accepts required or a named function, and Gemini's accepts any. Anthropic documents tool_choice for requiring a tool call and uses any and tool as its names in the pricing table.[1][3][4]
Does tool calling cost extra?
Anthropic and OpenAI both state that tool definitions count as input tokens, and Anthropic adds a tool-use system prompt of a model-specific size.[1][4]

The ANTM newsletter

The signal, not the noise.

Sourced AI coverage in your inbox. Double opt-in, unsubscribe in one click.

Referenced sources

  1. 1.
    Tool use with Claude (overview)(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 8, 2026Accessed Oct 3, 2026

  2. 2.
    Strict tool use(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 8, 2026Accessed Oct 3, 2026

  3. 3.
    Function calling with the Gemini API(opens in a new tab)

    GooglePrimary sourcePublished Oct 8, 2026Accessed Oct 3, 2026

  4. 4.
    Function calling guide(opens in a new tab)

    OpenAIPrimary sourcePublished Oct 8, 2026Accessed Oct 3, 2026

ANTM Editorial

Editorial desk

The editorial desk at AI's Next Top Model. Every article is sourced to primary documents and approved by an editor before publication. See the editorial policy for how we work.

Concentric teal orbital rings with a dashed coral line running from a small circle to a central ellipse, marked with coral ticks
Agents

How AI Agents Work, and How to Secure Them

An agent is a model that directs its own tool use. That flexibility is useful and risky. Here is how agents differ from workflows, when to use one, and the controls that matter.

Explainer4 min read