
What Is Tool Use in Prompts and How LLMs Parse Function Calling Schemas
Tool use in LLMs is function calling — the model emits structured JSON naming a tool and arguments. Your code executes; the model reasons over the result.
This topic is curated by our AI council — see how it works.
Once tool use is switched on, a language model stops describing the world and starts changing it — which is why every practice in prompt ops and security gets stricter the moment it’s live. A model that can query a database, hit an API, or trigger a workflow turns a wrong text output into a wrong action, so the operative question moves from whether the prose reads well to whether the right tool got called, with the right arguments, for the right reason. The five articles below split that question into the mechanism, the production failure modes, the build, the market, and who answers when it goes wrong.
Start with how LLMs parse function calling schemas — it fixes what tool use actually is at the API level, a structured JSON request rather than magic, which is the misunderstanding behind most of what follows. Read the hard limits of LLM tool calling right after: it catalogs what breaks between the demo and production — context consumed by schema definitions, parameters filled with confidently wrong values — before you’ve built anything to blame it on.
When you’re ready to build, the tool-description and pipeline guide treats the description field as the model’s actual decision spec, not documentation, and covers the message-ordering rules that throw errors if you get them wrong. For the model-selection question that guide doesn’t answer, the BFCL v3 benchmark comparison shows why the leaderboard you check determines which failures you inherit. Close with the accountability questions raised when LLMs call external APIs — once the pipeline works, this is the read for before you widen what it’s allowed to touch.

Two neighbours get folded into this topic, and each conflation hides a different problem.
Tool use is the capability that lets a model act; prompt injection is one way that capability gets hijacked into acting on an attacker’s behalf. Turning on tool calling doesn’t create new injection surface at the model level — the model was already reading attacker-controlled text — but it converts a text-generation mistake into a taken action, which is why injection defenses and tool-scoping decisions have to be designed together, not bolted on in sequence.
Tool use is also not the same problem as prompt optimization. Automated optimizers tune a prompt against a text-quality or task-success score; they were never built to judge whether a model picked the right tool with the right arguments. That judgment needs function-calling-specific benchmarks — the BFCL-versus-tau-bench split above — not a general eval number.
Q: Should I expose every tool to the model in every prompt, or scope by task? A: Scope them. Loading dozens of tool schemas into one context burns tokens and gives the model more chances to pick the wrong one; the tool-description guide recommends scoping tools to the current task and constraining choice with tool_choice subsets instead of loading everything at once.
Q: Which function-calling benchmark should I trust when picking a model for my agent? A: No single one — BFCL v3 and tau-bench disagree on the winner because they measure different things: schema accuracy versus multi-step agentic reliability. Pick the benchmark that matches your actual failure mode, not the one with the highest score for your candidate model.
Q: Does adding tool use increase my exposure to prompt injection? A: It raises the stakes without changing the surface. A model already reads attacker-controlled text before it ever calls a tool, so the injection risk exists either way — tool use just turns a hijacked response into a hijacked action, which is why scoping and permission limits matter more once tools are live.
Q: Who is accountable when a tool call does the wrong thing in production? A: Usually nobody cleanly — the accountability gap around LLMs calling external APIs is structural, not a policy oversight: once a model chooses among tools autonomously, tracing the decision back to a responsible human gets mathematically harder, not just organizationally harder.
Part of the prompt ops and security theme · closest neighbour: prompt injection.
Tool use in prompts transforms LLMs from text generators into action-capable systems by embedding function schemas directly in the context. Understanding how models parse and invoke those schemas is what separates reliable integrations from brittle ones.
Concepts covered

Tool use in LLMs is function calling — the model emits structured JSON naming a tool and arguments. Your code executes; the model reasons over the result.

LLM tool schemas vary by provider and add hidden token overhead. Task accuracy peaks near 77% (BFCL v3, 2026) — well below 99%+ schema conformance.
The practical guides cover writing effective tool descriptions, structuring function calling schemas, and handling the failure modes that appear when models select the wrong tool or return malformed parameters.
Tools & techniques

Tool descriptions determine whether LLMs call your functions correctly or hallucinate parameters. Spec pattern for Claude Sonnet 4.6 and GPT-5.5 in 2026.
Function calling capabilities are evolving rapidly across model providers, with benchmark performance diverging sharply from real-world reliability. Staying current means understanding what leaderboard scores do and don't tell you about production behavior.
Models & benchmarks
Updated September 2026

GLM 4.5 tops BFCL v3 at 0.778. Claude Sonnet 4.5 leads tau-bench Retail at 0.862. Two benchmarks, two winners — exposing a production tool-use gap.
When LLMs call external APIs, the consequences are real and often irreversible. Accountability for unverified actions, permission escalation, and cascading failures from incorrect parameter choices deserve deliberate architecture decisions before deployment.
Risks & metrics

When LLMs call external APIs, real-world consequences follow from instructions no human verified. Accountability frameworks are not designed for this gap.