Tool Use in Prompts

Authors 5 articles 59 min total read

This topic is curated by our AI council — see how it works.

Once tool use is switched on, a language model stops describing the world and starts changing it — which is why every practice in prompt ops and security gets stricter the moment it’s live. A model that can query a database, hit an API, or trigger a workflow turns a wrong text output into a wrong action, so the operative question moves from whether the prose reads well to whether the right tool got called, with the right arguments, for the right reason. The five articles below split that question into the mechanism, the production failure modes, the build, the market, and who answers when it goes wrong.

  • Tool use turns a text-generation model into a system that can trigger real side effects — writing tool descriptions well and validating parameters matters more than prompt wording alone.
  • Function-calling leaderboards don’t agree on a winner: BFCL v3 and tau-bench measure different things, so picking a model by one score alone is a bet on the wrong axis.
  • Wrong-parameter tool calls are a documented, hard limit of current tool calling, not a rare edge case — the fix is architectural (validation, scoping, gating), not a better system prompt.
  • Once a model can act, accountability for what it did becomes a design decision, not an afterthought — governance has not caught up to autonomous tool calling.

The tool-use reading path: mechanism, limits, then the stakes

Start with how LLMs parse function calling schemas — it fixes what tool use actually is at the API level, a structured JSON request rather than magic, which is the misunderstanding behind most of what follows. Read the hard limits of LLM tool calling right after: it catalogs what breaks between the demo and production — context consumed by schema definitions, parameters filled with confidently wrong values — before you’ve built anything to blame it on.

When you’re ready to build, the tool-description and pipeline guide treats the description field as the model’s actual decision spec, not documentation, and covers the message-ordering rules that throw errors if you get them wrong. For the model-selection question that guide doesn’t answer, the BFCL v3 benchmark comparison shows why the leaderboard you check determines which failures you inherit. Close with the accountability questions raised when LLMs call external APIs — once the pipeline works, this is the read for before you widen what it’s allowed to touch.

MONA asks: 'If the schema validates, why did the tool call still do the wrong thing?' MAX answers: 'Validation checks shape, not judgment — the model picked the wrong tool or the wrong moment, and no schema catches that.' — comic dialog.
A schema-valid call can still be a bad decision.

How tool use differs from prompt injection and prompt optimization

Two neighbours get folded into this topic, and each conflation hides a different problem.

Tool use is the capability that lets a model act; prompt injection is one way that capability gets hijacked into acting on an attacker’s behalf. Turning on tool calling doesn’t create new injection surface at the model level — the model was already reading attacker-controlled text — but it converts a text-generation mistake into a taken action, which is why injection defenses and tool-scoping decisions have to be designed together, not bolted on in sequence.

Tool use is also not the same problem as prompt optimization. Automated optimizers tune a prompt against a text-quality or task-success score; they were never built to judge whether a model picked the right tool with the right arguments. That judgment needs function-calling-specific benchmarks — the BFCL-versus-tau-bench split above — not a general eval number.

Common questions about tool use in prompts

Q: Should I expose every tool to the model in every prompt, or scope by task? A: Scope them. Loading dozens of tool schemas into one context burns tokens and gives the model more chances to pick the wrong one; the tool-description guide recommends scoping tools to the current task and constraining choice with tool_choice subsets instead of loading everything at once.

Q: Which function-calling benchmark should I trust when picking a model for my agent? A: No single one — BFCL v3 and tau-bench disagree on the winner because they measure different things: schema accuracy versus multi-step agentic reliability. Pick the benchmark that matches your actual failure mode, not the one with the highest score for your candidate model.

Q: Does adding tool use increase my exposure to prompt injection? A: It raises the stakes without changing the surface. A model already reads attacker-controlled text before it ever calls a tool, so the injection risk exists either way — tool use just turns a hijacked response into a hijacked action, which is why scoping and permission limits matter more once tools are live.

Q: Who is accountable when a tool call does the wrong thing in production? A: Usually nobody cleanly — the accountability gap around LLMs calling external APIs is structural, not a policy oversight: once a model chooses among tools autonomously, tracing the decision back to a responsible human gets mathematically harder, not just organizationally harder.

Part of the prompt ops and security theme · closest neighbour: prompt injection.

1

Understand the Fundamentals

Tool use in prompts transforms LLMs from text generators into action-capable systems by embedding function schemas directly in the context. Understanding how models parse and invoke those schemas is what separates reliable integrations from brittle ones.

2

Build with Tool Use in Prompts

The practical guides cover writing effective tool descriptions, structuring function calling schemas, and handling the failure modes that appear when models select the wrong tool or return malformed parameters.

4

Risks and Considerations

When LLMs call external APIs, the consequences are real and often irreversible. Accountability for unverified actions, permission escalation, and cascading failures from incorrect parameter choices deserve deliberate architecture decisions before deployment.