1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Tool Surface Design

When an action deserves its own tool instead of bash, how to write a description that raises selection accuracy, and how to scale past a few dozen tools without destroying the prompt cache.

Last updated

After this section you can

  • Decide when an action deserves its own tool, using reversibility, UI, audit and parallelism as tests
  • Write a tool description and schema that raise selection accuracy and narrow invalid arguments
  • Scale a large tool surface with tool search and deferred loading without breaking the prompt cache
07

Tool Surface Design: What Deserves a Tool

One bash tool can do almost anything, which is exactly why your harness can do almost nothing with it. Which actions get their own tool, how you describe them and how many you load decide what the model picks and what your code can control.

Key idea

The tool surface is an interface in two directions. The model reads descriptions and schemas to choose a call; your harness reads the call to decide whether to allow it. Design every tool for both readers.

Same delete, two tool surfaces. Only the dedicated tool lets your harness see what is about to happen.
VIA BASH · ONE OPAQUE STRING bash(cmd="psql -c 'DELETE FROM orders WHERE id = 42'") Your harness sees a string cannot require approval for deletes only cannot render a real confirmation cannot tell a safe read from a write, so runs all in series DEDICATED TOOL · TYPED ARGUMENTS delete_order(order_id="42") Your harness sees an action and its arguments gate it behind human approval show “Delete order 42?” with the order details log it for audit; run read-only tools in parallel Start broad for reach. Promote an action to its own tool when you need to gate, render, audit or parallelise it.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium