technique
Designing tools
Naming, describing, and shaping a tool so a model picks it correctly and calls it with valid arguments: input and output schemas, structured content, and errors.
Before this
This page assumes you are comfortable with:
- techniqueAnatomy of an MCP requestHost, client, and server roles, and what every MCP request and result carries in the stateless 2026-07-28 protocol: method, params, _meta, and resultType.
- prerequisiteJSON SchemaHow a schema describes what valid JSON looks like, so a model knows what arguments a tool takes and a server can reject bad ones.
Why you need this
A tool is the main thing an MCP server (the program offering capabilities) gives to the model (the language model inside the host, the app a person uses). The model never sees your code. It sees a name, a description, and two schemas, and from those alone it decides whether to call the tool and what to pass. A tool that works perfectly in your tests can still be called at the wrong time, with the wrong arguments, by every model that ever reads it. This is stage 2 of the server's life, and most of the long-term cost lives here.
The idea
A tool is a function plus a description a model reads. Each part of the definition does one job:
| Field | Who reads it | Job |
|---|---|---|
name |
the model, the host, logs | A stable identifier. The model writes it back when it calls the tool. |
title |
people, in the host's interface | Optional display name. |
description |
the model | When to use the tool, when not to, and what it returns. |
inputSchema |
the model and the server | A JSON Schema (a description of which JSON values are valid) for the arguments. |
outputSchema |
the host and the client | Optional JSON Schema for the structured result. |
annotations |
the host | Hints about behavior, such as "this only reads". |
The Tools section of the 2026-07-28 specification asks that names be 1 to 128 characters drawn from ASCII letters, digits, underscore, hyphen, and dot, with no spaces, and unique within the server. Those are rules for the wire. The design rules are about the reader.
Name for the intent. get_daily_forecast tells the model what it gets. weather or query tells it almost nothing. Verb plus object is a good default.
Describe when to use it, and when not to. The model chooses between your tool, other tools on your server, tools on other servers, and answering from memory. A description that only restates the name ("Weather tool.") gives it nothing to choose with. One sentence on what it does, one on when to use it, one on when not to.
One tool per user intent, not per API endpoint. If your backend needs three calls (look up the city id, fetch the forecast, convert units), the model should still see one tool. Every extra step you expose is a step the model can do in the wrong order or forget.
Constrain arguments in the schema, not in prose. If units can only be metric or imperial, say so with an enum (a fixed list of allowed values). If days must be 1 to 7, use minimum and maximum. The model reads the schema, and the server can reject a bad call before it does any work.
Return structure, not just prose. A result has two places to put data. content is a list of blocks (text, images, audio, links to resources) meant for the model to read. structuredContent is one JSON value that matches your outputSchema, meant for programs: the host can render a table from it, and a client can check it. The specification says that if an output schema is given, the server's structured results must conform to it, and that for backward compatibility a tool returning structured content should also put the serialized JSON in a text block.
Two kinds of failure. A protocol error is a JSON-RPC error object: the request itself was wrong (malformed, unknown method). A tool execution error is a normal result with isError: true and a message in content. The model reads the second kind and can try again with better arguments, so anything the model could fix ("unknown city", "date must be in the future") belongs there. The specification lists an unknown tool name as a protocol error; the SDK version used below reports it as an isError result instead, so a client should handle both.
Annotations are hints, not guarantees. readOnlyHint, destructiveHint, idempotentHint (calling twice has the same effect as once), and openWorldHint (talks to things outside the server) help a host decide when to ask the person for confirmation. A malicious server can lie in them, so the specification tells clients to treat annotations from untrusted servers as untrusted. They never replace checks in your own code.
Keep tools/list in a deterministic order. The specification asks servers to return tools in the same order when the set has not changed. Hosts put tool definitions near the start of the model's input, and a stable order lets both the client cache and the model provider's prompt cache reuse earlier work.
Worked example
Here is a bad tool and its rewrite, both registered on one server with the official Python SDK. In Python the SDK uses snake_case attribute names (input_schema, structured_content), while the wire uses camelCase (inputSchema, structuredContent). The printed definitions below are the wire form.
# forecast_tools.py
import asyncio
import json
from typing import Annotated, Literal
from pydantic import BaseModel, Field
from mcp import Client
from mcp.server import MCPServer
from mcp.server.mcpserver.exceptions import ToolError
from mcp_types import ToolAnnotations
server = MCPServer("weather")
KNOWN_CITIES = {"Des Moines": (12.0, 4.0), "Denver": (9.0, -1.0)}
# Before: vague name, one free-text argument, free-text output.
@server.tool()
def weather(q: str) -> str:
"""Weather tool."""
return f"It looks like 12 and cloudy for {q}, maybe rain later."
# After.
class DayForecast(BaseModel):
date: str = Field(description="ISO date, YYYY-MM-DD")
high: float
low: float
conditions: Literal["clear", "cloudy", "rain", "snow"]
class Forecast(BaseModel):
city: str
units: Literal["metric", "imperial"]
days: list[DayForecast]
@server.tool(annotations=ToolAnnotations(read_only_hint=True, open_world_hint=True))
def get_daily_forecast(
city: Annotated[str, Field(description="City name, e.g. 'Des Moines'. One city per call.")],
days: Annotated[int, Field(ge=1, le=7, description="Number of days, starting today.")] = 3,
units: Literal["metric", "imperial"] = "metric",
) -> Forecast:
"""Daily high, low, and conditions for one city, for up to 7 days.
Use this for questions about upcoming weather in a named city.
Do not use it for past weather or for current conditions right now.
"""
if city not in KNOWN_CITIES:
raise ToolError(f"Unknown city '{city}'. Try a nearby larger city, spelled out in full.")
high, low = KNOWN_CITIES[city]
if units == "imperial":
high, low = high * 9 / 5 + 32, low * 9 / 5 + 32
out = [DayForecast(date=f"2026-10-0{2 + i}", high=high, low=low, conditions="cloudy") for i in range(days)]
return Forecast(city=city, units=units, days=out)
def show(label, obj):
print("===", label)
print(json.dumps(obj.model_dump(by_alias=True, exclude_none=True, mode="json"), indent=2))
async def main():
async with Client(server) as client:
for tool in (await client.list_tools()).tools:
show(f"tool {tool.name}", tool)
show("ok", await client.call_tool("get_daily_forecast", {"city": "Denver", "days": 1}))
show("unknown city", await client.call_tool("get_daily_forecast", {"city": "Atlantis"}))
show("days out of range", await client.call_tool("get_daily_forecast", {"city": "Denver", "days": 10}))
show("no such tool", await client.call_tool("get_weather", {"city": "Denver"}))
asyncio.run(main())
Client(server) connects to the server in the same process, with no subprocess or network. Running python forecast_tools.py prints each definition as tools/list returns it. The bad one, in full:
{
"name": "weather",
"description": "Weather tool.",
"inputSchema": {
"type": "object",
"properties": {
"q": { "title": "Q", "type": "string" }
},
"required": ["q"],
"title": "weatherArguments"
},
"outputSchema": {
"properties": {
"result": { "title": "Result", "type": "string" }
},
"required": ["result"],
"title": "weatherOutput",
"type": "object"
}
}
Count what the model does not know. What goes in q: a city, a zip code, a sentence? Which days? Which units? The output is one string, so the host cannot draw a table and the model has to parse prose like "12 and cloudy, maybe rain later" (12 what?).
The rewrite, with the long output schema trimmed:
{
"name": "get_daily_forecast",
"description": "Daily high, low, and conditions for one city, for up to 7 days.\n\nUse this for questions about upcoming weather in a named city.\nDo not use it for past weather or for current conditions right now.\n",
"inputSchema": {
"type": "object",
"properties": {
"city": { "description": "City name, e.g. 'Des Moines'. One city per call.", "title": "City", "type": "string" },
"days": { "default": 3, "description": "Number of days, starting today.", "maximum": 7, "minimum": 1, "title": "Days", "type": "integer" },
"units": { "default": "metric", "enum": ["metric", "imperial"], "title": "Units", "type": "string" }
},
"required": ["city"],
"title": "get_daily_forecastArguments"
},
"outputSchema": {
"$defs": { "DayForecast": { ... } },
"properties": {
"city": { "title": "City", "type": "string" },
"units": { "enum": ["metric", "imperial"], "title": "Units", "type": "string" },
"days": { "items": { "$ref": "#/$defs/DayForecast" }, "title": "Days", "type": "array" }
},
"required": ["city", "units", "days"],
"title": "Forecast",
"type": "object"
},
"annotations": { "readOnlyHint": true, "openWorldHint": true }
}
Every change came from the Python: the docstring became the description, Literal[...] became an enum, Field(ge=1, le=7) became minimum and maximum, the default values became default, and the return annotation became the output schema.
Now three calls. A good one, {"city": "Denver", "days": 1}, returns both forms of the data (the text block is the same JSON as a string, trimmed here):
{
"content": [{ "type": "text", "text": "{\n \"city\": \"Denver\", ..." }],
"structuredContent": {
"city": "Denver",
"units": "metric",
"days": [{ "date": "2026-10-02", "high": 9.0, "low": -1.0, "conditions": "cloudy" }]
},
"isError": false,
"resultType": "complete"
}
An unknown city hits the ToolError, which becomes a result the model can read and act on:
{
"content": [{ "type": "text", "text": "Error executing tool get_daily_forecast: Unknown city 'Atlantis'. Try a nearby larger city, spelled out in full." }],
"isError": true,
"resultType": "complete"
}
And "days": 10 never reaches the function body: the SDK validates the arguments against the input schema first and returns isError: true with the text Error executing tool get_daily_forecast: 1 validation error for get_daily_forecastArguments\ndays\n Input should be less than or equal to 7 .... Calling get_weather, which does not exist, returned isError: true with Unknown tool: get_weather. (Each printed result also carried _meta with io.modelcontextprotocol/serverInfo, omitted here.)
In a server's life
- Design the surface (stage 2). This page. Decide the tool list, names, descriptions, and schemas before writing handlers.
- Build and connect (stage 3). The SDK turns type hints and docstrings into the JSON above; Building a server in Python shows the full flow.
- Secure it (stage 4). Annotations drive confirmation prompts, and descriptions are an attack surface; see Security threats and defenses.
- Maintain it (stage 6). The name and schemas are a contract. Renaming a tool or adding a required argument breaks clients; see Evolving a server without breaking clients.
When a tool needs something only the person can supply mid-call, such as a confirmation, that is Multi round-trip requests.
Common mistakes
- One tool per endpoint. The model calls
get_city_id, then forgets to callget_forecast, and answers from memory. Merge steps the model should never do separately. - Free-text arguments for closed choices. The model sends
"units": "celsius"and your code silently falls through to the default. Use anenum. - Raising a plain exception for a predictable failure. The SDK treats it as a crash, and the model sees only
Error executing tool get_daily_forecastwith no hint how to fix the call. RaiseToolErrorwith a message that says what to change. - Descriptions that oversell. "Use this for any weather question" makes the model call it for last week's weather and report a forecast as history. Say what it does not do.
- Trusting annotations. A host that skips confirmation because a server said
destructiveHint: falsecan be steered into deleting data. Confirm based on what you know, not what the server claims. - Shuffled tool order. Building the list from a set or a dictionary with unstable order changes the model's input on every request, so prompt caching never hits and costs rise.
Cost
Every tool definition is sent to the model as part of its input on every turn where the tool is available, so a definition costs tokens (the chunks of text a model reads and is billed for) on every request, whether or not the tool is called. Measured as compact JSON, the good definition is 774 characters without its output schema and the bad one 189, about four times larger. That is a good trade: one wrong call, with its retry and its confused answer, costs more than the extra tokens. The pressure comes from count. A host connected to several servers with dozens of tools each can spend a large share of the model's context window (the most text it can take in at once) on definitions before the person has typed anything, and every similar-looking tool is one more wrong choice available. Keep the list short, merge tools by intent, and split a large server into focused ones a host can enable separately. Results cost tokens too, and a structured result carries its data twice (once in structuredContent, once as text), so return only what the caller needs. Engineering time is front-loaded: a good name and description take minutes to write and much longer to change once clients depend on them.
Going further
- The Tools section of the 2026-07-28 specification, especially "Error Handling" and the non-normative "Stateful Tools" note on passing explicit handles between calls.
- JSON Schema keywords beyond the basics:
pattern,format, andoneOffor arguments with several valid shapes. - Evaluating tool descriptions with real models: give a model a set of questions and count how often it picks the right tool.
- The
x-mcp-headerschema extension, which mirrors a chosen argument into an HTTP header for routing.
Leads to
- techniqueBuilding a server in PythonA complete MCP server with the official Python SDK: registering tools, resources, and prompts on MCPServer, schemas from type hints, and running it over stdio.
- techniqueEvolving a server without breaking clientsChanging tools, schemas, and resources after people depend on them: which changes are safe, how to deprecate, how to tell clients the list changed, and how to follow the spec's own revisions.
- techniqueMulti round-trip requestsHow a stateless server asks for more input mid-call: returning input_required, the client gathering answers through elicitation, and the retry that completes the request.
- techniqueSecurity threats and defensesWhat goes wrong when a model is the caller: prompt injection through tool results, poisoned tool descriptions, confused deputies, over-broad permissions, and the defenses on each side.