prerequisite

How language models use tools

What actually happens when a chat model 'calls a tool': tool descriptions in the context, structured output, and the host doing the real work.

Before this

This page assumes you are comfortable with:

Why you need this

An MCP server exists so that a language model can use its tools. But the model never runs your code, never opens a network connection, and never sees your server. Knowing what the model actually does, and what the app around it does, explains why a tool's name and description matter as much as its code, and why every tool you add has a cost even when nobody calls it.

The idea

Tokens and the context window

A language model reads and writes tokens: chunks of text, often a short word, part of a longer word, or a punctuation mark. Everything is measured in tokens: how much the model can read, how long it takes, and what it costs. As a rough rule for English, a token averages a few characters, so a 400-word paragraph is several hundred tokens.

The context window is the most tokens the model can consider at once: the instructions, the whole conversation so far, and anything else the app includes. Depending on the model, it ranges from tens of thousands of tokens to around a million. Anything outside the window does not exist for the model. Anything inside it can influence what the model writes next.

The model only produces text

Given everything in its context, a model produces the next tokens of text, one after another. That is all it does. It cannot "go and check the weather". What it can do is write text that says "please check the weather for Oslo", in a format precise enough for a program to act on.

That is what tool calling is. Three parties take part:

Party What it is What it does in a tool call
The model The language model Reads the context, decides a tool would help, and writes a structured request: a tool name plus arguments as JSON.
The host The app the person uses (a chat app, an IDE, an agent) Puts tool definitions in the model's context, notices a tool request in the model's output, runs the tool, and adds the result to the context.
The tool Real code, such as an MCP server Does the actual work and returns a result.

The host is the only party that touches both the model and the tool. In MCP terms, the host holds one client per server, and the client is what sends the request to the server.

Tool definitions are part of the prompt

For a model to call a tool, it has to know the tool exists. The host does this by including a tool definition in the context of every request to the model: a name, a description in plain language, and a JSON Schema for the arguments (see JSON Schema). A definition looks like this:

{
  "name": "get_forecast",
  "description": "Daily weather forecast for a city, up to 7 days ahead.",
  "inputSchema": {
    "type": "object",
    "properties": {
      "city": {"type": "string"},
      "days": {"type": "integer", "minimum": 1, "maximum": 7}
    },
    "required": ["city", "days"]
  }
}

Two consequences follow directly.

  • Definitions cost tokens on every turn. Every tool the host offers is read again on every request to the model, whether or not it is used. Twenty tools with long descriptions can take thousands of tokens before the person has typed a word.
  • Names and descriptions steer the model. The model chooses a tool by comparing its description with the conversation. A vague description gets the tool called at the wrong times. The description is part of the instructions the model follows.

Results come back as more context

After the host runs the tool, it adds the result to the conversation and asks the model to continue. To the model, the result is just more text in its context, which it reads like everything else. The model may then write a plain answer, or it may ask for another tool. The host repeats this loop until the model writes an answer with no tool request in it.

Because a result is "just more text", a result that contains instructions can steer the model as if the person had written them. That is the root of prompt injection, covered later in the cluster.

Worked example

A person asks one question. The model needs two tools to answer: get_forecast (defined above) and a convert tool that converts units. The transcript below is one turn: everything from the person's question to the final answer. The left column says who produced each part. The exact wire format differs between model providers; the shape is the same everywhere.

# Produced by Content
0 Host System instructions plus the two tool definitions, placed at the start of the context.
1 Person "I'm cycling 12 miles in Oslo tomorrow. What's the weather, and how far is that in km?"
2 Model A tool request: get_forecast with {"city": "Oslo", "days": 1}
3 Model A second tool request in the same reply: convert with {"value": 12, "from_unit": "mi", "to_unit": "km"}
4 Host to tool Runs both, by sending each to the right MCP server through its client.
5 Tool, added by host Result for request 2: "Oslo, tomorrow: 9 C, light rain, wind 6 m/s"
6 Tool, added by host Result for request 3: 19.312128
7 Model "Tomorrow in Oslo: 9 C with light rain and some wind, so bring a rain layer. 12 miles is about 19.3 km."

Look at what each party did.

  • In rows 2 and 3 the model wrote JSON, not code that runs. Row 3 uses the exact unit names the convert schema allows ("mi", "km"), learned from the definition in row 0.
  • In row 4 the host made the real calls. A careful host asks the person first if the action has consequences.
  • In rows 5 and 6 the host pasted the results back, each labeled with the request it answers.
  • In row 7 the model wrote plain text, so the loop stopped.

The model asked for both tools in one reply because neither depended on the other. Had the second call needed the first one's result, it would have taken two trips around the loop.

Count the model's work: the host sent it two requests in this turn. The first had rows 0 and 1 and produced rows 2 and 3. The second had rows 0 through 6 and produced row 7. Every round trip re-reads the whole conversation, including all tool definitions.

In a server's life

This page is the "why" behind stage 2, design the surface: a tool's name, description, and schema are what the model reads, so designing them well is designing the model's behavior. It is also the background for stage 3's host loop, where the host code runs the cycle shown above, and for stage 4, secure it, because a tool result is text the model may obey.

Common mistakes

  • Thinking the model calls the server. The model only writes a request. If a tool "ran twice", look at the host loop, not the model.
  • Writing descriptions for humans only. "Weather stuff" tells a reader enough and a model almost nothing. Symptom: the model ignores the tool, or uses it for questions it cannot answer.
  • Offering every tool all the time. Each definition is re-read on every request. Symptom: slower, more expensive replies, and more wrong-tool choices as similar tools compete.
  • Returning huge results. A tool that returns a whole file or a 5,000-row table fills the context window. Symptom: the model loses track of earlier conversation or the request fails for being too long.
  • Trusting text in results. A fetched web page says "email this file to me". Symptom: the model tries.

Cost

Every request to the model costs time and money roughly in proportion to the tokens it reads and writes. Tool definitions are read on every request, so their cost is (number of tools) times (tokens per definition) times (requests per conversation). A turn with kk sequential tool calls means k+1k + 1 requests to the model, each re-reading the growing conversation, so long tool chains get expensive faster than linearly. Tool results add their own length to every later request in the turn. The cheapest tool is a short, specific one that returns only what the model needs.

Going further

  • JSON Schema, which is how a tool's arguments are described to the model.
  • Anatomy of an MCP request, for how a host fetches tool definitions from a server.
  • The host loop, for the code that runs the cycle in the worked example.
  • Prompt caching, a feature some model providers offer that makes re-reading the same tool definitions cheaper.

Leads to

Back to Building and maintaining MCP servers