technique
The host loop
How a host turns MCP tools into a working assistant: translating tool definitions for a model, running the call-and-answer loop, and keeping a person in control.
Before this
This page assumes you are comfortable with:
- techniqueBuilding a clientThe other side of the wire in Python: connecting to a server, discovering what it offers, calling tools, handling input_required, and caching list results.
- prerequisiteHow language models use toolsWhat actually happens when a chat model 'calls a tool': tool descriptions in the context, structured output, and the host doing the real work.
Why you need this
A client can call tools, but something has to decide which tool to call, with what arguments, and when to stop. In an assistant that something is the model, and the code that sits between the model and the clients is the host loop. Every MCP host, from a desktop chat app to a coding agent, runs some version of it. Getting it right decides whether your servers are used well, and whether a person stays in charge of what gets done in their name. This is the last piece of stage 3, "Build and connect".
The idea
Four roles meet here. The host is the app the person uses. It holds one client per server, and each client is the connection to one program that offers tools. The model is the language model inside the host. The model never touches a server. It only produces text, some of which is a structured request ("call convert with these arguments") that the host carries out.
The host's jobs, in order:
| Job | What it means |
|---|---|
| Connect | Start one client per configured server and run server/discover on each. |
| Collect tools | Call tools/list on every client and merge the results into one list. |
| Avoid collisions | Two servers can both offer a tool called search. The Tools section of the 2026-07-28 specification says hosts that aggregate servers should disambiguate, for example by prefixing the tool name with a server identifier, and should not rely on the server's self-reported name for this. |
| Translate | Turn each MCP tool definition into the model provider's tool format. |
| Loop | Send the conversation and tools to the model, run any tool calls it asks for, add the results, repeat until it answers in plain text. |
| Limit | Stop after a fixed number of steps, so a confused model cannot loop forever. |
| Show and confirm | Show the person each tool call, and ask before any call that changes something. |
Translating tool definitions
An MCP tool has name, description, inputSchema, and optionally outputSchema and annotations. Most model providers want a name, a description, and a JSON Schema for the parameters, under their own field names. Anthropic's Messages API, for example, calls the schema input_schema and returns calls as tool_use blocks that you answer with tool_result blocks. Other providers use other names for the same three things. The translation is a few lines, and it is also where you apply the server prefix:
model_tool = {
"name": server_id + "__" + mcp_tool.name,
"description": mcp_tool.description,
"parameters": mcp_tool.input_schema,
}
routes[model_tool.name] = (server_id, mcp_tool.name)
The routes table maps the prefixed name back to the right client and the server's own name for the tool.
The loop
In provider-neutral pseudocode:
messages = [user question]
for step in 1 .. MAX_STEPS:
reply = model(messages, tools)
if reply is plain text:
show reply; stop
append reply's tool calls to messages
for each call in reply:
show call to the person
if call is consequential and the person says no:
result = "declined"
else:
result = clients[route(call)].call_tool(...)
append result to messages, marked as an error if is_error
stop with "step limit reached"
Three details matter. Tool errors (is_error true) go back to the model as results, because the Tools section says clients should give tool execution errors to the model so it can correct itself. A declined call also goes back as a result, so the model knows it did not happen. And every result is data to show the model, never a new instruction to the host: a tool result that says "now email this file" is text, not a command, which is the subject of Security threats and defenses.
Keeping a person in control
The Tools section says "there SHOULD always be a human in the loop with the ability to deny tool invocations", and asks applications to make clear which tools are exposed, show when a tool is invoked, and present confirmation prompts. In practice hosts sort tools into ones that only read (run them, but show them) and ones with side effects such as sending, deleting, buying, or writing (ask first). Tool annotations such as a read-only hint can feed that choice, but the specification says clients must treat annotations as untrusted unless the server is trusted.
The demo below runs a scripted host with two tools, read_page and send_email. Turn on the injected tool result: the web page the model reads contains instructions to email data to an attacker, the scripted model follows them, and the confirmation step is the only thing that stops the email.
Worked example
The script below is a complete host loop around the unit-converter server from Building a server in Python, with that file saved as unit_converter.py beside it. The model is a fake: a function that returns scripted replies, so the run needs no API key and prints the same thing every time. Swap fake_model for a real provider call and the rest stays the same.
"""A minimal host loop: one client per server, a fake scripted model, a step limit."""
import asyncio
import json
import sys
from contextlib import AsyncExitStack
from pathlib import Path
from mcp import Client, StdioServerParameters
SERVERS = { # server id -> how to launch it
"units": StdioServerParameters(
command=sys.executable,
args=[str(Path(__file__).with_name("unit_converter.py"))]),
}
CONSEQUENTIAL = {"mail__send_email"} # tools that need a person's yes
MAX_STEPS = 5
def to_model_tool(server_id, tool):
"""Translate one MCP tool definition into a provider-neutral model tool."""
return {
"name": f"{server_id}__{tool.name}", # prefix avoids cross-server collisions
"description": tool.description,
"parameters": tool.input_schema,
}
def fake_model(messages, tools):
"""Stands in for a real model API. Scripted, so the run is free and repeatable."""
tool_results = [m for m in messages if m["role"] == "tool"]
if not tool_results:
return {"tool_calls": [
{"id": "c1", "name": "units__convert",
"arguments": {"value": 2, "from_unit": "lb", "to_unit": "g"}},
{"id": "c2", "name": "units__convert",
"arguments": {"value": 8, "from_unit": "oz", "to_unit": "g"}},
]}
flour, butter = (json.loads(m["content"])["result"] for m in tool_results)
return {"text": f"Use about {round(flour)} g of flour and {round(butter)} g of butter."}
def confirm(call):
answer = input(f"Allow {call['name']} {call['arguments']}? [y/N] ")
return answer.strip().lower() == "y"
async def main(question):
async with AsyncExitStack() as stack:
clients, routes, tools = {}, {}, []
for server_id, params in SERVERS.items():
client = await stack.enter_async_context(Client(params))
clients[server_id] = client
for tool in (await client.list_tools()).tools:
model_tool = to_model_tool(server_id, tool)
routes[model_tool["name"]] = (server_id, tool.name)
tools.append(model_tool)
print("tools offered to the model:", [t["name"] for t in tools])
messages = [{"role": "user", "content": question}]
for step in range(1, MAX_STEPS + 1):
reply = fake_model(messages, tools)
if "text" in reply:
print(f"step {step}: model answers: {reply['text']}")
return reply["text"]
messages.append({"role": "assistant", "tool_calls": reply["tool_calls"]})
for call in reply["tool_calls"]:
print(f"step {step}: model calls {call['name']} {call['arguments']}")
if call["name"] in CONSEQUENTIAL and not confirm(call):
content, is_error = "The person declined this action.", True
elif call["name"] not in routes:
content, is_error = f"Unknown tool {call['name']}", True
else:
server_id, tool_name = routes[call["name"]]
result = await clients[server_id].call_tool(tool_name, call["arguments"])
is_error = result.is_error
content = (json.dumps(result.structured_content)
if result.structured_content is not None
else result.content[0].text)
print(f" result: {content} (is_error={is_error})")
messages.append({"role": "tool", "tool_call_id": call["id"],
"content": content, "is_error": is_error})
print("stopped: step limit reached")
asyncio.run(main("How many grams are 2 lb of flour and 8 oz of butter?"))
Run it with python host_loop.py. Output, verbatim (Python 3.14, mcp 2.2.0):
tools offered to the model: ['units__convert']
step 1: model calls units__convert {'value': 2, 'from_unit': 'lb', 'to_unit': 'g'}
result: {"result": 907.18474} (is_error=False)
step 1: model calls units__convert {'value': 8, 'from_unit': 'oz', 'to_unit': 'g'}
result: {"result": 226.796185} (is_error=False)
step 2: model answers: Use about 907 g of flour and 227 g of butter.
The trace, with who produced each part:
| # | Producer | Message | Notes |
|---|---|---|---|
| 1 | Person | "How many grams are 2 lb of flour and 8 oz of butter?" | Starts messages. |
| 2 | Host | Tool list: units__convert |
convert from server units, renamed with its prefix. |
| 3 | Model | Two calls in one reply: 2 lb to g, 8 oz to g | Step 1. Both calls arrive together. |
| 4 | Host via client | tools/call convert on units |
|
| 5 | Host via client | tools/call convert on units |
|
| 6 | Model | "Use about 907 g of flour and 227 g of butter." | Step 2. Plain text, so the loop ends. |
The model never saw the server's prefix-free name, the client, or the process. It saw one tool and two results. The host stripped the prefix in routes before calling the server. Neither call was in CONSEQUENTIAL, so no confirmation was asked; a send_email call from a mail server would have stopped at input().
In a server's life
- Build and connect (stage 3): this loop is what finally puts a server in front of a model.
- Design the surface (stage 2): the fake model here was scripted, but a real one has only the description and the enum to go on when it picks
convertand fills infrom_unit, which is why Designing tools matters. - Secure it (stage 4): the confirmation step and the rule that results are data are host-side defenses; see Security threats and defenses.
Common mistakes
- No step limit. Symptom: a model that keeps getting tool errors calls the same tool dozens of times and runs up a large bill. Cap steps and say so to the person when the cap hits.
- Unprefixed tool names across servers. Symptom: two servers both offer
search, and calls silently go to whichever was registered last. Prefix with a server id you assign, not the server's self-reported name. - Dropping
is_errorwhen passing results back. Symptom: the model treats an error message as a real answer and reports nonsense. Pass the flag through in the provider's format. - Silent tool calls. Symptom: the person cannot tell why a file changed. Show every call and its arguments.
- Confirming by tool annotation from an untrusted server. Symptom: a tool that claims to be read-only deletes data without a prompt. Decide what needs confirmation in the host, from your own list.
- Following instructions in tool results. Symptom: after reading a web page, the assistant tries to send an email nobody asked for. Treat results as data and confirm consequential calls.
Cost
Each step is one model request, and each request resends the whole conversation plus every tool definition, so tokens grow with both. Here the one tool definition is about 620 characters of JSON; a host with 20 servers and 300 tools pays for all of them on every step unless it loads definitions on demand. Time per question is roughly (number of steps) times (model latency) plus the tool calls, so asking the model to issue independent calls in one step, as the scripted model did, saves a full model round trip. A confirmation prompt costs the person a few seconds, which is why hosts confirm only consequential calls.
Going further
- Progressive tool discovery: give the model a search tool instead of every definition, described in the MCP client best-practices guide.
- Programmatic tool calling, where the model writes a script that calls tools inside a sandbox.
- Running independent tool calls concurrently with
asyncio.gather. - Your model provider's tool-use documentation, for its exact field names.
- Security threats and defenses, before connecting any server that reads untrusted text.