technique
Building a server in Python
A complete MCP server with the official Python SDK: registering tools, resources, and prompts on MCPServer, schemas from type hints, and running it over stdio.
Before this
This page assumes you are comfortable with:
- techniqueDesigning toolsNaming, describing, and shaping a tool so a model picks it correctly and calls it with valid arguments: input and output schemas, structured content, and errors.
- techniqueResources and promptsThe two primitives besides tools: resources expose data the host can read by URI, and prompts are reusable templates a person picks, plus caching and completion for both.
- prerequisitePython essentials for MCPJust enough Python to read and run every sample in this cluster: functions, type hints, decorators, async and await, and virtual environments.
- prerequisiteProcesses and standard streamsWhat a running program is, and how stdin, stdout, and stderr let one program talk to another through pipes.
Why you need this
Designing a tool on paper tells you what the model should see. This page turns that design into a program that answers real MCP requests. With the official Python SDK, a working server is one file: you write ordinary Python functions, mark them with decorators, and the SDK builds every schema, parses every request, and writes every reply. This is stage 3 of a server's life, "Build and connect".
The idea
An MCP server is a program that offers tools, resources, and prompts. A host (the app a person uses) runs one client per server, and the client sends the server requests on behalf of the model, the language model inside the host. On the stdio transport the host launches your server as a child process and the two exchange one JSON message per line over the child's standard input and standard output.
The SDK splits that work in two:
| You write | The SDK does |
|---|---|
| A Python function with type hints and a docstring | Builds the tool's inputSchema and outputSchema from the hints, and its description from the docstring |
@server.tool(), @server.resource(uri), @server.prompt() |
Registers the function under its name so tools/list, resources/list, and prompts/list include it |
return a value |
Wraps it as content (text for the model) plus structuredContent (JSON matching outputSchema) |
raise ToolError("...") |
Returns a result with isError: true and your message, so the model can read it and try again |
server.run() |
Reads requests from stdin, answers server/discover, dispatches to your functions, writes replies to stdout |
The class that holds all this is MCPServer, imported with from mcp.server import MCPServer. SDK version 1 called this class FastMCP; version 2 renamed it.
Setting up
You need Python 3.12 or newer and a virtual environment, a private folder of packages for one project. Every sample on this page ran against mcp version 2.2.0.
python -m venv .venv
.venv/Scripts/python -m pip install "mcp==2.2.0"
On macOS and Linux the second line starts .venv/bin/python. With the uv tool, uv add mcp does both steps.
How type hints become a schema
A type hint is the : float or -> float you write after a parameter or a function. The SDK reads them when the decorator runs:
| Python | JSON Schema in the tool definition |
|---|---|
value: float |
{"type": "number"}, listed in required |
x: int = 3 |
{"type": "integer", "default": 3}, not required |
unit: Literal["m", "km"] |
{"type": "string", "enum": ["m", "km"]} |
-> float |
outputSchema with one property, result, of type number |
-> SomeTypedDict or a Pydantic model |
outputSchema with that object's fields |
A bare float return becomes {"result": 8.04672} in structuredContent: the SDK wraps a plain value in an object with one field, result, which also leaves room to add fields later without breaking callers. Python attribute names are snake_case (tool.input_schema, result.structured_content) while the wire format is camelCase (inputSchema, structuredContent). You will see both on this page, each in its own context.
Errors: ToolError versus a crash
A tool can fail in two ways, and the model sees them very differently.
- You raise
ToolError(from mcp.server.mcpserver.exceptions import ToolError). The result hasisError: trueand the textError executing tool convert: cannot convert mass (kg) to length (m). The model reads that and can fix its arguments. - Anything else escapes, such as a
ZeroDivisionError. The SDK treats it as a crash: the model sees onlyError executing tool divide, and the full traceback goes to the server's log on stderr. That is deliberate, since a traceback can leak file paths or secrets, but it means the model cannot learn what went wrong.
So raise ToolError for every failure the caller could correct, and let real bugs crash. Raising the SDK's MCPError instead produces a JSON-RPC protocol error rather than a tool result; save it for a request that is malformed, not for a bad input value.
Async handlers
A tool can be async def. The SDK awaits it on the event loop. A plain def tool runs in a worker thread, so a slow blocking call does not freeze other requests. Use async def when the body awaits something, such as an HTTP client or a database driver.
Logging goes to stderr
On stdio, stdout is the protocol channel. The stdio section of the 2026-07-28 specification says the server "MUST NOT write anything to its stdout that is not a valid MCP message". A stray print("debug") breaks that rule. When we tried it, the SDK's own client logged Failed to parse JSONRPC message from server and skipped the line; a stricter client may drop the connection. Use the logging module, which writes to stderr by default, or print(..., file=sys.stderr).
Worked example
The server below converts lengths and masses. It offers one tool (convert), one resource (units://all), and one prompt (recipe_to_metric). It is the same server the client and testing pages drive.
"""A small MCP server that converts between units of length and mass."""
from typing import Literal
from mcp.server import MCPServer
from mcp.server.mcpserver.exceptions import ToolError
Unit = Literal["mm", "cm", "m", "km", "in", "ft", "mi", "g", "kg", "oz", "lb"]
# How many of the base unit (meters for length, grams for mass) one unit is.
FACTORS = {
"mm": ("length", 0.001), "cm": ("length", 0.01), "m": ("length", 1.0),
"km": ("length", 1000.0), "in": ("length", 0.0254), "ft": ("length", 0.3048),
"mi": ("length", 1609.344),
"g": ("mass", 1.0), "kg": ("mass", 1000.0), "oz": ("mass", 28.349523125),
"lb": ("mass", 453.59237),
}
server = MCPServer("unit-converter")
@server.tool()
def convert(value: float, from_unit: Unit, to_unit: Unit) -> float:
"""Convert a value between two units of the same kind (length or mass)."""
kind_from, factor_from = FACTORS[from_unit]
kind_to, factor_to = FACTORS[to_unit]
if kind_from != kind_to:
raise ToolError(f"cannot convert {kind_from} ({from_unit}) to {kind_to} ({to_unit})")
return round(value * factor_from / factor_to, 6)
@server.resource("units://all")
def list_units() -> str:
"""Every unit this server knows, grouped by kind."""
lines = []
for kind in ("length", "mass"):
names = [u for u, (k, _) in FACTORS.items() if k == kind]
lines.append(f"{kind}: {', '.join(names)}")
return "\n".join(lines)
@server.prompt()
def recipe_to_metric(recipe: str) -> str:
"""Ask the model to rewrite a recipe's quantities in metric units."""
return f"Rewrite every quantity in this recipe in metric units, using the convert tool:\n\n{recipe}"
if __name__ == "__main__":
server.run()
Save it as unit_converter.py. server.run() with no argument uses stdio, so python unit_converter.py starts it and then waits silently for requests on stdin. A host normally launches it for you.
Step by step, what the decorators did:
MCPServer("unit-converter")set the name the server reports in_metaasio.modelcontextprotocol/serverInfo.@server.tool()registeredconvert. Its name is the function name; its description is the docstring.Unit = Literal[...]turned into anenumof 11 strings for bothfrom_unitandto_unit, so a model cannot ask for"stone"without the call being rejected before your code runs.-> floatproduced anoutputSchemawith one required number,result.@server.resource("units://all")and@server.prompt()registered the other two primitives the same way.
To see exactly what a host receives, we sent the server this line on stdin (a server/discover went first; every request carries the same _meta):
{
"jsonrpc": "2.0",
"id": 2,
"method": "tools/list",
"params": {
"_meta": {
"io.modelcontextprotocol/protocolVersion": "2026-07-28",
"io.modelcontextprotocol/clientInfo": {"name": "raw-probe", "version": "0.1.0"},
"io.modelcontextprotocol/clientCapabilities": {}
}
}
}
The server wrote back this line on stdout, shown pretty-printed:
{
"jsonrpc": "2.0",
"id": 2,
"result": {
"cacheScope": "private",
"resultType": "complete",
"tools": [
{
"description": "Convert a value between two units of the same kind (length or mass).",
"inputSchema": {
"type": "object",
"properties": {
"value": {"title": "Value", "type": "number"},
"from_unit": {
"enum": ["mm", "cm", "m", "km", "in", "ft", "mi", "g", "kg", "oz", "lb"],
"title": "From Unit",
"type": "string"
},
"to_unit": {
"enum": ["mm", "cm", "m", "km", "in", "ft", "mi", "g", "kg", "oz", "lb"],
"title": "To Unit",
"type": "string"
}
},
"required": ["value", "from_unit", "to_unit"],
"title": "convertArguments"
},
"name": "convert",
"outputSchema": {
"properties": {"result": {"title": "Result", "type": "number"}},
"required": ["result"],
"title": "convertOutput",
"type": "object"
}
}
],
"ttlMs": 0,
"_meta": {
"io.modelcontextprotocol/serverInfo": {"name": "unit-converter", "version": ""}
}
}
}
Every field you see came from the Python above. resultType: "complete" means this is a final answer, not a request for more input. ttlMs: 0 with cacheScope: "private" are the SDK's default caching hints: treat the list as stale at once, and do not share it between users. A server whose tools never change per user could pass cache_hints={"tools/list": CacheHint(ttl_ms=60000, scope="public")} to MCPServer (with from mcp.server.caching import CacheHint); this file keeps the defaults. version is empty because the file never passed version= to MCPServer.
Calling convert with {"value": 5, "from_unit": "mi", "to_unit": "km"} returned "content": [{"type": "text", "text": "8.04672"}], "structuredContent": {"result": 8.04672}, and "isError": false. Calling it with kilograms to meters returned "isError": true with the ToolError text, and the server wrote one line to stderr:
Tool 'convert' failed: 'Error executing tool convert: cannot convert mass (kg) to length (m)'
In a server's life
- Build and connect (stage 3) is this page. The transports page covers
server.run("streamable-http", host="127.0.0.1", port=8000), the same file served over HTTP. - Design the surface (stage 2) happens in the type hints and docstrings. A vague docstring is a vague tool description; see Designing tools.
- Test and ship (stage 5) drives this exact file through the SDK's client; see Testing MCP servers.
- Maintain it (stage 6): renaming a parameter renames a schema property, which breaks every caller. The decorator makes that easy to do by accident.
Common mistakes
print()in a stdio server. Symptom: the client logs a JSON parse failure, or disconnects, the first time the tool runs. Log to stderr instead.- Letting expected failures crash. Symptom: the model keeps retrying the same bad call because it only ever sees
Error executing tool convert. RaiseToolErrorwith a message that says what to change. - No type hints. Symptom:
inputSchemahas no useful types, so the model passes strings where you wanted numbers. Annotate every parameter. - A plain
strwhere a closed set belongs. Symptom: the model invents values like"miles"or"kilo". UseLiteral[...]so the schema carries anenum. - Writing
@server.toolwithout parentheses. Symptom: aTypeErrorat import, telling you to use@tool(). - Blocking calls inside
async def. Symptom: one slow tool call stalls every other request. Either await an async library or make the tool a plaindefso it runs in a thread.
Cost
Engineering time is small: the server above is 48 lines, and adding a tool is one decorated function. Each request costs one function call plus argument validation, small next to any real work the tool does. The larger cost is tokens: the whole convert definition, 620 characters of JSON when written compactly, is sent to the model on every turn, so long docstrings and wide enums are paid for again and again. A stdio server also costs one process per host that uses it, which is fine on a laptop and is the reason remote servers use HTTP instead.
Going further
- Serving the same file over Streamable HTTP and what changes on the wire.
Annotated[..., Resolve(...)]withElicit, the SDK's way to ask the person a question mid-call through a multi round-trip request.- Returning a Pydantic model or
TypedDictfor richerstructuredContent. - The
lifespanargument toMCPServerfor opening a database connection once at startup. - The MCP Inspector, to click through this server's tools by hand.