technique

Building a server in Python

A complete MCP server with the official Python SDK: registering tools, resources, and prompts on MCPServer, schemas from type hints, and running it over stdio.

Before this

This page assumes you are comfortable with:

Why you need this

Designing a tool on paper tells you what the model should see. This page turns that design into a program that answers real MCP requests. With the official Python SDK, a working server is one file: you write ordinary Python functions, mark them with decorators, and the SDK builds every schema, parses every request, and writes every reply. This is stage 3 of a server's life, "Build and connect".

The idea

An MCP server is a program that offers tools, resources, and prompts. A host (the app a person uses) runs one client per server, and the client sends the server requests on behalf of the model, the language model inside the host. On the stdio transport the host launches your server as a child process and the two exchange one JSON message per line over the child's standard input and standard output.

The SDK splits that work in two:

You write The SDK does
A Python function with type hints and a docstring Builds the tool's inputSchema and outputSchema from the hints, and its description from the docstring
@server.tool(), @server.resource(uri), @server.prompt() Registers the function under its name so tools/list, resources/list, and prompts/list include it
return a value Wraps it as content (text for the model) plus structuredContent (JSON matching outputSchema)
raise ToolError("...") Returns a result with isError: true and your message, so the model can read it and try again
server.run() Reads requests from stdin, answers server/discover, dispatches to your functions, writes replies to stdout

The class that holds all this is MCPServer, imported with from mcp.server import MCPServer. SDK version 1 called this class FastMCP; version 2 renamed it.

Setting up

You need Python 3.12 or newer and a virtual environment, a private folder of packages for one project. Every sample on this page ran against mcp version 2.2.0.

python -m venv .venv
.venv/Scripts/python -m pip install "mcp==2.2.0"

On macOS and Linux the second line starts .venv/bin/python. With the uv tool, uv add mcp does both steps.

How type hints become a schema

A type hint is the : float or -> float you write after a parameter or a function. The SDK reads them when the decorator runs:

Python JSON Schema in the tool definition
value: float {"type": "number"}, listed in required
x: int = 3 {"type": "integer", "default": 3}, not required
unit: Literal["m", "km"] {"type": "string", "enum": ["m", "km"]}
-> float outputSchema with one property, result, of type number
-> SomeTypedDict or a Pydantic model outputSchema with that object's fields

A bare float return becomes {"result": 8.04672} in structuredContent: the SDK wraps a plain value in an object with one field, result, which also leaves room to add fields later without breaking callers. Python attribute names are snake_case (tool.input_schema, result.structured_content) while the wire format is camelCase (inputSchema, structuredContent). You will see both on this page, each in its own context.

Errors: ToolError versus a crash

A tool can fail in two ways, and the model sees them very differently.

  • You raise ToolError (from mcp.server.mcpserver.exceptions import ToolError). The result has isError: true and the text Error executing tool convert: cannot convert mass (kg) to length (m). The model reads that and can fix its arguments.
  • Anything else escapes, such as a ZeroDivisionError. The SDK treats it as a crash: the model sees only Error executing tool divide, and the full traceback goes to the server's log on stderr. That is deliberate, since a traceback can leak file paths or secrets, but it means the model cannot learn what went wrong.

So raise ToolError for every failure the caller could correct, and let real bugs crash. Raising the SDK's MCPError instead produces a JSON-RPC protocol error rather than a tool result; save it for a request that is malformed, not for a bad input value.

Async handlers

A tool can be async def. The SDK awaits it on the event loop. A plain def tool runs in a worker thread, so a slow blocking call does not freeze other requests. Use async def when the body awaits something, such as an HTTP client or a database driver.

Logging goes to stderr

On stdio, stdout is the protocol channel. The stdio section of the 2026-07-28 specification says the server "MUST NOT write anything to its stdout that is not a valid MCP message". A stray print("debug") breaks that rule. When we tried it, the SDK's own client logged Failed to parse JSONRPC message from server and skipped the line; a stricter client may drop the connection. Use the logging module, which writes to stderr by default, or print(..., file=sys.stderr).

Worked example

The server below converts lengths and masses. It offers one tool (convert), one resource (units://all), and one prompt (recipe_to_metric). It is the same server the client and testing pages drive.

"""A small MCP server that converts between units of length and mass."""
from typing import Literal

from mcp.server import MCPServer
from mcp.server.mcpserver.exceptions import ToolError

Unit = Literal["mm", "cm", "m", "km", "in", "ft", "mi", "g", "kg", "oz", "lb"]

# How many of the base unit (meters for length, grams for mass) one unit is.
FACTORS = {
    "mm": ("length", 0.001), "cm": ("length", 0.01), "m": ("length", 1.0),
    "km": ("length", 1000.0), "in": ("length", 0.0254), "ft": ("length", 0.3048),
    "mi": ("length", 1609.344),
    "g": ("mass", 1.0), "kg": ("mass", 1000.0), "oz": ("mass", 28.349523125),
    "lb": ("mass", 453.59237),
}

server = MCPServer("unit-converter")


@server.tool()
def convert(value: float, from_unit: Unit, to_unit: Unit) -> float:
    """Convert a value between two units of the same kind (length or mass)."""
    kind_from, factor_from = FACTORS[from_unit]
    kind_to, factor_to = FACTORS[to_unit]
    if kind_from != kind_to:
        raise ToolError(f"cannot convert {kind_from} ({from_unit}) to {kind_to} ({to_unit})")
    return round(value * factor_from / factor_to, 6)


@server.resource("units://all")
def list_units() -> str:
    """Every unit this server knows, grouped by kind."""
    lines = []
    for kind in ("length", "mass"):
        names = [u for u, (k, _) in FACTORS.items() if k == kind]
        lines.append(f"{kind}: {', '.join(names)}")
    return "\n".join(lines)


@server.prompt()
def recipe_to_metric(recipe: str) -> str:
    """Ask the model to rewrite a recipe's quantities in metric units."""
    return f"Rewrite every quantity in this recipe in metric units, using the convert tool:\n\n{recipe}"


if __name__ == "__main__":
    server.run()

Save it as unit_converter.py. server.run() with no argument uses stdio, so python unit_converter.py starts it and then waits silently for requests on stdin. A host normally launches it for you.

Step by step, what the decorators did:

  1. MCPServer("unit-converter") set the name the server reports in _meta as io.modelcontextprotocol/serverInfo.
  2. @server.tool() registered convert. Its name is the function name; its description is the docstring.
  3. Unit = Literal[...] turned into an enum of 11 strings for both from_unit and to_unit, so a model cannot ask for "stone" without the call being rejected before your code runs.
  4. -> float produced an outputSchema with one required number, result.
  5. @server.resource("units://all") and @server.prompt() registered the other two primitives the same way.

To see exactly what a host receives, we sent the server this line on stdin (a server/discover went first; every request carries the same _meta):

{
  "jsonrpc": "2.0",
  "id": 2,
  "method": "tools/list",
  "params": {
    "_meta": {
      "io.modelcontextprotocol/protocolVersion": "2026-07-28",
      "io.modelcontextprotocol/clientInfo": {"name": "raw-probe", "version": "0.1.0"},
      "io.modelcontextprotocol/clientCapabilities": {}
    }
  }
}

The server wrote back this line on stdout, shown pretty-printed:

{
  "jsonrpc": "2.0",
  "id": 2,
  "result": {
    "cacheScope": "private",
    "resultType": "complete",
    "tools": [
      {
        "description": "Convert a value between two units of the same kind (length or mass).",
        "inputSchema": {
          "type": "object",
          "properties": {
            "value": {"title": "Value", "type": "number"},
            "from_unit": {
              "enum": ["mm", "cm", "m", "km", "in", "ft", "mi", "g", "kg", "oz", "lb"],
              "title": "From Unit",
              "type": "string"
            },
            "to_unit": {
              "enum": ["mm", "cm", "m", "km", "in", "ft", "mi", "g", "kg", "oz", "lb"],
              "title": "To Unit",
              "type": "string"
            }
          },
          "required": ["value", "from_unit", "to_unit"],
          "title": "convertArguments"
        },
        "name": "convert",
        "outputSchema": {
          "properties": {"result": {"title": "Result", "type": "number"}},
          "required": ["result"],
          "title": "convertOutput",
          "type": "object"
        }
      }
    ],
    "ttlMs": 0,
    "_meta": {
      "io.modelcontextprotocol/serverInfo": {"name": "unit-converter", "version": ""}
    }
  }
}

Every field you see came from the Python above. resultType: "complete" means this is a final answer, not a request for more input. ttlMs: 0 with cacheScope: "private" are the SDK's default caching hints: treat the list as stale at once, and do not share it between users. A server whose tools never change per user could pass cache_hints={"tools/list": CacheHint(ttl_ms=60000, scope="public")} to MCPServer (with from mcp.server.caching import CacheHint); this file keeps the defaults. version is empty because the file never passed version= to MCPServer.

Calling convert with {"value": 5, "from_unit": "mi", "to_unit": "km"} returned "content": [{"type": "text", "text": "8.04672"}], "structuredContent": {"result": 8.04672}, and "isError": false. Calling it with kilograms to meters returned "isError": true with the ToolError text, and the server wrote one line to stderr:

Tool 'convert' failed: 'Error executing tool convert: cannot convert mass (kg) to length (m)'

In a server's life

  • Build and connect (stage 3) is this page. The transports page covers server.run("streamable-http", host="127.0.0.1", port=8000), the same file served over HTTP.
  • Design the surface (stage 2) happens in the type hints and docstrings. A vague docstring is a vague tool description; see Designing tools.
  • Test and ship (stage 5) drives this exact file through the SDK's client; see Testing MCP servers.
  • Maintain it (stage 6): renaming a parameter renames a schema property, which breaks every caller. The decorator makes that easy to do by accident.

Common mistakes

  • print() in a stdio server. Symptom: the client logs a JSON parse failure, or disconnects, the first time the tool runs. Log to stderr instead.
  • Letting expected failures crash. Symptom: the model keeps retrying the same bad call because it only ever sees Error executing tool convert. Raise ToolError with a message that says what to change.
  • No type hints. Symptom: inputSchema has no useful types, so the model passes strings where you wanted numbers. Annotate every parameter.
  • A plain str where a closed set belongs. Symptom: the model invents values like "miles" or "kilo". Use Literal[...] so the schema carries an enum.
  • Writing @server.tool without parentheses. Symptom: a TypeError at import, telling you to use @tool().
  • Blocking calls inside async def. Symptom: one slow tool call stalls every other request. Either await an async library or make the tool a plain def so it runs in a thread.

Cost

Engineering time is small: the server above is 48 lines, and adding a tool is one decorated function. Each request costs one function call plus argument validation, small next to any real work the tool does. The larger cost is tokens: the whole convert definition, 620 characters of JSON when written compactly, is sent to the model on every turn, so long docstrings and wide enums are paid for again and again. A stdio server also costs one process per host that uses it, which is fine on a laptop and is the reason remote servers use HTTP instead.

Going further

  • Serving the same file over Streamable HTTP and what changes on the wire.
  • Annotated[..., Resolve(...)] with Elicit, the SDK's way to ask the person a question mid-call through a multi round-trip request.
  • Returning a Pydantic model or TypedDict for richer structuredContent.
  • The lifespan argument to MCPServer for opening a database connection once at startup.
  • The MCP Inspector, to click through this server's tools by hand.

Leads to

Back to Building and maintaining MCP servers