technique

Testing MCP servers

How to know a server works before a model touches it: unit tests for handlers, an in-process client for protocol tests, schema contract tests, and an interactive inspector.

Before this

This page assumes you are comfortable with:

Why you need this

A server bug rarely shows up as a crash on your machine. It shows up later, as a model that keeps calling a tool wrong, or a client that stops working after you renamed one argument. Tests catch those before any model or person sees them. This is stage 5 of a server's life, "Test and ship": the last check before the server runs somewhere other than your laptop.

The idea

An MCP server offers tools, resources, and prompts. A client is one connection to it, usually run by a host (the app a person uses) on behalf of the model (the language model inside the host). A server has three things worth testing, and each needs a different kind of test.

Layer What it checks How
Handler Your function computes the right answer and raises the right errors Call the Python function directly, no MCP at all
Protocol The SDK turns your function into the right messages: schemas, structuredContent, isError Drive the server with the SDK's own client, in the same process or over stdio
Contract What clients depend on (tool names, argument names, types, output shape) has not changed by accident Save the tools/list result to a file and compare against it on every run

Handler tests

With the official Python SDK, @server.tool() registers your function and hands it back unchanged. So convert(5, "mi", "km") is still an ordinary call you can test like any function. These tests are fast and give the clearest failure messages. They cannot tell you anything about schemas or what the wire looks like.

Protocol tests

The SDK's Client class accepts a server object directly: Client(server) connects in memory, with no subprocess and no network. Every request still goes through the SDK's real request handling: argument validation against inputSchema, error wrapping, resultType. To test the server the way a desktop host will run it, pass StdioServerParameters(command=..., args=[...]) instead and the client launches the server as a child process. The stdio version is slower and catches different bugs, such as a stray print() that corrupts stdout. Run most tests in memory and a few over stdio.

Python attribute names are snake_case (result.structured_content, result.is_error) while the wire format is camelCase (structuredContent, isError). Tests use the Python names.

Snapshot tests of tools/list

A snapshot test saves an output once, after you have checked it by eye, and fails whenever the output differs. Applied to tools/list, it turns the tool definitions into a contract. Rename an argument, add a required one, change a return type, or reword a description, and the test fails with a diff. The failure is not a verdict that the change is wrong. It forces a decision: if the change is intended, you rewrite the snapshot and commit it, and the diff in code review shows everyone what clients will see. Evolving a server covers which of those diffs break clients.

Error-path tests

Two error shapes reach a client. A tool execution error is a normal result with isError: true and a message the model can read. A protocol error is a JSON-RPC error object, for a malformed request. Test that every failure a caller could fix produces the first kind, with a message that says what to change. A message the model cannot act on is a bug even though nothing crashed.

Manual exploration

The MCP Inspector connects to a server over stdio or HTTP, lists everything it offers, and lets you call each tool with arguments you type, showing the raw messages. Use it while building, and to reproduce a bug report. It does not replace tests, because nothing it shows you is checked again tomorrow.

Evaluation is not testing

Whether a real model picks the right tool, with the right arguments, for a real question is a question about descriptions, not code. Answering it means running prompts through a model and grading the outcomes, which is slow, costs tokens, and gives a different answer on each run. That is evaluation: a score over many runs, not a pass or fail. It is still worth doing, because the tool description is the one part of the server the model actually reads, and a description change can shift how often the right tool is chosen without any test noticing. Keep a small set of realistic questions with the tool call you expect, run it when descriptions change, and track the success rate over time.

Worked example

The server under test is the unit-converter from Building a server in Python: one tool, convert(value, from_unit, to_unit), that raises ToolError when the two units measure different things. The file unit_converter.py sits next to the test file, unchanged.

"""Three tests for unit_converter.py, one per layer."""
import json
import os
from pathlib import Path

import pytest
from mcp import Client
from mcp.server.mcpserver.exceptions import ToolError

from unit_converter import convert, server

SNAPSHOT = Path(__file__).with_name("tools_list.snapshot.json")


def test_convert_as_a_plain_function():
    assert convert(5, "mi", "km") == 8.04672
    assert convert(1, "lb", "g") == 453.59237
    with pytest.raises(ToolError, match="cannot convert mass"):
        convert(1, "kg", "m")


@pytest.mark.asyncio
async def test_convert_over_the_protocol():
    async with Client(server) as client:
        ok = await client.call_tool("convert", {"value": 5, "from_unit": "mi", "to_unit": "km"})
        bad = await client.call_tool("convert", {"value": 1, "from_unit": "kg", "to_unit": "m"})
    assert ok.result_type == "complete"
    assert not ok.is_error
    assert ok.structured_content == {"result": 8.04672}
    assert bad.is_error
    assert bad.content[0].text == "Error executing tool convert: cannot convert mass (kg) to length (m)"


@pytest.mark.asyncio
async def test_tools_list_matches_snapshot():
    async with Client(server) as client:
        listing = await client.list_tools()
    tools = [t.model_dump(mode="json", by_alias=True, exclude_none=True) for t in listing.tools]
    if os.environ.get("UPDATE_SNAPSHOTS") == "1":
        SNAPSHOT.write_text(json.dumps(tools, indent=2) + "\n")
        pytest.skip("snapshot rewritten; review the diff before committing it")
    assert tools == json.loads(SNAPSHOT.read_text())

What each test does:

  1. Handler. Plain function calls. Five miles is 8.04672 km, one pound is 453.59237 g, and kilograms to meters raises ToolError.
  2. Protocol. The same two conversions through Client(server). The success has result_type == "complete" and structured_content == {"result": 8.04672}, because the SDK wraps a bare float return in an object with one field, result. The failure has is_error set and the SDK's prefix Error executing tool convert: in front of the message. The asserts sit after the async with block, so a failure prints a plain assertion instead of being wrapped in an exception group from the client's shutdown.
  3. Contract. model_dump(by_alias=True) turns each tool back into its camelCase wire form. The first run, with UPDATE_SNAPSHOTS=1, writes tools_list.snapshot.json; every later run compares against it.

Install pytest and pytest-asyncio next to the SDK, write the snapshot once, then run the tests:

python -m pip install pytest pytest-asyncio
UPDATE_SNAPSHOTS=1 python -m pytest -q test_unit_converter.py
python -m pytest -v test_unit_converter.py

On Windows PowerShell, set the variable with $env:UPDATE_SNAPSHOTS = "1" first. The second command printed:

collected 3 items

test_unit_converter.py::test_convert_as_a_plain_function PASSED          [ 33%]
test_unit_converter.py::test_convert_over_the_protocol PASSED            [ 66%]
test_unit_converter.py::test_tools_list_matches_snapshot PASSED          [100%]

============================== 3 passed in 0.94s ==============================

The snapshot file holds the one tool's full definition: its name, description, the inputSchema with both 11-unit enums and "required": ["value", "from_unit", "to_unit"], and the outputSchema with one required number, result.

To see the contract test earn its keep, we added a required argument, digits: int, to convert in a copy of the server and ran only that test:

>       assert tools == json.loads(SNAPSHOT.read_text())
E       AssertionError: assert [{'name': 'co...': 'object'}}] == [{'name': 'co...': 'object'}}]
E         
E         At index 0 diff: {... 'digits': {'title': 'Digits', 'type': 'integer'}}, 'required': ['value', 'from_unit', 'to_unit', 'digi...
...
FAILED test_unit_converter.py::test_tools_list_matches_snapshot - AssertionEr...
1 failed in 1.02s

The handler and protocol tests would still pass after that change, since both supply every argument. Only the snapshot noticed that every existing client, which sends three arguments, would now be rejected.

In a server's life

  • Test and ship (stage 5) is this page. Run the tests on every change, before deploying the server.
  • Maintain it (stage 6): the snapshot is where a breaking change first becomes visible. Evolving a server explains how to tell a safe diff from a breaking one.
  • Design the surface (stage 2): evaluation with real models is how you find out a description is unclear, which no unit test can.

Common mistakes

  • Only testing handlers. Symptom: every test passes, but a host gets an argument named fromUnit where the schema says from_unit, or no structuredContent at all. Add at least one protocol test per tool.
  • Rewriting snapshots without reading them. Symptom: a release breaks clients even though "all tests passed". Treat a snapshot diff like a code diff and review it.
  • Asserting inside async with Client(...). Symptom: a failure shows up as a long ExceptionGroup traceback instead of a one-line assertion. Collect results inside the block, assert after it.
  • Testing only the happy path. Symptom: the model loops on the same bad call because the error text is Error executing tool convert with no reason. Assert on the exact error message.
  • Calling real downstream services in tests. Symptom: tests fail when the network is down, or slowly drain an API quota. Replace the downstream call with a stub at the handler layer.
  • Counting evaluation as a test. Symptom: CI fails at random because a model chose differently today. Run evaluation on its own schedule and track a rate.

Cost

The three tests above ran in about one second in total. With pytest's --durations=0 flag, the protocol test took 0.08 seconds and the other two under 0.005 seconds each; the rest of the second went to collecting the tests, which includes importing the SDK. Hundreds of in-memory tests are affordable on every save. Stdio tests add a process launch each, roughly the time to import Python and the SDK, so keep them few. The snapshot costs one file per server and a minute of review whenever the tool list changes. Evaluation is the expensive layer: every question costs at least one model call, each carrying every tool definition as input tokens, and you need dozens of questions run several times each before a change in success rate means anything.

Going further

  • Parametrized pytest tests, to run every unit pair through the handler in one test function.
  • Running the same protocol test over stdio with StdioServerParameters.
  • Snapshot-testing resources/list and prompts/list the same way.
  • The MCP Inspector, for clicking through a server by hand.
  • Building a small evaluation set: realistic questions, the expected tool call, and a success rate per release.

Leads to

Back to Building and maintaining MCP servers