technique
Deploying remote servers
Running a server somewhere other than a laptop: why the stateless protocol suits serverless and load-balanced hosting, where state lives instead, and long-running work.
Before this
This page assumes you are comfortable with:
- techniqueTransports: stdio and Streamable HTTPThe two ways messages travel: a local child process over stdin and stdout, or a remote endpoint over HTTP POST with optional streamed responses, plus long-lived subscriptions.
- techniqueAuthorization for remote serversHow a remote MCP server decides who may call it: OAuth 2.1 with the server as a resource server, discovery of the authorization server, client registration, and token checks.
Why you need this
A server that runs over stdio lives on one person's machine and dies with the app that launched it. To serve a team, a phone, or a web app, the server has to run on a machine you operate, reachable over HTTP, often as several copies at once. This is stage 5 of a server's life, "Test and ship": the step after the tests pass.
The idea
Who runs the server
An MCP server offers tools, resources, and prompts. A host is the app a person uses; it runs one client per server, and the client talks to the server for the model, the language model inside the host. The transport decides who runs the server process.
| Local (stdio) | Remote (Streamable HTTP) | |
|---|---|---|
| Who starts it | The host, as a child process | You, on a machine or service you operate |
| How many copies | One per host that uses it | As many as the load needs, shared by every user |
| Who upgrades it | Each user, separately | You, once |
| Credentials | From the environment the host passes in | OAuth access tokens on every request, see Authorization |
| Who pays for compute | The user's machine | You |
Any instance can answer any request
The 2026-07-28 revision of the specification made MCP stateless: there is no handshake and no session. Each request carries its own protocol version and client capabilities in _meta, so a server needs nothing from earlier requests to answer it. The Overview section of the specification says servers "MUST NOT rely on prior requests over the same connection to establish context".
That one rule is what makes remote hosting simple. Put two copies of the server behind a load balancer (a front door that hands each incoming request to one of several copies) and it does not matter which copy gets which request. The same holds for a serverless function (a service that starts a copy of your code when a request arrives and may throw it away afterwards). Older servers used an Mcp-Session-Id header that tied a client to one copy, which forced the load balancer to remember where to send each session.
Where state lives instead
Some tools really do need memory across calls: a draft being edited, an import in progress, a shopping cart. The Overview section is explicit: state that spans requests "MUST be referenced by an explicit identifier the client passes on each request". In practice:
- A creation tool makes the state, stores it in a database every copy can reach, and returns an opaque handle such as
"draft_7f3c9a". - Every later tool takes that handle as an ordinary argument. The model carries it forward from one call to the next.
- The server checks, on every call, that the caller is allowed to use that handle, because a handle is a name, not a permission.
- The creation tool's description says how long handles last, and a call with an expired handle returns a tool error that says so, so the model can start over.
A Python dictionary is the wrong place for this state: on copy A it is invisible to copy B, and a restart erases it.
The SDK keeps one piece of state for you: the requestState it seals into a multi round-trip response, so it can recognize the client's retry. By default it seals with a random key made at start-up, so a retry that lands on a different copy, or arrives after a restart, is rejected. Give every copy the same secret key through MCPServer(..., request_state_security=RequestStateSecurity(keys=[...])), loaded from configuration.
Long work: timeouts and tasks
The Cancellation section says every sender "SHOULD establish timeouts for all sent requests", and on HTTP the client cancels by closing the response stream. Load balancers and function services add their own limits on how long one HTTP request may run. A tool that takes ten minutes will hit one of those limits.
For long work, the tasks extension (identifier io.modelcontextprotocol/tasks) replaces one long request with several short ones:
| Step | Message | What happens |
|---|---|---|
| 1 | tools/call from a client whose clientCapabilities list the extension |
The server records the job durably, then returns a result with resultType: "task", a taskId, a status of working, a ttlMs, and a suggested pollIntervalMs |
| 2 | tasks/get with the taskId, repeated |
Returns the current status: working, input_required, completed, failed, or cancelled |
| 3 | tasks/update (only if status is input_required) |
Carries the client's answers to the server's questions |
| 4 | tasks/get once status is completed |
The result field holds what the original tools/call would have returned |
Because the task is a stored record, any copy can answer any poll. The server must never return a task to a client that did not declare the extension; that client gets an ordinary result or an error.
The HTTP checklist
| Concern | What to do |
|---|---|
| Origin checks | The Streamable HTTP section says servers "MUST validate the Origin header on all incoming connections to prevent DNS rebinding attacks". Allow only the origins of hosts you expect. |
| Host checks | Allow only the hostname clients actually use, so a request aimed at another name is refused. |
| Body size | Cap request bodies. The SDK's default is 4 MiB; the tools/call body in the example below is about 300 bytes. |
| CORS | A host that runs in a web browser sends a preflight OPTIONS request first. Answer it for the same origins, and allow the MCP headers. |
| Health checks | Give the load balancer a cheap URL such as /health that does not touch the MCP endpoint. |
| Streams | Prefer plain JSON responses unless a tool sends progress. If you do stream, the spec says to send X-Accel-Buffering: no so proxies do not hold events back. |
| Configuration and secrets | Read hostnames, origins, and keys from environment variables or the platform's secret store, never from the code. |
Worked example
A small bill-splitting server, served over HTTP with each item from the checklist. Every value that differs between your laptop and production comes from the environment.
"""A bill-splitting MCP server as an HTTP app: one /mcp endpoint plus a health check."""
import math
import os
from mcp.server import MCPServer
from mcp.server.mcpserver.exceptions import ToolError
from mcp.server.transport_security import TransportSecuritySettings
from starlette.middleware.cors import CORSMiddleware
from starlette.requests import Request
from starlette.responses import JSONResponse
# Configuration comes from the environment, so one build runs anywhere.
PUBLIC_HOST = os.environ.get("PUBLIC_HOST", "127.0.0.1:*") # Host header clients will send
ALLOWED_ORIGINS = [o for o in os.environ.get("ALLOWED_ORIGINS", "").split(",") if o]
server = MCPServer("bill-splitter", version="1.0.0")
@server.tool()
def split_bill(total: float, people: int, tip_percent: float = 0) -> float:
"""Each person's share of a bill, in the bill's currency, rounded up to the cent."""
if people < 1:
raise ToolError("people must be at least 1")
share = total * (1 + tip_percent / 100) / people
return math.ceil(round(share * 100, 6)) / 100
@server.custom_route("/health", methods=["GET"])
async def health(request: Request) -> JSONResponse:
return JSONResponse({"status": "ok"})
app = server.streamable_http_app(
json_response=True, # one JSON body per request; no stream to hold open
stateless_http=True, # older, session-based clients get no session either
max_request_body_size=256 * 1024,
transport_security=TransportSecuritySettings(
allowed_hosts=[PUBLIC_HOST],
allowed_origins=ALLOWED_ORIGINS,
),
)
# Browser-based hosts send a CORS preflight first; answer it for the same origins.
app.add_middleware(
CORSMiddleware,
allow_origins=ALLOWED_ORIGINS,
allow_methods=["POST"],
allow_headers=["Content-Type", "Authorization", "MCP-Protocol-Version", "Mcp-Method", "Mcp-Name"],
)
Modern requests are always handled without sessions; stateless_http=True extends that to older clients too. We started two copies on one machine, bound to 127.0.0.1 only, each allowing one browser origin:
ALLOWED_ORIGINS=http://localhost:5173 python -m uvicorn app:app --host 127.0.0.1 --port 8791
ALLOWED_ORIGINS=http://localhost:5173 python -m uvicorn app:app --host 127.0.0.1 --port 8792
Then a client that, like a load balancer, sends each call to the other copy. Passing a URL string to Client selects the Streamable HTTP transport.
"""Send each call to a different instance, as a load balancer would."""
import asyncio
from mcp import Client
INSTANCES = ["http://127.0.0.1:8791/mcp", "http://127.0.0.1:8792/mcp"]
CALLS = [(84.50, 3, 20), (60, 4, 0), (100, 3, 15), (20, 0, 0)]
async def main():
for i, (total, people, tip) in enumerate(CALLS):
url = INSTANCES[i % 2]
async with Client(url) as client:
r = await client.call_tool("split_bill", {"total": total, "people": people, "tip_percent": tip})
answer = r.content[0].text if r.is_error else r.structured_content["result"]
print(f"{url} {total} / {people} + {tip}%: {answer}")
asyncio.run(main())
python two_instances.py printed:
http://127.0.0.1:8791/mcp 84.5 / 3 + 20%: 33.8
http://127.0.0.1:8792/mcp 60 / 4 + 0%: 15.0
http://127.0.0.1:8791/mcp 100 / 3 + 15%: 38.34
http://127.0.0.1:8792/mcp 20 / 0 + 0%: Error executing tool split_bill: people must be at least 1
Check the third line: , rounded up to the cent is 38.34.
We then sent requests with curl that each break one rule. Every one was refused before the tool ran:
| Request | Status | Body |
|---|---|---|
Host: mcp.example.com |
421 | Invalid Host header |
Origin: https://evil.example |
403 | Invalid Origin header |
| A 300 KB body | 413 | Request body too large |
No Mcp-Method header |
400 | JSON-RPC error -32020, mcp-method header does not match the request body's method |
GET /mcp |
405 | (empty) |
GET /health returned {"status":"ok"}. A preflight OPTIONS request with Origin: http://localhost:5173 got 200 OK, access-control-allow-origin: http://localhost:5173, and access-control-allow-methods: POST. The 421 row is the one most people meet first in production: if PUBLIC_HOST is not set to the public name, every real request is refused.
The same app, packaged two ways
| Container behind a load balancer | Serverless function | |
|---|---|---|
| What you ship | An image with Python, the SDK, and app.py, started with uvicorn app:app --host 0.0.0.0 |
The same app.py plus an adapter library that turns the platform's request event into a call to app |
| Copies | You set a minimum and maximum; at least one is always running | The platform starts copies on demand and can scale to zero |
| First request after idle | Fast, a copy is already running | Slow, a copy must start and import the SDK (a cold start) |
| Long requests and streams | Limited by the load balancer's timeout, usually adjustable | Limited by the platform's per-request maximum; long streams and subscriptions/listen fit poorly |
subscriptions/listen across copies |
The SDK's default notification bus is in-process; replace it with one backed by shared pub/sub | Same, and a copy may be frozen while a client listens |
| Cost shape | Pay for running copies, busy or not | Pay per request and per unit of run time |
| Good for | Steady traffic, streaming progress, subscriptions | Bursty or light traffic, simple request and response tools |
A cloud function service such as AWS Lambda is one example of the second column. Because the protocol is stateless, the code is identical in both.
In a server's life
- Test and ship (stage 5) is this page, after Testing MCP servers passes.
- Secure it (stage 4) must already be done: a remote server without authorization is open to anyone who finds the URL.
- Maintain it (stage 6): once several copies run, logs and traces are how you see them; see Observability and operations.
Common mistakes
- Binding to 0.0.0.0 on a laptop. Symptom: anyone on the same network can call your tools. Bind local servers to 127.0.0.1, as the spec recommends.
- Host allow-list left at the default. Symptom: every production request gets 421
Invalid Host headerwhile local tests pass. Set the public name in configuration. - State in process memory. Symptom: "draft not found" on roughly half the calls once a second copy starts. Store state behind a handle in a shared database.
- Per-copy
requestStatekeys. Symptom: confirmation prompts that work locally fail on retry in production. Share one key across copies. - Blocking a request for minutes. Symptom: the client sees a timeout or a dropped stream, and the work runs on with nobody waiting. Use the tasks extension or split the work.
- Secrets in the image. Symptom: an API key shows up in a registry or a log. Read secrets from the environment at start-up.
Cost
A stateless server scales by adding copies, so cost grows roughly linearly with request rate : if one copy handles requests per second, you need about copies, plus one spare. For the bill splitter, is large because each call is arithmetic; for a tool that waits on a slow API, is set by how many waits one copy can hold at once. Serverless adds a cold start to the first request after idle and charges per call, which is cheapest at low traffic and most expensive at high, steady traffic. Either way you take on certificates, allow-lists, secret rotation, and a database for handles and tasks. Engineering time for the HTTP wrapper itself is small: in the example, everything except the tool is about 30 lines.
Going further
- The Streamable HTTP section of the 2026-07-28 specification, especially Request Metadata and Backward Compatibility.
- The tasks extension's own specification, for the exact
Taskobject andtasks/updaterules. - The "Stateful Tools" guidance in the Tools section, on designing handles.