technique
Resources and prompts
The two primitives besides tools: resources expose data the host can read by URI, and prompts are reusable templates a person picks, plus caching and completion for both.
Before this
This page assumes you are comfortable with:
Why you need this
An MCP server (the program offering capabilities) can offer three kinds of things. Tools get most of the attention, but two others solve problems tools handle badly: handing the host (the app a person uses) a piece of data to show or attach, and handing the person a ready-made request they can pick from a menu. Choosing the right primitive is part of stage 2, designing the surface, and it changes who is in control of when each thing is used.
The idea
The three primitives differ mainly in who decides when they are used:
| Primitive | Controlled by | Typical interface | Example |
|---|---|---|---|
| Tool | the model (the language model inside the host) | The model calls it mid-answer | delete_note |
| Resource | the application (the host) | A file picker, an "attach" menu, automatic context | note://12 |
| Prompt | the user | A slash command or a menu item | "Summarize note" |
Resources
A resource is a piece of data identified by a URI (a string like note://12 or file:///project/README.md: a scheme, a colon, then whatever identifies the thing). Servers may use standard schemes such as file:// or invent their own, like note://.
resources/listreturns the resources the server knows about, each with auri, aname, and optionally atitle,description,mimeType(the format, such astext/plainorimage/png), andsize.resources/templates/listreturns resource templates: URIs with holes, such asnote://{note_id}, for when there are too many resources to list one by one.resources/readtakes oneuriand returnscontents, a list of items. Each item has eithertext(for text) orblob(binary data encoded as base64, a way of writing bytes as letters and digits).
If the resource does not exist, the Resources section of the 2026-07-28 specification says the server returns a JSON-RPC error with code -32602 (Invalid params), and never an empty contents list, since an empty list could mean "exists but empty". Older servers used -32002 for this, so clients should accept both.
Prompts
A prompt is a named template the server fills in and returns as messages, ready to send to the model. prompts/list returns each prompt's name, description, and arguments (each with a name and whether it is required). prompts/get takes a name and argument values and returns messages, each with a role (user or assistant) and content (text, an image, audio, or an embedded resource). Unknown prompt names and missing required arguments are both -32602 errors.
Caching
The results of tools/list, prompts/list, resources/list, resources/templates/list, resources/read, and server/discover are cacheable. The Caching section of the specification says servers must include two hints on them:
| Field | Meaning |
|---|---|
ttlMs |
How many milliseconds the client may treat the result as fresh. 0 means stale at once. |
cacheScope |
"public": no user-specific data, any cache may share it. "private": reuse only for the same authorization context (the same access token). |
A client records when it received the result and treats it as fresh while . With ttlMs of 60000 received at 10:00:00, the cached copy serves every read until 10:01:00 and the next read after that goes back to the server. TTL is a freshness hint, not a promise that the data will not change.
Pagination
The four list methods can return a page at a time. A result carries nextCursor when there is more; the client sends it back as params.cursor. Cursors are opaque: the client must not parse or build them, and even an empty string counts as a cursor. A missing nextCursor means the end. An invalid cursor should get -32602. Each page is cached separately with its own ttlMs.
Completion
When a person fills in a prompt argument or a template variable, the host can ask the server for suggestions with completion/complete, naming the prompt (ref/prompt) or template (ref/resource), the argument, and what has been typed so far. The server returns up to 100 values, with an optional total and hasMore.
Resource or read-only tool?
Both can return the same data. The difference is who asks. Use a resource when the person or the host should choose what goes into the conversation ("attach this note"), when the data is naturally addressed by a URI, or when caching matters. Use a read-only tool when the model should decide, mid-answer, that it needs the data ("let me look up note 12"). Many servers offer both over the same function.
Worked example
A notes server with one resource template, one prompt, and completion for note ids, using the official Python SDK. In Python, attribute names are snake_case (ttl_ms, cache_scope); on the wire they are camelCase (ttlMs, cacheScope).
# notes_server.py
import asyncio
import json
from mcp import Client
from mcp.server import MCPServer
from mcp.server.caching import CacheHint
from mcp.server.mcpserver.exceptions import ResourceNotFoundError
from mcp_types import Completion, PromptReference, ResourceTemplateReference
NOTES = {
"7": "Buy flour, eggs, and milk. Pick up the bike from the shop on Friday.",
"12": "Lab report due Monday. Ask Sam for the titration data first.",
"15": "Call the dentist to move the cleaning to next week.",
}
server = MCPServer(
"notes",
cache_hints={"resources/read": CacheHint(ttl_ms=60000, scope="private")},
)
@server.resource("note://{note_id}", mime_type="text/plain")
def read_note(note_id: str) -> str:
"""One note's text, by its numeric id."""
if note_id not in NOTES:
raise ResourceNotFoundError(f"No note with id {note_id}")
return NOTES[note_id]
@server.prompt()
def summarize_note(note_id: str, style: str = "bullets") -> str:
"""Summarize one note in the chosen style (bullets or sentence)."""
return f"Summarize note {note_id} as {style}. The note says:\n\n{NOTES.get(note_id, '(missing)')}"
@server.completion()
async def complete(ref, argument, context):
if argument.name == "note_id":
matches = sorted(k for k in NOTES if k.startswith(argument.value))
return Completion(values=matches, total=len(matches), has_more=False)
if argument.name == "style":
return Completion(values=[s for s in ("bullets", "sentence") if s.startswith(argument.value)])
return None
def show(label, obj):
print("===", label)
print(json.dumps(obj.model_dump(by_alias=True, exclude_none=True, mode="json"), indent=2))
async def main():
async with Client(server) as client:
show("resources/templates/list", await client.list_resource_templates())
show("resources/read note://12", await client.read_resource("note://12"))
try:
await client.read_resource("note://99")
except Exception as exc:
print("=== resources/read note://99")
print(type(exc).__name__, exc.error.code, exc.error.message)
show("prompts/list", await client.list_prompts())
show("prompts/get", await client.get_prompt("summarize_note", {"note_id": "12"}))
show("completion/complete prompt arg",
await client.complete(PromptReference(name="summarize_note"), {"name": "note_id", "value": "1"}))
show("completion/complete template arg",
await client.complete(ResourceTemplateReference(uri="note://{note_id}"), {"name": "note_id", "value": "1"}))
asyncio.run(main())
python notes_server.py prints these results (every result also carried _meta with io.modelcontextprotocol/serverInfo, removed here). The template list:
{
"ttlMs": 0,
"cacheScope": "private",
"resourceTemplates": [
{
"name": "read_note",
"uriTemplate": "note://{note_id}",
"description": "One note's text, by its numeric id.",
"mimeType": "text/plain"
}
],
"resultType": "complete"
}
Reading note://12. On the wire the request is {"method": "resources/read", "params": {"uri": "note://12", "_meta": {...}}} and the result is:
{
"ttlMs": 60000,
"cacheScope": "private",
"contents": [
{
"uri": "note://12",
"mimeType": "text/plain",
"text": "Lab report due Monday. Ask Sam for the titration data first."
}
],
"resultType": "complete"
}
The 60000 came from the server's CacheHint. The list result shows the SDK's default: ttlMs of 0 and cacheScope of "private", the safe choice when you have not thought about it.
Reading a note that does not exist printed MCPError -32602 No note with id 99: a protocol error, as the specification requires. Raising a plain ValueError there instead produced -32603 (internal error) with a generic message, which tells the client the server crashed rather than that the note is missing.
The prompt list and one prompts/get:
{
"prompts": [
{
"name": "summarize_note",
"description": "Summarize one note in the chosen style (bullets or sentence).",
"arguments": [
{ "name": "note_id", "required": true },
{ "name": "style", "required": false }
]
}
],
...
}
{
"description": "Summarize one note in the chosen style (bullets or sentence).",
"messages": [
{
"role": "user",
"content": {
"type": "text",
"text": "Summarize note 12 as bullets. The note says:\n\nLab report due Monday. Ask Sam for the titration data first."
}
}
],
"resultType": "complete"
}
The server did not call a model. It returned a message; the host decides what to do with it, normally placing it in the conversation as if the person had typed it.
Completion, with the person having typed 1 into the note_id box, returned the same answer for the prompt argument and for the template variable:
{
"completion": { "values": ["12", "15"], "total": 2, "hasMore": false },
"resultType": "complete"
}
In a server's life
- Design the surface (stage 2). Decide for each piece of data whether the model, the host, or the person should pull it in.
- Build and connect (stage 3). The SDK decorators shown here are covered with the rest of a server in Building a server in Python; a client honoring
ttlMsis in Building a client. - Maintain it (stage 6). A host can open
subscriptions/listento hearnotifications/resources/list_changedornotifications/resources/updatedfor chosen URIs; see Transports and Evolving a server without breaking clients.
Common mistakes
- Empty
contentsfor a missing resource. The host shows a blank attachment and the person thinks the note is empty. Return-32602. "public"on per-user data. A shared gateway serves Alice's note to Bob. Use"private"for anything that depends on who is asking, and still check access on every read.- A long
ttlMson data that changes. The host keeps showing yesterday's note after an edit. Pick a TTL you can live with, or send update notifications. - Parsing cursors. A client that decodes
nextCursorto "skip ahead" breaks when the server changes the cursor format. Treat cursors as opaque. - Path traversal in
file://resources. A request forfile:///project/../../etc/passwdreads outside the folder. Normalize and check every path. - Expecting a prompt to run itself.
prompts/getreturns text, not an answer. If nothing happens, the host never sent the messages to the model.
Cost
A resources/read is one request and one result; with a positive ttlMs, repeated reads within the window cost nothing on the network. The token cost is the contents, paid only when the host actually attaches them, which is the main saving over a tool the model calls speculatively. Prompts cost tokens only when the person picks one. Completion requests arrive as the person types, so with keystrokes you can receive up to requests per field; hosts should debounce them and servers should rate-limit them, and a suggestion search over items is per request unless you index it. List pagination trades one large response for round trips for items at page size : 250 resources at 100 per page is 3 requests. Engineering cost is low: each primitive is a decorated function, and the real work is choosing which one.
Going further
- The Resources, Prompts, Caching, Completion, and Pagination sections of the 2026-07-28 specification.
- Resource annotations (
audience,priority,lastModified) for hinting how a host should use data. - Embedded resources inside tool results and prompt messages.
- URI templates (RFC 6570) for templates with more than one variable.