LLM Agent Programming from Scratch: The MCP Edition
1. What Is MCP (Model Context Protocol)
MCP (Model Context Protocol) is a communication protocol launched and open-sourced by Anthropic in 2024, designed to solve the connection problem between large language models (LLMs) and external data sources and tools. It defines the protocol for communication between the Model and external interfaces/data/Prompts. Tool/resource providers only need to implement the MCP protocol to connect with an LLM APP that implements an MCP client. During runtime, the LLM APP automatically retrieves the tool list/Prompt/resource list returned by the protocol from the MCP server via JsonRpc.
A simple example: integrating Amap so the model can query weather/route/surrounding map information through Amap’s API.
- Without MCP: you need to implement the Tools that call Amap’s OpenAPI yourself, and write the Prompt that organizes requests to the LLM yourself.
- With MCP: just use an MCP Client, fill in the Endpoint and Key of the Amap MCP Service, and the LLM will proactively query Amap-related resources through MCP during runtime, and feed back to the LLM using the Prompt that Amap has already organized.
2. What MCP defines:
The primitives MCP defines:
- Tools: FunctionCall
- Resource: resources
- Prompts: provide structured templates
- Sampling: allows the server to request that the client call the LLM
2.1 Tools:
People often compare FunctionCall and MCP, and some even lament that “why must both FunctionCall and MCP exist?” Personally, I don’t think FunctionCall and MCP conflict. FunctionCall is actually a subset of MCP, and MCP also supports FunctionCall — it’s just that MCP additionally supports definitions like Resource/Prompt, and imposes explicit protocol constraints on their retrieval/invocation/update at the protocol layer.
MCP defines the tool retrieval and invocation protocol at the protocol layer:
tools/listRetrieve the list of all tools currently provided by the MCP, mainly metadata: the tool description, the required parameters, and the schema of the output.tools/callPerform the action of invoking a tool and obtain the result.notifications/tools/list_changedUpdate the Tools information cached on the Client side via Push over a long-lived connection. MCP Server -> Client
Using it in an Agent generally follows the same approach as traditional FunctionCall: after connecting to the MCP and retrieving the Meta of all Tools, you can simply render them into the SystemPrompt.
2.2 Resources
A Resource in MCP is an application-controlled primitive that lets a server expose data and content to a client that can be read, and that content can be used as context for LLM (large language model) interactions. Resources are similar to the definition of a resource in a RESTful interface; it can be a file / database record / API response / log file.
MCP requires every resource entity to have a unique URI, in standard URL formatprotocol://host/path, and the handler has a URI — for example, if you want to expose a table in postgres as a resource, its URI ispostgres://<host>:5432/<schema>/<database>/<table>.
In MCP, the metadata of a resource is defined as:
1export interface Resource {
2 uri: string;
3 name: string;
4 description?: string;
5 mimeType?: string;
6 annotations?: Annotations;
7 size?: number;
8}
A resource is actually fetched as a whole, and that is the biggest difference from FunctionCall. For example, in the database scenario, if I implement a QueryTools it can also achieve an effect similar to a resource, but a resource puts more emphasis on returning all the resource’s information in one shot, while Tools emphasize the result obtained by performing some action, and the action may be a read or a write.
2.3 Prompt
The Prompt templates specific to different MCP Servers — as long as they are Prompts tailored to the functionality the current MCP provides — generally allow Prompts to quickly enable the LLM to better invoke the capabilities of the tools in the MCP Server.
For example, a code-refactoring Agent:
- Resource: The resources provided by its MCP Server are generally local code files, as well as standard files for code conventions.
- Tools: These are generally the results of local Lint tools checking specific files. For example, analyze_code_complexity/check_code_standards/check_code_type
- Prompt: This generally provides unique code-convention Prompts, as well as how the large model should use the tools. For example, it can provide the model with a sequence to call during refactoring: check_code_standards->check_code_standards->check_code_type
2.4 Sampling
Sampling is an MCP feature that allows a server to request an LLM completion from the client. This stands in sharp contrast to the traditional interaction pattern: normally the client requests data or functionality from the server, but with Sampling the server can proactively ask the client to invoke an LLM to generate text or perform reasoning.
Put simply, Sampling lets an MCP server “reverse” the use of the language model connected to the client, enabling more complex AI agent behavior while maintaining security and privacy controls.
The Sampling workflow follows these steps:
- Server initiates the request: the server sends a
sampling/createMessagerequest to the client - Client review: the client inspects the request and may modify it
- LLM invocation: the client calls the LLM and obtains the completion
- Client reviews the result: the client inspects the content generated by the LLM
- Return the result: the client returns the result to the server.
When requesting Sampling, the server can supply various parameters to fine-tune the LLM’s behavior:
- temperature: controls randomness (0.0 to 1.0)
- maxTokens: the maximum number of tokens to generate
- stopSequences: an array of sequences that stop generation
- metadata: additional provider-specific parameters
The server can also use the modelPreferences object to specify model selection preferences, and the systemPrompt field to request a particular system prompt, but the client ultimately decides which model to use and whether to honor the system prompt.
Sampling is especially useful in scenarios that require “agentic behavior”, that is, where the server needs the LLM’s help to complete a task. Typical use cases include:
- Git service tools: request the LLM to write a commit message based on a code diff.
- Data analysis services: request the LLM to explain data analysis results and provide insights.
- Content generation: generate text content for domain-specific tools, such as email drafts or document summaries.
- Complex decisions: request the LLM to make decision recommendations based on domain-specific data provided by the server.
A Case
sequenceDiagram
autonumber
actor User as Developer
participant Client as Git MCP Client
participant Server as Git MCP Server
participant LLM as Language Model
Note over User: Modify code and stage changes
User->>Client: Request commit message generation
Client->>Server: Call generate_commit_message tool
Note over Server: Collect context information
Server->>Server: Get staged changes (git diff --cached)
Server->>Server: Get modified file list (git status)
Server->>Server: Get recent commit history (git log)
rect rgb(240, 248, 255)
Note over Server, Client: Sampling flow begins
Server->>Client: sampling/createMessage request
Note right of Server: Includes code diff, file list<br/>Commit history and other context
Client->>User: Show sampling request and ask for confirmation
User->>Client: Confirm allow use of LLM
Client->>LLM: Call LLM API
LLM->>Client: Return generated commit message
Client->>User: Show generated commit message
User->>Client: Accept/Edit/Reject message
Client->>Server: Return sampling result (Final commit message)
Note over Server, Client: Sampling flow ends
end
Server->>Client: Return generated commit message
Client->>User: Show generated message and ask whether to commit
alt User confirms commit
User->>Client: Confirm commit
Client->>Server: Call commit_changes tool
Server->>Server: Execute git commit command
Server->>Client: Return commit result
Client->>User: Show commit success message
else User cancels
User->>Client: Cancel commit
Client->>User: Show operation cancelled
end
3. A Modern MCP Example
3.1 Sequential Thinking MCP + Using MCP Function for State Intervention
Sequential Thinking MCP is a standard MCP example. It guides the LLM through functions to think step by step and reach a conclusion. Its Tool Prompt template is:
1A detailed tool for dynamic and reflective problem-solving through thoughts.
2This tool helps analyze problems through a flexible thinking process that can adapt and evolve.
3Each thought can build on, question, or revise previous insights as understanding deepens.
4
5When to use this tool:
6- Breaking down complex problems into steps
7- Planning and design with room for revision
8- Analysis that might need course correction
9- Problems where the full scope might not be clear initially
10- Problems that require a multi-step solution
11- Tasks that need to maintain context over multiple steps
12- Situations where irrelevant information needs to be filtered out
13
14Key features:
15- You can adjust total_thoughts up or down as you progress
16- You can question or revise previous thoughts
17- You can add more thoughts even after reaching what seemed like the end
18- You can express uncertainty and explore alternative approaches
19- Not every thought needs to build linearly - you can branch or backtrack
20- Generates a solution hypothesis
21- Verifies the hypothesis based on the Chain of Thought steps
22- Repeats the process until satisfied
23- Provides a correct answer
24
25Parameters explained:
26- thought: Your current thinking step, which can include:
27* Regular analytical steps
28* Revisions of previous thoughts
29* Questions about previous decisions
30* Realizations about needing more analysis
31* Changes in approach
32* Hypothesis generation
33* Hypothesis verification
34- next_thought_needed: True if you need more thinking, even if at what seemed like the end
35- thought_number: Current number in sequence (can go beyond initial total if needed)
36- total_thoughts: Current estimate of thoughts needed (can be adjusted up/down)
37- is_revision: A boolean indicating if this thought revises previous thinking
38- revises_thought: If is_revision is true, which thought number is being reconsidered
39- branch_from_thought: If branching, which thought number is the branching point
40- branch_id: Identifier for the current branch (if any)
41- needs_more_thoughts: If reaching end but realizing more thoughts needed
42
43You should:
441. Start with an initial estimate of needed thoughts, but be ready to adjust
452. Feel free to question or revise previous thoughts
463. Don't hesitate to add more thoughts if needed, even at the "end"
474. Express uncertainty when present
485. Mark thoughts that revise previous thinking or branch into new paths
496. Ignore information that is irrelevant to the current step
507. Generate a solution hypothesis when appropriate
518. Verify the hypothesis based on the Chain of Thought steps
529. Repeat the process until satisfied with the solution
5310. Provide a single, ideally correct answer as the final output
5411. Only set next_thought_needed to false when truly done and a satisfactory answer is reached
Below, the classic “Weak-minded Bar” question — “You can’t drink it directly, and you can’t eat an apple directly, so why is it that once you wash the apple with water that you can’t drink directly, you can eat it?” — is used as the query to demonstrate the whole guidance process.
sequenceDiagram
participant User as User
participant LLM as Large Language Model
participant MCP as fwZGTnwmPNgMz0ViR9Umm8 (MCP)
User->>LLM: Question: Water cannot be drunk directly, apples cannot be eaten directly,<br/>Why can apples be eaten after washing them with water that cannot be drunk directly?
LLM->>LLM: Internal thinking: "This question involves understanding the properties of things and logical reasoning...<br/>...the reason apples can be eaten after being washed with water that cannot be drunk directly is that the purpose of washing is to clean the apples, not to change the nature of the water."
Note over LLM,MCP: Start calling the fwZGTnwmPNgMz0ViR9Umm8 tool for step-by-step thinking
LLM->>MCP: Call the fwZGTnwmPNgMz0ViR9Umm8 tool<br/>{<br/> "thought": "The main problem with water that cannot be drunk directly is that it contains substances unsuitable for direct drinking,<br/>these substances will not transfer to the apples, and the washed apples can still be safely eaten.",<br/> "nextThoughtNeeded": true,<br/> "thoughtNumber": 1,<br/> "totalThoughts": 3,<br/> "isRevision": false,<br/> "revisesThought": null,<br/> "branchFromThought": null,<br/> "branchId": null,<br/> "needsMoreThoughts": true<br/>}
MCP->>LLM: Return thinking step status<br/>{<br/> "thoughtNumber": 1,<br/> "totalThoughts": 3,<br/> "nextThoughtNeeded": true,<br/> "branches": [],<br/> "thoughtHistoryLength": 1<br/>}
LLM->>MCP: Call the fwZGTnwmPNgMz0ViR9Umm8 tool<br/>{<br/> "thought": "The main problem with water that cannot be drunk directly is that it contains substances unsuitable for direct drinking,<br/>these substances will not transfer to the apples, and the washed apples can still be safely eaten.",<br/> "nextThoughtNeeded": true,<br/> "thoughtNumber": 1, <br/> "totalThoughts": 3,<br/> "isRevision": false,<br/> "revisesThought": null,<br/> "branchFromThought": null,<br/> "branchId": null,<br/> "needsMoreThoughts": true<br/>}
MCP->>LLM: Return thinking step status<br/>{<br/> "thoughtNumber": 2,<br/> "totalThoughts": 3,<br/> "nextThoughtNeeded": true,<br/> "branches": [],<br/> "thoughtHistoryLength": 2<br/>}
LLM->>MCP: Call the fwZGTnwmPNgMz0ViR9Umm8 tool<br/>{<br/> "thought": "The main problem with water that cannot be drunk directly is that it contains substances unsuitable for direct drinking,<br/>these substances will not transfer to the apples, and the washed apples can still be safely eaten.",<br/> "nextThoughtNeeded": true,<br/> "thoughtNumber": 1,<br/> "totalThoughts": 3,<br/> "isRevision": false,<br/> "revisesThought": null,<br/> "branchFromThought": null,<br/> "branchId": null,<br/> "needsMoreThoughts": true<br/>}
MCP->>LLM: Return thinking step status<br/>{<br/> "thoughtNumber": 3,<br/> "totalThoughts": 3,<br/> "nextThoughtNeeded": false,<br/> "branches": [],<br/> "thoughtHistoryLength": 3<br/>}
Note over LLM,MCP: Thinking process completed
LLM->>User: Provide final answer: "By using the fwZGTnwmPNgMz0ViR9Umm8 tool for step-by-step thinking...<br/>...the final answer is that the water used to wash the apples does not become directly drinkable water, but the contaminants on the apple surface are removed, making them safe to eat."
4. Transport Layer
The Agent system communicates with McpClient via JsonRpc: Two modes:
- stdio pipe: The overall logic of the protocol comes from the Language Server Protocol. When the Agent starts, it launches the MCP Client by starting a subprocess, and the Agent communicates with the MCP service by sending JsonRpc messages through the stdio pipe.
- HTTP-SSE/Streamable-HTTP service: Allows remote communication over HTTP Stream. The old protocol used HTTP-SSE (HTML5), which allows bidirectional communication with the MCP Server through a long-lived HTTP Stream connection to the server. Starting in April 2025, Streamable-HTTP is supported (https://github.com/modelcontextprotocol/modelcontextprotocol/pull/206), deprecates the previous HTTP-SSE protocol. Streamable-HTTP is better optimized for compute forms like FC.
Why use such a strange HTTP-SSE approach with separate event endpoint and message endpoint:
The separation of the session establishment and messaging endpoints is intended to simplify Cross-Origin Resource Sharing (CORS). By > providing a ‘simple’ HTTP POST endpoint for message exchange, CORS preflight requests can be avoided
- MCP’s main use case is in the browser. Without separating the endpoints, the SessionID information would be carried in the HTTP headers, which does not satisfy the browser’s Simple Request requirement and would require a CORS preflight check [OPTIONS]
- By separating the endpoints, all requests can become “simple requests” and will not trigger an OPTIONS check
- Why Stream-HTTP later abandoned this approach:
a. It was decided that the performance impact of CORS preflight in modern web development is no longer a major issue b. The implementation complexity introduced by endpoint separation outweighed the benefit of avoiding preflight c. It provides a clearer session management mechanism (via the Mcp-Session-Id header) Cloudflare introduced: https://github.com/modelcontextprotocol/modelcontextprotocol/pull/206
HTTP-SSE Client Python implementation:
- An implementation using thread-synchronous programming needs a separate thread to establish an HTTP-Stream long connection with the Msg Endpoint to receive the JsonRpc return events, and to notify the main thread to harvest events via a callback function plus a queue.
sequenceDiagram
%% Define participant styles
participant Client as Client
participant MsgEndpoint as Server message endpoint
participant SSEEndpoint as Server SSE event endpoint
rect rgb(240, 240, 255)
Note over Client,MsgEndpoint: Get Endpoint
Client->>MsgEndpoint: Initiate HTTP GET request to establish connection(Stream=True)
MsgEndpoint-->>Client: Return HTTP 200 response, return the Endpoint address, URL contains unique SessionId
end
rect rgb(240, 240, 255)
Note over Client,SSEEndpoint: Initialize, get MCP metadata from server
Client->>SSEEndpoint: Rpc Post request,{"method": "initialize", "jsonrpc": "2.0", "id": 1}
SSEEndpoint-->>Client: Return HTTP 200
MsgEndpoint-->>Client: Return JsonRpc Response {"jsonrpc": "2.0", "id": 1, "result": ...}
Client->>SSEEndpoint: NotifyRPC, no return needed{"method": "notifications/initialized", "jsonrpc": "2.0"}
end
rect rgb(240, 240, 255)
Note over Client,SSEEndpoint: Get tool list
Client->>SSEEndpoint: Rpc Post request,{"method": "tools/list", "jsonrpc": "2.0", "id": 2}
SSEEndpoint-->>Client: Return HTTP 200
MsgEndpoint-->>Client: Return JsonRpc Response {"jsonrpc": "2.0", "id": 1, "result": ...}
end
rect rgb(240, 240, 255)
Note over Client,SSEEndpoint: Call tool
Client ->> SSEEndpoint: Rpc Post request,{"method": "tools/call, "jsonrpc": "2.0", "id": 3, params: {name: "get_weather",arguments: {"location": "Beijing"}}
SSEEndpoint-->>Client: Return HTTP 200
MsgEndpoint-->>Client: Return JsonRpc Response {"jsonrpc": "2.0", "id": 3, "result": ...}
end
A simple Python implementation
1import os
2import queue
3import threading
4import logging
5from typing import Any, Callable, Literal
6
7import json
8
9from urllib.parse import urljoin
10from pydantic import BaseModel, ConfigDict
11import requests
12
13logger = logging.getLogger(__name__)
14
15
16class Message(BaseModel):
17 event_type: Literal["message"]
18 data: str
19
20
21class JsonRpcHeader(BaseModel):
22 method: str
23 jsonrpc: str = "2.0"
24
25
26class JsonRpcRequest(JsonRpcHeader):
27 params: dict | None = None
28 id: int
29
30
31class JsonRpcNotify(JsonRpcHeader): ...
32
33
34type JsonRpcMessage = JsonRpcHeader | JsonRpcNotify
35
36
37class JSONRPCResponse(BaseModel):
38 """A successful (non-error) response to a request."""
39
40 jsonrpc: str
41 id: int
42 result: dict[str, Any]
43 model_config = ConfigDict(extra="allow")
44
45
46class SSEClient(threading.Thread):
47 def __init__(self, url: str) -> None:
48 super().__init__()
49 self._url = url
50 self._queue = queue.Queue()
51 self._ready_event = threading.Event()
52 self._endpoint = ""
53 self._callback = {}
54
55 def run(self) -> None:
56 print("run")
57 self._event_loop()
58
59 def wait_ready_for_endpoint(self) -> str:
60 self._ready_event.wait()
61 return self._endpoint
62
63 def register_callback(
64 self, message_type: str, callback: Callable[[Message], None]
65 ) -> None:
66 self._callback[message_type] = callback
67
68 def _event_loop(self) -> None:
69 response = requests.get(
70 self._url,
71 stream=True,
72 )
73
74 if response.status_code not in (200, 202):
75 raise Exception(f"Failed to connect to server: {response.status_code}")
76
77 event_type = None
78 event_data = None
79
80 for line in response.iter_lines(chunk_size=2, decode_unicode=True):
81 print(line)
82 if line.startswith("event:"):
83 event_type = line[6:].strip()
84 if line.startswith("data:"):
85 event_data = line[5:].strip()
86
87 if event_data is not None and event_type is not None:
88 match event_type:
89 case "message":
90 msg = Message(event_type=event_type, data=event_data)
91 if "message" in self._callback:
92 try:
93 self._callback["message"](msg)
94 except Exception as e:
95 logger.error(e)
96 else:
97 self._queue.put(msg)
98
99 case "endpoint":
100 event_data = event_data.strip()
101 self._endpoint = event_data
102 self._ready_event.set()
103
104 event_type, event_data = None, None
105
106
107class McpClient:
108 def __init__(
109 self, sse: SSEClient, endpoint: str, session: requests.Session
110 ) -> None:
111 self._endpoint = endpoint
112 self._sse = sse
113 self._sess = session
114 self.capabilities = {}
115 self._id = 0
116 self._callback_record = {}
117 self._sse.register_callback("message", self.notify_callback)
118
119 def initialize(self) -> None:
120 result: dict = self._send_request(
121 method="initialize",
122 params={
123 "protocolVersion": "2024-11-05",
124 "capabilities": {"tools": {"call": True}, "resources": {"read": True}},
125 "clientInfo": dict(name="mcp", version="0.1.0"),
126 },
127 )
128 self.capabilities = result.get("capabilities", {})
129 self._send_request(method="notifications/initialized")
130
131 def get_tools(self) -> dict:
132 return self._send_request(method="tools/list")
133
134 def call_tool(self, tool_name: str, params: dict) -> dict:
135 return self._send_request(
136 method="tools/call",
137 params={
138 "name": tool_name,
139 "arguments": params,
140 },
141 )
142
143 def notify_callback(self, message: Message) -> None:
144 response = JSONRPCResponse.model_validate(json.loads(message.data))
145 if response.id in self._callback_record:
146 self._callback_record[response.id](response)
147
148 def _send_request(self, method: str, params: dict | None = None) -> dict | None:
149
150 event = threading.Event()
151 result = None
152 if "notifications" in method:
153 data = JsonRpcNotify(
154 method=method,
155 )
156 else:
157 data = JsonRpcRequest(
158 method=method,
159 params=params,
160 id=self._id,
161 )
162
163 def get_result(item):
164 nonlocal result
165 event.set()
166 result = item.result
167
168 self._callback_record[self._id] = get_result
169 self._id += 1
170
171 print(f"request: url:{self._endpoint} data:{data.json()}")
172
173 response = self._sess.post(
174 url=self._endpoint,
175 json=data.model_dump(
176 mode="json",
177 exclude_none=True,
178 ),
179 )
180
181 if response.status_code not in (200, 202):
182 raise Exception(f"Connect Error: {response.text}")
183
184 if "notifications" in method:
185 return None
186
187 event.wait()
188 return result
189
190
191
192
193def main() -> None:
194 base = "https://mcp.amap.com/sse?key=" + os.environ["AMAP_KEY"]
195
196 sse_client = SSEClient(url=base)
197 sse_client.start()
198 endpoint = sse_client.wait_ready_for_endpoint()
199
200 mcp = McpClient(
201 sse=sse_client, endpoint=urljoin(base, endpoint), session=requests.Session()
202 )
203 mcp.initialize()
204 print("init Ok!")
205 tools = mcp.get_tools()
206 for tool in tools['tools']:
207 print(f"tool: {tool['name']}")
208 print(f" description: {tool['description']}")
209 print(f" parameters: {json.dumps(tool['inputSchema'], indent=2, ensure_ascii=False)}")
210
211 print("get tools Ok!")
212
213 # Call a tool
214 result = mcp.call_tool("maps_weather", {"city": "北京"})
215 print("call_tool Ok!")
216 print(f"result: {json.dumps(result, indent=2, ensure_ascii=False)}")
217
218 exit()
219
220
221if __name__ == "__main__":
222 main()
References
- RFC Streamable HTTP
- MCP begins supporting Streamable
- How to deploy MCP on Cloudflare Workers
- MCP Server