Chapter 9
Chapter 9 . RAG and MCP . MCP basics

MCP: one protocol, every AI client

An LLM cannot read your Jira board or your test-case export on its own, and pasting data into the chat goes stale and burns the context window. The Model Context Protocol fixes the plumbing: you write one server, and any MCP client can discover and call it. This chapter is the concepts; chapter 10 builds the server.

14
Sections in MCP.md
From the problem MCP solves to a failure-mode table.
3
Server primitives
Tools the model calls, resources the app reads, prompts the user picks.
2
Transports
stdio for local servers, Streamable HTTP for remote ones. Same messages on both.
40 to 13
Integrations
5 AI clients x 8 data sources need 40 without MCP and 5 + 8 with it.

01Why MCP, and why a tester cares

"Paste more into the prompt" fails in three ways that every QA team has seen:

FailureWhat it looks like in QA work
Context burnA 5,000-row test-case CSV fills the context window before the model has answered anything.
StalenessThe pasted copy is frozen. Someone edits the sheet, and the answer is now wrong without any warning.
N x M integrations5 AI clients and 8 data sources make 40 one-off integrations, each rewritten when either side changes.

MCP is an open protocol for the connection between an LLM application and a data source or tool. Write one server and every MCP client (Claude Desktop, Claude Code, Cursor or your own app) can use it, so N x M becomes N + M. The chapter compares it to the Language Server Protocol, which did the same for editors and programming languages.

  • Your data becomes queryable, not pasted. The model asks for 3 matching test cases instead of reading 5,000.
  • It stays current. The server reads the live source on every call.
  • It is a testable surface. A server has inputs, outputs and error paths, so every tool, resource and prompt is something you can write test cases against.

02Host, client, server

MCP splits an AI application into three roles. The split is what makes the security model work.

flowchart LR
  subgraph HOST["Host: Claude Desktop, Claude Code, Cursor"]
    H["Host: conversation, consent, the LLM"]
    H --> C1[Client 1]
    H --> C2[Client 2]
    H --> C3[Client 3]
  end
  C1 -->|1:1| S1["Server 1: files and Git"]
  C2 -->|1:1| S2["Server 2: test-case CSV"]
  C3 -->|1:1 over the internet| S3["Server 3: Jira or another API"]
  S1 --- D1[(local data A)]
  S2 --- D2[(local data B)]
  S3 --- D3[(remote data C)]
Client-host-server: the host owns the conversation and the consent prompts; each client talks to exactly one server, which is what keeps servers from reading each other.
RoleWhat it isWhat it is responsible for
HostThe AI application the person usesCreates and manages clients, enforces security policy, asks for consent, runs the LLM, combines context
ClientA connector inside the hostTalks to exactly one server, routes its messages, keeps servers isolated from each other
ServerYour programExposes tools, resources and prompts. Can be a local subprocess or a remote service
The 1:1 rule is the security boundary. A server sees only what it is handed. The full conversation stays with the host, and one server cannot see into another. The chapter lists this among the spec's design principles, next to "servers should be extremely easy to build" and "highly composable".

03Live demo: step through a real session

This is a real MCP session between a small client and the chapter 10 server (server.py on fastmcp 3.4.4), recorded line by line over stdio. Press Next message to send each one, click any message to read it, and pick which tool the client calls. Then switch Server stdout to see what one stray print() does.

A recorded MCP session, one message at a timeNo API key needed
data-testid=mcp-tooldata-testid=mcp-stdoutdata-testid=mcp-stepdata-testid=mcp-rundata-testid=mcp-resetdata-testid=mcp-msg-1data-testid=mcp-message
Messages (click one to inspect it)
    Server stderr (timestamps removed)
    
        
    Message, pretty-printed
    
          
    The same message as one line on the wire
    
        
    Requests and replies pair up by id. Select a reply and its request gets a dashed outline. The notification has no id because nothing answers it.

    04Tools, resources, prompts: who pulls the trigger

    This is the distinction people get wrong. Ask who decides to use the capability:

    PrimitiveWho triggers itLikeExample from the course
    ToolThe model, mid-conversationa function callsearch_test_cases(query, limit)
    ResourceThe application, by URIa file read or a GETtestcases://schema
    PromptThe user, from a menua slash command/review_test_case VWO-1001
    flowchart LR
      M["The LLM decides, mid-conversation"] -->|tools/call| S[MCP server]
      A["The client app fetches by URI"] -->|resources/read| S
      U["The user picks from a menu"] -->|prompts/get| S
      S --> D[(Your data)]
    One question sorts every capability: who triggers it? The model calls tools, the application reads resources, the user picks prompts.
    • Tools take arguments the model works out, and may have side effects. The model reads the tool's name, description and input schema to decide whether to call it, so those fields are functional code. A vague description gives you a tool the model never calls, or calls wrongly.
    • Resources are context the application pulls by URI, like reading a file: no side effects. One templated URI can serve many documents, so pair it with a plain resource that lists the valid values.
    • Prompts are templates the user picks, usually as slash commands: "review this test case", "build a regression suite for this module".
    MCP.md (resource URIs)
    testcases://schema                 # a fixed resource
    testcases://modules                # a fixed resource: the list of valid module names
    testcases://module/{name}          # a templated resource: 17 modules, one declaration
    Rule of thumb from the chapter: if it needs arguments the model has to work out, it is a tool.

    05The wire: JSON-RPC 2.0 over a transport

    MCP messages are JSON-RPC 2.0 in UTF-8. If you have tested a REST API this is familiar, with different plumbing: instead of a URL path and a verb you send a method and a params object, and you match the reply to the request by id. The chapter's example of a tool call:

    MCP.md (a tools/call request)
    {
      "jsonrpc": "2.0",
      "id": 7,
      "method": "tools/call",
      "params": {
        "name": "search_test_cases",
        "arguments": { "query": "invite user", "limit": 3 }
      }
    }
    MCP.md (the reply)
    {
      "jsonrpc": "2.0",
      "id": 7,
      "result": {
        "content": [{ "type": "text", "text": "[{\"id\": \"VWO-1042\", ...}]" }]
      }
    }

    Only two directions exist: the client sends requests and notifications, and the server sends responses and notifications. A server never starts a request of its own. The transport only decides how those messages are framed:

    TransportHow messages travelUse it for
    stdioThe client launches the server as a subprocess. One JSON-RPC message per line on stdin and stdout; no embedded newlines. Shutdown: the client closes stdin.Local servers. Chapter 10 uses it.
    Streamable HTTPEach message is an HTTP POST to one MCP endpoint; the reply is JSON or a request-scoped SSE stream.Remote or shared servers, several users, network boundaries
    sequenceDiagram
      participant C as Client (inside the host)
      participant S as server.py over stdio
      C->>S: initialize (id 1)
      S-->>C: result id 1 with capabilities and serverInfo
      C-)S: notifications/initialized (no id, no reply)
      C->>S: tools/list (id 2)
      S-->>C: result id 2 with three tool schemas
      C->>S: tools/call (id 3)
      S-->>C: result id 3 with content and isError
      C-)S: close stdin, the server exits on EOF
    The session the demo replays, as FastMCP 3.4.4 speaks it: a handshake, discovery, one call, and a shutdown by closing stdin.
    Never write to stdout in a stdio server. stdout is the JSON-RPC channel. What a stray, flushed print() did in runs recorded for this page with the MCP Python SDK client: a line on its own was logged as a parse error and skipped; text without a newline glued itself to the next reply, so that reply was lost and the call hung until the client timed out. Without flush=True the text sat in Python's stdout buffer during the run, which only makes the bug show up later. The chapter warns that clients can also disconnect. Log to stderr, which the spec permits.

    06Capabilities, protocol eras and asking for input

    Neither side guesses what the other supports. In the recorded session the server's initialize reply declares tools (with listChanged: true), resources, prompts and logging, plus an extensions entry for io.modelcontextprotocol/ui. A client must not call what a server never advertised.

    The chapter's notes describe two protocol eras. They describe spec revision 2026-07-28 as stateless: every request carries its own protocol version and client capabilities (in _meta), and a client discovers a server with server/discover. Earlier revisions open a session with an initialize handshake. The chapter also notes that the FastMCP version pinned in chapter 10 still speaks the session-based model, and that is what the demo shows: the server answered initialize with protocol version 2025-11-25.

    sequenceDiagram
      participant Host
      participant Client
      participant Server
      Client->>Server: tools/call with version and capabilities in _meta
      Server-->>Client: InputRequiredResult (needs user input)
      Client->>Host: forward to the user or the LLM
      Host-->>Client: the answer
      Client->>Server: the original request, now with the input
      Server-->>Client: result
    Servers never send requests, so a server that needs input asks inside a reply; the client gets the answer and re-sends the original request. This is the flow as the chapter describes it.

    Clients can offer features back to the server: elicitation (ask the user) and, in earlier revisions, sampling (ask the host's LLM) and roots (filesystem boundaries). Optional extensions add Tasks for long-running work, Skills over MCP and MCP Apps for inline UI; both sides must opt in.

    07Security is a test plan

    MCP gives a model data access and code execution, and the protocol cannot enforce safety on its own. The chapter lists three obligations on implementers, and each one is a test case:

    • User consent and control. Does the host ask before it runs a tool, and can the user see and refuse what is shared?
    • Data privacy. Does the host get consent before it hands user data to a server, and keep resource data from going anywhere else?
    • Tool safety. Tools are arbitrary code. Descriptions and annotations from an untrusted server are untrusted input: test what the host does with a tool whose description tries to give orders.

    Common failure modes

    SymptomCauseFix
    Client logs a JSON parse error, a call never answers, or the client disconnectsSomething wrote to stdout that was not an MCP messageSend every log line to stderr
    The model never calls your toolA vague name, description or input schemaRewrite them: the model reads them to decide
    Wrong values for a templated resourceNo companion resource that lists the valid valuesPublish testcases://modules next to testcases://module/{name}
    The server seems to hang on launchCorrect: a stdio server waits on stdin for a clientDrive it with the Inspector, not a bare terminal
    A bad ID returns a stack traceAn unhandled exception instead of a typed errorRaise typed errors whose message names the valid values
    A new client cannot talk to an old serverProtocol eras differ (stateless vs the initialize handshake)Probe with server/discover, fall back to initialize

    08What is in the chapter, and how to try it

    PieceWhat it isHow to use it
    chapter_09_MCP_Basics/MCP.mdThe whole chapter: 14 sections on the problem, the roles, the wire format, the primitives, transports, security and failure modesRead it before chapter 10
    MCP InspectorA browser UI that speaks MCP, like Postman for MCP servers (Node.js only)npx -y @modelcontextprotocol/inspector <command that starts your server>
    The chapter 10 servertestcase-creator-mcp/server.py: 3 tools, 4 resources, 2 prompts over 5,000 test casesThe server this page's demo recorded; build it in chapter 10

    Chapter 9 has nothing to install. To drive a server by hand you need Node.js for the MCP Inspector:

    terminal
    # any MCP server: pass the command that starts it
    npx -y @modelcontextprotocol/inspector <command to start your server>
    
    # the chapter 10 server, from chapter_10_MCP_Creation_VIBE/testcase-creator-mcp
    npx -y @modelcontextprotocol/inspector uv run --directory "$(pwd)" python server.py

    Open the printed localhost:6274 URL, press Connect, then walk the Tools, Resources and Prompts tabs. Run the happy path and the error path for each, the same way you would for any API.