AI Tester Blueprint RAG and MCP Build an MCP server
Chapter 10
Chapter 10 . RAG and MCP . Build an MCP server

Build an MCP server: your test cases as tools

Chapter 9 explained the protocol. Here you build the server: one FastMCP file that reads a 5,000-row VWO test-case export once and exposes it as three tools the model can call, four resources the app can read, and two prompts the user can pick. It was generated from a structured prompt, then checked by hand in the MCP Inspector.

3 / 4 / 2
Tools, resources, prompts
All three MCP primitives over one CSV, in a 292-line server.py.
5,000
Test cases served
14 columns, keys VWO-1001 to VWO-6000, read once and cached.
11
Inspector checks
The README's click-through checklist, error paths included.
3.4.4
fastmcp version
Pinned in pyproject.toml. The prompt asked for 2.x; the build verified 3.x.

01What you build, and why a tester should

Pasting a 5,000-row CSV into a chat burns the context window and goes stale the moment the file changes. An MCP server lets the model ask for exactly what it needs ("three invite-user cases in User Management") and works in every MCP client without changes. For a tester it is also a new kind of system under test: inputs, outputs, schemas and error paths you can check like any API.

flowchart LR
  CSV[(vwo_5000_test_cases.csv)] -->|read once, cached| S[FastMCP server: vwo-testcases]
  M[The model decides] -->|tools/call| S
  A[The client app] -->|resources/read| S
  U[The user picks] -->|prompts/get| S
  S --> T[3 tools: search, get, stats]
  S --> R[4 resources: schema, all, modules, module/NAME]
  S --> P[2 prompts: review, regression suite]
One server, one dataset, all three primitives. stdio is the transport, so a client starts server.py as a subprocess.
KindNameWhat it does
Toolsearch_test_cases(query, module, test_type, priority, limit=20)Case-insensitive text search over Summary, Description, Steps, Expected Result, Labels and Preconditions, with optional filters; limit 1 to 200
Toolget_test_case(test_id)One case by key; 1001 and vwo-1001 also resolve to VWO-1001
Tooltest_case_stats(group_by)Counts by module, priority, status, test_type, browser or device
Resourcetestcases://schemaRow count, primary key, and each column's distinct count and values
Resourcetestcases://allEvery row as JSON: bulk data belongs in a resource, not a tool reply
Resourcetestcases://modulesThe 17 module names with counts: the valid values for the template below
Resource (templated)testcases://module/{name}All cases for one module, matched case-insensitively (Reports has 333)
Promptreview_test_case(test_id)A senior-QA review rubric on four axes, with the case embedded as JSON
Promptgenerate_regression_suite(module)The first 40 cases of a module, with "Showing 40 of N" stated in the text
Read the real header first. The column names are not the obvious guesses: the key is Issue Key (not ID), the title is Summary (not Title) and the module is Component (not Module). The build prompt makes the assistant show the schema and wait before it writes any code.

02Live demo: the tool-call playground

Pick a tool, try a case or type your own arguments, and send the tools/call. The tool definitions are the server's real tools/list reply. The answers come from a JavaScript port of server.py and FastMCP's argument checks, run over the same 5,000 rows; on 669 recorded calls to the real server (fastmcp 3.4.4) it produced the same reply, byte for byte.

Watch two columns disagree. The inputSchema check is what a client could verify before sending; the reply is what the server actually does. They are not the same rules.

Tool-call playground: the vwo-testcases serverNo API key needed
Tool definition from tools/list

      Try a case
      
data-testid=ms-tooldata-testid=ms-preset-1data-testid=ms-argsdata-testid=ms-senddata-testid=ms-checkdata-testid=ms-responsedata-testid=ms-verdict
inputSchema check:
    Request on stdin
    
          Reply on stdout 
          
    
          The code that answered (server.py)
          
    
          Server stderr (abridged)
          
    
        
    Two kinds of error, one shape. Bad types and unknown arguments are rejected by FastMCP before your function runs. Business rules (the 1 to 200 limit, an unknown module) are rejected by your own ToolError. Both reach the client as "isError": true with a readable message, while the traceback stays on stderr.

    03How server.py is built

    The whole server is one file. Five decisions carry it. First, nothing but JSON-RPC may reach stdout, so logging is pointed at stderr before anything else runs:

    server.py
    logging.basicConfig(
        level=logging.INFO,
        stream=sys.stderr,
        format="%(asctime)s %(levelname)s [%(name)s] %(message)s",
    )
    log: Final = logging.getLogger("vwo-testcases")

    Second, the CSV path is resolved, never hard-coded: an environment override first, then three places relative to the file. The data is read once and cached in _CASES with a key index in _BY_ID.

    server.py
    def _resolve_csv_path() -> Path:
        """Locate the dataset via the env override, then paths relative to this file."""
        override = os.environ.get(CSV_ENV_VAR)
        if override:
            path = Path(override).expanduser()
            if not path.is_file():
                raise FileNotFoundError(f"{CSV_ENV_VAR}={override!r} is not a readable file")
            return path
        for candidate in _CANDIDATES:
            if candidate.is_file():
                return candidate
        searched = ", ".join(str(c) for c in _CANDIDATES)
        raise FileNotFoundError(f"{CSV_FILENAME} not found (looked in: {searched}); set {CSV_ENV_VAR} to override")

    Third, one decorator per primitive, and the docstring and type hints are functional: FastMCP turns them into the description and the inputSchema the model reads.

    server.py: the signature tools/list: the inputSchema the model reads def search_test_cases( query: str, module: str | None = None, test_type: str | None = None, priority: str | None = None, limit: int = 20, ) -> list[dict[str, Any]]: "required": ["query"], "query": {"type": "string"}, "module": {"anyOf": [string, null], "default": null}, test_type, priority: the same "limit": {"type": "integer", "default": 20}, "additionalProperties": false No **kwargs in the signature, so any argument the function does not name is rejected.
    FastMCP builds each tool's inputSchema from the type hints and defaults, and its description from the docstring. This is the real schema from the server's tools/list reply.

    Fourth, errors are typed and name the valid values, so a wrong guess corrects itself in one more call:

    server.py
    @mcp.tool
    def get_test_case(test_id: str) -> dict[str, Any]:
        """Return one test case by its issue key, for example VWO-1001."""
        row = _lookup(test_id)
        if row is None:
            raise ToolError(
                f"unknown test_id {test_id!r}; expected an issue key such as "
                f"{next(iter(_BY_ID))} (dataset holds {len(_BY_ID)} cases)"
            )
        return _expand(row)

    Fifth, a templated resource is paired with a plain resource that lists what {name} accepts:

    server.py
    @mcp.resource("testcases://modules", mime_type="application/json")
    def modules_resource() -> list[ResourceContent]:
        """The valid module names accepted by testcases://module/{name}, with case counts."""
        counts = Counter(row[COL_MODULE] for row in _cases())
        return _json_resource([{"module": name, "count": n} for name, n in counts.most_common()])
    
    
    @mcp.resource("testcases://module/{name}", mime_type="application/json")
    def module_resource(name: str) -> list[ResourceContent]:
        """All test cases belonging to one module, matched case-insensitively."""
        hits = _module_rows(name)
        if not hits:
            raise ResourceError(f"unknown module {name!r}; read testcases://modules for the valid list")
        return _json_resource([_expand(row) for row in hits])

    Prompts are templates that the server fills in; nothing reaches a model until a client sends the text. The entry point starts stdio with the FastMCP banner turned off, because the banner would go to stdout:

    server.py
    if __name__ == "__main__":
        try:
            log.info("startup: %d test cases cached", len(_cases()))
        except ToolError as exc:
            log.error("startup: %s", exc)
        mcp.run(show_banner=False)

    04Resources and prompts, as the server answers them

    The playground covers tools. Resources and prompts follow the same request and reply pattern; these replies were recorded from the server over stdio.

    resources/read testcases://modules (first 3 of 17)
    [
      {"module": "Reports", "count": 333},
      {"module": "Personalization", "count": 310},
      {"module": "Feature Rollout", "count": 309},
      ...
    ]
    prompts/get generate_regression_suite {"module": "Reports"} (first two lines)
    You are building a regression suite for the Reports module.
    Showing 40 of 333 available cases.
    • resources/list returns the three fixed resources. The templated one appears in resources/templates/list as testcases://module/{name}.
    • Every resource reply carries "mimeType": "application/json", because each resource returns [ResourceContent(..., mime_type="application/json")].
    • Resource and prompt errors come back as JSON-RPC errors, not results: reading testcases://module/xyz returned {"code": 0, "message": "unknown module 'xyz'; read testcases://modules for the valid list"}. Tool errors come back as results with isError. Test both shapes.
    • The regression-suite prompt states its sample ("Showing 40 of 333") instead of cutting silently, so the model knows it is working from part of the data.

    05Built from a prompt: Prompt.md

    The server was vibe-coded from a structured brief in the course's RICE-POT shape: Role, Instructions, Context, Example, Parameters, Output and Tone, then a gated PROCESS. (The repo README calls it the "RISE-CEPT brief"; the file's own headings spell R-I-C-E-P-O-T.) The instructions read like acceptance criteria:

    Prompt.md (Instructions)
    Build ONE runnable MCP server in Python using FastMCP that exposes all three MCP
    primitives over a local dataset file, vwo_5000_test_cases.csv.
    
    [Critical] Before writing any code, read the CSV header and first 5 rows, then
      show me the detected schema (column names + inferred types) and WAIT for my
      confirmation. Do not invent or assume column names.
    [Mandatory] Load the CSV once at server startup into memory and reuse it across
      calls. Do not re-read the file on every request.
    [Mandatory] Expose at least 3 TOOLS, e.g.:
      - search_test_cases(query: str, module: str | None, limit: int) -> list[dict]
      - get_test_case(test_id: str) -> dict
      - test_case_stats(group_by: str) -> dict   (counts by module/priority/status)
    [Mandatory] Expose at least 3 RESOURCES, including one templated URI:
      - testcases://schema        -> column names and types
      - testcases://all           -> the full dataset (JSON)
      - testcases://module/{name} -> all cases for one module (templated)
    [Mandatory] Expose at least 2 PROMPTS:
      - review_test_case(test_id)      -> asks the LLM to critique coverage/clarity
      - generate_regression_suite(module) -> builds a suite from that module's cases
    flowchart LR
      P1[Phase 1: read the CSV header, show the schema, propose the primitives] -->|you approve| P2[Phase 2: write one file at a time]
      P2 --> P3[Phase 3: run it and walk the Inspector checklist]
    The PROCESS section of Prompt.md: the assistant stops for approval after the schema, so it cannot build on guessed column names.
    Verify the library, not the prompt. The brief asks for "FastMCP (latest 2.x)", but the build pins fastmcp==3.4.4, the latest release when it was written. On 3.x a resource that returns list[dict] raises, and on both versions a bare str silently becomes text/plain. The course notes record smoke-testing every API shape on the installed version before writing the real file.

    06Run it, inspect it, register it

    Python 3.11+ and uv run the server; Node.js is only needed for the Inspector. No API key: the server is local and read-only.

    terminal
    cd chapter_10_MCP_Creation_VIBE/testcase-creator-mcp
    uv sync                          # creates .venv with fastmcp==3.4.4
    uv run python server.py          # waits on stdin; logs go to stderr
    
    npx -y @modelcontextprotocol/inspector uv run --directory "$(pwd)" python server.py
    # port 6277 busy? use other ports:
    CLIENT_PORT=6284 SERVER_PORT=6287 npx -y @modelcontextprotocol/inspector uv run --directory "$(pwd)" python server.py
    stderr at startup (from the README)
    INFO [vwo-testcases] loaded 5000 test cases from .../resource/vwo_5000_test_cases.csv
    INFO [vwo-testcases] startup: 5000 test cases cached
    PieceWhat it doesCommand
    server.pyThe stdio MCP server. It waits on stdin, so it looks like it hangs in a terminal.uv run python server.py
    MCP InspectorA browser UI to call every tool, resource and prompt by handnpx -y @modelcontextprotocol/inspector uv run --directory "$(pwd)" python server.py
    Claude CodeRegisters the server with the CLIclaude mcp add vwo-testcases -- uv run --directory "$(pwd)" python server.py
    Claude DesktopRegisters the server in claude_desktop_config.jsonEdit the file shown below, then quit and reopen the app
    VWO_TESTCASES_CSVPoints the server at another CSV without editing codeVWO_TESTCASES_CSV=/path/to/other.csv uv run python server.py
    Prompt.mdThe RICE-POT brief the server was generated fromPaste it into your coding assistant and approve each phase

    To register it in Claude Desktop, edit ~/Library/Application Support/Claude/claude_desktop_config.json on macOS. Use absolute paths: the app does not inherit your shell PATH, so a bare "command": "uv" fails with spawn uv ENOENT. Find your uv with which uv, then quit and reopen the app.

    claude_desktop_config.json
    {
      "mcpServers": {
        "vwo-testcases": {
          "command": "/absolute/path/to/uv",
          "args": [
            "run",
            "--directory",
            "/absolute/path/to/chapter_10_MCP_Creation_VIBE/testcase-creator-mcp",
            "python",
            "server.py"
          ]
        }
      }
    }

    07What to watch

    • The schema and the server disagree, both ways. limit has no minimum or maximum in the schema, so 0 passes any client-side check and only the function body stops it. In the other direction FastMCP accepts "3", 3.0 and even true for an integer. Test the server, not the schema.
    • Search does not look at the key. search_test_cases("VWO-3400") finds nothing, because Issue Key is not in the searched fields. That is what get_test_case is for.
    • Never print. stdout is the JSON-RPC channel. In recorded runs a stray line made the client log a parse error, and stray text without a newline swallowed the next reply. Log with logging to stderr.
    • FastMCP 3.x resources. Return list[ResourceContent]. A list[dict] raises TypeError: contents[0] must be ResourceContent, got dict.; a bare str forces text/plain.
    • A server that seems to hang is fine. A stdio server waits for a client on stdin. Drive it with the Inspector, not a bare terminal.
    • Absolute paths for Claude Desktop, and a full quit and restart after every config change.
    • Expected errors leave tracebacks on stderr. The README leaves that noise in on purpose: muting it would also hide real bugs. The client only sees the clean message.