What this class covered
- A demo of QABuddy, locally and deployed
- What QABuddy is for, and taking it to work as a proof of concept
- The phase 2 plan: auto-ingestion, pipeline tests, deployment
- The course plan, from MCP to DeepEval
- MCP as a standard for connecting agents to tools
- Adding an MCP server to your editor
- Local servers over stdio, remote servers over a URL
- Host, client and server
- The three-phase lifecycle
- MCP Inspector, with Playwright MCP and the Jira MCP
- Tools, resources and prompts
- MCP versus API, and MCP versus RAG
- Two tasks for the week
QABuddy, demonstrated
The last class's project is finished and deployed: a hybrid RAG over 500 test cases, 11 Jira tickets, meeting notes, company documents, diagrams, the SRS, Jenkins logs and both frameworks, built on Qdrant, Qwen3 4B embeddings, a BGE re-ranker and an LLM on Groq. The notes from that class cover how it is built. This class showed it answering, both locally and on the hosted copy:
- Bug triage: asked whether VWO-33 is a duplicate, it found that it duplicates VWO-26, which matches the Jira exports.
- Test design: asked for heatmap test cases, it drew on the PRD, which describes heatmaps and session recordings.
- Framework coding help: asked to write a test, it followed the team's own conventions, such as the custom visual step annotation from the Playwright framework and the wait helper in the Selenium one.
Every answer cites its sources. The instructor reported testing it with 25,000 test cases without trouble.
One fix from the last task came up first: Astra DB does work with Ollama's Nomic embeddings, as long as the collection is created with Nomic's dimension, 768.
Take it to work. The class's advice was to build QABuddy as a proof of concept for your own team, using open-source models or whatever LLM access you can get approved, and to do it soon, before someone else does.
The phase 2 plan
The roadmap was written in class and pushed as Todo_List.md:
| Phase | Items |
|---|---|
| 1, done | multi-source RAG support; the UI, running on open-source tools |
| 2 | auto-ingestion every hour; caching and fast retrieval; testing the RAG pipeline; deployment to AWS, DigitalOcean, GCP or Azure |
| 3, later | Figma support; screenshot support |
Auto-ingestion is the priority. Whenever a test case, the repository, the PRD or a PDF changes, the index should update on its own. A cron file for hourly ingestion is already in the repo, though the detailed task list, PENDING_TASKS.md, says it needs fixing before it is installed. That list breaks the open work into 24 prioritised tasks, each with a "done when", starting with security clean-up and the first real deployment.
Could the same system be rebuilt in n8n or Langflow? Yes, but it would be far more complex, so the course keeps the coded version.
The course plan, from MCP to DeepEval
The batch repo now has a folder for each remaining chapter:
| Chapter | Topic |
|---|---|
| 13 | MCP basics, this class |
| 14 | Creating an MCP server, with AI-assisted coding |
| 15 | Python and pytest |
| 16 | AI agents with CrewAI |
| 17 | AI agents with LangChain |
| 18 | AI agents with LangGraph |
| 19 | LLM evaluation |
| 20 | DeepEval basics |
| 21 | DeepEval framework |
From chapter 15 on, the work is code. The class picked DeepEval over Ragas because one open-source framework covers RAG, AI agents and MCP, where Ragas grew up around evaluating RAG.
MCP: a standard for connecting agents to tools
A large language model is the brain. Add memory and tools and it becomes an agent. Before MCP, every tool an agent talked to had its own interface, so each connection was built by hand. The Model Context Protocol (MCP) is an open standard, created by Anthropic, for connecting AI applications to external systems. On one side are the AI applications: Claude Code, Codex, Cursor, Windsurf, Kiro, GitHub Copilot. On the other are the external systems: Jira, Slack, Playwright and so on. Think of it as a USB port for agents.
Adding an MCP server to your editor
- Claude Code and Command Code: the
/mcpcommand manages servers. - GitHub Copilot in VS Code: the tools button in the chat panel lists MCP servers, and each can be started and stopped there. The quickest way to add one is the "Install in VS Code" link on the server's own page, such as Playwright MCP's. Either way, VS Code writes the configuration to
.vscode/mcp.json:
{
"servers": {
"playwright": {
"type": "stdio",
"command": "npx",
"args": ["@playwright/mcp@latest"]
}
}
}
When a model refuses your request. With several MCP servers switched on, one demo request carried more than 128 tools, and the model rejected it. Switching off the servers the task did not need fixed it. Enable only the MCP servers you are using.
Local servers and remote servers
The configuration itself tells you which kind of server you have:
- Local servers run on your machine. Their configuration has
"type": "stdio"(standard input and output) and a command, usuallynpx. Playwright, Selenium and API-testing MCP servers are local, because the browser or the requests run on your machine. - Remote servers are hosted by the vendor. Their configuration has a URL, they talk over streamable HTTP, and they usually ask you to sign in. Jira (hosted by Atlassian) and GitHub are remote.
What happens with a local server. Ask Copilot to open a site, and there is no Playwright server running yet. The client reads the configuration, runs npx @playwright/mcp@latest, and that command starts the server. The client connects, and the server drives Playwright, which drives the browser.
Host, client and server
| Part | What it is | In the demo |
|---|---|---|
| Host | the application where the AI model runs | VS Code with GitHub Copilot |
| Client | lives inside the host and makes the connection to a server | Copilot's MCP client |
| Server | exposes an external system's capabilities as tools | Playwright MCP, Jira MCP |
The three-phase lifecycle
Every MCP connection runs in three phases:
- Initialization. The client connects, the server states its capabilities, the client confirms, and the server sends its list of tools.
- Operation. For each request, the LLM decides which tool to call, the client calls it, and the server returns the result.
- Shutdown. The connection is closed.
A learner summed it up in class, and the instructor agreed: the agent is just running functions the server provides, and the LLM decides which one to call.
MCP Inspector
MCP Inspector is a debugging tool for seeing exactly what a server offers. It starts with:
npx @modelcontextprotocol/inspector
With Playwright MCP, connecting showed the lifecycle in its history panel: the initialize exchange, the notification that client and server are connected, and the tool list. The Tools tab lists every tool, such as browser_navigate, browser_close and browser_resize. Running browser_navigate with app.vwo.com opened the page, which is exactly what the agent does when you ask it in plain words.
With the Jira MCP, the configuration has a URL, so it is remote, and connecting asked for an Atlassian sign-in before anything else. Once connected, the tools included creating an issue, fetching one, searching, and looking up account details. The class created a bug in the VWO project from the Inspector. The first attempt failed until the cloud ID field was given the Jira site's URL. The connection details also show the transport: streamable HTTP for a remote server, stdio for a local one.
Keep tokens off screen. The Inspector shows the connection's token and client details. They are as good as a password, so close that panel before you share your screen or a screenshot.
Tools, resources and prompts
Tools are only one of three things an MCP server can provide:
| Primitive | What it is |
|---|---|
| Tools | functions the agent can execute, such as navigate or create an issue |
| Resources | read-only information, such as documentation or test cases |
| Prompts | reusable prompt templates for set interactions |
Resources and prompts get a demo in the next class.
MCP versus API, and versus RAG
This one is asked in interviews, and the class said people get rejected over it: MCP is not an API.
| API | MCP | |
|---|---|---|
| Designed for | web applications | AI agents |
| Interaction | request and response | a session with a lifecycle |
| Discovering what it can do | no, you read the documentation | yes, the server lists its tools |
| State | stateless | stateful for the session |
| Responses | fixed formats | rich content, and an agent's results can differ run to run |
And do not confuse either with RAG. RAG is retrieval: fetching knowledge to ground an answer. MCP is a protocol connecting an agent to external systems. An API is a protocol connecting applications, for web clients rather than agents.
Tasks and announcements
Task 1, inspect five MCP servers. Open each with MCP Inspector and share a screenshot of its tools: Playwright, Selenium, GitHub, Jira, and either Slack or Figma.
Task 2, run QABuddy. Install QABuddy and run it locally with the open-source tools, then deploy it to Vercel with dummy data. Never put company data in a public deployment.
- QABuddy's code is in the batch repo, with the phase roadmap and the pending tasks beside it.
- The fixed Langflow flows from the last class are in the repo too: the fixed naive RAG and the advanced RAG with its Cohere re-ranker.
- Advanced MCP sessions run this week.
- Next class: MCP resources and prompts, creating an MCP server (chapter 14), and the start of Python.