The Testing Academy · Class Notes Sunday, 27 September (IST)
Live class · study guide

Rebuilding the naive RAG in Langflow, choosing a vector database, and the MCP certification plan

The naive RAG from the previous class, rebuilt in Langflow: running Langflow locally or on a server, which vector databases and free embeddings are available, and the component flow from file to answer. The live build did not finish, because the vector database step failed, and the fix was deferred. The class also set out the MCP certification plan and the route into Python.

By Pramod Dutta, The Testing Academy. Study notes from the live AI Tester Blueprint 4x class, built from the session recording. The batch repository had no commit for this class when this page was written; the class said the Langflow flow would be exported and shared. The Eraser deck was not reachable while this page was written. The live build stopped at the vector database step, and this page says so rather than presenting a finished flow.

01

What this class covered

  • The surprise test, and why tests are mandatory
  • The MCP plan: part two certification, then building an MCP in Python
  • The road ahead: the QA buddy project, then Python and pytest
  • Who had built the naive RAG in n8n, and why it is not optional
  • The same RAG in Langflow: install options and versions
  • Running Langflow locally or on a cloud server
  • Vector database options in Langflow, and what newer versions hide
  • Free embedding options
  • The Langflow flow, component by component, and why it needs a type converter
  • Where the live build failed
  • Why the class builds live instead of using ready-made templates
  • Next week's project: a QA co-pilot over everything your team owns
02

The MCP plan, and the road to Python

Surprise tests are mandatory. Their job is to show you where you stand before an interviewer does.

MCP, part two. The earlier session, Introduction to MCP, came with a certification; complete it if you have not. Part two, MCP Advanced Concepts, is a mandatory extra session with its own certification, on Friday at 8 PM IST. It covers how MCP tooling works, how to debug an MCP, and using Jira, Playwright and Slack MCPs together. After that, a session only for this batch on building your own MCP in Python.

Then Python. After the RAG project comes Python and pytest. Python is the language behind building MCPs, DeepEval for LLM evaluation, CrewAI, LangChain and, last, LangGraph, so it has to come first. Extra masterclasses on Cursor, Codex and a Claude Code revision are planned alongside.

03

The naive RAG is not optional

Several people had rebuilt the n8n naive RAG from the previous class. Many had not. RAG sits under almost every QA agent you will build: a bug-log finder, a test case analyser, a knowledge base. That agent is what gets noticed at work and asked about in interviews. Build it, then put it on GitHub as proof.

Companies differ in which tools they allow: some n8n, some Langflow, some only code. The concept is the same in each, which is why the class rebuilt it in a second tool.

04

Installing and hosting Langflow

Langflow can be installed three ways: the desktop app, Docker, or the Python package. The class was on version 1.11 and updated to 1.12, the latest at the time. Updates change the interface and the available components, which caused problems later in the session.

SAME FLOW, TWO PLACES TO RUN IT Local desktop app, Docker, or the Python package free runs only while your computer is on only you can use it right for learning Cloud a rented virtual server (a VPS), Ubuntu, Python package a small monthly cost runs around the clock shareable through the server's address what companies do, on their own servers DigitalOcean, AWS, Google Cloud and Azure all rent these servers. The flow you build is identical in both.
Learn locally. Your company will ask for the cloud version, hosted on its own infrastructure.
05

Vector databases in Langflow

Any vector database does the same job; the choice is about where it runs and what your company allows.

Database Where it runs Note from class
Chroma DB locally, on disk set a persist directory and a collection name
Pinecone cloud used in the n8n version
Qdrant local or cloud another common choice
Astra DB cloud, with a free tier can embed text for you with a built-in NVIDIA model

Newer Langflow versions hide some databases. In the updated version, Chroma DB and Pinecone did not appear in the component list at all; an older installation (around 1.9) still showed both. They can be added back by installing a bundle. If a tutorial shows a component you cannot find, check your version first.

Embeddings. OpenAI embeddings cost money. Free options shown in class: NVIDIA's hosted embedding models, which need a free API key from build.nvidia.com (a Mistral model is also available through NVIDIA, on the same key), and Ollama running locally, if you are on the desktop app. The same rule as last class applies: ingestion and retrieval must use the same embedding model.

06

The Langflow flow

The concept is identical to the n8n version: ingestion on one side, retrieval, augmentation and generation on the other, sharing only the vector database.

INGESTION Read Fileraw content Split Text1000, overlap 0 Vector store+ embedding model the live build failed here RETRIEVAL, AUGMENTATION, GENERATION Chat Input Vector storesearch Type ConvertJSON to text Prompt Template{context} {question} Groq Chat Output the question also fills the template
Context comes from the top matches in the vector store; the question comes from the user. The model answers from both.

Ingestion. Read File loads the test case file and emits its raw content. Split Text chunks it at 1000 characters with no overlap; its separator field is optional and can be left empty. The chunks go into the vector store component, with an embedding model attached.

Retrieval. The question from Chat Input searches the same vector store. The Prompt Template has two slots: the context, filled with the top matching chunks, and the question, filled with what the user asked. The completed prompt goes to the model (Groq in class), and the answer to Chat Output.

Why Type Convert sits in the middle. The vector store's search returns structured data, JSON, and the Prompt Template's context slot takes text. Langflow will not connect two ports of different types. The Type Convert component turns the results into text, and then the connection is allowed.

This is still a naive RAG: no loop and no one-document-per-test-case cleaning like the fixed n8n version. The goal was to get data in and answers out first, then improve it.

07

Where the live build failed

The retrieval side was assembled, but ingestion never completed, and each attempt failed for a different reason:

  1. Chroma DB did not appear in the updated Langflow, and the attempt that was made reported that it could not open the database.
  2. NVIDIA embeddings asked for an additional package to be installed.
  3. Astra DB databases stayed in a pending state, and then did not appear in Langflow's database list once active.
  4. Using Astra DB's built-in embeddings instead produced an unexpected Google API key error, then Langflow crashed and had to be restarted.

The fix was deferred: the instructor said he would debug it afterwards, share a working version, and cover it in a later session. Chroma DB running locally was the combination he expected to work, since it had been tested before, and the new databases were the source of the trouble. The practical lesson: version updates and unfamiliar databases are where these tools break, and a working flow on one version is not evidence it works on the next.

Why build live instead of using a template? A ready-made flow would have run first time, and nobody would have learned how to debug it. The problems you watch being worked through are the ones you will hit at work. The instructor had a working template available and chose to build from scratch deliberately.

08

Tasks and announcements

  • Complete the Introduction to MCP certification if you have not.
  • Attend MCP Advanced Concepts, Friday at 8 PM IST. Mandatory, with a certification.
  • Rebuild the n8n naive RAG if you have not, and upload it to GitHub as proof.
  • Try the Langflow flow locally once it is shared, ideally with Chroma DB, and report whether it runs.
  • Next week: the QA co-pilot, a RAG close to production that answers questions across your Playwright and Selenium code, thousands of test cases, PRDs, Jira tickets, Confluence pages and meeting notes. The 3x batch's version also ingested Jenkins logs and design documents. The aim is to run it entirely locally, swapping the hosted model for a local one.