What this class covered
- Where this sits: RAG, MCP, agents, evaluation and LangChain as one agentic QA stack
- The three kinds of RAG a tester needs to know: naive, advanced and hybrid
- What a test-case RAG is useful for
- The build plan: ingestion and retrieval as two separate phases sharing one database
- Ingestion in n8n: form upload, loader, splitter, embeddings and Pinecone
- Choosing an embedding model, and why ingestion and retrieval must use the same one
- Chunk size and overlap, and one test case per chunk
- Retrieval: a chat trigger, an AI agent, memory, and Pinecone as a tool
- A grounded system prompt with anti-hallucination and citation rules
- The bug: broken test cases in the answers, and finding the real cause
- The fix: a cleaning step that builds one document per test case
Where this sits
The concepts class covered what RAG is. This one builds it. The framing: RAG, MCP, AI agents, LLM evaluation with DeepEval, and LangChain with LangGraph are the skills that make up agentic QA, and they are mandatory for testers now, whether manual or automation.
Tools differ from company to company. Some allow n8n, some only allow LangFlow, others want code. The concepts do not change, so the course builds the same RAG more than one way, and this class is the n8n version. And the claim you will read online that RAG is outdated does not match what companies are doing: most are only now implementing it.
The three kinds of RAG to know
There is a long list of RAG variants: naive, advanced, modular, graph, agentic, self, corrective, hybrid, multimodal, contextual. For a tester, three matter:
- Naive RAG: the baseline. Load the documents, split them into chunks, embed the chunks, store them in a vector database, and at question time hand the nearest chunks to the model. This class.
- Advanced RAG: the same with improvements such as re-ranking. Next class.
- Hybrid RAG: stores code as well as text, and already includes agentic behaviour.
One limit worth knowing: a video cannot go into a RAG usefully. Its transcript can, and so can images and designs, but models could not yet read what is happening in the video itself.
What a test-case RAG is for
The example that runs through the class: you have thousands of test cases in a CSV or spreadsheet. A RAG over them gives you:
- A regression suggester. "I changed the login feature, which twenty test cases matter most?"
- A knowledge base to check before writing new test cases, so you do not duplicate ones that already exist for a feature.
- Failure analysis, if you store your test results (
results.json): the top failing modules or tests over the last month.
These are starting points. The method is what transfers: your company's version is whatever question your team keeps asking.
Two phases, one shared database
The two phases do not connect to each other. Ingestion runs once (and later on a schedule, when new test cases arrive). Retrieval is the part users touch. The Pinecone index is the only thing they share. Evaluation is a third phase, covered later with DeepEval.
The test data was 100 login test cases in Jira format, valid and invalid, generated with ChatGPT and saved as a CSV. It is in the repo as Wingify_Login_100_Jira_Test_Cases.csv, with IDs from WING-LOGIN-TC-001.
Ingestion in n8n
The ingestion flow, in n8n's trigger, action, result pattern:
| Node | Setting used |
|---|---|
| Form trigger | one file field, "Testcase Path" |
| Default data loader | binary input (a CSV file), custom splitting |
| Recursive character text splitter | chunk size 1000, overlap 0 |
| OpenAI embeddings | text-embedding-3-large |
| Pinecone vector store | insert mode, into the new index |
Pinecone is a production vector database with a free starter tier; you need an API key from the console (they start with pcsk). In Pinecone an index is a database, a named collection, not a position number. Index names cannot contain underscores, and the index is created with a fixed dimension that must match the embedding model: 3072 for text-embedding-3-large.
The embedding model. n8n gives new accounts some free OpenAI credit, so OpenAI's embeddings were the practical choice. Of OpenAI's models, large has more dimensions than small, and more dimensions carry finer shades of meaning, which helps match a question to the right test case. Small works too; it is a trade-off, not a requirement.
The same embedding model on both sides, always. A query embedded with a different model from the stored chunks lands in a different space, and the similarity search returns noise. Write down the model and dimension you used when you create the index.
Chunking. The aim is one test case per chunk. Split a test case down the middle and you separate its summary from its steps, and the meaning goes with it, the same way splitting a person's first name from their surname loses who they are. The class asked ChatGPT to check the settings, pasting in a screenshot of the workflow for context, and it recommended 1000 characters with no overlap. Overlap exists to carry context across a chunk boundary; when each chunk is meant to be one complete record, you do not want any.
Embedding batch size is not chunk size. It is how many documents are sent to the embedding model per request, a throughput setting. It has nothing to do with how the documents are split, or how many test cases you have. Leave it at the default.
Uploading the CSV through the form ran the flow, and Pinecone's record count went from zero to 300. The browser view shows a preview of each record's text; the stored thing is the vector.
Text is small, vectors are not. The class estimated 15,000 test cases at about 28 MB, and that matches the CSV (187 KB for 100 cases). But that is the size of the text. What a vector database stores is the embedding: 3072 numbers at 4 bytes each is about 12 KB per chunk, so 15,000 chunks is roughly 180 MB before metadata. Still small, but size a vector database by the vectors, not by the source file.
Retrieval
The second flow is what a user talks to:
| Node | Setting used |
|---|---|
| Chat trigger | when a chat message is received |
| AI agent | model gpt-5-mini |
| Simple memory | a window of the last five messages |
| Pinecone vector store | retrieve-as-tool mode, top 10 results |
| OpenAI embeddings | text-embedding-3-large, the same as ingestion |
Top K is how many chunks come back per question. More is not better: a higher number pulls in weaker matches and dilutes the answer. Re-ranking, which reorders results by relevance, comes in the next class.
The first system prompt was short. The class had ChatGPT improve it, again with a screenshot for context, into a retrieval-grounded prompt. The version in the repo opens like this:
You are a retrieval-grounded assistant for a test-case knowledge base. Answer
factual questions using only evidence returned by the connected Pinecone
retrieval tool during the current turn.
1. Retrieve before answering
- Call the Pinecone retrieval tool before a
and goes on to rules for handling missing fields, keeping each test case intact, never inventing values, respecting the retrieval limit, and citing the test case each answer came from.
The bug: test cases came back broken
The first question, "give me the test cases related to the login page", returned ten results, but they were incomplete. The debugging went in two steps, and the second is the lesson.
Step one, the prompt. The answers said the title was missing. The CSV has no Title column; it has Summary. Telling the agent to use the summary fixed that part, and at first it looked like a prompt problem.
Step two, the real cause. Some results were still missing their summary or their ID, and two of the "ten" results were fragments of the same test case. Handing ChatGPT the workflow JSON and the response pointed at the data loader: it was splitting the CSV by characters, and a test case could straddle a chunk boundary. The numbers confirm it. The raw rows average 1,877 characters across 18 columns, so at 1000 characters per chunk almost every test case was cut into pieces. That is also why 100 test cases became 300 records.
Where would you look first? When a RAG's answers are wrong, the instinct is to rewrite the prompt. Check what went into the database before you touch it. If the chunks are broken, no prompt can reassemble them.
The fix: one document per test case
The corrected workflow, 01_Naive_RAG - Complete Test Cases.json in the repo, puts a cleaning step in front of the loader:
- Normalize CSV Upload: check exactly one CSV came in.
- Extract CSV Rows: parse it properly, with a header row, instead of treating it as a block of text.
- Build Test Case Documents: one document per test case, from eleven chosen fields, with the ID, summary, category and priority also saved as metadata.
- Loop One Test Case: insert the documents into Pinecone one at a time.
- Ingestion Complete: report how many were stored.
The heart of step three builds the text and refuses to continue if a test case would not fit in one chunk:
const text = fields.filter(key => get(key)).map(key => `${key}: ${get(key)}`).join('\n\n');
if (text.length > 1000) {
throw new Error(`${tcId} has ${text.length} characters. To retain a complete case, raise this limit AND the splitter chunkSize together before ingestion. No truncation was performed.`);
}
Running that logic over the CSV: all 100 test cases pass, with 100 unique IDs, and the documents run from 750 to 969 characters. Every one fits inside a 1000-character chunk, so the splitter never cuts one. That is also why 1000 was the right chunk size for this data, and why it would need raising together with that guard for longer test cases.
The class re-ingested into a fresh index (the old one held the broken chunks), and retrieval then returned complete test cases with their IDs and summaries, for questions about the login page and about the free trial. The cleaning code itself was generated with ChatGPT; the link to that conversation is in the repo's Prompt.md.
Tasks and announcements
- Build this naive RAG from scratch in n8n, including the cleaning step. Doing it yourself is where the details the demo skipped show up.
- Finish the RAG Explorer from the previous class before moving on to the LangFlow version.
- Next class: keeping the index current when test cases are added or deleted (re-running ingestion on a schedule), re-ranking and advanced RAG, and a vibe-coded QA co-pilot that stores code, Jira tickets, Confluence pages, meeting transcripts and Figma designs.
- Evaluating the RAG comes later, with DeepEval, after MCP, CrewAI and LangChain.
Your build checklist. A Pinecone index at dimension 3072. The same embedding model in both flows. A cleaning step that makes one document per test case. Chunk size at least as large as your longest document. Then ask it three questions and check every result is a whole test case.