The Testing Academy · Class Notes Saturday, 26 September (IST)
Live class · study guide

Building a naive RAG in n8n and Pinecone, and the chunking bug that split test cases in half

The RAG from the concepts class, built for real: a knowledge base of 100 test cases in n8n, stored in Pinecone with OpenAI embeddings, and a chat agent that retrieves from it. The first version returned broken test cases, and the cause was chunking, not the prompt. The fix, a step that builds one clean document per test case, is the most reusable thing on this page.

By Pramod Dutta, The Testing Academy. Study notes from the live AI Tester Blueprint 4x class, built from the session recording and the batch repository, which received both n8n workflows, the test-case CSV and a link to the prompt conversation during the class (commit "feat: add chapter 11 naive RAG on Pinecone"). The workflow settings on this page were read from the exported JSON, and the chunk-length figures were computed by running the workflow's own document-building logic over the CSV. n8n and Pinecone themselves were not run. The Eraser deck was not reachable while this page was written.

01

What this class covered

  • Where this sits: RAG, MCP, agents, evaluation and LangChain as one agentic QA stack
  • The three kinds of RAG a tester needs to know: naive, advanced and hybrid
  • What a test-case RAG is useful for
  • The build plan: ingestion and retrieval as two separate phases sharing one database
  • Ingestion in n8n: form upload, loader, splitter, embeddings and Pinecone
  • Choosing an embedding model, and why ingestion and retrieval must use the same one
  • Chunk size and overlap, and one test case per chunk
  • Retrieval: a chat trigger, an AI agent, memory, and Pinecone as a tool
  • A grounded system prompt with anti-hallucination and citation rules
  • The bug: broken test cases in the answers, and finding the real cause
  • The fix: a cleaning step that builds one document per test case
02

Where this sits

The concepts class covered what RAG is. This one builds it. The framing: RAG, MCP, AI agents, LLM evaluation with DeepEval, and LangChain with LangGraph are the skills that make up agentic QA, and they are mandatory for testers now, whether manual or automation.

Tools differ from company to company. Some allow n8n, some only allow LangFlow, others want code. The concepts do not change, so the course builds the same RAG more than one way, and this class is the n8n version. And the claim you will read online that RAG is outdated does not match what companies are doing: most are only now implementing it.

03

The three kinds of RAG to know

There is a long list of RAG variants: naive, advanced, modular, graph, agentic, self, corrective, hybrid, multimodal, contextual. For a tester, three matter:

  • Naive RAG: the baseline. Load the documents, split them into chunks, embed the chunks, store them in a vector database, and at question time hand the nearest chunks to the model. This class.
  • Advanced RAG: the same with improvements such as re-ranking. Next class.
  • Hybrid RAG: stores code as well as text, and already includes agentic behaviour.

One limit worth knowing: a video cannot go into a RAG usefully. Its transcript can, and so can images and designs, but models could not yet read what is happening in the video itself.

04

What a test-case RAG is for

The example that runs through the class: you have thousands of test cases in a CSV or spreadsheet. A RAG over them gives you:

  • A regression suggester. "I changed the login feature, which twenty test cases matter most?"
  • A knowledge base to check before writing new test cases, so you do not duplicate ones that already exist for a feature.
  • Failure analysis, if you store your test results (results.json): the top failing modules or tests over the last month.

These are starting points. The method is what transfers: your company's version is whatever question your team keeps asking.

05

Two phases, one shared database

PHASE 1, INGESTION: ONE-TIME SETUP form upload data loader splitter embeddings Pinecone index the only thing they share PHASE 2, RETRIEVAL: WHAT USERS SEE chat message AI agent + memoryanswers from what it finds retrieval tool, same embeddings eval later
In the class's own analogy: eating and what comes out later are separate activities, and the only thing they share is the stomach. Here the stomach is the vector database. Neither phase knows the other exists.

The two phases do not connect to each other. Ingestion runs once (and later on a schedule, when new test cases arrive). Retrieval is the part users touch. The Pinecone index is the only thing they share. Evaluation is a third phase, covered later with DeepEval.

The test data was 100 login test cases in Jira format, valid and invalid, generated with ChatGPT and saved as a CSV. It is in the repo as Wingify_Login_100_Jira_Test_Cases.csv, with IDs from WING-LOGIN-TC-001.

06

Ingestion in n8n

The ingestion flow, in n8n's trigger, action, result pattern:

Node Setting used
Form trigger one file field, "Testcase Path"
Default data loader binary input (a CSV file), custom splitting
Recursive character text splitter chunk size 1000, overlap 0
OpenAI embeddings text-embedding-3-large
Pinecone vector store insert mode, into the new index

Pinecone is a production vector database with a free starter tier; you need an API key from the console (they start with pcsk). In Pinecone an index is a database, a named collection, not a position number. Index names cannot contain underscores, and the index is created with a fixed dimension that must match the embedding model: 3072 for text-embedding-3-large.

The embedding model. n8n gives new accounts some free OpenAI credit, so OpenAI's embeddings were the practical choice. Of OpenAI's models, large has more dimensions than small, and more dimensions carry finer shades of meaning, which helps match a question to the right test case. Small works too; it is a trade-off, not a requirement.

The same embedding model on both sides, always. A query embedded with a different model from the stored chunks lands in a different space, and the similarity search returns noise. Write down the model and dimension you used when you create the index.

Chunking. The aim is one test case per chunk. Split a test case down the middle and you separate its summary from its steps, and the meaning goes with it, the same way splitting a person's first name from their surname loses who they are. The class asked ChatGPT to check the settings, pasting in a screenshot of the workflow for context, and it recommended 1000 characters with no overlap. Overlap exists to carry context across a chunk boundary; when each chunk is meant to be one complete record, you do not want any.

Embedding batch size is not chunk size. It is how many documents are sent to the embedding model per request, a throughput setting. It has nothing to do with how the documents are split, or how many test cases you have. Leave it at the default.

Uploading the CSV through the form ran the flow, and Pinecone's record count went from zero to 300. The browser view shows a preview of each record's text; the stored thing is the vector.

Text is small, vectors are not. The class estimated 15,000 test cases at about 28 MB, and that matches the CSV (187 KB for 100 cases). But that is the size of the text. What a vector database stores is the embedding: 3072 numbers at 4 bytes each is about 12 KB per chunk, so 15,000 chunks is roughly 180 MB before metadata. Still small, but size a vector database by the vectors, not by the source file.

07

Retrieval

The second flow is what a user talks to:

Node Setting used
Chat trigger when a chat message is received
AI agent model gpt-5-mini
Simple memory a window of the last five messages
Pinecone vector store retrieve-as-tool mode, top 10 results
OpenAI embeddings text-embedding-3-large, the same as ingestion

Top K is how many chunks come back per question. More is not better: a higher number pulls in weaker matches and dilutes the answer. Re-ranking, which reorders results by relevance, comes in the next class.

The first system prompt was short. The class had ChatGPT improve it, again with a screenshot for context, into a retrieval-grounded prompt. The version in the repo opens like this:

Text
You are a retrieval-grounded assistant for a test-case knowledge base. Answer
factual questions using only evidence returned by the connected Pinecone
retrieval tool during the current turn.

1. Retrieve before answering

- Call the Pinecone retrieval tool before a

and goes on to rules for handling missing fields, keeping each test case intact, never inventing values, respecting the retrieval limit, and citing the test case each answer came from.

08

The bug: test cases came back broken

The first question, "give me the test cases related to the login page", returned ten results, but they were incomplete. The debugging went in two steps, and the second is the lesson.

Step one, the prompt. The answers said the title was missing. The CSV has no Title column; it has Summary. Telling the agent to use the summary fixed that part, and at first it looked like a prompt problem.

Step two, the real cause. Some results were still missing their summary or their ID, and two of the "ten" results were fragments of the same test case. Handing ChatGPT the workflow JSON and the response pointed at the data loader: it was splitting the CSV by characters, and a test case could straddle a chunk boundary. The numbers confirm it. The raw rows average 1,877 characters across 18 columns, so at 1000 characters per chunk almost every test case was cut into pieces. That is also why 100 test cases became 300 records.

BEFORE: THE LOADER CUT TEST CASES APART raw row: 18 columns, 1,877 chars cut at 1000 chunk 1: ID, summary, preconditions chunk 2: steps, expected result 10 results is not 10 test cases AFTER: ONE CLEAN DOCUMENT PER TEST CASE 11 chosen fields, 750 to 969 characters every document fits inside one 1000-character chunk every result is a whole test case
The prompt was never the real problem. A RAG can only retrieve what ingestion stored, and ingestion had stored halves.

Where would you look first? When a RAG's answers are wrong, the instinct is to rewrite the prompt. Check what went into the database before you touch it. If the chunks are broken, no prompt can reassemble them.

09

The fix: one document per test case

The corrected workflow, 01_Naive_RAG - Complete Test Cases.json in the repo, puts a cleaning step in front of the loader:

  1. Normalize CSV Upload: check exactly one CSV came in.
  2. Extract CSV Rows: parse it properly, with a header row, instead of treating it as a block of text.
  3. Build Test Case Documents: one document per test case, from eleven chosen fields, with the ID, summary, category and priority also saved as metadata.
  4. Loop One Test Case: insert the documents into Pinecone one at a time.
  5. Ingestion Complete: report how many were stored.

The heart of step three builds the text and refuses to continue if a test case would not fit in one chunk:

JavaScript
const text = fields.filter(key => get(key)).map(key => `${key}: ${get(key)}`).join('\n\n');
if (text.length > 1000) {
  throw new Error(`${tcId} has ${text.length} characters. To retain a complete case, raise this limit AND the splitter chunkSize together before ingestion. No truncation was performed.`);
}

Running that logic over the CSV: all 100 test cases pass, with 100 unique IDs, and the documents run from 750 to 969 characters. Every one fits inside a 1000-character chunk, so the splitter never cuts one. That is also why 1000 was the right chunk size for this data, and why it would need raising together with that guard for longer test cases.

The class re-ingested into a fresh index (the old one held the broken chunks), and retrieval then returned complete test cases with their IDs and summaries, for questions about the login page and about the free trial. The cleaning code itself was generated with ChatGPT; the link to that conversation is in the repo's Prompt.md.

10

Tasks and announcements

  • Build this naive RAG from scratch in n8n, including the cleaning step. Doing it yourself is where the details the demo skipped show up.
  • Finish the RAG Explorer from the previous class before moving on to the LangFlow version.
  • Next class: keeping the index current when test cases are added or deleted (re-running ingestion on a schedule), re-ranking and advanced RAG, and a vibe-coded QA co-pilot that stores code, Jira tickets, Confluence pages, meeting transcripts and Figma designs.
  • Evaluating the RAG comes later, with DeepEval, after MCP, CrewAI and LangChain.

Your build checklist. A Pinecone index at dimension 3072. The same embedding model in both flows. A cleaning step that makes one document per test case. Chunk size at least as large as your longest document. Then ask it three questions and check every result is a whole test case.