Chapter 7
Chapter 7 . RAG and MCP . RAG basics

RAG: answers from your own docs

An LLM only knows what it was trained on. Retrieval-Augmented Generation hands it the right part of your PRD or test suite at question time: split the document into chunks, find the chunks that match the question, and send only those with a rule to answer from them. This chapter builds that pipeline four ways, from a React app that shows every stage to a hybrid pipeline over 5,000 test cases.

4
Builds of one idea
RAG Explorer, an n8n workflow, two LangFlow flows and the Advanced RAG Flask app.
7
Chunks from the PRD
The 6-page VWO PRD at the default CHUNK_SIZE=1200 and CHUNK_OVERLAP=200.
5,000
Test cases in Advanced RAG
Seeded VWO cases, one chunk per row, dense and sparse vectors in embedded Qdrant.
768
Embedding dimensions
nomic-embed-text through local Ollama. bge-m3 dense vectors in Advanced RAG have 1024.

01Why RAG, and why a tester cares

Ask a general-purpose LLM what the VWO PRD lists as success metrics and it will guess, fluently. RAG fixes that by putting the right part of the document into the prompt at question time. The model is told to answer only from those chunks, and to say so when they do not cover the question.

For a tester, the useful part is that a RAG system has seams you can check one at a time:

  • Chunking. Did the fact survive the split, or is half of the list in another chunk?
  • Retrieval. Did the chunk that holds the fact make the top-k?
  • The prompt. Does it carry those chunks and an answer-only-from-context rule?
  • Generation. Did the model use the context, cite it, and refuse when the context is silent?

When an answer is wrong, walk the seams in that order. RAG Explorer puts each one on screen: chunk counts, a sample vector, the retrieved chunks with their similarity, and the exact prompt sent to Groq.

flowchart TB
  subgraph Ingest
    direction LR
    P[PRD PDF in data/] --> X[pdf-parse: text]
    X --> C[chunkText 1200 / 200]
    C --> E[nomic-embed-text: 768-d]
    E --> S[(ChromaDB vwo_prd)]
  end
  subgraph Ask
    direction LR
    Q[Question] --> QE[embed, same model]
    QE --> R[query: top-k by cosine]
    R --> B[buildPrompt]
    B --> G[Groq gpt-oss-120b]
    G --> A[answer + chunk numbers]
  end
  Ingest -->|searched by| Ask
RAG Explorer: ingest once (read, chunk, embed, store), then embed each question with the same model and send only the top-k chunks to Groq.

02Live demo: chunk, rank, ground

Run the pipeline by hand. The chunker is a line-for-line port of the chapter's chunk.js, and the prompt is built the same way as groq.js. One stand-in: instead of 768-dimension Nomic vectors, each chunk gets a TF-IDF score against the question, so the ranking runs offline and is identical every time.

Start with What is the goal of this PRD? at 1200 / 200, then set the chunk size to 300 and the overlap to 50 and watch the evidence count drop. Switch to the test cases to see what windows do to records.

Chunk, rank, ground: the RAG pipeline by handNo API key needed
document chunkText score (TF-IDF) top-k buildPrompt Groq LLM
data-testid=rag-docdata-testid=rag-textdata-testid=rag-sizedata-testid=rag-overlapdata-testid=rag-topkdata-testid=rag-splitterdata-testid=rag-questiondata-testid=rag-suggest-1data-testid=rag-rundata-testid=rag-compare
Chunks: 0
    Ranked chunks (the top-k are sent)
      Evidence in the context:
      Augmented prompt (what the LLM would receive)
      
          
      Reading the numbers. The score is a TF-IDF cosine between the question and each chunk, from 0 to 1. RAG Explorer shows a different number: 1 - cosine distance between Nomic embeddings, as "% match". The ranking idea is the same; the values are not comparable.
      The evidence check looks for the facts a complete answer needs inside the chunks that were sent. A grounded model cannot use a fact that is not in its context, so this is how chunk size changes the answer before the LLM is even called.

      03Four builds of one idea

      The same pipeline, built with code, without code, and then upgraded for a real corpus. Every row below is in chapter_07_RAG/.

      BuildWhereWhat it showsRun it
      RAG Explorer (React + Express)Basic_RAG/rag-explorerThe VWO PRD through PDF, chunk, Nomic embed, ChromaDB, retrieve and Groq. Pipeline lights, chunk stats, a sample vector, the top-k chunks with % match, the augmented prompt, and a Vector Store tab with a heatmap of every stored 768-d vector.npm run dev
      n8n Basic RAG (no code)n8n_BASIC_RAG/AI3X_Basic_RAG.jsonPhase 1: an upload form, a recursive character splitter (overlap 200), OpenAI embeddings, a Pinecone index. Phase 2: a chat trigger and a RAG Agent (gpt-5-mini, window memory) that uses Pinecone as a retrieval tool with top K 3.Import the JSON, reconnect your own OpenAI and Pinecone credentials
      LangFlow naive RAGLangFlow_RAG/AI_3X_Naive RAG.jsonThe whole 500-row CSV as one text, split on new lines at 1000 / 200, OpenAI text-embedding-3-large into Chroma, 10 results into a prompt for gpt-4o-mini.Import into LangFlow, reconnect credentials, upload the CSV again
      LangFlow improved chunkingLangFlow_RAG/AI_3X_Naive RAG_Imporve_Chunk.jsonEach CSV row rendered as a labelled record (Scenario TID, Description, PreCondition, Test Steps, Expected Result) before a 1000 / 0 split.Same as above
      Advanced RAG Explorer (Flask)Advance_RAG/app.pyHybrid retrieval over 5,000 test cases: bge-m3 dense + sparse, embedded Qdrant, query rewriting, RRF, a cross-encoder reranker, Answer and Generate modes. Tables for dense, sparse, fused and reranked hits per question.python app.py (port 5050)
      CLI ingestAdvance_RAG/ingest.pyThe same ingest without the UI: pick text and metadata columns, chunk, embed in batches of 16, upsert into Qdrant.python ingest.py testcase/vwo_5000_test_cases.csv
      Corpus generatorAdvance_RAG/testcase/generate_testcases.pyWrites the 5,000 seeded test cases (VWO-1001 to VWO-6000, seed 20260705) next to the script.python testcase/generate_testcases.py
      Explainer pageAdvance_RAG/Advanced_RAG_Explained.htmlA standalone animated walkthrough of hybrid embeddings, RRF, the cross-encoder and query rewriting.Open the file in a browser

      The no-code builds are worth one run each: the n8n workflow shows retrieval as a tool an agent decides to call, while LangFlow wires retrieval straight into the prompt. The exported files carry no keys: after import you connect your own credentials.

      04RAG Explorer, line by line

      Three small server files hold the parts you test most: the chunker, the retriever and the prompt. First the chunker: character windows with overlap that try to end on a clean boundary.

      server/lib/chunk.js
      export function chunkText(text, { size = 1200, overlap = 200 } = {}) {
        const clean = text.replace(/\r\n/g, '\n').replace(/\n{3,}/g, '\n\n').trim()
        const chunks = []
        let start = 0
      
        while (start < clean.length) {
          let end = Math.min(start + size, clean.length)
      
          // If we're not at the very end, try to end on a nice boundary within the
          // last 30% of the window (paragraph break > sentence end > space).
          if (end < clean.length) {
            const window = clean.slice(start, end)
            const floor = Math.floor(size * 0.7)
            const para = window.lastIndexOf('\n\n')
            const sentence = Math.max(window.lastIndexOf('. '), window.lastIndexOf('.\n'))
            const space = window.lastIndexOf(' ')
            const cut = para > floor ? para : sentence > floor ? sentence + 1 : space > floor ? space : -1
            if (cut > 0) end = start + cut
          }
      
          const piece = clean.slice(start, end).trim()
          if (piece) {
            chunks.push({
              index: chunks.length,
              text: piece,
              charStart: start,
              charEnd: end,
              length: piece.length,
            })
          }
      
          if (end >= clean.length) break
          start = Math.max(end - overlap, start + 1)
        }
      
        return chunks
      }
      chunkText(text, { size: 1200, overlap: 200 }) chunk 1: start .. start + 1200 last 30% floor = 840 cut 200 chunk 2: cut - 200 .. + 1200 Inside the last 30% of the window: cut at a paragraph break, else a sentence end, else a space. No boundary past the floor: cut at exactly 1200. The next chunk starts 200 characters before the cut.
      One step of chunkText: look for a clean cut only in the last 30% of the window, then step back by the overlap so the next chunk repeats the end of this one.

      Retrieval embeds the question with the same model as the chunks, asks Chroma for the nearest k, and turns cosine distance into the "% match" the UI shows:

      server/lib/chroma.js
      // Embeds the query, retrieves top-k, returns normalized results + the query vector.
      export async function retrieve(collection, queryText, k = 4) {
        const queryEmbedding = await embedQuery(queryText)
        const res = await collection.query({
          queryEmbeddings: [queryEmbedding],
          nResults: k,
          include: ['documents', 'metadatas', 'distances'],
        })
        const docs = res.documents?.[0] || []
        const metas = res.metadatas?.[0] || []
        const dists = res.distances?.[0] || []
        const ids = res.ids?.[0] || []
      
        const results = docs.map((text, i) => {
          const distance = dists[i]
          return {
            id: ids[i],
            text,
            metadata: metas[i] || {},
            distance,
            // cosine distance -> similarity in [0,1] for display
            similarity: typeof distance === 'number' ? Math.max(0, 1 - distance) : null,
          }
        })
        return { results, queryEmbedding }
      }

      The prompt is the product. The system prompt fixes the refusal sentence; buildPrompt numbers the chunks in retrieval order, so [Chunk 1] is the best match, not the first page:

      server/lib/groq.js
      const SYSTEM_PROMPT =
        'You are a precise assistant answering questions about a Product Requirements Document (PRD). ' +
        'Answer ONLY from the provided context chunks. If the answer is not in the context, say so plainly ' +
        '("The document does not cover that."). Be concise, cite which chunk numbers you used, and never invent facts.'
      
      // Builds the augmented prompt from retrieved chunks. Returned so the UI can show it.
      export function buildPrompt(question, chunks) {
        const context = chunks
          .map((c, i) => `[Chunk ${i + 1}]\n${c.text}`)
          .join('\n\n---\n\n')
        return `Context from the PRD:\n\n${context}\n\n---\n\nQuestion: ${question}\n\nAnswer using only the context above.`
      }
      • runIngest() in server/index.js resets the collection on every ingest, so the store always holds one document set.
      • Each chunk is stored with metadata (file, index, charStart, charEnd, length), which is what lets the UI show where a chunk came from.
      • Embedding is one Ollama call per chunk, in sequence, so ingest time grows with the number of chunks.

      05Advanced RAG: hybrid search, fusion and a reranker

      A single dense vector captures meaning but blurs exact tokens such as VWO-3400 or a module name. Advanced RAG keeps both signals: bge-m3 returns a dense vector and sparse lexical weights from one pass, Qdrant stores both, and every query runs both searches.

      flowchart TB
        subgraph Ingest
          direction LR
          CSV[5,000-row CSV] --> BD[build_documents]
          BD --> CK[chunk_text 1000 / 150]
          CK --> M[bge-m3: dense + sparse]
          M --> QD[(Qdrant, embedded)]
        end
        subgraph Chat
          direction LR
          Q[Question] --> RW[rewrite_query x3]
          RW --> DS[dense top 20]
          RW --> SS[sparse top 20]
          DS --> F[rrf_fuse, keep 12]
          SS --> F
          F --> RR[reranker, keep 4]
          RR --> G[generate_answer]
        end
        Ingest -->|searched by| Chat
      Advanced RAG: one bge-m3 pass gives each chunk a dense and a sparse vector. Both searches run for the question and each of its three rewrites, RRF merges the lists, and a cross-encoder picks the final four. detect_mode() only chooses the answer style.

      The two rankings use different score scales, so they are merged by rank, not by score. Reciprocal Rank Fusion adds 1 / (k + rank) from each list:

      rag_core.py
      def rrf_fuse(dense_hits, sparse_hits, k=60, limit=None):
          """Reciprocal Rank Fusion of two ranked lists keyed by point id."""
          scores, meta = {}, {}
          for rank, h in enumerate(dense_hits):
              scores[h["id"]] = scores.get(h["id"], 0.0) + 1.0 / (k + rank + 1)
              meta[h["id"]] = h
          for rank, h in enumerate(sparse_hits):
              scores[h["id"]] = scores.get(h["id"], 0.0) + 1.0 / (k + rank + 1)
              meta.setdefault(h["id"], h)
          fused = sorted(scores.items(), key=lambda kv: -kv[1])
          out = [{"id": pid, "rrf": round(s, 5), "payload": meta[pid]["payload"]} for pid, s in fused]
          return out[:limit] if limit else out
      ChunkDense rankSparse rankRRF score (k = 60)
      A621/66 + 1/62 = 0.03128
      B1not returned1/61 = 0.01639

      A chunk that both searches agree on beats one that only a single search loved. The fused list (12 candidates) then goes to a cross-encoder that reads the question and each chunk together and keeps the best four:

      app.py (the /chat route)
          mode = rag.detect_mode(question)
          t0 = time.time()
      
          # 1) query rewriting
          rewrites = rag.rewrite_query(question, 3) if REWRITE_ENABLED else [question]
          queries = [question] + [r for r in rewrites if r and r != question]
      
          # 2) hybrid retrieval over all query variants
          dense_all, sparse_all = {}, {}
          for qv in queries:
              d, s = rag.embed([qv], batch_size=1)
              for h in store().dense_search(d[0], TOP_N_HYBRID):
                  dense_all[h["id"]] = h if h["id"] not in dense_all else dense_all[h["id"]]
              for h in store().sparse_search(s[0], TOP_N_HYBRID):
                  sparse_all[h["id"]] = h if h["id"] not in sparse_all else sparse_all[h["id"]]
          dense_hits = list(dense_all.values())[:TOP_N_HYBRID]
          sparse_hits = list(sparse_all.values())[:TOP_N_HYBRID]
      
          # 3) RRF fusion
          fused = rag.rrf_fuse(dense_hits, sparse_hits, RRF_K, limit=max(TOP_K_RERANK * 3, 12))
      
          # 4) cross-encoder rerank
          reranked = rag.rerank(question, [dict(f) for f in fused], TOP_K_RERANK)
          STATE["last_chat_ids"] = [c["id"] for c in reranked]
      
          # 5) generation
          answer = rag.generate_answer(question, reranked, mode)
      Two modes. detect_mode() switches to Generate when a request matches (create|generate|write|draft|add) ... (test case|test|scenario): the answer becomes a structured test case (Title, Preconditions, Steps, Expected Result, Priority, Tags) modelled on the retrieved cases.

      06Set up and run

      RAG Explorer needs Node.js 20+, Ollama with the Nomic model, the ChromaDB CLI and a free Groq key.

      terminal
      cd chapter_07_RAG/Basic_RAG/rag-explorer
      npm install
      cp .env.example .env          # set GROQ_API_KEY
      ollama pull nomic-embed-text
      pip install chromadb          # provides the chroma command
      npm run dev                   # ChromaDB :8000, API :8787, UI :5175

      Open the UI, ingest the PRD, then ask one of the four suggested questions. Prefer separate terminals? npm run chroma, npm run server and npm run client start the three parts one by one.

      Advanced RAG needs Python 3.10+. The first ingest and chat download bge-m3 (about 2.3 GB) and the reranker (about 570 MB) once; later runs reuse them.

      terminal
      cd chapter_07_RAG/Advance_RAG
      python3 -m venv .venv && source .venv/bin/activate
      pip install -r requirements.txt
      cp .env.example .env          # set GROQ_API_KEY
      python app.py                 # http://127.0.0.1:5050
      
      # optional headless ingest: stop app.py first (embedded Qdrant has one writer)
      python ingest.py testcase/vwo_5000_test_cases.csv \
        --text-cols "Summary,Description,Steps,Expected Result" \
        --meta-cols "Issue Key,Priority,Component,Test Type"
      BuildVariables (names only)Defaults worth knowing
      RAG ExplorerGROQ_API_KEY (required), GROQ_MODEL, OLLAMA_URL, EMBED_MODEL, CHROMA_URL, CHROMA_COLLECTION, DATA_DIR, CHUNK_SIZE, CHUNK_OVERLAP, TOP_K, PORTChunks 1200 / 200; UI on 5175, API on 8787, ChromaDB on 8000, Ollama on 11434
      Advanced RAGGROQ_API_KEY (or LLM_API_KEY), LLM_BASE_URL, LLM_MODEL, EMBED_MODEL, RERANK_MODEL, BGE_USE_FP16, QDRANT_PATH or QDRANT_URL, QDRANT_COLLECTION, CHUNK_SIZE, CHUNK_OVERLAP, TOP_N_HYBRID, TOP_K_RERANK, RRF_K, REWRITE_ENABLED, INGEST_BATCH, PORTChunks 1000 / 150, 20 hits per search, RRF k 60, 4 chunks after rerank, port 5050
      Keys stay local. Only the answer step (and query rewriting in Advanced RAG) calls Groq. Embeddings, the vector store and the reranker all run on your machine.

      07What to watch

      • Which top-k are you running? The README and the code default say 4 (process.env.TOP_K || 4), but .env.example sets TOP_K=3. Copy the example and you retrieve three chunks.
      • Top-k always returns k chunks. There is no similarity floor in RAG Explorer, so the third chunk can share nothing with the question (the demo's goal question sends a 0.000 chunk at k = 3). Chapter 8 adds a confidence cutoff.
      • Citations are ranks. [Chunk 2] in an answer is the second retrieved chunk, not chunk index 2 in the Vector Store tab.
      • Overlap can start mid-word. The next start is a raw offset (end - overlap), so a chunk preview can begin in the middle of a word. Harmless for retrieval, confusing when you read chunks.
      • The prompt always says PRD. buildPrompt and the system prompt are written for a PRD, so an uploaded test-case file is still introduced as "Context from the PRD".
      • Labelled rows are not row-safe chunks. The improved LangFlow flow labels every row, then splits at 1000 characters with no overlap on line breaks, so test cases are still cut across chunks (an offline emulation found about 1 in 481 chunks starting at a row). Chunk one row at a time, as Advanced RAG does.
      • RAG cannot count. "Which modules have the most security test cases?" needs all 5,000 rows; four chunks cannot answer it. The CSV says Goals & Metrics and Integrations, 30 each. Use a query over the data for aggregates.
      • Generate mode misfires. "write me a summary of test coverage" matches the Generate regex because it contains write and test.
      • One writer at a time. Embedded Qdrant locks qdrant_data/: do not run app.py and ingest.py together.
      • Reranker wrapper bug. FlagEmbedding 1.4's reranker crashes with the current tokenizer, so rag_core.py runs bge-reranker-v2-m3 through transformers directly (see learnings/2026-07-12-flagembedding-reranker-bypass.md).