The Testing Academy · Class Notes Saturday, 3 October (IST)
Live class · study guide

Embedding dimensions, advanced RAG with a re-ranker and HyDE, modular and graph RAG, and QABuddy, a hybrid RAG for QA

Why last week's Langflow RAG failed (an embedding dimension that did not match the collection), the move from naive to advanced RAG with a re-ranker and HyDE, the other RAG types from modular to graph and agentic, and the final project: QABuddy, a self-hosted hybrid RAG that answers QA questions with citations across test cases, Jira, documents, logs and code.

By Pramod Dutta, The Testing Academy. Study notes from the live AI Tester Blueprint 4x class, built from the session recording and the batch repository, which received the QABuddy project that afternoon (commit "feat: add chapter 12 QABuddy hybrid RAG"). The recording covers the class up to the break; the build after it was not available when this page was written, so the QABuddy sections describe the project as pushed. Its 18 regression tests were run and pass, and its hosted demo was checked live. The retrieval evaluation was not re-run: the figures shown are the repo's recorded run. The Eraser deck was not reachable while this page was written.

01

What this class covered

  • Why last week's Langflow RAG failed: the embedding and collection dimensions
  • The dimension rule of thumb, with the common embedding models
  • A recap of naive RAG: ingestion and retrieval
  • Advanced RAG: a re-ranker after retrieval
  • Advanced RAG: HyDE before retrieval
  • Modular RAG, with a smart router
  • Graph RAG for source code, and agentic RAG
  • Self-RAG, corrective RAG, and BM25 as an interview topic
  • The final project: QABuddy, a hybrid RAG for QA
  • The QABuddy project as pushed to the repo
02

Why the naive RAG failed: dimensions

Last week's Langflow build stopped at the vector store. The class traced it to one setting: the embedding model's dimension did not match the collection's dimension.

An embedding model turns every chunk into a fixed-length list of numbers, its dimension. A collection in a vector database is created with a dimension too, and every vector stored in it must have exactly that length. Use a 3072-dimension model with a collection created at 1024, and nothing lands.

THE ONE RULE: MODEL DIMENSION = COLLECTION DIMENSION embedding model 3072 numbers per chunk collection dimension 3072 stored and searchable the vectors fit the collection the same model 3072 numbers per chunk collection dimension 1024 nothing lands whatever the collection is called Retrieval embeds the question with the same model, so the same dimension applies on the way out.
A collection's name can say anything. Only the dimension it was created with decides which embedding model it accepts.

Three more things to check when a Langflow RAG will not run:

  • The database token is valid, and the database itself has actually been created.
  • Langflow Desktop is slow to list databases and collections: anywhere from a couple of minutes to a quarter of an hour. Waiting is sometimes the fix.
  • Ingestion and retrieval use the same embedding model.

The fixed flow reads the file, structures the raw content with a parser, splits it into 1,000-character chunks, and stores them in Astra DB. The fixed flow is in the batch repo as Fixed_AI4X_Naive_RAG.json, beside the 500 test cases. The advanced RAG flow from this class is there too, as Ai4x_Advance_RAG.json: it adds a Cohere re-rank component and a second language model to the retrieval side.

The rule of thumb: look up your embedding model's dimension first, then create the collection with exactly that number.

Embedding model Dimension
OpenAI text-embedding-3-large 3072
OpenAI text-embedding-3-small 1536
OpenAI text-embedding-ada-002 1536
Astra DB's NVIDIA default 1024
nomic-embed-text (Ollama, free) 768
all-MiniLM-L6-v2 (free) 384

One correction. The class gave text-embedding-3-small as 1024 dimensions. Its default is 1536, the same as ada-002. Given that a dimension mismatch was the very bug being fixed, check the number for your model before you create the collection.

Any vector database works the same way: Pinecone, Qdrant, ChromaDB, Milvus. Free embedding models include BGE, Nomic and Qwen.

03

Naive RAG, recapped

Naive RAG has two halves. Ingestion: read the documents, parse and clean them, split them into chunks, embed each chunk, and store the vectors. Retrieval: embed the question with the same model, fetch the closest chunks, and pass them to the LLM with the question to write the answer.

Advanced RAG keeps ingestion exactly as it is. Both of its additions are on the retrieval side.

NAIVE RETRIEVAL question embed vector database top 3 chunks LLM answer ADVANCED RETRIEVAL question HyDE (LLM) better search text embed vector database top 20 chunks re-ranker keeps the best 3-4 LLM answer The two shaded steps are the whole difference. Ingestion, the embedding model and the database stay the same.
HyDE improves what goes into the search. The re-ranker improves what comes out of it.
04

Advanced RAG: the re-ranker

A vector database returns the chunks most similar to the question, the top K, say 20. They are similar, but not ranked by how well they answer the question. A re-ranker reads the question and each of those 20 chunks together, scores them, and passes only the best three or four to the LLM as context.

  • Paid: Cohere Rerank is the best-known; a free trial key comes from the Cohere dashboard. Pinecone also sells one, and some vector databases have re-ranking built in.
  • Free: BGE re-rankers, Qwen re-rankers, and models served through Ollama.
05

Advanced RAG: HyDE

A question like "give login TC" is a poor search query. The vector database has little to match it against. The class's step before retrieval was an LLM that turns the weak question into better search text, such as "test cases for the login feature, including valid and invalid credentials". In the class's words, it is a better prompter than you, so ask it first.

The class's analogy. Wikipedia's definition of REST is accurate, but hard to follow. A good teacher explains the same idea in plain words, constraint by constraint. HyDE does the same for your question before it reaches the database.

Strictly, HyDE (hypothetical document embeddings) has the LLM write a hypothetical answer to the question, and searches with that answer's embedding, because an answer looks more like the stored chunks than a question does. Rewriting one question into several better ones has its own name, multi-query or query expansion. Both happen before retrieval, for the same reason, and an interviewer may ask you to tell them apart.

06

Modular RAG: a smart router

Ingestion is unchanged here too. Modular RAG keeps several collections, such as API, UI and performance test cases, and puts an LLM in front of them as a router. The router reads the question and sends it to the right collection. The Langflow demo routed between two collections, positive and negative test cases, using a smart router component. You choose the model that does the routing, such as GPT, Gemma or DeepSeek.

07

Graph RAG and agentic RAG

  • Graph RAG is for source code. Code does not chunk well as plain text, so it is parsed into an abstract syntax tree (AST) and stored in a graph database such as Neo4j, where its structure can be searched.
  • Agentic RAG puts an agent in place of the router. The agent decides when to retrieve, which collection to search, and when it has enough information, following the ReAct pattern of reasoning, then acting. It returns later in the course, when RAG systems are tested with DeepEval.

The class described self-RAG and corrective RAG as rarely used in practice. The type you will meet most is hybrid RAG, and the final project is one.

BM25 is the keyword side, not the semantic side. The class flagged BM25 as a common interview topic and described it as a semantic match. It is the opposite. BM25 ranks chunks by keyword overlap: how often the query's words appear, weighted by how rare they are. Hybrid search pairs BM25 with semantic vector search, because each catches what the other misses. The QABuddy results below show it.

08

The final project: QABuddy

The plan from class: a QABuddy, or QA copilot, that ingests Jira, PDFs, the Selenium and Playwright repositories and meeting transcripts, and answers questions about all of them in one chat. It combines normal RAG for the documents with graph RAG for the code, which is what the class meant by a hybrid RAG. Keeping it up to date, and stopping old and new versions of a ticket from conflicting, is handled with metadata.

The class built it with Claude. Command Code, OpenCode, or Antigravity with its free Gemini quota work as well.

The build started from a brief written in class, pushed as QABUDDY_AI_BUILD_.prompt.md. In it, the instructor asks the agent to recommend an open-source embedding model, an open-source vector database, the chunk sizes and overlaps, and a plan for a chatbot running on a DigitalOcean machine. The brief also lists the ten data sources and puts hourly auto-ingestion in phase two. A cleaned-up version, QABUDDY_Improved_.prompt.md, sits beside it.

09

QABuddy as pushed to the repo

The result is in the batch repo as chapter_12_RAG_QA_BuddyAI: a Python pipeline, a React chat UI, deployment files and an evaluation set. It is self-hosted: the embedding model and the vector database run on your own machine or server.

INGEST 10 data folders tests, Jira, docs, logs, code source-aware chunkers Qwen3-Embedding 4B, 1024-d code-aware BM25 Qdrant: one collection a dense and a BM25 vector per chunk ASK question + mode (RCA, RTM...) exact ID lookup + hybrid search, RRF re-ranker bge-reranker-v2-m3 selection quotas, token budget LLM cite or refuse answer [n] citations The LLM is gpt-oss-120b on Groq. Every answer carries numbered citations and a trace of how each chunk was ranked. Code goes through tree-sitter: chunks follow classes, methods and test blocks, so a method is never cut in half.
Ingestion stores two kinds of vector per chunk. Asking combines them, re-ranks, and lets the LLM answer only from what it was given.

The repo's README records the five decisions the brief asked for:

Decision QABuddy's choice Why, per the README
Embedding model Qwen3-Embedding 4B through Ollama, cut to 1024 dimensions strong on retrieval benchmarks, handles code, long inputs, Apache 2.0
Vector database Qdrant one collection holding dense and BM25 vectors, with native fusion and filters
Chunking along the unit a QA engineer asks about one test case per chunk, one Jira comment, one heading-aware section, one code method
Preprocessing normalise, redact, add context secrets redacted before indexing; every chunk starts with where it came from
Architecture ingest, then exact-ID lookup, hybrid search, re-rank, select, answer the diagram above

A few chunking choices worth copying, from the README's per-source table:

Source One chunk is Why
Test cases one row, about 190 tokens, no overlap a test case is already atomic, and splitting one glued half of it to the next in chapter 11
Jira tickets the summary and description, then each comment separately comments carry the root cause, so they must be findable on their own
Requirement documents one heading-aware section, about 500 tokens answers live in a section
Jenkins logs a build summary, plus one chunk per failure window most of a log is noise, and repeated failures are de-duplicated
Source code one class, method or test block a method is never cut in half

Two differences from the plan. The graph RAG shows up as code-aware chunking: the code is parsed by structure with tree-sitter and stored in the same Qdrant collection, not in a separate graph database. And "hybrid" gains a second meaning, hybrid search: every question runs dense and BM25 search together and fuses the two rankings.

10

How well it retrieves, and what an answer looks like

The repo includes a 24-question evaluation set. The README's table comes from the recorded run, eval/last_run.json, and the two match:

Retriever hit@6 MRR
dense only (Qwen3) 100% 0.83
BM25 only 100% 0.88
hybrid (Qdrant RRF fusion) 100% 0.91
full pipeline 100% 0.91

Hit@6 asks whether the right chunk appears anywhere in the top six. MRR (mean reciprocal rank) rewards putting it near the top: 1.0 means always first. Every retriever found the answer, so the difference is in the ranking. BM25 beats dense search here, which is the point of the earlier note. Take the README's example "Is testLoginPositiveVWO flaky?": by meaning, the ticket that answers it ranks 31st; by keyword, 2nd.

QABuddy answering a failure-analysis question. The answer lists the related tickets: QAB-101, CI login tests failing after a Jenkins agent migration; VWO-26, a customer report of the same symptom; and VWO-33, marked a duplicate of VWO-26. It then gives the fix and next steps: add the CI agent's IP address to the test account's allowlist, add a pre-flight login check stage to the regression pipeline, replace a fixed five-second wait in LoginPage.loginToVWOLoginValidCreds with an explicit Selenium wait, and rerun build 142. Small numbered badges cite eight source cards below the answer: a Jira ticket, the login-failure triage meeting notes, and several Jenkins log windows. A status line shows 5.0 seconds in total, 1.1 seconds of re-ranking, the first token at 920 milliseconds, and 3,301 tokens in and 687 out. The header shows the stack: Qdrant with 838 chunks, qwen3-embedding 4b at 1024 dimensions, bge-reranker-v2-m3, and gpt-oss-120b on Groq. The sidebar lists the six modes and the knowledge sources with their chunk counts.
A failure-analysis answer, from the repo's own screenshots. Every claim carries a numbered citation to a source card, so a human can check it.

The repo also shows how that answer was assembled. The retrieval trace lists each candidate chunk with its rank from dense search, from BM25 and after fusion, then the re-ranker's score and the final decision:

The retrieval trace behind a QABuddy answer: a table with columns for source, chunk title, dense rank, BM25 rank, fused RRF rank, re-ranker score and decision. Rows scoring between 0.915 and 0.955 are marked selected, one of them through the Jenkins quota. They include Jenkins log windows for two failing login tests, build 142's failure and summary, and the login-failure triage meeting notes, which ranked 14th, 12th and 13th in the first stage. The TestSuite results summary, ranked first by dense, BM25 and fusion alike, scored 0.810 and is marked beyond top-k, as are lower rows such as the CI pipeline diagram, a later build, stand-up notes and onboarding guide sections.
The re-ranker at work. The triage meeting notes ranked 13th after fusion and made the answer; the summary every first-stage retriever ranked first did not.

Try it: the hosted demo is at qabuddy-ai-nine.vercel.app. A public page cannot reach a self-hosted Qdrant, Ollama or re-ranker, so the demo replays answers recorded through the full local pipeline, retrieval traces included. New questions are answered with BM25 in the browser plus a serverless function, limited to 20 questions an hour per visitor.

Run it: ./run.sh starts Qdrant and Ollama, ingests on first run, and opens the UI locally. ./run.sh test runs the 18 regression tests, each pinning a bug found while building; they need no database or model and pass as pushed. ./run.sh eval re-runs the retrieval evaluation.

The README is frank about its limits. Groq's free tier serves about two answers a minute. The 4B embedding model is slow on a CPU-only server. The Docker deployment has not been run end to end. And grounding rules reduce model errors without removing them, which is why every answer cites its sources.

11

Tasks and announcements

Task, the fixed naive RAG. Rebuild last week's Langflow RAG today, using the fixed flow in the batch repo as a reference. Check the embedding model's dimension against the collection's before you ingest.

  • QABuddy: the code, sample data and prompts are in the batch repo. Clone it with --recurse-submodules to fetch the two framework repositories it indexes.
  • The advanced MCP session is being organised, probably this week.