What this class covered
- Why last week's Langflow RAG failed: the embedding and collection dimensions
- The dimension rule of thumb, with the common embedding models
- A recap of naive RAG: ingestion and retrieval
- Advanced RAG: a re-ranker after retrieval
- Advanced RAG: HyDE before retrieval
- Modular RAG, with a smart router
- Graph RAG for source code, and agentic RAG
- Self-RAG, corrective RAG, and BM25 as an interview topic
- The final project: QABuddy, a hybrid RAG for QA
- The QABuddy project as pushed to the repo
Why the naive RAG failed: dimensions
Last week's Langflow build stopped at the vector store. The class traced it to one setting: the embedding model's dimension did not match the collection's dimension.
An embedding model turns every chunk into a fixed-length list of numbers, its dimension. A collection in a vector database is created with a dimension too, and every vector stored in it must have exactly that length. Use a 3072-dimension model with a collection created at 1024, and nothing lands.
Three more things to check when a Langflow RAG will not run:
- The database token is valid, and the database itself has actually been created.
- Langflow Desktop is slow to list databases and collections: anywhere from a couple of minutes to a quarter of an hour. Waiting is sometimes the fix.
- Ingestion and retrieval use the same embedding model.
The fixed flow reads the file, structures the raw content with a parser, splits it into 1,000-character chunks, and stores them in Astra DB. The fixed flow is in the batch repo as Fixed_AI4X_Naive_RAG.json, beside the 500 test cases. The advanced RAG flow from this class is there too, as Ai4x_Advance_RAG.json: it adds a Cohere re-rank component and a second language model to the retrieval side.
The rule of thumb: look up your embedding model's dimension first, then create the collection with exactly that number.
| Embedding model | Dimension |
|---|---|
| OpenAI text-embedding-3-large | 3072 |
| OpenAI text-embedding-3-small | 1536 |
| OpenAI text-embedding-ada-002 | 1536 |
| Astra DB's NVIDIA default | 1024 |
| nomic-embed-text (Ollama, free) | 768 |
| all-MiniLM-L6-v2 (free) | 384 |
One correction. The class gave text-embedding-3-small as 1024 dimensions. Its default is 1536, the same as ada-002. Given that a dimension mismatch was the very bug being fixed, check the number for your model before you create the collection.
Any vector database works the same way: Pinecone, Qdrant, ChromaDB, Milvus. Free embedding models include BGE, Nomic and Qwen.
Naive RAG, recapped
Naive RAG has two halves. Ingestion: read the documents, parse and clean them, split them into chunks, embed each chunk, and store the vectors. Retrieval: embed the question with the same model, fetch the closest chunks, and pass them to the LLM with the question to write the answer.
Advanced RAG keeps ingestion exactly as it is. Both of its additions are on the retrieval side.
Advanced RAG: the re-ranker
A vector database returns the chunks most similar to the question, the top K, say 20. They are similar, but not ranked by how well they answer the question. A re-ranker reads the question and each of those 20 chunks together, scores them, and passes only the best three or four to the LLM as context.
- Paid: Cohere Rerank is the best-known; a free trial key comes from the Cohere dashboard. Pinecone also sells one, and some vector databases have re-ranking built in.
- Free: BGE re-rankers, Qwen re-rankers, and models served through Ollama.
Advanced RAG: HyDE
A question like "give login TC" is a poor search query. The vector database has little to match it against. The class's step before retrieval was an LLM that turns the weak question into better search text, such as "test cases for the login feature, including valid and invalid credentials". In the class's words, it is a better prompter than you, so ask it first.
The class's analogy. Wikipedia's definition of REST is accurate, but hard to follow. A good teacher explains the same idea in plain words, constraint by constraint. HyDE does the same for your question before it reaches the database.
Strictly, HyDE (hypothetical document embeddings) has the LLM write a hypothetical answer to the question, and searches with that answer's embedding, because an answer looks more like the stored chunks than a question does. Rewriting one question into several better ones has its own name, multi-query or query expansion. Both happen before retrieval, for the same reason, and an interviewer may ask you to tell them apart.
Modular RAG: a smart router
Ingestion is unchanged here too. Modular RAG keeps several collections, such as API, UI and performance test cases, and puts an LLM in front of them as a router. The router reads the question and sends it to the right collection. The Langflow demo routed between two collections, positive and negative test cases, using a smart router component. You choose the model that does the routing, such as GPT, Gemma or DeepSeek.
Graph RAG and agentic RAG
- Graph RAG is for source code. Code does not chunk well as plain text, so it is parsed into an abstract syntax tree (AST) and stored in a graph database such as Neo4j, where its structure can be searched.
- Agentic RAG puts an agent in place of the router. The agent decides when to retrieve, which collection to search, and when it has enough information, following the ReAct pattern of reasoning, then acting. It returns later in the course, when RAG systems are tested with DeepEval.
The class described self-RAG and corrective RAG as rarely used in practice. The type you will meet most is hybrid RAG, and the final project is one.
BM25 is the keyword side, not the semantic side. The class flagged BM25 as a common interview topic and described it as a semantic match. It is the opposite. BM25 ranks chunks by keyword overlap: how often the query's words appear, weighted by how rare they are. Hybrid search pairs BM25 with semantic vector search, because each catches what the other misses. The QABuddy results below show it.
The final project: QABuddy
The plan from class: a QABuddy, or QA copilot, that ingests Jira, PDFs, the Selenium and Playwright repositories and meeting transcripts, and answers questions about all of them in one chat. It combines normal RAG for the documents with graph RAG for the code, which is what the class meant by a hybrid RAG. Keeping it up to date, and stopping old and new versions of a ticket from conflicting, is handled with metadata.
The class built it with Claude. Command Code, OpenCode, or Antigravity with its free Gemini quota work as well.
The build started from a brief written in class, pushed as QABUDDY_AI_BUILD_.prompt.md. In it, the instructor asks the agent to recommend an open-source embedding model, an open-source vector database, the chunk sizes and overlaps, and a plan for a chatbot running on a DigitalOcean machine. The brief also lists the ten data sources and puts hourly auto-ingestion in phase two. A cleaned-up version, QABUDDY_Improved_.prompt.md, sits beside it.
QABuddy as pushed to the repo
The result is in the batch repo as chapter_12_RAG_QA_BuddyAI: a Python pipeline, a React chat UI, deployment files and an evaluation set. It is self-hosted: the embedding model and the vector database run on your own machine or server.
The repo's README records the five decisions the brief asked for:
| Decision | QABuddy's choice | Why, per the README |
|---|---|---|
| Embedding model | Qwen3-Embedding 4B through Ollama, cut to 1024 dimensions | strong on retrieval benchmarks, handles code, long inputs, Apache 2.0 |
| Vector database | Qdrant | one collection holding dense and BM25 vectors, with native fusion and filters |
| Chunking | along the unit a QA engineer asks about | one test case per chunk, one Jira comment, one heading-aware section, one code method |
| Preprocessing | normalise, redact, add context | secrets redacted before indexing; every chunk starts with where it came from |
| Architecture | ingest, then exact-ID lookup, hybrid search, re-rank, select, answer | the diagram above |
A few chunking choices worth copying, from the README's per-source table:
| Source | One chunk is | Why |
|---|---|---|
| Test cases | one row, about 190 tokens, no overlap | a test case is already atomic, and splitting one glued half of it to the next in chapter 11 |
| Jira tickets | the summary and description, then each comment separately | comments carry the root cause, so they must be findable on their own |
| Requirement documents | one heading-aware section, about 500 tokens | answers live in a section |
| Jenkins logs | a build summary, plus one chunk per failure window | most of a log is noise, and repeated failures are de-duplicated |
| Source code | one class, method or test block | a method is never cut in half |
Two differences from the plan. The graph RAG shows up as code-aware chunking: the code is parsed by structure with tree-sitter and stored in the same Qdrant collection, not in a separate graph database. And "hybrid" gains a second meaning, hybrid search: every question runs dense and BM25 search together and fuses the two rankings.
How well it retrieves, and what an answer looks like
The repo includes a 24-question evaluation set. The README's table comes from the recorded run, eval/last_run.json, and the two match:
| Retriever | hit@6 | MRR |
|---|---|---|
| dense only (Qwen3) | 100% | 0.83 |
| BM25 only | 100% | 0.88 |
| hybrid (Qdrant RRF fusion) | 100% | 0.91 |
| full pipeline | 100% | 0.91 |
Hit@6 asks whether the right chunk appears anywhere in the top six. MRR (mean reciprocal rank) rewards putting it near the top: 1.0 means always first. Every retriever found the answer, so the difference is in the ranking. BM25 beats dense search here, which is the point of the earlier note. Take the README's example "Is testLoginPositiveVWO flaky?": by meaning, the ticket that answers it ranks 31st; by keyword, 2nd.
The repo also shows how that answer was assembled. The retrieval trace lists each candidate chunk with its rank from dense search, from BM25 and after fusion, then the re-ranker's score and the final decision:
Try it: the hosted demo is at qabuddy-ai-nine.vercel.app. A public page cannot reach a self-hosted Qdrant, Ollama or re-ranker, so the demo replays answers recorded through the full local pipeline, retrieval traces included. New questions are answered with BM25 in the browser plus a serverless function, limited to 20 questions an hour per visitor.
Run it: ./run.sh starts Qdrant and Ollama, ingests on first run, and opens the UI locally. ./run.sh test runs the 18 regression tests, each pinning a bug found while building; they need no database or model and pass as pushed. ./run.sh eval re-runs the retrieval evaluation.
The README is frank about its limits. Groq's free tier serves about two answers a minute. The 4B embedding model is slow on a CPU-only server. The Docker deployment has not been run end to end. And grounding rules reduce model errors without removing them, which is why every answer cites its sources.
Tasks and announcements
Task, the fixed naive RAG. Rebuild last week's Langflow RAG today, using the fixed flow in the batch repo as a reference. Check the embedding model's dimension against the collection's before you ingest.
- QABuddy: the code, sample data and prompts are in the batch repo. Clone it with
--recurse-submodulesto fetch the two framework repositories it indexes. - The advanced MCP session is being organised, probably this week.