An LLM only knows what it was trained on. Retrieval-Augmented Generation hands it the right part of your PRD or test suite at question time: split the document into chunks, find the chunks that match the question, and send only those with a rule to answer from them. This chapter builds that pipeline four ways, from a React app that shows every stage to a hybrid pipeline over 5,000 test cases.
RAG Explorer, an n8n workflow, two LangFlow flows and the Advanced RAG Flask app.
7
Chunks from the PRD
The 6-page VWO PRD at the default CHUNK_SIZE=1200 and CHUNK_OVERLAP=200.
5,000
Test cases in Advanced RAG
Seeded VWO cases, one chunk per row, dense and sparse vectors in embedded Qdrant.
768
Embedding dimensions
nomic-embed-text through local Ollama. bge-m3 dense vectors in Advanced RAG have 1024.
01Why RAG, and why a tester cares
Ask a general-purpose LLM what the VWO PRD lists as success metrics and it will guess, fluently. RAG fixes that by putting the right part of the document into the prompt at question time. The model is told to answer only from those chunks, and to say so when they do not cover the question.
For a tester, the useful part is that a RAG system has seams you can check one at a time:
Chunking. Did the fact survive the split, or is half of the list in another chunk?
Retrieval. Did the chunk that holds the fact make the top-k?
The prompt. Does it carry those chunks and an answer-only-from-context rule?
Generation. Did the model use the context, cite it, and refuse when the context is silent?
When an answer is wrong, walk the seams in that order. RAG Explorer puts each one on screen: chunk counts, a sample vector, the retrieved chunks with their similarity, and the exact prompt sent to Groq.
flowchart TB
subgraph Ingest
direction LR
P[PRD PDF in data/] --> X[pdf-parse: text]
X --> C[chunkText 1200 / 200]
C --> E[nomic-embed-text: 768-d]
E --> S[(ChromaDB vwo_prd)]
end
subgraph Ask
direction LR
Q[Question] --> QE[embed, same model]
QE --> R[query: top-k by cosine]
R --> B[buildPrompt]
B --> G[Groq gpt-oss-120b]
G --> A[answer + chunk numbers]
end
Ingest -->|searched by| Ask
RAG Explorer: ingest once (read, chunk, embed, store), then embed each question with the same model and send only the top-k chunks to Groq.
02Live demo: chunk, rank, ground
Run the pipeline by hand. The chunker is a line-for-line port of the chapter's chunk.js, and the prompt is built the same way as groq.js. One stand-in: instead of 768-dimension Nomic vectors, each chunk gets a TF-IDF score against the question, so the ranking runs offline and is identical every time.
Start with What is the goal of this PRD? at 1200 / 200, then set the chunk size to 300 and the overlap to 50 and watch the evidence count drop. Switch to the test cases to see what windows do to records.
Chunk, rank, ground: the RAG pipeline by handNo API key needed
Ranked chunks (the top-k are sent)Evidence in the context:
Augmented prompt (what the LLM would receive)
Same question, four chunking settings
Setting
Chunks
Sent (top-k)
Evidence
Prompt chars
Reading the numbers. The score is a TF-IDF cosine between the question and each chunk, from 0 to 1. RAG Explorer shows a different number: 1 - cosine distance between Nomic embeddings, as "% match". The ranking idea is the same; the values are not comparable.
The evidence check looks for the facts a complete answer needs inside the chunks that were sent. A grounded model cannot use a fact that is not in its context, so this is how chunk size changes the answer before the LLM is even called.
03Four builds of one idea
The same pipeline, built with code, without code, and then upgraded for a real corpus. Every row below is in chapter_07_RAG/.
Build
Where
What it shows
Run it
RAG Explorer (React + Express)
Basic_RAG/rag-explorer
The VWO PRD through PDF, chunk, Nomic embed, ChromaDB, retrieve and Groq. Pipeline lights, chunk stats, a sample vector, the top-k chunks with % match, the augmented prompt, and a Vector Store tab with a heatmap of every stored 768-d vector.
npm run dev
n8n Basic RAG (no code)
n8n_BASIC_RAG/AI3X_Basic_RAG.json
Phase 1: an upload form, a recursive character splitter (overlap 200), OpenAI embeddings, a Pinecone index. Phase 2: a chat trigger and a RAG Agent (gpt-5-mini, window memory) that uses Pinecone as a retrieval tool with top K 3.
Import the JSON, reconnect your own OpenAI and Pinecone credentials
LangFlow naive RAG
LangFlow_RAG/AI_3X_Naive RAG.json
The whole 500-row CSV as one text, split on new lines at 1000 / 200, OpenAI text-embedding-3-large into Chroma, 10 results into a prompt for gpt-4o-mini.
Import into LangFlow, reconnect credentials, upload the CSV again
LangFlow improved chunking
LangFlow_RAG/AI_3X_Naive RAG_Imporve_Chunk.json
Each CSV row rendered as a labelled record (Scenario TID, Description, PreCondition, Test Steps, Expected Result) before a 1000 / 0 split.
Same as above
Advanced RAG Explorer (Flask)
Advance_RAG/app.py
Hybrid retrieval over 5,000 test cases: bge-m3 dense + sparse, embedded Qdrant, query rewriting, RRF, a cross-encoder reranker, Answer and Generate modes. Tables for dense, sparse, fused and reranked hits per question.
python app.py (port 5050)
CLI ingest
Advance_RAG/ingest.py
The same ingest without the UI: pick text and metadata columns, chunk, embed in batches of 16, upsert into Qdrant.
python ingest.py testcase/vwo_5000_test_cases.csv
Corpus generator
Advance_RAG/testcase/generate_testcases.py
Writes the 5,000 seeded test cases (VWO-1001 to VWO-6000, seed 20260705) next to the script.
python testcase/generate_testcases.py
Explainer page
Advance_RAG/Advanced_RAG_Explained.html
A standalone animated walkthrough of hybrid embeddings, RRF, the cross-encoder and query rewriting.
Open the file in a browser
The no-code builds are worth one run each: the n8n workflow shows retrieval as a tool an agent decides to call, while LangFlow wires retrieval straight into the prompt. The exported files carry no keys: after import you connect your own credentials.
04RAG Explorer, line by line
Three small server files hold the parts you test most: the chunker, the retriever and the prompt. First the chunker: character windows with overlap that try to end on a clean boundary.
server/lib/chunk.js
exportfunctionchunkText(text, { size = 1200, overlap = 200 } = {}) {
const clean = text.replace(/\r\n/g, '\n').replace(/\n{3,}/g, '\n\n').trim()
const chunks = []
let start = 0while (start < clean.length) {
let end = Math.min(start + size, clean.length)
// If we're not at the very end, try to end on a nice boundary within the// last 30% of the window (paragraph break > sentence end > space).if (end < clean.length) {
const window = clean.slice(start, end)
const floor = Math.floor(size * 0.7)
const para = window.lastIndexOf('\n\n')
const sentence = Math.max(window.lastIndexOf('. '), window.lastIndexOf('.\n'))
const space = window.lastIndexOf(' ')
const cut = para > floor ? para : sentence > floor ? sentence + 1 : space > floor ? space : -1if (cut > 0) end = start + cut
}
const piece = clean.slice(start, end).trim()
if (piece) {
chunks.push({
index: chunks.length,
text: piece,
charStart: start,
charEnd: end,
length: piece.length,
})
}
if (end >= clean.length) break
start = Math.max(end - overlap, start + 1)
}
return chunks
}
One step of chunkText: look for a clean cut only in the last 30% of the window, then step back by the overlap so the next chunk repeats the end of this one.
Retrieval embeds the question with the same model as the chunks, asks Chroma for the nearest k, and turns cosine distance into the "% match" the UI shows:
The prompt is the product. The system prompt fixes the refusal sentence; buildPrompt numbers the chunks in retrieval order, so [Chunk 1] is the best match, not the first page:
server/lib/groq.js
const SYSTEM_PROMPT =
'You are a precise assistant answering questions about a Product Requirements Document (PRD). ' +
'Answer ONLY from the provided context chunks. If the answer is not in the context, say so plainly ' +
'("The document does not cover that."). Be concise, cite which chunk numbers you used, and never invent facts.'// Builds the augmented prompt from retrieved chunks. Returned so the UI can show it.exportfunctionbuildPrompt(question, chunks) {
const context = chunks
.map((c, i) => `[Chunk ${i + 1}]\n${c.text}`)
.join('\n\n---\n\n')
return`Context from the PRD:\n\n${context}\n\n---\n\nQuestion: ${question}\n\nAnswer using only the context above.`
}
runIngest() in server/index.js resets the collection on every ingest, so the store always holds one document set.
Each chunk is stored with metadata (file, index, charStart, charEnd, length), which is what lets the UI show where a chunk came from.
Embedding is one Ollama call per chunk, in sequence, so ingest time grows with the number of chunks.
05Advanced RAG: hybrid search, fusion and a reranker
A single dense vector captures meaning but blurs exact tokens such as VWO-3400 or a module name. Advanced RAG keeps both signals: bge-m3 returns a dense vector and sparse lexical weights from one pass, Qdrant stores both, and every query runs both searches.
flowchart TB
subgraph Ingest
direction LR
CSV[5,000-row CSV] --> BD[build_documents]
BD --> CK[chunk_text 1000 / 150]
CK --> M[bge-m3: dense + sparse]
M --> QD[(Qdrant, embedded)]
end
subgraph Chat
direction LR
Q[Question] --> RW[rewrite_query x3]
RW --> DS[dense top 20]
RW --> SS[sparse top 20]
DS --> F[rrf_fuse, keep 12]
SS --> F
F --> RR[reranker, keep 4]
RR --> G[generate_answer]
end
Ingest -->|searched by| Chat
Advanced RAG: one bge-m3 pass gives each chunk a dense and a sparse vector. Both searches run for the question and each of its three rewrites, RRF merges the lists, and a cross-encoder picks the final four. detect_mode() only chooses the answer style.
The two rankings use different score scales, so they are merged by rank, not by score. Reciprocal Rank Fusion adds 1 / (k + rank) from each list:
rag_core.py
defrrf_fuse(dense_hits, sparse_hits, k=60, limit=None):
"""Reciprocal Rank Fusion of two ranked lists keyed by point id."""
scores, meta = {}, {}
for rank, h inenumerate(dense_hits):
scores[h["id"]] = scores.get(h["id"], 0.0) + 1.0 / (k + rank + 1)
meta[h["id"]] = h
for rank, h inenumerate(sparse_hits):
scores[h["id"]] = scores.get(h["id"], 0.0) + 1.0 / (k + rank + 1)
meta.setdefault(h["id"], h)
fused = sorted(scores.items(), key=lambda kv: -kv[1])
out = [{"id": pid, "rrf": round(s, 5), "payload": meta[pid]["payload"]} for pid, s in fused]
return out[:limit] if limit else out
Chunk
Dense rank
Sparse rank
RRF score (k = 60)
A
6
2
1/66 + 1/62 = 0.03128
B
1
not returned
1/61 = 0.01639
A chunk that both searches agree on beats one that only a single search loved. The fused list (12 candidates) then goes to a cross-encoder that reads the question and each chunk together and keeps the best four:
app.py (the /chat route)
mode = rag.detect_mode(question)
t0 = time.time()
# 1) query rewriting
rewrites = rag.rewrite_query(question, 3) if REWRITE_ENABLED else [question]
queries = [question] + [r for r in rewrites if r and r != question]
# 2) hybrid retrieval over all query variants
dense_all, sparse_all = {}, {}
for qv in queries:
d, s = rag.embed([qv], batch_size=1)
for h instore().dense_search(d[0], TOP_N_HYBRID):
dense_all[h["id"]] = h if h["id"] notin dense_all else dense_all[h["id"]]
for h instore().sparse_search(s[0], TOP_N_HYBRID):
sparse_all[h["id"]] = h if h["id"] notin sparse_all else sparse_all[h["id"]]
dense_hits = list(dense_all.values())[:TOP_N_HYBRID]
sparse_hits = list(sparse_all.values())[:TOP_N_HYBRID]
# 3) RRF fusion
fused = rag.rrf_fuse(dense_hits, sparse_hits, RRF_K, limit=max(TOP_K_RERANK * 3, 12))
# 4) cross-encoder rerank
reranked = rag.rerank(question, [dict(f) for f in fused], TOP_K_RERANK)
STATE["last_chat_ids"] = [c["id"] for c in reranked]
# 5) generation
answer = rag.generate_answer(question, reranked, mode)
Two modes.detect_mode() switches to Generate when a request matches (create|generate|write|draft|add) ... (test case|test|scenario): the answer becomes a structured test case (Title, Preconditions, Steps, Expected Result, Priority, Tags) modelled on the retrieved cases.
06Set up and run
RAG Explorer needs Node.js 20+, Ollama with the Nomic model, the ChromaDB CLI and a free Groq key.
terminal
cd chapter_07_RAG/Basic_RAG/rag-explorer
npm install
cp .env.example .env # set GROQ_API_KEYollama pull nomic-embed-text
pip install chromadb # provides the chroma commandnpm run dev # ChromaDB :8000, API :8787, UI :5175
Open the UI, ingest the PRD, then ask one of the four suggested questions. Prefer separate terminals? npm run chroma, npm run server and npm run client start the three parts one by one.
Advanced RAG needs Python 3.10+. The first ingest and chat download bge-m3 (about 2.3 GB) and the reranker (about 570 MB) once; later runs reuse them.
terminal
cd chapter_07_RAG/Advance_RAG
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env # set GROQ_API_KEYpython app.py # http://127.0.0.1:5050# optional headless ingest: stop app.py first (embedded Qdrant has one writer)python ingest.py testcase/vwo_5000_test_cases.csv \
--text-cols "Summary,Description,Steps,Expected Result" \
--meta-cols "Issue Key,Priority,Component,Test Type"
Chunks 1200 / 200; UI on 5175, API on 8787, ChromaDB on 8000, Ollama on 11434
Advanced RAG
GROQ_API_KEY (or LLM_API_KEY), LLM_BASE_URL, LLM_MODEL, EMBED_MODEL, RERANK_MODEL, BGE_USE_FP16, QDRANT_PATH or QDRANT_URL, QDRANT_COLLECTION, CHUNK_SIZE, CHUNK_OVERLAP, TOP_N_HYBRID, TOP_K_RERANK, RRF_K, REWRITE_ENABLED, INGEST_BATCH, PORT
Chunks 1000 / 150, 20 hits per search, RRF k 60, 4 chunks after rerank, port 5050
Keys stay local. Only the answer step (and query rewriting in Advanced RAG) calls Groq. Embeddings, the vector store and the reranker all run on your machine.
07What to watch
Which top-k are you running? The README and the code default say 4 (process.env.TOP_K || 4), but .env.example sets TOP_K=3. Copy the example and you retrieve three chunks.
Top-k always returns k chunks. There is no similarity floor in RAG Explorer, so the third chunk can share nothing with the question (the demo's goal question sends a 0.000 chunk at k = 3). Chapter 8 adds a confidence cutoff.
Citations are ranks.[Chunk 2] in an answer is the second retrieved chunk, not chunk index 2 in the Vector Store tab.
Overlap can start mid-word. The next start is a raw offset (end - overlap), so a chunk preview can begin in the middle of a word. Harmless for retrieval, confusing when you read chunks.
The prompt always says PRD.buildPrompt and the system prompt are written for a PRD, so an uploaded test-case file is still introduced as "Context from the PRD".
Labelled rows are not row-safe chunks. The improved LangFlow flow labels every row, then splits at 1000 characters with no overlap on line breaks, so test cases are still cut across chunks (an offline emulation found about 1 in 481 chunks starting at a row). Chunk one row at a time, as Advanced RAG does.
RAG cannot count. "Which modules have the most security test cases?" needs all 5,000 rows; four chunks cannot answer it. The CSV says Goals & Metrics and Integrations, 30 each. Use a query over the data for aggregates.
Generate mode misfires. "write me a summary of test coverage" matches the Generate regex because it contains write and test.
One writer at a time. Embedded Qdrant locks qdrant_data/: do not run app.py and ingest.py together.
Reranker wrapper bug. FlagEmbedding 1.4's reranker crashes with the current tokenizer, so rag_core.py runs bge-reranker-v2-m3 through transformers directly (see learnings/2026-07-12-flagembedding-reranker-bypass.md).
DDrills for the chapter
Demo drills use the live demo on the Page tab (turn on Show locator badges to see every data-testid). Code drills use chapter_07_RAG/.
Count the chunksPlaywright
Automate it: load the page, read rag-chunk-count, then fill rag-size and rag-overlap with 600 / 100 and 300 / 50.
Hint
The count updates on every input event; a web-first toHaveText waits for it.
Expected result
7, then 15, then 28 on this page's copy of the PRD text. RAG Explorer also reports 7 chunks for the PDF at the default. The smaller settings can shift slightly with a different PDF extraction, because whitespace differs: text from pdftotext gives 15 and 29.
Watch the goal answer lose its evidence
Ask What is the goal of this PRD? with top-k 3 at 1200 / 200, then at 300 / 50. Read the ranked list each time.
Expected result
4 of 4 at 1200 / 200: all four goals sit in chunk #1. At 300 / 50 it drops to 1 of 4: the title chunk wins on "PRD", a chunk that mentions "Custom goals" takes a slot, and only "Improve conversion rates" reaches the context.
Keep the ID with its recordPlaywright
Pick the 8 test cases, set 600 / 100 and ask the traffic allocation question. Then switch rag-splitter to blank-line blocks with size 1200.
Hint
Assert on rag-facts and rag-chunk-count.
Expected result
0 of 3 with windows: the Issue Key lines land in different chunks from the traffic allocation text. With one record per chunk: 8 chunks and 3 of 3 (VWO-1130, VWO-1227, VWO-1378).
Find the limit of keyword scoring
At the defaults, ask What are the key features described? Which words matched, and how many feature areas reached the context?
Expected result
Only 2 of 5: Behavioral Insights and Personalization, which share a chunk with "key pages". The ranking keyed on "key" ("key pages", "key user funnels") and on the table header "Feature"; "described" appears nowhere in the PRD. A vague question like this needs meaning, not shared words, which is why the real apps use embeddings and Advanced RAG runs dense and sparse search together.
Read the citation numbers
In groq.js, what does [Chunk 1] refer to: the first chunk of the document or something else?
Expected result
The best-ranked retrieved chunk. buildPrompt numbers chunks by their position in the retrieved list (i + 1), so a citation of [Chunk 2] means the second retrieved chunk.
Which top-k runs?
The README says top 4. You copied .env.example to .env. How many chunks does RAG Explorer retrieve, and why?
Expected result
Three. .env.example sets TOP_K=3, and the server reads Number(process.env.TOP_K || 4), so 4 is only the fallback.
Do the RRF arithmetic
With k = 60, chunk A is 6th in the dense list and 2nd in the sparse list; chunk B is 1st in the dense list only. Which ranks higher after rrf_fuse?
Expected result
A: 1/66 + 1/62 = 0.03128, against 1/61 = 0.01639 for B. The code adds 1.0 / (k + rank + 1) with a 0-based rank.
Catch the Generate-mode false positive
Run detect_mode on "Create a test case for VWO-3400 heatmap privacy masking" and on "write me a summary of test coverage".
Expected result
Both return generate. The second is wrong: the regex only needs a verb such as write followed later by test. Tighten it, for example by requiring "test case" or anchoring the verb at the start.
Assert what the prompt carriesPlaywright
Set rag-topk to 1, click the What are the success metrics? suggestion, and check the prompt and the evidence.
Hint
The suggestion buttons are rag-suggest-1 to rag-suggest-4; success metrics is the fourth.
Expected result
rag-prompt contains [Chunk 1] and no [Chunk 2], and rag-facts reads 5 of 5: the KPI list fits in one chunk at 1200 / 200.
SSolutions: test the demo, then read the pipeline
The Playwright spec drives the demo on this page and passes as written. The other tabs are the exact files from the course repo that the demo ports.
tests/rag-demo.spec.ts
import { test, expect } from'@playwright/test';
const URL = 'https://app.thetestingacademy.com/ai/blueprint/learn/rag.html';
test('chunk size and overlap change how many chunks the PRD becomes', async ({ page }) => {
await page.goto(URL);
awaitexpect(page.getByTestId('rag-chunk-count')).toHaveText('7');
await page.getByTestId('rag-size').fill('600');
await page.getByTestId('rag-overlap').fill('100');
awaitexpect(page.getByTestId('rag-chunk-count')).toHaveText('15');
await page.getByTestId('rag-size').fill('300');
await page.getByTestId('rag-overlap').fill('50');
awaitexpect(page.getByTestId('rag-chunk-count')).toHaveText('28');
});
test('small chunks lose the evidence for the goal question', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('rag-suggest-1').click(); // What is the goal of this PRD?awaitexpect(page.getByTestId('rag-facts')).toHaveText('4 of 4');
awaitexpect(page.getByTestId('rag-hit-1')).toContainText('chunk #1 score');
awaitexpect(page.getByTestId('rag-prompt')).toContainText('Answer using only the context above.');
await page.getByTestId('rag-size').fill('300');
await page.getByTestId('rag-overlap').fill('50');
awaitexpect(page.getByTestId('rag-facts')).toHaveText('1 of 4');
});
test('one record per chunk keeps each test case ID with its content', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('rag-doc').selectOption('cases');
await page.getByTestId('rag-size').fill('600');
await page.getByTestId('rag-overlap').fill('100');
await page.getByTestId('rag-suggest-1').click(); // traffic allocation in A/B Testingawaitexpect(page.getByTestId('rag-facts')).toHaveText('0 of 3');
await page.getByTestId('rag-splitter').selectOption('blocks');
await page.getByTestId('rag-size').fill('1200');
awaitexpect(page.getByTestId('rag-chunk-count')).toHaveText('8');
awaitexpect(page.getByTestId('rag-facts')).toHaveText('3 of 3');
});
Basic_RAG/rag-explorer/server/lib/chunk.js
// Character-based splitter with overlap. Tries to break on paragraph/sentence// boundaries so chunks stay readable instead of cutting mid-word.exportfunctionchunkText(text, { size = 1200, overlap = 200 } = {}) {
const clean = text.replace(/\r\n/g, '\n').replace(/\n{3,}/g, '\n\n').trim()
const chunks = []
let start = 0while (start < clean.length) {
let end = Math.min(start + size, clean.length)
// If we're not at the very end, try to end on a nice boundary within the// last 30% of the window (paragraph break > sentence end > space).if (end < clean.length) {
const window = clean.slice(start, end)
const floor = Math.floor(size * 0.7)
const para = window.lastIndexOf('\n\n')
const sentence = Math.max(window.lastIndexOf('. '), window.lastIndexOf('.\n'))
const space = window.lastIndexOf(' ')
const cut = para > floor ? para : sentence > floor ? sentence + 1 : space > floor ? space : -1if (cut > 0) end = start + cut
}
const piece = clean.slice(start, end).trim()
if (piece) {
chunks.push({
index: chunks.length,
text: piece,
charStart: start,
charEnd: end,
length: piece.length,
})
}
if (end >= clean.length) break
start = Math.max(end - overlap, start + 1)
}
return chunks
}
Basic_RAG/rag-explorer/server/lib/groq.js
// Groq chat completion. "OpenGPT 120B" => openai/gpt-oss-120b, OpenAI-compatible API.const GROQ_URL = 'https://api.groq.com/openai/v1/chat/completions'const GROQ_MODEL = process.env.GROQ_MODEL || 'openai/gpt-oss-120b'const SYSTEM_PROMPT =
'You are a precise assistant answering questions about a Product Requirements Document (PRD). ' +
'Answer ONLY from the provided context chunks. If the answer is not in the context, say so plainly ' +
'("The document does not cover that."). Be concise, cite which chunk numbers you used, and never invent facts.'// Builds the augmented prompt from retrieved chunks. Returned so the UI can show it.exportfunctionbuildPrompt(question, chunks) {
const context = chunks
.map((c, i) => `[Chunk ${i + 1}]\n${c.text}`)
.join('\n\n---\n\n')
return`Context from the PRD:\n\n${context}\n\n---\n\nQuestion: ${question}\n\nAnswer using only the context above.`
}
exportasyncfunctiongenerateAnswer(question, chunks) {
const apiKey = process.env.GROQ_API_KEY
if (!apiKey) thrownewError('GROQ_API_KEY is not set. Add it to .env (see .env.example).')
const userPrompt = buildPrompt(question, chunks)
const res = awaitfetch(GROQ_URL, {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${apiKey}`,
},
body: JSON.stringify({
model: GROQ_MODEL,
temperature: 0.2,
messages: [
{ role: 'system', content: SYSTEM_PROMPT },
{ role: 'user', content: userPrompt },
],
}),
})
if (!res.ok) {
const detail = await res.text().catch(() => '')
thrownewError(`Groq request failed (${res.status}): ${detail || res.statusText}`)
}
const data = await res.json()
const answer = data.choices?.[0]?.message?.content?.trim() || '(no answer returned)'return {
answer,
prompt: userPrompt,
model: GROQ_MODEL,
usage: data.usage || null,
}
}
exportconst groqInfo = { model: GROQ_MODEL, provider: 'groq' }
Advance_RAG/rag_core.py
# ---- fusion + rerank ------------------------------------------------------defrrf_fuse(dense_hits, sparse_hits, k=60, limit=None):
"""Reciprocal Rank Fusion of two ranked lists keyed by point id."""
scores, meta = {}, {}
for rank, h inenumerate(dense_hits):
scores[h["id"]] = scores.get(h["id"], 0.0) + 1.0 / (k + rank + 1)
meta[h["id"]] = h
for rank, h inenumerate(sparse_hits):
scores[h["id"]] = scores.get(h["id"], 0.0) + 1.0 / (k + rank + 1)
meta.setdefault(h["id"], h)
fused = sorted(scores.items(), key=lambda kv: -kv[1])
out = [{"id": pid, "rrf": round(s, 5), "payload": meta[pid]["payload"]} for pid, s in fused]
return out[:limit] if limit else out
defrerank(query, candidates, top_k=4):
ifnot candidates:
return []
tok, model, torch = get_reranker()
pairs = [[query, c["payload"].get("text", "")] for c in candidates]
with torch.no_grad():
inputs = tok(pairs, padding=True, truncation=True, max_length=512, return_tensors="pt")
logits = model(**inputs).logits.view(-1).float()
scores = torch.sigmoid(logits).tolist() # normalize to 0..1 for displayfor c, s inzip(candidates, scores):
c["rerank"] = float(s)
ranked = sorted(candidates, key=lambda c: -c["rerank"])
return ranked[:top_k]