QA questions mix exact identifiers (VWO-2002, NoSuchElementException, build 128) with fuzzy intent ("why is checkout flaky?"). QABuddy.ai answers both from ten sources at once: framework code, 5,000 test cases, Jira, docs, meeting notes, flow exports, PRDs and Jenkins logs. It searches by keyword and by meaning, fuses the two rankings, reranks, refuses when the evidence is weak, and cites every claim.
Two framework repos, test cases, Jira, docs, Figma (Phase 2), meetings, Lucid flows, PRDs, Jenkins logs.
12
Unit tests
Chunkers and loaders on tiny fixtures; all 12 pass offline in about a second.
0.22
Confidence gate
Best rerank score below this and QABuddy says "not in the knowledge base" instead of calling the LLM.
13
Golden questions
A retrieval-only eval: the expected source must appear in the top 6 chunks.
01What QA Buddy is, and why testers need citations
The chapter's capstone turns the chapter 7 hybrid stack into a self-hosted "QA knowledge brain". A tester asks one question ("why is the checkout coupon test flaky?") and gets one answer that cites a meeting note, a Jira ticket and a log, each with an exact reference. An answer you cannot trace is an answer you cannot use in a bug report, so citations are the feature, not a garnish.
Every source is chunked along its natural unit, so a citation points at something a person can open:
Folder
Source
Chunk unit
Citation label
In the demo
01, 02
Selenium and Playwright framework repos
one method or class, with line numbers
repo/path:start-end
no (cloned at setup)
03
5,000 test cases (CSV)
one row
test case VWO-1002
4 rows
04
Jira tickets (JSON)
one ticket with its comments
JIRA VWO-2002
3 tickets
05
Company docs
heading sections, 2048 chars, 300 overlap
sample_qa_process.md
4 sections
06
Figma designs
Phase 2 (designed, not built)
07
Meeting notes
speaker-turn windows, 3600 / 480
meeting sample_sprint42_planning
1
08
Lucid chart exports
doc windows
sample_checkout_flow.txt
1
09
PRD / SRS / BRD / FRD
one PDF page at a time
...VWO.com.pdf p.3
6 pages
10
Jenkins logs
failure blocks plus a build summary
Jenkins build #128
2
One collection, ten sources. Every chunk carries a source_type in its payload. The UI's source checkboxes only add a Qdrant filter, and cross-source answers (a ticket, a log and a meeting note together) come for free.
02One question, one cited answer: the pipeline
QABuddy reuses chapter 7's choices (bge-m3, Qdrant, RRF with k = 60, the transformers reranker, Groq) and spends its new code on what a team tool needs: follow-up questions, several sources, a refusal rule and citations.
flowchart LR
Q[Question and chat history] --> C[condense_question]
C --> M{detect_mode}
C --> RW[rewrite_query x3]
RW --> DS[dense search, top 20]
RW --> SS[sparse search, top 20]
F[source checkboxes] -.->|source_type filter| DS
F -.-> SS
DS --> RRF[rrf_fuse, k 60, keep 12]
SS --> RRF
RRF --> RR[rerank, keep 6]
RR --> G{best score under 0.22?}
G -->|yes| NA[NO_ANSWER, no LLM call]
G -->|no| P[build_messages]
M --> P
P --> L[Groq gpt-oss-120b, streamed]
L --> A[answer that cites every chunk]
One question through app/retrieval.py. The server streams it to the UI as SSE events: meta (mode), citations, then token chunks and done.
Condense turns a follow-up ("and why does it fail on Safari?") into a standalone question using the last 3 turns.
Rewrite asks the LLM for 3 alternate phrasings; each variant runs both searches.
Fuse and rerank: 20 dense and 20 sparse hits per variant, RRF down to 12, a cross-encoder down to 6 (about 2.5 to 3k prompt tokens, according to the plan).
Gate: if the best rerank score is below relevance_threshold (0.22 in config.yaml), the fixed "not in the knowledge base" message streams and no LLM call is made.
03Live demo: rank, fuse, rerank, gate, cite
This runs the retrieval half of the pipeline on a miniature knowledge base: 21 chunks that the chapter's own loaders produced from its sample files (3 Jira tickets, 4 handbook sections, the sprint 42 notes, the checkout flow, 6 PRD pages, the build 128 log and 4 test-case rows). Their dense, sparse and rerank scores were computed once, offline, with the chapter's models (bge-m3 and bge-reranker-v2-m3) for the questions in the list. The page then does exactly what retrieve() does: rank, fuse, rerank, gate and build the prompt. There are no rewrites, as in scripts/eval.py without --rewrites, and no LLM call.
Read the columns left to right. Keyword search wins on exact IDs (VWO-2002, LoginTest); semantic search wins on meaning ("quarantine policy"). The fused list rewards agreement, and the reranker decides what is actually cited. Names in the sample files are replaced with roles on this page.
04Keyword, meaning, and the fusion between them
bge-m3 returns two representations of every chunk in one pass: a dense vector (meaning) and sparse lexical weights (which exact tokens matter). Qdrant stores both as named vectors, and every query searches both. The two score scales cannot be compared, so the lists are merged by rank:
The demo's numbers for "What is JIRA VWO-2002 about?". Keyword search puts the exact ticket first; the dense model prefers its neighbour VWO-2003. Fusion rewards agreement, so VWO-2003 edges ahead until the reranker reads both.
app/core/fusion.py
"""Reciprocal Rank Fusion of dense + sparse result lists."""defrrf_fuse(dense_hits, sparse_hits, k=60, limit=None):
scores, meta = {}, {}
for rank, h inenumerate(dense_hits):
scores[h["id"]] = scores.get(h["id"], 0.0) + 1.0 / (k + rank + 1)
meta[h["id"]] = h
for rank, h inenumerate(sparse_hits):
scores[h["id"]] = scores.get(h["id"], 0.0) + 1.0 / (k + rank + 1)
meta.setdefault(h["id"], h)
fused = sorted(scores.items(), key=lambda kv: -kv[1])
out = [{"id": pid, "rrf": round(s, 5), "payload": meta[pid]["payload"]} for pid, s in fused]
return out[:limit] if limit else out
The fused list is then cut to 12 and handed to the cross-encoder, which reads the question and each chunk together. In the demo, "What failed in Jenkins build 128?" fuses the handbook's bug-triage section to the top (it mentions Jenkins build numbers), and the reranker puts both build 128 chunks first.
app/retrieval.py
defretrieve(question, source_types=None, rewrites=None):
top_n = C.cfg("retrieval.top_n_hybrid", 20)
rrf_k = C.cfg("retrieval.rrf_k", 60)
n_cand = C.cfg("retrieval.rerank_candidates", 12)
top_k = C.cfg("retrieval.top_k", 6)
t0 = time.time()
if rewrites isNone:
rewrites = rewrite_query(question, C.cfg("retrieval.rewrite_n", 3)) \
if C.cfg("retrieval.rewrite_enabled", True) else []
queries = [question] + [r for r in rewrites if r and r != question]
t_rewrite = time.time() - t0
dense_all, sparse_all = {}, {}
for qv in queries:
d, s = embed([qv], batch_size=1)
for h instore().dense_search(d[0], top_n, source_types):
dense_all.setdefault(h["id"], h)
for h instore().sparse_search(s[0], top_n, source_types):
sparse_all.setdefault(h["id"], h)
dense_hits = list(dense_all.values())[:top_n]
sparse_hits = list(sparse_all.values())[:top_n]
t_search = time.time() - t0 - t_rewrite
fused = rrf_fuse(dense_hits, sparse_hits, rrf_k, limit=n_cand)
reranked = rerank(question, [dict(f) for f in fused], top_k)
return {
"candidates": reranked, "fused": fused,
"dense": dense_hits, "sparse": sparse_hits, "rewrites": rewrites,
"timings": {"rewrite": round(t_rewrite, 2), "search": round(t_search, 2),
"rerank": round(time.time() - t0 - t_rewrite - t_search, 2)},
}
05Trust: the gate, the citations and the modes
Three small pieces of code carry the trust story. First the refusal: one fixed sentence, used whenever the best rerank score is under the threshold.
app/retrieval.py
NO_ANSWER = ("This is not in the QABuddy knowledge base yet (best match scored below the ""confidence threshold). Try rephrasing with exact ids (ticket key, test case id, ""file name), widening the source filter, or ingest the missing document.")
app/retrieval.py
best = max((c.get("rerank", 0.0) for c in cands), default=0.0)
threshold = C.cfg("retrieval.relevance_threshold", 0.22)
ifnot cands or best < threshold:
yield {"type": "token", "text": NO_ANSWER}
yield {"type": "done", "answer": NO_ANSWER, "no_answer": True,
"elapsed": round(time.time() - t0, 2)}
return
Then the grounding contract every mode starts with, and the label each chunk carries into the prompt, so [2] resolves to JIRA VWO-2002 or repo/path:12-40:
app/prompts.py
BASE_RULES = (
"You are QABuddy.ai, the internal assistant for the QA team. You answer using ONLY the ""retrieved context chunks below. Every claim must cite its chunk as [n] (e.g. [1], [2]). ""Use exact identifiers from the context (file paths, line numbers, ticket keys, test case ids). ""If the context does not contain the answer, say what is missing instead of guessing."
)
app/prompts.py
defchunk_label(payload):
"""Short human ref used inside the prompt so the model cites naturally."""
st = payload.get("source_type", "")
if st in ("selenium_framework", "playwright_framework"):
loc = f"{payload.get('repo', '')}/{payload.get('path', '')}"if payload.get("line_start"):
loc += f":{payload['line_start']}-{payload['line_end']}"return loc
if st == "test_cases":
returnf"test case {payload.get('tc_id', '?')}"if st == "jira_tickets":
returnf"JIRA {payload.get('ticket_key', '?')}"if st == "jenkins_logs":
b = payload.get("build_id")
returnf"Jenkins build #{b}"if b elsef"Jenkins {payload.get('path', '')}"if st == "meeting_notes":
returnf"meeting {payload.get('title', payload.get('path', ''))}"if st in ("company_docs", "prd_docs", "lucid_charts"):
ref = payload.get("title") or payload.get("path", "")
if payload.get("page"):
ref += f" p.{payload['page']}"return ref
return payload.get("path", st)
Finally the mode router: plain regexes, checked in order, so a question that matches two modes gets the first one. The UI can override the detected mode.
app/retrieval.py
MODE_PATTERNS = [
("generate", re.compile(r"\b(create|generate|write|draft|design|add)\b.{0,60}\b(test ?cases?|tests?|scenarios?|test ?plan)\b", re.I)),
("rca", re.compile(r"\b(root ?cause|rca|why\s+(is|did|does|was).{0,60}(fail|flaky|break)|analy[sz]e.{0,40}(failure|log|build))\b", re.I)),
("review", re.compile(r"\b(review|missing|gaps?|coverage|critique|what.{0,20}not covered)\b", re.I)),
]
defdetect_mode(q):
for mode, rx in MODE_PATTERNS:
if rx.search(q or""):
return mode
return"answer"
The glossary goes into the prompt, not the embeddings.glossary.yaml holds 8 company terms (VWO, RTM, RCA, SDET, TC, POM, flaky test, ATB) that build_messages appends to the system prompt as "Company terminology".
06Ingestion you can re-run every hour
Each source has its own loader and chunker. Every chunk gets a stable id (sha256 of source, path and text, and a Qdrant point id derived from it with uuid5), and every file gets a signature. Re-ingesting compares signatures with a manifest, so only changed files are re-embedded:
flowchart LR
D[data/01 to data/10] --> L[per-source loader and chunker]
L --> U[Doc uid: sha256 of source, path, text]
U --> S[one signature per file]
S --> X{compare with the manifest}
X -->|changed| CH[delete by path, embed, upsert]
X -->|removed| RM[delete by path]
X -->|unchanged| SK[skip]
CH --> Q[(Qdrant: qabuddy_kb)]
RM --> Q
Idempotent ingestion: data/ is the source of truth and Qdrant is a derived index, so a re-run only touches files whose chunks changed.
app/ingestion/pipeline.py
manifest = _load_manifest()
src_man = manifest.get(num, {})
changed = {p: ds for p, ds in by_path.items()
if force or src_man.get(p) != _sig(ds)}
removed = [p for p in src_man if p notin by_path]
unchanged = len(by_path) - len(changed)
report("diff", changed=len(changed), unchanged=unchanged, removed=len(removed))
for p in removed:
store.delete_where(source_type=spec["source_type"], path=p)
for p in changed:
store.delete_where(source_type=spec["source_type"], path=p)
todo = [d for p insorted(changed) for d in changed[p]]
done = 0for b inrange(0, len(todo), batch_size):
batch = todo[b:b + batch_size]
dense, sparse = embed([d.text for d in batch], batch_size=batch_size)
store.upsert_docs(batch, dense, sparse)
done += len(batch)
report("embed", done=done, total=len(todo))
ifnot limit: # a limited demo run must not poison the manifestfor p in removed:
src_man.pop(p, None)
for p, ds in changed.items():
src_man[p] = _sig(ds)
manifest[num] = src_man
_save_manifest(manifest)
The CLI prints the same diff: [05] changed=1 unchanged=0 removed=0. A --limit run never writes the manifest, so a quick demo ingest cannot poison the next real one. This is what makes the Phase 2 plan (an hourly cron, and a QABuddy MCP server exposing qa_search to IDE copilots, see chapters 9 and 10) a scheduling job rather than new code.
07Set up, run, test, evaluate, deploy
Python 3.11+ (the quickstart uses uv with 3.13; the Dockerfile uses 3.12). The models download on first use (about 4.6 GB for bge-m3 and the reranker, per the deploy notes). Only chat needs a key: ingest, tests and the eval run without one.
terminal
cd chapter_08_QABuddyAI
uv venv .venv --python 3.13 && uv pip install -p .venv/bin/python -r requirements.txt
cp .env.example .env # set GROQ_API_KEY (chat only)
./scripts/fetch_repos.sh # clone the two framework repos
./scripts/setup_fixtures.sh # test-case CSV and PRD from chapter 7
.venv/bin/python -m app.ingestion.cli ingest --all
.venv/bin/python -m pytest tests -q # 12 unit tests
.venv/bin/python scripts/eval.py # golden-question hit rate
./scripts/dev.sh # UI on http://127.0.0.1:5080
Piece
What it does
Command
scripts/fetch_repos.sh
Clones (or fast-forwards) the two framework repos into data/01 and data/02
./scripts/fetch_repos.sh
scripts/setup_fixtures.sh
Copies the chapter 7 test-case CSV and the VWO PRD into data/03 and data/09
./scripts/setup_fixtures.sh
Ingestion CLI
Loads, diffs against the manifest, embeds and upserts; stats and reset too
12 tests on chunkers and loaders with tiny fixtures; no models or keys
.venv/bin/python -m pytest tests -q
Retrieval eval
13 golden questions; a HIT needs the expected source in the top 6; exits 1 under 70%
.venv/bin/python scripts/eval.py
Chat server
Flask UI and API on port 5080: /api/chat (SSE), /api/search, /api/ingest, health and stats
./scripts/dev.sh
scripts/jira_fetch.py
Pulls tickets with REST and JQL into data/04 in the QABuddy schema
.venv/bin/python scripts/jira_fetch.py
scripts/backup.sh
Nightly Qdrant snapshot (server mode) or tar of qdrant_data, plus the manifest
./scripts/backup.sh
Docker stack
Qdrant server, the app on gunicorn, Caddy with TLS and basic auth
docker compose up -d --build
Environment variables (names only): GROQ_API_KEY (or LLM_API_KEY), LLM_BASE_URL, LLM_MODEL, EMBED_MODEL, RERANK_MODEL, BGE_USE_FP16, QDRANT_PATH or QDRANT_URL, QDRANT_COLLECTION, PORT, and for Jira pulls JIRA_BASE_URL, JIRA_EMAIL, JIRA_API_TOKEN, JIRA_JQL. Tunables such as chunk sizes, top_k and the threshold live in config.yaml.
README figures are claims, not measurements on this page: a 13/13 golden hit rate on the seeded 5,531-chunk corpus, about $0.001 per question on Groq, and about $55 to $75 a month for an 8 GB droplet. Re-measure them on your own data.
08What to watch
The eval does not test the gate.scripts/eval.py counts a HIT when the right source is in the top 6, and only prints the best rerank score. In the demo, "What is JIRA VWO-2002 about?" retrieves the ticket but the best score is 0.188, so the gate answers "not in the knowledge base". Assert on the answer, not only on retrieval.
One process owns embedded Qdrant. While dev.sh runs, ingest from the UI panel, not the CLI, or run Qdrant as a server (QDRANT_URL).
A mislabelled failure. The build 128 failure chunk is labelled failure: Test, not LoginTest.testInvalidPassword: the test-name regex takes the first match in the block, and the context lines include [Pipeline] stage (Test).
Old logs are skipped. Jenkins logs older than retention_days (90) by file modification time are not ingested.
The plan and the code differ, on purpose.Plan.md names tree-sitter for code chunking; the build chose a small signature and brace-depth parser with a line-window fallback instead (no native dependencies), as the course notes record.
Streaming through a proxy. The Caddyfile sets flush_interval -1 on the reverse proxy so streamed tokens reach the browser as they arrive.
One gunicorn worker, eight threads (-w 1 --threads 8): every extra worker process would load its own copy of the models.
Unpinned image.docker-compose.yml uses qdrant/qdrant:latest; pin it before you depend on it.
DDrills for the chapter
Playwright drills target the live demo on the Page tab: turn on Show locator badges to see every data-testid. The other drills use the chapter's code and CLI.
Keyword vs meaning on an IDPlaywright
Select "What is JIRA VWO-2002 about?" in qb-question and assert the first item of qb-sparse and of qb-dense.
Hint
Use getByRole('listitem').first() inside each list.
Expected result
Keyword: JIRA VWO-2002 (sparse 0.1757). Semantic: JIRA VWO-2003 (dense 0.4197), the ticket that reads most alike. The answer box shows the "not in the knowledge base" message.
Move the gatePlaywright
With the same question, change qb-threshold from 0.22 to 0.15.
Expected result
qb-status flips from Gated to Answerable, and the prompt starts its context with [1] (meeting sample_sprint42_planning), the best reranked chunk at 0.188.
Filter a sourcePlaywright
Ask "Why is the checkout coupon test flaky?", then uncheck qb-src-meeting_notes.
Expected result
The mode is rca (detected). The first citation changes from meeting sample_sprint42_planning (0.923) to JIRA VWO-2002 (0.864), and the gate still passes.
Why rerank at all?
Ask "What failed in Jenkins build 128?". Compare the top of the fused list with the first citation.
Expected result
Fused #1 is the handbook's Bug triage section (it mentions Jenkins build numbers). The reranker puts the two Jenkins build #128 chunks first (0.931 and 0.910) and drops Bug triage to 0.010.
Do the fusion by hand
At k = 60, VWO-2003 is 1st in the semantic list and 2nd in the keyword list; VWO-2002 is 4th and 1st. Compute both RRF scores.
Expected result
VWO-2003: 1/61 + 1/62 = 0.03252. VWO-2002: 1/64 + 1/61 = 0.03202. Agreement beats a single first place.
Route the question
Which mode does detect_mode pick for: (a) "Why is the checkout coupon test flaky?" (b) "What failed in Jenkins build 128?" (c) "Review coverage gaps for the payment module" (d) "Design a test plan and review gaps"?
Expected result
(a) rca, (b) answer, (c) review, (d) generate: the patterns are checked in order, so generate wins over review.
Fix the failure label
Run the loaders on the sample build 128 log and read the failure chunk's first line. Why does it say failure: Test, and why does test_split_log_failures still pass?
Expected result
TEST_RE takes the first match in the block, and the 12 context lines include [Pipeline] stage (Test). The unit-test fixture has no such line. Match the test name on the trigger line first, then fall back to the block.
Re-ingest one file
After a full ingest, fix a typo in data/05_company_docs/sample_qa_process.md and run ingest --source 05.
Expected result
[05] changed=1 unchanged=0 removed=0: the file's signature changed, so its 4 heading chunks are deleted by path and re-embedded. Run it with --limit and the manifest is left alone.
Hit, or answer?
Run scripts/eval.py and look at the best column. Which golden questions could pass as HIT and still be refused by the 0.22 gate?
Expected result
Any question whose best score is under 0.22. In the demo, "What is JIRA VWO-2002 about?" retrieves the ticket (a HIT by the eval's rule) but tops out at 0.188, so the answer is the refusal. Scores on your full knowledge base can differ, so read the column.
SSolutions: test the demo, then read the pipeline
The Playwright spec drives the demo on this page and passes as written. The other tabs are exact files from the course repo.
tests/qa-buddy-demo.spec.ts
import { test, expect } from'@playwright/test';
const URL = 'https://app.thetestingacademy.com/ai/blueprint/learn/qa-buddy.html';
test('keyword search finds the ticket ID, semantic search prefers its neighbour', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('qb-question').selectOption({ label: 'What is JIRA VWO-2002 about?' });
awaitexpect(page.getByTestId('qb-sparse').getByRole('listitem').first()).toContainText('JIRA VWO-2002');
awaitexpect(page.getByTestId('qb-dense').getByRole('listitem').first()).toContainText('JIRA VWO-2003');
awaitexpect(page.getByTestId('qb-answer')).toContainText('This is not in the QABuddy knowledge base yet');
});
test('the confidence gate is a threshold you can move', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('qb-question').selectOption({ label: 'What is JIRA VWO-2002 about?' });
awaitexpect(page.getByTestId('qb-status')).toContainText('Gated');
await page.getByTestId('qb-threshold').fill('0.15');
awaitexpect(page.getByTestId('qb-status')).toContainText('Answerable');
awaitexpect(page.getByTestId('qb-prompt')).toContainText('[1] (meeting sample_sprint42_planning)');
});
test('the source filter changes what gets cited', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('qb-question').selectOption({ label: 'Why is the checkout coupon test flaky?' });
awaitexpect(page.getByTestId('qb-mode-badge')).toHaveText('rca (detected)');
awaitexpect(page.getByTestId('qb-final').getByRole('listitem').first()).toContainText('meeting sample_sprint42_planning');
await page.getByTestId('qb-src-meeting_notes').uncheck();
awaitexpect(page.getByTestId('qb-final').getByRole('listitem').first()).toContainText('JIRA VWO-2002');
});
app/retrieval.py
"""Ask pipeline: condense -> rewrite -> hybrid search -> RRF -> rerank -> cited answer."""import re
import time
from . import config as C
from . import llm, prompts
from .core.embedder import embed
from .core.fusion import rrf_fuse
from .core.reranker import rerank
from .core.store import Store
_store = Nonedefstore():
global _store
if _store isNone:
_store = Store()
return _store
# ---- modes -----------------------------------------------------------------
MODE_PATTERNS = [
("generate", re.compile(r"\b(create|generate|write|draft|design|add)\b.{0,60}\b(test ?cases?|tests?|scenarios?|test ?plan)\b", re.I)),
("rca", re.compile(r"\b(root ?cause|rca|why\s+(is|did|does|was).{0,60}(fail|flaky|break)|analy[sz]e.{0,40}(failure|log|build))\b", re.I)),
("review", re.compile(r"\b(review|missing|gaps?|coverage|critique|what.{0,20}not covered)\b", re.I)),
]
defdetect_mode(q):
for mode, rx in MODE_PATTERNS:
if rx.search(q or""):
return mode
return"answer"# ---- query shaping ---------------------------------------------------------defcondense_question(question, history):
"""Rewrite a follow-up into a standalone question using recent turns."""ifnot history:
return question
turns = history[-2 * C.cfg("retrieval.history_turns", 3):]
convo = "\n".join(f"{h.get('role')}: {str(h.get('content'))[:400]}"for h in turns)
try:
out = llm.chat([
{"role": "system", "content":
"Rewrite the user's last message as ONE standalone search question, resolving ""pronouns from the conversation. Return only the question."},
{"role": "user", "content": f"Conversation:\n{convo}\n\nLast message: {question}"},
], temperature=0.0, max_tokens=120)
return out.strip() or question
except Exception:
return question
defrewrite_query(query, n=3):
try:
out = llm.chat([
{"role": "system", "content":
"You rewrite a search query into alternate phrasings for retrieval. "f"Return exactly {n} rewrites, one per line, no numbering, no extra text."},
{"role": "user", "content": query},
], temperature=0.4, max_tokens=200)
lines = [re.sub(r"^\s*[-*\d.]+\s*", "", l).strip() for l in out.splitlines() if l.strip()]
return lines[:n] if lines else []
except Exception:
return []
# ---- retrieval -------------------------------------------------------------defretrieve(question, source_types=None, rewrites=None):
top_n = C.cfg("retrieval.top_n_hybrid", 20)
rrf_k = C.cfg("retrieval.rrf_k", 60)
n_cand = C.cfg("retrieval.rerank_candidates", 12)
top_k = C.cfg("retrieval.top_k", 6)
t0 = time.time()
if rewrites isNone:
rewrites = rewrite_query(question, C.cfg("retrieval.rewrite_n", 3)) \
if C.cfg("retrieval.rewrite_enabled", True) else []
queries = [question] + [r for r in rewrites if r and r != question]
t_rewrite = time.time() - t0
dense_all, sparse_all = {}, {}
for qv in queries:
d, s = embed([qv], batch_size=1)
for h instore().dense_search(d[0], top_n, source_types):
dense_all.setdefault(h["id"], h)
for h instore().sparse_search(s[0], top_n, source_types):
sparse_all.setdefault(h["id"], h)
dense_hits = list(dense_all.values())[:top_n]
sparse_hits = list(sparse_all.values())[:top_n]
t_search = time.time() - t0 - t_rewrite
fused = rrf_fuse(dense_hits, sparse_hits, rrf_k, limit=n_cand)
reranked = rerank(question, [dict(f) for f in fused], top_k)
return {
"candidates": reranked, "fused": fused,
"dense": dense_hits, "sparse": sparse_hits, "rewrites": rewrites,
"timings": {"rewrite": round(t_rewrite, 2), "search": round(t_search, 2),
"rerank": round(time.time() - t0 - t_rewrite - t_search, 2)},
}
# ---- citations -------------------------------------------------------------
SOURCE_LABELS = {
"selenium_framework": "Selenium repo", "playwright_framework": "Playwright repo",
"test_cases": "Test case", "jira_tickets": "JIRA", "company_docs": "Doc",
"meeting_notes": "Meeting", "lucid_charts": "Lucid", "prd_docs": "PRD",
"jenkins_logs": "Jenkins", "figma_designs": "Figma",
}
defbuild_citation(n, cand):
p = cand["payload"]
return {
"n": n,
"source_type": p.get("source_type"),
"label": SOURCE_LABELS.get(p.get("source_type"), p.get("source_type")),
"ref": prompts.chunk_label(p),
"path": p.get("path"),
"url": p.get("url"),
"rerank": round(cand.get("rerank", 0.0), 3),
"rrf": cand.get("rrf"),
"snippet": (p.get("text") or"")[:500],
}
# ---- ask (streaming event generator) --------------------------------------
NO_ANSWER = ("This is not in the QABuddy knowledge base yet (best match scored below the ""confidence threshold). Try rephrasing with exact ids (ticket key, test case id, ""file name), widening the source filter, or ingest the missing document.")
defask_events(question, sources=None, mode=None, history=None):
"""Yields: meta -> citations -> token* -> done | error."""
t0 = time.time()
try:
q = condense_question(question, history)
mode = mode if mode in prompts.MODE_SYS elsedetect_mode(q)
yield {"type": "meta", "mode": mode, "question": q}
r = retrieve(q, source_types=sources orNone)
cands = r["candidates"]
citations = [build_citation(i + 1, c) for i, c inenumerate(cands)]
yield {"type": "citations", "items": citations, "rewrites": r["rewrites"],
"timings": r["timings"]}
best = max((c.get("rerank", 0.0) for c in cands), default=0.0)
threshold = C.cfg("retrieval.relevance_threshold", 0.22)
ifnot cands or best < threshold:
yield {"type": "token", "text": NO_ANSWER}
yield {"type": "done", "answer": NO_ANSWER, "no_answer": True,
"elapsed": round(time.time() - t0, 2)}
return
messages = prompts.build_messages(mode, q, cands, history)
answer = ""try:
for delta in llm.chat_stream(messages):
answer += delta
yield {"type": "token", "text": delta}
except Exception:
answer = llm.chat(messages) # streaming unsupported -> one shotyield {"type": "token", "text": answer}
yield {"type": "done", "answer": answer, "elapsed": round(time.time() - t0, 2)}
except Exception as e:
yield {"type": "error", "message": str(e)}
defask(question, sources=None, mode=None, history=None):
"""Non-streaming wrapper: returns {answer, citations, mode, ...}."""
out = {"question": question, "citations": [], "answer": "", "mode": None}
for ev inask_events(question, sources=sources, mode=mode, history=history):
if ev["type"] == "meta":
out["mode"] = ev["mode"]
elif ev["type"] == "citations":
out["citations"] = ev["items"]
out["rewrites"] = ev.get("rewrites")
out["timings"] = ev.get("timings")
elif ev["type"] == "done":
out["answer"] = ev["answer"]
out["elapsed"] = ev["elapsed"]
out["no_answer"] = ev.get("no_answer", False)
elif ev["type"] == "error":
out["error"] = ev["message"]
return out
app/prompts.py
"""Prompt templates per mode + glossary injection + citation labeling."""from . import config as C
BASE_RULES = (
"You are QABuddy.ai, the internal assistant for the QA team. You answer using ONLY the ""retrieved context chunks below. Every claim must cite its chunk as [n] (e.g. [1], [2]). ""Use exact identifiers from the context (file paths, line numbers, ticket keys, test case ids). ""If the context does not contain the answer, say what is missing instead of guessing."
)
MODE_SYS = {
"answer": "Answer concisely and practically, like a senior QA engineer helping a teammate.",
"generate": (
"You are a senior SDET. Using the retrieved test cases, requirements and code as ""templates, produce well-structured NEW test case(s) for the request. Format each with ""exactly these headers: Title, Preconditions, Steps (numbered), Expected Result, ""Priority, Tags. Ground every design choice in the retrieved chunks and cite them [n]."),
"review": (
"You are reviewing test coverage. Compare the retrieved test cases against the retrieved ""requirements/PRD sections and code. List: covered areas, GAPS (missing test cases with ""a one-line suggested case each), and risky assumptions. Cite evidence [n] for every gap."),
"rca": (
"You are doing root cause analysis of a test failure. From the retrieved logs, tickets, ""code and test cases: state the most probable root cause, the evidence chain (cite [n] ""for each step), whether it looks like a product bug / test bug / flaky infra, and the ""concrete next fix. If evidence is thin, say what log or data is missing."),
}
def_glossary_block():
ifnot C.GLOSSARY:
return""
lines = "\n".join(f"- {k}: {v}"for k, v in C.GLOSSARY.items())
returnf"\n\nCompany terminology:\n{lines}"defchunk_label(payload):
"""Short human ref used inside the prompt so the model cites naturally."""
st = payload.get("source_type", "")
if st in ("selenium_framework", "playwright_framework"):
loc = f"{payload.get('repo', '')}/{payload.get('path', '')}"if payload.get("line_start"):
loc += f":{payload['line_start']}-{payload['line_end']}"return loc
if st == "test_cases":
returnf"test case {payload.get('tc_id', '?')}"if st == "jira_tickets":
returnf"JIRA {payload.get('ticket_key', '?')}"if st == "jenkins_logs":
b = payload.get("build_id")
returnf"Jenkins build #{b}"if b elsef"Jenkins {payload.get('path', '')}"if st == "meeting_notes":
returnf"meeting {payload.get('title', payload.get('path', ''))}"if st in ("company_docs", "prd_docs", "lucid_charts"):
ref = payload.get("title") or payload.get("path", "")
if payload.get("page"):
ref += f" p.{payload['page']}"return ref
return payload.get("path", st)
defbuild_messages(mode, question, chunks, history=None):
context = "\n\n---\n\n".join(
f"[{i + 1}] ({chunk_label(c['payload'])})\n{c['payload'].get('text', '')}"for i, c inenumerate(chunks))
system = BASE_RULES + "\n\n" + MODE_SYS.get(mode, MODE_SYS["answer"]) + _glossary_block()
messages = [{"role": "system", "content": system}]
for h in (history or [])[-2 * C.cfg("retrieval.history_turns", 3):]:
if h.get("role") in ("user", "assistant") and h.get("content"):
messages.append({"role": h["role"], "content": str(h["content"])[:600]})
messages.append({"role": "user",
"content": f"Context:\n{context}\n\n---\n\nRequest: {question}"})
return messages
config.yaml
# QABuddy.ai tunables. Sizes are characters (~4 chars per token).chunking:
code:
max_chars: 2000# one method/class unit; bigger units get split min_chars: 250# tiny units merge into neighbors fallback_size: 1600# whole-file split when no units detected fallback_overlap: 200 testcases:
size: 1400# 1 row = 1 chunk; only absurdly long rows spill overlap: 0 jira:
size: 4800# 1 ticket = 1 chunk; long comment threads spill overlap: 0 docs: # PDFs, MD, company docs, PRD/SRS, lucid exports size: 2048# ~512 tokens overlap: 300# ~15% transcripts:
size: 3600# ~900 tokens overlap: 480 logs:
max_block_chars: 2400# one failure block (stack trace + context) context_lines: 12 retention_days: 90retrieval:
top_n_hybrid: 20# candidates per dense / sparse search rerank_candidates: 12# fused list fed to cross-encoder top_k: 6# final chunks sent to the LLM rrf_k: 60 rewrite_enabled: true
rewrite_n: 3 relevance_threshold: 0.22# max rerank score below this -> "not in KB" history_turns: 3llm:
temperature: 0.2 max_tokens: 1400