AI Tester Blueprint RAG and MCP QA Buddy copilot
Chapter 8
Chapter 8 . RAG and MCP . QA Buddy copilot

QA Buddy: one question, one cited answer

QA questions mix exact identifiers (VWO-2002, NoSuchElementException, build 128) with fuzzy intent ("why is checkout flaky?"). QABuddy.ai answers both from ten sources at once: framework code, 5,000 test cases, Jira, docs, meeting notes, flow exports, PRDs and Jenkins logs. It searches by keyword and by meaning, fuses the two rankings, reranks, refuses when the evidence is weak, and cites every claim.

10
Knowledge sources
Two framework repos, test cases, Jira, docs, Figma (Phase 2), meetings, Lucid flows, PRDs, Jenkins logs.
12
Unit tests
Chunkers and loaders on tiny fixtures; all 12 pass offline in about a second.
0.22
Confidence gate
Best rerank score below this and QABuddy says "not in the knowledge base" instead of calling the LLM.
13
Golden questions
A retrieval-only eval: the expected source must appear in the top 6 chunks.

01What QA Buddy is, and why testers need citations

The chapter's capstone turns the chapter 7 hybrid stack into a self-hosted "QA knowledge brain". A tester asks one question ("why is the checkout coupon test flaky?") and gets one answer that cites a meeting note, a Jira ticket and a log, each with an exact reference. An answer you cannot trace is an answer you cannot use in a bug report, so citations are the feature, not a garnish.

Every source is chunked along its natural unit, so a citation points at something a person can open:

FolderSourceChunk unitCitation labelIn the demo
01, 02Selenium and Playwright framework reposone method or class, with line numbersrepo/path:start-endno (cloned at setup)
035,000 test cases (CSV)one rowtest case VWO-10024 rows
04Jira tickets (JSON)one ticket with its commentsJIRA VWO-20023 tickets
05Company docsheading sections, 2048 chars, 300 overlapsample_qa_process.md4 sections
06Figma designsPhase 2 (designed, not built)
07Meeting notesspeaker-turn windows, 3600 / 480meeting sample_sprint42_planning1
08Lucid chart exportsdoc windowssample_checkout_flow.txt1
09PRD / SRS / BRD / FRDone PDF page at a time...VWO.com.pdf p.36 pages
10Jenkins logsfailure blocks plus a build summaryJenkins build #1282
One collection, ten sources. Every chunk carries a source_type in its payload. The UI's source checkboxes only add a Qdrant filter, and cross-source answers (a ticket, a log and a meeting note together) come for free.

02One question, one cited answer: the pipeline

QABuddy reuses chapter 7's choices (bge-m3, Qdrant, RRF with k = 60, the transformers reranker, Groq) and spends its new code on what a team tool needs: follow-up questions, several sources, a refusal rule and citations.

flowchart LR
  Q[Question and chat history] --> C[condense_question]
  C --> M{detect_mode}
  C --> RW[rewrite_query x3]
  RW --> DS[dense search, top 20]
  RW --> SS[sparse search, top 20]
  F[source checkboxes] -.->|source_type filter| DS
  F -.-> SS
  DS --> RRF[rrf_fuse, k 60, keep 12]
  SS --> RRF
  RRF --> RR[rerank, keep 6]
  RR --> G{best score under 0.22?}
  G -->|yes| NA[NO_ANSWER, no LLM call]
  G -->|no| P[build_messages]
  M --> P
  P --> L[Groq gpt-oss-120b, streamed]
  L --> A[answer that cites every chunk]
One question through app/retrieval.py. The server streams it to the UI as SSE events: meta (mode), citations, then token chunks and done.
  • Condense turns a follow-up ("and why does it fail on Safari?") into a standalone question using the last 3 turns.
  • Rewrite asks the LLM for 3 alternate phrasings; each variant runs both searches.
  • Fuse and rerank: 20 dense and 20 sparse hits per variant, RRF down to 12, a cross-encoder down to 6 (about 2.5 to 3k prompt tokens, according to the plan).
  • Gate: if the best rerank score is below relevance_threshold (0.22 in config.yaml), the fixed "not in the knowledge base" message streams and no LLM call is made.

03Live demo: rank, fuse, rerank, gate, cite

This runs the retrieval half of the pipeline on a miniature knowledge base: 21 chunks that the chapter's own loaders produced from its sample files (3 Jira tickets, 4 handbook sections, the sprint 42 notes, the checkout flow, 6 PRD pages, the build 128 log and 4 test-case rows). Their dense, sparse and rerank scores were computed once, offline, with the chapter's models (bge-m3 and bge-reranker-v2-m3) for the questions in the list. The page then does exactly what retrieve() does: rank, fuse, rerank, gate and build the prompt. There are no rewrites, as in scripts/eval.py without --rewrites, and no LLM call.

Ask QA Buddy offline: rank, fuse, rerank, gate, citeNo API key needed
Sources (the source_type filter)
data-testid=qb-questiondata-testid=qb-modedata-testid=qb-kdata-testid=qb-thresholddata-testid=qb-src-meeting_notesdata-testid=qb-rundata-testid=qb-sparsedata-testid=qb-densedata-testid=qb-fuseddata-testid=qb-finaldata-testid=qb-answer
Mode:
1. Keyword (sparse)
    2. Semantic (dense)
      3. Fused with RRF (top 12)
        4. Reranked: the 6 chunks the answer would cite
          5. The prompt (build_messages)
          
              
          Read the columns left to right. Keyword search wins on exact IDs (VWO-2002, LoginTest); semantic search wins on meaning ("quarantine policy"). The fused list rewards agreement, and the reranker decides what is actually cited. Names in the sample files are replaced with roles on this page.

          04Keyword, meaning, and the fusion between them

          bge-m3 returns two representations of every chunk in one pass: a dense vector (meaning) and sparse lexical weights (which exact tokens matter). Qdrant stores both as named vectors, and every query searches both. The two score scales cannot be compared, so the lists are merged by rank:

          Keyword (sparse) ranking Semantic (dense) ranking RRF, k = 60 1 JIRA VWO-2002 2 JIRA VWO-2003 3 PRD p.1 4 sprint 42 notes 1 JIRA VWO-2003 2 PRD p.1 3 sprint 42 notes 4 JIRA VWO-2002 VWO-2003 .03252 VWO-2002 .03202 PRD p.1 .03200 notes .03150 VWO-2003: 1/(60+1) + 1/(60+2) = 0.03252 VWO-2002: 1/(60+1) + 1/(60+4) = 0.03202 Ranks only, never raw scores: a chunk both searches agree on beats one that a single search ranked first.
          The demo's numbers for "What is JIRA VWO-2002 about?". Keyword search puts the exact ticket first; the dense model prefers its neighbour VWO-2003. Fusion rewards agreement, so VWO-2003 edges ahead until the reranker reads both.
          app/core/fusion.py
          """Reciprocal Rank Fusion of dense + sparse result lists."""
          
          
          def rrf_fuse(dense_hits, sparse_hits, k=60, limit=None):
              scores, meta = {}, {}
              for rank, h in enumerate(dense_hits):
                  scores[h["id"]] = scores.get(h["id"], 0.0) + 1.0 / (k + rank + 1)
                  meta[h["id"]] = h
              for rank, h in enumerate(sparse_hits):
                  scores[h["id"]] = scores.get(h["id"], 0.0) + 1.0 / (k + rank + 1)
                  meta.setdefault(h["id"], h)
              fused = sorted(scores.items(), key=lambda kv: -kv[1])
              out = [{"id": pid, "rrf": round(s, 5), "payload": meta[pid]["payload"]} for pid, s in fused]
              return out[:limit] if limit else out

          The fused list is then cut to 12 and handed to the cross-encoder, which reads the question and each chunk together. In the demo, "What failed in Jenkins build 128?" fuses the handbook's bug-triage section to the top (it mentions Jenkins build numbers), and the reranker puts both build 128 chunks first.

          app/retrieval.py
          def retrieve(question, source_types=None, rewrites=None):
              top_n = C.cfg("retrieval.top_n_hybrid", 20)
              rrf_k = C.cfg("retrieval.rrf_k", 60)
              n_cand = C.cfg("retrieval.rerank_candidates", 12)
              top_k = C.cfg("retrieval.top_k", 6)
          
              t0 = time.time()
              if rewrites is None:
                  rewrites = rewrite_query(question, C.cfg("retrieval.rewrite_n", 3)) \
                      if C.cfg("retrieval.rewrite_enabled", True) else []
              queries = [question] + [r for r in rewrites if r and r != question]
              t_rewrite = time.time() - t0
          
              dense_all, sparse_all = {}, {}
              for qv in queries:
                  d, s = embed([qv], batch_size=1)
                  for h in store().dense_search(d[0], top_n, source_types):
                      dense_all.setdefault(h["id"], h)
                  for h in store().sparse_search(s[0], top_n, source_types):
                      sparse_all.setdefault(h["id"], h)
              dense_hits = list(dense_all.values())[:top_n]
              sparse_hits = list(sparse_all.values())[:top_n]
              t_search = time.time() - t0 - t_rewrite
          
              fused = rrf_fuse(dense_hits, sparse_hits, rrf_k, limit=n_cand)
              reranked = rerank(question, [dict(f) for f in fused], top_k)
              return {
                  "candidates": reranked, "fused": fused,
                  "dense": dense_hits, "sparse": sparse_hits, "rewrites": rewrites,
                  "timings": {"rewrite": round(t_rewrite, 2), "search": round(t_search, 2),
                              "rerank": round(time.time() - t0 - t_rewrite - t_search, 2)},
              }

          05Trust: the gate, the citations and the modes

          Three small pieces of code carry the trust story. First the refusal: one fixed sentence, used whenever the best rerank score is under the threshold.

          app/retrieval.py
          NO_ANSWER = ("This is not in the QABuddy knowledge base yet (best match scored below the "
                       "confidence threshold). Try rephrasing with exact ids (ticket key, test case id, "
                       "file name), widening the source filter, or ingest the missing document.")
          app/retrieval.py
                  best = max((c.get("rerank", 0.0) for c in cands), default=0.0)
                  threshold = C.cfg("retrieval.relevance_threshold", 0.22)
                  if not cands or best < threshold:
                      yield {"type": "token", "text": NO_ANSWER}
                      yield {"type": "done", "answer": NO_ANSWER, "no_answer": True,
                             "elapsed": round(time.time() - t0, 2)}
                      return

          Then the grounding contract every mode starts with, and the label each chunk carries into the prompt, so [2] resolves to JIRA VWO-2002 or repo/path:12-40:

          app/prompts.py
          BASE_RULES = (
              "You are QABuddy.ai, the internal assistant for the QA team. You answer using ONLY the "
              "retrieved context chunks below. Every claim must cite its chunk as [n] (e.g. [1], [2]). "
              "Use exact identifiers from the context (file paths, line numbers, ticket keys, test case ids). "
              "If the context does not contain the answer, say what is missing instead of guessing."
          )
          app/prompts.py
          def chunk_label(payload):
              """Short human ref used inside the prompt so the model cites naturally."""
              st = payload.get("source_type", "")
              if st in ("selenium_framework", "playwright_framework"):
                  loc = f"{payload.get('repo', '')}/{payload.get('path', '')}"
                  if payload.get("line_start"):
                      loc += f":{payload['line_start']}-{payload['line_end']}"
                  return loc
              if st == "test_cases":
                  return f"test case {payload.get('tc_id', '?')}"
              if st == "jira_tickets":
                  return f"JIRA {payload.get('ticket_key', '?')}"
              if st == "jenkins_logs":
                  b = payload.get("build_id")
                  return f"Jenkins build #{b}" if b else f"Jenkins {payload.get('path', '')}"
              if st == "meeting_notes":
                  return f"meeting {payload.get('title', payload.get('path', ''))}"
              if st in ("company_docs", "prd_docs", "lucid_charts"):
                  ref = payload.get("title") or payload.get("path", "")
                  if payload.get("page"):
                      ref += f" p.{payload['page']}"
                  return ref
              return payload.get("path", st)

          Finally the mode router: plain regexes, checked in order, so a question that matches two modes gets the first one. The UI can override the detected mode.

          app/retrieval.py
          MODE_PATTERNS = [
              ("generate", re.compile(r"\b(create|generate|write|draft|design|add)\b.{0,60}\b(test ?cases?|tests?|scenarios?|test ?plan)\b", re.I)),
              ("rca", re.compile(r"\b(root ?cause|rca|why\s+(is|did|does|was).{0,60}(fail|flaky|break)|analy[sz]e.{0,40}(failure|log|build))\b", re.I)),
              ("review", re.compile(r"\b(review|missing|gaps?|coverage|critique|what.{0,20}not covered)\b", re.I)),
          ]
          
          
          def detect_mode(q):
              for mode, rx in MODE_PATTERNS:
                  if rx.search(q or ""):
                      return mode
              return "answer"
          The glossary goes into the prompt, not the embeddings. glossary.yaml holds 8 company terms (VWO, RTM, RCA, SDET, TC, POM, flaky test, ATB) that build_messages appends to the system prompt as "Company terminology".

          06Ingestion you can re-run every hour

          Each source has its own loader and chunker. Every chunk gets a stable id (sha256 of source, path and text, and a Qdrant point id derived from it with uuid5), and every file gets a signature. Re-ingesting compares signatures with a manifest, so only changed files are re-embedded:

          flowchart LR
            D[data/01 to data/10] --> L[per-source loader and chunker]
            L --> U[Doc uid: sha256 of source, path, text]
            U --> S[one signature per file]
            S --> X{compare with the manifest}
            X -->|changed| CH[delete by path, embed, upsert]
            X -->|removed| RM[delete by path]
            X -->|unchanged| SK[skip]
            CH --> Q[(Qdrant: qabuddy_kb)]
            RM --> Q
          Idempotent ingestion: data/ is the source of truth and Qdrant is a derived index, so a re-run only touches files whose chunks changed.
          app/ingestion/pipeline.py
              manifest = _load_manifest()
              src_man = manifest.get(num, {})
              changed = {p: ds for p, ds in by_path.items()
                         if force or src_man.get(p) != _sig(ds)}
              removed = [p for p in src_man if p not in by_path]
              unchanged = len(by_path) - len(changed)
              report("diff", changed=len(changed), unchanged=unchanged, removed=len(removed))
          
              for p in removed:
                  store.delete_where(source_type=spec["source_type"], path=p)
              for p in changed:
                  store.delete_where(source_type=spec["source_type"], path=p)
          
              todo = [d for p in sorted(changed) for d in changed[p]]
              done = 0
              for b in range(0, len(todo), batch_size):
                  batch = todo[b:b + batch_size]
                  dense, sparse = embed([d.text for d in batch], batch_size=batch_size)
                  store.upsert_docs(batch, dense, sparse)
                  done += len(batch)
                  report("embed", done=done, total=len(todo))
          
              if not limit:  # a limited demo run must not poison the manifest
                  for p in removed:
                      src_man.pop(p, None)
                  for p, ds in changed.items():
                      src_man[p] = _sig(ds)
                  manifest[num] = src_man
                  _save_manifest(manifest)

          The CLI prints the same diff: [05] changed=1 unchanged=0 removed=0. A --limit run never writes the manifest, so a quick demo ingest cannot poison the next real one. This is what makes the Phase 2 plan (an hourly cron, and a QABuddy MCP server exposing qa_search to IDE copilots, see chapters 9 and 10) a scheduling job rather than new code.

          07Set up, run, test, evaluate, deploy

          Python 3.11+ (the quickstart uses uv with 3.13; the Dockerfile uses 3.12). The models download on first use (about 4.6 GB for bge-m3 and the reranker, per the deploy notes). Only chat needs a key: ingest, tests and the eval run without one.

          terminal
          cd chapter_08_QABuddyAI
          uv venv .venv --python 3.13 && uv pip install -p .venv/bin/python -r requirements.txt
          cp .env.example .env                    # set GROQ_API_KEY (chat only)
          ./scripts/fetch_repos.sh                # clone the two framework repos
          ./scripts/setup_fixtures.sh             # test-case CSV and PRD from chapter 7
          .venv/bin/python -m app.ingestion.cli ingest --all
          .venv/bin/python -m pytest tests -q     # 12 unit tests
          .venv/bin/python scripts/eval.py        # golden-question hit rate
          ./scripts/dev.sh                        # UI on http://127.0.0.1:5080
          PieceWhat it doesCommand
          scripts/fetch_repos.shClones (or fast-forwards) the two framework repos into data/01 and data/02./scripts/fetch_repos.sh
          scripts/setup_fixtures.shCopies the chapter 7 test-case CSV and the VWO PRD into data/03 and data/09./scripts/setup_fixtures.sh
          Ingestion CLILoads, diffs against the manifest, embeds and upserts; stats and reset too.venv/bin/python -m app.ingestion.cli ingest --all
          Unit tests12 tests on chunkers and loaders with tiny fixtures; no models or keys.venv/bin/python -m pytest tests -q
          Retrieval eval13 golden questions; a HIT needs the expected source in the top 6; exits 1 under 70%.venv/bin/python scripts/eval.py
          Chat serverFlask UI and API on port 5080: /api/chat (SSE), /api/search, /api/ingest, health and stats./scripts/dev.sh
          scripts/jira_fetch.pyPulls tickets with REST and JQL into data/04 in the QABuddy schema.venv/bin/python scripts/jira_fetch.py
          scripts/backup.shNightly Qdrant snapshot (server mode) or tar of qdrant_data, plus the manifest./scripts/backup.sh
          Docker stackQdrant server, the app on gunicorn, Caddy with TLS and basic authdocker compose up -d --build

          Environment variables (names only): GROQ_API_KEY (or LLM_API_KEY), LLM_BASE_URL, LLM_MODEL, EMBED_MODEL, RERANK_MODEL, BGE_USE_FP16, QDRANT_PATH or QDRANT_URL, QDRANT_COLLECTION, PORT, and for Jira pulls JIRA_BASE_URL, JIRA_EMAIL, JIRA_API_TOKEN, JIRA_JQL. Tunables such as chunk sizes, top_k and the threshold live in config.yaml.

          terminal (24x7 on a VPS)
          docker compose up -d --build
          docker compose exec app python -m app.ingestion.cli ingest --all
          docker compose exec app python scripts/eval.py
          docker compose exec app curl -s localhost:5080/api/health
          README figures are claims, not measurements on this page: a 13/13 golden hit rate on the seeded 5,531-chunk corpus, about $0.001 per question on Groq, and about $55 to $75 a month for an 8 GB droplet. Re-measure them on your own data.

          08What to watch

          • The eval does not test the gate. scripts/eval.py counts a HIT when the right source is in the top 6, and only prints the best rerank score. In the demo, "What is JIRA VWO-2002 about?" retrieves the ticket but the best score is 0.188, so the gate answers "not in the knowledge base". Assert on the answer, not only on retrieval.
          • One process owns embedded Qdrant. While dev.sh runs, ingest from the UI panel, not the CLI, or run Qdrant as a server (QDRANT_URL).
          • A mislabelled failure. The build 128 failure chunk is labelled failure: Test, not LoginTest.testInvalidPassword: the test-name regex takes the first match in the block, and the context lines include [Pipeline] stage (Test).
          • Old logs are skipped. Jenkins logs older than retention_days (90) by file modification time are not ingested.
          • The plan and the code differ, on purpose. Plan.md names tree-sitter for code chunking; the build chose a small signature and brace-depth parser with a line-window fallback instead (no native dependencies), as the course notes record.
          • Streaming through a proxy. The Caddyfile sets flush_interval -1 on the reverse proxy so streamed tokens reach the browser as they arrive.
          • One gunicorn worker, eight threads (-w 1 --threads 8): every extra worker process would load its own copy of the models.
          • Unpinned image. docker-compose.yml uses qdrant/qdrant:latest; pin it before you depend on it.