Chapters 4 and 5 built agents on a canvas. Here you write them in Python with LangChain 1.x: a bare model call, then create_agent, streaming, a system prompt, tools, typed output and an agent that drives a real browser. The chapter ends with a pipeline that reads a Jira ticket, checks the existing test library, plans, runs Playwright and reports.
From 001, a raw model call, to 013, the full QA pipeline.
32
Playwright tools
In 5 bundles; they return error strings instead of raising.
3
Model providers
Groq, Gemini and DeepSeek, swapped by one line.
20
Library test cases
The local RAG corpus the pipeline checks before it plans.
01From a model call to an agent
Script 001 calls a chat model directly: llm.invoke(question) returns one message. From 002 on, the scripts build an agent with create_agent: a model plus an optional system prompt plus tools plus the loop that lets the model call those tools until it has an answer. The result is the whole message list, and the answer is its last message.
001_Hello_LC.py
defmain():
llm = ChatGroq(model=os.getenv("LLM_MODEL"),temperature=1)
query = input(" Enter the question ")
response = llm.invoke(query)
print(response.content)
002_Hello_Gemini.py
defmain():
agent = create_agent(model=os.environ["GEMINI_LLM_MODEL"])
query = input("Ask your question: ")
result = agent.invoke({"messages": [("user", query)]})
print(result["messages"][-1].text)
flowchart LR
ENV[".env: GEMINI_LLM_MODEL + GOOGLE_API_KEY"] --> CA["create_agent(model)"]
SP["system_prompt (optional role)"] -.-> CA
T["tools (optional)"] -.-> CA
CA --> AG[agent]
AG -->|"invoke({messages: [...]})"| RES["result['messages']"]
RES --> LAST["messages[-1].text = the answer"]
An agent is a model, a role, tools and a loop. You read the answer from the last message.
Provider-prefixed model strings:"google_genai:gemini-flash-lite-latest" picks the provider and the model in one string. Change the string to switch provider.
.text versus .content: Gemini returns a list of content blocks. message.text joins the text blocks; .content can show the raw list.
invoke or stream:invoke for CI and batch runs, stream(..., stream_mode="messages") when a person watches tokens arrive.
02Live demo: watch the agent loop
Three panels, each built from the chapter's real code. Tool call grows the agent's message list one message at a time while the calculator tool runs a faithful port of the script's safe arithmetic. Duplicate finder runs the keyword fallback of the local RAG over the real 20-case library, exactly as script 013 queries it. Two models replays the recorded browser runs: same tools, same prompt, different outcome.
Agent loop: tools, retrieval and two modelsNo API key needed
qa_metric_calculator(expression: str) -> str
Calculate a QA metric from an arithmetic expression.
Use this for pass rate, defect density, automation coverage, defect leakage
or execution time. Input must be plain arithmetic with no words, for example
'(438/500)*100' for a pass rate or '18/12' for defects per KLOC.
The tool runs a real port of the script's safe arithmetic: it parses the expression and allows only numbers, + - * / ** and unary minus. Anything else returns the script's error string, so model text is never executed.
A replay of the two recorded runs. The tool return strings are the exact templates from playwright_tools.py; the arguments and the screenshot name are illustrative.
The final AI messages in panel 1 are illustrative wording, because the model writes its own. Everything else (the tool calls, the tool results, the retrieval scores and the tool return strings) is computed or copied from the repo.
03Script by script
Each script adds one idea. Run them in order from chapter_17_LangChain/src/chapters.
#
Script
What it teaches
Model
Tools
001
001_Hello_LC.py
A bare chat-model call: llm.invoke() returns one message
Groq
none
002
002_Hello_Gemini.py
The first agent: create_agent(model="provider:model")
Gemini
none
003
003_Hello_Gemini_Steam.py
Streaming tokens with stream_mode="messages"
Gemini
none
004
004_SP.py
A system prompt gives the agent a role
Gemini
none
005
005_Agent_Parallel_Vs_Sequential.py
Four questions at once with asyncio.gather
Gemini
none
006
006_Tool.py
One @tool: the docstring is the description the model reads
Gemini
qa_metric_calculator
007
007_MultiTool.py
Two tools; the model picks one per question
Gemini
search, fetch page
008
008_Structure_output.py
Typed output with response_format= and Pydantic
Gemini
none
009
009_Playwright_Agent_Orch.py
An agent that drives a real browser
Groq
32 Playwright tools
010
010_Playwright_Agent_Orch_Deepseek.py
Same agent, different model: tool discipline is a model property
DeepSeek
32 Playwright tools
011
011_FULL_E2E_Playwright_Agent_Orch_Deepseek.py
A full purchase flow with verify-as-you-go rules
DeepSeek
32 Playwright tools
012
012_Fetch_JIRA_QA_Orch.py
Jira ticket to typed plan to browser run, with two human gates
DeepSeek
Playwright tools
013
013_FULL_E2E_Fetch_JIRA_Local_RAG_QA_Orch.py
Jira, local RAG, plan, browser, and a Slack summary
Groq + DeepSeek
Playwright + Slack tools
Streaming (003)
003_Hello_Gemini_Steam.py
for token, metadata in agent.stream(
{
"messages": [
{"role": "system", "content": "You are a software testing instructor."},
{"role": "user", "content": "Explain LLM Eval in 3000 words blog"},
]
},
stream_mode="messages",
):
print(token.text, end="", flush=True)
Spot the bug: 003 asks input("Ask your question: ") and then ignores the answer: it always streams the hard-coded blog prompt. Drill 1 fixes it.
A system prompt (004)
004_SP.py
agent = create_agent(
model="google_genai:gemini-flash-lite-latest",
system_prompt=(
"You are a helpful AI assistant specialized in automation testing ""and software development. Be clear, concise, and always think step by step."
),
)
Several questions at once (005)
005_Agent_Parallel_Vs_Sequential.py
asyncdefmain():
print("Creating LangChain agent...\n")
start = time.perf_counter()
# asyncio.gather runs every agent.ainvoke() call CONCURRENTLY
results = await asyncio.gather(*[
agent.ainvoke({"messages": [{"role": "user", "content": q["prompt"]}]})
for q in questions
])
asyncio.gather over agent.ainvoke is Python's Promise.all: all four calls are in flight together and the results come back in the order of the questions. Despite the file name, only the parallel path exists; writing the sequential loop is a drill.
04Tools: how the model decides
A tool is a Python function plus metadata. @tool takes the name from the function, the description from the docstring and the input schema from the type hints. The model reads the description, so the docstring is part of your prompt. It decides on its own whether a question needs the tool: the third query in 006 has no maths, so the agent answers directly.
006_Tool.py
@tooldefqa_metric_calculator(expression: str) -> str:
"""Calculate a QA metric from an arithmetic expression.
Use this for pass rate, defect density, automation coverage, defect leakage
or execution time. Input must be plain arithmetic with no words, for example
'(438/500)*100' for a pass rate or '18/12' for defects per KLOC.
"""try:
result = _safe_eval(ast.parse(expression, mode="eval").body)
returnf"{round(result, 2)}"except Exception:
return"Error: give me a plain arithmetic expression, e.g. '(438/500)*100'"
agent = create_agent(
model="google_genai:gemini-flash-lite-latest",
tools=[qa_metric_calculator], # list of available tools
system_prompt=(
"You are a QA metrics assistant for a test automation team. ""Use the qa_metric_calculator tool whenever a number must be computed. ""State the formula you used, then the result with its unit."
),
)
flowchart LR
Q[User question] --> M["Model (tools registered)"]
M --> D{Need a tool?}
D -->|no| A[Answer directly]
D -->|yes| C["Tool call + arguments"]
C --> R["Python runs the function"]
R -->|"result as a ToolMessage"| M
The tool loop: the model asks, Python runs, the result goes back to the model.
Never eval() model text. The calculator walks the parsed expression and allows only numbers and arithmetic operators, so __import__('os').system('ls') gets an error string back instead of running. Try it in panel 1 of the demo.
Typed output (008)
response_format=TestCaseList makes the agent return a validated Pydantic object in result["structured_response"]: no JSON parsing, and priority can only be High, Medium or Low.
008_Structure_output.py
# 1) Define the response schema with PydanticclassTestCase(BaseModel):
title: str
description: str
steps: list[str]
expected_result: str
priority: Literal["High", "Medium", "Low"]
tags: list[str]
classTestCaseList(BaseModel):
# Tip: always wrap arrays inside an object (better compatibility, especially with Anthropic)
test_cases: list[TestCase]
# 2) Create the agent with response_format
agent = create_agent(
model="google_genai:gemini-flash-lite-latest",
# model="anthropic:claude-haiku-4-5", # optional
system_prompt=(
"You are a senior automation testing engineer. Always return test cases ""in the exact structured JSON format requested. Be detailed and professional."
),
response_format=TestCaseList, # enforce the schema on the final output
)
05A browser agent with 32 Playwright tools
From 009 on, the agent gets PLAYWRIGHT_TOOLS from playwright_tools.py: 32 async tools over one shared browser, in five bundles (core 12, interaction 10, assertion 5, diagnostic 2, advanced 3). They are designed for an agent, not a person:
Errors come back as strings, never as exceptions, because an agent cannot catch an exception. "Error: Browser not launched. Call launch_browser first." lets the model correct itself.
Look before acting:snapshot_page returns the page's ARIA snapshot, so the model reads real roles and names instead of guessing selectors.
Assertions return PASS or FAIL in plain text, and the diagnostic tools report console errors and failed requests.
playwright_tools.py
@tool@_needs_pageasyncdefassert_visible(selector: str, should_be_visible: bool = True) -> str:
"""Check whether an element is visible. Returns PASS or FAIL, never raises.
This is the verification step of a test - call it after an action to decide
whether the scenario actually passed.
"""
actual = await _page.locator(selector).first.is_visible()
ok = actual == should_be_visible
return (f"{'PASS' if ok else 'FAIL'}: {selector} visible={actual}, "f"expected visible={should_be_visible}")
Same tools, same prompt, different model
Run
Model
What happened (recorded in the repo README)
009
Groq, openai/gpt-oss-120b, temperature 1
Called navigate_to before launch_browser despite rule 1, got the error string back, launched, snapshotted a blank page, then gave up after 3 steps.
010
DeepSeek, deepseek-chat, temperature 0
Followed the order and went further than asked: verified the post-login URL and title, asserted text=Products was visible, took a screenshot, closed cleanly. PASSED.
011
DeepSeek, temperature 0
Full purchase flow: 11 steps, order confirmed.
Tool discipline is a property of the model. Script 011 tightens the rules for a long flow:
011_FULL_E2E_Playwright_Agent_Orch_Deepseek.py
SYSTEM_PROMPT="""
You are an end-to-end automation testing agent that controls a real browser
using Playwright.
Rules:
1. Always call launch_browser FIRST, before any other action.
2. Call snapshot_page whenever you land on a new page. It returns every role and
name on that page, so you never have to guess a selector.
3. Prefer click_by_role and type_by_label over raw CSS. This app also exposes
stable [data-test="..."] attributes if you need a CSS selector.
4. Verify as you go with assert_visible and assert_text_contains. Do not assume
a step worked because the previous tool call returned no error.
5. Never invent a value you were not given, and never report a price you did not
actually read off the page.
6. Before closing, call get_console_errors and take_screenshot.
7. Report every step, every assertion with its PASS/FAIL, and one final verdict:
TEST PASSED or TEST FAILED."""
06From a Jira ticket to a tested result
012 turns a ticket into a typed test plan and runs the automatable cases in Chromium, with a confirmation gate before planning and another before the browser. 013 adds retrieval: before planning, it searches the team's existing test cases so the plan can say which new cases duplicate, extend or add to the library, and a reporter agent "posts" the verdict through a dummy Slack MCP that only prints.
flowchart LR
J["Jira REST v3 (VWO-114)"] -->|offline fallback| F["fixtures/VWO-114.json"]
J --> T[summary + description]
F --> T
T --> R["Local RAG: 20 cases"]
R -->|"top-k similar cases"| P["Planner agent (Groq)<br/>typed TestPlan"]
P --> G1{Human gate}
G1 -->|yes| X["Executor agent (DeepSeek)<br/>32 Playwright tools"]
X --> S["Reporter agent<br/>dummy Slack MCP"]
Script 013: retrieval runs in plain Python before the planner, so the same ticket always sees the same context.
The plan schema is where the guardrails live: a case the target app cannot support must say so instead of inventing UI.
012_Fetch_JIRA_QA_Orch.py
classTestCase(BaseModel):
id: str = Field(description="e.g. TC-01")
title: str
priority: Literal["P0", "P1", "P2"]
steps: list[str] = Field(description="Concrete UI steps against the target app")
expected_result: str
automatable: bool = Field(
description="False when the ticket describes something absent from the target app")
reason_if_not: str = Field(default="", description="Why it cannot be automated here")
classTestPlan(BaseModel):
ticket_key: str
scope: str = Field(description="2-3 sentences: what is tested and what is not")
risks: list[str]
test_cases: list[TestCase]
What the recorded runs found
VWO-49 (add Passkey and SSO login) yields 0 automatable cases against TTACart, correctly, because those buttons do not exist there. The script stops with "Nothing automatable against the target app."
VWO-114 ('Invalid credentials' message) found that the app's wording differs from the ticket (TTACart says "Epic sadface: Username and password do not match any user in this service") and that a username with surrounding spaces still logs in.
Two runs of the same ticket produced 6 cases (all passed) and 7 cases (2 passed, 5 failed). Generated plans are not reproducible, so pin a plan before it gates a build.
The retrieval behind the duplicate check
tta_rag.py
def_keyword_search(self, query: str, k: int) -> list[tuple[float, dict]]:
"""Jaccard overlap fallback. Crude, but it never fails to install."""deftoks(s: str) -> set[str]:
return {w for w in re.findall(r"[a-z0-9]+", s.lower()) iflen(w) > 2}
q = toks(query)
scored = [(len(q & toks(d)) / max(len(q | toks(d)), 1), c)
for d, c inzip(self.docs, self.cases)]
returnsorted(scored, key=lambda x: -x[0])[:k]
With fastembed installed, the same interface uses BAAI/bge-small-en-v1.5 embeddings cached in .index.npz; without it, this keyword fallback still works offline. The demo's panel 2 is this function.
07Set up and run
One virtual environment for the chapter (Python 3.11+, LangChain 1.x: create_agent does not exist in 0.3).
Keys go in chapter_17_LangChain/.env (never committed). Names only:
Variable
Used by
Notes
GEMINI_LLM_MODEL, GOOGLE_API_KEY
002 to 008
Model string such as google_genai:gemini-flash-lite-latest. A Gemini key starts with AIza.
LLM_MODEL, GROQ_API_KEY
001, 009, the 013 planner
For example openai/gpt-oss-120b.
DEEPSEEK_API, DEEPSEEK_MODEL
010 to 013
Passed explicitly: the SDK itself looks for DEEPSEEK_API_KEY.
JIRA_URL, JIRA_EMAIL, JIRA_API_TOKEN
012, 013
Optional. Without them the scripts read fixtures/<KEY>.json.
TARGET_APP_URL, TARGET_APP_USER, TARGET_APP_PASS
012, 013
Default to TTACart and its public demo login.
SLACK_CHANNEL
013
Default #qa-automation. The Slack tools are a dummy: nothing is sent.
08Gotchas from real runs
Use temperature 0 for agents that drive tools. A sampled CSS selector is a flaky test you wrote on purpose.
Return errors as strings from tools. That is what let the Groq run recover from calling navigate_to too early.
load_dotenv() before os.getenv, and prefer os.environ["NAME"] for required values: a named KeyError beats a None passed down.
DeepSeek's thinking mode rejects a forced tool choice.response_format pins one function, so the 012 planner disables thinking through extra_body.
Keep grounding in code, not in a tool. 013 runs retrieval in Python so every run of the same ticket sees the same library context.
Stub the integration at one seam. The dummy Slack MCP replaces only call_tool; the agent code is identical to the real integration.
012_Fetch_JIRA_QA_Orch.py
# Stage 2 is NOT. response_format pins tool_choice to one function, and# deepseek-flash rejects that while thinking is on:# 400 "Thinking mode does not support this tool_choice"# extra_body reaches the raw request body; model_kwargs does not.
planner_llm = ChatDeepSeek(
model=MODEL,
api_key=os.environ["DEEPSEEK_API"],
temperature=0,
extra_body={"thinking": {"type": "disabled"}},
)
DDrills for the chapter
Code drills run from chapter_17_LangChain/src/chapters. Playwright drills target the live demo on this page; turn on Show locator badges to see every data-testid.
Stream the real question
Make 003 stream the question you type instead of the hard-coded blog prompt.
Hint
Only the user message content changes.
Expected result
Replace the hard-coded content with query; the streamed answer now matches your question.
Write the sequential path
Add a plain for loop with agent.invoke to 005 and time both versions.
Expected result
The parallel run takes roughly as long as the slowest single call; the sequential run takes roughly the sum. Exact times vary.
Compute a new metric
Ask 006 for automation coverage when 180 of 240 cases are automated.
Error: give me a plain arithmetic expression, e.g. '(438/500)*100', and nothing executes.
Extend the schema
Add preconditions: list[str] to TestCase in 008 and run it.
Expected result
Every generated case now has a preconditions list; no parsing code changed.
Read the retrieval
Run tta_rag.py without fastembed installed.
Expected result
backend: keyword; "SSO and passkey login buttons" ranks TTA-020 first at 0.114.
Drive the tool loopPlaywright
In panel 1, run the pass-rate question and assert the message list and the tool result.
Hint
lc-query, lc-tool-run, lc-messages (rows have data-kind).
Expected result
Four messages; the ToolMessage row (data-kind="ok") contains 87.6.
Prove the tool refuses codePlaywright
Click lc-tool-attack and assert the error row.
Expected result
A data-kind="bad" row with the exact error string.
Assert the retrieval rankingPlaywright
In panel 2, select VWO-114 and then VWO-49 and assert the first hit each time.
Hint
Hit rows carry data-case.
Expected result
VWO-114: TTA-002 at 0.211. VWO-49: TTA-020 at 0.181.
Compare the two modelsPlaywright
In panel 3, run Groq, then DeepSeek, and assert both verdicts.
Expected result
"Gave up after 3 steps. Not logged in." then "TEST PASSED. ..."
SSolutions and the key files
The Playwright spec passes against this page as written. The Python files are the chapter's own, verbatim.
tests/langchain-demo.spec.ts
import { test, expect } from'@playwright/test';
const URL = 'https://app.thetestingacademy.com/ai/blueprint/learn/langchain.html';
test('the agent calls the calculator tool for a pass rate', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('lc-query').selectOption('pass-rate');
await page.getByTestId('lc-tool-run').click();
const messages = page.getByTestId('lc-messages').locator('li');
awaitexpect(messages).toHaveCount(4);
awaitexpect(messages.nth(1)).toContainText('qa_metric_calculator');
awaitexpect(page.getByTestId('lc-messages').locator('li[data-kind="ok"]')).toContainText('87.6');
});
test('a question without maths gets no tool call', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('lc-query').selectOption('no-maths');
await page.getByTestId('lc-tool-run').click();
awaitexpect(page.getByTestId('lc-messages').locator('li')).toHaveCount(2);
awaitexpect(page.getByTestId('lc-messages')).not.toContainText('tool_calls');
});
test('the tool refuses code instead of running it', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('lc-tool-attack').click();
awaitexpect(page.getByTestId('lc-messages').locator('li[data-kind="bad"]'))
.toContainText("Error: give me a plain arithmetic expression, e.g. '(438/500)*100'");
});
test('retrieval ranks the closest existing test case first', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('lc-mode-rag').click();
await page.getByTestId('lc-ticket').selectOption('VWO-114');
awaitexpect(page.getByTestId('lc-hits').locator('li').first()).toHaveAttribute('data-case', 'TTA-002');
awaitexpect(page.getByTestId('lc-rag-status')).toContainText('TTA-002 0.211');
await page.getByTestId('lc-ticket').selectOption('VWO-49');
awaitexpect(page.getByTestId('lc-hits').locator('li').first()).toHaveAttribute('data-case', 'TTA-020');
});
test('the same tools with two models end differently', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('lc-mode-models').click();
await page.getByTestId('lc-model').selectOption('groq');
await page.getByTestId('lc-replay-run').click();
awaitexpect(page.getByTestId('lc-verdict')).toContainText('Gave up after 3 steps');
await page.getByTestId('lc-model').selectOption('deepseek');
await page.getByTestId('lc-replay-run').click();
awaitexpect(page.getByTestId('lc-verdict')).toContainText('TEST PASSED');
awaitexpect(page.getByTestId('lc-trace').locator('li[data-tool="close_browser"]')).toHaveCount(1);
});
006_Tool.py
import ast
import operator
from dotenv import load_dotenv
from langchain.agents import create_agent
from langchain.tools import tool # create custom toolsload_dotenv()
# eval() on text a model produced is arbitrary code execution. This walks the# parsed expression instead, so only arithmetic can ever run.
_OPS = {
ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul,
ast.Div: operator.truediv, ast.Pow: operator.pow, ast.USub: operator.neg,
}
def_safe_eval(node):
ifisinstance(node, ast.Constant) andisinstance(node.value, (int, float)):
return node.value
ifisinstance(node, ast.BinOp) andtype(node.op) in _OPS:
return _OPS[type(node.op)](_safe_eval(node.left), _safe_eval(node.right))
ifisinstance(node, ast.UnaryOp) andtype(node.op) in _OPS:
return _OPS[type(node.op)](_safe_eval(node.operand))
raiseValueError("only arithmetic is allowed")
@tooldefqa_metric_calculator(expression: str) -> str:
"""Calculate a QA metric from an arithmetic expression.
Use this for pass rate, defect density, automation coverage, defect leakage
or execution time. Input must be plain arithmetic with no words, for example
'(438/500)*100' for a pass rate or '18/12' for defects per KLOC.
"""try:
result = _safe_eval(ast.parse(expression, mode="eval").body)
returnf"{round(result, 2)}"except Exception:
return"Error: give me a plain arithmetic expression, e.g. '(438/500)*100'"
agent = create_agent(
model="google_genai:gemini-flash-lite-latest",
tools=[qa_metric_calculator], # list of available tools
system_prompt=(
"You are a QA metrics assistant for a test automation team. ""Use the qa_metric_calculator tool whenever a number must be computed. ""State the formula you used, then the result with its unit."
),
)
queries = [
# maths -> the agent must call the tool"We ran 500 regression tests and 438 passed. What is the pass rate?",
"A module of 12 KLOC has 18 defects. What is the defect density per KLOC?",
# no maths -> the agent answers directly, no tool call"What is the difference between smoke testing and sanity testing?",
]
for q in queries:
print(f"\nQuestion: {q}")
result = agent.invoke({"messages": [{"role": "user", "content": q}]})
print("Agent:", result["messages"][-1].text)
008_Structure_output.py
from typing import Literal
from dotenv import load_dotenv
from pydantic import BaseModel
from langchain.agents import create_agent
load_dotenv()
# 1) Define the response schema with PydanticclassTestCase(BaseModel):
title: str
description: str
steps: list[str]
expected_result: str
priority: Literal["High", "Medium", "Low"]
tags: list[str]
classTestCaseList(BaseModel):
# Tip: always wrap arrays inside an object (better compatibility, especially with Anthropic)
test_cases: list[TestCase]
# 2) Create the agent with response_format
agent = create_agent(
model="google_genai:gemini-flash-lite-latest",
# model="anthropic:claude-haiku-4-5", # optional
system_prompt=(
"You are a senior automation testing engineer. Always return test cases ""in the exact structured JSON format requested. Be detailed and professional."
),
response_format=TestCaseList, # enforce the schema on the final output
)
# 3) Provide a user story and invoke the agent
user_story = ("As a logged-in user, I want to add items to my shopping cart ""so that I can purchase them later.")
result = agent.invoke({
"messages": [{
"role": "user",
"content": f"Generate 2 detailed test cases for this user story: {user_story}",
}]
})
# 4) Work with the structured result: no manual parsingprint("=== Structured Output ===")
test_cases = result["structured_response"].test_cases
for tc in test_cases:
print(tc.model_dump_json(indent=2))
print(f"\nSuccessfully generated {len(test_cases)} test cases!")
tta_rag.py
"""A very small local RAG over the TTACart test-case library.
No server, no database, no API key. 20 test cases in a JSON file, embedded once
with fastembed (ONNX, ~130MB model, downloads on first run) and cached to disk as
a .npy matrix. Retrieval is a cosine dot product over 20 rows, which is instant.
Why this is enough: RAG is retrieval + generation. At 20 documents a vector
"database" is a numpy array, and anything heavier is ceremony. Swap the search()
body for a Qdrant call when the corpus outgrows memory, the interface is the same.
If fastembed is missing the module degrades to keyword overlap scoring, so a
class demo still works offline with no install.
"""import hashlib
import json
import re
from pathlib import Path
CORPUS_PATH = Path(__file__).parent / "rag_corpus" / "tta_testcases.json"
INDEX_PATH = Path(__file__).parent / "rag_corpus" / ".index.npz"
MODEL_NAME = "BAAI/bge-small-en-v1.5"# 384 dims, ~130MBdef_document(case: dict) -> str:
"""What actually gets embedded. Title and tags carry most of the signal."""return (f"{case['title']}. Area: {case['area']}. Type: {case['type']} test case. "f"Tags: {', '.join(case['tags'])}. "f"Steps: {' '.join(case['steps'])} Expected: {case['expected']}")
class TestCaseRAG:
def__init__(self, corpus_path: Path = CORPUS_PATH):
self.cases: list[dict] = json.loads(corpus_path.read_text())
self.docs = [_document(c) for c inself.cases]
self._vectors = Noneself.backend = "keyword"self._try_load_vectors()
# ---------------------------------------------------------------- indexdef_corpus_fingerprint(self) -> str:
return hashlib.sha256("\n".join(self.docs).encode()).hexdigest()[:16]
def_try_load_vectors(self) -> None:
try:
import numpy as np
from fastembed import TextEmbedding
except ImportError:
return# stay on keyword scoring
fp = self._corpus_fingerprint()
if INDEX_PATH.exists():
cached = np.load(INDEX_PATH, allow_pickle=False)
ifstr(cached["fingerprint"]) == fp: # corpus unchanged -> reuseself._vectors, self.backend = cached["vectors"], f"embeddings ({MODEL_NAME}, cached)"return
model = TextEmbedding(model_name=MODEL_NAME)
vectors = np.array(list(model.embed(self.docs)), dtype="float32")
vectors /= np.linalg.norm(vectors, axis=1, keepdims=True) # unit -> dot == cosine
np.savez(INDEX_PATH, vectors=vectors, fingerprint=fp)
self._vectors, self.backend = vectors, f"embeddings ({MODEL_NAME}, built)"# --------------------------------------------------------------- searchdefsearch(self, query: str, k: int = 5) -> list[tuple[float, dict]]:
"""Return the k most similar test cases as (score, case), best first."""ifself._vectors isNone:
returnself._keyword_search(query, k)
import numpy as np
from fastembed import TextEmbedding
q = np.array(next(iter(TextEmbedding(model_name=MODEL_NAME).embed([query]))), dtype="float32")
q /= np.linalg.norm(q)
scores = self._vectors @ q
order = np.argsort(-scores)[:k]
return [(float(scores[i]), self.cases[i]) for i in order]
def_keyword_search(self, query: str, k: int) -> list[tuple[float, dict]]:
"""Jaccard overlap fallback. Crude, but it never fails to install."""deftoks(s: str) -> set[str]:
return {w for w in re.findall(r"[a-z0-9]+", s.lower()) iflen(w) > 2}
q = toks(query)
scored = [(len(q & toks(d)) / max(len(q | toks(d)), 1), c)
for d, c inzip(self.docs, self.cases)]
returnsorted(scored, key=lambda x: -x[0])[:k]
defformat_for_prompt(hits: list[tuple[float, dict]]) -> str:
"""Render retrieved cases as context the planner can cite by id."""ifnot hits:
return"No existing test cases matched."
out = []
for score, c in hits:
out.append(
f"[{c['id']}] ({c['type'].upper()}, {c['priority']}, {c['area']}, "f"{c['status']}, last run: {c['last_result']}, similarity {score:.3f})\n"f" {c['title']}\n"f" Expected: {c['expected']}\n"f" Tags: {', '.join(c['tags'])}")
return"\n".join(out)
if __name__ == "__main__":
rag = TestCaseRAG()
print(f"corpus: {len(rag.cases)} cases | backend: {rag.backend}\n")
for q in ["SSO and passkey login buttons",
"wrong password shows an error",
"tax and total calculation at checkout"]:
print(f"QUERY: {q}")
for s, c in rag.search(q, 3):
print(f" {s:.3f} {c['id']} {c['title']}")
print()
012_Fetch_JIRA_QA_Orch.py
# ---------------------------------------------------------------- stage 2classTestCase(BaseModel):
id: str = Field(description="e.g. TC-01")
title: str
priority: Literal["P0", "P1", "P2"]
steps: list[str] = Field(description="Concrete UI steps against the target app")
expected_result: str
automatable: bool = Field(
description="False when the ticket describes something absent from the target app")
reason_if_not: str = Field(default="", description="Why it cannot be automated here")
classTestPlan(BaseModel):
ticket_key: str
scope: str = Field(description="2-3 sentences: what is tested and what is not")
risks: list[str]
test_cases: list[TestCase]
PLANNER_PROMPT = f"""You are a senior QA engineer writing a test plan from a Jira ticket.
The test cases will be executed by a browser agent against THIS app, which may be
a different product from the one the ticket describes:
Target app: {TARGET_URL}
Login: {TARGET_USER} / {TARGET_PASS}
Rules:
1. Write 5-8 test cases covering the ticket's acceptance criteria.
2. Set automatable=true ONLY when every step can be performed on the target app
above. If the ticket describes a feature the target app does not have (for
example SSO or passkey buttons that are not on this login page), set
automatable=false and say so in reason_if_not. Do NOT invent UI.
3. Steps must be concrete and clickable: name the field, button or text to check.
4. Never invent credentials beyond the ones given."""# ---------------------------------------------------------------- stage 3
EXECUTOR_PROMPT = """You are an end-to-end automation testing agent driving a real
browser with Playwright.
Rules:
1. Call launch_browser FIRST.
2. Call snapshot_page on every new page instead of guessing selectors.
3. Prefer click_by_role and type_by_label over raw CSS.
4. Verify with assert_visible / assert_text_contains. Never assume a step worked.
5. If a case's element genuinely does not exist, report it FAILED with what you
saw. Do not pretend it passed and do not substitute a different element.
6. Before closing, call get_console_errors and take_screenshot.
7. Report each test case id with PASS or FAIL, then a final summary line."""