Chapter 17
Chapter 17 . Agent frameworks . LangChain

LangChain: agents in plain Python

Chapters 4 and 5 built agents on a canvas. Here you write them in Python with LangChain 1.x: a bare model call, then create_agent, streaming, a system prompt, tools, typed output and an agent that drives a real browser. The chapter ends with a pipeline that reads a Jira ticket, checks the existing test library, plans, runs Playwright and reports.

13
Scripts
From 001, a raw model call, to 013, the full QA pipeline.
32
Playwright tools
In 5 bundles; they return error strings instead of raising.
3
Model providers
Groq, Gemini and DeepSeek, swapped by one line.
20
Library test cases
The local RAG corpus the pipeline checks before it plans.

01From a model call to an agent

Script 001 calls a chat model directly: llm.invoke(question) returns one message. From 002 on, the scripts build an agent with create_agent: a model plus an optional system prompt plus tools plus the loop that lets the model call those tools until it has an answer. The result is the whole message list, and the answer is its last message.

001_Hello_LC.py
def main():
    llm = ChatGroq(model=os.getenv("LLM_MODEL"),temperature=1)
    query = input(" Enter the question ")
    response = llm.invoke(query)
    print(response.content)
flowchart LR
  ENV[".env: GEMINI_LLM_MODEL + GOOGLE_API_KEY"] --> CA["create_agent(model)"]
  SP["system_prompt (optional role)"] -.-> CA
  T["tools (optional)"] -.-> CA
  CA --> AG[agent]
  AG -->|"invoke({messages: [...]})"| RES["result['messages']"]
  RES --> LAST["messages[-1].text = the answer"]
An agent is a model, a role, tools and a loop. You read the answer from the last message.
  • Provider-prefixed model strings: "google_genai:gemini-flash-lite-latest" picks the provider and the model in one string. Change the string to switch provider.
  • .text versus .content: Gemini returns a list of content blocks. message.text joins the text blocks; .content can show the raw list.
  • invoke or stream: invoke for CI and batch runs, stream(..., stream_mode="messages") when a person watches tokens arrive.

02Live demo: watch the agent loop

Three panels, each built from the chapter's real code. Tool call grows the agent's message list one message at a time while the calculator tool runs a faithful port of the script's safe arithmetic. Duplicate finder runs the keyword fallback of the local RAG over the real 20-case library, exactly as script 013 queries it. Two models replays the recorded browser runs: same tools, same prompt, different outcome.

Agent loop: tools, retrieval and two modelsNo API key needed
data-testid=lc-querydata-testid=lc-tool-stepdata-testid=lc-tool-rundata-testid=lc-tool-attackdata-testid=lc-messages
result["messages"] (0)
    The tool the model can call
    qa_metric_calculator(expression: str) -> str
    
    Calculate a QA metric from an arithmetic expression.
    
    Use this for pass rate, defect density, automation coverage, defect leakage
    or execution time. Input must be plain arithmetic with no words, for example
    '(438/500)*100' for a pass rate or '18/12' for defects per KLOC.

    The tool runs a real port of the script's safe arithmetic: it parses the expression and allows only numbers, + - * / ** and unary minus. Anything else returns the script's error string, so model text is never executed.

    The final AI messages in panel 1 are illustrative wording, because the model writes its own. Everything else (the tool calls, the tool results, the retrieval scores and the tool return strings) is computed or copied from the repo.

    03Script by script

    Each script adds one idea. Run them in order from chapter_17_LangChain/src/chapters.

    #ScriptWhat it teachesModelTools
    001001_Hello_LC.pyA bare chat-model call: llm.invoke() returns one messageGroqnone
    002002_Hello_Gemini.pyThe first agent: create_agent(model="provider:model")Gemininone
    003003_Hello_Gemini_Steam.pyStreaming tokens with stream_mode="messages"Gemininone
    004004_SP.pyA system prompt gives the agent a roleGemininone
    005005_Agent_Parallel_Vs_Sequential.pyFour questions at once with asyncio.gatherGemininone
    006006_Tool.pyOne @tool: the docstring is the description the model readsGeminiqa_metric_calculator
    007007_MultiTool.pyTwo tools; the model picks one per questionGeminisearch, fetch page
    008008_Structure_output.pyTyped output with response_format= and PydanticGemininone
    009009_Playwright_Agent_Orch.pyAn agent that drives a real browserGroq32 Playwright tools
    010010_Playwright_Agent_Orch_Deepseek.pySame agent, different model: tool discipline is a model propertyDeepSeek32 Playwright tools
    011011_FULL_E2E_Playwright_Agent_Orch_Deepseek.pyA full purchase flow with verify-as-you-go rulesDeepSeek32 Playwright tools
    012012_Fetch_JIRA_QA_Orch.pyJira ticket to typed plan to browser run, with two human gatesDeepSeekPlaywright tools
    013013_FULL_E2E_Fetch_JIRA_Local_RAG_QA_Orch.pyJira, local RAG, plan, browser, and a Slack summaryGroq + DeepSeekPlaywright + Slack tools

    Streaming (003)

    003_Hello_Gemini_Steam.py
        for token, metadata in agent.stream(
            {
                "messages": [
                    {"role": "system", "content": "You are a software testing instructor."},
                    {"role": "user", "content": "Explain LLM Eval in 3000 words blog"},
                ]
            },
            stream_mode="messages",
        ):
            print(token.text, end="", flush=True)
    Spot the bug: 003 asks input("Ask your question: ") and then ignores the answer: it always streams the hard-coded blog prompt. Drill 1 fixes it.

    A system prompt (004)

    004_SP.py
    agent = create_agent(
        model="google_genai:gemini-flash-lite-latest",
        system_prompt=(
            "You are a helpful AI assistant specialized in automation testing "
            "and software development. Be clear, concise, and always think step by step."
        ),
    )

    Several questions at once (005)

    005_Agent_Parallel_Vs_Sequential.py
    async def main():
        print("Creating LangChain agent...\n")
        start = time.perf_counter()
    
        # asyncio.gather runs every agent.ainvoke() call CONCURRENTLY
        results = await asyncio.gather(*[
            agent.ainvoke({"messages": [{"role": "user", "content": q["prompt"]}]})
            for q in questions
        ])

    asyncio.gather over agent.ainvoke is Python's Promise.all: all four calls are in flight together and the results come back in the order of the questions. Despite the file name, only the parallel path exists; writing the sequential loop is a drill.

    04Tools: how the model decides

    A tool is a Python function plus metadata. @tool takes the name from the function, the description from the docstring and the input schema from the type hints. The model reads the description, so the docstring is part of your prompt. It decides on its own whether a question needs the tool: the third query in 006 has no maths, so the agent answers directly.

    006_Tool.py
    @tool
    def qa_metric_calculator(expression: str) -> str:
        """Calculate a QA metric from an arithmetic expression.
    
        Use this for pass rate, defect density, automation coverage, defect leakage
        or execution time. Input must be plain arithmetic with no words, for example
        '(438/500)*100' for a pass rate or '18/12' for defects per KLOC.
        """
        try:
            result = _safe_eval(ast.parse(expression, mode="eval").body)
            return f"{round(result, 2)}"
        except Exception:
            return "Error: give me a plain arithmetic expression, e.g. '(438/500)*100'"
    
    
    agent = create_agent(
        model="google_genai:gemini-flash-lite-latest",
        tools=[qa_metric_calculator],          # list of available tools
        system_prompt=(
            "You are a QA metrics assistant for a test automation team. "
            "Use the qa_metric_calculator tool whenever a number must be computed. "
            "State the formula you used, then the result with its unit."
        ),
    )
    flowchart LR
      Q[User question] --> M["Model (tools registered)"]
      M --> D{Need a tool?}
      D -->|no| A[Answer directly]
      D -->|yes| C["Tool call + arguments"]
      C --> R["Python runs the function"]
      R -->|"result as a ToolMessage"| M
    The tool loop: the model asks, Python runs, the result goes back to the model.
    Never eval() model text. The calculator walks the parsed expression and allows only numbers and arithmetic operators, so __import__('os').system('ls') gets an error string back instead of running. Try it in panel 1 of the demo.

    Typed output (008)

    response_format=TestCaseList makes the agent return a validated Pydantic object in result["structured_response"]: no JSON parsing, and priority can only be High, Medium or Low.

    008_Structure_output.py
    # 1) Define the response schema with Pydantic
    class TestCase(BaseModel):
        title: str
        description: str
        steps: list[str]
        expected_result: str
        priority: Literal["High", "Medium", "Low"]
        tags: list[str]
    
    class TestCaseList(BaseModel):
        # Tip: always wrap arrays inside an object (better compatibility, especially with Anthropic)
        test_cases: list[TestCase]
    
    # 2) Create the agent with response_format
    agent = create_agent(
        model="google_genai:gemini-flash-lite-latest",
        # model="anthropic:claude-haiku-4-5",     # optional
        system_prompt=(
            "You are a senior automation testing engineer. Always return test cases "
            "in the exact structured JSON format requested. Be detailed and professional."
        ),
        response_format=TestCaseList,             # enforce the schema on the final output
    )

    05A browser agent with 32 Playwright tools

    From 009 on, the agent gets PLAYWRIGHT_TOOLS from playwright_tools.py: 32 async tools over one shared browser, in five bundles (core 12, interaction 10, assertion 5, diagnostic 2, advanced 3). They are designed for an agent, not a person:

    • Errors come back as strings, never as exceptions, because an agent cannot catch an exception. "Error: Browser not launched. Call launch_browser first." lets the model correct itself.
    • Look before acting: snapshot_page returns the page's ARIA snapshot, so the model reads real roles and names instead of guessing selectors.
    • Assertions return PASS or FAIL in plain text, and the diagnostic tools report console errors and failed requests.
    playwright_tools.py
    @tool
    @_needs_page
    async def assert_visible(selector: str, should_be_visible: bool = True) -> str:
        """Check whether an element is visible. Returns PASS or FAIL, never raises.
    
        This is the verification step of a test - call it after an action to decide
        whether the scenario actually passed.
        """
        actual = await _page.locator(selector).first.is_visible()
        ok = actual == should_be_visible
        return (f"{'PASS' if ok else 'FAIL'}: {selector} visible={actual}, "
                f"expected visible={should_be_visible}")

    Same tools, same prompt, different model

    RunModelWhat happened (recorded in the repo README)
    009Groq, openai/gpt-oss-120b, temperature 1Called navigate_to before launch_browser despite rule 1, got the error string back, launched, snapshotted a blank page, then gave up after 3 steps.
    010DeepSeek, deepseek-chat, temperature 0Followed the order and went further than asked: verified the post-login URL and title, asserted text=Products was visible, took a screenshot, closed cleanly. PASSED.
    011DeepSeek, temperature 0Full purchase flow: 11 steps, order confirmed.

    Tool discipline is a property of the model. Script 011 tightens the rules for a long flow:

    011_FULL_E2E_Playwright_Agent_Orch_Deepseek.py
    SYSTEM_PROMPT="""
    You are an end-to-end automation testing agent that controls a real browser
    using Playwright.
    
    Rules:
    1. Always call launch_browser FIRST, before any other action.
    2. Call snapshot_page whenever you land on a new page. It returns every role and
       name on that page, so you never have to guess a selector.
    3. Prefer click_by_role and type_by_label over raw CSS. This app also exposes
       stable [data-test="..."] attributes if you need a CSS selector.
    4. Verify as you go with assert_visible and assert_text_contains. Do not assume
       a step worked because the previous tool call returned no error.
    5. Never invent a value you were not given, and never report a price you did not
       actually read off the page.
    6. Before closing, call get_console_errors and take_screenshot.
    7. Report every step, every assertion with its PASS/FAIL, and one final verdict:
       TEST PASSED or TEST FAILED."""

    06From a Jira ticket to a tested result

    012 turns a ticket into a typed test plan and runs the automatable cases in Chromium, with a confirmation gate before planning and another before the browser. 013 adds retrieval: before planning, it searches the team's existing test cases so the plan can say which new cases duplicate, extend or add to the library, and a reporter agent "posts" the verdict through a dummy Slack MCP that only prints.

    flowchart LR
      J["Jira REST v3 (VWO-114)"] -->|offline fallback| F["fixtures/VWO-114.json"]
      J --> T[summary + description]
      F --> T
      T --> R["Local RAG: 20 cases"]
      R -->|"top-k similar cases"| P["Planner agent (Groq)<br/>typed TestPlan"]
      P --> G1{Human gate}
      G1 -->|yes| X["Executor agent (DeepSeek)<br/>32 Playwright tools"]
      X --> S["Reporter agent<br/>dummy Slack MCP"]
    Script 013: retrieval runs in plain Python before the planner, so the same ticket always sees the same context.

    The plan schema is where the guardrails live: a case the target app cannot support must say so instead of inventing UI.

    012_Fetch_JIRA_QA_Orch.py
    class TestCase(BaseModel):
        id: str = Field(description="e.g. TC-01")
        title: str
        priority: Literal["P0", "P1", "P2"]
        steps: list[str] = Field(description="Concrete UI steps against the target app")
        expected_result: str
        automatable: bool = Field(
            description="False when the ticket describes something absent from the target app")
        reason_if_not: str = Field(default="", description="Why it cannot be automated here")
    
    
    class TestPlan(BaseModel):
        ticket_key: str
        scope: str = Field(description="2-3 sentences: what is tested and what is not")
        risks: list[str]
        test_cases: list[TestCase]

    What the recorded runs found

    • VWO-49 (add Passkey and SSO login) yields 0 automatable cases against TTACart, correctly, because those buttons do not exist there. The script stops with "Nothing automatable against the target app."
    • VWO-114 ('Invalid credentials' message) found that the app's wording differs from the ticket (TTACart says "Epic sadface: Username and password do not match any user in this service") and that a username with surrounding spaces still logs in.
    • Two runs of the same ticket produced 6 cases (all passed) and 7 cases (2 passed, 5 failed). Generated plans are not reproducible, so pin a plan before it gates a build.

    The retrieval behind the duplicate check

    tta_rag.py
        def _keyword_search(self, query: str, k: int) -> list[tuple[float, dict]]:
            """Jaccard overlap fallback. Crude, but it never fails to install."""
            def toks(s: str) -> set[str]:
                return {w for w in re.findall(r"[a-z0-9]+", s.lower()) if len(w) > 2}
            q = toks(query)
            scored = [(len(q & toks(d)) / max(len(q | toks(d)), 1), c)
                      for d, c in zip(self.docs, self.cases)]
            return sorted(scored, key=lambda x: -x[0])[:k]

    With fastembed installed, the same interface uses BAAI/bge-small-en-v1.5 embeddings cached in .index.npz; without it, this keyword fallback still works offline. The demo's panel 2 is this function.

    07Set up and run

    One virtual environment for the chapter (Python 3.11+, LangChain 1.x: create_agent does not exist in 0.3).

    terminal
    cd chapter_17_LangChain/src
    python3 -m venv .venv
    .venv/bin/pip install -U langchain langchain-google-genai langchain-groq langchain-deepseek python-dotenv requests
    .venv/bin/pip install -U playwright && .venv/bin/playwright install chromium   # for 009 to 013
    .venv/bin/pip install -U fastembed                                            # optional, for 013's RAG
    cd chapters
    ../.venv/bin/python 002_Hello_Gemini.py
    ../.venv/bin/python 013_FULL_E2E_Fetch_JIRA_Local_RAG_QA_Orch.py VWO-114 --dry-run

    Keys go in chapter_17_LangChain/.env (never committed). Names only:

    VariableUsed byNotes
    GEMINI_LLM_MODEL, GOOGLE_API_KEY002 to 008Model string such as google_genai:gemini-flash-lite-latest. A Gemini key starts with AIza.
    LLM_MODEL, GROQ_API_KEY001, 009, the 013 plannerFor example openai/gpt-oss-120b.
    DEEPSEEK_API, DEEPSEEK_MODEL010 to 013Passed explicitly: the SDK itself looks for DEEPSEEK_API_KEY.
    JIRA_URL, JIRA_EMAIL, JIRA_API_TOKEN012, 013Optional. Without them the scripts read fixtures/<KEY>.json.
    TARGET_APP_URL, TARGET_APP_USER, TARGET_APP_PASS012, 013Default to TTACart and its public demo login.
    SLACK_CHANNEL013Default #qa-automation. The Slack tools are a dummy: nothing is sent.

    08Gotchas from real runs

    • Use temperature 0 for agents that drive tools. A sampled CSS selector is a flaky test you wrote on purpose.
    • Return errors as strings from tools. That is what let the Groq run recover from calling navigate_to too early.
    • load_dotenv() before os.getenv, and prefer os.environ["NAME"] for required values: a named KeyError beats a None passed down.
    • DeepSeek's thinking mode rejects a forced tool choice. response_format pins one function, so the 012 planner disables thinking through extra_body.
    • Keep grounding in code, not in a tool. 013 runs retrieval in Python so every run of the same ticket sees the same library context.
    • Stub the integration at one seam. The dummy Slack MCP replaces only call_tool; the agent code is identical to the real integration.
    012_Fetch_JIRA_QA_Orch.py
    # Stage 2 is NOT. response_format pins tool_choice to one function, and
    # deepseek-flash rejects that while thinking is on:
    #     400 "Thinking mode does not support this tool_choice"
    # extra_body reaches the raw request body; model_kwargs does not.
    planner_llm = ChatDeepSeek(
        model=MODEL,
        api_key=os.environ["DEEPSEEK_API"],
        temperature=0,
        extra_body={"thinking": {"type": "disabled"}},
    )