Where this sits
DeepEval is behind the batch, and this is the last of the big topics: LangChain, in Python. The framing was practical. Ninety percent of what a QA needs from an agent framework, LangChain does on its own; the rest of the family is for when you outgrow it.
Asked in class: how is this different from CrewAI, which the batch already used? Both build agents with tools, memory and MCP. The honest answer given was that they are two pizzas from two shops, and the choice depends on the job: LangChain for a small agent that works on its own, CrewAI for several agents that have to talk to each other. Real teams mix them, including the instructor's own and others he named.
The four names, sorted
This was the part the room was told to remember, and the analogy is the reason it sticks:
The class's own way of choosing: basic agent work, LangChain. Approvals, loops, several agents passing state between them (a planner writes the test plan, a reviewer picks it up), LangGraph. Wanting to know whether the answer was any good, LangSmith or LangFuse.
LangGraph and LangSmith are not part of this course. They were offered as extra sessions for anyone who wants them, and the instructor mentioned having built a LangGraph RAG pipeline at work, where the stack is LangGraph plus LangFuse rather than LangSmith.
Setup
A virtual environment and a handful of packages:
mkdir langchain && cd langchain
pip install virtualenv
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install langchain langchain-groq langchain-google-genai langchain-ollama python-dotenv pydantic
One package per provider: Groq, Google Gemini and Ollama each ship their own, and so do OpenAI, Anthropic, OpenRouter and Mistral. Keys go in .env, never in the file.
GROQ_API_KEY=...
LLM_MODEL=openai/gpt-oss-120b
GEMINI_LLM_MODEL=google_genai:gemini-flash-lite-latest
Gemini's key must be named GOOGLE_API_KEY in the environment. The model name can live under any variable you like, but the library looks for that exact key name and fails confusingly if it is called something else.
The first agent, two ways
The direct way: build the model, invoke it, print the reply.
from dotenv import load_dotenv
from langchain_groq import ChatGroq
import os
load_dotenv()
def main():
llm = ChatGroq(model=os.getenv("LLM_MODEL"), temperature=1)
query = input(" Enter the question ")
response = llm.invoke(query)
print(response.content)
if __name__ == "__main__":
main()
Ask it what two plus two is and it answers. That is the whole file, and the reaction in class was the right one: it is simpler than people expect.
The agent way, which is what the rest of the chapter uses:
from dotenv import load_dotenv
from langchain.agents import create_agent
import os
load_dotenv()
def main():
agent = create_agent(model=os.environ["GEMINI_LLM_MODEL"])
query = input("Ask your question: ")
result = agent.invoke({"messages": [("user", query)]})
print(result["messages"][-1].text)
The difference matters later: create_agent is what takes tools, a system prompt and a response format. ChatGroq on its own is just the model.
Two __main__ typos cost a few minutes live, and they are the classic ones: = instead of ==, and the underscores around main. Worth recognising quickly, because the error message (name 'main' is not defined) does not point at the real problem.
The system prompt
Same shape as the TypeScript version of this material: the system prompt is the job description, the user message is the task.
agent = create_agent(
model="google_genai:gemini-flash-lite-latest",
system_prompt=(
"You are a helpful AI assistant specialized in automation testing "
"and software development. Be clear, concise, and always think step by step."
),
)
result = agent.invoke({"messages": [{"role": "user", "content": "Who are you?"}]})
print(result["messages"][-1].content)
Note the model string: "google_genai:gemini-flash-lite-latest". Provider and model in one string, so switching providers is a one-line change.
Four questions at once
invoke waits. ainvoke does not, and asyncio.gather runs a batch of them concurrently:
import asyncio, time
from langchain.agents import create_agent
agent = create_agent(
model="google_genai:gemini-flash-lite-latest",
system_prompt="You are an experienced software testing instructor. Explain each concept in 10 bullet points",
)
questions = [
{"id": 0, "topic": "LLM Eval", "prompt": "What is llm Eval?"},
{"id": 1, "topic": "Smoke Testing", "prompt": "What is smoke testing?"},
{"id": 2, "topic": "Sanity Testing", "prompt": "What is sanity testing?"},
{"id": 3, "topic": "Regression Testing", "prompt": "What is regression testing?"},
]
async def main():
start = time.perf_counter()
# asyncio.gather runs every agent.ainvoke() call CONCURRENTLY
results = await asyncio.gather(*[
agent.ainvoke({"messages": [{"role": "user", "content": q["prompt"]}]})
for q in questions
])
# results come back in the same order as the questions list
for q, result in zip(questions, results):
print(f"{q['id']}. {q['topic']}")
print(result["messages"][-1].content)
asyncio.run(main())
Results arrive in the order of the input list, which is why zip lines them back up. The use case named in class: summarising a hundred Jira tickets, where doing it one at a time is the difference between a minute and twenty.
Tools: the agent decides
A tool is a function plus its metadata plus a schema, and the docstring is the metadata: it is what the model reads to decide whether this tool is the right one.
from langchain.tools import tool
@tool
def qa_metric_calculator(expression: str) -> str:
"""Calculate a QA metric from an arithmetic expression.
Use this for pass rate, defect density, automation coverage, defect leakage
or execution time. Input must be plain arithmetic with no words, for example
'(438/500)*100' for a pass rate or '18/12' for defects per KLOC.
"""
result = _safe_eval(ast.parse(expression, mode="eval").body)
return f"{round(result, 2)}"
agent = create_agent(
model="google_genai:gemini-flash-lite-latest",
tools=[qa_metric_calculator],
system_prompt=(
"You are a QA metrics assistant for a test automation team. "
"Use the qa_metric_calculator tool whenever a number must be computed. "
"State the formula you used, then the result with its unit."
),
)
Three questions go in: a pass rate from 438 of 500, a defect density from 18 defects in 12 KLOC, and "what is the difference between smoke and sanity testing". The first two call the tool. The third does not, because there is no arithmetic in it, and that is the point of the demo: the agent chooses.
The repo version is safer than the obvious one. A calculator tool is usually written with eval(), and eval() on a string a model produced is arbitrary code execution. The version that landed walks the parsed expression tree with a fixed table of arithmetic operators, so nothing but maths can ever run. Worth copying that habit rather than the shortcut.
A second script gives the agent two tools, a DuckDuckGo search and a page fetch, and lets the system prompt steer which one gets used. Same lesson, more tools.
Structured output with Pydantic
Prose is hard to use. A schema turns the same call into typed objects:
from typing import Literal
from pydantic import BaseModel
from langchain.agents import create_agent
class TestCase(BaseModel):
title: str
description: str
steps: list[str]
expected_result: str
priority: Literal["High", "Medium", "Low"]
tags: list[str]
class TestCaseList(BaseModel):
# Tip: always wrap arrays inside an object (better compatibility, especially with Anthropic)
test_cases: list[TestCase]
agent = create_agent(
model="google_genai:gemini-flash-lite-latest",
system_prompt=("You are a senior automation testing engineer. Always return test cases "
"in the exact structured JSON format requested."),
response_format=TestCaseList,
)
result = agent.invoke({"messages": [{"role": "user", "content":
"Generate 2 detailed test cases for this user story: As a logged-in user, "
"I want to add items to my shopping cart so that I can purchase them later."}]})
for tc in result["structured_response"].test_cases:
print(tc.model_dump_json(indent=2))
response_format enforces the schema, and result["structured_response"] hands back real Python objects with no parsing. Wrap the list inside an object rather than returning a bare array: the comment in the repo says it plainly, compatibility is better that way, Anthropic especially.
The closing act: an agent that drives a browser
The last build is an independent agent with Playwright tools, and it logs into TTACart.
llm = ChatGroq(model=os.getenv("LLM_MODEL"), temperature=1)
SYSTEM_PROMPT = """
You are an automation testing agent that controls a real browser using Playwright.
Rules:
1. Always start by launching the browser before performing any action.
2. Perform the actions in the order they are requested.
3. Use clear CSS selectors to interact with elements on the page.
4. Take a screenshot before closing the browser to capture the final state.
5. At the end, report every step you took and whether the test PASSED or FAILED."""
TASK = """ Test the login functionality on https://app.thetestingacademy.com/playwright/ttacart/
1. Launch the browser
2. Navigate to the login page
3. Enter the username "standard_user"
4. Enter the password "tta_secret"
5. Click the login button
6. Close the browser
"""
async def main():
agent = create_agent(model=llm, tools=PLAYWRIGHT_TOOLS, system_prompt=SYSTEM_PROMPT)
# Playwright's async API needs ainvoke: every tool call runs on the same event loop
result = await agent.ainvoke({"messages": [{"role": "user", "content": TASK}]})
print("\n===== Steps taken =====")
for msg in result["messages"]:
if msg.type == "ai" and msg.tool_calls:
for call in msg.tool_calls:
print(f"-> {call['name']}({call['args']})")
elif msg.type == "tool":
print(f" {msg.content}")
asyncio.run(main())
Asked in class: how is this different from Playwright MCP? MCP is driven from your editor, by Copilot or Claude Code, with you in the chat. This is a standalone agent: it can be triggered by a Jira ticket or a schedule and run without anyone watching. Same browser, different owner. The instructor's own team runs agents autonomously through long post-login flows this way.
On speed: Groq was noticeably slow on this task and DeepSeek was faster, which is a model choice rather than a framework one. And the honest caveat given about the whole approach: agent-driven execution is slower, costs tokens, and is not yet what you would put in a regression suite. The prediction attached to it was six months.
Tasks and announcements
Next
- A mandatory live test covering everything through this LangChain material, replacing the next session.
- The notes from this class are being shared, and the chapter code is in the repository.
- Hackathon winners: funds are going out over the next day, once UPI details are in.
Coming up in the course
Memory, multi-agent setups, MCP and RAG, all on top of what was built today.
Offered as extra, outside the course
LangGraph and LangSmith, for anyone who wants the house and the neighbour rather than just the bricks.
The same material exists in TypeScript on the practice site, as the LangChain agent guide, ending in the same kind of Playwright agent. Useful if your team is a JavaScript shop, or as a second pass over the same ideas in a language you already know.
Repository: AITesterBlueprint3x received chapter_17_LangChain/ with the chapter scripts 001 to 010 (hello world on Groq and Gemini, streaming, system prompt, parallel calls, one tool, multiple tools, structured output, and the Playwright orchestrator on both Groq and DeepSeek), plus playwright_tools.py holding the browser tools and an HTML notes file.