The Testing Academy

AI agents for testers, study guide

LangChain with TypeScript

Build AI agents for test automation in TypeScript. Seven core chapters, four advanced ones (memory, multi-agent workflows, RAG over your test docs, MCP servers) and a final project: a Playwright agent that runs a login test from a plain-English task.

Think of it this way
LangChain = build with bricksmodels, prompts, tools, agents
LangGraph = connect the bricks into workflowsstate, branching, loops, approvals
LangSmith = inspect, evaluate and improvetraces, tests, monitoring

What is LangChain?

LangChain is a framework for building applications powered by large language models. It gives you the building blocks (models, prompts, tools, memory, retrieval) and a single createAgent() function that wires them into an agent.

A LangChain agent is an LLM-powered system that can understand user input, reason, and generate responses. Give it tools and it can also take action: run a calculation, call an API, fetch a page, generate test cases in a fixed format.

Agent = Model + System prompt + Tools + a loop. The model reads the messages, decides whether to answer or call a tool, and repeats until the task is done. Everything in this guide is a variation of that one idea.

LangChain vs LangGraph vs LangSmith

The LangChain ecosystem: purpose, features and differences. This course lives mostly in the first column, and uses LangGraph for memory (chapter 08) and multi-agent workflows (chapter 09).

LangChain

Build LLM applications easily

A framework for developing LLM-powered applications. It provides the building blocks to chain components together and create end-to-end apps.

Key capabilities

  • Prompt templates and few-shot examples
  • LLM integrations (OpenAI, Anthropic, Gemini, Ollama)
  • Chains: LLMChain, SequenceChain, RouterChain
  • Agents: tool calling, MRKL, ReAct agents
  • Memory: conversation buffer, summary, vector store
  • Retrieval: document loaders, text splitters, vector stores
  • Tools and functions: integrate external tools and APIs
  • Easy to start with pre-built components

Best for

Quickly building LLM apps, chatbots, RAG pipelines and AI agents with minimal orchestration complexity.

Example use cases

  • Chatbots
  • RAG apps
  • Tool-using agents
  • Q&A over documents

LangGraph

Build stateful, multi-agent workflows

Builds on LangChain to help you create stateful, cyclic and deterministic workflows (graphs) for complex multi-step AI applications.

Key capabilities

  • Graph-based workflow orchestration (nodes and edges)
  • Stateful execution with persistent memory
  • Cycles, loops, branching, conditional routing
  • Human-in-the-loop and interrupts
  • Multi-agent collaboration patterns
  • Durable execution and checkpointing
  • Streaming and real-time updates
  • Works with LangChain components inside nodes

Best for

Complex, production-ready AI workflows that need control, state management, branching and cycles.

Example use cases

  • Multi-step AI workflows
  • Multi-agent systems
  • Human-in-the-loop
  • Long-running stateful apps

LangSmith

Debug, test, evaluate and monitor

A developer platform for observability and evaluation of LLM applications. It helps you debug, test, evaluate and monitor your chains and agents.

Key capabilities

  • Tracing: see every LLM call, chain, tool and sub-step
  • Debugging: inspect inputs, outputs, prompts, errors
  • Evaluation: test datasets, expected outputs, scoring
  • Experimentation: compare prompts, models, strategies
  • Monitoring: production logs, latency, tokens, cost
  • Datasets: manage test sets and golden datasets
  • Collaboration: share traces, feedback, annotations
  • Integrates with LangChain and LangGraph

Best for

Observability, debugging, evaluation and monitoring of LLM applications in development and production.

Example use cases

  • Debug failing agents
  • Evaluate outputs
  • Compare models and prompts
  • Monitor production
LangChainLangGraphLangSmith
Primary focusBuilding blocks and components for LLM applicationsOrchestrating complex, stateful workflowsObservability, evaluation and monitoring
What it isFramework / libraryWorkflow orchestration frameworkDeveloper platform (SaaS / cloud)
Handles state?Limited (via memory components)Yes, built-in state management and persistenceN/A (observability layer)
Core strengthEasy to build LLM apps fastControl flow, branching, cycles, multi-agentVisibility, testing, evaluation, monitoring
Works withLLMs, prompts, chains, tools, memory, retrieversLangChain components (inside nodes)LangChain and LangGraph applications

Use LangChain when

  • You want to quickly build LLM apps, chains, RAG pipelines or simple agents.
  • Great for prototypes and straightforward use cases.

Use LangGraph when

  • Your application has multiple steps, decisions, loops or human approvals, or needs robust stateful workflows.
  • Great for multi-agent systems and enterprise workflows.

Use LangSmith when

  • Throughout the lifecycle: debug, evaluate, test and monitor your LLM apps in development and production.
  • Essential for quality, reliability and iteration.

Setup: Node.js and .env

Five minutes, once. Every chapter after this is a single file you run with npx tsx. Chapters 08 to 12 each add one or two packages, listed where they are used.

  1. Create a project

    terminal
    mkdir langchain-typescript && cd langchain-typescript
    npm init -y
  2. Install LangChain, the providers and the TypeScript runner

    terminal
    npm install langchain @langchain/core @langchain/google-genai @langchain/anthropic @langchain/ollama @langchain/langgraph dotenv zod axios
    npm install -D typescript tsx @types/node

    langchain gives you createAgent and tool. Each model provider ships as its own package. zod defines tool inputs and structured output, dotenv loads your keys, tsx runs TypeScript files directly, axios is used by the web tools in chapter 06.

  3. Create a .env file for your API keys

    Keys live in .env, never in code. Get a free Gemini key from Google AI Studio. Anthropic keys come from console.anthropic.com, LangSmith keys from smith.langchain.com.

    .env
    # Cloud models
    GOOGLE_API_KEY=your_gemini_api_key_here
    ANTHROPIC_API_KEY=your_anthropic_key_here
    
    # Optional: LangSmith tracing (see every call your agent makes)
    # LANGSMITH_TRACING=true
    # LANGSMITH_API_KEY=your_langsmith_key_here
    # LANGSMITH_PROJECT=langchain-typescript-course
    .gitignore
    .env
    node_modules/
    memory.db
    screenshots/
  4. Optional: local models with Ollama (offline, free)

    terminal
    # install from https://ollama.com, then pull a small chat model
    ollama pull gemma2:2b
    # chapter 10 (RAG) uses a local embedding model
    ollama pull nomic-embed-text
  5. Project layout and how to run a chapter

    project structure
    langchain-typescript/
    ├── .env                          # API keys (never commit)
    ├── .gitignore
    ├── package.json
    ├── documents/
    │   └── login-requirements.txt    # chapter 10
    ├── mcp-servers/
    │   └── qa-utils-server.ts        # chapter 11
    └── lectures/
        ├── 01-first-agent.ts
        ├── 02-invoke-stream.ts
        ├── 03-system-prompt.ts
        ├── 04-parallel-calls.ts
        ├── 05-first-tool.ts
        ├── 06-multiple-tools.ts
        ├── 07-structured-output.ts
        ├── 08-memory/                # 5 small files
        ├── 09-multi-agent.ts
        ├── 10-rag-test-docs.ts
        ├── 11-mcp-agent.ts
        └── 12-playwright/            # playwright-tools.ts, playwright-agent.ts
    terminal
    npx tsx lectures/01-first-agent.ts

Model strings: one line to switch providers

Every agent takes a model as "provider:model". Swap the string, keep the rest of the code.

Model stringNeedsWhy
google-genai:gemini-flash-lite-latestGOOGLE_API_KEYFast, cheap, free tier. Used in chapters 01 to 06.
anthropic:claude-haiku-4-5ANTHROPIC_API_KEYStrong at tool use and structured output. Used from chapter 07 onward.
ollama:gemma2:2bOllama running locallyRuns on your laptop, no internet, no key.
01

Creating your first LangChain agent

Set up an agent with createAgent(), send it a message, print the reply.

User "What is the capital of India?" LangChain Agent LLM (Model) model reply comes back as the last message
You send a list of messages (role + content). The agent forwards them to the model and gives you back the full conversation, with the reply at the end.
lectures/01-first-agent.ts
import "dotenv/config";                  // load environment variables
import { createAgent } from "langchain";  // import createAgent

async function main() {
  console.log("Creating your first LangChain agent...\n");

  const agent = createAgent({
    model: "google-genai:gemini-flash-lite-latest",
    // model: "ollama:gemma2:2b",         // use local model (offline)
  });

  const result = await agent.invoke({
    messages: [
      {
        role: "user",
        content: "What is the capital of India?",
      },
    ],
  });

  // Print only the final response (clean output)
  const finalMessage = result.messages.at(-1);
  console.log("Agent Response:");
  console.log(finalMessage?.content);
}

main().catch(console.error);

Run: npx tsx lectures/01-first-agent.ts

Output
Creating your first LangChain agent...

Agent Response:
The capital of India is New Delhi.
dotenv/config

Loads variables from the .env file, so the API key is in the environment and not in your code.

createAgent()

Builds the agent. The model string picks the provider. Comment one line, uncomment the other, to switch cloud and local.

.messages.at(-1)

The conversation history is an array. .at(-1) is the agent's reply; ?.content is safe access in case it is undefined.

Key takeaways

  • Set up a basic agent with createAgent()
  • Pass messages to the agent as an array of role + content
  • Extract and display the final response
  • Use .env for API keys
  • Switch between cloud and local models with one line

Exercises

  1. Set up your .env file with your Gemini API key and run the file.
  2. Run the code with both Gemini and Ollama models and compare speed and quality.
  3. Change the question to something from your own project and observe the results.
02

invoke vs stream

Two ways to get an answer, and the system + user message pattern that makes agents behave.

At a glance

invoke

  • Get the complete response at once
  • Simple and easy
  • Best for scripts and batch processing
VS

stream

  • Get the response token by token
  • Feels faster (real-time)
  • Best for chat UIs and live tools

A system message sets the behaviour and expertise of the agent. The user message is the actual question. This pair is the foundation of every good agent.

lectures/02-invoke-stream.ts
import "dotenv/config";
import { createAgent } from "langchain";

async function main() {
  console.log("Understanding invoke and stream with multiple roles...\n");

  const agent = createAgent({
    model: "google-genai:gemini-flash-lite-latest",
  });

  const input = {
    messages: [
      {
        role: "system",
        content: "You are a software testing instructor.",
      },
      {
        role: "user",
        content: "Explain automation testing in two sentences.",
      },
    ],
  };

  // 1) invoke: wait for the full answer
  const result = await agent.invoke(input);
  console.log("Agent Response:");
  console.log(result.messages.at(-1)?.content);

  // 2) stream: print each token as it is generated
  console.log("\nStreaming Response:");
  for await (const [token] of await agent.stream(
    {
      messages: [
        { role: "system", content: "You are a software testing instructor." },
        { role: "user", content: "Explain smoke testing in two sentences." },
      ],
    },
    { streamMode: "messages" }
  )) {
    process.stdout.write(token.text ?? "");
  }
  console.log();
}

main().catch(console.error);

Run: npx tsx lectures/02-invoke-stream.ts

Output
Agent Response:
Automation testing is the process of using tools and scripts to test software
applications automatically. It helps find bugs faster, saves time, and improves
software quality.

Streaming Response:
Smoke testing is a quick check that the most important features of a build
work... (appears word by word)
Featureinvokestream
Response styleFull response at onceToken by token
Use caseScripts, batch processingChat UIs, live demos
Speed perceptionWaits until the endFeels faster (typing effect)
Memory usageLower for short responsesSlightly higher
ComplexityVery simpleSlightly more code

streamMode: "messages" yields [token, metadata] pairs for LLM tokens. Use streamMode: "updates" instead when you want one event per agent step (useful once your agent calls tools, as in chapter 12).

In automation testing agents we often use both: stream for developer-facing tools where someone is watching, invoke for backend workflows such as CI jobs.

Key takeaways

  • invoke() gets the complete response (most common for starters)
  • stream() gets the response in real time, token by token
  • Always use a system message to define the agent's role
  • Proper message structure (role + content) is the foundation of good agents
  • These two methods are used in 95% of LangChain implementations

Exercises

  1. Ask the same question with invoke and with stream. Which one fits a CI pipeline, and which a chat UI?
  2. Change the system message to "You are a senior Playwright engineer" and compare the answer.
  3. Loop over result.messages and print each message's type and content to see the full history.
03

System prompts: giving your agent personality and instructions

A system prompt is a job description you hand the agent before the conversation starts.

At a glance

User message

  • What the user asks
  • The actual question or request
  • Changes every time
VS

System prompt

  • Defines who the agent is
  • Sets behaviour, tone and rules
  • Stays constant
  • Huge impact on response quality

It tells the agent who it is, how it should behave, what tone to use and what rules to follow. Think of hiring: system prompt = job description, user message = task given, agent response = work delivered.

lectures/03-system-prompt.ts
import "dotenv/config";
import { createAgent } from "langchain";

async function main() {
  const agent = createAgent({
    model: "google-genai:gemini-flash-lite-latest",
    systemPrompt: `You are a helpful AI assistant specialized in automation testing
      and software development. Be clear, concise, and always think step by step.`,
  });

  const result = await agent.invoke({
    messages: [{ role: "user", content: "Who are you?" }],
  });

  console.log("Agent Response:");
  console.log(result.messages.at(-1)?.content);
}

main().catch(console.error);

Run: npx tsx lectures/03-system-prompt.ts

Output
Agent Response:
I am an AI assistant specialized in automation testing and software development.
I help with test automation, frameworks, code, and best practices.

Without a good system prompt (for example systemPrompt: "You are AI assistant.") the response is generic, less useful and not focused on your domain. A small change in the prompt is a big improvement in output quality.

Best practices for system prompts

RuleExample
Define the role clearly"You are a senior automation testing engineer..."
Set behaviour rules"Be clear, concise, think step by step, always use tools when needed."
Name the domain expertiseMention Playwright, Appium, test frameworks, API testing.
Give output guidelinesAsk for structured, professional or JSON output when needed.
Keep it conciseVery long prompts can reduce performance.

Real testing example

an improved system prompt
systemPrompt: `You are a senior automation testing engineer with 10+ years of experience.
  You specialize in Playwright, Appium, and API testing.
  Always think step by step and provide practical, production-ready solutions.`

Result: more professional, focused and practical responses.

Key takeaways

  • System prompts define your agent's identity and behaviour
  • They are one of the most powerful tools for controlling output quality
  • Small changes in the prompt can create big improvements

Exercises

  1. Create 3 different system prompts: general assistant, testing expert, code reviewer.
  2. Compare the responses for the same question.
  3. Try making a prompt specialised in generating Playwright locators.
04

Parallel agent calls with Promise.all()

Work smarter, not slower. Run many queries at the same time.

At a glance

Sequential calls

  • One by one execution
  • Takes more time
  • Blocks the next call
  • Not efficient for batch work
VS

Parallel calls (Promise.all)

  • All at the same time
  • Much faster
  • Independent execution
  • Perfect for batch processing
Sequential Smoke Sanity Regression time: 3 units Parallel Smoke Sanity Regression time: 1 unit all three start together
Generate test cases for several user stories, explain several concepts, or analyse several bugs: running them one by one is slow. In parallel you get all responses much faster.
lectures/04-parallel-calls.ts
import "dotenv/config";
import { createAgent } from "langchain";

interface AgentQuestion {
  id: number;
  topic: string;
  prompt: string;
}

async function main() {
  console.log("Creating LangChain agent...\n");

  // one agent with a strong system prompt, shared across all calls
  const agent = createAgent({
    model: "google-genai:gemini-flash-lite-latest",
    systemPrompt:
      "You are an experienced software testing instructor. " +
      "Explain each concept in exactly two sentences using simple language.",
  });

  const questions: AgentQuestion[] = [
    { id: 1, topic: "Smoke Testing", prompt: "What is smoke testing?" },
    { id: 2, topic: "Sanity Testing", prompt: "What is sanity testing?" },
    { id: 3, topic: "Regression Testing", prompt: "What is regression testing?" },
    // ... more questions
  ];

  const start = Date.now();

  // Promise.all() runs all agent.invoke() calls CONCURRENTLY
  const results = await Promise.all(
    questions.map((q) =>
      agent.invoke({ messages: [{ role: "user", content: q.prompt }] })
    )
  );

  // Process results in order using the original array
  results.forEach((result, index) => {
    const q = questions[index];
    console.log(`${q.id}. ${q.topic}`);
    console.log(result.messages.at(-1)?.content);
    console.log("-".repeat(70));
  });

  console.log(`Finished ${questions.length} calls in ${(Date.now() - start) / 1000}s`);
}

main().catch(console.error);

Run: npx tsx lectures/04-parallel-calls.ts

Output
1. Smoke Testing
Smoke testing is a type of testing to verify that the basic functionality of an
application works as expected. It is done at the build level to ensure stability.
----------------------------------------------------------------------
2. Sanity Testing
Sanity testing is performed to check that recent changes or bug fixes are working
as expected. It is a narrower and more focused type of testing.
----------------------------------------------------------------------
3. Regression Testing
...
Finished 3 calls in 2.1s

Shortcut: agent.batch([input1, input2, input3]) runs an array of inputs in parallel in one call. Promise.all is the pattern to know when you mix agents with other async work (Playwright, HTTP calls, databases).

Key concepts

  • Promise.all() runs multiple promises in parallel
  • Each agent.invoke() is independent but uses the same agent config
  • The system prompt is shared across all calls
  • Results arrive in the order of the original array, so the index lines them up

Real testing use cases: generate test cases for multiple user stories at once, explain multiple testing concepts, analyse multiple bugs in parallel.

Exercises

  1. Add 3 more testing concepts to the questions array.
  2. Try running 8 to 10 questions at once.
  3. Measure the time with and without Promise.all().
05

Creating your first tool and agent

Give your agent a tool and watch it do real work instead of guessing.

At a glance

Without tools

  • Agent can only reply using its knowledge
  • Cannot perform calculations or external tasks
  • Limited to text responses
VS

With tools

  • Agent can call tools to do real work
  • Gets accurate and dynamic results
  • Much more powerful and useful

A tool = function + metadata + input schema:

implementation

The function that runs when the tool is called.

name + description

Metadata that helps the agent understand when and how to use the tool.

schema (Zod)

The expected input shape, validated before your function runs.

lectures/05-first-tool.ts
import "dotenv/config";                        // load environment variables (.env)
import { createAgent, tool } from "langchain"; // LangChain core
import * as z from "zod";                      // for schema validation

const calculatorTool = tool(
  ({ expression }) => {
    try {
      return String(eval(expression));
      // Demo only: use math.js or any safe math library in real projects
    } catch (e) {
      return "Error: Invalid mathematical expression";
    }
  },
  {
    name: "calculator",
    description:
      "Useful for performing mathematical calculations. " +
      "Input should be a valid math expression like '25 * 4' or '100 / 2'.",
    schema: z.object({
      expression: z.string().describe("The mathematical expression to evaluate"),
    }),
  }
);

async function main() {
  const agent = createAgent({
    model: "google-genai:gemini-flash-lite-latest",
    tools: [calculatorTool],
    systemPrompt: `You are a helpful automation testing assistant.
      Use the calculator tool whenever math is required.
      Always be helpful and accurate.`,
  });

  const queries = [
    "What is 15 multiplied by 24?",   // Math  -> agent uses calculator
    "Calculate 100 divided by 4",     // Math  -> agent uses calculator
    "What is the capital of India?",  // General -> should NOT use calculator
  ];

  for (const q of queries) {
    console.log(`\nQuestion: ${q}`);
    const result = await agent.invoke({
      messages: [{ role: "user", content: q }],
    });
    console.log("Agent:", result.messages.at(-1)?.content);
  }
}

main().catch(console.error);

Run: npx tsx lectures/05-first-tool.ts

Output
Question: What is 15 multiplied by 24?
Agent: The result of 15 multiplied by 24 is 360.

Question: Calculate 100 divided by 4
Agent: The result of 100 divided by 4 is 25.

Question: What is the capital of India?
Agent: The capital of India is New Delhi.

The agent intelligently decides whether to use the tool or answer from its knowledge. To watch that decision, print the history: result.messages.forEach(m => console.log(m.getType(), m.content)). A tool message appears only for the math questions.

Key takeaways

  • Tools extend agent capabilities beyond LLM knowledge
  • A tool = function + metadata + input schema
  • The agent decides when to call the tool
  • Clear descriptions and schemas help the agent use tools correctly
  • This is the foundation for building powerful automation testing agents

Exercises

  1. Add a new tool (for example a DateTime tool or a string manipulation tool).
  2. Modify the calculator to handle advanced operations.
  3. Ask the agent more questions and observe tool usage.
06

Creating multiple tools and agents

More tools, more power. The agent picks the right one for each request.

At a glance

Single-tool agent

  • Can do one type of task
  • Limited capabilities
  • Good for learning basics
VS

Multi-tool agent

  • Can do multiple types of tasks
  • More capable and flexible
  • Ready for real-world automation
lectures/06-multiple-tools.ts
import "dotenv/config";                        // loads .env variables (API keys)
import { createAgent, tool } from "langchain"; // helps create custom tools
import * as z from "zod";                      // for input validation
import axios from "axios";                     // to call external APIs

// Tool 1: Search Tool - uses DuckDuckGo and returns the page title
const searchTool = tool(
  async ({ query }) => {
    try {
      const res = await axios.get(
        `https://duckduckgo.com/?q=${encodeURIComponent(query)}`
      );
      return res.data.split("<title>")[1].split("</title>")[0];
    } catch (e) {
      return "Error performing search";
    }
  },
  {
    name: "search_web",
    description: "Search the web for a query and return the results page title",
    schema: z.object({
      query: z.string().describe("The search query"),
    }),
  }
);

// Tool 2: Web Content Tool - fetches a URL and returns plain text (first 1000 chars)
const webContentTool = tool(
  async ({ url }) => {
    try {
      const res = await axios.get(url);
      const text = res.data.replace(/<[^>]*>/g, "");
      return text.substring(0, 1000);
    } catch (e) {
      return "Error fetching web content";
    }
  },
  {
    name: "get_web_content",
    description: "Fetch the content of a web page from a URL as plain text",
    schema: z.object({
      url: z.string().describe("The full URL to fetch"),
    }),
  }
);

async function main() {
  const agent = createAgent({
    model: "google-genai:gemini-flash-lite-latest",
    tools: [searchTool, webContentTool],   // we pass both tools in the tools array
    systemPrompt: `You are a helpful automation testing assistant.
      Use the available tools to search the web and fetch content when needed.`,
  });

  const queries = [
    "What is Playwright?",                          // general info (may use search)
    "Search for Appium official website",           // needs search
    "Get content from https://playwright.dev docs", // needs web content tool
  ];

  for (const q of queries) {
    const result = await agent.invoke({
      messages: [{ role: "user", content: q }],
    });
    console.log(`\nQ: ${q}`);
    console.log(`Agent: ${result.messages.at(-1)?.content}`);
  }
}

main().catch(console.error);

Run: npx tsx lectures/06-multiple-tools.ts

Output
Q: What is Playwright?
Agent: Playwright is an open-source framework for end-to-end testing of web apps...

Q: Search for Appium official website
Agent: The search results page title is "Appium official website at DuckDuckGo"...

Q: Get content from https://playwright.dev docs
Agent: The page introduces Playwright, its installation via npm init playwright...

Tools available to the agent

searchTool: search the web and get the page title
webContentTool: get content from any URL

Why multiple tools?

Expand agent capabilities, handle real-world automation tasks, combine multiple data sources, make the agent more powerful and useful.

Key takeaways

  • Build each tool with a clear input schema (Zod)
  • Give good descriptions so the agent knows when to use each tool
  • Pass all tools in the tools: [] array
  • The agent will choose the right tool automatically
  • This is the foundation for advanced agents

Exercises

  1. Add a Date/Time tool and test it.
  2. Build two QA tools: generate_test_data(type) returning a test email, name or phone number, and suggest_locator(element) returning a data-testid locator.
  3. Ask 5 different queries and observe which tool is used.
  4. Build a tool that calculates a pass percentage from passed and total tests.
07

Structured output with Zod schemas

Get clean, validated JSON from the LLM and generate test cases automatically.

At a glance

Without structured output

  • LLM returns plain text
  • Difficult to parse
  • Inconsistent format
  • Manual effort to extract information
VS

With structured output (Zod schema)

  • Clean, validated JSON
  • Consistent and reliable
  • Easy to use in automation frameworks
  • Saves time and reduces errors
lectures/07-structured-output.ts
import "dotenv/config";
import { createAgent } from "langchain";
import * as z from "zod";

// Anthropic requires the top-level response schema to be an object,
// so wrap the array of test cases inside an object.
const TestCaseSchema = z.object({
  testCases: z.array(
    z.object({
      title: z.string(),
      description: z.string(),
      steps: z.array(z.string()),
      expectedResult: z.string(),
      priority: z.enum(["High", "Medium", "Low"]),
      tags: z.array(z.string()),
    })
  ),
});

async function main() {
  const agent = createAgent({
    model: "anthropic:claude-haiku-4-5",
    // model: "ollama:llama3.2",             // optional
    systemPrompt: `You are a senior automation testing engineer.
      Always return test cases in the exact structured JSON format requested.
      Be detailed and professional.`,
    responseFormat: TestCaseSchema,          // enforce the Zod schema on the output
  });

  const userStory = `As a logged-in user, I want to add items to my shopping cart
    so that I can purchase them later.`;

  const result = await agent.invoke({
    messages: [
      {
        role: "user",
        content: `Generate 2 detailed test cases for this user story: ${userStory}`,
      },
    ],
  });

  console.log("=== Structured Output ===");
  console.dir(result.structuredResponse, { depth: null });

  const testCases = result.structuredResponse.testCases;
  console.log(`\nSuccessfully generated ${testCases.length} test cases!`);
}

main().catch(console.error);

Run: npx tsx lectures/07-structured-output.ts

Output
=== Structured Output ===
{
  testCases: [
    {
      title: 'Add Item to Cart - Successful Flow',
      description: 'Verify that a user can add an item to the shopping cart successfully.',
      steps: [ 'Open the application', 'Login with valid credentials', 'Select a product', 'Add to cart' ],
      expectedResult: 'Item should be added to the cart and cart count should increase.',
      priority: 'High',
      tags: [ 'cart', 'functional', 'smoke' ]
    },
    ...
  ]
}

Successfully generated 2 test cases!

The schema defines

title, description, steps (array), expected result, priority (enum), tags (array).

responseFormat

Tells LangChain to enforce the Zod schema on the final output. result.structuredResponse gives you fully typed, validated JavaScript objects with no manual parsing.

Powerful use cases in automation testing

Test cases from user stories

Story in, reviewed and tagged test cases out.

Test data in structured format

Users, orders, payloads that match your fixtures.

Locators or selectors

Recommended locator, alternatives, strategy, confidence.

API request and response schemas

Contracts you can validate against.

Bug reports, logs and reports

Consistent fields for Jira and dashboards.

Key takeaways

  • Use Zod to define the exact output structure
  • Wrap arrays in an object for compatibility
  • Pass the schema via responseFormat in createAgent()
  • Get reliable, ready-to-use JSON output
  • Perfect for automation, APIs and frameworks

Exercises

  1. Generate 5 test cases for a new user story.
  2. Modify the schema: add preconditions and test data.
  3. Build a LocatorSchema (element description, recommended locator, alternative locators, confidence 0 to 100, strategy enum) and generate locators for a login form.
  4. Use the output directly in a Playwright test.
08

Memory: making your agent remember

A checkpointer plus a thread id turns one-off questions into a real conversation.

At a glance

Without memory

  • Every invoke starts from zero
  • "What test cases should I write?" has no subject
  • You must resend the whole history yourself
VS

With memory

  • The agent recalls earlier turns
  • Follow-up questions just work
  • LangChain stores and replays the history for you

Memory needs two things: a checkpointer (where the conversation is saved) and a thread id (which conversation to load). Different thread ids are different conversations, exactly like separate chat tabs.

Turn 1 "I need to test login..." Turn 2 "What test cases should I write?" Agent thread_id: test-session-1 save load Checkpointer MemorySaver (RAM) SqliteSaver (memory.db)
Both turns run against the same thread. The agent loads the saved history before answering and writes the new messages back after.

Part 1: the problem, an agent without memory

lectures/08-memory/agent-without-memory.ts
import "dotenv/config";
import { createAgent } from "langchain";

async function main() {
  const agent = createAgent({
    model: "anthropic:claude-haiku-4-5",
    systemPrompt: "You are a helpful automation testing assistant.",
  });

  console.log("=== Turn 1 ===");
  let result = await agent.invoke({
    messages: [
      { role: "user", content: "I need to test the login functionality using Playwright typescript." },
    ],
  });
  console.log("Agent:", result.messages.at(-1)?.content);

  console.log("\n=== Turn 2 (No Memory) ===");
  result = await agent.invoke({
    messages: [{ role: "user", content: "What test cases should I write?" }],
  });
  console.log("Agent:", result.messages.at(-1)?.content);
}

main().catch(console.error);
Output: turn 2 has lost the context
=== Turn 1 ===
Agent: Great, for Playwright with TypeScript you'll want to start with...

=== Turn 2 (No Memory) ===
Agent: I'd be happy to help! Could you tell me what feature or application
you're testing? Test cases depend on the functionality you have in mind.

Part 2: add a checkpointer and a thread id

MemorySaver keeps the conversation in RAM. It lasts as long as the process runs, which is perfect for a chat session or a script with several turns.

lectures/08-memory/agent-with-memory.ts
import "dotenv/config";
import { createAgent } from "langchain";
import { MemorySaver } from "@langchain/langgraph";

async function main() {
  const checkpointer = new MemorySaver();

  const agent = createAgent({
    model: "anthropic:claude-haiku-4-5",
    systemPrompt: "You are a helpful automation testing assistant.",
    checkpointer,                                  // 1) where the history is stored
  });

  const threadId = "test-session-1";               // 2) which conversation to use

  console.log("=== Turn 1 ===");
  let result = await agent.invoke(
    { messages: [{ role: "user", content: "I need to test the login functionality using Playwright typescript." }] },
    { configurable: { thread_id: threadId } }
  );
  console.log("Agent:", result.messages.at(-1)?.content);

  console.log("\n=== Turn 2 (with Memory) ===");
  result = await agent.invoke(
    { messages: [{ role: "user", content: "What test cases should I write?" }] },
    { configurable: { thread_id: threadId } }
  );
  console.log("Agent:", result.messages.at(-1)?.content);
}

main().catch(console.error);

Run: npx tsx lectures/08-memory/agent-with-memory.ts

Output: turn 2 now knows we are talking about Playwright login tests
=== Turn 1 ===
Agent: Great, for Playwright with TypeScript you'll want to start with...

=== Turn 2 (with Memory) ===
Agent: For the login functionality you mentioned, I'd write these Playwright
test cases: valid login, invalid password, empty fields, account lockout
after 3 attempts, session persistence and logout...

Part 3: persistent memory that survives a restart

MemorySaver is gone when the process exits. SqliteSaver writes to a file, so the agent still remembers tomorrow.

terminal
npm install @langchain/langgraph-checkpoint-sqlite
lectures/08-memory/create-memory.ts (session 1: three turns, saved to disk)
import "dotenv/config";
import { createAgent } from "langchain";
import { SqliteSaver } from "@langchain/langgraph-checkpoint-sqlite";
import { ChatAnthropic } from "@langchain/anthropic";

async function main() {
  console.log("Step 1: Creating a new session with memory...");

  const checkpointer = SqliteSaver.fromConnString("./memory.db");

  const model = new ChatAnthropic({
    model: "claude-haiku-4-5",
    temperature: 0.9,         // 0 to 1 - use 0.1 to 0.3 for precise, factual answers
    maxTokens: 1000,
  });

  const agent = createAgent({
    model,
    systemPrompt: `You are a helpful automation testing assistant.
      Always remember the previous conversation and provide relevant suggestions.`,
    checkpointer,
  });

  const threadId = "demo-persistent-memory";

  const conversation = [
    "I need to test the login functionality using Playwright typescript.",
    "What test cases should I write?",
    "Can you provide me with a sample test case for the login functionality?",
  ];

  for (const [index, message] of conversation.entries()) {
    console.log(`\n=== Turn ${index + 1} ===`);
    console.log("You:", message);

    const result = await agent.invoke(
      { messages: [{ role: "user", content: message }] },
      { configurable: { thread_id: threadId } }
    );

    console.log("Agent:", result.messages.at(-1)?.content);
  }
}

main().catch(console.error);
lectures/08-memory/continue-memory.ts (session 2: a new process, same thread)
import "dotenv/config";
import { createAgent } from "langchain";
import { SqliteSaver } from "@langchain/langgraph-checkpoint-sqlite";
import { ChatAnthropic } from "@langchain/anthropic";

async function main() {
  console.log("Step 2: Continuing from the previous session with memory...");

  const checkpointer = SqliteSaver.fromConnString("./memory.db");

  const model = new ChatAnthropic({
    model: "claude-haiku-4-5",
    temperature: 0.1,
    maxTokens: 1000,
  });

  const agent = createAgent({
    model,
    systemPrompt: `You are a helpful automation testing assistant.
      Always remember the previous conversation and provide relevant suggestions.`,
    checkpointer,
  });

  const threadId = "demo-persistent-memory";   // same thread as session 1

  const followUp =
    "What were we discussing earlier? Can you summarize the previous conversation and provide the next steps?";

  console.log("You:", followUp);

  const result = await agent.invoke(
    { messages: [{ role: "user", content: followUp }] },
    { configurable: { thread_id: threadId } }
  );

  console.log("\nAgent remembers:\n", result.messages.at(-1)?.content);
  console.log("\nMemory is persistent across sessions!\n");
}

main().catch(console.error);

Run both, in order: npx tsx lectures/08-memory/create-memory.ts  then  npx tsx lectures/08-memory/continue-memory.ts

Output of session 2 (a brand new process)
Step 2: Continuing from the previous session with memory...
You: What were we discussing earlier? Can you summarize the previous conversation...

Agent remembers:
Earlier we discussed testing the login functionality with Playwright and
TypeScript. I suggested test cases covering valid login, invalid credentials,
empty fields and account lockout, then shared a sample spec file.
Next steps: add the lockout test, move credentials into fixtures, and wire
the spec into your CI pipeline.

Memory is persistent across sessions!

Part 4: clearing a thread

Memory grows with every turn, and every turn is resent to the model. Clear a thread when a test session is over.

lectures/08-memory/clear-memory.ts
import { SqliteSaver } from "@langchain/langgraph-checkpoint-sqlite";

async function clear() {
  const checkpointer = SqliteSaver.fromConnString("./memory.db");

  await checkpointer.deleteThread("demo-persistent-memory");

  console.log("Memory cleared for thread: demo-persistent-memory");
}

clear().catch(console.error);
CheckpointerStored inSurvives restart?Use it for
MemorySaverRAMNoDemos, single scripts, one chat session
SqliteSavermemory.db fileYesLocal tools you reopen, long-running test sessions
Postgres / Redis saversA real databaseYesProduction, multiple users, shared threads

Two traps: forgetting configurable: { thread_id } on an invoke silently starts a fresh conversation, and letting one thread grow forever makes every call slower and more expensive. Use one thread per test session, clear it when done.

Key takeaways

  • Memory = a checkpointer plus a thread_id
  • MemorySaver lives in RAM, SqliteSaver writes to a file
  • Pass the thread id on every invoke of that conversation
  • Different thread ids are completely separate conversations
  • Delete threads you no longer need to control cost and latency

Exercises

  1. Run agent-without-memory.ts and agent-with-memory.ts back to back and compare turn 2.
  2. Run two threads in the same script, "login-tests" and "api-tests", and confirm they do not leak into each other.
  3. Build a small terminal chat loop with readline: read a line, invoke with a fixed thread id, print the reply, repeat until the user types exit.
  4. Run create-memory.ts, stop the process, then run continue-memory.ts and check the summary is accurate. Then clear the thread and run it again to see the difference.
  5. Give the memory agent the calculator tool from chapter 05 and ask "add 20 to the number I gave you earlier" to see memory and tools working together.
09

Multi-agent workflows

One agent writes the test cases, another reviews them, a supervisor decides who runs next.

At a glance

Single agent

  • One prompt does everything
  • Quality drops as the task grows
  • Hard to see which step went wrong
VS

Multi-agent workflow

  • Each agent has one job and one prompt
  • A supervisor routes the work
  • Every step is visible and testable

This is where LangGraph comes in. Three pieces: State (shared memory every node reads and writes), Nodes (the agents, plain async functions), and Edges (who runs next, fixed or conditional).

START Supervisor decides the next step no test cases yet no review yet Generator agent writes 3 test cases Reviewer agent improves them END both done
The supervisor is called after every worker. It looks at the state and picks the next node, or ends the run.
lectures/09-multi-agent.ts
import "dotenv/config";
import { Annotation, END, START, StateGraph } from "@langchain/langgraph";
import { ChatAnthropic } from "@langchain/anthropic";

// ================
// 1. Shared State
// ================
const MultiAgentState = Annotation.Root({
  request: Annotation<string>,
  testCases: Annotation<string>,
  review: Annotation<string>,
  finalOutput: Annotation<string>,
  next: Annotation<string>,
});

// ================
// 2. Model
// ================
const model = new ChatAnthropic({
  model: "claude-haiku-4-5",
  temperature: 0.2,
  maxTokens: 1500,
});

// ================
// 3. Worker: Generator Agent
// ================
async function generatorNode(state: typeof MultiAgentState.State) {
  console.log("Generator Worker: Creating test cases...");

  const prompt = `You are a Test Case Generator Agent.

  User request: ${state.request}

  Generate 3 clear test cases.
  For each test case include:
  - Title
  - Type (Positive/Negative/Edge Case)
  - Preconditions
  - Steps
  - Expected Result`;

  const response = await model.invoke(prompt);
  const content =
    typeof response.content === "string"
      ? response.content
      : JSON.stringify(response.content);

  return { testCases: content, next: "supervisor" };
}

// ================
// 4. Worker: Reviewer Agent
// ================
async function reviewerNode(state: typeof MultiAgentState.State) {
  console.log("Reviewer Worker: Reviewing test cases...");

  const prompt = `You are a Test Case Reviewer Agent.

  Review the following test cases generated by the Generator Agent and improve them.
  Focus on:
  - missing negative scenarios
  - clarity of steps
  - expected results
  - practical QA quality

  Test Cases:
  ${state.testCases}

  Return the improved test cases in a clear format.`;

  const response = await model.invoke(prompt);
  const content =
    typeof response.content === "string"
      ? response.content
      : JSON.stringify(response.content);

  return { review: content, finalOutput: content, next: "supervisor" };
}

// ================
// 5. Supervisor Agent
// ================
async function supervisorNode(state: typeof MultiAgentState.State) {
  console.log("Supervisor: Deciding next step...");

  // If test cases are not generated yet, go to generator
  if (!state.testCases) {
    return { next: "generator" };
  }
  // If review is not done yet, go to reviewer
  if (!state.review) {
    return { next: "reviewer" };
  }

  console.log("Supervisor: All steps completed. Final output ready.");
  return {
    next: "end",
    finalOutput: state.finalOutput || state.review || state.testCases,
  };
}

// ================
// 6. Routing Function
// ================
function routeSupervisor(state: typeof MultiAgentState.State) {
  if (state.next === "generator") return "generator";
  if (state.next === "reviewer") return "reviewer";
  return END;
}

// ================
// 7. Build the Graph
// ================
const workflow = new StateGraph(MultiAgentState)
  .addNode("supervisor", supervisorNode)
  .addNode("generator", generatorNode)
  .addNode("reviewer", reviewerNode)

  .addEdge(START, "supervisor")
  .addConditionalEdges("supervisor", routeSupervisor, {
    generator: "generator",
    reviewer: "reviewer",
    [END]: END,
  })

  .addEdge("generator", "supervisor")
  .addEdge("reviewer", "supervisor");

const graph = workflow.compile();

async function main() {
  console.log("Starting Multi-Agent Workflow...");

  const result = await graph.invoke({
    request:
      "Create test cases for a login feature for a web application. Include positive, negative, and edge cases.",
    // empty initial state for the other fields
    testCases: "",
    review: "",
    finalOutput: "",
    next: "",
  });

  console.log("Final Output from Multi-Agent Workflow:");
  console.log(result.finalOutput || result.review || result.testCases);
}

main().catch(console.error);

Run: npx tsx lectures/09-multi-agent.ts

Output
Starting Multi-Agent Workflow...
Supervisor: Deciding next step...
Generator Worker: Creating test cases...
Supervisor: Deciding next step...
Reviewer Worker: Reviewing test cases...
Supervisor: Deciding next step...
Supervisor: All steps completed. Final output ready.
Final Output from Multi-Agent Workflow:

TC-01: Successful login with valid credentials (Positive)
Preconditions: Registered user exists, account not locked
Steps: 1. Open /login  2. Enter valid username  3. Enter valid password  4. Click Login
Expected Result: User is redirected to the dashboard and a session is created

TC-02: Login fails with invalid password (Negative)
...
TC-03: Account locks after 3 failed attempts (Edge Case)
...
State

Every node returns a partial state object. LangGraph merges it, so testCases written by the generator is readable by the reviewer.

Supervisor

Plain if logic here, which is deterministic and cheap. You can also let an LLM decide the route when the workflow is less predictable.

Conditional edges

addConditionalEdges takes a routing function that returns a node name. This is how branching and loops are built.

Handoff notes: a common upgrade is to have the generator also write a short note for the reviewer ("focus on missing negative scenarios"), stored in a handoffNote field. The reviewer reads the note along with the test cases, which makes the review far more targeted.

Key takeaways

  • Split a big job into small agents, each with one clear prompt
  • State is the communication channel between agents
  • A supervisor node plus conditional edges decides who runs next
  • Every node logs, so you can see exactly where a bad output came from
  • This is the base pattern for requirements, generation, review and execution pipelines

Exercises

  1. Add a third worker, requirementsNode, that turns a raw feature description into clear acceptance criteria before the generator runs. Route it first in the supervisor.
  2. Add a handoffNote field to the state. Make the generator write the note and the reviewer read it.
  3. Add a score field. The reviewer scores the test cases from 1 to 10, and the supervisor sends them back to the generator if the score is below 7, with a max of 3 attempts so it cannot loop forever.
  4. Make the reviewer return structured output with a Zod schema from chapter 07, so the final output is JSON your framework can import.
  5. Give the whole graph a checkpointer from chapter 08 and run two different requests on two thread ids.
10

RAG over your test docs

Answer questions from your own requirement documents instead of the model's general knowledge.

At a glance

Plain LLM

  • Answers from training data
  • Knows nothing about your SRS or Jira
  • Invents plausible details
VS

RAG pipeline

  • Retrieves the matching chunks from your docs
  • Answers only from that context
  • Says "not in the document" when it is missing

RAG = Retrieval Augmented Generation. Load the document, split it into chunks, turn each chunk into a vector, find the chunks closest to the question, and hand only those to the LLM.

Indexing (once) login-requirements.txt Split intochunks (300) Embeddings(nomic-embed) Vector storein memory Asking (every question) "What happens after3 failed attempts?" Retrievertop k = 2 chunks Prompt with{context} + {input} LLM Answer
Indexing runs once. Retrieval runs on every question and only the two closest chunks reach the model.

Extra setup for this chapter

terminal
npm install @langchain/classic @langchain/textsplitters
ollama pull nomic-embed-text        # free local embedding model

The document we will ask about

documents/login-requirements.txt
Login Feature Requirements

The login feature allows registered users to access the application.

Acceptance Criteria:
1. The user must be able to enter their username and password.
2. The system must validate the credentials against the database.
3. The system must display an error message for invalid credentials.
4. After 3 failed login attempts, the user account should be locked for 15 minutes.
5. On successful login, the user should be redirected to the dashboard.
6. Passwords must be stored securely using hashing algorithms.
7. Password must be at least 8 characters long.

Business Rules:
- Guest users cannot access the login feature.
- The system should log all login attempts for security auditing.
- The login feature must comply with the company's security policies and standards.
- Login session expires after 30 minutes of inactivity.
- Forgot password link should be available on the login page to reset the password.
lectures/10-rag-test-docs.ts
import "dotenv/config";

// Represents a document object that LangChain understands
import { Document } from "@langchain/core/documents";
// Splits large text into smaller overlapping chunks
import { RecursiveCharacterTextSplitter } from "@langchain/textsplitters";
// Converts text into numerical vectors (embeddings) using Ollama - local FREE model
import { OllamaEmbeddings } from "@langchain/ollama";
// In-memory vector store to hold embeddings and perform similarity search
import { MemoryVectorStore } from "@langchain/classic/vectorstores/memory";
// Helps us create structured prompts for LLMs
import { ChatPromptTemplate } from "@langchain/core/prompts";
// Combines retrieved documents into a single input for the LLM
import { createStuffDocumentsChain } from "@langchain/classic/chains/combine_documents";
// Connects retriever + document chain to form a complete RAG pipeline
import { createRetrievalChain } from "@langchain/classic/chains/retrieval";
import { ChatAnthropic } from "@langchain/anthropic";
import * as fs from "fs";

async function main() {
  console.log("RAG Pipeline Example");

  // 1. Load the document
  const sampleDocument = fs.readFileSync("./documents/login-requirements.txt", "utf-8");

  const documents = [
    new Document({
      pageContent: sampleDocument,
      metadata: { source: "login-requirements.txt" },
    }),
  ];
  console.log("Document Loaded");

  // 2. Split the document into chunks
  const textSplitter = new RecursiveCharacterTextSplitter({
    chunkSize: 300,      // each chunk has a maximum of 300 characters
    chunkOverlap: 50,    // each chunk overlaps the next by 50 characters
  });

  const splits = await textSplitter.splitDocuments(documents);
  console.log("Document split into", splits.length, "chunks.");

  // 3. Create embeddings + vector store
  const embeddings = new OllamaEmbeddings({
    model: "nomic-embed-text",
    baseUrl: "http://localhost:11434",
  });

  const vectorStore = await MemoryVectorStore.fromDocuments(splits, embeddings);
  console.log("--- Vector store created using Ollama embeddings ---");

  // 4. Create the retriever
  const retriever = vectorStore.asRetriever({ k: 2 });  // top 2 most relevant chunks

  // 5. Create the LLM
  const model = new ChatAnthropic({
    model: "claude-haiku-4-5",
    temperature: 0.2,
    maxTokens: 1500,
  });

  // 6. Create the prompt template
  // {context} is replaced with the retrieved chunks, {input} with the user question
  const prompt = ChatPromptTemplate.fromTemplate(`
    You are a helpful assistant for software testing.
    Answer the question based only on the following context provided from the document.
    If the answer is not contained within the context, say
    "I don't have enough information in the document".

    Context:
    {context}

    Question: {input}`);

  // 7. Create the document chain
  const documentChain = await createStuffDocumentsChain({
    llm: model,
    prompt: prompt,
  });

  // 8. Create the full RAG chain
  // Retriever (find relevant chunks) + Document chain (send chunks to LLM)
  const ragChain = await createRetrievalChain({
    retriever,
    combineDocsChain: documentChain,
  });

  // 9. Ask questions
  const questions = [
    "What are the acceptance criteria for the login feature?",
    "What happens after 3 failed login attempts?",
    "Is password hashing required?",
    "Can guest users access the login feature?",
    "What is the maximum file upload size?",   // not in the document
  ];

  for (const question of questions) {
    console.log("\n--------------------------------------------------");
    console.log(`Question: ${question}`);
    const result = await ragChain.invoke({ input: question });
    console.log("Answer:", result.answer);
  }
}

main().catch(console.error);

Run: npx tsx lectures/10-rag-test-docs.ts  (Ollama must be running)

Output
RAG Pipeline Example
Document Loaded
Document split into 6 chunks.
--- Vector store created using Ollama embeddings ---

--------------------------------------------------
Question: What happens after 3 failed login attempts?
Answer: After 3 failed login attempts, the user account should be locked for
15 minutes.

--------------------------------------------------
Question: Is password hashing required?
Answer: Yes. Passwords must be stored securely using hashing algorithms, and
passwords must be at least 8 characters long.

--------------------------------------------------
Question: What is the maximum file upload size?
Answer: I don't have enough information in the document.

That last answer is the point of RAG. The model does not guess, because the prompt tells it to answer only from the retrieved context. For a tester this is the difference between a usable assistant and a confident liar.

chunkSize / overlap

300 characters with 50 overlap keeps related sentences together. Too small and a rule gets cut in half, too large and retrieval loses precision.

k = 2

How many chunks reach the model. Raise it for broad questions, keep it low to stay cheap and focused.

PDF instead of text

Swap fs.readFileSync for new PDFLoader("./documents/SRS.pdf").load() from @langchain/community. Everything after that is identical.

Agentic RAG: instead of a fixed chain, wrap the retriever in a tool and hand it to createAgent. The agent then decides when to search the docs, can search twice with different wording, and can combine the result with other tools.

Key takeaways

  • RAG = load, split, embed, retrieve, then answer from the retrieved context
  • Ollama embeddings keep indexing free and local, which matters for internal documents
  • The prompt is what forces the model to stay inside the document
  • Chunk size, overlap and k are the three knobs worth tuning
  • Wrapping the retriever in a tool turns a fixed chain into an agent that decides when to search

Exercises

  1. Add your own requirement file to documents/ and ask 5 questions about it. Include one question the document cannot answer and confirm the agent refuses.
  2. Change chunkSize to 100 and then 1000, and k to 1 and then 5. Note where the answers get worse and why.
  3. Load a PDF SRS with PDFLoader instead of the text file and rerun the same questions.
  4. Combine chapters 07 and 10: make the RAG pipeline return test cases in the Zod TestCaseSchema, generated only from the acceptance criteria in your document.
  5. Turn the retriever into a tool and give it to createAgent along with the calculator. Ask "how many minutes is the lockout plus the session timeout?" and watch it use both.
11

MCP servers: tools that live outside your agent

Write a QA utilities server once, then use it from any agent, Claude Code, or VS Code.

At a glance

Local tools (chapters 05 and 06)

  • Defined inside the agent file
  • Only that agent can use them
  • Copy-paste to share them
VS

MCP server tools

  • A separate process with its own tools
  • Any MCP client can connect: your agent, Claude Code, VS Code
  • Write once, reuse across your whole team

MCP (Model Context Protocol) is a standard way for a program to expose tools to any AI client. An MCP server has three parts: create the server, register tools, start it on a transport (stdio for local servers).

Extra setup for this chapter

terminal
npm install @modelcontextprotocol/sdk @langchain/mcp-adapters

Part 1: build the MCP server

mcp-servers/qa-utils-server.ts
/*
 3 important parts:
 1. Create the server
 2. Create tools / list tools / handle tool calls
 3. Start the server with stdio transport
*/

import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import { z } from "zod";

// 1. Create the MCP server
const server = new McpServer({
  name: "QA Utils Server",
  version: "1.0.0",
});

// 2. Register tools

// Tool 1: Generate a random test email address
server.registerTool(
  "generate_test_email",
  {
    description: "Generate a random test email address useful for automation testing.",
    inputSchema: z.object({}),
  },
  async () => {
    const random = Math.floor(Math.random() * 10000);
    const email = `test.user_${random}@example.com`;
    return { content: [{ type: "text", text: email }] };
  }
);

// Tool 2: Generate a strong test password
server.registerTool(
  "generate_test_password",
  {
    description: "Generate a strong test password useful for automation testing.",
    inputSchema: z.object({}),
  },
  async () => {
    const random = Math.floor(Math.random() * 10000);
    const password = `Test@${random}Password!`;
    return { content: [{ type: "text", text: password }] };
  }
);

// Tool 3: Generate a mobile number
server.registerTool(
  "generate_mobile_number",
  {
    description: "Generate a random mobile number useful for automation testing.",
    inputSchema: z.object({}),
  },
  async () => {
    const random = Math.floor(Math.random() * 10000);
    const mobileNumber = `+1-555-000-${random.toString().padStart(4, "0")}`;
    return { content: [{ type: "text", text: mobileNumber }] };
  }
);

// Tool 4: Add numbers
server.registerTool(
  "add_numbers",
  {
    description: "Add two numbers together and return the sum.",
    inputSchema: z.object({
      a: z.number().describe("First number to add"),
      b: z.number().describe("Second number to add"),
    }),
  },
  async ({ a, b }) => {
    const sum = a + b;
    return { content: [{ type: "text", text: sum.toString() }] };
  }
);

// 3. Start the server with stdio transport
async function main() {
  const transport = new StdioServerTransport();
  await server.connect(transport);
  console.error("QA Utils Server is running. Waiting for requests...");
}

main().catch((error) => {
  console.error("Error starting the server:", error);
  process.exit(1);
});

Log with console.error, never console.log, in a stdio MCP server. Standard output carries the protocol messages, so a stray console.log corrupts the stream and the client fails to connect.

Part 2: connect an agent to the server

The client starts the server as a child process, loads whatever tools it exposes, and hands them to createAgent like any other tool array.

lectures/11-mcp-agent.ts
import "dotenv/config";
import { createAgent } from "langchain";
import { MultiServerMCPClient } from "@langchain/mcp-adapters";
import path from "path";

async function main() {
  console.log("------- Using MCP tools with a LangChain agent -------");

  // 1. Connect to the MCP server
  const client = new MultiServerMCPClient({
    qaUtils: {
      transport: "stdio",
      command: "npx",
      args: ["tsx", path.resolve("./mcp-servers/qa-utils-server.ts")],
    },
  });

  // Load all tools exposed by the MCP server
  const tools = await client.getTools();
  console.log("Tools loaded from MCP server:", tools.map((t) => t.name));

  // 2. Create the agent with the MCP tools
  const agent = createAgent({
    model: "anthropic:claude-haiku-4-5",
    tools,
    systemPrompt: `You are a helpful QA Automation Assistant.
      You have access to MCP tools for generating test emails, strong passwords,
      mobile numbers, and adding numbers.

      Rules:
      1. When the user asks for test data, use the MCP tools.
      2. When the user asks a general testing question, answer from your knowledge.
      3. Be clear and concise.`,
  });

  // 3. Demo questions
  const questions = [
    "Generate a test email for me",
    "Generate a strong password for me",
    "Generate a mobile number for me",
    "Add 123 and 456 for me",
    "What is the best way to test a web application?",
  ];

  for (const question of questions) {
    console.log("\n----------------------------");
    console.log("Q:", question);
    const result = await agent.invoke({
      messages: [{ role: "user", content: question }],
    });
    console.log("Agent:", result.messages.at(-1)?.content);
  }

  await client.close();   // close the connection to the MCP server
}

main().catch(console.error);

Run: npx tsx lectures/11-mcp-agent.ts

Output
------- Using MCP tools with a LangChain agent -------
Tools loaded from MCP server: [
  'generate_test_email', 'generate_test_password',
  'generate_mobile_number', 'add_numbers'
]

----------------------------
Q: Generate a test email for me
Agent: Here is a test email you can use: test.user_4821@example.com

----------------------------
Q: Add 123 and 456 for me
Agent: 123 + 456 = 579

----------------------------
Q: What is the best way to test a web application?
Agent: Start with a risk-based approach: cover critical user journeys with
end-to-end tests, push the rest down to API and unit level... (no tool used)

Part 3: use the same server in your editor

The same server works in VS Code and other MCP clients. Add a config file and your editor's AI gets your QA tools.

.vscode/mcp.json
{
  "servers": {
    "qa-utils-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["tsx", "./mcp-servers/qa-utils-server.ts"]
    }
  }
}

Why this matters for QA: your team already has scripts for test data, environment resets, log parsing and Jira lookups. Wrap them once as an MCP server and every agent, editor and teammate can call them with no copy-paste.

Key takeaways

  • An MCP server = create server, register tools, connect a transport
  • MultiServerMCPClient.getTools() returns normal LangChain tools
  • One server can serve your agent, VS Code and Claude Code at the same time
  • Use console.error for logs in a stdio server
  • Always client.close() so the child process exits

Exercises

  1. Add two tools to the server: generate_test_user returning a full JSON user object, and validate_email(email) returning whether it is well formed.
  2. Connect a second MCP server in the same client (for example a filesystem server) and ask a question that needs a tool from each.
  3. Register the server in .vscode/mcp.json and call one of your tools from your editor's AI chat.
12

Final project: a Playwright agent that runs your test

Describe the test in plain English. The agent opens a real browser, navigates, fills, clicks, verifies, takes a screenshot and reports the result.

At a glance

Scripted Playwright test

  • You write every step and selector in code
  • Runs the same way every time
  • Best for regression suites in CI
VS

Playwright agent

  • You describe the task, the agent picks the tools and the order
  • Reasons about what it sees (flash messages, page text)
  • Best for exploratory checks and smoke runs
Task in plain English "Test the login on..." LangChain agent Reason → Act → Observe repeat until the task is done Final report: pass or fail launch_browser navigate_to type_text click_element get_text take_screenshot close_browser Playwright tools (one shared browser) Chromium visible, slowMo
Everything from chapters 01 to 11 in one place: a system prompt with rules, seven tools that share one browser, and the reasoning loop that calls them in the right order.

Extra setup for this project

terminal
npm install playwright
npx playwright install chromium       # downloads the browser Playwright will drive

File 1: the Playwright tools

Seven small tools that share one browser and one page through module-level variables. Every tool returns a short string, because that string is what the agent reads next.

lectures/12-playwright/playwright-tools.ts
import { tool } from "langchain";
import { z } from "zod";
import { chromium, Browser, Page } from "playwright";

// Global variables (shared across tools)
let browser: Browser | null = null;
let page: Page | null = null;

// ============================================
// Playwright Tools
// ============================================

export const launchBrowser = tool(
  async () => {
    if (browser) return "Browser is already open.";

    browser = await chromium.launch({
      headless: false,
      slowMo: 700,
    });

    const context = await browser.newContext();
    page = await context.newPage();

    return "Browser launched successfully (visible mode).";
  },
  {
    name: "launch_browser",
    description: "Launch a visible Chromium browser. Always call this first.",
    schema: z.object({}),
  }
);

export const navigateTo = tool(
  async ({ url }) => {
    if (!page) return "Error: Browser not launched. Call launch_browser first.";
    await page.goto(url, { waitUntil: "domcontentloaded" });
    return `Navigated to ${url}`;
  },
  {
    name: "navigate_to",
    description: "Navigate to a URL",
    schema: z.object({
      url: z.string().url().describe("Full URL to open"),
    }),
  }
);

export const typeText = tool(
  async ({ selector, text }) => {
    if (!page) return "Error: Browser not launched.";
    await page.fill(selector, text);
    return `Typed "${text}" into ${selector}`;
  },
  {
    name: "type_text",
    description: "Type text into an input field",
    schema: z.object({
      selector: z.string().describe("CSS selector"),
      text: z.string().describe("Text to type"),
    }),
  }
);

export const clickElement = tool(
  async ({ selector }) => {
    if (!page) return "Error: Browser not launched.";
    await page.click(selector);
    return `Clicked on ${selector}`;
  },
  {
    name: "click_element",
    description: "Click on an element",
    schema: z.object({
      selector: z.string().describe("CSS selector of element to click"),
    }),
  }
);

export const getText = tool(
  async ({ selector }) => {
    if (!page) return "Error: Browser not launched.";
    const text = await page.textContent(selector);
    return text || "No text found";
  },
  {
    name: "get_text",
    description: "Extract text from an element",
    schema: z.object({
      selector: z.string().describe("CSS selector"),
    }),
  }
);

export const takeScreenshot = tool(
  async ({ filename = "screenshot.png" }) => {
    if (!page) return "Error: Browser not launched.";
    const path = `./screenshots/${filename}`;
    await page.screenshot({ path, fullPage: true });
    return `Screenshot saved as ${path}`;
  },
  {
    name: "take_screenshot",
    description: "Take a full page screenshot",
    schema: z.object({
      filename: z.string().optional().describe("Filename for screenshot"),
    }),
  }
);

export const closeBrowser = tool(
  async () => {
    if (browser) {
      await browser.close();
      browser = null;
      page = null;
      return "Browser closed successfully.";
    }
    return "No browser is open.";
  },
  {
    name: "close_browser",
    description: "Close the browser",
    schema: z.object({}),
  }
);

File 2: the agent

The system prompt carries the rules (launch first, act in order, screenshot before closing, explain each action). The task is the test case, written the way you would write it for a colleague.

lectures/12-playwright/playwright-agent.ts
import "dotenv/config";
import { createAgent } from "langchain";
import { ChatAnthropic } from "@langchain/anthropic";

import {
  launchBrowser,
  navigateTo,
  typeText,
  clickElement,
  getText,
  closeBrowser,
  takeScreenshot,
} from "./playwright-tools.js";

async function main() {
  const model = new ChatAnthropic({
    model: "claude-haiku-4-5",
    temperature: 0.2,
    maxTokens: 1500,
  });

  const agent = createAgent({
    model,
    tools: [
      launchBrowser,
      navigateTo,
      typeText,
      clickElement,
      getText,
      closeBrowser,
      takeScreenshot,
    ],
    systemPrompt: `You are an automation testing agent that controls a real browser
      using Playwright.

      Rules:
      1. Always start with launching a browser before performing any actions.
      2. Perform actions in the order they are requested.
      3. Take a screenshot before closing the browser to capture the final state.
      4. Use clear CSS selectors for interacting with elements on the page.
      5. Explain each action you take, and report whether the test PASSED or FAILED.`,
  });

  const task = `Test the login functionality on https://the-internet.herokuapp.com/login:

    1. Launch Browser
    2. Navigate to the login page
    3. Enter username "tomsmith"
    4. Enter password "SuperSecretPassword!"
    5. Click the login button
    6. Verify successful login by reading the flash message
    7. Take a screenshot named "login_test_result.png"
    8. Close Browser`;

  const result = await agent.invoke({
    messages: [{ role: "user", content: task }],
  });

  // Show each tool call and what the browser answered
  console.log("\n===== Steps taken =====");
  for (const msg of result.messages) {
    const type = msg.getType();
    if (type === "ai" && msg.tool_calls?.length) {
      for (const call of msg.tool_calls) {
        console.log(`-> ${call.name}(${JSON.stringify(call.args)})`);
      }
    } else if (type === "tool") {
      console.log(`   ${msg.content}`);
    }
  }

  console.log("\n===== Final Result =====");
  console.log(result.messages.at(-1)?.content);
}

main().catch(console.error);

Run: npx tsx lectures/12-playwright/playwright-agent.ts

Output (a Chromium window opens and you watch the test run)
===== Steps taken =====
-> launch_browser({})
   Browser launched successfully (visible mode).
-> navigate_to({"url":"https://the-internet.herokuapp.com/login"})
   Navigated to https://the-internet.herokuapp.com/login
-> type_text({"selector":"#username","text":"tomsmith"})
   Typed "tomsmith" into #username
-> type_text({"selector":"#password","text":"SuperSecretPassword!"})
   Typed "SuperSecretPassword!" into #password
-> click_element({"selector":"button[type='submit']"})
   Clicked on button[type='submit']
-> get_text({"selector":"#flash"})
   You logged into a secure area! ×
-> take_screenshot({"filename":"login_test_result.png"})
   Screenshot saved as ./screenshots/login_test_result.png
-> close_browser({})
   Browser closed successfully.

===== Final Result =====
Test PASSED. I launched the browser, opened the login page, entered the username
and password, clicked Login and read the flash message "You logged into a secure
area!", which confirms a successful login. The final state is saved in
./screenshots/login_test_result.png and the browser is closed.

How it works

shared state

All tools read the same page. Launch creates it, close clears it, every other tool checks it exists before acting.

short return strings

The tool's return value is what the agent reads to decide the next step. "Clicked on #login" is enough; a huge DOM dump is not.

the loop

The agent calls a tool, reads the result, decides the next call. Eight steps in the task become eight tool calls plus a report.

Where this fits in real testing: agents are great for exploratory checks, smoke runs and "does this flow still work?" questions. For your regression suite keep scripted tests: they are cheaper, faster and deterministic. The strongest combination is to let the agent explore, then save what worked as a normal Playwright spec.

Key takeaways

  • Tools can drive anything, including a real browser
  • Return short, clear strings from tools: they are the agent's eyes
  • Rules in the system prompt make multi-step runs reliable
  • Read the message history to see every step the agent took
  • Screenshot before closing, so a failed run still leaves evidence

Exercises

  1. Change the task to a different flow, for example log in to https://www.saucedemo.com, add the first product to the cart and verify the cart badge shows 1.
  2. Add three tools: get_page_title, select_option(selector, value) and wait_for_text(selector, text).
  3. Write a negative test: enter a wrong password and verify the error message. Make the agent report FAILED when the flash message does not match.
  4. Combine with chapter 07: pass responseFormat with a TestReport schema (steps, status, screenshotPath) so the final report is JSON.
  5. Combine with chapter 08: add a checkpointer and a thread id, run the test, then ask "which selector did you use for the login button?" in a second invoke.
  6. Combine with chapter 10: have the agent read your requirements document, generate the test steps from the acceptance criteria, then execute them in the browser.

What comes next

You can now build an agent, control it with a system prompt, run calls in parallel, give it tools, get typed output, keep memory across sessions, coordinate several agents, ground answers in your own documents, share tools over MCP and drive a real browser.

LangGraph in depth

Conditional edges, cycles with a max-attempts guard, and human-in-the-loop approval with interrupt() before a destructive step runs.

Agentic RAG

Wrap the retriever as a tool so the agent decides when to search, and can search again with better wording.

Playwright MCP

Replace your hand-written browser tools with the official Playwright MCP server and get accessibility-tree based interaction.

LangSmith

Turn on the three tracing lines in your .env and inspect every LLM call, tool call and token your agents make.

Evaluation

Build a golden dataset of user stories and expected test cases, then score your agent's output on every change.

Put it in CI

Run the generator agent on every new Jira story and post the draft test cases as a comment for review.