The Testing Academy · Playwright + AI

Playwright AI agents, all of them

The three agents that ship inside Playwright, the 89 tool server that drives them, the skills and the command line that need no server at all, and the third-party agents people confuse them with. Run end to end against TTACart, including the generated assertion that failed and the two lines that fixed it.

planner, generator, healer89 tools6 editorsa real failure

Everything in sections 1 to 8 was produced by running it: Playwright 1.63.0, driven over stdio and from the terminal against the live TTACart. Tool counts, file listings, logs, generated code, the failure and the passing run are copied from that session. The third-party tools in section 10 need paid API keys and were not executed; their versions come from the package registries and their signatures from the type definitions shipped in the packages.

01

Three kinds of agent, and which one you want

The word covers three products that fail in different ways. Sort them first, compare second.

THREE FAMILIES, THREE DIFFERENT JOBS 1. They write the tests You end up with .spec.ts in git. planner explores, writes a plan generator replays it, emits code healer runs it, fixes what broke Ships with Playwright. npx playwright init-agents Output is reviewable, diffable, and runs without the model. 2. They drive the browser You end up with an outcome. Stagehand act / observe / extract browser-use Python, goal driven Midscene JS SDK and extension The model is in the loop every run, so every run costs tokens and can take a different path. 3. They sit inside a test One step, not the whole spec. auto-playwright auto('click login', { page }) ZeroStep ai('click login', { page, test }) Use for the one step that keeps breaking, not for the whole file. Needs a key at run time, so CI pays per execution.
The word "agent" covers three different products. Decide which one you are buying before you compare them.

When somebody says Playwright AI agent they mean one of three things, and the three have almost nothing in common except the word.

What it producesWhen the model runsWhat you review
Writes testsA committed .spec.tsOnce, while authoringA normal pull request
Drives the browserAn outcome, plus a traceEvery single runA transcript, if you kept one
Lives inside a testOne resolved stepEvery run, for that stepThe step's prompt

The distinction that matters in CI. Family one runs with no model and no API key once the file is written, so it costs nothing per run and fails deterministically. Families two and three call a model on every execution, so a green build depends on a vendor being up and on the same prompt resolving the same way twice. That is a different risk profile, not a better or worse one.

This page walks family one in depth, because it is the one that ships inside Playwright and the one most teams should start with, then covers the other two honestly enough that you can pick.

02

The three agents that ship with Playwright

One command writes them. They are markdown files with a tool allowlist, not a service.

Playwright 1.63 has three agent-related commands of its own. Nothing here needs a third-party package:

npx playwright --help, trimmed to the agent commands
mcp [options]           run the Playwright MCP server
cli                     run playwright cli commands from terminal
init-agents [options]   Initialize repository agents
init-skills [options]   Install Playwright agent skills

Running init-agents against a project scaffolds the team:

terminal
$ npx playwright init-agents --loop=claude

 playwright Using project "chromium" as a primary project
 specs/README.md - directory for test plans
 tests/seed.spec.ts - default environment seed file
 .claude/agents/playwright-test-generator.md - agent definition
 .claude/agents/playwright-test-healer.md - agent definition
 .claude/agents/playwright-test-planner.md - agent definition
 .mcp.json - mcp configuration
 Done.

Each agent is a markdown file: frontmatter naming the model and the exact tools it may call, then a prompt. The interesting part is the allowlist, because it is where the safety lives.

AgentModelCan write filesSignature tools
plannersonnetNoplanner_setup_page, planner_save_plan
generatorsonnetNogenerator_setup_page, generator_read_log, generator_write_test, the four browser_verify_* tools
healersonnetYes, Edit MultiEdit Writetest_run, test_debug, test_list, browser_generate_locator

Only the healer can edit your repository. The planner and the generator have no write tools at all; they hand their output to planner_save_plan and generator_write_test, which write to paths the tool controls. If you are nervous about pointing an agent at a repo, that asymmetry is the thing to notice, and it is why the healer is the one to watch in review.

Two lines from the healer prompt are worth reading before you let it near a red build:

.claude/agents/playwright-test-healer.md, quoted
If the error persists and you have high level of confidence that the test is correct,
mark this test as test.fixme() so that it is skipped during the execution.

Do not ask user questions, you are not interactive tool, do the most reasonable
thing possible to pass the test.

This is the one to put a policy around. An agent whose instruction is to make the suite pass, which cannot ask you anything, and which is allowed to mark a test test.fixme(), has a legal way to turn a failure into a skip. That is sometimes correct. It is never correct silently, so grep every healer diff for fixme before merging.

03

The engine underneath: a second MCP server

Not @playwright/mcp. A bigger, test-aware server with the runner attached.

The .mcp.json that init-agents writes does not point at the Playwright MCP server you may already have installed:

.mcp.json
{
  "mcpServers": {
    "playwright-test": {
      "command": "npx",
      "args": ["playwright", "run-test-mcp-server"]
    }
  }
}

Asked over stdio what it is and what it can do, it answers:

initialize, then tools/list
server     : {'name': 'Playwright Test Runner', 'version': '1.63.0'}
protocol   : 2024-11-05
tool count : 89
FamilyCountWhat it is for
browser_*80Driving and reading the page, including cookies, storage, routing, tracing and video
planner_*3setup_page, save_plan, submit_plan
generator_*3setup_page, read_log, write_test
test_*3list, run, debug, the actual Playwright runner
ONE SERVER, THREE AGENTS, TWO ARTEFACTS planner explore, then describe generator replay, then emit healer run, then repair test MCP server npx playwright run-test-mcp-server owns the browser owns the test runner 89 tools specs/login.plan.md markdown, human readable reviewed before any code exists tests/login/*.spec.ts ordinary Playwright runs with no model and no API key plan feeds the generator
The planner and the generator never touch your filesystem directly. Both artefacts are plain text, and the second one runs without a model.

The one API difference that will bite you. Every browser_* tool on this server requires an intent string. Call one without it and you get invalid_type ... path: ["intent"]. The browser-only @playwright/mcp server has no such field. It is not bureaucracy: the intent string is what ends up as the comment above each step in the generated test, which is why these specs read like the plan instead of like a macro recording.

04

The loop, run end to end against TTACart

Real plan, real log, real spec. Driven tool by tool so every line here is output, not illustration.

The agents are prompts, so to show what the machinery actually does I drove the MCP server directly and played the agents by hand against TTACart. Everything below is copied from that session.

Step 1, the planner explores and saves a plan

planner_setup_page runs your seed test and pauses at the end of it, leaving the browser exactly where the seed left it:

planner_setup_page
Running 1 test using 1 worker
### Paused at end of test. ready for interaction

### Page state
- Page URL: about:blank

An empty seed gives you a blank page, and that is the point. The generated tests/seed.spec.ts contains only a comment, so the first run starts at about:blank. The seed file is where sign-in, base URL and any fixture setup belong. Put the navigation there once and every planned scenario starts from the same known state.

tests/seed.spec.ts, after making it useful
import { test, expect } from '@playwright/test';

test.describe('Test group', () => {
  test('seed', async ({ page }) => {
    await page.goto('https://app.thetestingacademy.com/playwright/ttacart/');
  });
});

With that in place the snapshot comes back as an accessibility tree, not pixels, which is what the planner reasons over:

browser_snapshot
- generic [ref=e2]:
  - heading "TTACart" [level=1] [ref=e3]
  - generic [ref=e5]:
    - textbox "Username" [ref=e7]
    - textbox "Password" [ref=e9]
    - button "Login" [ref=e10] [cursor=pointer]

planner_save_plan then writes markdown. This is the file a human reviews, before a single line of test code exists:

specs/login.plan.md, written by the tool
# TTACart Login

## Application Overview

TTACart is a demo storefront used by The Testing Academy. The landing page is a
login form with a username field, a password field and a Login button.

## Test Scenarios

### 1. Login

**Seed:** `tests/seed.spec.ts`

#### 1.1. Standard user can sign in

**File:** `tests/login/standard-user-can-sign-in.spec.ts`

**Steps:**
  1. Fill the Username field with "standard_user"
  2. Fill the Password field with "tta_secret"
  3. Click the "Login" button
    - expect: The inventory page is shown
    - expect: The page lists six products

Step 2, the generator replays the plan and reads its own log

The generator does not write code from the plan. It performs each step against the live page, then reads back a log of what actually happened. Acting:

browser_click, with its required intent
intent : Click the "Login" button
target : [data-test="login-button"]

### Ran Playwright code
await page.locator('[data-test="login-button"]').click();
### Page
- Page URL: https://app.thetestingacademy.com/playwright/ttacart/inventory
- Page Title: TTACart - Products

And then generator_read_log, which is the whole trick:

generator_read_log, trimmed
# Steps

### Fill the Username field with "standard_user" and the Password field with "tta_secret"
await page.locator('[data-test="username"]').fill('standard_user');
await page.locator('[data-test="password"]').fill('tta_secret');

### Click the "Login" button
await page.locator('[data-test="login-button"]').click();

### The inventory page is shown
await expect(page.locator('[data-test="title"]')).toBeVisible();

# Best practices
- Do not improvise, do not add directives that were not asked for
- Use reliable locators from this log
- Use local variables for locators that are used multiple times
- NEVER! use page.waitForLoadState()
- NEVER! use page.waitForNavigation()
- NEVER! use page.waitForTimeout()
- NEVER! use page.evaluate()

This is why the output is usually decent. The model is not inventing a selector from a screenshot or from training data. It is copying lines that already executed against your page, from a log the tool wrote, with four hard bans attached. Most of the reason AI-generated Playwright used to be full of waitForTimeout is that nobody handed the model this list.

Notice the third block: asking to verify the text Products produced [data-test="title"], a real locator the tool resolved, not the string I asked about.

05

Then the generated test failed

One assertion the tool emitted could not pass. Worth showing, because this is the whole argument for review.

The last expectation, the page lists six products, went through browser_verify_list_visible. The tool reported success:

browser_verify_list_visible
### Result
Done
### Ran Playwright code
await expect(page.locator('body')).toMatchAriaSnapshot(`
  - list:
    - listitem: "Test.allTheThings() T-Shirt (Red)"
    - listitem: "TTA Bike Light"
    - listitem: "TTA Bolt T-Shirt"
    - listitem: "TTA Fleece Jacket"
    - listitem: "TTA Junior Tester Onesie"
    - listitem: "TTA Practice Backpack"
  `);

That code went into the spec. Running the spec:

npx playwright test
1 failed
  [chromium] > tests/login/standard-user-can-sign-in.spec.ts:7:3 > Login > Standard user can sign in

The reason is checkable in one call. TTACart's inventory page has no list in it at all:

page.evaluate on the live inventory page
{ "ul": 0, "ol": 0, "li": 0, "article": 6,
  "[data-test=\"inventory-item-name\"]": 6 }

What happened. The live check and the emitted assertion are not the same thing. The tool confirmed those six strings were visible, which was true, then wrote an aria snapshot describing them as a list of listitems anchored at body. The products are six article elements, so nothing on the page can ever satisfy that shape. A tool that says Done is telling you the step worked, not that the code it wrote will pass.

The healer's job, done here by hand, is a two line change to an assertion that matches the DOM:

tests/login/standard-user-can-sign-in.spec.ts, healed
// expect: The page lists six products
const productNames = page.locator('[data-test="inventory-item-name"]');
await expect(productNames).toHaveText([
  'Test.allTheThings() T-Shirt (Red)',
  'TTA Bike Light',
  'TTA Bolt T-Shirt',
  'TTA Fleece Jacket',
  'TTA Junior Tester Onesie',
  'TTA Practice Backpack',
]);
npx playwright test
Running 1 test using 1 worker
  1 passed (1.5s)

Read that sequence again, because it is the honest version of the pitch. Plan: good. Recorded locators: good, and better than a human in a hurry. One generated assertion: wrong in a way that only showed up when the test ran. The loop is designed for exactly this, the healer exists because the generator is not trusted, and your review is the last stage of the same chain. Generated tests are a starting draft that happens to compile.

06

The same three agents, in whichever editor you use

Six loop targets. Same prompts, different file format, and one of them ships a broken placeholder.

--loop decides the output format. The prompts are identical, so the choice is only about where your team already works:

terminal
npx playwright init-agents --loop=claude
npx playwright init-agents --loop=vscode
npx playwright init-agents --loop=copilot
npx playwright init-agents --loop=codex
npx playwright init-agents --loop=opencode
npx playwright init-agents --loop=vscode-legacy
--loopAgent filesMCP config
claude.claude/agents/*.md.mcp.json
vscode.github/agents/*.agent.md.vscode/mcp.json
copilot.github/agents/*.agent.md plus a setup workflow.vscode/mcp.json
codex.codex/agents/*.tomlwritten into the TOML
opencode.opencode/prompts/*.mdopencode.json
vscode-legacy.github/chatmodes/*.chatmode.md.vscode/mcp.json

The three MCP config dialects differ enough to matter if you hand-edit them:

three shapes for the same server
// .mcp.json                (Claude Code)
{ "mcpServers": { "playwright-test": { "command": "npx", "args": [...] } } }

// .vscode/mcp.json         (VS Code, Copilot)
{ "servers": { "playwright-test": { "type": "stdio", "command": "npx", "args": [...] } },
  "inputs": [] }

// opencode.json            (opencode)
{ "mcp": { "playwright-test": { "type": "local", "command": ["npx", ...], "enabled": true } } }

The codex loop is the only one that writes a sandbox policy, and it matches the tool allowlists exactly: sandbox_mode = "read-only" for the planner and the generator, "workspace-write" for the healer. If you are introducing this to a team that is nervous about agents editing code, that file is the clearest thing to show them.

One real bug to fix on the way past. The copilot and vscode loops write .github/workflows/copilot-setup-steps.yml whose build step reads run: npx run build. That is not a valid command, npx would look for a package called run. It sits under a # Customize this step as needed comment, so it is a placeholder rather than a claim, but it will fail if you enable the workflow untouched. Change it to npm run build or delete the step.

07

Skills and playwright-cli, the surface with no MCP server at all

A second, lighter mechanism: reference docs plus a plain command line the agent shells out to.

init-skills is not a smaller init-agents. It installs a different thing:

terminal
$ npx playwright init-skills --loop=claude

 Skill installed to `.claude/skills/playwright-cli`.
 Skill installed to `.claude/skills/playwright-component-testing`.
 Skill installed to `.claude/skills/playwright-trace`.
SkillSizeWhat it teaches
playwright-cli12.7 KB plus 9 reference filesDriving a browser from the command line: snapshots, refs, mocking, storage state, tracing
playwright-component-testing10.4 KB plus 5 reference filesComponent tests as a story gallery, and migrating off the experimental ct packages
playwright-trace4.8 KBReading a trace zip from the terminal: actions, requests, console, errors

Agents versus skills, in one line. An agent is a subagent with a tool allowlist that talks to an MCP server. A skill is documentation that gets loaded when it becomes relevant, and its allowed-tools here is just Bash. No server, no protocol, no ports. The playwright-cli skill is a manual for a command.

That command is real and you can drive it yourself. Against TTACart:

terminal
$ npx playwright cli open https://app.thetestingacademy.com/playwright/ttacart/
### Browser `default` opened with pid 82586.
### Ran Playwright code
await page.goto('https://app.thetestingacademy.com/playwright/ttacart/');

$ npx playwright cli snapshot
- textbox "Username" [ref=e7]
- textbox "Password" [ref=e9]
- button "Login" [ref=e10] [cursor=pointer]

$ npx playwright cli fill '[data-test="username"]' standard_user
$ npx playwright cli fill '[data-test="password"]' tta_secret

$ npx playwright cli click e10
### Ran Playwright code
await page.locator('[data-test="login-button"]').click();
### Page
- Page URL: .../ttacart/inventory
- Page Title: TTACart - Products

$ npx playwright cli find "Fleece"
### Result
Found 4 matches for "Fleece":

$ npx playwright cli close
Browser 'default' closed

Look at what clicking e10 printed. The agent passed an opaque ref from the snapshot, and the CLI answered with await page.locator('[data-test="login-button"]').click();. The ref is for the model to aim with; the locator is for you to keep. That is the same bargain as the generator log, reached without a protocol.

Targets are a selector or a ref, never a description. Both playwright cli and browser_generate_locator reject a role and name string like link "TTA Bike Light" with Unexpected token "" while parsing css selector. Take a snapshot, use the ref.

08

Two Playwright MCP servers, and which one you actually want

They share a prefix and almost nothing else.

This trips people up constantly, because both are official and both are called the Playwright MCP server in conversation.

@playwright/mcprun-test-mcp-server
Installnpx @playwright/mcp@latest, separate package (0.0.81)Built into playwright 1.63, no extra install
Tools2489
Knows about your testsNoYes: test_list, test_run, test_debug
Writes filesNoYes, plans and specs
intent requiredNoYes, on every browser_* call
Use it forGeneral browsing, exploring, scraping, a chat agent that needs a browserAuthoring and maintaining a Playwright suite

Rule of thumb. If the output you want is a committed test file, you want the test server and init-agents. If the output you want is an answer, you want @playwright/mcp. Installing both is fine, they do not collide, but pointing a general browsing agent at a repo and hoping for a good suite is the usual mistake.

There is a full setup walkthrough for the browser one, including VS Code installation and a complete TTACart checkout driven from plain English, in the Playwright MCP tutorial, and a deeper treatment of the planner, generator and healer team in the MCP and AI agents setup guide.

09

Agents that live inside a test you already wrote

One plain-English step in the middle of an ordinary spec. Verified signatures, honest caveats.

These do not author anything. You call a function mid-test and a model resolves that one step at run time. Both signatures below are read from the published type definitions of the installed packages.

auto-playwright 1.16.1

the exported signature
export declare const auto: (
  task: string,
  config: { page: Page; test?: Test },
  options?: StepOptions
) => Promise<any>;

type StepOptions = {
  debug?: boolean;
  model?: string;
  openaiApiKey?: string;
  openaiBaseUrl?: string;
};
how it reads in a spec
import { test, expect } from '@playwright/test';
import { auto } from 'auto-playwright';

test('checkout', async ({ page }) => {
  await page.goto('https://app.thetestingacademy.com/playwright/ttacart/');

  // deterministic where it can be
  await page.locator('[data-test="username"]').fill('standard_user');
  await page.locator('[data-test="password"]').fill('tta_secret');
  await page.locator('[data-test="login-button"]').click();

  // and a model only for the step that keeps moving
  await auto('add the fleece jacket to the cart', { page, test });
});

@zerostep/playwright 0.1.5

the exported signature
export declare const ai: (
  task: string | string[],
  config: { page, test },
  options?: ExecutionOptions & StepOptions
) => Promise<any>;

export declare const aiFixture: (test) => { ai: ... };

type ExecutionOptions = {
  parallelism?: number;      // default 10
  failImmediately?: boolean;
};

The array form is the interesting part. ai() accepts a list of tasks and runs them in parallel, ten at a time by default. That is useful for a batch of independent assertions and actively wrong for a sequence of steps that depend on each other, since nothing guarantees order. Pass a single string when the order matters.

What this costs you. Every run of that spec calls a model, so the step is billed per execution, needs a key in CI, and can resolve differently on a bad day. Use it for the one genuinely unstable step, keep the rest as ordinary locators, and treat a plain-English step that has been stable for a month as a candidate for being rewritten as a real locator.

10

Agents that drive the browser instead of testing it

Goal in, outcome out. Built on Playwright, aimed at automation rather than at a suite.

This family takes a goal and works the page until it is done. Playwright is the engine underneath, but the deliverable is the outcome, not a spec file.

ToolVersionLanguageShape
Stagehand4.1.0TypeScriptact / observe / extract on top of Playwright
browser-use0.13.10PythonGoal driven agent loop over a browser
Midscene1.12.9JavaScriptSDK, Chrome extension and YAML scripting
Skyvern1.0.48PythonWorkflow automation over real sites

Stagehand is the one a Playwright user recognises fastest, because extract gives you a typed object back rather than a string:

@browserbasehq/stagehand 4.1.0, from the published types
Stagehand.create(input: StagehandCreateOptions): Promise<Stagehand>

stagehand.act(instruction: string, options?): Promise<ActResult>
stagehand.observe(instruction?: string, options?): Promise<ObserveResult>
stagehand.extract<Schema extends z.ZodType>(
  instruction: string, schema: Schema, options?
): Promise<ExtractResult<Schema>>
stagehand.close(): Promise<void>

Two corrections worth carrying, if you learned this from an older tutorial. In 4.x you build the instance with the static Stagehand.create() rather than new Stagehand() followed by init(), and act, observe and extract sit directly on the instance rather than on stagehand.page. There is also no agent export in 4.1.0. Blog posts showing stagehand.page.act(...) are describing an older major version.

Said plainly: I did not run the four tools in this section. Each needs a paid API key, and this page does not use one. Their versions come from the registries and their signatures from the type definitions shipped in the packages, both checked while writing. Everything in the earlier sections was executed. I would rather tell you where the line is than blur it.

For the pattern where you assemble the agent yourself, with your own tools and your own control over the loop, there is a full TypeScript walkthrough in the Playwright agent with LangChain guide.

11

What to review before you trust any of it

The loop is good. It is not self-certifying, and one of these gates caught a real bug today.

FOUR GATES BEFORE YOU MERGE A GENERATED SPEC 1. Does it run? A tool saying Done means the step worked, not that the code it wrote passes. npx playwright test This caught the real bug on this page. 2. Does it assert? A spec that navigates and checks a heading exists is green and worthless. Every plan step with an expect line should have a matching assertion. 3. Banned waits? waitForTimeout waitForLoadState waitForNavigation evaluate The log bans all four. Their presence means it stopped following it. 4. Skip or fix? The healer is allowed to mark a test as skipped when it cannot pass it. git diff | grep fixme A shrinking failure count is not always progress.
Nothing here is specific to AI. It is the review you would give a new joiner, written down because the author is faster than you are.

Four checks, in the order that catches the most for the least effort:

  1. Run it yourself. The failure on this page was invisible until the spec executed: the tool reported Done and wrote an assertion that could never pass. No amount of reading the diff would have been as fast as one npx playwright test.
  2. Check it asserts something. Compare the spec against the plan file, which is the point of the plan being markdown. Every expect: line in the plan should have a visible counterpart in the code.
  3. Grep for the four banned waits. The generator log forbids waitForTimeout, waitForLoadState, waitForNavigation and evaluate. If one appears, something went off the rails and the rest of that file deserves a closer read.
  4. Grep the healer's diff for fixme. It is permitted to convert a failure it cannot fix into a skip, and it will not tell you in words. A suite that went from nine failures to zero deserves one command before you celebrate.

Where the real leverage is. The planner is the stage worth your attention, and it is the one people skip because markdown feels less impressive than code. A wrong plan produces a suite that is green, well-written and testing the wrong journeys, and that costs far more than one bad locator. Read specs/*.plan.md properly. It is the cheapest review in the chain.

And the honest summary of the whole thing. This is a very good drafting tool with a repair loop attached. It removes the tedious half of test authoring, the part where you hunt for selectors and retype the same login. It does not remove the part where somebody who understands the product decides what is worth asserting. Teams that keep the second half do well with it. Teams that treat a green suite as proof end up with a lot of tests and not much testing.

12

Try it yourself

Half an hour, one repo, no API key needed for the first three.

  1. Scaffold the team. In any project with a playwright.config.ts, run npx playwright init-agents --loop=claude (or your editor's loop) and read all three files it writes before you run anything. Note which agent can write to disk.
  2. Fix the seed first. Put your app's navigation and sign-in into tests/seed.spec.ts. Everything downstream starts from where the seed stops, and an empty seed means the planner explores about:blank.
  3. Drive the CLI by hand. npx playwright cli open <your app>, then snapshot, then click something by its ref. Watching the ref come back as a real locator is the fastest way to understand what the agents are actually doing.
  4. Plan one journey, generate it, then break it deliberately. Rename a data-test attribute in your app and let the healer run. Read its diff and decide whether you would have made the same change.
  5. Count the tokens. Run the same journey twice, once authored by the agents and once by hand. The generated one runs free forever after; note how long the authoring took versus your own. That number, not a demo video, is what tells you whether this belongs in your workflow.

A question worth being able to answer in an interview. "You used an AI agent to write these tests. How do you know they test anything?" The good answer is not a defence of the tool, it is the four gates above, plus the observation that the plan is reviewed in markdown before any code exists. Being able to say that clearly is worth more than the time the agents save you.