Three kinds of agent, and which one you want
The word covers three products that fail in different ways. Sort them first, compare second.
When somebody says Playwright AI agent they mean one of three things, and the three have almost nothing in common except the word.
| What it produces | When the model runs | What you review | |
|---|---|---|---|
| Writes tests | A committed .spec.ts | Once, while authoring | A normal pull request |
| Drives the browser | An outcome, plus a trace | Every single run | A transcript, if you kept one |
| Lives inside a test | One resolved step | Every run, for that step | The step's prompt |
The distinction that matters in CI. Family one runs with no model and no API key once the file is written, so it costs nothing per run and fails deterministically. Families two and three call a model on every execution, so a green build depends on a vendor being up and on the same prompt resolving the same way twice. That is a different risk profile, not a better or worse one.
This page walks family one in depth, because it is the one that ships inside Playwright and the one most teams should start with, then covers the other two honestly enough that you can pick.
The three agents that ship with Playwright
One command writes them. They are markdown files with a tool allowlist, not a service.
Playwright 1.63 has three agent-related commands of its own. Nothing here needs a third-party package:
mcp [options] run the Playwright MCP server
cli run playwright cli commands from terminal
init-agents [options] Initialize repository agents
init-skills [options] Install Playwright agent skillsRunning init-agents against a project scaffolds the team:
$ npx playwright init-agents --loop=claude
playwright Using project "chromium" as a primary project
specs/README.md - directory for test plans
tests/seed.spec.ts - default environment seed file
.claude/agents/playwright-test-generator.md - agent definition
.claude/agents/playwright-test-healer.md - agent definition
.claude/agents/playwright-test-planner.md - agent definition
.mcp.json - mcp configuration
Done.Each agent is a markdown file: frontmatter naming the model and the exact tools it may call, then a prompt. The interesting part is the allowlist, because it is where the safety lives.
| Agent | Model | Can write files | Signature tools |
|---|---|---|---|
| planner | sonnet | No | planner_setup_page, planner_save_plan |
| generator | sonnet | No | generator_setup_page, generator_read_log, generator_write_test, the four browser_verify_* tools |
| healer | sonnet | Yes, Edit MultiEdit Write | test_run, test_debug, test_list, browser_generate_locator |
Only the healer can edit your repository. The planner and the generator have no write tools at all; they hand their output to planner_save_plan and generator_write_test, which write to paths the tool controls. If you are nervous about pointing an agent at a repo, that asymmetry is the thing to notice, and it is why the healer is the one to watch in review.
Two lines from the healer prompt are worth reading before you let it near a red build:
If the error persists and you have high level of confidence that the test is correct,
mark this test as test.fixme() so that it is skipped during the execution.
Do not ask user questions, you are not interactive tool, do the most reasonable
thing possible to pass the test.This is the one to put a policy around. An agent whose instruction is to make the suite pass, which cannot ask you anything, and which is allowed to mark a test test.fixme(), has a legal way to turn a failure into a skip. That is sometimes correct. It is never correct silently, so grep every healer diff for fixme before merging.
The engine underneath: a second MCP server
Not @playwright/mcp. A bigger, test-aware server with the runner attached.
The .mcp.json that init-agents writes does not point at the Playwright MCP server you may already have installed:
{
"mcpServers": {
"playwright-test": {
"command": "npx",
"args": ["playwright", "run-test-mcp-server"]
}
}
}Asked over stdio what it is and what it can do, it answers:
server : {'name': 'Playwright Test Runner', 'version': '1.63.0'}
protocol : 2024-11-05
tool count : 89| Family | Count | What it is for |
|---|---|---|
browser_* | 80 | Driving and reading the page, including cookies, storage, routing, tracing and video |
planner_* | 3 | setup_page, save_plan, submit_plan |
generator_* | 3 | setup_page, read_log, write_test |
test_* | 3 | list, run, debug, the actual Playwright runner |
The one API difference that will bite you. Every browser_* tool on this server requires an intent string. Call one without it and you get invalid_type ... path: ["intent"]. The browser-only @playwright/mcp server has no such field. It is not bureaucracy: the intent string is what ends up as the comment above each step in the generated test, which is why these specs read like the plan instead of like a macro recording.
The loop, run end to end against TTACart
Real plan, real log, real spec. Driven tool by tool so every line here is output, not illustration.
The agents are prompts, so to show what the machinery actually does I drove the MCP server directly and played the agents by hand against TTACart. Everything below is copied from that session.
Step 1, the planner explores and saves a plan
planner_setup_page runs your seed test and pauses at the end of it, leaving the browser exactly where the seed left it:
Running 1 test using 1 worker
### Paused at end of test. ready for interaction
### Page state
- Page URL: about:blankAn empty seed gives you a blank page, and that is the point. The generated tests/seed.spec.ts contains only a comment, so the first run starts at about:blank. The seed file is where sign-in, base URL and any fixture setup belong. Put the navigation there once and every planned scenario starts from the same known state.
import { test, expect } from '@playwright/test';
test.describe('Test group', () => {
test('seed', async ({ page }) => {
await page.goto('https://app.thetestingacademy.com/playwright/ttacart/');
});
});With that in place the snapshot comes back as an accessibility tree, not pixels, which is what the planner reasons over:
- generic [ref=e2]:
- heading "TTACart" [level=1] [ref=e3]
- generic [ref=e5]:
- textbox "Username" [ref=e7]
- textbox "Password" [ref=e9]
- button "Login" [ref=e10] [cursor=pointer]planner_save_plan then writes markdown. This is the file a human reviews, before a single line of test code exists:
# TTACart Login
## Application Overview
TTACart is a demo storefront used by The Testing Academy. The landing page is a
login form with a username field, a password field and a Login button.
## Test Scenarios
### 1. Login
**Seed:** `tests/seed.spec.ts`
#### 1.1. Standard user can sign in
**File:** `tests/login/standard-user-can-sign-in.spec.ts`
**Steps:**
1. Fill the Username field with "standard_user"
2. Fill the Password field with "tta_secret"
3. Click the "Login" button
- expect: The inventory page is shown
- expect: The page lists six productsStep 2, the generator replays the plan and reads its own log
The generator does not write code from the plan. It performs each step against the live page, then reads back a log of what actually happened. Acting:
intent : Click the "Login" button
target : [data-test="login-button"]
### Ran Playwright code
await page.locator('[data-test="login-button"]').click();
### Page
- Page URL: https://app.thetestingacademy.com/playwright/ttacart/inventory
- Page Title: TTACart - ProductsAnd then generator_read_log, which is the whole trick:
# Steps
### Fill the Username field with "standard_user" and the Password field with "tta_secret"
await page.locator('[data-test="username"]').fill('standard_user');
await page.locator('[data-test="password"]').fill('tta_secret');
### Click the "Login" button
await page.locator('[data-test="login-button"]').click();
### The inventory page is shown
await expect(page.locator('[data-test="title"]')).toBeVisible();
# Best practices
- Do not improvise, do not add directives that were not asked for
- Use reliable locators from this log
- Use local variables for locators that are used multiple times
- NEVER! use page.waitForLoadState()
- NEVER! use page.waitForNavigation()
- NEVER! use page.waitForTimeout()
- NEVER! use page.evaluate()This is why the output is usually decent. The model is not inventing a selector from a screenshot or from training data. It is copying lines that already executed against your page, from a log the tool wrote, with four hard bans attached. Most of the reason AI-generated Playwright used to be full of waitForTimeout is that nobody handed the model this list.
Notice the third block: asking to verify the text Products produced [data-test="title"], a real locator the tool resolved, not the string I asked about.
Then the generated test failed
One assertion the tool emitted could not pass. Worth showing, because this is the whole argument for review.
The last expectation, the page lists six products, went through browser_verify_list_visible. The tool reported success:
### Result
Done
### Ran Playwright code
await expect(page.locator('body')).toMatchAriaSnapshot(`
- list:
- listitem: "Test.allTheThings() T-Shirt (Red)"
- listitem: "TTA Bike Light"
- listitem: "TTA Bolt T-Shirt"
- listitem: "TTA Fleece Jacket"
- listitem: "TTA Junior Tester Onesie"
- listitem: "TTA Practice Backpack"
`);That code went into the spec. Running the spec:
1 failed
[chromium] > tests/login/standard-user-can-sign-in.spec.ts:7:3 > Login > Standard user can sign inThe reason is checkable in one call. TTACart's inventory page has no list in it at all:
{ "ul": 0, "ol": 0, "li": 0, "article": 6,
"[data-test=\"inventory-item-name\"]": 6 }What happened. The live check and the emitted assertion are not the same thing. The tool confirmed those six strings were visible, which was true, then wrote an aria snapshot describing them as a list of listitems anchored at body. The products are six article elements, so nothing on the page can ever satisfy that shape. A tool that says Done is telling you the step worked, not that the code it wrote will pass.
The healer's job, done here by hand, is a two line change to an assertion that matches the DOM:
// expect: The page lists six products
const productNames = page.locator('[data-test="inventory-item-name"]');
await expect(productNames).toHaveText([
'Test.allTheThings() T-Shirt (Red)',
'TTA Bike Light',
'TTA Bolt T-Shirt',
'TTA Fleece Jacket',
'TTA Junior Tester Onesie',
'TTA Practice Backpack',
]);Running 1 test using 1 worker
1 passed (1.5s)Read that sequence again, because it is the honest version of the pitch. Plan: good. Recorded locators: good, and better than a human in a hurry. One generated assertion: wrong in a way that only showed up when the test ran. The loop is designed for exactly this, the healer exists because the generator is not trusted, and your review is the last stage of the same chain. Generated tests are a starting draft that happens to compile.
The same three agents, in whichever editor you use
Six loop targets. Same prompts, different file format, and one of them ships a broken placeholder.
--loop decides the output format. The prompts are identical, so the choice is only about where your team already works:
npx playwright init-agents --loop=claude
npx playwright init-agents --loop=vscode
npx playwright init-agents --loop=copilot
npx playwright init-agents --loop=codex
npx playwright init-agents --loop=opencode
npx playwright init-agents --loop=vscode-legacy| --loop | Agent files | MCP config |
|---|---|---|
claude | .claude/agents/*.md | .mcp.json |
vscode | .github/agents/*.agent.md | .vscode/mcp.json |
copilot | .github/agents/*.agent.md plus a setup workflow | .vscode/mcp.json |
codex | .codex/agents/*.toml | written into the TOML |
opencode | .opencode/prompts/*.md | opencode.json |
vscode-legacy | .github/chatmodes/*.chatmode.md | .vscode/mcp.json |
The three MCP config dialects differ enough to matter if you hand-edit them:
// .mcp.json (Claude Code)
{ "mcpServers": { "playwright-test": { "command": "npx", "args": [...] } } }
// .vscode/mcp.json (VS Code, Copilot)
{ "servers": { "playwright-test": { "type": "stdio", "command": "npx", "args": [...] } },
"inputs": [] }
// opencode.json (opencode)
{ "mcp": { "playwright-test": { "type": "local", "command": ["npx", ...], "enabled": true } } }The codex loop is the only one that writes a sandbox policy, and it matches the tool allowlists exactly: sandbox_mode = "read-only" for the planner and the generator, "workspace-write" for the healer. If you are introducing this to a team that is nervous about agents editing code, that file is the clearest thing to show them.
One real bug to fix on the way past. The copilot and vscode loops write .github/workflows/copilot-setup-steps.yml whose build step reads run: npx run build. That is not a valid command, npx would look for a package called run. It sits under a # Customize this step as needed comment, so it is a placeholder rather than a claim, but it will fail if you enable the workflow untouched. Change it to npm run build or delete the step.
Skills and playwright-cli, the surface with no MCP server at all
A second, lighter mechanism: reference docs plus a plain command line the agent shells out to.
init-skills is not a smaller init-agents. It installs a different thing:
$ npx playwright init-skills --loop=claude
Skill installed to `.claude/skills/playwright-cli`.
Skill installed to `.claude/skills/playwright-component-testing`.
Skill installed to `.claude/skills/playwright-trace`.| Skill | Size | What it teaches |
|---|---|---|
playwright-cli | 12.7 KB plus 9 reference files | Driving a browser from the command line: snapshots, refs, mocking, storage state, tracing |
playwright-component-testing | 10.4 KB plus 5 reference files | Component tests as a story gallery, and migrating off the experimental ct packages |
playwright-trace | 4.8 KB | Reading a trace zip from the terminal: actions, requests, console, errors |
Agents versus skills, in one line. An agent is a subagent with a tool allowlist that talks to an MCP server. A skill is documentation that gets loaded when it becomes relevant, and its allowed-tools here is just Bash. No server, no protocol, no ports. The playwright-cli skill is a manual for a command.
That command is real and you can drive it yourself. Against TTACart:
$ npx playwright cli open https://app.thetestingacademy.com/playwright/ttacart/
### Browser `default` opened with pid 82586.
### Ran Playwright code
await page.goto('https://app.thetestingacademy.com/playwright/ttacart/');
$ npx playwright cli snapshot
- textbox "Username" [ref=e7]
- textbox "Password" [ref=e9]
- button "Login" [ref=e10] [cursor=pointer]
$ npx playwright cli fill '[data-test="username"]' standard_user
$ npx playwright cli fill '[data-test="password"]' tta_secret
$ npx playwright cli click e10
### Ran Playwright code
await page.locator('[data-test="login-button"]').click();
### Page
- Page URL: .../ttacart/inventory
- Page Title: TTACart - Products
$ npx playwright cli find "Fleece"
### Result
Found 4 matches for "Fleece":
$ npx playwright cli close
Browser 'default' closedLook at what clicking e10 printed. The agent passed an opaque ref from the snapshot, and the CLI answered with await page.locator('[data-test="login-button"]').click();. The ref is for the model to aim with; the locator is for you to keep. That is the same bargain as the generator log, reached without a protocol.
Targets are a selector or a ref, never a description. Both playwright cli and browser_generate_locator reject a role and name string like link "TTA Bike Light" with Unexpected token "" while parsing css selector. Take a snapshot, use the ref.
Two Playwright MCP servers, and which one you actually want
They share a prefix and almost nothing else.
This trips people up constantly, because both are official and both are called the Playwright MCP server in conversation.
| @playwright/mcp | run-test-mcp-server | |
|---|---|---|
| Install | npx @playwright/mcp@latest, separate package (0.0.81) | Built into playwright 1.63, no extra install |
| Tools | 24 | 89 |
| Knows about your tests | No | Yes: test_list, test_run, test_debug |
| Writes files | No | Yes, plans and specs |
intent required | No | Yes, on every browser_* call |
| Use it for | General browsing, exploring, scraping, a chat agent that needs a browser | Authoring and maintaining a Playwright suite |
Rule of thumb. If the output you want is a committed test file, you want the test server and init-agents. If the output you want is an answer, you want @playwright/mcp. Installing both is fine, they do not collide, but pointing a general browsing agent at a repo and hoping for a good suite is the usual mistake.
There is a full setup walkthrough for the browser one, including VS Code installation and a complete TTACart checkout driven from plain English, in the Playwright MCP tutorial, and a deeper treatment of the planner, generator and healer team in the MCP and AI agents setup guide.
Agents that live inside a test you already wrote
One plain-English step in the middle of an ordinary spec. Verified signatures, honest caveats.
These do not author anything. You call a function mid-test and a model resolves that one step at run time. Both signatures below are read from the published type definitions of the installed packages.
auto-playwright 1.16.1
export declare const auto: (
task: string,
config: { page: Page; test?: Test },
options?: StepOptions
) => Promise<any>;
type StepOptions = {
debug?: boolean;
model?: string;
openaiApiKey?: string;
openaiBaseUrl?: string;
};import { test, expect } from '@playwright/test';
import { auto } from 'auto-playwright';
test('checkout', async ({ page }) => {
await page.goto('https://app.thetestingacademy.com/playwright/ttacart/');
// deterministic where it can be
await page.locator('[data-test="username"]').fill('standard_user');
await page.locator('[data-test="password"]').fill('tta_secret');
await page.locator('[data-test="login-button"]').click();
// and a model only for the step that keeps moving
await auto('add the fleece jacket to the cart', { page, test });
});@zerostep/playwright 0.1.5
export declare const ai: (
task: string | string[],
config: { page, test },
options?: ExecutionOptions & StepOptions
) => Promise<any>;
export declare const aiFixture: (test) => { ai: ... };
type ExecutionOptions = {
parallelism?: number; // default 10
failImmediately?: boolean;
};The array form is the interesting part. ai() accepts a list of tasks and runs them in parallel, ten at a time by default. That is useful for a batch of independent assertions and actively wrong for a sequence of steps that depend on each other, since nothing guarantees order. Pass a single string when the order matters.
What this costs you. Every run of that spec calls a model, so the step is billed per execution, needs a key in CI, and can resolve differently on a bad day. Use it for the one genuinely unstable step, keep the rest as ordinary locators, and treat a plain-English step that has been stable for a month as a candidate for being rewritten as a real locator.
Agents that drive the browser instead of testing it
Goal in, outcome out. Built on Playwright, aimed at automation rather than at a suite.
This family takes a goal and works the page until it is done. Playwright is the engine underneath, but the deliverable is the outcome, not a spec file.
| Tool | Version | Language | Shape |
|---|---|---|---|
| Stagehand | 4.1.0 | TypeScript | act / observe / extract on top of Playwright |
| browser-use | 0.13.10 | Python | Goal driven agent loop over a browser |
| Midscene | 1.12.9 | JavaScript | SDK, Chrome extension and YAML scripting |
| Skyvern | 1.0.48 | Python | Workflow automation over real sites |
Stagehand is the one a Playwright user recognises fastest, because extract gives you a typed object back rather than a string:
Stagehand.create(input: StagehandCreateOptions): Promise<Stagehand>
stagehand.act(instruction: string, options?): Promise<ActResult>
stagehand.observe(instruction?: string, options?): Promise<ObserveResult>
stagehand.extract<Schema extends z.ZodType>(
instruction: string, schema: Schema, options?
): Promise<ExtractResult<Schema>>
stagehand.close(): Promise<void>Two corrections worth carrying, if you learned this from an older tutorial. In 4.x you build the instance with the static Stagehand.create() rather than new Stagehand() followed by init(), and act, observe and extract sit directly on the instance rather than on stagehand.page. There is also no agent export in 4.1.0. Blog posts showing stagehand.page.act(...) are describing an older major version.
Said plainly: I did not run the four tools in this section. Each needs a paid API key, and this page does not use one. Their versions come from the registries and their signatures from the type definitions shipped in the packages, both checked while writing. Everything in the earlier sections was executed. I would rather tell you where the line is than blur it.
For the pattern where you assemble the agent yourself, with your own tools and your own control over the loop, there is a full TypeScript walkthrough in the Playwright agent with LangChain guide.
What to review before you trust any of it
The loop is good. It is not self-certifying, and one of these gates caught a real bug today.
Four checks, in the order that catches the most for the least effort:
- Run it yourself. The failure on this page was invisible until the spec executed: the tool reported
Doneand wrote an assertion that could never pass. No amount of reading the diff would have been as fast as onenpx playwright test. - Check it asserts something. Compare the spec against the plan file, which is the point of the plan being markdown. Every
expect:line in the plan should have a visible counterpart in the code. - Grep for the four banned waits. The generator log forbids
waitForTimeout,waitForLoadState,waitForNavigationandevaluate. If one appears, something went off the rails and the rest of that file deserves a closer read. - Grep the healer's diff for
fixme. It is permitted to convert a failure it cannot fix into a skip, and it will not tell you in words. A suite that went from nine failures to zero deserves one command before you celebrate.
Where the real leverage is. The planner is the stage worth your attention, and it is the one people skip because markdown feels less impressive than code. A wrong plan produces a suite that is green, well-written and testing the wrong journeys, and that costs far more than one bad locator. Read specs/*.plan.md properly. It is the cheapest review in the chain.
And the honest summary of the whole thing. This is a very good drafting tool with a repair loop attached. It removes the tedious half of test authoring, the part where you hunt for selectors and retype the same login. It does not remove the part where somebody who understands the product decides what is worth asserting. Teams that keep the second half do well with it. Teams that treat a green suite as proof end up with a lot of tests and not much testing.
Try it yourself
Half an hour, one repo, no API key needed for the first three.
- Scaffold the team. In any project with a
playwright.config.ts, runnpx playwright init-agents --loop=claude(or your editor's loop) and read all three files it writes before you run anything. Note which agent can write to disk. - Fix the seed first. Put your app's navigation and sign-in into
tests/seed.spec.ts. Everything downstream starts from where the seed stops, and an empty seed means the planner exploresabout:blank. - Drive the CLI by hand.
npx playwright cli open <your app>, thensnapshot, then click something by its ref. Watching the ref come back as a real locator is the fastest way to understand what the agents are actually doing. - Plan one journey, generate it, then break it deliberately. Rename a
data-testattribute in your app and let the healer run. Read its diff and decide whether you would have made the same change. - Count the tokens. Run the same journey twice, once authored by the agents and once by hand. The generated one runs free forever after; note how long the authoring took versus your own. That number, not a demo video, is what tells you whether this belongs in your workflow.
A question worth being able to answer in an interview. "You used an AI agent to write these tests. How do you know they test anything?" The good answer is not a defence of the tool, it is the four gates above, plus the observation that the plan is reviewed in markdown before any code exists. Being able to say that clearly is worth more than the time the agents save you.