The Testing Academy · Playwright + AI

Playwright MCP, start to finish

What MCP is, how to install Playwright and the Playwright MCP server in VS Code, how to start it and prove it is running, the 24 browser tools it gives your agent, and a full TTACart checkout driven from plain English, ending in a spec you can commit.

VS Code install24 toolsTTACart checkoutSession to spec

Everything quoted here was produced by running it: @playwright/mcp 0.0.80 (the server reports Playwright 1.63.0-alpha), driven over stdio against the live TTACart. Tool names, argument shapes, snapshots, totals and generated code are copied from that run, not from memory.

01

What MCP actually is

One protocol, so any AI client can use any tool without a custom integration for each pair.

MCP, the Model Context Protocol, is a standard way for an AI application to discover and call tools that live outside it. Three words are worth keeping straight:

host

Where you type

VS Code with Copilot Chat, Claude Code, Cursor. The host holds the conversation and decides when a tool should run.

client

One per server

The host creates a client for each server it is configured with. You never write this part.

server

The tools

A separate process that says "here are my tools, here are their arguments" and runs them when asked.

Before MCP, every tool needed a bespoke integration with every AI app: N clients times M tools. With MCP, a tool is written once and every compliant client can use it. That is the entire point.

HOST: THE APP YOU TYPE IN VS Code + Copilot Chat agent mode, your prompt holds one MCP client per configured server client 1 client 2 client 3 SERVERS: SEPARATE PROCESSES JSON-RPC over stdio, one child process per server Playwright MCP servernpx @playwright/mcp, 24 tools filesystem serverread and write project files your own serverJira, test data, env reset A real browserChromium, Chrome, Edge, Firefox The model never touches the browser. It asks a tool, and the server does the driving. One server, many hosts: the same process serves VS Code today and Claude Code tomorrow.
MCP is the wiring, not the intelligence. The host holds a client per server; each server is its own process exposing a list of tools.

Why a tester should care. Your agent stops guessing. Instead of writing a locator from a screenshot or from memory, it asks the browser what is actually on the page, gets back a structured tree with real roles and names, and acts on that. The selectors it produces come from the DOM you are testing rather than from the model's imagination.

This page is the hands-on version. For the long architectural read, the agent team pattern (Planner, Generator, Healer) and the security chapter, see the Playwright MCP and AI agents guide. Here we install it and drive a checkout.

02

What the Playwright MCP server adds

A browser the agent can read as structured text, plus the Playwright code for every action it takes.

The Playwright MCP server is one npm package, @playwright/mcp, that exposes a real browser as MCP tools. Two design choices make it good for testing:

  • It reads the accessibility tree, not pixels. No vision model, no coordinate guessing. Elements arrive with a role, an accessible name and a stable reference.
  • Every result includes the Playwright code it just ran. Your session is not a black box; it is a transcript you can paste into a spec.
Your promptplain English The modelpicks one tool andits arguments call Playwright MCP server acts on the real page returns a fresh snapshot and the code it ran result The browserChromium by default the loop repeats until the task is done, or until you stop it THE WHOLE IDEA: ACT, LOOK AGAIN, DECIDE Every result carries the page state, so the next decision is grounded in what is on screen now.
Each turn is act, look again, decide. The snapshot that comes back is what the model reads before choosing its next call.

The 24 tools, grouped

Verified by asking the server itself over stdio with tools/list, on the version quoted at the top of this page.

GroupTools
Moving aroundbrowser_navigate, browser_navigate_back, browser_tabs, browser_resize, browser_close
Reading the pagebrowser_snapshot, browser_find, browser_take_screenshot, browser_console_messages
Actingbrowser_click, browser_type, browser_fill_form, browser_select_option, browser_hover, browser_press_key, browser_drag, browser_drop, browser_file_upload
Waiting and dialogsbrowser_wait_for, browser_handle_dialog
Networkbrowser_network_requests, browser_network_request
Escape hatchesbrowser_evaluate, browser_run_code_unsafe

The one argument to understand. Every acting tool takes target, and it is a string: either a ref from the latest snapshot, such as e10, or a plain CSS selector such as [data-test="login-button"]. Refs are what an agent uses while exploring. Selectors are what you use when you already know the page, and they are what make the examples below reproducible.

03

Before you install

Node, a project folder, and a browser. Five minutes, once.

Playwright MCP needs Node.js 18 or newer. Check it first, because every other step runs through npx.

terminal
node -v      # v18 or newer
npm -v

Make a folder for this tutorial. The MCP server writes its output (snapshots, screenshots, traces) relative to the workspace you open in VS Code, so a real folder keeps things tidy.

terminal
mkdir playwright-mcp-tta && cd playwright-mcp-tta
code .        # open this folder in VS Code

The code command not found? Open VS Code, press Cmd+Shift+P (Ctrl+Shift+P on Windows and Linux), run Shell Command: Install 'code' command in PATH, then reopen your terminal.

04

Install Playwright in VS Code

The extension gives you a test explorer, a record button and a locator picker. The project gives you a place for the tests the agent writes.

Two separate things share the name. Install both.

  1. The VS Code extension

    In VS Code open the Extensions view (Cmd+Shift+X), search Playwright Test for VSCode, publisher Microsoft, and install it. The identifier is ms-playwright.playwright. From the terminal:

    terminal
    code --install-extension ms-playwright.playwright

    You get a Testing flask icon in the activity bar with run and debug buttons per test, Record new test, and Pick locator.

  2. The project

    Scaffold a Playwright project in the folder you just opened. Accept TypeScript, keep the default tests folder, and let it install the browsers.

    terminal
    npm init playwright@latest

    It creates playwright.config.ts, a tests/ folder with an example spec, and downloads Chromium, Firefox and WebKit.

  3. Check it runs

    terminal
    npx playwright test          # runs the example spec
    npx playwright show-report   # opens the HTML report

    Or click the run arrow next to a test in the Testing view. If the flask icon lists your tests, the extension and the project are talking to each other.

Why install the project at all, when MCP drives its own browser? Because the point is not to click forever in a chat window. It is to end with a committed spec that runs in CI. The agent explores; the project is where the result lands.

05

Install the Playwright MCP server in VS Code

One command, or one small JSON file you can commit for the whole team.

The one-liner

This is the command from the package's own README. It registers the server with your VS Code installation.

terminal
code --add-mcp '{"name":"playwright","command":"npx","args":["@playwright/mcp@latest"]}'

On Windows PowerShell the quoting differs, so use double quotes outside and escape the inner ones, or skip this and use the file below.

The file your team can share

A workspace file is the better option for a batch or a team, because it travels with the repository. Create .vscode/mcp.json:

.vscode/mcp.json
{
  "servers": {
    "playwright": {
      "type": "stdio",
      "command": "npx",
      "args": ["@playwright/mcp@latest", "--test-id-attribute", "data-test"]
    }
  }
}

Why --test-id-attribute data-test is in there. Playwright's getByTestId looks for data-testid by default, and TTACart marks its elements with data-test. Pass this flag and the code the agent generates uses clean getByTestId('login-button') locators for our app instead of falling back to something brittle. Change the value to whatever your own app uses.

Other hosts, same server

The server does not care who is calling. Most clients take this shape:

standard config, works in most clients
{
  "mcpServers": {
    "playwright": {
      "command": "npx",
      "args": ["@playwright/mcp@latest"]
    }
  }
}
Claude Code
claude mcp add playwright npx @playwright/mcp@latest
06

Start it, and check it is really alive

You do not run the server yourself. The host starts it on first use, which is exactly where people get stuck.

There is no npm start here. VS Code launches npx @playwright/mcp@latest as a child process the first time a tool is needed, and shuts it down when you are done. Three checks, in order:

  1. Is the server registered?

    Command Palette, run MCP: List Servers. You should see playwright. From there you can start, stop and view its logs, which is where connection errors show up.

  2. Are you in agent mode?

    Open Chat, and switch the mode picker from Ask to Agent. Ask mode cannot call tools at all. This is the single most common reason "MCP does not work".

  3. Can you see the tools?

    Click the tools icon in the chat input. The Playwright tools should be listed and tickable. Then type a first command and watch the browser window open, because the server runs headed by default.

    you, in agent mode

    Open https://app.thetestingacademy.com/playwright/ttacart/ and tell me what is on the page.

If you want to prove the server works without any editor involved, run it directly. It will sit and wait for JSON-RPC on standard input, which is all the confirmation you need that the package downloads and starts:

terminal
npx @playwright/mcp@latest --help      # prints every flag
npx @playwright/mcp@latest --headless  # starts and waits for a client (Ctrl+C to stop)

One profile, one browser. The server keeps a persistent browser profile per workspace. Two clients pointed at the same workspace will fight over it. Give the second one --isolated or its own --user-data-dir.

07

Automate the TTACart checkout

Open the practice app, log in, add a product, check out, and confirm the order. Every response below is real output from the run described at the top of this page.

The target is TTACart, the practice shop on this site. It behaves like a small real store: six products, a cart badge, a three-field checkout form, a totals page and a confirmation. Credentials are printed on the login page itself, standard_user with the password tta_secret.

SIX PAGES, ONE ORDER 1 Loginlogin-button 2 Productsadd-to-cart-* 3 Cartcheckout 4 Your infofirstName, zip 5 Overviewtotal-label 6 Donecomplete-header standard_user / tta_secretsix products, $7.99 to $49.99tax is 8 percent of the item total every element carries data-test, which is why the server runs with --test-id-attribute data-test
The practice app at /playwright/ttacart/. The whole run below is these six pages, driven by prompts.

How to read the steps. The you box is what you type in agent mode. The tool call is what the model decided to send. The result is what the server sent back, verbatim. In a chat window you will see the same content rendered as a tool card.

1. Open the app and look at it

you

Open https://app.thetestingacademy.com/playwright/ttacart/index.html and show me the page structure.

tool call
browser_navigate { "url": "https://app.thetestingacademy.com/playwright/ttacart/index.html" }
result
### Ran Playwright code
```js
await page.goto('https://app.thetestingacademy.com/playwright/ttacart/index.html');
```
### Page
- Page URL: https://app.thetestingacademy.com/playwright/ttacart/
- Page Title: TTACart - Login
### Snapshot
- generic [ref=e2]:
  - heading "TTACart" [level=1] [ref=e3]
  - generic [ref=e5]:
    - textbox "Username" [ref=e7]
    - textbox "Password" [ref=e9]
    - button "Login" [ref=e10] [cursor=pointer]
  - generic [ref=e11]:
    - heading "Accepted usernames are:" [level=4] [ref=e13]
    - paragraph [ref=e14]: standard_user locked_out_user problem_user performance_glitch_user error_user visual_user
    - heading "Password for all users:" [level=4] [ref=e16]
    - paragraph [ref=e17]: tta_secret

That snapshot is the page as the model sees it: roles, accessible names, and a ref per element. No screenshot was taken and none was needed.

2. Log in

you

Log in as standard_user with password tta_secret.

tool call
browser_fill_form {
  "fields": [
    { "name": "Username", "type": "textbox", "target": "[data-test=\"username\"]", "value": "standard_user" },
    { "name": "Password", "type": "textbox", "target": "[data-test=\"password\"]", "value": "tta_secret" }
  ]
}
result
### Ran Playwright code
```js
await page.locator('[data-test="username"]').fill('standard_user');
await page.locator('[data-test="password"]').fill('tta_secret');
```
tool call
browser_click { "element": "Login button", "target": "[data-test=\"login-button\"]" }
result
### Ran Playwright code
```js
await page.locator('[data-test="login-button"]').click();
```
### Page
- Page URL: https://app.thetestingacademy.com/playwright/ttacart/inventory
- Page Title: TTACart - Products

browser_fill_form beats three browser_type calls. One round trip, one snapshot back, a third of the tokens. Reach for it whenever a form has more than one field.

3. Add a product and watch the badge

you

Add the TTA Practice Backpack to the cart.

tool call
browser_click {
  "element": "Add to cart, backpack",
  "target": "[data-test=\"add-to-cart-tta-practice-backpack\"]"
}
result: the badge appeared in the snapshot
### Ran Playwright code
```js
await page.locator('[data-test="add-to-cart-tta-practice-backpack"]').click();
```
### Page
- Page URL: https://app.thetestingacademy.com/playwright/ttacart/inventory
- Page Title: TTACart - Products
### Snapshot
  - banner [ref=f1e8]:
    - button "Open menu" [ref=f1e9] [cursor=pointer]
    - generic [ref=f1e12]: TTACart
    - link "Shopping cart" [ref=f1e13] [cursor=pointer]:
      - /url: ./cart.html
      - generic [ref=f1e125]: "1"

The "1" under the cart link is the badge. Nothing asserted it, but the agent can see it changed, which is how it knows the click worked. The six add buttons follow one pattern, so the rest of the shop is the same call with a different slug:

the six products and their buttons
Test.allTheThings() T-Shirt (Red)  $15.99  add-to-cart-test-allthethings-tshirt-red
TTA Bike Light                    $9.99   add-to-cart-tta-bike-light
TTA Bolt T-Shirt                  $15.99  add-to-cart-tta-bolt-tshirt
TTA Fleece Jacket                 $49.99  add-to-cart-tta-fleece-jacket
TTA Junior Tester Onesie          $7.99   add-to-cart-tta-junior-tester-onesie
TTA Practice Backpack             $29.99  add-to-cart-tta-practice-backpack

4. Open the cart and start checkout

you

Open the cart and click Checkout.

tool calls
browser_click { "element": "cart link", "target": "[data-test=\"shopping-cart-link\"]" }
browser_click { "element": "Checkout",  "target": "[data-test=\"checkout\"]" }
result
### Ran Playwright code
```js
await page.locator('[data-test="shopping-cart-link"]').click();
```
### Page
- Page URL: https://app.thetestingacademy.com/playwright/ttacart/cart
- Page Title: TTACart - Your Cart

### Ran Playwright code
```js
await page.locator('[data-test="checkout"]').click();
```
### Page
- Page URL: https://app.thetestingacademy.com/playwright/ttacart/checkout-step-one
- Page Title: TTACart - Checkout: Your Information
### Snapshot
  - generic [ref=f3e19]: "Checkout: Your Information"
  - main [ref=f3e21]:
    - textbox "First Name" [ref=f3e25]
    - textbox "Last Name" [ref=f3e26]
    - textbox "Zip/Postal Code" [ref=f3e27]
    - link "Cancel" [ref=f3e29]
    - button "Continue" [ref=f3e32] [cursor=pointer]

5. Fill the form and read the totals

you

Fill the checkout form with Pramod Dutta, zip 560001, then continue.

tool calls
browser_fill_form {
  "fields": [
    { "name": "First Name", "type": "textbox", "target": "[data-test=\"firstName\"]",  "value": "Pramod" },
    { "name": "Last Name",  "type": "textbox", "target": "[data-test=\"lastName\"]",   "value": "Dutta" },
    { "name": "Zip",        "type": "textbox", "target": "[data-test=\"postalCode\"]", "value": "560001" }
  ]
}
browser_click { "element": "Continue", "target": "[data-test=\"continue\"]" }
result: the money is in the snapshot
### Ran Playwright code
```js
await page.locator('[data-test="firstName"]').fill('Pramod');
await page.locator('[data-test="lastName"]').fill('Dutta');
await page.locator('[data-test="postalCode"]').fill('560001');
await page.locator('[data-test="continue"]').click();
```
### Page
- Page URL: https://app.thetestingacademy.com/playwright/ttacart/checkout-step-two
- Page Title: TTACart - Checkout: Overview
### Snapshot
  - generic [ref=f4e31]: TTA Practice Backpack
  - generic [ref=f4e33]: $29.99
  - heading "Payment Information:" [level=4] [ref=f4e36]
  - generic [ref=f4e37]: "TTACard #31337"
  - heading "Shipping Information:" [level=4] [ref=f4e39]
  - generic [ref=f4e40]: Free TTA Express Delivery!
  - generic [ref=f4e44]: "Item total: $29.99"
  - generic [ref=f4e45]: "Tax: $2.40"
  - generic [ref=f4e46]: "Total: $32.39"
  - button "Finish" [ref=f4e51] [cursor=pointer]

This is the moment worth stopping on. The agent now has the item total, the tax and the grand total as text. Tax is 8 percent of $29.99, which is $2.3992, shown rounded to $2.40, and the total is $32.39. That is a real assertion waiting to be written, and it came out of the page rather than out of a spreadsheet.

6. Finish, and keep the evidence

you

Click Finish, confirm the order went through, and take a full page screenshot.

tool calls
browser_click { "element": "Finish", "target": "[data-test=\"finish\"]" }
browser_take_screenshot { "filename": "tta-order-complete.png", "fullPage": true }
result
### Ran Playwright code
```js
await page.locator('[data-test="finish"]').click();
```
### Page
- Page URL: https://app.thetestingacademy.com/playwright/ttacart/checkout-complete
- Page Title: TTACart - Checkout: Complete!
### Snapshot
  - generic [ref=f5e18]: "Checkout: Complete!"
  - main [ref=f5e20]:
    - heading "Thank you for your order!" [level=2] [ref=f5e26]
    - paragraph [ref=f5e27]: Your order has been dispatched, and will arrive just as fast as the TTA Express pony can get there!
    - link "Back Home" [ref=f5e28] [cursor=pointer]

### Result
- [Screenshot of full page](./tta-order-complete.png)
### Ran Playwright code
```js
// Screenshot full page and save it as ./tta-order-complete.png
await page.screenshot({
  fullPage: true,
  path: './tta-order-complete.png',
  scale: 'css',
  type: 'png'
});
```

Six prompts, nine tool calls, one order. The whole run took under ten seconds of browser time.

08

Turn the session into a real spec

The transcript is most of a test already. What it is missing is the assertions, and that part is yours.

Collect every Ran Playwright code block from the run and you have the actions in order. Ask for exactly that:

you

Write the run we just did as a Playwright test at tests/ttacart-checkout.spec.ts. Use the data-test selectors you used, and add assertions for the page title after login, the cart badge, the item total, the tax and the grand total.

tests/ttacart-checkout.spec.ts
import { test, expect } from '@playwright/test';

const APP = 'https://app.thetestingacademy.com/playwright/ttacart';

test('standard_user can buy the backpack', async ({ page }) => {
  await page.goto(`${APP}/index.html`);

  // 1. log in
  await page.locator('[data-test="username"]').fill('standard_user');
  await page.locator('[data-test="password"]').fill('tta_secret');
  await page.locator('[data-test="login-button"]').click();
  await expect(page).toHaveTitle('TTACart - Products');

  // 2. add one product, and prove the badge reacted
  await page.locator('[data-test="add-to-cart-tta-practice-backpack"]').click();
  await expect(page.locator('[data-test="shopping-cart-badge"]')).toHaveText('1');

  // 3. cart, then checkout
  await page.locator('[data-test="shopping-cart-link"]').click();
  await expect(page.locator('[data-test="inventory-item-name"]')).toHaveText('TTA Practice Backpack');
  await page.locator('[data-test="checkout"]').click();

  // 4. the form
  await page.locator('[data-test="firstName"]').fill('Pramod');
  await page.locator('[data-test="lastName"]').fill('Dutta');
  await page.locator('[data-test="postalCode"]').fill('560001');
  await page.locator('[data-test="continue"]').click();

  // 5. the money, which is the part a screenshot cannot check for you
  await expect(page.locator('[data-test="subtotal-label"]')).toHaveText('Item total: $29.99');
  await expect(page.locator('[data-test="tax-label"]')).toHaveText('Tax: $2.40');
  await expect(page.locator('[data-test="total-label"]')).toHaveText('Total: $32.39');

  // 6. finish
  await page.locator('[data-test="finish"]').click();
  await expect(page.locator('[data-test="complete-header"]')).toHaveText('Thank you for your order!');
});
terminal
npx playwright test tests/ttacart-checkout.spec.ts

Read every line before you keep it. The actions come from a real browser and are usually right. The assertions are the agent's guess at what matters, and a guessed assertion is worse than no assertion: it turns green whatever the app does. The six expect calls above exist because a human decided the badge and the three money lines are what this test is for.

Two habits that make the output better

Ask for one flow at a time

"Buy one product and check the totals" produces a clean spec. "Test the shop" produces a 200-line file that tests nothing in particular.

Say which selectors to prefer

Tell it to use data-test attributes or roles. Left alone it may pick whatever matched first, including a CSS path that breaks on the next redesign.

09

The flags that actually change your day

The full list is long. These are the ones a tester reaches for.

FlagWhat it does
--test-id-attribute data-testMakes getByTestId target this app's attribute. Required for clean TTACart locators.
--headlessNo visible window. The server is headed by default, which is what you want while learning.
--browser chrome|msedge|firefox|webkitDrive a real branded browser instead of bundled Chromium.
--device "iPhone 15"Device emulation, straight from Playwright's device list.
--isolatedKeep the profile in memory. Use it for a clean session, or for a second client.
--storage-state auth.jsonStart already logged in, so the agent skips the login page every run.
--output-dir ./mcp-outputWhere snapshots, screenshots and traces are written.
--save-sessionWrite the whole session to the output directory for later review.
--caps vision,pdf,devtoolsAdd the optional tool groups. Off by default to keep the tool list small.
--timeout-action 5000Per-action timeout in milliseconds. Navigation has its own, default 60000.
--secrets .envLoad secrets from a dotenv file rather than pasting them into the chat.
--config path/to.jsonEverything above, as a JSON file you can commit.

Flags go in the args array next to the package name:

.vscode/mcp.json with a fuller setup
{
  "servers": {
    "playwright": {
      "type": "stdio",
      "command": "npx",
      "args": [
        "@playwright/mcp@latest",
        "--test-id-attribute", "data-test",
        "--output-dir", "./mcp-output",
        "--isolated"
      ]
    }
  }
}

Skip the login every time. Log in once by hand with Playwright, save the storage state, then start the server with --storage-state auth.json. Every session then opens on the products page, which saves two tool calls and a lot of tokens on a long exploration.

10

Using it properly

What separates a useful session from an expensive one.

Cost is the real constraint

Every tool result carries a page snapshot, and a snapshot of a busy app is thousands of tokens. Three habits keep that under control:

  • Batch with browser_fill_form instead of typing field by field.
  • Use browser_find when you want one element, rather than re-snapshotting the whole page.
  • Start from a deep link or a saved storage state. Do not make the agent click through login on every run.

Snapshot, not screenshot

The server's own tool description is blunt about it: you cannot act on a screenshot, use browser_snapshot. Screenshots are for a human looking at evidence afterwards. Asking a model to find a button in an image is slower, costlier and less accurate than reading the tree.

Explore with the agent, commit a normal test

An MCP session is not a test suite. It is non-deterministic, it costs money per run, and it needs a model to be available. Use it to explore an unfamiliar flow, to work out selectors, to reproduce a bug, or to heal a broken locator. Then commit a plain Playwright spec that runs in CI in two seconds for free.

Two safety rules

Point it at test environments

The agent will happily complete a purchase, because that is what you asked for. Give it a practice app such as TTACart, or a staging environment with disposable data, never production.

Keep credentials out of the chat

Anything you type becomes model context. Use --storage-state or --secrets. The TTACart password in this tutorial is printed on its own login page, which is exactly why a practice app is the right place to learn.

Treat page content as data, not instructions. The snapshot the agent reads comes from the page. A page that contains text like "ignore your previous instructions" is handing your agent a prompt. That is harmless on TTACart and worth thinking about the first time you point this at a site you do not control.

When it is the wrong tool

Regression suites, anything that must run on every commit, and any check where a stable result matters more than flexibility. Those are scripted tests. MCP earns its keep where a human would otherwise be clicking around and reading the DOM by hand.

11

Try it yourself

Six exercises on TTACart, in the order they get harder.

  1. The happy path, your way

    Buy the TTA Fleece Jacket at $49.99 instead. Confirm the tax the app shows and work out whether it is the same 8 percent.

  2. Two products

    Add the bike light and the onesie, then check the badge reads 2 and the item total is the sum of both prices.

  3. A negative test

    Log in as locked_out_user and ask the agent to report what happens. Then have it write the spec with an assertion on [data-test="error"].

  4. The sort dropdown

    Use browser_select_option on [data-test="product-sort-container"] with "Price (low to high)" and verify the first card is the $7.99 onesie.

  5. Skip the login

    Save a storage state after logging in, restart the server with --storage-state, and confirm the next session starts on the products page.

  6. Break something on purpose

    Write a spec with a wrong selector such as [data-test="user-name"], let it fail, then paste the failure into agent mode and ask it to find the right locator from the live page. That is self-healing with a human in the loop, which is the only kind worth having.

Where to go next. The Playwright MCP and AI agents guide covers the Planner, Generator and Healer pattern and the security chapter. The LangChain agent guide builds the other half: an agent you write yourself, with a Playwright agent as its final project. The Playwright overview explains the driver, the protocols and the browser, context and page model underneath all of this.