The Testing Academy · Masterclass · 3-Part Series
Playwright and AI, from the ground up.
Most testers meet this stack from the top down: they install a plugin, type a prompt, and a browser moves. Then an interviewer asks what the protocol actually is and the whole thing collapses. This series runs the other way. What a language model really does, why it needs an agent around it, what the Model Context Protocol is (and why calling it an API is the answer that ends interviews), one prompt traced all the way to a browser click, and the practice system behind people who actually made the jump.
The machine underneath
Before the plugins and the prompts: what a language model is actually doing when it answers you, where the numbers come from, and the hard ceiling that forces everything built on top of it to exist.
Concept · what a model does
The cat sits on the what.
Read this sentence and finish it: the cat will sit on the ___. Whatever you just did, that is the entire job.
Ask a room of testers to fill that blank and you get mat, floor, chair, table, sofa, terrace. Nobody consulted a database. You split the sentence into pieces, you paid more attention to cat, will and sit than to the, and you produced the word those pieces made most likely. A language model does exactly that and nothing more: it predicts the next token. The textbook definition, a neural network trained on a large corpus, is true and tells you nothing. The guessing game tells you everything, because once you accept that guessing is all it does, every limitation further down this page stops being surprising.
Why it mattersEverything downstream is a workaround for this one sentence. Agents, memory, tools and protocols all exist because prediction alone cannot click a button or read your Jira.
Concept · tokenization
Tokens are not words.
The model never sees your sentence. It sees a list of integers.
Before anything is predicted, the text is cut into tokens, and a token is not a word. It is whatever fragment the tokenizer decided was worth its own id, which means common words are one token and awkward ones split into several. Each token maps to an integer, and that list of integers is the only thing the model ever reads. This is not trivia. Token count is what you are billed for, what fills a context window, and the reason the same task can cost three times more through one route than another, which is exactly the trade-off in section ten.
# the sentence you typed "The cat will sit on the" # what the tokenizer hands the model ["The", " cat", " will", " sit", " on", " the"] [ 464, 3797, 481, 1650, 319, 262 ] # awkward words are not one token each "tokenization" -> ["token", "ization"] "Playwright" -> ["Play", "wright"] Rule of thumb for English prose: one token is about four characters, so roughly 0.75 words. A 1,000 word page is about 1,300 tokens before you have added a single instruction.
Why it mattersTokens are the unit of cost and the unit of memory. Every design decision about prompts, context and tool choice is really a decision about how many of these you are willing to spend.
Concept · attention and meaning
Why king sits next to queen.
Each token carries a long list of numbers, and tokens that mean similar things end up with similar lists.
Every token is turned into a vector, a row of numbers. The transformer architecture, introduced in the paper Attention Is All You Need, adds the part that made this work: the model weighs how much each token should care about every other token in the sentence, so cat and sit pull hard on the blank while the barely registers. The consequence is the useful bit. King and queen land near each other in that space because they appear in similar company, while banana sits far away. Finish the king loves his ___ and the geometry, not a rule, is what makes queen beat banana.
Why it mattersThis is why retrieval works at all. The same geometry that puts queen beside king is what lets a search find the chunk about authentication when the question said sign in.
Concept · weights
The model is a file full of numbers.
Open any open-weights model on Hugging Face and look at the file size. That is the model.
There is no reasoning engine hiding somewhere. A model that you can download is a set of files containing the learned numbers, the weights, and a config describing how to run them. A large open model runs to hundreds of gigabytes, and every one of those bytes is a number that was nudged during training. Loading the model means loading those numbers into memory; inference means pushing your tokens through them. Understanding this kills two confusions at once: why models have a knowledge cutoff (the numbers were fixed when training stopped) and why they cannot look anything up (there is nothing to look up, only weights to multiply).
# the repository is mostly one thing: numbers model-00001-of-000163.safetensors 4.3 GB model-00002-of-000163.safetensors 4.3 GB ... ... model.safetensors.index.json which shard holds which tensor config.json architecture, layer count, dims tokenizer.json text -> integers Total several hundred GB of learned weights. # two consequences fall straight out of this knowledge cutoff -> the numbers froze when training stopped cannot fetch -> there is no lookup step, only matrix maths
Why it mattersA frozen file cannot know your sprint. Everything that makes a model useful on your product is bolted on from outside, which is the whole of part two.
Concept · the ceiling
A brain with no hands.
It can describe a login flow beautifully. It cannot log in.
Put the last four sections together and the limit is obvious. A model reads and writes text and that is the end of its powers. It cannot open a browser, click a button, read a database, call your Jira, or check what time it is. Give it a question about something that happened after its weights froze and it will answer anyway, fluently and wrongly, because producing a plausible next token is the only behaviour it has. That failure mode is not a bug to be patched; it is the machine working as designed. If you want it to do anything, you have to hand it hands.
Why it mattersName the ceiling in an interview and the rest follows naturally. Candidates who can say why a model needs tools tend to explain agents and protocols correctly too, because they built the idea from the bottom.
The connective tissue
How a predictor becomes something that can drive a browser: what an agent adds, what the Model Context Protocol actually standardises, and one prompt followed all the way from your keyboard to a filled login form.
Concept · agents
Model plus memory plus tools.
An agent is not a smarter model. It is the same model with things wired around it.
Take the brain from section five and add two things. Memory, so the conversation survives past one reply and the thing can refer to what it did three steps ago. Tools, so it can act: functions it is allowed to call, whose results come back as more text it can read. That composite is an agent, and the loop it runs is unglamorous: read the situation, choose a tool, call it, read the result, decide whether it is done. The tools are where all the real capability lives. The model only ever decides which one to reach for.
This is also where a very common interview answer goes wrong. ChatGPT is not a language model. It is an application wrapped around one, with memory, a tool belt and a UI. GitHub Copilot, Claude Code, Cursor, Windsurf, Kiro and Codex are the same shape pointed at your editor: AI coding agents. Calling any of them a language model in a room full of engineers signals that the layers have never been separated in your head.
Why it mattersMost rejections here are vocabulary, not ability. Saying the word agent when you mean model, or model when you mean agent, reads as someone who has used the tools without ever opening them.
Protocol · the definition that matters
MCP is not an API.
This single sentence has ended more interviews than any locator question.
The Model Context Protocol is an open standard describing how an AI application and a tool provider talk to each other. It is not a service you call. The distinction is the same one as between HTTP and a weather endpoint: HTTP is the protocol, the endpoint is the thing you call, and confusing them tells everyone you have only used the endpoint. Before MCP, wiring a model into a tool meant bespoke glue for every pairing, so ten agents against ten tools meant a hundred integrations. MCP makes it ten plus ten. Write the server once and every compliant client can use it.
There is an APIs-are-involved caveat worth being precise about, because a good interviewer will push. An MCP server frequently wraps an API underneath, and MCP itself rides a transport. That does not make the protocol an API, any more than a REST service being carried over TCP makes REST a transport protocol. MCP standardises the conversation: how a client discovers what tools exist, how it describes them to the model, how a call is made, and how the result comes back.
An API a specific service you call, with its own shape. You read its docs. You write code against it. Every new pairing is new integration work. MCP a protocol both sides agree to speak, so tools can be discovered and called without bespoke glue. Write the server once, every client can drive it. # the arithmetic that explains why it exists without MCP 10 agents x 10 tools = 100 custom integrations with MCP 10 agents + 10 tools = 20 conformant pieces # the honest caveat, in case you are asked an MCP server often wraps an API underneath, and MCP rides a transport. Neither fact makes the protocol itself an API.
Why it mattersPrecision here is cheap and it is checkable. It takes one sentence to say correctly and it is the fastest signal an interviewer has that you understand the stack rather than the plugin.
Anatomy · where the pieces sit
Host, client, server.
Three names, and people mix up all three. They sit in a fixed order.
The host is the application you have open: VS Code, or whichever editor or desktop app you work in. Inside it runs the client, the part that speaks MCP and holds one connection per server. The server is the tool provider, a separate process that advertises what it can do. For browser work that server is @playwright/mcp, started for you by the client. When it is connected, the client asks it what tools it has and hands that list to the model as part of its context. The model does not magically know Playwright exists; it was told, in the prompt, by the client.
{
"servers": {
"playwright": {
"command": "npx",
"args": ["@playwright/mcp@latest"]
}
}
}
# that is the entire integration. The client starts the server,
# asks for its tool list, and puts that list in the model's context.
# Sanity check it from the agent chat with:
> list the Playwright MCP tools you can see
Why it mattersKnowing the model was told, not trained, removes the magic. It also explains why a server that fails to start leaves the agent cheerfully claiming it cannot browse: the tool list simply arrived empty.
Trace · one prompt, end to end
One prompt, traced.
Type a sentence, a browser fills in a login form. Here is every step in between.
Ask an agent to open a page, enter credentials and submit. What follows is not improvisation. Your sentence is tokenized. The model, already holding the tool list the client gave it, emits a tool call: a structured request naming browser_navigate and its arguments. The client executes it against the server, the browser moves, and the result comes back as text the model reads before deciding on the next call. It repeats: browser_snapshot to see the page, browser_fill_form to type, browser_click to submit. Watch the tool names scroll past in the agent panel and you are watching this loop.
> open https://practice.example.dev/login, sign in as admin / wrong-password, and tell me what happens ● Running browser_navigate url: "https://practice.example.dev/login" ● Running browser_snapshot -> textbox "Username" [ref=e12] -> textbox "Password" [ref=e13] -> button "Sign in" [ref=e14] ● Running browser_fill_form fields: [{ ref: "e12", value: "admin" }, { ref: "e13", value: "wrong-password" }] ● Running browser_click ref: "e14" ● Running browser_snapshot -> alert "Invalid username or password" The login failed as expected: the form returned the error "Invalid username or password" and stayed on /login.
Why it mattersThose lines in the panel are not logging noise, they are the answer. Being able to narrate this loop is the difference between having used a plugin and understanding a protocol.
Trade-off · the other route
The CLI route, and what it costs.
There is a second way to give an agent a browser, and it is cheaper in tokens. Cheaper is not the same as better.
The MCP route spends tokens on ceremony. Every tool in the server's list is described to the model in its context, on every turn, whether or not it gets used. That is the price of discovery. The alternative skips it: install a skill, a folder of written instructions that teaches the agent to shell out to the Playwright command line directly. No server, no tool list, no handshake, roughly a third of the tokens for equivalent work.
The mistake is to read that as a replacement. It is not. MCP gives you structured discovery, which is what you want when the agent must explore a page it has never seen and decide what to do next from a live snapshot. The CLI gives you cheap repetition, which is what you want when the steps are already known and you just need them run. Most real setups end up with both, and a candidate who says one killed the other has usually read a headline rather than used either.
# the command line the skill teaches the agent to drive npx @playwright/cli@latest --help # it is a normal binary, so the agent just runs commands playwright-cli open https://practice.example.dev/login # with the skill installed, this needs no MCP server at all > using the Playwright CLI skill, open the login page and sign in as admin Choose MCP when the agent must look at an unfamiliar page and decide what to do from what it sees. Choose CLI when the steps are known and you are paying for the same ceremony on every single turn.
Why it mattersToken cost is an engineering constraint, not a detail. Explaining when you would pay for discovery and when you would not is a senior answer; declaring a winner is a junior one.
Using it without being fooled
The uncomfortable half. What happens when a tester with no framework grounding lets an agent write a framework, why that is the actual reason interviews go wrong, and the unglamorous practice system that fixes it.
Warning · the trap
The framework it writes in ten minutes.
Drop a skill file into the project, describe your app, and a complete framework appears. That is the demo. It is also the trap.
A skill is a folder of reusable instructions: how a page object should look, what the folder structure is, which fixtures exist, what a good spec reads like. Hand one to a coding agent with a URL and credentials and you will get page objects, fixtures, a config and passing specs faster than you could scaffold them by hand. The output is genuinely good. The problem is not the code. The problem is the moment someone asks you about it.
Because the questions are not hard. Why is that await there and what breaks without it. What is the difference between a fixture and a beforeEach. Why is this locator getByRole instead of a CSS selector. Why does the config set fullyParallel, and what would happen to shared state if it did not. Every one of those has a short, teachable answer, and none of them can be bluffed by someone who watched a generator produce the file. This is the mechanism behind a pattern worth naming plainly: if your fundamentals are thin, AI does not fill the gap, it hides it until an interview or an incident exposes it. The agent is a force multiplier, and a multiplier applied to zero is still zero.
// the agent produced this in about ninety seconds test('user can sign in', async ({ page, loginPage }) => { await loginPage.goto(); await loginPage.signIn('admin', process.env.PW_PASSWORD!); await expect(page.getByRole('heading', { name: 'Dashboard' })) .toBeVisible(); }); Now answer, out loud, without opening the docs: 1. drop any one await. Which line fails, and why not the next one? 2. loginPage arrives as a fixture. What did that replace, and what does it give you that beforeEach does not? 3. why getByRole and not .dashboard-title? 4. the config sets fullyParallel. What breaks if two specs share a logged-in user? 5. that expect auto-waits. For how long, and set where? // Five answers means the code is yours. Fewer means you are // carrying code you cannot defend, which is worse than none.
Why it mattersThis is the actual reason strong-looking candidates get turned down. Not the missing tool, the missing why. The fix is not to stop using agents; it is to be able to defend every line they hand you.
System · the split
Thirty, thirty, thirty.
Ninety days, cut into three equal blocks, in an order that refuses to be rearranged.
The order is the whole point. The language first, because Playwright is a library and you cannot debug a library in a language you cannot read: async and await, promises, array methods, destructuring, modules, then types on top. Playwright second, locators through fixtures, page objects, API testing, up to a framework you assembled yourself. AI and delivery third, MCP, agents, a little retrieval, and the pipeline that runs it all. Put the third block first and you get section eleven: impressive output you cannot explain. The realistic floor is an hour a day, or four hours a week. Below that this does not work, and it is more honest to say so than to sell a shortcut.
Why it mattersMost people fail on sequencing, not capacity. They start with the exciting block, hit the first thing they cannot explain, and conclude they are bad at coding.
System · the real bottleneck
The bottleneck is not talent.
Not the roadmap, not the notes, not the course, and not a fear of coding. The gap is reps.
People who make this jump and people who do not are separated by one measurable thing: how much code they have typed. Roughly ten small exercises a day for ninety days puts you somewhere near four hundred, and four hundred is enough for recall to stop being a performance. You will not reach it, because nobody does, and that is fine: half of four hundred still changes what you can do under interview pressure, where the failure is almost never conceptual. It is sitting in front of an editor having written five queries this year and being asked for the sixth.
Two rules make the reps count. Every exercise goes to a public repository, because a commit history is the one claim on a resume a reviewer can check in ten seconds, and it is dated, continuous evidence rather than an assertion. And you write it before you ask an agent. Use the agent afterwards, to review, to explain what you missed, to show a cleaner version. That order keeps every tool from section two onward as a multiplier instead of a substitute.
# day 34 of 90 · topic: Playwright locators # rule: write it first, THEN ask the agent to review it 01 getByRole for every button on the practice page 02 the same page with getByLabel, note where it is better 03 getByText exact vs substring, prove the difference 04 chain a locator, then do it again with .filter() 05 first / last / nth on a repeated row, then explain the risk 06 a table cell by row text plus column header 07 make a deliberately strict-mode-violating locator, read the error 08 fix it three different ways 09 the same flow with a bad CSS selector, then break it on purpose 10 write the assertion that would have caught the break git commit -m "day 34: locators, 10 reps" git push # 90 days of this is a dated, checkable record of the one # claim on your resume that most candidates cannot evidence.
Why it mattersInterviewers are not looking for perfect, they are looking for reachable. Someone who gets sixty percent of the way to an answer unaided reads as someone who will get to a hundred on the job. Someone blank reads as someone who has never practised.
System · the evidence
What goes on the resume.
Exercises prove you practise. Projects prove you can finish. You need both, and they are not the same artefact.
Aim for ten to fifteen Playwright projects that each do one thing completely: a framework with page objects and fixtures, an API suite, a data-driven suite, visual checks, a sharded CI run, a custom reporter. Then four or five where AI is load-bearing, not decorative: an MCP-driven exploratory run, a retrieval bot answering over your own test cases, a triage agent that reads a failure and files a structured summary, a self-healing locator experiment. The second group is what makes a profile unusual right now, and each one is also a story you can tell for ten minutes about a decision you made and a trade-off you understood.
One language choice, since it comes up every time. TypeScript, not Python, for Playwright specifically. The concepts port either way and Python is a perfectly good language, but the majority of Playwright roles advertise TypeScript, the examples and the agent-generated code you will meet are written in it, and the point of this exercise is employability. Learn the other one later, when it costs you a weekend instead of a quarter.
Group A ten to fifteen, each finished, each with a README 01 framework: page objects, fixtures, config, parallel 02 API suite with request context and schema checks 03 data-driven suite from CSV and JSON 04 visual regression on a handful of stable pages 05 sharded CI run with a merged report 06 custom reporter that posts a summary somewhere 07 auth state reuse, storage state, and a session test ... Group B four or five where the AI is doing real work 01 MCP-driven exploratory run over an unfamiliar app 02 retrieval bot answering over your own test cases 03 triage agent: read a failure, file a structured summary 04 self-healing locator experiment, with the honest results 05 an agent that turns a ticket into a draft spec For each one, be ready to say in two minutes: what it does · one decision you made · one thing you would change · what broke while you were building it
Why it mattersGroup B is the differentiator today and will be table stakes shortly. Build them while they still read as unusual, and make sure every one has a trade-off you can discuss rather than a tool you can name.