Where this sits
The advanced framework is complete on the classic side: API levels, Cucumber, page objects, the custom reporter. Today adds the piece the instructor called the end game, an AI layer, and the point was made early: nobody in the room has LLM API access at work yet, and within a month or two everyone will. The person adding this layer to a company framework should be you.
Two rules before any code:
- A coding-agent subscription is not an API key. Copilot, Claude Code, Codex, Amazon Q, Kiro, Kilo, Antigravity: none of them hand you a key. You buy API access separately.
- Do not use any of this at work without approval, open-source model or not. Learn it here first.
Every LLM API is the same curl
The Groq playground was opened first because it shows the request in the raw. A POST to chat/completions, an API key in the header, a JSON body with the model, a messages list where your question sits under role: user, and knobs: temperature (1 allows creative answers), a max-token cap, stream, a reasoning effort. Ask "what is 2 plus 2", get 4.
Then OpenRouter, then DeepSeek: the same shape every time. Only the base URL and the key change. Why? OpenAI shipped the format first, so everyone copied it. That is the whole reason one client can cover every provider.
Getting a key, as covered in class: Groq is free at playground scale (not for big responses), OpenRouter needs a minimum top-up of about 5 dollars and has free, open-source and paid models, DeepSeek is cheap and paid, and the Claude API is paid. The instructor ran the demo on DeepSeek; Groq is enough for you to follow along. LLMs are not free; only the small tier is.
What the room was asked to imagine
With one client that can talk to a model, everything else is a prompt:
- One test result in, an RCA verdict out, attached to the custom report.
- Two runs in, the flaky tests named.
- A failure in, bug triage out: severity, priority, a comment.
- Instead of Faker, a data generator: "give me three bookings for this endpoint, JSON only". The playground did exactly that live.
Two honest limits, both from the instructor. It works from what the model knows plus what you send it, not from reading your whole codebase (asked by Kalyan), and it gets you 70 to 80 percent of the way, not 100. At work the same pattern already generates customers, vehicles and VIN numbers.
The build, in the order it happened
- The stub agents already sitting in
src/aiwere deleted so the class could watch the layer appear from nothing. (The brief records that this broke the build, because the reporter still imported them; rebuilding them was step one for the agent.) - A dictated one-paragraph request was handed to the coding agent with one instruction: make this prompt better and save it as
docs/ai-factory.prompt.md. The result separates the objective, the five providers, the seams the reporter already exposes, the factory contract, each agent's definition of done, hard constraints and a working agreement: plan first, one agent at a time, stop for review. - The agent asked for a key.
DEEPSEEK_API_KEYwent into.env. - First deliverable:
LLMClient, the provider registry,agentFactory, the data generator,src/tests/aiTest/CustomDataGen.spec.ts, and the generated bookings showing in the report's AI Data tab. - A Ponytail review over
src/ai, the way the batch reviews everything: it found 8 of 17 exports with no importer and 4 dead symbols, and trimmed them (asked by Neet: Ponytail is a code-optimisation tool that installs beside any coding agent; yes, it works for Python too, and yes, on an office laptop, it is open source). - RCA and flaky agents with demo specs, then the self-healing agent with its own report tab, then the README and
.env.example. A security review and an API review were queued as the session ended.
The factory contract
The only HTTP code in the layer is LLMClient, which speaks two dialects: the OpenAI /chat/completions shape for DeepSeek, OpenRouter, Groq and OpenAI, and Anthropic's /v1/messages. Selection is by environment, never by import:
AI_PROVIDER |
Default model | Key |
|---|---|---|
deepseek (default) |
deepseek-chat |
DEEPSEEK_API_KEY |
openrouter |
deepseek/deepseek-chat |
OPENROUTER_API_KEY |
groq |
llama-3.3-70b-versatile |
GROQ_API_KEY |
openai |
gpt-4o-mini |
OPENAI_API_KEY |
anthropic |
claude-sonnet-5 |
ANTHROPIC_API_KEY |
An agent is a config object. The factory adds the transport, the JSON extraction, the schema check with the same Ajv SchemaValidator from Wednesday's class, and one retry:
export type AgentResult<TOutput> =
| { available: true; data: TOutput; provider: string; latencyMs: number }
| { available: false; reason: string };
export function createAgent<TInput, TOutput>(spec: AgentSpec<TInput>, client = new LLMClient()) {
return {
async run(input: TInput): Promise<AgentResult<TOutput>> {
if (!client.isAvailable) {
// The normal path in CI. Not an error.
return { available: false, reason: 'no API key configured' };
}
let prompt = spec.buildPrompt(input);
for (let attempt = 1; attempt <= 2; attempt++) {
const completion = await client.complete({ system: spec.system, user: prompt, ... });
const parsed = extractJson(completion.text);
const { valid, errors } = SchemaValidator.validate(spec.schema, parsed);
if (valid) return { available: true, data: parsed as TOutput, ... };
if (attempt === 2) return { available: false, reason: `output failed schema: ${errors.join('; ')}` };
prompt = `${prompt}\n\nYour previous reply did not match the required schema:\n` +
errors.map((e) => `- ${e}`).join('\n') + `\nReply again with JSON only.`;
}
},
};
}
Two guarantees every caller leans on: data is schema-valid or available is false, with no third state; and a missing key returns "unavailable" instead of throwing, so the suite runs unchanged without one.
The retry is load-bearing, not decoration. On the first live run DeepSeek returned an additionalneeds string longer than the schema's 60-character cap. The factory fed the exact Ajv error back and the second attempt was clean. The same guard has since fired on the flaky summary's length cap and on the RCA fixes list. Without it a malformed payload would have reached the API and come back as a confusing 500.
extractJson exists because models wrap JSON in prose or in a ```json fence often enough that a bare JSON.parse is a bug: it strips the fence, finds the first { or [, and parses to the last } or ].
Agent 1: the data generator
Faker gives random values. This gives situated ones: each booking carries a scenario label saying what it exercises, the part a test author would otherwise invent by hand. The prompt is the rules, not the transport:
export const dataGenAgent = createAgent<DataGenInput, GeneratedBookings>({
name: 'data-generator',
system: 'You generate test data for an API test suite. You reply with JSON only: no prose, ' +
'no markdown fences, no explanation. Every value must pass a strict JSON Schema.',
schema: schema as object,
temperature: 0.7, // at 0 the model returns the same names every run
maxTokens: 1500,
buildPrompt: ({ count, focus }) => `
Generate ${count} booking payloads for the Restful Booker API.
Return exactly this JSON shape: {"bookings":[{"firstname":"","lastname":"","totalprice":0,"depositpaid":true,
"bookingdates":{"checkin":"YYYY-MM-DD","checkout":"YYYY-MM-DD"},"additionalneeds":"","scenario":""}]}
Rules:
- checkin and checkout are YYYY-MM-DD, and checkout is strictly after checkin.
- totalprice is a whole number between 0 and 100000.
- scenario names what the row exercises, e.g. "minimum price" or "long stay".
- Vary the rows. Do not return ${count} near-identical bookings.
${focus ? `- Focus on: ${focus}` : ''}`.trim(),
});
The spec asks for three bookings focused on "edge cases around price and stay length", re-validates against ai-booking-payload.schema.json, attaches the result as ai-data, then posts each one to the real POST /booking, expects 200, and deletes what it created. With no key the file skips, which is the CI path. The report the class saw: provider deepseek/deepseek-chat, 1,498 ms, rows labelled "minimum price, one-night stay" (a one-letter name at price 0) and a maximum-price stay at 100,000.
Nothing in the spec asserts on generated prose. Schema validity, checkout > checkin, HTTP status: those decide. A test whose result rides on a sampled token is not a test. The one check the schema cannot make, a reversed stay, is asserted explicitly because the API would not catch it either.
Agents 2 and 3: RCA and flaky, from the reporter
RCA. The reporter already rendered an AI Verdict tab from a RcaVerdict shape, so the agent fills that shape rather than inventing one: severity from critical | high | medium | low, priority from P0 to P3, a rootCause of 20 to 600 characters, one to four fixes. The system prompt is the interesting line: "You are a senior test automation engineer triaging a failed Playwright test... Never suggest deleting or skipping the test to make it pass." Triage is not a fifth agent: severity and priority ride in the same call, which is cheaper and keeps the rating consistent with the explanation that justified it.
Flaky. The build-to-build diff was already deterministic and already worked with no key. The model adds only the summary, and only when something actually flipped. Real flakiness cannot be demonstrated on demand, so demoState.ts counts runs per test and three tests fail on even-numbered runs: run one is ten green, run two is seven green and three red. The report the class saw compared two builds and showed 3 flaky, 4 failing, 14 total, each flaky test listed, and an AI summary observing that all three sit in one file, share a passed-to-failed transition and have names describing timing, so shared environment or timing instability is likelier than a product regression.
Read one verdict closely and the 70 percent point becomes concrete. For flaky: booking search returns stale results the verdict said the assertion at line 32 "expecting the booking search result count toBe(1) received 0 because the search index was stale". The real line is expect(run % 2).toBe(1) with that sentence as its message; there is no search index. The model narrated the failure message as fact, and the fixes it proposed (poll until the booking appears, seed data first, wait for an indexing event) are good advice for the bug it imagined. Useful as a starting point, never as a verdict a test depends on.
Agent 4: self-healing locators
The demo uses a realistic near-miss: [data-test="user-name"] on the login page, where the real attribute is username. Playwright captured the page, the interactive elements go to the model, and every suggestion is re-run against the live page before it reaches the report.
const DEAD = '[data-test="user-name"]'; // real attribute is "username"
const INTENT = 'the username input on the login form';
try {
await page.locator(DEAD).fill('standard_user', { timeout: 5_000 });
} catch (error) {
if (!isLocatorFailure(error)) throw error;
// The page is still open here, which is the only moment the DOM exists
// to check candidates against. The reporter runs too late.
const report = await healLocator(page, DEAD, INTENT, { requires: 'editable' });
await testInfo.attach('self-heal', { body: JSON.stringify(report, null, 2), contentType: 'application/json' });
// Assert on the verification, never on the model's wording.
expect(report.verified.length).toBeGreaterThan(0);
expect(report.verified[0].matchCount).toBe(1);
// Prove the suggestion is usable, not merely well-formed.
await page.locator(report.verified[0].selector).fill('standard_user');
await expect(page.locator(report.verified[0].selector)).toHaveValue('standard_user');
}
Candidates come back with a strategy from css | testid | role | text | label | placeholder and a reason each. Against the real login page the agent ranked [data-test="username"] first, reasoning that user-name in the dead selector is the element's id, not its data-test. The report then shows both lists, verified and rejected, with the fix applied inside the test. The spec on disk is untouched.
requires: 'editable' exists because of a real miss. The first version checked only "resolves to exactly one element", reported four verified candidates, and the fill still failed with Element is not an <input>. A heading matches uniquely too. Verification now rejects anything that cannot take the intended action, rather than ranking it lower.
Running it, and why CI has no key
# data generator, the only demo that is green by design
npx playwright test --project=ai CustomDataGen
# flakiness: run 1 is all green, run 2 turns exactly 3 tests red
rm -f reports/ai-demo-state.json
AI_DEMO=1 npx playwright test --project=ai FlakyDemo
AI_DEMO=1 npx playwright test --project=ai FlakyDemo
# a real failure for the RCA agent, and a dead locator for self-healing
AI_DEMO=1 npx playwright test --project=ai RcaDemo
AI_DEMO=1 npx playwright test --project=ai SelfHealDemo
The demos contain deliberate failures, so they hide behind AI_DEMO. Without it the suite is 32 passed, 13 skipped; with no key at all, 31 passed. Both exit 0. The config gained a third project, ai, with a 180-second timeout because a model round trip plus a retry can pass the 60-second default, and the chromium project now ignores aiTest/ the way it ignores apisTests/. CI sets no key on purpose: a key there bills per build and makes the result depend on a third party being up. The tabs render empty and the suite stays green.
What comes next: the quality gate
Everyone has Copilot and Claude Code now, and the instructor's observation from work is that people raise four or five PRs a day of AI-written code and quality falls. The fix is a quality gate that fails the PR: ESLint for the JavaScript and TypeScript rules, type checks, lint, custom rules such as no page class with 200-plus locators, Ponytail inside the gate (asked by Shiva), and a custom extension that enforces it on pull requests. "Raise one PR, but it should be good." That is Monday's class, the last morning session, after which the framework counts as complete and Deepak takes over for Jenkins and GitHub Actions. Venkat's "what is a PR and how do I raise one" is answered there too.
Asked at the end by Deepti: which design pattern does the framework use? The answer to give: the Page Object Model, and say model, not pattern. It is a simple model built on the OOP concepts the batch already covered; do not dress it up as a formal design pattern in an interview.
Tasks and announcements
Before the next class
- Create a key: Groq (free, groq.com) is enough; DeepSeek or OpenRouter (5 dollar minimum) are optional.
- Pull the repository, put the key in
.envas the provider's*_API_KEY, and runnpx playwright test --project=ai CustomDataGen. Open the custom report and find the AI Data tab. - Then add one agent of your own: a schema file, a prompt, a
createAgentcall. No transport code. - Do not point any of this at company code or data without approval.
Schedule
- Monday morning, short class (it is Ganesh Chaturthi): the quality gate. Last morning session.
- After that: Deepak runs the framework on Jenkins and GitHub Actions.
- Evening sessions, Friday or Sunday at 8 pm, starting next week: Playwright CLI, Playwright MCP, the Playwright AI agents (generator, self-healer, planner) revisited, Selenium to Playwright migration, plus three certification classes on MCP, MCP advanced and subagents. Around ten sessions across September and October. "We are not going anywhere."
- Cucumber videos are already in the batch material. Venkat's API tool question goes to Meeti and Deepak.
Batch news. Sanjay wrote in during the wrap-up: three offers in hand, after being stuck at the same organisation for years, and he put the credit on the Playwright and AI batch. Congratulations from the instructor and the room. This is what the LinkedIn habit and the framework work are for.
Repository: AdvancePlaywrightFramework2x received docs/ai-factory.prompt.md, src/ai/ (client, providers, factory, four agents), four output schemas, the aiTest project with four specs, src/utils/selfHeal.ts, the Self-Heal tab in the reporter, README sections for the AI layer, and .env.example entries for every provider key, all blank.