Where this sits
The previous class split the API layer into an API helper, a service object with fixtures, and specs that only assert, then added JSONPath for reading responses. One piece was deliberately left for today: a check on the shape of a response, run before any assertion touches it.
What a schema checks
A JSON schema is a JSON document that describes another JSON document: which fields exist, what type each one is, and which are required. It says nothing about values. It is the difference between "the response has a firstname string" and "the first name is Pramod".
The analogy from the session: humans have two eyes, one nose, one mouth and two legs. That is a schema. If something with four legs arrives, you do not need to check its name to know it is not what you asked for.
Everything in the response either matches the skeleton or the test stops there. That is why this is also called contract testing: the shape is a contract between the developer and you, and the test fails when the developer changes it without telling anyone.
The schema, and how it was made
Take a real response, paste it into an online JSON-to-schema converter (the one used in class was Liquid Technologies'), and you get a schema back. Tighten it, save it as a .schema.json file under test data. This is the one that shipped:
{
"$schema": "http://json-schema.org/draft-07/schema#",
"title": "CreateBookingResponse",
"type": "object",
"required": ["bookingid", "booking"],
"additionalProperties": false,
"properties": {
"bookingid": { "type": "integer" },
"booking": {
"type": "object",
"required": ["firstname", "lastname", "totalprice", "depositpaid", "bookingdates"],
"additionalProperties": false,
"properties": {
"firstname": { "type": "string" },
"lastname": { "type": "string" },
"totalprice": { "type": "number" },
"depositpaid": { "type": "boolean" },
"bookingdates": {
"type": "object",
"required": ["checkin", "checkout"],
"additionalProperties": false,
"properties": {
"checkin": { "type": "string", "format": "date" },
"checkout": { "type": "string", "format": "date" }
}
},
"additionalneeds": { "type": "string" }
}
}
}
}
Four keywords carry the whole thing:
| Keyword | What it enforces |
|---|---|
type |
string, number, integer, boolean, object, array, null |
required |
fields that must be present. An optional field is simply not listed here; additionalneeds is the example |
additionalProperties: false |
nothing the schema does not name may appear. Set at every level here, so an unannounced new field fails too |
properties |
the fields themselves, nested as deep as the response goes |
The library that checks a document against this is Ajv, "Another JSON schema validator", installed with ajv-formats so that "format": "date" actually does something.
These are the three hard checks, and each was run against the shipped schema before this page was published:
| Drift | Ajv's verdict |
|---|---|
An extra field, booking.vip |
/booking must NOT have additional properties |
A wrong type, totalprice: "Pramod" |
/booking/totalprice must be number |
A missing field, no firstname |
/booking must have required property 'firstname' |
Asked in class: if we already declare the type in a TypeScript interface, why a schema? Because an interface is erased at compile time. parseJsonResponse<CreateBookingResponse>() is a cast, not a check: a response with no bookingid still satisfies it. The schema is the runtime counterpart, and the repository's own comment says to pair them, not replace one with the other.
The validator and the three-line spec
The check lives in one shared util so every spec can use it:
const ajv = addFormats(new Ajv({ allErrors: true }));
export class SchemaValidator {
static validate(schema: object, data: unknown): ValidationResult {
const validate = ajv.compile(schema);
const valid = validate(data) as boolean;
return { valid, errors: (validate.errors ?? []).map(SchemaValidator.describe) };
}
static assertValid(schema: object, data: unknown, label = 'response'): void {
const { valid, errors } = SchemaValidator.validate(schema, data);
if (!valid) {
throw new Error(`[SchemaValidator] ${label} does not match schema:\n - ${errors.join('\n - ')}`);
}
}
}
allErrors: true reports every violation at once instead of stopping at the first, and describe splices the offending key back into the path, because Ajv's raw errors for missing and extra properties carry an empty path and are unreadable in a report.
The spec that uses it is the whole point of the earlier layering. It is three lines of work:
test('@P0 @schema Level 5 - POST /booking matches the create-booking schema', async ({ bookingApi }) => {
const body = await bookingApi.createBooking(buildBookingFromGenerator());
SchemaValidator.assertValid(createBookingSchema, body, 'POST /booking');
await bookingApi.deleteBooking(body.bookingid);
});
It passed against the live API before this page went up.
The gotcha, from the repository's README: a green schema test does not prove the schema is right. An empty schema {} validates everything. Confirm the schema fails on something at least once, which is exactly what the drift table above does.
The order API tests run in
Asked twice, so here it is as a rule.
Two questions from the floor. What is WireMock, and how is it different from Ajv? Not related. WireMock mocks an API, which you need when the backend is not built yet and the UI team cannot wait: database first, then backend, then API, then UI is the usual build order, and a mock lets the UI start early. Ajv validates a real response. Is retry supported? Yes, built into Playwright's config; nothing to add.
Building it with AI, and what had to be cut
The whole feature was scaffolded with a coding agent, and the method was the same one from the dotenv class: plan mode first, never auto-plan, read the plan, then auto-accept. The prompt named where each piece must live, the schema in test data, the validator in a shared util, the spec in a 05 folder, and asked the agent to show its changes before making them.
The agent's questions were answered as it planned (one schema, for create booking only; strict additionalProperties). Then the review found the usual excess:
- It wrote its own rejection handling with reasons per missing field. Cut. A failed assertion with Ajv's message is enough.
- It generated its own booking payload instead of using
buildBookingFromGeneratorfrom the previous class. Cut, and the instruction was explicit: only use the utilities that already exist.
That second cut is the recurring lesson of this batch. An agent that does not know your utilities will write new ones. Tell it what exists.
Ponytail
Ponytail is a code optimiser that trims what the session called AI slop: the class, the wrapper and the helper an agent writes when a function would do. It installs as a plugin (Copilot, Claude Code, Codex, Gemini and others), and the command is ponytail review with a file name.
Run against last class's 73-line lifecycle spec, the result is 33 lines with the same three tests. What went, and why each cut was safe, is recorded in the repository:
| Cut | Why it was safe |
|---|---|
test.step around single calls |
The trace already lists every request with timings |
testInfo.attach of the body |
The same trace already holds it |
Ten log.info lines |
Each restated its own step name |
buildBooking |
buildBookingFromGenerator gives ordered, today-relative dates |
An explicit bookerToken |
Omitting it uses the managed token, which re-auths on a 403 |
Those cuts are only safe because the config sets trace: 'on'. Turn tracing off and the attachments stop being redundant.
The shipped version of the middle test:
test('update the booking, then read it back', async ({ bookingApi }) => {
// No token argument: BookingApi re-auths on a 403 and retries. Passing one opts out.
const updated = await bookingApi.updateBooking(
bookingId,
buildBookingFromGenerator({ firstname: 'E2E', lastname: 'Updated', totalprice: 950 }),
);
expect(updated).toMatchObject({ lastname: 'Updated', totalprice: 950 });
// The GET, not the PUT echo, is what proves it persisted.
expect((await bookingApi.getBooking(bookingId)).lastname).toBe('Updated');
});
A correction made live, and worth keeping. The first optimised file collapsed create, update and delete into a single test, and the room was told that is not wanted. Then the instructor checked: that collapse was the agent's own doing, not Ponytail's. Ponytail keeps your structure. The final file keeps three tests, because describe.serial reporting each stage as its own pass or fail is a reporting decision, not machinery to trim. Use the tool sparingly regardless, and read what it removed.
Two related tips from the same stretch: to stop a coding agent naming itself in your commits, add a global rule that it must never mention itself in commit messages. And the tiered-model skill shown routes trivial file creation to a small model and hard reasoning to a large one, so you are not spending the expensive model on boilerplate.
Workers, RAM and the question that became a document
Someone asked how to run in parallel and rerun failures, and the answer grew into a measured guide that now lives in the repository. The numbers below are from that guide, measured on the batch framework, not estimated.
What a worker is. One isolated lane with its own browser. workers: 1 runs everything nose to tail. Leave workers unset with fullyParallel: true and Playwright picks a default: half the logical cores, capped at the number of tests. On the 16-core machine used here, one schema test ran on one worker and four end-to-end tests ran on four; the "four" the session mentioned was that cap, not a fixed number.
The session's rule of thumb was about 1 GB per worker, rising to 1.5 or 2 GB for complex tests. The measurement says 0.7 GB headed and 0.4 GB headless on top of roughly 1 GB fixed. Keep the rule of thumb as your safety margin; use the measurement to understand where the ceiling comes from.
| System RAM | Free for tests | RAM ceiling, headed | RAM ceiling, headless | Cores / 2 | Practical pick |
|---|---|---|---|---|---|
| 8 GB | about 4 GB | 4 | 7 | 2 to 4 | 2 headed, 4 headless |
| 16 GB | about 11 GB | 14 | 25 | 4 to 6 | 4 headed, 6 headless |
| 32 GB | about 26 GB | 35 | 62 | 6 to 8 | 8, RAM is not the limit |
Read it this way: RAM only binds on 8 GB machines. From 16 GB up you run out of cores first, so buying RAM to get more workers is the wrong upgrade. The formula behind the table:
workers = min(
(free RAM in GB - 1) / 0.7, # 0.4 if headless
logical cores / 2
)
Setting it, in the two forms worth knowing:
npx playwright test --workers=4 # one run
npx playwright test --workers=50% # a share of cores, portable across machines
npx playwright test --workers=1 # serial: the first thing to try when a test passes alone and fails in the suite
describe.serial blocks run on one lane no matter what --workers says. That is the design of the lifecycle spec, and it is why the three booking tests cannot be sped up by adding workers.
Production numbers, for scale. The TechCon suite is about 12,000 tests on ten Jenkins agents, each a 64 GB machine running eight workers, and a full regression takes three to four hours. Eight, not more, because complex tests cost more than the average. Someone in the room reported 1,300 tests taking 8 to 12 hours on six agents; the reply was that the tests themselves are slow and need splitting.
Sharding. --shard=k/n slices a suite across n machines and combines with workers rather than replacing them. This batch does not use it, because the usual Jenkins layout already splits by module: one job per module, each job fanned out to its own agents (the session used the older master and slave terms), each agent running its workers. The split happens before Playwright ever sees the tests.
Tasks and announcements
Next class: the last layer, the AI layer, added to the framework as version 2: root cause analysis on failures, self-healing, prompt to spec, a PR review bot, and smart tests. After that the framework supports API, page objects, Cucumber, headless, workers and AI in one place.
Next week: Jenkins parts 3 and 4 and GitHub Actions, led by Deepak, running this framework on both.
Being shared
- The push with the Ajv level, the Ponytail spec beside the original, and the worker guide.
- A separate video on keeping a coding agent's name out of your commits.
- The Cucumber videos are already in the platform for anyone adding BDD to the framework.