The Testing Academy · Class Notes Wednesday, 9 September (IST)
Live class · study guide

The shape check that runs before every assertion, and how many workers your laptop can really take

The API layer gets its last piece: a JSON schema check with Ajv that fails the moment a field is dropped, renamed or retyped, and a fixed order for API tests so that check runs before anything else. Then two things that only exist because AI wrote the first draft: what had to be cut, and what Ponytail does to a spec. Finally the question that turned into a measured document: how many workers can my machine actually run.

By Pramod Dutta, The Testing Academy. Study notes from the live Playwright 2x class, rebuilt from the session recording, the batch repository pushed during the session and the Eraser deck on screen. The schema, validator, spec and worker guide are quoted from the repository. Before publishing, the schema spec was run against the live RESTful Booker API, the strict schema was fed an extra field, a wrong type and a missing field to capture Ajv's exact messages, and the worker numbers were read from the repository's measured guide rather than from memory. Where the session's rules of thumb and the measured numbers differ, both are shown.

01

Where this sits

The previous class split the API layer into an API helper, a service object with fixtures, and specs that only assert, then added JSONPath for reading responses. One piece was deliberately left for today: a check on the shape of a response, run before any assertion touches it.

02

What a schema checks

A JSON schema is a JSON document that describes another JSON document: which fields exist, what type each one is, and which are required. It says nothing about values. It is the difference between "the response has a firstname string" and "the first name is Pramod".

The analogy from the session: humans have two eyes, one nose, one mouth and two legs. That is a schema. If something with four legs arrives, you do not need to check its name to know it is not what you asked for.

Everything in the response either matches the skeleton or the test stops there. That is why this is also called contract testing: the shape is a contract between the developer and you, and the test fails when the developer changes it without telling anyone.

03

The schema, and how it was made

Take a real response, paste it into an online JSON-to-schema converter (the one used in class was Liquid Technologies'), and you get a schema back. Tighten it, save it as a .schema.json file under test data. This is the one that shipped:

JSON
{
    "$schema": "http://json-schema.org/draft-07/schema#",
    "title": "CreateBookingResponse",
    "type": "object",
    "required": ["bookingid", "booking"],
    "additionalProperties": false,
    "properties": {
        "bookingid": { "type": "integer" },
        "booking": {
            "type": "object",
            "required": ["firstname", "lastname", "totalprice", "depositpaid", "bookingdates"],
            "additionalProperties": false,
            "properties": {
                "firstname": { "type": "string" },
                "lastname": { "type": "string" },
                "totalprice": { "type": "number" },
                "depositpaid": { "type": "boolean" },
                "bookingdates": {
                    "type": "object",
                    "required": ["checkin", "checkout"],
                    "additionalProperties": false,
                    "properties": {
                        "checkin": { "type": "string", "format": "date" },
                        "checkout": { "type": "string", "format": "date" }
                    }
                },
                "additionalneeds": { "type": "string" }
            }
        }
    }
}

Four keywords carry the whole thing:

Keyword What it enforces
type string, number, integer, boolean, object, array, null
required fields that must be present. An optional field is simply not listed here; additionalneeds is the example
additionalProperties: false nothing the schema does not name may appear. Set at every level here, so an unannounced new field fails too
properties the fields themselves, nested as deep as the response goes

The library that checks a document against this is Ajv, "Another JSON schema validator", installed with ajv-formats so that "format": "date" actually does something.

These are the three hard checks, and each was run against the shipped schema before this page was published:

Drift Ajv's verdict
An extra field, booking.vip /booking must NOT have additional properties
A wrong type, totalprice: "Pramod" /booking/totalprice must be number
A missing field, no firstname /booking must have required property 'firstname'

Asked in class: if we already declare the type in a TypeScript interface, why a schema? Because an interface is erased at compile time. parseJsonResponse<CreateBookingResponse>() is a cast, not a check: a response with no bookingid still satisfies it. The schema is the runtime counterpart, and the repository's own comment says to pair them, not replace one with the other.

04

The validator and the three-line spec

The check lives in one shared util so every spec can use it:

TypeScript
const ajv = addFormats(new Ajv({ allErrors: true }));

export class SchemaValidator {
    static validate(schema: object, data: unknown): ValidationResult {
        const validate = ajv.compile(schema);
        const valid = validate(data) as boolean;
        return { valid, errors: (validate.errors ?? []).map(SchemaValidator.describe) };
    }

    static assertValid(schema: object, data: unknown, label = 'response'): void {
        const { valid, errors } = SchemaValidator.validate(schema, data);
        if (!valid) {
            throw new Error(`[SchemaValidator] ${label} does not match schema:\n  - ${errors.join('\n  - ')}`);
        }
    }
}

allErrors: true reports every violation at once instead of stopping at the first, and describe splices the offending key back into the path, because Ajv's raw errors for missing and extra properties carry an empty path and are unreadable in a report.

The spec that uses it is the whole point of the earlier layering. It is three lines of work:

TypeScript
test('@P0 @schema Level 5 - POST /booking matches the create-booking schema', async ({ bookingApi }) => {
    const body = await bookingApi.createBooking(buildBookingFromGenerator());
    SchemaValidator.assertValid(createBookingSchema, body, 'POST /booking');
    await bookingApi.deleteBooking(body.bookingid);
});

It passed against the live API before this page went up.

The gotcha, from the repository's README: a green schema test does not prove the schema is right. An empty schema {} validates everything. Confirm the schema fails on something at least once, which is exactly what the drift table above does.

05

The order API tests run in

Asked twice, so here it is as a rule.

1Pingis the API up at all?fails: nothing else matters 2JSON schemais the shape still the onewe agreed? fails: stop here 3Individual testsget, post, put, patch,delete, one at a time 4End to endthe fixture-basedlifecycle specs WHY THIS ORDER If ping fails, every later failure is noise. If the schema fails, fields are missing or retyped, so asserting on them is pointless. Only when both pass is a field-level assertion telling you something about the product rather than about the plumbing.
Ping, then schema, then everything else. The schema check is a pre-check, not a replacement for assertions.

Two questions from the floor. What is WireMock, and how is it different from Ajv? Not related. WireMock mocks an API, which you need when the backend is not built yet and the UI team cannot wait: database first, then backend, then API, then UI is the usual build order, and a mock lets the UI start early. Ajv validates a real response. Is retry supported? Yes, built into Playwright's config; nothing to add.

06

Building it with AI, and what had to be cut

The whole feature was scaffolded with a coding agent, and the method was the same one from the dotenv class: plan mode first, never auto-plan, read the plan, then auto-accept. The prompt named where each piece must live, the schema in test data, the validator in a shared util, the spec in a 05 folder, and asked the agent to show its changes before making them.

The agent's questions were answered as it planned (one schema, for create booking only; strict additionalProperties). Then the review found the usual excess:

  • It wrote its own rejection handling with reasons per missing field. Cut. A failed assertion with Ajv's message is enough.
  • It generated its own booking payload instead of using buildBookingFromGenerator from the previous class. Cut, and the instruction was explicit: only use the utilities that already exist.

That second cut is the recurring lesson of this batch. An agent that does not know your utilities will write new ones. Tell it what exists.

07

Ponytail

Ponytail is a code optimiser that trims what the session called AI slop: the class, the wrapper and the helper an agent writes when a function would do. It installs as a plugin (Copilot, Claude Code, Codex, Gemini and others), and the command is ponytail review with a file name.

Run against last class's 73-line lifecycle spec, the result is 33 lines with the same three tests. What went, and why each cut was safe, is recorded in the repository:

Cut Why it was safe
test.step around single calls The trace already lists every request with timings
testInfo.attach of the body The same trace already holds it
Ten log.info lines Each restated its own step name
buildBooking buildBookingFromGenerator gives ordered, today-relative dates
An explicit bookerToken Omitting it uses the managed token, which re-auths on a 403

Those cuts are only safe because the config sets trace: 'on'. Turn tracing off and the attachments stop being redundant.

The shipped version of the middle test:

TypeScript
test('update the booking, then read it back', async ({ bookingApi }) => {
    // No token argument: BookingApi re-auths on a 403 and retries. Passing one opts out.
    const updated = await bookingApi.updateBooking(
        bookingId,
        buildBookingFromGenerator({ firstname: 'E2E', lastname: 'Updated', totalprice: 950 }),
    );
    expect(updated).toMatchObject({ lastname: 'Updated', totalprice: 950 });

    // The GET, not the PUT echo, is what proves it persisted.
    expect((await bookingApi.getBooking(bookingId)).lastname).toBe('Updated');
});

A correction made live, and worth keeping. The first optimised file collapsed create, update and delete into a single test, and the room was told that is not wanted. Then the instructor checked: that collapse was the agent's own doing, not Ponytail's. Ponytail keeps your structure. The final file keeps three tests, because describe.serial reporting each stage as its own pass or fail is a reporting decision, not machinery to trim. Use the tool sparingly regardless, and read what it removed.

Two related tips from the same stretch: to stop a coding agent naming itself in your commits, add a global rule that it must never mention itself in commit messages. And the tiered-model skill shown routes trivial file creation to a small model and hard reasoning to a large one, so you are not spending the expensive model on boilerplate.

08

Workers, RAM and the question that became a document

Someone asked how to run in parallel and rerun failures, and the answer grew into a measured guide that now lives in the repository. The numbers below are from that guide, measured on the batch framework, not estimated.

What a worker is. One isolated lane with its own browser. workers: 1 runs everything nose to tail. Leave workers unset with fullyParallel: true and Playwright picks a default: half the logical cores, capped at the number of tests. On the 16-core machine used here, one schema test ran on one worker and four end-to-end tests ran on four; the "four" the session mentioned was that cap, not a fixed number.

RUN TIME, FOUR END-TO-END TESTS, HEADED 1 worker13.1 s 2 workers9.4 s 4 workers7.8 s 1 to 2 saved 3.7 s. 2 to 4 saved 1.6 s. The gain flattens once tests outnumber lanes only slightly. PEAK MEMORY, WHOLE RUN 1 worker1.70 GB 4 workers3.77 GB about 1 GB fixed, then0.7 GB per headed worker,0.4 GB per headless worker
Measured, not guessed. Same four tests headless: 2.47 GB and 3.3 seconds, which is why headless is the bigger win.

The session's rule of thumb was about 1 GB per worker, rising to 1.5 or 2 GB for complex tests. The measurement says 0.7 GB headed and 0.4 GB headless on top of roughly 1 GB fixed. Keep the rule of thumb as your safety margin; use the measurement to understand where the ceiling comes from.

System RAM Free for tests RAM ceiling, headed RAM ceiling, headless Cores / 2 Practical pick
8 GB about 4 GB 4 7 2 to 4 2 headed, 4 headless
16 GB about 11 GB 14 25 4 to 6 4 headed, 6 headless
32 GB about 26 GB 35 62 6 to 8 8, RAM is not the limit

Read it this way: RAM only binds on 8 GB machines. From 16 GB up you run out of cores first, so buying RAM to get more workers is the wrong upgrade. The formula behind the table:

Text
workers = min(
    (free RAM in GB - 1) / 0.7,     # 0.4 if headless
    logical cores / 2
)

Setting it, in the two forms worth knowing:

Terminal
npx playwright test --workers=4        # one run
npx playwright test --workers=50%      # a share of cores, portable across machines
npx playwright test --workers=1        # serial: the first thing to try when a test passes alone and fails in the suite

describe.serial blocks run on one lane no matter what --workers says. That is the design of the lifecycle spec, and it is why the three booking tests cannot be sped up by adding workers.

Production numbers, for scale. The TechCon suite is about 12,000 tests on ten Jenkins agents, each a 64 GB machine running eight workers, and a full regression takes three to four hours. Eight, not more, because complex tests cost more than the average. Someone in the room reported 1,300 tests taking 8 to 12 hours on six agents; the reply was that the tests themselves are slow and need splitting.

Sharding. --shard=k/n slices a suite across n machines and combines with workers rather than replacing them. This batch does not use it, because the usual Jenkins layout already splits by module: one job per module, each job fanned out to its own agents (the session used the older master and slave terms), each agent running its workers. The split happens before Playwright ever sees the tests.

09

Tasks and announcements

Next class: the last layer, the AI layer, added to the framework as version 2: root cause analysis on failures, self-healing, prompt to spec, a PR review bot, and smart tests. After that the framework supports API, page objects, Cucumber, headless, workers and AI in one place.

Next week: Jenkins parts 3 and 4 and GitHub Actions, led by Deepak, running this framework on both.

Being shared

  • The push with the Ajv level, the Ponytail spec beside the original, and the worker guide.
  • A separate video on keeping a coding agent's name out of your commits.
  • The Cucumber videos are already in the platform for anyone adding BDD to the framework.