What data-driven testing actually means
The definition the class opened with is short: an external data source plus the application under test, where the same test runs many times against different rows.
The word doing the work is external. Once the data lives outside the spec, the same login test covers five accounts by editing a file rather than by copying a test block five times.
JSON, because everything else is a detour
JSON stands for JavaScript Object Notation, a lightweight format for storing and moving data. It is the format the class kept coming back to, and by the end the recommendation was blunt: use JSON unless the data source forces something else.
The syntax rules are worth memorising because a single trailing comma breaks the file:
- Keys must use double quotes.
- String values must use double quotes.
- Key and value are separated by
:. - Multiple properties are separated by
,. - Objects use
{ }, arrays use[ ]. - The last property must not have a comma.
Six data types are supported: string, number, boolean, null, object, array.
{
"name": "Pramod",
"experience": 15,
"isTrainer": true,
"middleName": null,
"skills": ["JavaScript", "Playwright", "Selenium"],
"address": {
"city": "Bangalore",
"country": "India"
}
}
A JSON file and a JavaScript object literal look almost identical, and the class caught the difference live. { name: "Pramod" } is JavaScript, valid because unquoted keys are legal in source code. { "name": "Pramod" } is JSON. The moment the data leaves your .ts file and lands in a .json file, the quotes stop being optional.
Two conversions carry all the traffic between the two worlds:
const jsonText = '{ "name": "Pramod", "role": "QA Engineer" }';
const user = JSON.parse(jsonText); // JSON text -> JS object
console.log(user.name); // 'Pramod'
const backToText = JSON.stringify(user); // JS object -> JSON text
console.log(backToText); // '{"name":"Pramod","role":"QA Engineer"}'
Three ways to load a JSON file, and the one to use
The class compared them side by side:
| Way | Looks like | Verdict |
|---|---|---|
import |
import loginData from './testdata/login.json' |
Preferred. Modern, shortest, typed |
require |
const user = require('./user.json') |
Works, older CommonJS style |
fs |
fs.readFileSync(...) then JSON.parse(...) |
Needed only when you must also write |
For plain test data, import wins. The file becomes an object at the top of your spec and there is nothing to parse:
import { test } from '@playwright/test';
import loginData from './testdata/login.json';
test('valid user can log in', async ({ page }) => {
await page.locator('#email').fill(loginData.validUser.username);
await page.locator('#password').fill(loginData.validUser.password);
await page.getByTestId('login-button').click();
});
Swapping validUser for invalidUser turns the same test into the negative case, which is the entire point of keeping the data outside.
Loading JSON with import needs "resolveJsonModule": true in tsconfig.json. Playwright's own starter config sets it, so this usually just works, and when it does not, that flag is the first thing to check.
Reading and writing with the fs module
fs is the Node file system module. It is not a Playwright package, and that distinction mattered enough in class to be worth repeating: Playwright ships test and expect, everything else here is plain Node or a third-party library.
Reading:
const fs = require('fs');
const path = require('path');
const filePath = path.resolve(__dirname, 'testdata', 'users.json');
const raw = fs.readFileSync(filePath, 'utf8');
const users = JSON.parse(raw);
console.log(users.username);
Writing is the mirror image, with JSON.stringify doing the conversion:
const user = { name: 'Pramod', role: 'QA Engineer' };
fs.writeFileSync(
path.resolve(__dirname, 'output.json'),
JSON.stringify(user, null, 2),
'utf8'
);
The two extra arguments to JSON.stringify are the ones people skip and then wonder why the output is one unreadable line. null is the replacer, meaning replace nothing. 2 is the indent width. Drop them and the file still writes, it is just unreadable.
Writing to a name that already exists overwrites it. Appending is a different function, and fs has a long list of them.
The path bug the class hit live
The first attempt used a relative path and failed with ENOENT: no such file or directory. Guesses from the room included quoting and permissions. Neither was it.
The cause is that a relative path resolves against the process working directory, which is wherever you ran the command from, not against the file that contains the code. Run the same script from a different folder and it breaks.
The fix is __dirname, the directory of the current file, injected by Node:
// brittle: depends on where you ran the command
fs.readFileSync('testdata/users.json', 'utf8');
// stable: always relative to this file
fs.readFileSync(path.resolve(__dirname, 'testdata', 'users.json'), 'utf8');
__dirname exists in CommonJS. If your project sets "type": "module" in package.json, or you are writing .mjs, it is not defined and you need import.meta.dirname instead. Playwright's default TypeScript setup compiles to CommonJS, which is why __dirname worked in class without ceremony.
Generating one test per row
Here is the shape that turns a data array into a test run. The array of objects is the simplest possible source, and while it lives in the spec file it is still not ideal practice, it demonstrates the mechanism without a file read in the way.
import { test } from '@playwright/test';
const loginData = [
{ description: 'valid user', username: 'admin@tta.dev', password: 'admin123' },
{ description: 'wrong password', username: 'admin@tta.dev', password: 'nope' },
{ description: 'unknown user', username: 'ghost@tta.dev', password: 'admin123' },
{ description: 'empty password', username: 'admin@tta.dev', password: '' },
{ description: 'empty username', username: '', password: 'admin123' },
];
test.describe('DDT sample', () => {
test.beforeEach(async ({ page }) => {
await page.goto('https://app.thetestingacademy.com/practice/login');
});
for (const data of loginData) {
test(`Login attempt: ${data.description}`, async ({ page }) => {
await page.locator('#email').fill(data.username);
await page.locator('#password').fill(data.password);
await page.getByTestId('login-button').click();
});
}
});
Three details carry the whole pattern:
test.beforeEachstays outside the loop. The URL never changes, so the navigation is shared setup, not per-row data.- The title is a template literal. Playwright rejects duplicate test titles inside a describe block, so
`Login attempt: ${data.description}`is not decoration, it is what makes five tests legal instead of one repeated five times. - The loop runs at collection time. It executes when the file is loaded, before any test starts, which is exactly why five separate tests appear in the report instead of one test that loops internally. Get this wrong and a failure on row 3 hides rows 4 and 5.
Run it and the reporter shows five tests, each with its own name, its own retry and its own trace.
These five passed in class only because no assertion had been added yet. A data-driven test with no expect is a data-driven way of proving nothing. The row objects usually carry the expectation too, a shouldPass or expectedMessage field, so the assertion varies with the data instead of being hardcoded.
Writing a CSV reader from scratch
CSV needed a reader because Node does not parse it for you. The class wrote one by hand, and was explicit about why: not because you should ship it, but because the interview question is common and the answer should be yours.
The reader itself:
import * as fs from 'fs';
import * as path from 'path';
export interface TestDataRow {
[key: string]: string;
}
export function readCSVFile(filePath: string): TestDataRow[] {
const fullPath = path.resolve(filePath);
const content = fs.readFileSync(fullPath, 'utf8');
const lines = content.trim().split('\n');
const headers = lines[0].split(',');
const data: TestDataRow[] = [];
for (let i = 1; i < lines.length; i++) {
const values = lines[i].split(',');
const row: TestDataRow = {};
for (let j = 0; j < headers.length; j++) {
row[headers[j].trim()] = values[j]?.trim() ?? '';
}
data.push(row);
}
return data;
}
Reading it top to bottom:
import * as fspulls in every export under one name. It is the same import asconst fs = require('fs'), written in module syntax.interface TestDataRow { [key: string]: string }is an index signature: any string key, string value. CSV has no types, everything arrives as text, and the interface says so honestly.content.trim()before splitting kills the trailing newline almost every editor adds. Without it the last "row" is an empty string and you get a junk record.lines[0]is the header line, which is why the data loop starts ati = 1.values[j]?.trim() ?? ''covers the ragged row. If a line has fewer fields than the header,values[j]isundefined,?.stops the crash, and??substitutes an empty string.
The nested loops are the same two-loop shape as the star-pattern programs: outer walks rows, inner walks columns.
Using it is then unremarkable, which is the goal:
import { test } from '@playwright/test';
import * as path from 'path';
import { readCSVFile } from '../utils/csv-reader';
const loginData = readCSVFile(path.resolve(__dirname, '../testdata/login-data.csv'));
for (const data of loginData) {
test(`Login: ${data.description}`, async ({ page }) => {
// data.username, data.password, data.expectedResult
});
}
This reader is deliberately naive, and knowing where it breaks is the better half of the interview answer. It does not handle quoted fields containing commas ("Doe, John"), escaped quotes, or newlines inside a field. Real CSV is a nastier format than it looks. That is the honest case for reaching for a library.
Where you should stop writing code
Everything above stops at CSV for a reason. For every other format the class advice was to install the library and move on:
| Source | Library |
|---|---|
Excel .xlsx |
ExcelJS |
| YAML | JS-YAML |
| MySQL | the MySQL driver |
| JSON | nothing, import is built in |
Nobody writes these from scratch any more, and interviewers know it. The answer that lands is: I wrote the CSV reader myself so I understand the mechanics, and I used ExcelJS, JS-YAML and the MySQL driver in the framework because reimplementing them is not engineering.
YAML is worth recognising on sight, since it carries the same key-value data as JSON with the punctuation removed:
validUser:
username: admin@tta.dev
password: admin123
invalidUser:
username: ghost@tta.dev
password: nope
Picking a format
The decision is nearly always made for you by where the data already lives:
- JSON by default. Simplest to read, no library, and
importmakes it one line. - Excel when the business team owns the test data and will not leave a spreadsheet.
- CSV when an export drops it in your lap.
- YAML when the project already uses it for config.
- SQL when the data has to be pulled live from a database.
A database query is not a special case: fetch the rows, hold them in memory, and feed them into the same loop as everything else.
A locator note that came up
A question about self-healing locators surfaced the .or() chain, which was flagged as a preview rather than covered:
page.getByRole('button', { name: 'Login' })
.or(page.getByTestId('login-button'))
.or(page.getByText('Login'));
This is multiple locator strategies, not self-healing in the strict sense: nothing learns or rewrites itself. Playwright waits until any one of the alternatives matches, and the first match wins. It got its own treatment two classes later, alongside the page object model.
Tasks and announcements
- Milestone reached: 300 exercises. The full set is in the class repository.
- Today's task, mandatory: complete all 300 exercises and push them to your GitHub repository by the end of the week.
- For interviews, walk the CSV reader once. Resolve the path, read the file, trim, split on newlines, take line zero as headers, loop the rest, zip header to value, push the row. If you can narrate that, the question is answered.
- Coming next: page objects and post-login navigation.
- Later in the framework classes: environment variables for credentials, calendars, window and tab handling, and the multi-locator strategy in full.
Interview forms of today's material: why does a data-driven test title need a template literal? What breaks if you call JSON.stringify(data) with no second and third argument? Why does a relative path work when you run from the project root and fail from anywhere else? Name one input your hand-rolled CSV reader parses wrongly.