The Testing Academy · Class Notes Wednesday, 30 September (IST)
Live class · study guide

AI engineering, how LLMs use tokens and parameters, the core AI glossary, and anti-hallucination rules

The AI groundwork before JavaScript starts: why testers learn AI engineering rather than machine learning, how a large language model turns tokens into a prediction using its parameters, running a model locally with Ollama, the core AI glossary through to RAG and hallucination, and the anti-hallucination rules that make a model ask for the PRD before it writes a test plan. No code yet: the tasks are research and one prompt exercise.

By Pramod Dutta, The Testing Academy. Study notes from the second class of Playwright 4x, built from the session recording. This class had no code, and the batch has no class repository yet. The Eraser deck was not reachable while this page was written, so the on-screen notes and the exact wording of the anti-hallucination rules are not reproduced here; the rules and the product document are in the task post.

01

What this class covered

  • Tools for the course, and no AI-written code for the first month
  • AI, machine learning and deep learning, and why testers learn AI engineering
  • Language models as autocomplete, and the transformer's attention
  • Parameters: the numbers a model learns in training
  • Open and closed models, and running one locally with Ollama or LM Studio
  • Tokens, and how they differ from parameters
  • The core glossary: generative AI, prompt, context window, base model, training, fine-tuning, inference, dataset
  • Three ways to give a model your company's data, one of them RAG
  • Hallucination, and anti-hallucination rules in a live demo
  • The glossary terms still to come
  • Tasks, threads and daily entries
02

Tools, and no AI-written code yet

VS Code is the course editor, free and open source. If you already have Cursor, Windsurf, OpenCode, Claude Code, Codex, Command Code, Kilo, Amazon Q or similar, you are free to use them. GitHub Copilot is built into VS Code, and the class suggested its paid plan for a month or two if you can; Command Code is a cheaper option the previous batch used happily. Google's Antigravity editor looks almost identical to VS Code, and signing in with a Google account gives a free quota of some Gemini models.

Two conditions come with all of them:

  • Get permission first if you work on a company machine. Some companies do not allow these tools.
  • No AI for writing code for at least the first month, six weeks or so. JavaScript and TypeScript come first, and AI-written code you do not understand causes more problems than it solves. Until then, the only thing AI does with your code is help you push it to GitHub.

The class set one more rule for the whole 90 days: "Pramod is not learning, we are learning." For every concept taught, you research it, practise it, and push the result to your GitHub account.

03

AI, machine learning, deep learning and AI engineering

Testers keep asking whether they need machine learning or deep learning. The answer was no, and the class drew the map to show why:

  • AI is the umbrella: machines doing tasks that usually need human intelligence, such as reasoning, classifying, predicting, generating and automating. Chatbots, recommendation systems, fraud detection, computer vision and code generation are all AI.
  • Machine learning is a subset of AI in which a system learns patterns from data: classification, regression, clustering, recommendations, spam filtering, price prediction, churn prediction. More data makes it better.
  • Deep learning is a subset of machine learning that uses neural networks on large, complex data such as images, audio, video and language. Speech recognition, image recognition, self-driving perception, and building large language models all belong here.
  • AI engineering is what this course teaches: building applications on top of existing foundation models. Testers will not train a model or build an LLM. Thousands already exist, and the job is to use them well.
WHERE AI ENGINEERING SITS AI machines doing tasks that need human intelligence Machine learning learns patterns from data: classification, regression, spam filtering Deep learning neural networks on images, audio, video and language foundation models, such as LLMs testers build on top of these AI engineering: where testers sit builds applications on top of existing foundation models trains no models of its own
Machine learning and deep learning are where models get made. AI engineering starts once a model exists.

"Generative AI tester" is not a real job title. The class was blunt about it: the term that means something, on a resume or in an interview, is AI engineering.

Language models existed well before ChatGPT. What changed with ChatGPT was the delivery: a model offered as a service, with an interface the general public could use.

04

How a language model predicts

A language model is, at heart, autocomplete. Say "the king loves the" and "queen" comes next; say "the cat sat on the" and "mat" comes next. Your own brain does this all the time, which is what researchers were trying to copy.

The breakthrough behind today's models is the transformer, from the paper Attention Is All You Need. It introduced attention: as the model reads, it weighs how much each piece of the text matters to every other piece before predicting what comes next.

A large language model (LLM) is a language model trained on a massive amount of data. The class defined "large" as more than a billion parameters, with anything below that a small language model (SLM). They are called language models because the early ones only handled language tasks, such as reversing, converting or transcribing text.

Treat the billion-parameter line as a rule of thumb. There is no official cutoff between small and large, and models of a few billion parameters are still often called small.

05

Parameters: the numbers a model learns

Parameters are the numbers a model learns during training, also called its weights. They are what the model uses to judge what comes next. They are set once, during training, and in general, the more parameters, the better the predictions. Companies with closed models do not publish their parameter counts.

The class's analogy: a hundred students. To guess things about a class of a hundred students, you might note each one's gender, whether they are a manual tester, an automation tester or a fresher, where they live, and whether they have a job, and give each trait a number. From those numbers you could guess who is likely to fear coding, or who has practised. The traits are the parameters, and the more of them you track, the better your guesses. A model's numbers are set the same way, during training, except that it has billions of them.

In class, "parameters" and "attention numbers" were used for the same idea. Strictly, the parameters are fixed once training ends, while attention scores are worked out fresh for every prompt, using those parameters.

06

Open and closed models, and running one locally

Closed models Open models
Examples GPT, Claude, Gemini DeepSeek, Kimi, Qwen, Gemma
Weights published? no yes, downloadable from Hugging Face
Parameter counts published? no yes

Hugging Face hosts an enormous number of free open models to try.

Ollama is not a model. It is an app that runs models on your own machine, and LM Studio does the same job. The class ran Gemma 3 at 1 billion parameters in Ollama and it replied to "hi" instantly, with nothing leaving the laptop. A machine with 8 GB of RAM handles a model that size; the class advised against the 4 billion version on 8 GB.

Personal laptop only. Running local models on a company machine needs permission first, especially somewhere like a bank. Model subscriptions are not part of the course, but there are free routes: the free tier of ChatGPT, Antigravity's quota, and local models.

07

Tokens, and how they differ from parameters

A token is not a word. It is the smallest unit a model reads: a short run of characters. Common short words are often a single token, while longer or rarer words split into several. OpenAI's tokenizer page shows exactly how any sentence splits.

The class flagged the difference between a token and a parameter as an interview question:

TOKENS IN, ONE PREDICTED TOKEN OUT The cat will sit on the ___ The cat will sit on the six tokens, and each one becomes a list of numbers the model billions of parameters, all learned in training mat chair: less likely Token: a piece of your input the smallest unit the model reads Parameter: part of the model a number it learned in training, used to weigh the tokens
Tokens change with every prompt. Parameters stay fixed once training ends. The model uses the second to predict what follows the first.
08

The core glossary

Term Meaning
Generative AI AI that generates something: text, images, code or audio. Asking ChatGPT for test cases is generative AI.
LLM a prediction engine trained on a huge amount of data
Token the smallest unit a model reads
Prompt the instruction or input you give a model
Context window how much a model can take into account at once, measured in tokens; K means thousand, so 256K is 256,000 tokens
Base model also called a foundation model: trained on a large dataset (text, images, speech and more), so it can answer questions, analyse sentiment, extract information, describe images and follow instructions
Training teaching the neural network from examples
Fine-tuning extra training for a particular task or domain, which produces a new model
Inference a trained model generating an answer, the response
Dataset the collection of examples used for training and evaluation
Hallucination a model confidently giving a wrong answer, or one about things that do not exist

The analogies from class. A context window is like your own attention: after about an hour of class, new material stops going in. Training is how a child learns that "bhau bhau" means a dog, from school all the way through. Fine-tuning is what happens after the degree: a graduate joins a company and gets trained as a manual tester. And a neural network is a program that learns, not a brain.

Companies with closed models never share their datasets, which are proprietary. Golden datasets, a different idea, come later in the course.

09

Three ways to give a model your company's data

An LLM knows nothing about your company. Ask it something only your records hold, and it cannot answer from training. Your company's knowledge lives in Jira, Confluence, PDFs and source code, and there are three ways to get it to the model:

Way What you do
Prompt put the information and instructions into the prompt itself. Prompt engineering is the next class.
Fine-tuning train the model further on your data, which gives you a new model that is specific to your company
RAG keep your data in an external store; when a question arrives, retrieve the relevant pieces and hand them to the model along with it

RAG stands for retrieval-augmented generation: retrieve the right data, augment the prompt with it, then generate the answer. How it finds the right pieces, using embeddings and a vector database, comes in a later class.

10

Hallucination, and anti-hallucination rules

Hallucination is a model giving you a wrong answer confidently. The class showed it live with two prompts.

Without rules: "Give me the test plan for app.vwo.com", sent to a lightweight model. It returned a complete test plan at once: objectives, scope, conversions and more. No PRD, feature list or ticket had been given, so a good part of that detail was invented.

With anti-hallucination rules: the same request, with a set of rules pasted first that begin by telling the model it is a QA specialist working under strict rules. This time the model replied that it must use only the PRD, and asked for the PRD, API documentation, logs or screenshots before writing anything. Once the PRD was pasted, it named the product correctly as a digital experience platform, wrote a test plan based on the document, and listed what was still missing. Told to "ask me one by one", it went through the gaps in turn, asking first for the API documentation and then for test credentials.

WITHOUT RULES WITH ANTI-HALLUCINATION RULES Give me the test plan for app.vwo.com a full test plan, at once objectives, scope, conversions and more including details no document ever gave it confident, and partly invented [rules] + create a test plan "I must use only the PRD." asks for the PRD, API docs, logs or screenshots first PRD pasted a plan built from the PRD missing details listed, then asked for one by one grounded, and it says what it does not know
The rules change the model's first move: without them it writes, with them it asks. What it writes afterwards comes from the document you gave it.

Tell the AI, do not ask it. Most people ask a model for things. The class's rule is to tell it: give it the role, the documents and the constraints. It is the difference between walking into a restaurant and saying "bring food", and ordering the exact dishes you want. The more of the gaps you close, the better the answer.

The class also touched on how models learn: first from a book, then from a great many books, and now, increasingly, by trying things and learning from the results.

11

The glossary terms still to come

Learners called out more terms during class, and they will be covered over the coming weeks: agentic AI, bias, embeddings, vector databases, schema (which just means structure), prompt engineering (next class), MCP, LangChain (a framework rather than a concept), skills, agents, weights, chatbots, harness, prompt injection, guardrails, evals, chunks, ground truth, golden datasets and temperature. Researching them is one of today's tasks.

12

Tasks and announcements

Task 1, open versus closed models. Research the difference between open source and closed source models: which models fall on each side, and what each side publishes.

Task 2, the AI glossary. Research the glossary terms above, starting with the ones in the last section, and add any others you meet.

Task 3, a test plan with anti-hallucination rules. Use the anti-hallucination rules and the VWO product requirements document, both attached to the task post, to generate a test plan for app.vwo.com. Answer the model's questions about missing details. Post your answer only in the dated task thread.

Also try: Ollama or LM Studio, on a personal laptop, not an office one.

  • Threads: task answers go only in that day's task thread, and new threads for answers will be deleted. Questions go in the doubt thread.
  • Recordings are under My Courses.
  • No SDET Club access yet? Fill in the required access form, with your WhatsApp number in the form rather than in the chat, and the team will call you.
  • Add a learning entry every day from the dashboard: Add learning, write what you learned or researched, then save. Entries earn points, and the points are checked.
  • Extra AI sessions run in some evenings, around 8 PM, and invites go out beforehand.
  • Next class, Friday: prompt engineering and setting up the GitHub MCP. JavaScript from the start follows straight after. Friday 2 October is not a holiday for the batch.