A model does not look answers up. It splits your prompt into tokens, weighs every token against every other one (attention), scores each possible next token and samples one. That is why the same prompt can give different test cases, and why changing one word in a prompt changes the answer.
The 2017 paper (arXiv:1706.03762), one idea per slide.
5
Preset sentences
In the calculator, plus any sentence you type. Same input, same numbers.
01Why a tester should care
This is a concept chapter: two HTML pages from the course repo that run in any browser, with no install and no key. Everything later in the course (a Jira agent, a RAG copilot, a CrewAI crew, a LangGraph pipeline) is a loop around the one operation you meet here: predict the next token.
The chapter README puts it plainly: a model is not a database lookup. It weighs every token against every other token and predicts the next one. Three questions from the README follow from that, and each changes how you test AI output or write prompts:
Question (chapter README)
Short answer
What you do about it
Why does the same prompt give different test cases each run?
The next token is sampled. A temperature above 0 spreads probability over several candidates, and floating-point effects add a little more.
Pin temperature=0 and state explicit constraints. Expect "deterministic-ish", not identical, output.
Why does adding "be thorough" rarely help?
A vague word adds weight without direction.
Replace it with measurable constraints, such as "cover boundary, negative, and security cases".
Do I need to read the original paper?
No. You need one idea from it: every token is weighed against every other token.
Cut irrelevant words from prompts. They take part in attention too and pollute the answer.
02From prompt to next token
The README sums the chapter up in one flow. Follow a single prompt token from left to right.
flowchart LR
P[Prompt tokens] --> E[Embeddings]
E --> A[Self-attention]
A --> W[Token-to-token weights]
W --> N[Next-token logits]
N --> S{Sampling}
S -->|temp = 0| D[Deterministic-ish output]
S -->|temp > 0| V[Variable output]
The chapter README's mental model: how one prompt token reaches the output.
Term
What it is
Where you see it
Token
A piece of text the model reads: a word, part of a word, or punctuation
Step 1 of the calculator; the chips in the demo
Token id
The integer a token maps to. Real vocabularies hold 30,000 to 200,000 tokens
The id= badges (the calculator derives them from a hash of the text)
Embedding
A vector of d_model numbers per token, learned during training
Step 2: matrix E
Attention
Weights that say how much each token takes from every other token; each row sums to 1
Steps 3 to 7, and the heatmap further down
Logit
A raw score for one possible next token
The Logit column in the demo
Softmax
Turns raw scores into probabilities that sum to 1
Step 6, and every probability in the demo
Temperature
Divides the logits before softmax. Low T sharpens, high T flattens, T = 0 always takes the top token
The demo slider
Sampling
Drawing the next token at random, weighted by those probabilities
Sample 20 runs in the demo
03Live demo: next-token candidates and temperature
Pick a prompt and look at the candidates for the next token. The first pair differs by one token (401 or 500); the second pair is the README's "be thorough" question, vague against measurable. Drag the temperature, then sample 20 runs with a fixed seed. The token chips and their ids come from the repo's own tokenize() and strHash(); highlighted chips are the ones the paired prompt does not have.
Next-token candidates: wording and temperatureNo API key needed
The logits are illustrative. They are hard-coded for this page to show the mechanism: a real model scores every token in its vocabulary, and its numbers depend on the model. What is real is the math: softmax, temperature and a seeded sampler, the same functions the repo page uses. Same prompt, temperature and seed always give the same 20 runs.
04Inside the attention calculator
attention_interactive.html runs scaled dot-product attention, softmax(QKT / √dk) · V, on real but small matrices in your browser. Type a sentence and press Enter or Compute Attention; eight cards show every intermediate matrix. The defaults are the cat will sit on the mat, dmodel 6, 2 heads and seed 42.
Step
What the page computes
Shape for the default sentence
1 Tokenize
Trim, lowercase, split on whitespace; an id badge per token (strHash(token) % 10000)
7 tokens
2 Embed
One vector per token from a PRNG seeded by the token text, so the same word always gets the same vector
E: 7 × 6
3 Project
Q = E·WQ, K = E·WK, V = E·WV, weights seeded by seed + h × 1000
7 × 3 each (dk = 6 / 2)
4 Raw scores
Q·KT: how much token i wants from token j
7 × 7
5 Scale
Divide by √dk (√3 ≈ 1.732) so softmax does not saturate
7 × 7
6 Softmax
Row by row, so each row sums to 1; drawn as a heatmap
7 × 7
7 Weighted sum
output = A·V, with a "Show calculation for" picker that writes one row out in full
7 × 3
8 Multi-head
The same steps per head, then the head outputs side by side
7 × 6 (2 heads × 3)
flowchart LR
S["sentence"] --> T["tokens"] --> E["E: n x d_model"]
E --> H1["head 1: own WQ, WK, WV, seed"]
E --> H2["head 2: own WQ, WK, WV, seed + 1000"]
H1 --> A1["softmax(Q.K / sqrt d_k) . V"]
H2 --> A2["softmax(Q.K / sqrt d_k) . V"]
A1 --> C["concat: n x heads*d_k"]
A2 --> C
C -.->|"in the heading, never applied"| W["x W_O"]
How runAttention() runs two heads: each has its own seeded weights, and the outputs are concatenated. The output projection WO from the paper is not in the code.
attention_interactive.html (lines 620-634)
const heads = [];
for (let h = 0; h < nHeads; h++) {
const rng = mulberry32(seed + h * 1000);
const WQ = makeMatrix(dModel, dK, rng);
const WK = makeMatrix(dModel, dK, rng);
const WV = makeMatrix(dModel, dK, rng);
const Q = matmul(E, WQ);
const K = matmul(E, WK);
const V = matmul(E, WV);
const scoresRaw = matmul(Q, transpose(K));
const scoresScaled = scale(scoresRaw, 1 / Math.sqrt(dK));
const weights = softmaxRows(scoresScaled);
const out = matmul(weights, V);
heads.push({ WQ, WK, WV, Q, K, V, scoresRaw, scoresScaled, weights, out });
}
What the page does, exactly. The README is slightly generous, so check it against the code:
Hover a heatmap cell, not a token, for a 4-decimal tooltip such as sit → cat = 0.1558.
It recomputes on Enter, the button, a preset chip, or a change to dmodel, heads or seed. It does not update on every keystroke.
The weight matrices are random and untrained, so attention is close to uniform: 0.082 to 0.237 per cell in head 1 for the default sentence.
There is no positional encoding, so the two the rows are identical in every matrix.
Step 8's heading shows · WO, but the code only concatenates the heads; no output projection is applied.
05Read a heatmap, then check the slide math
Slide 9 of attention_is_all_you_need.html hard-codes weights of the kind a trained model learns for "The cat will sit on the mat". Read it row by row: the row is the token doing the looking, the columns are what it looks at, and every row sums to 1.00.
Slide 9's simulated weights. Row = the token doing the looking; the outlined row is "sit", which puts 0.45 on "cat" and 0.25 on "mat". The calculator's random weights cannot learn this; these are written by hand to show what training produces.
Slides 6 and 7 show the raw scores for "sit" and the weights "after softmax". The raw scores are fine. The after-softmax numbers are rounded for the story, not computed: the real softmax gives "cat" 0.636 and "mat" 0.212, not 0.52 and 0.34. Dividing the same scores by a temperature first shows what the demo slider does.
Token
Raw score
Slide 7 shows
Real softmax (T = 1)
T = 0.5
T = 2
The
0.8
0.02
0.021
0.001
0.073
cat
4.2
0.52
0.636
0.885
0.397
will
1.5
0.04
0.043
0.004
0.103
sit
2.0
0.07
0.071
0.011
0.132
on
0.6
0.01
0.017
0.001
0.066
mat
3.1
0.34
0.212
0.098
0.229
attention_interactive.html (lines 514-521)
functionsoftmaxRows(m) {
return m.map(row => {
const max = Math.max(...row);
const exp = row.map(v => Math.exp(v - max));
const sum = exp.reduce((a, b) => a + b, 0);
return exp.map(v => v / sum);
});
}
Attention itself has no temperature: the table reuses the slide's scores only to show the math. Temperature is applied to the next-token logits at the end of the model, which is what the demo above simulates.
06Run the two pages
Both files are self-contained: inline CSS and JavaScript, with Google Fonts as the only network request. Open them straight from your clone.
File
What it is
How to use it
attention_interactive.html
The 8-step calculator. Five preset chips: the cat will sit on the mat, she gave him a red apple, the bank by the river, I love testing AI models, attention is all you need
Type a sentence of at least 2 words; pick dmodel 4, 6 or 8, heads 1, 2 or 3, and a seed from 1 to 999
attention_is_all_you_need.html
A 14-slide explainer: the problem with RNNs, tokenize, embed, Q/K/V, score, softmax, weighted sum, heatmap, multi-head, stacking, why it mattered, recap
Right Arrow or Space for the next slide, Left Arrow for the previous one, F for full screen
Notes.md
A one-line placeholder. The chapter's notes live in the README section
Read the README section "Chapter 01" instead
terminal
# clone the course repo, then open the pages in a browser: no build, no installopen chapter_01_LLM_Basics/attention_interactive.html
open chapter_01_LLM_Basics/attention_is_all_you_need.html
open is the macOS command from the README. On Linux use xdg-open, on Windows start, or drag the file into a browser window.
07What to watch
Temperature 0 is "deterministic-ish". It removes sampling, not every source of variation. Treat AI output as something to assert on structure and facts, not on exact wording.
The calculator's attention is random. Its weights are seeded, not trained, so do not read meaning into which word attends to which. Use it to learn the shapes and the order of operations.
Same word, same row. With no positional encoding, repeated words get identical rows. Real models add position information, so "the" at position 1 and position 6 differ there.
dk is rounded down.dK = Math.max(1, Math.floor(dModel / nHeads)): with dmodel 8 and 3 heads, each head gets 2 columns and the concatenated output has 6, not 8.
Check the numbers on slides. Slide 7's weights are illustrative. When a figure claims to be computed, recompute it: that habit is most of LLM testing.
Model names on slide 12 will date. The architecture point (one design behind every modern LLM) does not.
DDrills for the chapter
Drills 1 to 6 use the two pages in chapter_01_LLM_Basics. Drills 7 to 9 automate the demo on this page: every control has a data-testid (turn on Show locator badges to see them).
Compute a real softmax
Slide 7 lists raw scores for "sit": The 0.8, cat 4.2, will 1.5, sit 2.0, on 0.6, mat 3.1. Compute the softmax yourself (a spreadsheet is fine).
Hint
Subtract the largest score first, exponentiate, then divide by the sum. That is what softmaxRows() does.
Expected result
About 0.021, 0.636, 0.043, 0.071, 0.017, 0.212. The slide shows cat 0.52 and mat 0.34: they sum to 1, but they are rounded for the story, not computed.
Explain two identical rows
In the calculator with the default sentence, compare the two the rows in E, Q, K, V and the heatmap. Why are they identical?
Expected result
The embedding depends only on the token text (mulberry32(strHash(tok))) and the page adds no positional encoding, so identical tokens produce identical rows everywhere.
Change the head count
Set dmodel to 8 and heads to 3. What is dk, and how many columns does the final concatenated output have?
Expected result
dk = floor(8 / 3) = 2, so the output has 3 × 2 = 6 columns (h0d0 to h2d1), not 8, and no WO projection is applied.
Change the seed
Change the seed from 42 to 43. Which matrices change and which stay the same?
Expected result
Tokens, ids and E stay the same: they depend only on the text. WQ, WK, WV, the scores, the weights and the outputs all change.
Find what "sit" looks at
With the defaults, pick sit (row 3) in step 7. Which three tokens does it pay most attention to? Compare with the "sit" row on slide 9.
Expected result
sit (23.7%), mat (18.7%), will (17.0%). Slide 9's hand-written weights put 0.45 on cat: random, untrained weights cannot learn who does the sitting.
Rewrite a vague prompt
Rewrite "Write test cases for the login page. Be thorough." using the README's advice.
Expected result
Replace the vague word with measurable constraints and an output format, for example: "Cover boundary, negative, and security cases. Output CSV only with columns Scenario, TID, Priority. No preamble." Chapter 2 turns this into the RICE-POT framework.
Pin the temperaturePlaywright
Write a Playwright test: choose the 401 prompt, set the temperature to 0, sample, and assert that every run picked the same token.
Hint
Locators: llm-prompt (value auth401), llm-temp (a range input; fill('0') works), llm-run, llm-runs, llm-status.
Expected result
One row in llm-runs: token, "20 of 20 runs", and the status says 1 distinct token.
One token, a new top candidatePlaywright
Assert the top candidate for the 401 prompt, switch to the 500 prompt, and assert again. Then assert that the changed token chip is highlighted.
Hint
Token chips are llm-tok-1 to llm-tok-7 inside llm-tokens; a highlighted chip has class is-on.
Expected result
llm-top goes from token (46.6%) to server (43.2%). llm-tok-5 reads 500 and has class is-on.
See the noise in 20 runsPlaywright
Keep the 401 prompt at T = 1.0 and seed 42, and sample. Which token wins the most runs? Is it the most likely one?
Expected result
password wins 8 of 20 runs although token is the most likely (46.6%, 7 runs). Twenty draws are noisy: the reason one passing run of an AI test proves little. Change the seed and the split changes.
SSolutions: the demo test and the repo code
The Playwright spec for the demo, then the parts of the two repo pages worth reading line by line. The repo excerpts are verbatim.
tests/llm-basics-demo.spec.ts
import { test, expect } from'@playwright/test';
const URL = 'https://app.thetestingacademy.com/ai/blueprint/learn/llm-basics.html';
test('temperature 0 returns the top token on every run', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('llm-prompt').selectOption('auth401');
await page.getByTestId('llm-temp').fill('0');
awaitexpect(page.getByTestId('llm-temp-value')).toHaveText('T = 0.0');
await page.getByTestId('llm-run').click();
const runs = page.getByTestId('llm-runs').getByRole('listitem');
awaitexpect(runs).toHaveCount(1);
awaitexpect(runs.first()).toContainText('20 of 20 runs');
awaitexpect(page.getByTestId('llm-status')).toContainText('1 distinct token');
});
test('changing one token in the prompt moves the top candidate', async ({ page }) => {
await page.goto(URL);
awaitexpect(page.getByTestId('llm-top')).toHaveText('token (46.6%)');
await page.getByTestId('llm-prompt').selectOption('server500');
awaitexpect(page.getByTestId('llm-top')).toHaveText('server (43.2%)');
const changed = page.getByTestId('llm-tok-5');
awaitexpect(changed).toContainText('500');
awaitexpect(changed).toHaveClass(/is-on/);
});
test('a measurable prompt keeps its top token as temperature rises', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('llm-prompt').selectOption('specific');
awaitexpect(page.getByTestId('llm-prob-1')).toHaveText('78.1%');
await page.getByTestId('llm-temp').fill('2');
awaitexpect(page.getByTestId('llm-prob-1')).toHaveText('48.9%');
await page.getByTestId('llm-prompt').selectOption('vague');
awaitexpect(page.getByTestId('llm-prob-1')).toHaveText('20.4%');
awaitexpect(page.getByTestId('llm-spread')).toHaveText('6 candidates cover 90% of the probability');
});
// The demo's math, in the same style as the repo page (this page's code, not the repo's)
function probabilities(logits, T) {
if (T === 0) { // greedy: all weight on the top logit
const top = Math.max(...logits);
return logits.map((l) => (l === top ? 1 : 0));
}
const scaled = logits.map((l) => l / T); // temperature divides the logits
const max = Math.max(...scaled);
const exp = scaled.map((v) => Math.exp(v - max)); // same trick as softmaxRows()
const sum = exp.reduce((a, b) => a + b, 0);
return exp.map((v) => v / sum);
}
// one seeded stream for all 20 runs: mulberry32(seed) from attention_interactive.html
function sample(probs, rng) {
const u = rng();
let acc = 0;
for (let i = 0; i < probs.length; i++) { acc += probs[i]; if (u < acc) return i; }
return probs.length - 1;
}