Ask a model the same question twice and you get two different sentences, often both correct, so assertEquals fails on good answers. This chapter is the vocabulary and the method: ground truth, golden datasets, three kinds of evaluator, and a threshold that turns a score into a pass or a fail. DeepEval demonstrates it in chapters 15 and 16; the ideas outlive the tool.
Non-determinism, open-ended outputs, hallucination, safety and bias, cost and latency.
12
Key terms in the notes
From prompt to eval harness; nine have a written definition in the chapter README.
3
Evaluator families
Rule based, model based and LLM as judge.
1
File in the chapter folder
README.md. The runnable evals start in chapter 15.
01Why assertEquals breaks on LLM output
Classic automation compares an actual value with an expected one. That works when the system returns the same thing every time. A language model does not: it phrases the same correct answer differently on each run, and sometimes states a wrong one fluently. The chapter names five reasons the old assertion stops working.
What breaks
What it looks like in a test
What replaces the check
Non-determinism
The same prompt returns a different sentence on every run, so assertEquals flakes on correct answers.
A score compared with a threshold
Open-ended outputs
"Summarize this document" has no single correct string to compare against.
Several measured axes: relevancy, faithfulness, correctness
Hallucination
A fluent, confident answer states something that is not true.
Faithfulness and hallucination checks against a source of truth
Safety and bias
A toxic or biased reply, a jailbreak, a prompt injection.
Safety metrics and an attack library (chapter 16)
Cost and latency
A correct answer that is too slow or too expensive to ship.
Token cost, p95 latency and tool-call counts as test signals
chapter_14_LLM_Eval/README.md
# Why LLM Eval?
Traditional QA isn't enough for LLM-powered systems. Here's what breaks and why you need a new approach.
#### <u>Non-determinism</u>
Same prompt, different output. `<u>assertEquals</u>` breaks. You need probabilistic, score-based checks with thresholds.
#### <u>Open-ended outputs</u>
There's no single "correct" answer for "summarize this document." You measure quality along multiple axes instead.
#### <u>Hallucinations</u>
LLMs fabricate facts confidently. Fluent ≠ correct. You need faithfulness checks against a source of truth.
#### <u>Safety & bias</u>
Toxic, biased, or unsafe responses are failures. Jailbreaks and prompt injections are a new attack surface
#### Cost & latency
Quality is only one dimension. Token cost, p95 latency, and tool-call efficiency are first-class test signals.
Concepts first, tool second. The chapter opens with "Tool can be replaced!": DeepEval is how the course demonstrates evaluation, but the vocabulary on this page applies to any framework.
02The vocabulary
Ten terms carry the rest of the evaluation track. The middle column is the course's definition; the right column is the testing idea you already know, following the QA mapping on chapter 15's slide.
Term
What it means in the course
In testing terms
Prompt
The input: system instructions, user message, and often retrieved context
The request payload
Completion / response
The output being evaluated
The actual result
Ground truth
The human-written correct answer, used as the reference
The expected result
Golden dataset
Curated input to expected-output pairs
Your regression suite and its test data
Evaluator / judge
Whatever assigns the score: a rule, a model, or another LLM
The test oracle
LLM-as-judge
A strong LLM that scores another LLM on criteria such as relevancy or correctness
An oracle that is itself non-deterministic
Hallucination
A fluent, confident statement not grounded in the facts or the provided context
A wrong value presented as right
Faithfulness
Does the answer stick to the retrieved context without inventing things
Output consistent with the input data it was given
Relevancy
Does the answer address the question that was asked
The response matches the request
Context precision / recall
Did retrieval fetch the right chunks, and did it fetch all of them
Precision and recall of a search
flowchart TD
GD[(Golden dataset<br/>input to expected output)] --> RUN
P[Prompt<br/>system + user + retrieved context] --> RUN[Run the LLM]
RUN --> C[Completion]
C --> J{Evaluator}
GT[(Ground truth<br/>human written)] --> J
CTX[(Retrieved context)] --> J
J -->|rule based| M1[Exact / regex / schema]
J -->|model based| M2[Semantic similarity]
J -->|LLM as judge| M3[Relevancy, faithfulness,<br/>correctness]
M1 --> SC[Scores vs thresholds]
M2 --> SC
M3 --> SC
SC -->|pass| OK[Ship]
SC -->|below threshold| FAIL[Fail the build<br/>+ show the failing case]
C --> COST[Cost + p95 latency<br/>first-class signals]
From a prompt to a graded result: the flow from the course README. Every evaluator ends in the same place, a score compared with a threshold.
03Live demo: one question, six answers, three checks
The question and reference answer are the first golden case of the chapter 16 chatbot. Answers 1 to 3 are real: the same bot gave all three in one recorded run, which is non-determinism you can read. Answers 4 to 6 were written for this page, each to break a different check. Pick an answer, move the threshold, or type your own, and compare each verdict with the human label.
Score one answer three waysNo API key needed
Question
What is your refund window?
Reference answer (the golden case)
Refunds are processed within 7 business days of receiving the returned item. Returns must be initiated within 30 days of delivery.
Human labelLLM judge (recorded in chapter 16)How the scores were computed
Scoreboard: all six answers at the current threshold (amber cells disagree with the human label)
#
Answer
Human
Exact match
Token F1
Keyword coverage
Recorded reply (Answer Relevancy card)
Correct
Recorded reply (Faithfulness card)
Correct
Recorded reply (Hallucination card)
Correct
Numbers swapped
Wrong
Correct paraphrase
Correct
Right phrases, wrong meaning
Wrong
Disagreements with the human label
What to notice: exact match fails every correct answer. Token F1 gives the swapped-numbers answer a perfect 1.00, so no threshold makes it agree with a person. Keyword coverage catches the swap but passes answer 6, which uses the right phrases to say the wrong thing. And the recorded LLM judge failed answer 1 on relevancy even though its facts are right. Every evaluator has a blind spot; you choose which one you can live with.
04Three kinds of evaluator
The course groups evaluators into three families. Use the cheapest one that can see the failure you care about, and know what each one cannot see.
its own bias; scores move between runs and judge models
a paid model call per metric, per case
flowchart LR
Q[Question] --> REL{{Relevancy}}
A[Answer] --> REL
A --> FAI{{Faithfulness}}
RC[(Retrieved context)] --> FAI
A --> HAL{{"Hallucination / correctness"}}
GT[(Ground truth)] --> HAL
RC --> CPR{{"Context precision / recall"}}
GT --> CPR
REL --> R1[Did it answer what was asked?]
FAI --> R2[Did it invent anything beyond the context?]
HAL --> R3[Is it true?]
CPR --> R4[Right chunks, and all of them?]
What each check compares. The pair of inputs is the metric: the same answer can pass one check and fail another.
Faithful is not the same as true. A RAG bot that answers perfectly from an out-of-date chunk scores high on faithfulness, because it stuck to what it was given. Only a check against ground truth notices. Chapter 16 runs both for this reason.
05Swap the assertion, not the harness
Arrange and act do not change: you still load a golden case and call the system under test. Only the last line changes, from equality to a scored metric with a threshold. The course README shows the shape with DeepEval. rag_pipeline stands for your own system and is not defined in the repo, so read it as a pattern, not a script.
The harness stays; the assertion becomes a score, a threshold and a written reason.
README.md (chapter 14 section)
from deepeval import assert_test
from deepeval.metrics import AnswerRelevancyMetric, FaithfulnessMetric
from deepeval.test_case import LLMTestCase
# One case out of a golden dataset. retrieval_context is what RAG actually# fetched, so faithfulness can check the answer against it rather than vibes.case = LLMTestCase(
input="Does the cart total include the SAVE20 discount?",
actual_output=rag_pipeline("Does the cart total include the SAVE20 discount?"),
retrieval_context=["SAVE20 applies 20% off the cart subtotal before tax."],
)
# Scores, not equality. A threshold is a product decision, so write it down.assert_test(case, [
AnswerRelevancyMetric(threshold=0.7), # did it answer the question askedFaithfulnessMetric(threshold=0.8), # did it stay inside the context
])
A threshold is a product decision. Choosing 0.7 or 0.8 is the same conversation as "response time under 2 seconds": someone owns the number, and it is written down next to the test.
06Where each idea becomes code
The chapter folder holds one file, chapter_14_LLM_Eval/README.md, and nothing in it runs. This table maps each concept on this page to the code that runs it in the next two chapters.
A golden row in practice: the case the demo uses, from the chapter 16 dataset. The input, the human-written expected output, and the ground-truth context that the hallucination metric checks against.
datasets/chatbot_goldens.py
CHATBOT_GOLDENS: list[ChatbotGolden] = [
ChatbotGolden(
input="What is your refund window?",
expected_output=(
"Refunds are processed within 7 business days of receiving the returned item. ""Returns must be initiated within 30 days of delivery."
),
context=[
"Refunds are processed within 7 business days of receiving the returned item.",
"Items can be returned within 30 days of delivery in original condition.",
],
categories=["policy", "refund"],
),
The first runnable evaluation is chapter 15's single test. It needs a judge model configured first (a Groq or OpenAI key), which the next chapter walks through.
The judge is an LLM too. It is non-deterministic and biased. The course's rule: pin the judge model, keep a human-written golden dataset as the reference, and treat a score as a signal, not proof.
Relevancy is not truth. A confidently wrong answer can be perfectly on topic. The course README says it in its chapter 15 sample: a confidently wrong answer "can still score 1.0" on relevancy.
Fluent is not correct. Hallucination checks need a source of truth: ground truth for correctness, retrieved context for faithfulness.
Do not pay a judge to check a format. A JSON schema, a required field or an allowed value is a rule. Rules are free and deterministic; save the judge for meaning.
Cost and latency are results. In chapter 16's recorded run the judge spent 38,251 tokens grading answers that cost 17,510 to produce, 2.18 times as much.
Two listed terms are not defined in the notes. Traces and spans, and the eval harness, appear in the chapter's term list without a definition. The harness is what chapters 15 and 16 build: pytest plus a dashboard on one metric catalogue.
DDrills for the chapter
Concept drills need only this page and the course repo. Playwright drills target the live demo: every control has a data-testid (turn on Show locator badges to see them).
Equality against three real replies
In the demo, load answers 1, 2 and 3: the same bot answering the same question in one run. Would assertEquals(reference, reply) pass any of them?
Expected result
None. Exact match is 0.00 on all three, yet each states both facts: 7 business days to process, 30 days to return.
Why does a wrong answer score 1.00?
Load answer 4. Token F1 gives it a perfect score. Explain why before you open the detail panel.
Expected result
It uses exactly the reference's 21 words with 7 and 30 swapped. F1 counts shared words and ignores order, so precision and recall are both 21/21.
Find the best threshold for token F1
Slide the threshold from 0 to 1 and watch the disagreement count under the token F1 column. What is the lowest it gets, and can it reach zero?
Hint
Answers 5 and 6 score 0.49 and 0.62, and only one of them is correct.
Expected result
Two is the floor (at 0.00 to 0.45, and at 0.65). It never reaches zero: answer 4 passes at every threshold, and the wrong answer 6 outscores the correct answer 5, so no single threshold separates them. At 0.00 every check passes everything: a gate that never fails.
Break keyword coverage
Answer 6 passes keyword coverage. Sketch a rule that would fail it, then name a correct answer your rule would also fail.
Expected result
For example, fail any answer with "not" before a required fact. It catches answer 6, but it also fails the correct "Refunds take no more than 7 business days, and returns are not accepted after 30 days." Rules check form; meaning needs a model or a person.
Pick the evaluator family
The reply must be JSON with a priority field of High, Medium or Low. Rule, model, or LLM judge?
Expected result
A rule: a schema check. It is free and deterministic; an LLM judge adds cost and noise for no gain.
Faithful or true?
A RAG bot answers perfectly from a chunk that is out of date. Which check notices: faithfulness, or a check against ground truth?
Expected result
Only the check against ground truth (hallucination or correctness). Faithfulness compares the answer with the chunk it was given, so it can still score high.
Write a golden row
Write the input, expected output and context for "How long does standard shipping take?" in the shape of chatbot_goldens.py.
Expected result
The repo's row: expected output "Standard shipping is free on orders over $50 and takes 5-7 business days inside the US.", context "Standard shipping (free on orders over $50): 5-7 business days inside the US."
Assert the swapped-numbers trapPlaywright
Write a Playwright test: select answer 4 and assert that token F1 is 1.00 and PASS while keyword coverage is 0.00.
ev-score-f1 reads 1.00, ev-verdict-f1 reads PASS, ev-score-keywords reads 0.00, and ev-human starts with Wrong.
Drive the threshold and type an answerPlaywright
Fill ev-threshold with 0.4 and assert the label and the scoreboard. Then type the reference answer into the answer box and assert that exact match passes.
Hint
A range input accepts fill('0.4'). Scoreboard cells are ev-cell-{row}-{check}, for example ev-cell-5-f1; the answer box is labelled "Answer text".
Expected result
ev-threshold-value reads 0.40, ev-cell-5-f1 reads 0.49 PASS and ev-disagree-f1 drops to 2. After typing, ev-candidate switches to custom and ev-score-exact reads 1.00.
SSolutions: test the demo, then reproduce the checks
The Playwright spec drives the demo on this page. The Python file reproduces the three checks so you can see the numbers come out the same without a browser. Both pass as written.
tests/llm-evaluation-demo.spec.ts
import { test, expect } from'@playwright/test';
const URL = 'https://app.thetestingacademy.com/ai/blueprint/learn/llm-evaluation.html';
test('token F1 gives the swapped-numbers answer a perfect score', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('ev-candidate').selectOption('4');
awaitexpect(page.getByTestId('ev-score-f1')).toHaveText('1.00');
awaitexpect(page.getByTestId('ev-verdict-f1')).toHaveText('PASS');
awaitexpect(page.getByTestId('ev-score-keywords')).toHaveText('0.00');
awaitexpect(page.getByTestId('ev-human')).toContainText('Wrong.');
awaitexpect(page.getByTestId('ev-status')).toContainText('token F1');
});
test('the threshold decides how often token F1 disagrees with a human', async ({ page }) => {
await page.goto(URL);
awaitexpect(page.getByTestId('ev-cell-5-f1')).toHaveText('0.49 FAIL');
awaitexpect(page.getByTestId('ev-disagree-f1')).toHaveText('3');
await page.getByTestId('ev-threshold').fill('0.4');
awaitexpect(page.getByTestId('ev-threshold-value')).toHaveText('0.40');
awaitexpect(page.getByTestId('ev-cell-5-f1')).toHaveText('0.49 PASS');
awaitexpect(page.getByTestId('ev-cell-4-f1')).toHaveText('1.00 PASS');
awaitexpect(page.getByTestId('ev-disagree-f1')).toHaveText('2');
});
test('the reference copied exactly is the only way to pass exact match', async ({ page }) => {
await page.goto(URL);
awaitexpect(page.getByTestId('ev-verdict-exact')).toHaveText('FAIL');
await page.getByLabel('Answer text').fill(
'Refunds are processed within 7 business days of receiving the returned item. Returns must be initiated within 30 days of delivery.');
awaitexpect(page.getByTestId('ev-candidate')).toHaveValue('custom');
awaitexpect(page.getByTestId('ev-score-exact')).toHaveText('1.00');
awaitexpect(page.getByTestId('ev-verdict-exact')).toHaveText('PASS');
awaitexpect(page.getByTestId('ev-human')).toContainText('Not labelled.');
});
rule_checks.py
# rule_checks.py: the demo's three checks in plain Python, with pytest tests.# Written for this page (not in the course repo). Run: pip install pytest && pytest rule_checks.pyimport re
REFERENCE = ("Refunds are processed within 7 business days of receiving the returned item. ""Returns must be initiated within 30 days of delivery.")
FACTS = ["7 business days", "30 days"]
defwords(text):
return re.sub(r"[^a-z0-9]+", " ", text.lower()).split()
defexact_match(answer, reference=REFERENCE):
return1.0if answer.strip() == reference else0.0deftoken_f1(answer, reference=REFERENCE):
a, r = words(answer), words(reference)
bag = {}
for w in r:
bag[w] = bag.get(w, 0) + 1
common = 0for w in a:
if bag.get(w, 0) > 0:
common += 1
bag[w] -= 1ifnot common:
return0.0
precision, recall = common / len(a), common / len(r)
return2 * precision * recall / (precision + recall)
defkeyword_coverage(answer, facts=FACTS):
padded = " " + " ".join(words(answer)) + " "returnsum(f" {f} "in padded for f in facts) / len(facts)
SWAPPED = ("Refunds are processed within 30 business days of receiving the returned item. ""Returns must be initiated within 7 days of delivery.")
PARAPHRASE = ("Send it back within 30 days of delivery. Once it reaches us, ""your money is returned within 7 business days.")
deftest_exact_match_only_passes_a_copy():
assertexact_match(REFERENCE) == 1.0assertexact_match(PARAPHRASE) == 0.0deftest_token_f1_cannot_see_swapped_numbers():
asserttoken_f1(SWAPPED) == 1.0# same 21 words, so a perfect scoredeftest_token_f1_punishes_a_correct_paraphrase():
assertround(token_f1(PARAPHRASE), 2) == 0.49deftest_keyword_coverage_catches_the_swap():
assertkeyword_coverage(SWAPPED) == 0.0assertkeyword_coverage(PARAPHRASE) == 1.0
Written for this page, not in the course repo. Expected: four passing tests. The numbers match the demo because both use the same rules: lowercase, punctuation removed, word order ignored.
chapter_14_LLM_Eval/README.md
# Key Terms used in the LLM Evaluation
1. What is a prompt?
2. Completion and response.
3. Ground truth.
4. Golden dataset.
5. What is a judge or evaluator?
6. What is an LLM as a judge?
7. What is hallucination?
8. What is faithfulness and groundness?
9. What is the relevancy?
10. What is the context precision, and many more?
11. What are traces and span?
12. What is the eval and harness?
---
**<u>Prompt</u>**
The input to the LLM. Includes system instructions, user message, and often retrieved context<u>.</u>
**<u>Completion / Response</u>**
The LLM's output that we evaluate.
**<u>Ground Truth</u>**
The "correct" or expected answer -> human-written, used as a reference.
**<u>Golden Dataset</u>**
A curated set of _input → expected output_ pairs. Your regression suite. ( TestData).
**<u>Evaluator / Judge (LLM, OR Human)</u>**
The component that scores a response. Can be rule-based, a model, or another LLM.
chapter_14_LLM_Eval/README.md (continued)
**<u>LLM-as-Judge</u>**
- Using a strong LLM (often GPT-4-class) to score another LLM's outputs on criteria like relevancy or correctness.
**<u>Hallucination</u>**
| |
| ----- |
| <u>A fluent, confident statement that isn't grounded in facts or the provided context.</u> |
---
**<u>Relevancy</u>**
```
Does the answer actually address the user's question?
```
```
**Faithfulness - Does the answer stick to the retrieved context,
without inventing things?**
```
```
Context Precision / Recall
```