Chapter 14
Chapter 14 . LLM evaluation . Concepts

LLM evaluation: scores, not string equality

Ask a model the same question twice and you get two different sentences, often both correct, so assertEquals fails on good answers. This chapter is the vocabulary and the method: ground truth, golden datasets, three kinds of evaluator, and a threshold that turns a score into a pass or a fail. DeepEval demonstrates it in chapters 15 and 16; the ideas outlive the tool.

5
Reasons classic QA breaks
Non-determinism, open-ended outputs, hallucination, safety and bias, cost and latency.
12
Key terms in the notes
From prompt to eval harness; nine have a written definition in the chapter README.
3
Evaluator families
Rule based, model based and LLM as judge.
1
File in the chapter folder
README.md. The runnable evals start in chapter 15.

01Why assertEquals breaks on LLM output

Classic automation compares an actual value with an expected one. That works when the system returns the same thing every time. A language model does not: it phrases the same correct answer differently on each run, and sometimes states a wrong one fluently. The chapter names five reasons the old assertion stops working.

What breaksWhat it looks like in a testWhat replaces the check
Non-determinismThe same prompt returns a different sentence on every run, so assertEquals flakes on correct answers.A score compared with a threshold
Open-ended outputs"Summarize this document" has no single correct string to compare against.Several measured axes: relevancy, faithfulness, correctness
HallucinationA fluent, confident answer states something that is not true.Faithfulness and hallucination checks against a source of truth
Safety and biasA toxic or biased reply, a jailbreak, a prompt injection.Safety metrics and an attack library (chapter 16)
Cost and latencyA correct answer that is too slow or too expensive to ship.Token cost, p95 latency and tool-call counts as test signals
chapter_14_LLM_Eval/README.md
# Why LLM Eval?
Traditional QA isn't enough for LLM-powered systems. Here's what breaks  and why you need a new approach.

#### <u>Non-determinism</u>
Same prompt, different output. `<u>assertEquals</u>` breaks. You need probabilistic, score-based checks with thresholds.

#### <u>Open-ended outputs</u>
There's no single "correct" answer for "summarize this document." You measure quality along multiple axes instead.

#### <u>Hallucinations</u>
LLMs fabricate facts confidently. Fluent ≠ correct. You need faithfulness checks against a source of truth.

#### <u>Safety & bias</u>
Toxic, biased, or unsafe responses are failures. Jailbreaks and prompt injections are a new attack surface

#### Cost & latency
Quality is only one dimension. Token cost, p95 latency, and tool-call efficiency are first-class test signals.
Concepts first, tool second. The chapter opens with "Tool can be replaced!": DeepEval is how the course demonstrates evaluation, but the vocabulary on this page applies to any framework.

02The vocabulary

Ten terms carry the rest of the evaluation track. The middle column is the course's definition; the right column is the testing idea you already know, following the QA mapping on chapter 15's slide.

TermWhat it means in the courseIn testing terms
PromptThe input: system instructions, user message, and often retrieved contextThe request payload
Completion / responseThe output being evaluatedThe actual result
Ground truthThe human-written correct answer, used as the referenceThe expected result
Golden datasetCurated input to expected-output pairsYour regression suite and its test data
Evaluator / judgeWhatever assigns the score: a rule, a model, or another LLMThe test oracle
LLM-as-judgeA strong LLM that scores another LLM on criteria such as relevancy or correctnessAn oracle that is itself non-deterministic
HallucinationA fluent, confident statement not grounded in the facts or the provided contextA wrong value presented as right
FaithfulnessDoes the answer stick to the retrieved context without inventing thingsOutput consistent with the input data it was given
RelevancyDoes the answer address the question that was askedThe response matches the request
Context precision / recallDid retrieval fetch the right chunks, and did it fetch all of themPrecision and recall of a search
flowchart TD
    GD[(Golden dataset<br/>input to expected output)] --> RUN
    P[Prompt<br/>system + user + retrieved context] --> RUN[Run the LLM]
    RUN --> C[Completion]
    C --> J{Evaluator}
    GT[(Ground truth<br/>human written)] --> J
    CTX[(Retrieved context)] --> J
    J -->|rule based| M1[Exact / regex / schema]
    J -->|model based| M2[Semantic similarity]
    J -->|LLM as judge| M3[Relevancy, faithfulness,<br/>correctness]
    M1 --> SC[Scores vs thresholds]
    M2 --> SC
    M3 --> SC
    SC -->|pass| OK[Ship]
    SC -->|below threshold| FAIL[Fail the build<br/>+ show the failing case]
    C --> COST[Cost + p95 latency<br/>first-class signals]
From a prompt to a graded result: the flow from the course README. Every evaluator ends in the same place, a score compared with a threshold.

03Live demo: one question, six answers, three checks

The question and reference answer are the first golden case of the chapter 16 chatbot. Answers 1 to 3 are real: the same bot gave all three in one recorded run, which is non-determinism you can read. Answers 4 to 6 were written for this page, each to break a different check. Pick an answer, move the threshold, or type your own, and compare each verdict with the human label.

Score one answer three waysNo API key needed
Question

What is your refund window?

Reference answer (the golden case)

Refunds are processed within 7 business days of receiving the returned item. Returns must be initiated within 30 days of delivery.

data-testid=ev-candidatedata-testid=ev-answerdata-testid=ev-thresholddata-testid=ev-reset
Three checks on this answer
CheckScoreVerdict
Exact match
Token F1
Keyword coverage
Human label
LLM judge (recorded in chapter 16)
How the scores were computed

    
Scoreboard: all six answers at the current threshold (amber cells disagree with the human label)
#AnswerHumanExact matchToken F1Keyword coverage
Recorded reply (Answer Relevancy card)Correct
Recorded reply (Faithfulness card)Correct
Recorded reply (Hallucination card)Correct
Numbers swappedWrong
Correct paraphraseCorrect
Right phrases, wrong meaningWrong
Disagreements with the human label
What to notice: exact match fails every correct answer. Token F1 gives the swapped-numbers answer a perfect 1.00, so no threshold makes it agree with a person. Keyword coverage catches the swap but passes answer 6, which uses the right phrases to say the wrong thing. And the recorded LLM judge failed answer 1 on relevancy even though its facts are right. Every evaluator has a blind spot; you choose which one you can live with.

04Three kinds of evaluator

The course groups evaluators into three families. Use the cheapest one that can see the failure you care about, and know what each one cannot see.

EvaluatorExamplesGood atBlind toCost
Rule basedexact match, regex, JSON schema, required keywordsformats, fields, values that must appearmeaning: negation, paraphrase, swapped factsfree and deterministic
Model basedsemantic similarity between embeddingsparaphrases that keep the meaningsmall factual changes inside similar wordingcheap; stable for a fixed embedding model
LLM as judgerelevancy, faithfulness, correctness, G-Eval rubricsmeaning, with a written reasonits own bias; scores move between runs and judge modelsa paid model call per metric, per case
flowchart LR
    Q[Question] --> REL{{Relevancy}}
    A[Answer] --> REL
    A --> FAI{{Faithfulness}}
    RC[(Retrieved context)] --> FAI
    A --> HAL{{"Hallucination / correctness"}}
    GT[(Ground truth)] --> HAL
    RC --> CPR{{"Context precision / recall"}}
    GT --> CPR
    REL --> R1[Did it answer what was asked?]
    FAI --> R2[Did it invent anything beyond the context?]
    HAL --> R3[Is it true?]
    CPR --> R4[Right chunks, and all of them?]
What each check compares. The pair of inputs is the metric: the same answer can pass one check and fail another.
Faithful is not the same as true. A RAG bot that answers perfectly from an out-of-date chunk scores high on faithfulness, because it stuck to what it was given. Only a check against ground truth notices. Chapter 16 runs both for this reason.

05Swap the assertion, not the harness

Arrange and act do not change: you still load a golden case and call the system under test. Only the last line changes, from equality to a scored metric with a threshold. The course README shows the shape with DeepEval. rag_pipeline stands for your own system and is not defined in the repo, so read it as a pattern, not a script.

Arrange and act stay the same; only the assertion changes Arrange golden case Act call the model assertEquals(exp, act) the old last line score >= threshold the new last line FAIL on a paraphrase PASS or FAIL plus a reason
The harness stays; the assertion becomes a score, a threshold and a written reason.
README.md (chapter 14 section)
from deepeval import assert_test
from deepeval.metrics import AnswerRelevancyMetric, FaithfulnessMetric
from deepeval.test_case import LLMTestCase

# One case out of a golden dataset. retrieval_context is what RAG actually
# fetched, so faithfulness can check the answer against it rather than vibes.
case = LLMTestCase(
    input="Does the cart total include the SAVE20 discount?",
    actual_output=rag_pipeline("Does the cart total include the SAVE20 discount?"),
    retrieval_context=["SAVE20 applies 20% off the cart subtotal before tax."],
)

# Scores, not equality. A threshold is a product decision, so write it down.
assert_test(case, [
    AnswerRelevancyMetric(threshold=0.7),   # did it answer the question asked
    FaithfulnessMetric(threshold=0.8),      # did it stay inside the context
])
A threshold is a product decision. Choosing 0.7 or 0.8 is the same conversation as "response time under 2 seconds": someone owns the number, and it is written down next to the test.

06Where each idea becomes code

The chapter folder holds one file, chapter_14_LLM_Eval/README.md, and nothing in it runs. This table maps each concept on this page to the code that runs it in the next two chapters.

Concept on this pageWhere it runs in the repoRun it
The concept noteschapter_14_LLM_Eval/README.mdRead it on GitHub or in your editor
A threshold instead of equalitychapter_15_DeepEval/test_01_Anwser_Relevancy.pypytest test_01_Anwser_Relevancy.py
Subject model and judge modelchapter_15_DeepEval/test_02_Groq_Qwen_Vs_GROQ_OpenAI.pypytest test_02_Groq_Qwen_Vs_GROQ_OpenAI.py
Golden datasetchapter_16_DeepEval_Framwork/03_DeepFramework/datasets/Imported by the suite and the dashboard
Relevancy, faithfulness, hallucination, correctness at scale03_DeepFramework/tests/chatbot/ and tests/rag/venv/bin/python -m pytest -m quality
Context precision03_DeepFramework/tests/rag/test_01_rag_contextual_precision.pyvenv/bin/python -m pytest tests/rag/test_01_rag_contextual_precision.py
Cost as a test signal03_DeepFramework/token_meter.pyShown on every dashboard card

A golden row in practice: the case the demo uses, from the chapter 16 dataset. The input, the human-written expected output, and the ground-truth context that the hallucination metric checks against.

datasets/chatbot_goldens.py
CHATBOT_GOLDENS: list[ChatbotGolden] = [
    ChatbotGolden(
        input="What is your refund window?",
        expected_output=(
            "Refunds are processed within 7 business days of receiving the returned item. "
            "Returns must be initiated within 30 days of delivery."
        ),
        context=[
            "Refunds are processed within 7 business days of receiving the returned item.",
            "Items can be returned within 30 days of delivery in original condition.",
        ],
        categories=["policy", "refund"],
    ),

The first runnable evaluation is chapter 15's single test. It needs a judge model configured first (a Groq or OpenAI key), which the next chapter walks through.

terminal
cd chapter_15_DeepEval
python3 -m venv venv && source venv/bin/activate
pip install -U deepeval requests
pytest test_01_Anwser_Relevancy.py

07What to watch

  • The judge is an LLM too. It is non-deterministic and biased. The course's rule: pin the judge model, keep a human-written golden dataset as the reference, and treat a score as a signal, not proof.
  • Relevancy is not truth. A confidently wrong answer can be perfectly on topic. The course README says it in its chapter 15 sample: a confidently wrong answer "can still score 1.0" on relevancy.
  • Fluent is not correct. Hallucination checks need a source of truth: ground truth for correctness, retrieved context for faithfulness.
  • Do not pay a judge to check a format. A JSON schema, a required field or an allowed value is a rule. Rules are free and deterministic; save the judge for meaning.
  • Cost and latency are results. In chapter 16's recorded run the judge spent 38,251 tokens grading answers that cost 17,510 to produce, 2.18 times as much.
  • Two listed terms are not defined in the notes. Traces and spans, and the eval harness, appear in the chapter's term list without a definition. The harness is what chapters 15 and 16 build: pytest plus a dashboard on one metric catalogue.