Chapter 15
Chapter 15 . LLM evaluation . DeepEval

DeepEval: a threshold where assertEquals was

DeepEval is a pytest plugin. You put the question, the model's answer and the known facts into an LLMTestCase, a second model called the judge scores it, and assert_test fails the test when a score misses its threshold. The chapter has two tests: a hard-coded 2+2 case, and a live Groq model graded by a judge from a different model family.

2
Test files
test_01 scores a hard-coded answer; test_02 asks a live Groq model first.
2
Metrics
AnswerRelevancyMetric and HallucinationMetric.
3
Judge options
OpenAI gpt-4o-mini, Groq openai/gpt-oss-120b, or a local Ollama model.
2
Key names to set
OPENAI_API_KEY for the subject call, LOCAL_MODEL_API_KEY for the judge.

01From assertEquals to assert_test

Arrange and act stay ordinary pytest. Only the assertion changes: assert_test(case, [metric]) sits where assert actual == expected used to, so CI needs no new runner. What is new is the test case itself. The chapter's slide maps each DeepEval parameter to the QA idea it replaces.

DeepEval parameterQA / SDET equivalent (chapter slide)In this chapter
inputTest input / request payloadBoth tests: "What is 2+2?"
actual_outputActual result from the system under testHard-coded in test_01, a live model answer in test_02
expected_outputExpected result / golden answer"4" in both tests
contextGround truth data (known facts)Read by HallucinationMetric in test_02
retrieval_contextWhat the retriever actually returnedNot used here; the chapter 16 RAG tests use it
thresholdAcceptance criteria (for example, response time under 2s)0.9 and 0.8 for relevancy, 0.3 for hallucination
model (judge LLM)Test oracle / validation serviceSet once per machine; test_02 also names it per metric
criteria, evaluation_params (G-Eval)The assertion rubric, and which fields to assert onChapter 16
additional_metadataTest tags / labels for filteringNot used
The judge is the oracle, and it is another LLM. DeepEval does not compare strings. A second model reads the test case and returns a score from 0 to 1 with a written reason, and the threshold turns that score into a pass or a fail.

02The two tests

The chapter folder holds two test files, setup notes and two slides.

FileWhat it doesRun
test_01_Anwser_Relevancy.pyOne hard-coded case ("What is 2+2?" answered "4") scored by AnswerRelevancyMetric(threshold=0.9). Only the judge model is called.pytest test_01_Anwser_Relevancy.py
test_02_Groq_Qwen_Vs_GROQ_OpenAI.pyAsks qwen/qwen3.8-27b on Groq at temperature 0, then a judge from another family, openai/gpt-oss-120b, scores relevancy (0.8) and hallucination (0.3).pytest test_02_Groq_Qwen_Vs_GROQ_OpenAI.py
Notes.mdInstall checklist, verify command, judge set-up for Groq and OpenAI, run command, two watch-outs.Read first
.env.sampleThe two key names: OPENAI_API_KEY (subject call in test_02) and LOCAL_MODEL_API_KEY (the judge).cp .env.sample .env
image-1.png, image.pngTwo slides: DeepEval parameters as QA concepts, and "normal vs inverted thresholds".Open in the repo

test_01 is the whole idea in one function. Nothing calls a model under test: the answer is typed in, and only the judge is called.

test_01_Anwser_Relevancy.py
# Setup: DeepEval needs a "judge" LLM to score the output. Pick one.
# Groq (cheap, OpenAI-compatible endpoint -> registered as a local model):
#   deepeval set-local-model --model openai/gpt-oss-120b \
#       --base-url "https://api.groq.com/openai/v1" --format json --prompt-api-key
# OpenAI:
#   export OPENAI_API_KEY=your_key_here
#   deepeval set-openai --model gpt-4o-mini
# Note: `deepeval set-grok` is xAI's Grok, NOT Groq.com.
# Run: pytest test_01_Anwser_Relevancy.py


from deepeval.test_case import LLMTestCase
from deepeval import assert_test
from deepeval.metrics import AnswerRelevancyMetric


def test_hello_world():

    test = LLMTestCase(
        input="What is 2+2?",
        actual_output="4",
        expected_output="4",
        context=["Basic arithmetice perform and give result"]
    )
    metric = [     AnswerRelevancyMetric(threshold=0.9)]
    assert_test(test, metric)

test_02 adds a real subject model. Note where the call happens: at module level, so pytest asks Groq while it is still collecting the file.

test_02_Groq_Qwen_Vs_GROQ_OpenAI.py
question = "What is 2+2? Reply with just the number."
answer = llm_response(question)
print(f"\n[Groq {GROQ_MODEL}] → {answer!r}\n")



def test_l4_with_judge_qwen():
    case = LLMTestCase(
        input=question,
        actual_output=answer,
        expected_output="4",
        # HallucinationMetric scores actual_output against this grounding text.
        context=["Basic arithmetic fact: 2 + 2 = 4"],
    )

    metrics = [
        AnswerRelevancyMetric(threshold=0.8, model=JUDGE_MODEL),
        HallucinationMetric(threshold=0.3, model=JUDGE_MODEL),
    ]
    assert_test(case, metrics)
sequenceDiagram
    participant P as pytest
    participant S as Subject qwen3.8-27b on Groq
    participant D as DeepEval
    participant J as Judge gpt-oss-120b on Groq
    P->>S: at import, ask the question at temperature 0
    S-->>P: raw answer text
    P->>D: assert_test with the LLMTestCase and two metrics
    D->>J: Answer Relevancy prompts
    J-->>D: score and reason
    D->>J: Hallucination prompts, against the context
    J-->>D: score and reason
    D-->>P: pass only if every score meets its threshold
test_02: the subject answers once, then a judge from a different model family scores the answer with two metrics.

03Live demo: build a test case and run it

Pick a case, choose which fields go into the LLMTestCase, tick metrics and set thresholds, then press Run pytest. The pass and fail rules are real: a metric errors when a field it reads is missing, and a score passes when it is at or above its threshold. The scores are fixed per case. The repo records no scores for this chapter, so they are illustrative, except Faithfulness and Correctness on the refund case, which chapter 16 recorded for that exact reply.

Build an LLMTestCase and run itNo API key needed
LLMTestCase fields (tick to include)
input
actual_output
Metrics and thresholds
reads input, actual_output
reads retrieval_context too
reads context too
reads expected_output too
data-testid=de-casedata-testid=de-use-expecteddata-testid=de-m-relevancydata-testid=de-t-relevancydata-testid=de-judgedata-testid=de-ruledata-testid=de-run
Results
    pytest output (simplified)
    
          The test you would write
          
    
          Judge set-up, once per machine
          
    
        
    Try this order: run test_01 as it is, then the wrong "5" (relevancy still passes), then tick Faithfulness on test_01 (a missing field is an error, not a low score), then switch the pass rule on test_02 to see what the old slide does to a clean answer.

    04Wire the judge

    The judge is configured once per machine. DeepEval defaults to OpenAI. Groq speaks the OpenAI API, so the course registers it as a "local model" and grades with openai/gpt-oss-120b for a fraction of the cost.

    flowchart TD
        T["pytest test_01_Anwser_Relevancy.py"] --> TC["LLMTestCase<br/>input + actual_output + context"]
        TC --> M["AnswerRelevancyMetric&#40;threshold=0.9&#41;"]
        M --> R{"Which judge?<br/>.deepeval provider flag"}
        R -->|USE_OPENAI_MODEL| O["api.openai.com<br/>gpt-4o-mini"]
        R -->|USE_LOCAL_MODEL| G["api.groq.com/openai/v1<br/>openai/gpt-oss-120b"]
        R -->|USE_OLLAMA_MODEL| L["localhost:11434<br/>free, slower"]
        O --> SC["Score 0.0 - 1.0"]
        G --> SC
        L --> SC
        SC -->|">= 0.9"| PASS["Test passes"]
        SC -->|"< 0.9"| FAIL["Test fails<br/>+ the judge's reason"]
    Which judge scores the test: the provider flag DeepEval saves decides where every metric call goes. From the course README.
    Notes.md
    ### Pick the judge LLM
    DeepEval scores your output with a *second* LLM. Configure it once per machine.
    
    **Groq (cheap, free tier).** Groq speaks the OpenAI API, so register it as a "local model".
    `deepeval set-grok` is xAI's Grok, not Groq.com, so do not use it.
    ```
    deepeval set-local-model \
      --model openai/gpt-oss-120b \
      --base-url "https://api.groq.com/openai/v1" \
      --format json \
      --prompt-api-key
    ```
    
    **OpenAI.**
    ```
    export OPENAI_API_KEY=sk-...
    deepeval set-openai --model gpt-4o-mini
    ```
    
    Add `--save dotenv:.env.local` to either command to persist the key across sessions.
    Switch back with `deepeval unset-local-model`.
    JudgeCostUse it whenKey it reads
    OpenAI gpt-4o-miniPaid, cheapYou want DeepEval's defaults and best-tested pathOPENAI_API_KEY
    Groq openai/gpt-oss-120bPaid, very cheap, free tierLocal runs and learning; fast, OpenAI-compatibleLOCAL_MODEL_API_KEY
    Ollama, localFreeNo key at all, offline; slower and less consistent scoresNone
    deepeval set-grok is xAI's Grok, not Groq. The names differ by one letter. Groq is wired in with set-local-model and its OpenAI-compatible base URL.
    One Groq key, two variable names. test_02's subject call reads OPENAI_API_KEY, which here holds a Groq key (it starts with gsk_). The judge reads LOCAL_MODEL_API_KEY. Leave the second one unset and the judge fails with "Local API key is not configured" while the subject call still works.

    05Threshold direction: the version trap

    The chapter's second slide teaches that Hallucination, Toxicity and Bias are inverted: lower is better and the threshold is a ceiling. That is the DeepEval 3.x convention. The course's later notes record that DeepEval 4.x made every metric pass on score >= threshold, so 1.0 is always the good end, and that with 4.x installed, test_02's HallucinationMetric(threshold=0.3) is a weak floor rather than a 30% ceiling.

    MetricSlide: good / bad scoreSlide rule (3.x)DeepEval 4.x rule
    AnswerRelevancy0.95 / 0.30threshold 0.7, must be >= 0.7Same: score >= threshold
    Faithfulness0.98 / 0.20threshold 0.8, must be >= 0.8Same
    Summarization0.85 / 0.25threshold 0.6, must be >= 0.6Same
    ContextualPrecision0.90 / 0.10threshold 0.7, must be >= 0.7Same
    ContextualRecall0.95 / 0.15threshold 0.7, must be >= 0.7Same
    Hallucination0.05 / 0.90threshold 0.5, must be <= 0.5Flipped: score >= threshold; 1.00 means nothing contradicts ground truth
    Toxicity0.02 / 0.85threshold 0.5, must be <= 0.5Flipped: 1.00 means no toxic or demeaning language
    Bias0.03 / 0.75threshold 0.5, must be <= 0.5Flipped: 1.00 means no bias detected in the reply
    The same threshold of 0.3 read with the old rule and with the DeepEval 4.x rule HallucinationMetric(threshold=0.3), test_02 Old slide (3.x) pass at score <= 0.3 pass fail DeepEval 4.x pass at score >= 0.3 fail pass 0.0 0.3 1.0 "5" scores 0.0 "4" scores 1.0
    One threshold, two readings. With 4.x scores a clean "4" scores 1.0 and passes; read with the old slide rule, the same 1.0 fails and a wrong "5" passes. Scores are illustrative.
    Confirm the rule from the installed source. The course's check: print BiasMetric.is_successful and read the comparison your version runs. Trust that over any tutorial or slide.
    terminal
    python -c "import inspect; from deepeval.metrics import BiasMetric; print(inspect.getsource(BiasMetric.is_successful))"

    06Set up and run

    From Notes.md and the course README. Python 3.11 or newer, one virtual environment, and a judge configured before the first run.

    terminal
    cd chapter_15_DeepEval
    python3 -m venv venv
    source venv/bin/activate
    pip install --upgrade pip
    pip install -U deepeval requests
    pip install openai python-dotenv   # test_02 imports both; Notes.md does not list them
    python -c "import deepeval, requests; print(deepeval.__version__, requests.__version__)"
    
    cp .env.sample .env                # set OPENAI_API_KEY and LOCAL_MODEL_API_KEY
    pytest test_01_Anwser_Relevancy.py
    pytest test_02_Groq_Qwen_Vs_GROQ_OpenAI.py
    deepeval test run test_01_Anwser_Relevancy.py   # DeepEval's own runner, same test
    Every metric is a paid call. The course README's arithmetic: a three-metric test on 100 golden cases is 300 paid calls. Treat that as a floor: one metric can make several judge calls (chapter 16 recorded three for a single Answer Relevancy case). Keep the golden set small, and never point metrics at a production key.

    The repo stores no run output for these two files. .env, .env.local, venv/ and .deepeval/ are gitignored; never commit a key.

    07Gotchas

    • test_02 calls Groq when it is imported. The question is asked at module level, so even pytest --collect-only runs that call: collecting the file needs a key and network access.
    • A 200 OK is not an answer. While test_02 was being built, the configured model id was a prompt-guard classifier. It returned '0.0003637653135228902' with no error: a jailbreak probability, not a reply. Assert on the value, not the status.
    • A variable name does not identify the provider. OPENAI_API_KEY held a Groq key. Read the prefix of the value, not the name of the variable.
    • temperature=0 pins only the subject. The judge is still non-deterministic, so the same test can score differently on two runs.
    • Different model families on purpose. qwen answers and gpt-oss grades, because "a judge scoring its own family inflates scores" (test_02's header).
    • Relevancy is not correctness. A confidently wrong "5" is on topic. Add a metric that reads context or expected_output whenever the facts matter.
    • Type the file name as it is. The first test is test_01_Anwser_Relevancy.py in the repo, spelling included.