DeepEval is a pytest plugin. You put the question, the model's answer and the known facts into an LLMTestCase, a second model called the judge scores it, and assert_test fails the test when a score misses its threshold. The chapter has two tests: a hard-coded 2+2 case, and a live Groq model graded by a judge from a different model family.
test_01 scores a hard-coded answer; test_02 asks a live Groq model first.
2
Metrics
AnswerRelevancyMetric and HallucinationMetric.
3
Judge options
OpenAI gpt-4o-mini, Groq openai/gpt-oss-120b, or a local Ollama model.
2
Key names to set
OPENAI_API_KEY for the subject call, LOCAL_MODEL_API_KEY for the judge.
01From assertEquals to assert_test
Arrange and act stay ordinary pytest. Only the assertion changes: assert_test(case, [metric]) sits where assert actual == expected used to, so CI needs no new runner. What is new is the test case itself. The chapter's slide maps each DeepEval parameter to the QA idea it replaces.
DeepEval parameter
QA / SDET equivalent (chapter slide)
In this chapter
input
Test input / request payload
Both tests: "What is 2+2?"
actual_output
Actual result from the system under test
Hard-coded in test_01, a live model answer in test_02
expected_output
Expected result / golden answer
"4" in both tests
context
Ground truth data (known facts)
Read by HallucinationMetric in test_02
retrieval_context
What the retriever actually returned
Not used here; the chapter 16 RAG tests use it
threshold
Acceptance criteria (for example, response time under 2s)
0.9 and 0.8 for relevancy, 0.3 for hallucination
model (judge LLM)
Test oracle / validation service
Set once per machine; test_02 also names it per metric
criteria, evaluation_params (G-Eval)
The assertion rubric, and which fields to assert on
Chapter 16
additional_metadata
Test tags / labels for filtering
Not used
The judge is the oracle, and it is another LLM. DeepEval does not compare strings. A second model reads the test case and returns a score from 0 to 1 with a written reason, and the threshold turns that score into a pass or a fail.
02The two tests
The chapter folder holds two test files, setup notes and two slides.
File
What it does
Run
test_01_Anwser_Relevancy.py
One hard-coded case ("What is 2+2?" answered "4") scored by AnswerRelevancyMetric(threshold=0.9). Only the judge model is called.
pytest test_01_Anwser_Relevancy.py
test_02_Groq_Qwen_Vs_GROQ_OpenAI.py
Asks qwen/qwen3.8-27b on Groq at temperature 0, then a judge from another family, openai/gpt-oss-120b, scores relevancy (0.8) and hallucination (0.3).
pytest test_02_Groq_Qwen_Vs_GROQ_OpenAI.py
Notes.md
Install checklist, verify command, judge set-up for Groq and OpenAI, run command, two watch-outs.
Read first
.env.sample
The two key names: OPENAI_API_KEY (subject call in test_02) and LOCAL_MODEL_API_KEY (the judge).
cp .env.sample .env
image-1.png, image.png
Two slides: DeepEval parameters as QA concepts, and "normal vs inverted thresholds".
Open in the repo
test_01 is the whole idea in one function. Nothing calls a model under test: the answer is typed in, and only the judge is called.
test_01_Anwser_Relevancy.py
# Setup: DeepEval needs a "judge" LLM to score the output. Pick one.# Groq (cheap, OpenAI-compatible endpoint -> registered as a local model):# deepeval set-local-model --model openai/gpt-oss-120b \# --base-url "https://api.groq.com/openai/v1" --format json --prompt-api-key# OpenAI:# export OPENAI_API_KEY=your_key_here# deepeval set-openai --model gpt-4o-mini# Note: `deepeval set-grok` is xAI's Grok, NOT Groq.com.# Run: pytest test_01_Anwser_Relevancy.pyfrom deepeval.test_case import LLMTestCase
from deepeval import assert_test
from deepeval.metrics import AnswerRelevancyMetric
deftest_hello_world():
test = LLMTestCase(
input="What is 2+2?",
actual_output="4",
expected_output="4",
context=["Basic arithmetice perform and give result"]
)
metric = [ AnswerRelevancyMetric(threshold=0.9)]
assert_test(test, metric)
test_02 adds a real subject model. Note where the call happens: at module level, so pytest asks Groq while it is still collecting the file.
test_02_Groq_Qwen_Vs_GROQ_OpenAI.py
question = "What is 2+2? Reply with just the number."
answer = llm_response(question)
print(f"\n[Groq {GROQ_MODEL}] → {answer!r}\n")
deftest_l4_with_judge_qwen():
case = LLMTestCase(
input=question,
actual_output=answer,
expected_output="4",
# HallucinationMetric scores actual_output against this grounding text.
context=["Basic arithmetic fact: 2 + 2 = 4"],
)
metrics = [
AnswerRelevancyMetric(threshold=0.8, model=JUDGE_MODEL),
HallucinationMetric(threshold=0.3, model=JUDGE_MODEL),
]
assert_test(case, metrics)
sequenceDiagram
participant P as pytest
participant S as Subject qwen3.8-27b on Groq
participant D as DeepEval
participant J as Judge gpt-oss-120b on Groq
P->>S: at import, ask the question at temperature 0
S-->>P: raw answer text
P->>D: assert_test with the LLMTestCase and two metrics
D->>J: Answer Relevancy prompts
J-->>D: score and reason
D->>J: Hallucination prompts, against the context
J-->>D: score and reason
D-->>P: pass only if every score meets its threshold
test_02: the subject answers once, then a judge from a different model family scores the answer with two metrics.
03Live demo: build a test case and run it
Pick a case, choose which fields go into the LLMTestCase, tick metrics and set thresholds, then press Run pytest. The pass and fail rules are real: a metric errors when a field it reads is missing, and a score passes when it is at or above its threshold. The scores are fixed per case. The repo records no scores for this chapter, so they are illustrative, except Faithfulness and Correctness on the refund case, which chapter 16 recorded for that exact reply.
Resultspytest output (simplified)The test you would writeJudge set-up, once per machine
Try this order: run test_01 as it is, then the wrong "5" (relevancy still passes), then tick Faithfulness on test_01 (a missing field is an error, not a low score), then switch the pass rule on test_02 to see what the old slide does to a clean answer.
04Wire the judge
The judge is configured once per machine. DeepEval defaults to OpenAI. Groq speaks the OpenAI API, so the course registers it as a "local model" and grades with openai/gpt-oss-120b for a fraction of the cost.
flowchart TD
T["pytest test_01_Anwser_Relevancy.py"] --> TC["LLMTestCase<br/>input + actual_output + context"]
TC --> M["AnswerRelevancyMetric(threshold=0.9)"]
M --> R{"Which judge?<br/>.deepeval provider flag"}
R -->|USE_OPENAI_MODEL| O["api.openai.com<br/>gpt-4o-mini"]
R -->|USE_LOCAL_MODEL| G["api.groq.com/openai/v1<br/>openai/gpt-oss-120b"]
R -->|USE_OLLAMA_MODEL| L["localhost:11434<br/>free, slower"]
O --> SC["Score 0.0 - 1.0"]
G --> SC
L --> SC
SC -->|">= 0.9"| PASS["Test passes"]
SC -->|"< 0.9"| FAIL["Test fails<br/>+ the judge's reason"]
Which judge scores the test: the provider flag DeepEval saves decides where every metric call goes. From the course README.
Notes.md
### Pick the judge LLM
DeepEval scores your output with a *second* LLM. Configure it once per machine.
**Groq (cheap, free tier).** Groq speaks the OpenAI API, so register it as a "local model".
`deepeval set-grok` is xAI's Grok, not Groq.com, so do not use it.
```
deepeval set-local-model \
--model openai/gpt-oss-120b \
--base-url "https://api.groq.com/openai/v1" \
--format json \
--prompt-api-key
```
**OpenAI.**
```
export OPENAI_API_KEY=sk-...
deepeval set-openai --model gpt-4o-mini
```
Add `--save dotenv:.env.local` to either command to persist the key across sessions.
Switch back with `deepeval unset-local-model`.
Judge
Cost
Use it when
Key it reads
OpenAI gpt-4o-mini
Paid, cheap
You want DeepEval's defaults and best-tested path
OPENAI_API_KEY
Groq openai/gpt-oss-120b
Paid, very cheap, free tier
Local runs and learning; fast, OpenAI-compatible
LOCAL_MODEL_API_KEY
Ollama, local
Free
No key at all, offline; slower and less consistent scores
None
deepeval set-grok is xAI's Grok, not Groq. The names differ by one letter. Groq is wired in with set-local-model and its OpenAI-compatible base URL.
One Groq key, two variable names. test_02's subject call reads OPENAI_API_KEY, which here holds a Groq key (it starts with gsk_). The judge reads LOCAL_MODEL_API_KEY. Leave the second one unset and the judge fails with "Local API key is not configured" while the subject call still works.
05Threshold direction: the version trap
The chapter's second slide teaches that Hallucination, Toxicity and Bias are inverted: lower is better and the threshold is a ceiling. That is the DeepEval 3.x convention. The course's later notes record that DeepEval 4.x made every metric pass on score >= threshold, so 1.0 is always the good end, and that with 4.x installed, test_02's HallucinationMetric(threshold=0.3) is a weak floor rather than a 30% ceiling.
Metric
Slide: good / bad score
Slide rule (3.x)
DeepEval 4.x rule
AnswerRelevancy
0.95 / 0.30
threshold 0.7, must be >= 0.7
Same: score >= threshold
Faithfulness
0.98 / 0.20
threshold 0.8, must be >= 0.8
Same
Summarization
0.85 / 0.25
threshold 0.6, must be >= 0.6
Same
ContextualPrecision
0.90 / 0.10
threshold 0.7, must be >= 0.7
Same
ContextualRecall
0.95 / 0.15
threshold 0.7, must be >= 0.7
Same
Hallucination
0.05 / 0.90
threshold 0.5, must be <= 0.5
Flipped: score >= threshold; 1.00 means nothing contradicts ground truth
Toxicity
0.02 / 0.85
threshold 0.5, must be <= 0.5
Flipped: 1.00 means no toxic or demeaning language
Bias
0.03 / 0.75
threshold 0.5, must be <= 0.5
Flipped: 1.00 means no bias detected in the reply
One threshold, two readings. With 4.x scores a clean "4" scores 1.0 and passes; read with the old slide rule, the same 1.0 fails and a wrong "5" passes. Scores are illustrative.
Confirm the rule from the installed source. The course's check: print BiasMetric.is_successful and read the comparison your version runs. Trust that over any tutorial or slide.
terminal
python -c "import inspect; from deepeval.metrics import BiasMetric; print(inspect.getsource(BiasMetric.is_successful))"
06Set up and run
From Notes.md and the course README. Python 3.11 or newer, one virtual environment, and a judge configured before the first run.
terminal
cd chapter_15_DeepEval
python3 -m venv venv
source venv/bin/activate
pip install --upgrade pip
pip install -U deepeval requests
pip install openai python-dotenv # test_02 imports both; Notes.md does not list thempython -c "import deepeval, requests; print(deepeval.__version__, requests.__version__)"cp .env.sample .env # set OPENAI_API_KEY and LOCAL_MODEL_API_KEYpytest test_01_Anwser_Relevancy.py
pytest test_02_Groq_Qwen_Vs_GROQ_OpenAI.py
deepeval test run test_01_Anwser_Relevancy.py # DeepEval's own runner, same test
Every metric is a paid call. The course README's arithmetic: a three-metric test on 100 golden cases is 300 paid calls. Treat that as a floor: one metric can make several judge calls (chapter 16 recorded three for a single Answer Relevancy case). Keep the golden set small, and never point metrics at a production key.
The repo stores no run output for these two files. .env, .env.local, venv/ and .deepeval/ are gitignored; never commit a key.
07Gotchas
test_02 calls Groq when it is imported. The question is asked at module level, so even pytest --collect-only runs that call: collecting the file needs a key and network access.
A 200 OK is not an answer. While test_02 was being built, the configured model id was a prompt-guard classifier. It returned '0.0003637653135228902' with no error: a jailbreak probability, not a reply. Assert on the value, not the status.
A variable name does not identify the provider.OPENAI_API_KEY held a Groq key. Read the prefix of the value, not the name of the variable.
temperature=0 pins only the subject. The judge is still non-deterministic, so the same test can score differently on two runs.
Different model families on purpose. qwen answers and gpt-oss grades, because "a judge scoring its own family inflates scores" (test_02's header).
Relevancy is not correctness. A confidently wrong "5" is on topic. Add a metric that reads context or expected_output whenever the facts matter.
Type the file name as it is. The first test is test_01_Anwser_Relevancy.py in the repo, spelling included.
DDrills for the chapter
Code drills use the two test files in chapter_15_DeepEval with a configured judge; judge scores and reasons vary from run to run, so the answers describe what holds, not exact numbers. Playwright drills target the live demo on this page (turn on Show locator badges to see every data-testid).
Relevancy versus truth
In test_01 change actual_output to "5" and run. Does it fail? Which metric would catch it?
Expected result
Relevancy can still pass: "5" answers the question that was asked, and the course notes that a confidently wrong answer can score 1.0 here. Add HallucinationMetric with test_02's context, ["Basic arithmetic fact: 2 + 2 = 4"]. test_01's own context states no fact to contradict.
An off-topic reply
Change actual_output to "Paris is the capital of France." and run test_01.
Expected result
Answer Relevancy drops far below 0.9 and the test fails with the judge's reason. The exact score and wording depend on the judge and the run.
Missing judge key
Register Groq with set-local-model but leave LOCAL_MODEL_API_KEY unset, then run test_02.
Expected result
The subject call works and the judge fails with "Local API key is not configured", as recorded in the course's build notes.
Read the value, not the status
Why is an HTTP 200 from the subject model not proof that the test is wired correctly?
Expected result
The prompt-guard classifier returned '0.0003637653135228902' with no error. Assert on the content (for example that the answer contains "4") before you trust any score built on it.
Which way does 0.3 point?
With DeepEval 4.x installed, what does HallucinationMetric(threshold=0.3) in test_02 require?
Expected result
A score of at least 0.3. It is a weak floor, not "at most 30% hallucination". Making it strict means raising it, which the course leaves as the test author's call.
Count the cost
How many paid calls does a three-metric test over 100 golden cases make, and why can the real bill be higher?
Expected result
300 by the course README's count, one per metric per case. Treat that as a floor: chapter 16 recorded three judge calls for a single Answer Relevancy case, so the number of requests can be several times higher.
Catch the wrong answer in the demoPlaywright
Select the "confidently wrong 5" case, run it, and assert the status and both result rows.
Hint
Locators: de-case, de-run, de-status, and de-results with li[data-kind="ok"] and li[data-kind="bad"].
Expected result
de-status reads 1 failed: Hallucination did not meet its threshold. One PASS row (Answer Relevancy 1.00 >= 0.90) and one FAIL row (Hallucination 0.00 >= 0.30).
Trigger a missing-field errorPlaywright
Load test_01, tick Faithfulness, and run. Assert that the run fails for the right reason.
Hint
The checkbox is de-m-faithfulness; test_01 has no retrieval_context, so de-use-retrieval is disabled.
Expected result
de-status contains Faithfulness needs retrieval_context and the results hold exactly one ERROR row.
Flip the pass rulePlaywright
Run test_02, then choose the old slide rule with getByLabel('Pass rule') and run again.
Expected result
First 1 passed: every metric met its threshold., then 1 failed: Hallucination did not meet its threshold. with a note that the old rule flipped Hallucination.
SSolutions: the demo spec and the chapter's files
The Playwright spec passes against this page as written. The other tabs are the chapter's files, verbatim from the course repo.
tests/deepeval-demo.spec.ts
import { test, expect } from'@playwright/test';
const URL = 'https://app.thetestingacademy.com/ai/blueprint/learn/deepeval.html';
test('relevancy passes a confidently wrong answer; hallucination fails it', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('de-case').selectOption('wrong');
await page.getByTestId('de-run').click();
awaitexpect(page.getByTestId('de-status')).toHaveText('1 failed: Hallucination did not meet its threshold.');
const results = page.getByTestId('de-results');
awaitexpect(results.locator('li[data-kind="ok"]')).toContainText('Answer Relevancy 1.00 >= 0.90');
awaitexpect(results.locator('li[data-kind="bad"]')).toContainText('Hallucination 0.00 >= 0.30');
});
test('a metric errors when the field it reads is missing', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('de-case').selectOption('t01');
await page.getByTestId('de-m-faithfulness').check();
await page.getByTestId('de-run').click();
awaitexpect(page.getByTestId('de-status')).toContainText('Faithfulness needs retrieval_context');
awaitexpect(page.getByTestId('de-results').locator('li[data-kind="bad"]')).toHaveCount(1);
});
test('reading a 4.x score with the old slide rule flips the verdict', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('de-case').selectOption('t02');
await page.getByTestId('de-run').click();
awaitexpect(page.getByTestId('de-status')).toHaveText('1 passed: every metric met its threshold.');
await page.getByLabel('Pass rule').selectOption('slide');
await page.getByTestId('de-run').click();
awaitexpect(page.getByTestId('de-status')).toContainText('1 failed: Hallucination did not meet its threshold.');
awaitexpect(page.getByTestId('de-status')).toContainText('flipped Hallucination');
});
test_02_Groq_Qwen_Vs_GROQ_OpenAI.py
#Flow:# 1. Send "What is 2+2?" to Groq Qwen3.8-27b (the model under test).# 2. Capture the raw answer.# 3. Hand input + answer + context to DeepEval.# 4. openai/gpt-oss-120b as judge -> scores AnswerRelevancy + Hallucination.# Subject != judge on purpose: a judge scoring its own family inflates scores.
GROQ_MODEL = "qwen/qwen3.8-27b"
JUDGE_MODEL = "openai/gpt-oss-120b"import os
from dotenv import load_dotenv
from openai import OpenAI
from deepeval.metrics import AnswerRelevancyMetric, HallucinationMetric
from deepeval.test_case import LLMTestCase
from deepeval import assert_test
load_dotenv()
# Groq speaks the OpenAI API, so the official openai SDK works unchanged:# just point base_url at Groq. The key in .env is a Groq key (gsk_...) that# lives under OPENAI_API_KEY because that is the name deepeval's# set-local-model path reads when talking to an OpenAI-compatible endpoint.
groq = OpenAI(
api_key=os.getenv("OPENAI_API_KEY"),
base_url="https://api.groq.com/openai/v1",
)
defllm_response(question):
"""Ask GROQ_MODEL one question and return its raw answer text."""
response = groq.chat.completions.create(
model=GROQ_MODEL,
messages=[{"role": "user", "content": question}],
# temperature=0 keeps the answer stable so the score below is# reproducible. The judge is still non-deterministic; this only# pins the thing under test.
temperature=0,
)
return response.choices[0].message.content.strip()
question = "What is 2+2? Reply with just the number."
answer = llm_response(question)
print(f"\n[Groq {GROQ_MODEL}] → {answer!r}\n")
deftest_l4_with_judge_qwen():
case = LLMTestCase(
input=question,
actual_output=answer,
expected_output="4",
# HallucinationMetric scores actual_output against this grounding text.
context=["Basic arithmetic fact: 2 + 2 = 4"],
)
metrics = [
AnswerRelevancyMetric(threshold=0.8, model=JUDGE_MODEL),
HallucinationMetric(threshold=0.3, model=JUDGE_MODEL),
]
assert_test(case, metrics)
Notes.md
# Deepeval
### Pre Req.
* Python Installed
* pytest Basics knowledge
* LLM Brain - openai key or deepseek api key or groq api key, claude api ($5),
* FREE LLM Brain
* nvidia api free
* amd api free (limited)

### Installation of Deepeval
* ```python3 -m venv venv```
* ```source venv/bin/activate```
* ```pip install --upgrade pip```
* ```pip install -U deepeval requests```
### Verify
```
python -c "import deepeval, requests; print(deepeval.__version__, requests.__version__)"
```

**Deactivate later:**
```
deactivate
```
### Pick the judge LLM
DeepEval scores your output with a *second* LLM. Configure it once per machine.
**Groq (cheap, free tier).** Groq speaks the OpenAI API, so register it as a "local model".
`deepeval set-grok` is xAI's Grok, not Groq.com, so do not use it.
```
deepeval set-local-model \
--model openai/gpt-oss-120b \
--base-url "https://api.groq.com/openai/v1" \
--format json \
--prompt-api-key
```
**OpenAI.**
```
export OPENAI_API_KEY=sk-...
deepeval set-openai --model gpt-4o-mini
```
Add `--save dotenv:.env.local` to either command to persist the key across sessions.
Switch back with `deepeval unset-local-model`.
### Run
```
pytest test_01_Anwser_Relevancy.py
```
### Watch out
* Every metric assertion is a real, paid LLM call. Keep the golden dataset small.
* `.env`, `.env.local`, `venv/` and `.deepeval/` are gitignored. Never commit a key.