Where this sits
The previous session covered why LLM evaluation exists and the vocabulary. This one turns that into a running test.
The short answer to why any of this is needed: LLM output is not deterministic. Open-ended responses, hallucinations, safety and bias concerns, and cost and latency that vary run to run. None of that survives a plain assertEquals.
The framing from the room, and the best one-line definition offered: evaluation is a deterministic way to validate a probabilistic response.
Vocabulary worth having straight before the practical half: prompt and completion, ground truth and golden dataset, LLM as a judge, hallucination, faithfulness, relevancy, contextual precision, and traces and spans, which are logs. Plus harness, which simply means control: the thing that keeps the model's output in a shape you can work with.
A threshold is a decision, not a lookup
In ordinary QA every assertion is binary. The status code is 200 or it is not.
LLM metrics are not binary. Every metric in DeepEval produces a score between 0 and 1. You make it binary by drawing a line, and that line is the threshold.
The question that follows is the one that matters: who decides the number? You do. There is no correct value to look up. It is acceptance criteria, and acceptance criteria are a decision about what is good enough for your product.
| DeepEval | The QA word for it |
|---|---|
threshold | Acceptance criteria |
model | What is under test |
evaluation_params | Which field you are asserting on |
GEval | A custom assertion |
input | Test input |
actual_output | Actual result |
expected_output | The golden answer |
context | Ground truth |
retrieval_context | What the retrieval step actually returned |
| Additional metadata | Tags and labels |
Read that table once and DeepEval stops looking like a new discipline. It is the same test anatomy with different field names.
The trap, and how DeepEval actually resolves it
The session asked for full attention here, and it is worth it, though the resolution is the opposite of what you might expect.
Ask it as the session did. The threshold is 0.75 and you scored 0.81. Did you pass?
The reasoning offered was that it depends on the metric: for relevancy a high score is good, but for hallucination, toxicity or bias a high score means you have a lot of a bad thing, so 0.81 would be a failure. That is a real distinction in LLM evaluation generally, and it is how several tools work.
DeepEval is not one of them. Every metric in DeepEval scores the good direction, so the answer is simply yes, you passed.
Read it from the installed source rather than taking anyone's word, this being version 4.2.1:
# HallucinationMetric._calculate_score
factually_aligned_count = 0
for verdict in self.verdicts:
if verdict.verdict.strip().lower() != "no":
factually_aligned_count += 1
score = factually_aligned_count / number_of_verdicts
It counts the claims that are factually aligned. ToxicityMetric counts non_toxic_count. BiasMetric counts unbiased_count. And every metric, without exception, decides pass the same way:
self.success = self.score >= self.threshold
This page originally said the opposite, and the correction matters more than the mistake. If you set a hallucination threshold expecting a ceiling, you will write 0.2 meaning "allow up to twenty percent hallucination" and DeepEval will read it as "at least twenty percent of claims must be factually aligned", which passes almost anything. The threshold is a floor, always. For hallucination you want it high, not low.
The confusion is not invented: the concept exists, other evaluation tools do invert, and DeepEval itself scored some metrics the other way before 4.x. Confirm the direction from the installed source, not from a tutorial, because this changed under everyone's feet.
What threshold should I start with? The guidance given: start around 0.5, run it against real output, look at the distribution you actually get, then calibrate. Tighten as the code moves toward production: loosest locally, tighter in QA, highest in production. A threshold set before you have seen any scores is a guess. And because the threshold is always a floor, tightening always means moving it up.
Setting up
Prerequisites, and the third one is the one that stops people:
- Python installed.
- pytest basics, which the batch has already covered.
- An LLM API key. DeepEval uses a model as the judge, so it needs a brain of its own.
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install deepeval requests
Verify by importing rather than trusting the installer:
python -c "import deepeval; print(deepeval.__version__)"
deactivate closes the environment when you are done.
Same rule as the other session today: a subscription is not an API key. A Claude plan needs API credit added separately. GitHub Copilot gives you nothing here. DeepEval needs a judge model and that judge costs money or quota.
Free or cheap options: Groq is the meaningfully free one, NVIDIA and AMD give limited free quota, and if you are going to spend a few dollars the recommendation was OpenRouter rather than any single vendor, because one balance reaches many models.
When the judge has no credit
The demo hit a 429 on OpenAI: the key was valid, the account had no credits.
Worth knowing what this looks like, because it is not subtle and it happens at the wrong moment. On DeepEval 4.2.1, a metric with no key configured fails the instant you construct it:
DeepEvalError: OpenAI API key is not configured.
Set OPENAI_API_KEY in your environment or pass api_key to OpenAIModel(...).
Not when the test runs. When the metric object is created. So a missing key surfaces as a collection error, which is confusing the first time.
The fix used in class was to point DeepEval at Groq instead:
deepeval set-local-model
It prompts for the base URL, the model name and the response format, which is JSON. Any OpenAI-compatible endpoint works through the same command, which is what makes Groq, OpenRouter, NVIDIA and a local model interchangeable here.
The first test
from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import AnswerRelevancyMetric
def test_answer_relevancy():
test_case = LLMTestCase(
input="What is Playwright used for?",
actual_output="Playwright is a framework for end to end browser automation testing.",
)
metric = AnswerRelevancyMetric(threshold=0.9)
assert_test(test_case, [metric])
Run it with DeepEval's own runner rather than plain pytest:
deepeval test run test_answer_relevancy.py
The demo scored 1.0 against a threshold of 0.9, and passed.
Two things about that shape are worth naming. The test case is data, an input and an output, with no assertions in it. The metric is the assertion, and the threshold is its acceptance criteria. That separation is why one test case can be run against several metrics at once.
AnswerRelevancyMetric asks a narrow question: does the output actually answer the input? It says nothing about whether the answer is true. That is FaithfulnessMetric, and it needs context to check against.
Asked in class: should the judge be a better model than the one under test? Broadly yes. The judge has to be capable enough to assess the output it is grading, so grading a strong model with a weak judge tells you more about the judge than the model.
Seeing the results
deepeval view pushes a run to the Confident AI dashboard, which is free for a single project and paid beyond that. It gives you scores, reasons and history without building anything.
The plan mentioned in the session is to build a local dashboard instead, to avoid the paid tier. Worth knowing both routes exist: the hosted one to see it working today, your own if this becomes part of a real pipeline.
Tasks and announcements
Today's task: write your own first answer relevancy test case, get it running, and share a screenshot of it passing.
- Hackathon: prize money for places one to five goes out this evening. Participation certificates for all roughly 65 to 70 entrants follow. The next hackathon is planned for late September or early October, at an advanced level.
- Certifications this month: Introduction to MCP, Sub-Agents, AI Capabilities and Limitations, and Advanced MCP. Four to five is the target.
- Extra sessions continue, usually 8 PM IST on a weekday, and stay free for 3x members through the end of the year.
- A new Playwright batch starts in the third week of September, from zero coding: JavaScript, TypeScript, Playwright, API, Cucumber, an advanced framework and an agent factory. It is the last one this year, with early bird pricing to be announced.
- A tutorial on using the free NVIDIA and AMD APIs is being recorded, which will make the judge model free for anyone following along.