The Testing Academy · Class Notes Saturday, 5 September (IST)
Live class · study guide

Thresholds, and why every DeepEval score points the same way

From the concepts to a passing test. Why a threshold is a decision you make rather than a number you look up, the direction question that trips everyone up and how DeepEval quietly settles it, then the practical half: virtual environment, install, a judge model that is not OpenAI, and a first answer relevancy test that scores 1.0.

By Pramod Dutta, The Testing Academy. Study notes from the live AI Tester Blueprint 3x workshop, rebuilt from the session recording. DeepEval was installed and inspected before publishing, on version 4.2.1, so the package names, CLI commands and metric families here were checked rather than recalled. One point taught in the session, that hallucination, toxicity and bias score in the opposite direction, does not hold in DeepEval 4.x, and section 3 shows the source that settles it.

01

Where this sits

The previous session covered why LLM evaluation exists and the vocabulary. This one turns that into a running test.

The short answer to why any of this is needed: LLM output is not deterministic. Open-ended responses, hallucinations, safety and bias concerns, and cost and latency that vary run to run. None of that survives a plain assertEquals.

The framing from the room, and the best one-line definition offered: evaluation is a deterministic way to validate a probabilistic response.

Vocabulary worth having straight before the practical half: prompt and completion, ground truth and golden dataset, LLM as a judge, hallucination, faithfulness, relevancy, contextual precision, and traces and spans, which are logs. Plus harness, which simply means control: the thing that keeps the model's output in a shape you can work with.

02

A threshold is a decision, not a lookup

In ordinary QA every assertion is binary. The status code is 200 or it is not.

LLM metrics are not binary. Every metric in DeepEval produces a score between 0 and 1. You make it binary by drawing a line, and that line is the threshold.

The question that follows is the one that matters: who decides the number? You do. There is no correct value to look up. It is acceptance criteria, and acceptance criteria are a decision about what is good enough for your product.

DeepEvalThe QA word for it
thresholdAcceptance criteria
modelWhat is under test
evaluation_paramsWhich field you are asserting on
GEvalA custom assertion
inputTest input
actual_outputActual result
expected_outputThe golden answer
contextGround truth
retrieval_contextWhat the retrieval step actually returned
Additional metadataTags and labels

Read that table once and DeepEval stops looking like a new discipline. It is the same test anatomy with different field names.

03

The trap, and how DeepEval actually resolves it

The session asked for full attention here, and it is worth it, though the resolution is the opposite of what you might expect.

Ask it as the session did. The threshold is 0.75 and you scored 0.81. Did you pass?

The reasoning offered was that it depends on the metric: for relevancy a high score is good, but for hallucination, toxicity or bias a high score means you have a lot of a bad thing, so 0.81 would be a failure. That is a real distinction in LLM evaluation generally, and it is how several tools work.

DeepEval is not one of them. Every metric in DeepEval scores the good direction, so the answer is simply yes, you passed.

Read it from the installed source rather than taking anyone's word, this being version 4.2.1:

Code
# HallucinationMetric._calculate_score
factually_aligned_count = 0
for verdict in self.verdicts:
    if verdict.verdict.strip().lower() != "no":
        factually_aligned_count += 1
score = factually_aligned_count / number_of_verdicts

It counts the claims that are factually aligned. ToxicityMetric counts non_toxic_count. BiasMetric counts unbiased_count. And every metric, without exception, decides pass the same way:

Code
self.success = self.score >= self.threshold
Every metric, the same direction 0 1 threshold 0.75 pass fail Obviously good things relevancy, faithfulness, recall the score is how much you have Apparently bad things hallucination, toxicity, bias the score is how much you avoided DeepEval flips the bad ones for you, so the threshold is always a floor. 0.81 passes either way.
The distinction is real in the abstract. DeepEval removes it by always scoring the direction you want more of.

This page originally said the opposite, and the correction matters more than the mistake. If you set a hallucination threshold expecting a ceiling, you will write 0.2 meaning "allow up to twenty percent hallucination" and DeepEval will read it as "at least twenty percent of claims must be factually aligned", which passes almost anything. The threshold is a floor, always. For hallucination you want it high, not low.

The confusion is not invented: the concept exists, other evaluation tools do invert, and DeepEval itself scored some metrics the other way before 4.x. Confirm the direction from the installed source, not from a tutorial, because this changed under everyone's feet.

What threshold should I start with? The guidance given: start around 0.5, run it against real output, look at the distribution you actually get, then calibrate. Tighten as the code moves toward production: loosest locally, tighter in QA, highest in production. A threshold set before you have seen any scores is a guess. And because the threshold is always a floor, tightening always means moving it up.

04

Setting up

Prerequisites, and the third one is the one that stops people:

  1. Python installed.
  2. pytest basics, which the batch has already covered.
  3. An LLM API key. DeepEval uses a model as the judge, so it needs a brain of its own.
Code
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install deepeval requests

Verify by importing rather than trusting the installer:

Code
python -c "import deepeval; print(deepeval.__version__)"

deactivate closes the environment when you are done.

Same rule as the other session today: a subscription is not an API key. A Claude plan needs API credit added separately. GitHub Copilot gives you nothing here. DeepEval needs a judge model and that judge costs money or quota.

Free or cheap options: Groq is the meaningfully free one, NVIDIA and AMD give limited free quota, and if you are going to spend a few dollars the recommendation was OpenRouter rather than any single vendor, because one balance reaches many models.

05

When the judge has no credit

The demo hit a 429 on OpenAI: the key was valid, the account had no credits.

Worth knowing what this looks like, because it is not subtle and it happens at the wrong moment. On DeepEval 4.2.1, a metric with no key configured fails the instant you construct it:

Code
DeepEvalError: OpenAI API key is not configured.
Set OPENAI_API_KEY in your environment or pass api_key to OpenAIModel(...).

Not when the test runs. When the metric object is created. So a missing key surfaces as a collection error, which is confusing the first time.

The fix used in class was to point DeepEval at Groq instead:

Code
deepeval set-local-model

It prompts for the base URL, the model name and the response format, which is JSON. Any OpenAI-compatible endpoint works through the same command, which is what makes Groq, OpenRouter, NVIDIA and a local model interchangeable here.

06

The first test

Code
from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import AnswerRelevancyMetric


def test_answer_relevancy():
    test_case = LLMTestCase(
        input="What is Playwright used for?",
        actual_output="Playwright is a framework for end to end browser automation testing.",
    )
    metric = AnswerRelevancyMetric(threshold=0.9)
    assert_test(test_case, [metric])

Run it with DeepEval's own runner rather than plain pytest:

Code
deepeval test run test_answer_relevancy.py

The demo scored 1.0 against a threshold of 0.9, and passed.

LLMTestCase input actual_output data only, no assertion AnswerRelevancyMetric threshold=0.9 this is the assertion, and its acceptance criteria The judge a second LLM reads the output and scores it needs its own API key score 1.0, so it passes a number between 0 and 1 One test case can be run against several metrics at once, which is why they are separate.
The test case carries the data, the metric carries the judgement. Splitting them is what lets one case be scored several ways.

Two things about that shape are worth naming. The test case is data, an input and an output, with no assertions in it. The metric is the assertion, and the threshold is its acceptance criteria. That separation is why one test case can be run against several metrics at once.

AnswerRelevancyMetric asks a narrow question: does the output actually answer the input? It says nothing about whether the answer is true. That is FaithfulnessMetric, and it needs context to check against.

Asked in class: should the judge be a better model than the one under test? Broadly yes. The judge has to be capable enough to assess the output it is grading, so grading a strong model with a weak judge tells you more about the judge than the model.

07

Seeing the results

deepeval view pushes a run to the Confident AI dashboard, which is free for a single project and paid beyond that. It gives you scores, reasons and history without building anything.

The plan mentioned in the session is to build a local dashboard instead, to avoid the paid tier. Worth knowing both routes exist: the hosted one to see it working today, your own if this becomes part of a real pipeline.

08

Tasks and announcements

Today's task: write your own first answer relevancy test case, get it running, and share a screenshot of it passing.

  • Hackathon: prize money for places one to five goes out this evening. Participation certificates for all roughly 65 to 70 entrants follow. The next hackathon is planned for late September or early October, at an advanced level.
  • Certifications this month: Introduction to MCP, Sub-Agents, AI Capabilities and Limitations, and Advanced MCP. Four to five is the target.
  • Extra sessions continue, usually 8 PM IST on a weekday, and stay free for 3x members through the end of the year.
  • A new Playwright batch starts in the third week of September, from zero coding: JavaScript, TypeScript, Playwright, API, Cucumber, an advanced framework and an agent factory. It is the last one this year, with early bird pricing to be announced.
  • A tutorial on using the free NVIDIA and AMD APIs is being recorded, which will make the judge model free for anyone following along.