The Testing Academy · Class Notes Sunday, 30 August (IST)
Live class · study guide

AI is not tested, it is evaluated

The hackathon results, then the topic that changes how you think about the job. An LLM is probabilistic, so assertEquals has nothing to hold on to. This session sets up the vocabulary of evaluation: ground truth, golden dataset, LLM-as-judge, faithfulness, and the threshold that turns a score back into a pass or a fail.

By Pramod Dutta, The Testing Academy. Study notes from the live AI Tester Blueprint 3x class, rebuilt from the session recording and cross-checked against chapter_14_LLM_Eval in the AITesterBlueprint3x repository, pushed during the session. Hackathon winners are named as announced publicly; individual feedback given privately in the session is deliberately not reproduced.

01

Hackathon results

Around sixty submissions on time, shortlisted to ten, and six recognised.

Place Project
1st AI-powered QA pipeline
2nd AI Detective
3rd Work Pulse
4th TraceFix
5th Parity Scope
Special mention Vision AI, submitted late

Prize money is being refunded by UPI once details are collected, and every participant gets a certificate. The next hackathon lands mid-September and the theme is LLM evaluations, which is exactly what the rest of this session sets up.

02

The term everyone uses wrongly

"AI testing" is the wrong phrase, and the correction is the spine of the whole session.

Testing compares an actual result against an expected one. You know what should come back, and you assert on it. Evaluation scores a response against criteria and compares that score to a threshold. There is no single right answer, so there is nothing to equal.

Testing, deterministic expected actual assertEquals same input, same output, every time Evaluation, probabilistic response judge scores it 0.82 >= 0.7 ? pass The threshold is what turns a probability back into a pass or a fail, so a suite can still go red.
You cannot assert on a probabilistic system. You score it, then draw a line.

Five reasons ordinary QA stops working, taken from the chapter notes:

Problem What breaks
Non-determinism Same prompt, different output. assertEquals has nothing to hold
Open-ended outputs No single correct summary. You measure along several axes
Hallucination Fluent is not correct. You need a faithfulness check against a source
Safety and bias Toxic or biased output is a failure, and jailbreaks are a new attack surface
Cost and latency Token cost and p95 latency are first-class signals, not afterthoughts

The example used: a fast-food chatbot that cheerfully wrote Python when asked for chicken nuggets. Nothing was broken in the traditional sense. There were simply no guardrails, and nobody had evaluated for it.

03

The three ways AI arrives in your product

Worth knowing because it tells you what you will be asked to evaluate:

  1. A chatbot
  2. An AI agent
  3. A RAG pipeline

All three are wrappers around an LLM, and all three need evaluating. The distinction restated for anyone shaky on it: an agent has a brain, memory and tool access. A chatbot may have none of those beyond the brain.

04

The vocabulary

This is the part to memorise, because every tool uses these words.

Term Means
Prompt The input: system instructions, user message, and often retrieved context
Completion The output you are evaluating
Ground truth The correct answer, written by a human, used as the reference
Golden dataset A curated set of input to expected-output pairs. Your regression suite, in other words your test data
Evaluator or judge Whatever scores the response: a rule, a model, or another LLM
LLM-as-judge Using a strong model to score another model's output
Hallucination Confidently stated, factually wrong
Faithfulness Does the answer stay true to the source it was given
Relevancy Does it actually answer what was asked
Contextual precision and recall Did retrieval fetch the right things, and did it fetch enough of them
Traces and spans The step-by-step record of what ran
Red teaming Deliberately trying to break it: jailbreaks, prompt injection, unsafe requests

The one QA people latch onto fastest: a golden dataset is just test data with a better name. Input, expected output, curated, versioned, run on every change. You have built these your whole career.

Use a different model as the judge than the one under test. Asking a model to grade its own homework builds in the bias you are trying to measure.

05

The tool, and why it is not the point

DeepEval is the course's choice: open source, more than twenty-five metrics, custom dashboards, and it wires into CI so evaluations run with your build.

Ragas was named as the main alternative, with the honest caveat that it is concentrated on RAG pipelines, so it is narrower. MLflow and LangSmith also came up.

The framing on the chapter's first line is the one to keep: the tool does not matter, the concept matters, and the tool can be replaced. The same sentence has now been said about n8n, about CrewAI providers, and about this. It is the most transferable thing in the course.

DeepEval is built on pytest, which the batch already knows. Markers, fixtures and the flags from the pytest reference carry over unchanged. An evaluation is a test with a scorer in place of an assertion.

06

Tasks and announcements

  • Next session: DeepEval hands on, covering setup, thresholds, parameters, and how it compares to Ragas.
  • Next hackathon is mid-September, themed on LLM evaluations, after DeepEval and LangChain.
  • Winners: a form is coming to collect UPI details for the cashback, and the full participant list will be shared.
  • Extra certification sessions are running Tuesday and Friday at 8 PM IST.
  • Chapter notes: chapter_14_LLM_Eval in the AITesterBlueprint3x repository.
  • Related: the DeepEval masterclass and the LLM evaluation cheat sheet.
  • Previous: the CrewAI bug triage crew.