Hackathon results
Around sixty submissions on time, shortlisted to ten, and six recognised.
| Place | Project |
|---|---|
| 1st | AI-powered QA pipeline |
| 2nd | AI Detective |
| 3rd | Work Pulse |
| 4th | TraceFix |
| 5th | Parity Scope |
| Special mention | Vision AI, submitted late |
Prize money is being refunded by UPI once details are collected, and every participant gets a certificate. The next hackathon lands mid-September and the theme is LLM evaluations, which is exactly what the rest of this session sets up.
The term everyone uses wrongly
"AI testing" is the wrong phrase, and the correction is the spine of the whole session.
Testing compares an actual result against an expected one. You know what should come back, and you assert on it. Evaluation scores a response against criteria and compares that score to a threshold. There is no single right answer, so there is nothing to equal.
Five reasons ordinary QA stops working, taken from the chapter notes:
| Problem | What breaks |
|---|---|
| Non-determinism | Same prompt, different output. assertEquals has nothing to hold |
| Open-ended outputs | No single correct summary. You measure along several axes |
| Hallucination | Fluent is not correct. You need a faithfulness check against a source |
| Safety and bias | Toxic or biased output is a failure, and jailbreaks are a new attack surface |
| Cost and latency | Token cost and p95 latency are first-class signals, not afterthoughts |
The example used: a fast-food chatbot that cheerfully wrote Python when asked for chicken nuggets. Nothing was broken in the traditional sense. There were simply no guardrails, and nobody had evaluated for it.
The three ways AI arrives in your product
Worth knowing because it tells you what you will be asked to evaluate:
- A chatbot
- An AI agent
- A RAG pipeline
All three are wrappers around an LLM, and all three need evaluating. The distinction restated for anyone shaky on it: an agent has a brain, memory and tool access. A chatbot may have none of those beyond the brain.
The vocabulary
This is the part to memorise, because every tool uses these words.
| Term | Means |
|---|---|
| Prompt | The input: system instructions, user message, and often retrieved context |
| Completion | The output you are evaluating |
| Ground truth | The correct answer, written by a human, used as the reference |
| Golden dataset | A curated set of input to expected-output pairs. Your regression suite, in other words your test data |
| Evaluator or judge | Whatever scores the response: a rule, a model, or another LLM |
| LLM-as-judge | Using a strong model to score another model's output |
| Hallucination | Confidently stated, factually wrong |
| Faithfulness | Does the answer stay true to the source it was given |
| Relevancy | Does it actually answer what was asked |
| Contextual precision and recall | Did retrieval fetch the right things, and did it fetch enough of them |
| Traces and spans | The step-by-step record of what ran |
| Red teaming | Deliberately trying to break it: jailbreaks, prompt injection, unsafe requests |
The one QA people latch onto fastest: a golden dataset is just test data with a better name. Input, expected output, curated, versioned, run on every change. You have built these your whole career.
Use a different model as the judge than the one under test. Asking a model to grade its own homework builds in the bias you are trying to measure.
The tool, and why it is not the point
DeepEval is the course's choice: open source, more than twenty-five metrics, custom dashboards, and it wires into CI so evaluations run with your build.
Ragas was named as the main alternative, with the honest caveat that it is concentrated on RAG pipelines, so it is narrower. MLflow and LangSmith also came up.
The framing on the chapter's first line is the one to keep: the tool does not matter, the concept matters, and the tool can be replaced. The same sentence has now been said about n8n, about CrewAI providers, and about this. It is the most transferable thing in the course.
DeepEval is built on pytest, which the batch already knows. Markers, fixtures and the flags from the pytest reference carry over unchanged. An evaluation is a test with a scorer in place of an assertion.
Tasks and announcements
- Next session: DeepEval hands on, covering setup, thresholds, parameters, and how it compares to Ragas.
- Next hackathon is mid-September, themed on LLM evaluations, after DeepEval and LangChain.
- Winners: a form is coming to collect UPI details for the cashback, and the full participant list will be shared.
- Extra certification sessions are running Tuesday and Friday at 8 PM IST.
- Chapter notes:
chapter_14_LLM_Evalin the AITesterBlueprint3x repository. - Related: the DeepEval masterclass and the LLM evaluation cheat sheet.
- Previous: the CrewAI bug triage crew.