The Testing Academy · Class Notes Saturday, 29 August (IST)
Live class · study guide

Three agents for the meeting nobody wants to attend

From a two-agent toy to something you could defend in a planning meeting. A bug triage crew of three agents that classifies, finds root cause, and recommends the tests that should have caught it, with Jira wired in as a tool and a fallback for when the model provider fails mid-run.

By Pramod Dutta, The Testing Academy. Study notes from the live AI Tester Blueprint 3x class, rebuilt from the session recording. The costing argument was worked through live and the presenter corrected it on the spot and called the numbers illustrative; the arithmetic is redone correctly here and still labelled illustrative. No API keys from the session are reproduced.

01

Where this sits

Last session built a first CrewAI agent, hardcoded and single-purpose. This one goes from that toy to a crew you could put in front of a manager.

The remaining path was restated: finish CrewAI, then LangChain, then DeepEval, then the extras.

02

Warming up: two agents, in order

Before the real thing, a smaller crew: a researcher that finds the top bugs in a class of web application and ranks them, and a technical writer that turns the findings into a prevention checklist.

Python
from crewai import Crew, Process

crew = Crew(
    agents=[researcher, writer],
    tasks=[research_task, write_task],
    process=Process.sequential,
)

The point of sequential here is not tidiness. The researcher must finish first, because the writer consumes its output. Order is the dependency.

Asked live: what happens if you do not set process? Sequential is the default, so it behaves the same, and the class sets it explicitly anyway so the intent is visible in the file. Process.hierarchical is the other option, where a manager agent decides who runs and in what order.

03

The meeting this replaces

Bug triage, defined plainly: set priority and severity, work out the root cause, decide what happens next. It is a bug lifecycle activity that manual testers do every day.

Then the argument that makes it worth automating, which is the part to reuse when you pitch this internally.

Today 30 people, 30 minutes, 20 days a month 300 person-hours a month on one recurring meeting With a crew doing the first pass the crew classifies and drafts, then a smaller group reviews. Five in the room, not thirty. 30 x 30 = 900 person-minutes per meeting. x 20 days = 18,000 person-minutes. / 60 = 300 person-hours. Illustrative numbers, and they were called that in the session. Put your own team's in. The humans do not disappear. The review stays. The room gets smaller.
The pitch is time returned to a team, not headcount removed from it.

The arithmetic was worked out live and corrected on the spot, so here it is done once, cleanly. Thirty people for thirty minutes is 900 person-minutes per meeting. Across twenty working days that is 18,000 person-minutes, which is 300 person-hours a month, not thirty. At a notional thirty dollars an hour it is around nine thousand dollars a month. The session called these dummy numbers and asked nobody to argue the figures, which is the right framing: the method is the point, and you should substitute your own team's numbers before you say any of this out loud.

04

The crew: three agents, three jobs

bug report from Jira 1. classify2. root cause3. recommend priority andseverity why did thisactually happen the tests thatwould have caught it human review Each agent gets exactly one job, and each one reads what the previous agent produced. The last box is not optional. The crew drafts; a person still signs.
Agent three is the interesting one: it turns a bug into the test that should have existed.

One agent, one task. Stated as a policy, and it matters more the bigger the workflow gets. An agent asked to classify and investigate and recommend does all three vaguely. Three agents with one responsibility each produce output you can actually review, because you can see which step went wrong.

05

Wiring Jira in

The bug does not arrive by hand. A plain Python function that calls the Jira API becomes a tool the crew can use, and the class deliberately started with pass as the body so the shape was clear before the implementation.

Asked live: could you use the Jira MCP server instead? Yes, if you have access to it. The function is the version that works everywhere.

Once one tool is attached, the others follow the same pattern:

Attach To get
Jira API or Jira MCP The ticket itself
Repository access What the code actually does
RAG over your docs Your team's history and standards
Slack trigger Triage that starts itself when a bug lands
06

When the provider falls over

Groq hit token and key problems during the session, for the second class running. The fix was not to retry harder: a DeepSeek fallback was added, with switchable provider logic so the crew keeps running when one provider refuses.

This is worth stealing as a habit rather than treating as a mishap. A crew that depends on one free provider is a crew that stops on the day you demo it. Make the provider a setting, not a hardcoded value, and you can move a failing run onto another brain without touching the agents. The same failure and the same fix showed up in the BLAST and n8n session a few hours earlier.

07

The bigger pipeline

The session closed on the shape of the full QA pipeline the crew grows into: analyse the Jira ticket, produce a test plan, produce test cases, generate Playwright code, put a UI on it, allow downloads, and wrap it in guardrails.

That last word is the one to hold on to. Every step above generates something a human will be asked to trust, and guardrails are what stop a confident wrong answer becoming a merged one.

08

Tasks and announcements

  • Build the bug triage crew. Three agents, one job each, Jira wired in as a tool.
  • Make the provider switchable while you are in there, so a rate limit does not end the run.
  • The code, prompts, documentation and results from the session will be shared separately.
  • The three free certifications are still the ask, and all three have notes here: Claude 101, Claude Code 101, AI Fluency.
  • Hackathon: 60-plus submissions, narrowed to a top 10 and then a top 5.
  • Coming next: LangChain, then DeepEval.
  • Previous: your first CrewAI agent. Reference: the CrewAI cheat sheet.