CrewAI writes the person behind a prompt as code: an Agent has a role, a goal and a backstory, a Task says what to deliver, and a Crew runs the tasks in order. The chapter goes from one test-analyst agent to a three-agent crew that triages a Jira bug, finds the likely root cause and plans the tests, and shows how to keep it inside a free Groq key's 8,000-token request limit.
01 to 05. 03 is the in-class skeleton; 04 is its finished version.
3
Specialists in the triage crew
Triage, root cause and test strategy, chained with context.
8,000
Groq tokens per request
Free tier: prompt plus max_tokens together.
3,200
Groq budget for agent 3
8,000 minus its ~4.8k prompt. The last agent has the least room.
01What CrewAI gives a tester
Up to now a prompt was something you pasted into a chat. CrewAI lets you write the expert behind the prompt once, as code, and reuse it on every ticket, in a pipeline or on a schedule. Five names cover everything this chapter uses:
Concept
In CrewAI
In the triage crew (04)
LLM
LLM(model, api_key, base_url, max_tokens)
One per agent, each with its own token budget
Agent
Agent(role, goal, backstory, llm): who does the work
A senior triage analyst whose backstory carries the S0-S4 and P0-P4 rulebook
Task
Task(description, expected_output, agent, context): what to deliver, and in what shape
"Triage this bug report", with a 400-word HARD LIMIT
Crew
Crew(agents, tasks, process): the team and the order
Three specialists, Process.sequential
kickoff()
Runs the crew and returns a result object
result.tasks_output holds all three reports
The four objects every crew in this chapter uses. The agent is the who, the task is the what, and the crew decides the order.
Why a triage crew? The docstring of 04_Build_QABugTriageCrew_Prod.py makes the business case: a daily triage meeting of 30 minutes for 30 people is about 300 hours a month. A crew that pre-triages every ticket before the meeting buys most of that back, and because each verdict is asked to quote the line of the bug report that drove it, a person can check it quickly.
02Your first agent in four steps
01_test_analyst_Agent.py is the smallest useful crew: set up the brain, define the agent, give it a task, add both to a crew and kick it off. The comments in the file name the steps the same way.
01_test_analyst_Agent.py (lines 29-67)
# Step 0 - Set up the Brain (Groq LLM)
load_dotenv() # reads the .env file in this folder
# Groq exposes an OpenAI-compatible API, so we use the "openai/" provider
# prefix with a custom base_url pointing at Groq.
# GROQ_MODEL in .env = openai/gpt-oss-120b (exact ID from groq.com console)
groq_llm = LLM(
model=f"openai/{os.getenv('GROQ_MODEL')}",
api_key=os.getenv("GROQ_API_KEY"),
base_url=os.getenv("GROQ_BASE_URL"),
)
# Step 1. - Define the Agent (identity)
qa_agent = Agent(
role="QA Enginner",
goal="Analyse the feature or the requirements, and create 5-10 test cases out of it.",
backstory="You are a senior QA engineer with 15 years of experience in test planning and testcases creation",
llm = groq_llm,
verbose=True
)
# Step 2 - Give the Task to the Agent
test_case_task = Task(
description="Create 5-10 test cases",
expected_output="A numbered list of 5-10 test cases with brief descriptions for a app.vwo.com Login page with the username, password and submit button with remember me functionality",
agent=qa_agent
)
# Step 3. Add them to the Crew
crew = Crew(
agents=[qa_agent],
tasks=[test_case_task],
verbose=True
)
# Step 4. kickOff
if __name__ == "__main__":
result = crew.kickoff()
print(result)
The brain is Groq. CrewAI defaults to OpenAI. Groq speaks the same API, so the code sets base_url to Groq and prefixes the model with openai/.
The agent is the who.role, goal and backstory become the persona the model is told to play. One agent can take many tasks.
The task is the what. Here description is only "Create 5-10 test cases" and the page details sit in expected_output. Drill 1 moves them where they belong.
Script 01 from configuration to output. The root README lists what a run covers: valid login, empty fields, password masking, remember-me persistence, injection protection and accessibility.
03Live demo: kick off a crew
Pick a crew and press Kick off, or use Run next task to go one task at a time. Each task card shows the agent and the exact context wiring from the script, and the panel on the right shows what the next task receives, including the earlier outputs. For the 04 crew, change a Groq budget or a key and watch the FallbackLLM rules from the script decide what happens.
Crew runner: sequential tasks, context and the token squeezeNo API key needed
Groq request size: prompt + max_tokens against the 8,000 capInputs to the current taskTask outputs (illustrative text, not a recorded run)
What is real here: the agents, tasks, context lists, budgets, the 8,000 cap, the approximate prompt sizes (from the repo's learnings note) and the fallback logic. What is not: the task outputs. The repo keeps no recorded run, so they are illustrative text written for this page. A real run words, rates and scores things differently every time.
04Sequential crews and context
With Process.sequential the tasks run in list order. What each task receives from the earlier ones depends on context:
No context set (script 02): the sequential process passes the earlier output along, so the writer sees the research without any wiring.
context=[...] set (script 04): the task receives exactly the outputs of the tasks you list. The root cause task gets the triage verdict; the test strategy task gets both.
04_Build_QABugTriageCrew_Prod.py (lines 595-637)
test_task = Task(
description=f"""Using the triage verdict and the root cause analysis from the
previous tasks, design the test strategy for this bug:
{bug_report}
Deliver:
1. THE MISSING TEST: name the single test that should have existed and would have
caught this before production, and explain why the current suite missed it.
2. VERIFICATION TEST: the one test that proves the fix works. State its level.
3. REGRESSION SET: 3-5 cases that stop this specific bug returning, each pinned to
the cheapest level that can catch it (unit / API / E2E) with a one-line reason.
4. BOUNDARY AND EDGE CASES: apply boundary value analysis to the failing threshold
in this report, plus the negative and state-transition cases.
5. PLAYWRIGHT TYPESCRIPT: runnable code for the E2E tests worth automating,
following your standards (web-first assertions, no hard waits, role/testid
locators, API-seeded state, route interception where it isolates the layer).
6. WHAT TO KEEP MANUAL and why.
7. EXIT CRITERIA: what must be green before this ticket is closed.""",
expected_output="""A test strategy in markdown: the Missing Test, a table of tests
(ID, Level, Type, What it proves, Priority), runnable Playwright TypeScript code
blocks for the E2E cases, a Manual-only section with justification, and Exit
Criteria.
HARD LIMIT: 800 words including the code. You are the last agent and have
the least room, so budget it: keep the tables narrow, write ONE complete
Playwright spec file rather than several half-finished ones, and make sure
you reach Exit Criteria. A complete short plan beats a truncated long one.""",
agent=test_recommender,
context=[triage_task, root_cause_task], # Uses both outputs
)
# ---------------------------------------------------------------------------
# Step 3 + 4 - Assemble the Crew and kick it off
# ---------------------------------------------------------------------------
crew = Crew(
agents=[bug_analyst, root_cause_agent, test_recommender],
tasks=[triage_task, root_cause_task, test_task],
process=Process.sequential,
verbose=True,
)
flowchart TD
ID["BUG_ID, default VWO-48"] --> J["requests.get: /rest/api/3/issue/BUG_ID"]
J -->|"200"| ADF["_adf_to_text(): flat bug text"]
J -->|"any error"| C["bug_cache/VWO-48.txt"]
ADF --> T1
C --> T1
T1["Task 1: triage_task<br/>Bug Triage Analyst"] -->|"context"| T2["Task 2: root_cause_task<br/>Root Cause Investigator"]
T1 -->|"context"| T3["Task 3: test_task<br/>Test Strategy Advisor"]
T2 -->|"context"| T3
T3 --> R["result.tasks_output: three reports"]
The 04 crew. Jira first, the local cache when Jira fails, then three specialists, each reading the verdicts before it.
Read every agent, not just the last. Printing crew.kickoff() shows only the final task's output. 04 loops over result.tasks_output to print all three reports, labelled by agent.
05The token squeeze: budget every agent
On Groq's free tier one request may not exceed 8,000 tokens, and that counts the prompt and max_tokens together. In a sequential crew the prompt grows at every step, because each task carries the earlier outputs. So the agent with the most to write has the least room to write it. The repo's learnings note measured the shape:
Agent (LLM)
Prompt size, approx.
Room under 8,000
Groq max_tokens
DeepSeek max_tokens
Word limit
Bug Triage Analyst (triage_llm)
~2.3k
~5.7k
4000
4000
400
Root Cause Investigator (rca_llm)
~3.0k
~5.0k
4000
5000
650
Test Strategy Advisor (sdet_llm)
~4.8k
~3.2k
3200
6000
800
04_Build_QABugTriageCrew_Prod.py (lines 137-171)
def make_llm(groq_tokens, deepseek_tokens):
"""Build one LLM per agent, with the other provider wired in as a parachute.
Each provider gets its own budget because their limits differ: Groq must
stay under GROQ_TPM_LIMIT (prompt + output), DeepSeek can breathe.
"""
has_groq = bool(os.getenv("GROQ_API_KEY"))
has_deepseek = bool(os.getenv("DEEPSEEK_API_KEY"))
groq = make_groq_llm(groq_tokens) if has_groq else None
deepseek = make_deepseek_llm(deepseek_tokens) if has_deepseek else None
if PRIMARY_PROVIDER == "groq":
primary, secondary = groq, deepseek
else:
primary, secondary = deepseek, groq
primary = primary or secondary
if primary is None:
raise RuntimeError("Set GROQ_API_KEY and/or DEEPSEEK_API_KEY in .env")
if secondary is primary:
secondary = None
if secondary is None:
return primary
return FallbackLLM(
model=f"{PRIMARY_PROVIDER}-with-fallback", primary=primary, secondary=secondary
)
# Budgets: (groq, deepseek). The Groq numbers keep prompt + output under
# GROQ_TPM_LIMIT; the DeepSeek numbers are what these reports actually need.
triage_llm = make_llm(4000, 4000) # smallest prompt, runs first
rca_llm = make_llm(4000, 5000) # prompt also carries task 1's output
sdet_llm = make_llm(3200, 6000) # prompt also carries task 1 + task 2 output
Two fixes work together. Each agent gets its own LLM with a budget that fits its position in the chain (make_llm(groq_tokens, deepseek_tokens)), and every expected_output ends with a HARD LIMIT word count and the reason for it ("every word you spend costs them room"), because a wordy first agent steals room from the last one.
How the author found it: raising max_tokens to 8000 turned silent mid-sentence truncation into a loud error, Limit 8000, Requested 10421. The note also records that every model on the key had the same 8,000 limit, so switching models was not an escape: the budget had to be engineered.
06A fallback that also catches empty answers
A free key will fail mid-demo: a 429 rate limit, a 413 request too large, an expired key, an outage. FallbackLLM wraps two providers and re-issues the failed call on the second one, so the crew keeps going. The detail that matters is the first branch: an empty string counts as a failure too.
04_Build_QABugTriageCrew_Prod.py (lines 77-95)
def call(self, *args, **kwargs):
try:
response = self.primary.call(*args, **kwargs)
if not self._is_empty(response):
return response
reason = "returned an EMPTY response"
except Exception as exc:
reason = f"raised {type(exc).__name__}: {str(exc)[:150]}"
if self.secondary is None:
raise
if self.secondary is None:
raise RuntimeError(f"Primary LLM {reason} and no fallback is configured")
print(f"\n[llm] PRIMARY {reason}\n[llm] falling back to secondary provider...\n")
response = self.secondary.call(*args, **kwargs)
if self._is_empty(response):
raise RuntimeError("Both primary and fallback LLMs returned empty responses")
return response
flowchart TD
A["agent needs a completion"] --> P["primary.call()"]
P -->|"text"| OK["return the answer"]
P -->|"empty string"| R["reason: returned an EMPTY response"]
P -->|"exception"| E["reason: raised Type: message"]
R --> S{"secondary configured?"}
E --> S
S -->|"no"| X["raise"]
S -->|"yes"| F["print [llm] PRIMARY reason, call the secondary"]
F -->|"text"| OK
F -->|"empty"| B["RuntimeError: both returned empty"]
FallbackLLM.call(). Without the empty check, DeepSeek's empty answer is not an exception, so the wrapper never fires and CrewAI fails a few steps later.
Why wrap, not subclass: the docstring explains that CrewAI's LLM is a factory that returns a provider-specific object, so the code wraps BaseLLM and delegates.
Groq leads, DeepSeek catches. DeepSeek has no 8,000 cap but sometimes returns empty or short answers on long prompts; Groq is tight but predictable. LLM_PROVIDER=deepseek flips the order.
The cost: with the wrapper in place, CrewAI's usage counter sees the wrapper, so 04 reports that token usage is not tracked.
07The scripts, and how to run them
Five scripts, a build prompt and a cached bug. Python 3.10 or later; the root README installs only crewai and python-dotenv for this chapter.
File
What it does
How to run it
01_test_analyst_Agent.py
One agent, one task: 5-10 test cases for a login page with remember me, on Groq.
python 01_test_analyst_Agent.py
02_Research_Write_AI_Agent.py
Two agents in sequence: a researcher lists bug categories, a writer turns them into a pre-PR checklist. Runs as soon as the file is imported (no main guard).
python 02_Research_Write_AI_Agent.py
03_Build_QABugTriageCrew.py
In-class skeleton: the business case, a fetch_jira_ticket stub and empty Agent() / Task() calls to fill in. Not runnable as written.
Exercise; 04 is the answer
04_Build_QABugTriageCrew_Prod.py
The finished triage crew: Jira REST fetch with a local cache fallback, three specialists chained with context, per-agent budgets and FallbackLLM.
python 04_Build_QABugTriageCrew_Prod.py
05_QAPipeline_TP_TC_TW_PW.py
Four agents: a Jira analyst using the mcp-atlassian MCP server, then test plan, test cases and Playwright writers that save Markdown files to output/.
Defines run_crew(ticket_id) but never calls it: add the call yourself
Final_QA_Pipeline.md
The RICE-POT build prompt that specifies the chapter 13 Streamlit app.
Read it before chapter 13
bug_cache/VWO-48.txt
A cached copy of bug VWO-48, used by 04 when Jira cannot be reached.
Read by 04 automatically
terminal
cd chapter_12_CrewAI
python3 -m pip install crewai python-dotenv
cp .env.sample .env # add GROQ_API_KEY; DEEPSEEK_API_KEY and the Jira values are optionalpython 01_test_analyst_Agent.py
python 02_Research_Write_AI_Agent.py
BUG_ID=VWO-48 python 04_Build_QABugTriageCrew_Prod.py
LLM_PROVIDER=deepseek python 04_Build_QABugTriageCrew_Prod.py
Environment variables (names only): GROQ_API_KEY, GROQ_MODEL, GROQ_BASE_URL, DEEPSEEK_API_KEY, DEEPSEEK_BASE_URL, DEEPSEEK_MODEL, JIRA_EMAIL, JIRA_API_TOKEN, JIRA_BASE_URL (all in .env.sample), plus LLM_PROVIDER and BUG_ID (read by 04) and JIRA_URL / JIRA_USERNAME (read by 05).
05 hands the Jira analyst tools from an MCP server. It imports MCPServerAdapter from crewai_tools and launches uvx mcp-atlassian, so it also needs the crewai-tools package and uv. It keeps only two of the server's tools, because every tool schema is added to the prompt:
05_QAPipeline_TP_TC_TW_PW.py (lines 368-386)
ALLOWED = {"jira_get_issue", "jira_search"}
filtered_tools = [t for t in mcp_tools if t.name in ALLOWED]
print(f"🔧 Using {len(filtered_tools)} tool(s) to stay under TPM limit: "
f"{[t.name for t in filtered_tools]}\n")
# ── Create Agents (inject MCP tools into Agent 1) ─────────
agents = create_agents(filtered_tools, ticket_id)
# ── Create Tasks (wired sequentially) ─────────────────────
tasks = create_tasks(agents, ticket_id)
# ── Assemble the Crew ─────────────────────────────────────
crew = Crew(
agents=list(agents),
tasks=list(tasks),
process=Process.sequential,
verbose=True,
max_rpm=4, # Rate limit for Groq free tier
)
When 04 cannot reach Jira, it reads the cached bug instead. This is the text the crew triages:
bug_cache/VWO-48.txt (lines 10-26)
Environment: Production, Chrome 120, Windows 11
Severity (Reporter): High
Steps to Reproduce:
1. Add 3+ items to shopping cart (total > $50)
2. Apply discount code "SAVE20" (20% off)
3. Observe the cart total
Actual Result: Cart total shows $0.00 instead of discounted price
Expected Result: Cart total should show original price minus 20%
Additional Info:
- Happens only when cart has 3+ items
- Works fine with 1-2 items
- Started after last Friday's deployment (v2.4.1)
- No errors in browser console
- API response shows correct discounted amount
08Gotchas
Double model prefix.GROQ_MODEL is openai/gpt-oss-120b and the code adds another openai/, so CrewAI receives openai/openai/gpt-oss-120b. The outer prefix picks the OpenAI-compatible client; base_url points it at Groq. Chapter 13's .env.example shows the same id.
02 runs on import. It has no if __name__ == "__main__": guard, so importing it from another file starts the crew.
03 is not runnable.Agent() and Task() with no arguments fail CrewAI's required fields. It is the exercise; 04 is the worked answer.
05 never runs its crew.run_crew() is defined but not called, and the docstring's "fallback to API" is not implemented (requests is imported but unused). Chapter 13 builds that fallback properly.
The 8,000 limit is per request. A higher max_tokens does not buy room; it moves you from truncated output to a 413.
Point Jira at your own site. The Jira URL defaults in 04 and 05 are the course's own workspace. Set JIRA_BASE_URL (04) or JIRA_URL (05) to something like https://your-site.atlassian.net.
The first role string has a typo. 01's role is "QA Enginner". Harmless, but it is part of the prompt the model reads.
Model output is a draft. A triage verdict is a suggestion with quoted evidence, not a decision. The crew is useful because a person can check each claim against the bug report quickly.
DDrills for the chapter
Code drills use the scripts in chapter_12_CrewAI and need a Groq key. Playwright drills target the live demo on the Page tab; every control has a data-testid (turn on Show locator badges to see them).
Separate the what from the shape
In 01 the login-page detail sits in expected_output and description is only "Create 5-10 test cases". Move the page detail into description, keep expected_output for the format, and add a second task that lists negative cases only.
Hint
A second Task needs its own description, expected_output and agent, and must be added to tasks=[...].
Expected result
A two-task crew whose second output lists only negative cases. The rule of thumb: description says what to do, expected_output says what the answer must look like. The wording of the cases changes every run.
Make the hand-off explicit
In 02 add context=[researcher_task] to writing_task and run again. Does the checklist change?
Expected result
The wiring does not change: in a two-task sequential crew the writer already receives the research, so the checklist is built from the same input (its wording still varies run to run). The dependency is now written down, so it survives reordering or a third task added later.
Do the token arithmetic
Agent 3 in 04 starts with a prompt of about 4.8k tokens. What is the largest Groq max_tokens it can request on the free tier?
Expected result
8,000 minus 4,800 = 3,200, which is exactly the Groq budget in sdet_llm = make_llm(3200, 6000).
Read the 413
With max_tokens=8000 the author's run failed with Limit 8000, Requested 10421. How big was the prompt, and why did a bigger max_tokens make things worse?
Expected result
10,421 minus 8,000 = 2,421 prompt tokens. The cap counts prompt and max_tokens together, so a request for 8,000 output tokens can never fit: the silent truncation became a hard error.
Fire the fallback on purpose
Put an invalid GROQ_API_KEY in .env next to a valid DEEPSEEK_API_KEY and run 04.
Hint
Both variables are set, so make_llm builds a FallbackLLM for every agent.
Expected result
Each call prints [llm] PRIMARY raised ... and [llm] falling back to secondary provider..., and the crew completes on DeepSeek. The run ends with the note that token usage is not tracked through the fallback wrapper.
Run without Jira
Remove JIRA_EMAIL and JIRA_API_TOKEN from .env and run 04.
Expected result
[jira] live fetch failed (...) -> using cached VWO-48, then the crew triages the text in bug_cache/VWO-48.txt.
Assert the context wiringPlaywright
Kick off the 04 crew in the demo and assert that all three tasks finished and that task 3 lists both earlier tasks as context.
crew-status contains 3 of 3 tasks, crew-task-3 has data-state="done" and shows context=[triage_task, root_cause_task].
Find the edge of the capPlaywright
Set crew-max-3 to 3200, then to 3201, and assert the text of crew-bar-3 each time.
Expected result
3200: 8,000 of 8,000 and Fits, 0 tokens to spare. 3201: 8,001 of 8,000 and Over the cap by 1.
Break the parachutePlaywright
Set crew-max-3 to 3300, tick crew-ds-key and crew-ds-empty, then kick off.
Expected result
Tasks 1 and 2 finish on Groq. Task 3 hits the 413, falls back, gets an empty answer from DeepSeek and fails: crew-status contains Both primary and fallback LLMs returned empty responses.
SSolutions: the demo spec and the triage crew
The Playwright spec passes against the demo as written. The other tabs are the exact code from the course repo: the triage crew's tasks and wiring, its budgets and fallback, the first agent, and the persona prompt that does most of the triage work.
tests/crewai-demo.spec.ts
import { test, expect } from'@playwright/test';
const URL = 'https://app.thetestingacademy.com/ai/blueprint/learn/crewai.html';
test('the triage crew runs three tasks in order and passes context forward', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('crew-kickoff').click();
awaitexpect(page.getByTestId('crew-status')).toContainText('3 of 3 tasks');
awaitexpect(page.getByTestId('crew-task-3')).toHaveAttribute('data-state', 'done');
awaitexpect(page.getByTestId('crew-task-3')).toContainText('context=[triage_task, root_cause_task]');
awaitexpect(page.getByTestId('crew-out-3')).toContainText('answered by groq');
});
test('agent 3 over the 8,000-token cap stops the crew when there is no fallback', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('crew-max-3').fill('3300');
awaitexpect(page.getByTestId('crew-bar-3')).toContainText('8,100 of 8,000');
await page.getByTestId('crew-kickoff').click();
awaitexpect(page.getByTestId('crew-task-2')).toHaveAttribute('data-state', 'done');
awaitexpect(page.getByTestId('crew-task-3')).toHaveAttribute('data-state', 'fail');
awaitexpect(page.getByTestId('crew-status')).toContainText('413');
});
test('with a DeepSeek key the FallbackLLM finishes the run', async ({ page }) => {
await page.goto(URL);
await page.getByTestId('crew-max-3').fill('3300');
await page.getByTestId('crew-ds-key').check();
await page.getByTestId('crew-kickoff').click();
awaitexpect(page.getByTestId('crew-log')).toContainText('falling back to secondary provider');
awaitexpect(page.getByTestId('crew-out-3')).toContainText('answered by deepseek');
awaitexpect(page.getByTestId('crew-status')).toContainText('3 of 3 tasks');
});
04_Build_QABugTriageCrew_Prod.py (lines 520-637)
# ---------------------------------------------------------------------------
# Task 1: Bug Triage
# ---------------------------------------------------------------------------
triage_task = Task(
description=f"""Triage this bug report from Jira:
{bug_report}
Work through your triage checklist and deliver:
1. SEVERITY: one of S0-S4, with the exact line of the report that justifies it.
2. PRIORITY: one of P0-P4, with the business reasoning (users affected, money at
risk, workaround available or not). If severity and priority differ, explain why.
3. Do you AGREE with the reporter's severity and the Jira priority as filed? If not,
state your override and the reason.
4. CATEGORY: exactly one from your taxonomy.
5. AFFECTED COMPONENT / MODULE and the most likely system layer.
6. BUSINESS IMPACT: what it costs if this ships as-is.
7. MISSING INFORMATION: what you would ask the reporter for, and the assumption you
triaged under in the meantime.
8. CONFIDENCE: High / Medium / Low.""",
expected_output="""A structured triage verdict in markdown with these headed
sections: Severity, Priority, Reporter Disagreement, Category, Affected Component,
Business Impact, Missing Information, Confidence. Each rating must quote the
evidence line from the bug report that drove it.
HARD LIMIT: 400 words. This verdict is handed to two more agents, so every
word you spend costs them room. Use compact bullets, not wide tables.
Finishing every section matters more than depth in any one of them.""",
agent=bug_analyst,
)
# ---------------------------------------------------------------------------
# Task 2: Root Cause (uses triage output as context)
# ---------------------------------------------------------------------------
root_cause_task = Task(
description=f"""Using the triage verdict from the previous task, investigate WHY
this bug happens:
{bug_report}
Deliver:
1. DIFFERENTIAL DIAGNOSIS: what is different between the working case and the
failing case? Start here.
2. PRIMARY ROOT CAUSE HYPOTHESIS with a confidence percentage, the system layer it
lives in, the specific function or config that is suspect, and the evidence line
from the report supporting it.
3. TWO ALTERNATE HYPOTHESES, ranked, each with confidence and evidence.
4. KILL TEST for each of the three: the single fastest check that would disprove it.
5. LAYER VERDICT: frontend, backend, data, infra or third-party, and what the
evidence lets you EXONERATE.
6. INVESTIGATION STEPS in order, naming the exact log file and grep string, the
dashboard and metric, the SQL query, the DevTools network call, and the git
command to diff the suspect release.
7. BLAST RADIUS: what else shares this code path and is probably broken too.
8. WHAT YOU COULD NOT DETERMINE and the one piece of data that would settle it.""",
expected_output="""A root cause analysis in markdown with: Differential Diagnosis,
Primary Hypothesis (with confidence % and kill test), Alternate Hypotheses 2 and 3,
Layer Verdict including what is exonerated, ordered Investigation Steps with named
artifacts, Blast Radius, and Open Questions.
HARD LIMIT: 650 words. This analysis is handed to the test strategist, so
every word you spend costs them room. Use compact bullets, not wide
markdown tables. Finishing every section matters more than depth in any one
of them: a complete short report beats a detailed one that gets cut off.""",
agent=root_cause_agent,
context=[triage_task], # Receives output from triage
)
# ---------------------------------------------------------------------------
# Task 3: Test Recommendation (uses both previous outputs)
# ---------------------------------------------------------------------------
test_task = Task(
description=f"""Using the triage verdict and the root cause analysis from the
previous tasks, design the test strategy for this bug:
{bug_report}
Deliver:
1. THE MISSING TEST: name the single test that should have existed and would have
caught this before production, and explain why the current suite missed it.
2. VERIFICATION TEST: the one test that proves the fix works. State its level.
3. REGRESSION SET: 3-5 cases that stop this specific bug returning, each pinned to
the cheapest level that can catch it (unit / API / E2E) with a one-line reason.
4. BOUNDARY AND EDGE CASES: apply boundary value analysis to the failing threshold
in this report, plus the negative and state-transition cases.
5. PLAYWRIGHT TYPESCRIPT: runnable code for the E2E tests worth automating,
following your standards (web-first assertions, no hard waits, role/testid
locators, API-seeded state, route interception where it isolates the layer).
6. WHAT TO KEEP MANUAL and why.
7. EXIT CRITERIA: what must be green before this ticket is closed.""",
expected_output="""A test strategy in markdown: the Missing Test, a table of tests
(ID, Level, Type, What it proves, Priority), runnable Playwright TypeScript code
blocks for the E2E cases, a Manual-only section with justification, and Exit
Criteria.
HARD LIMIT: 800 words including the code. You are the last agent and have
the least room, so budget it: keep the tables narrow, write ONE complete
Playwright spec file rather than several half-finished ones, and make sure
you reach Exit Criteria. A complete short plan beats a truncated long one.""",
agent=test_recommender,
context=[triage_task, root_cause_task], # Uses both outputs
)
# ---------------------------------------------------------------------------
# Step 3 + 4 - Assemble the Crew and kick it off
# ---------------------------------------------------------------------------
crew = Crew(
agents=[bug_analyst, root_cause_agent, test_recommender],
tasks=[triage_task, root_cause_task, test_task],
process=Process.sequential,
verbose=True,
)
04_Build_QABugTriageCrew_Prod.py (lines 50-171)
class FallbackLLM(BaseLLM):
"""Try the primary LLM, fall back to the secondary when it fails.
Why this exists: a free-tier key WILL let you down mid-demo. Rate limit
(429), request too large (413), expired key (401), provider outage (5xx) -
any of them kills the crew halfway through and you lose the whole run.
This wrapper catches the failure on that one call and re-issues it against
the backup provider, so the crew keeps going instead of crashing.
CrewAI's `LLM` is a factory (it returns a provider-specific object), so you
cannot subclass it. Wrap it instead and delegate.
"""
primary: Any = None
secondary: Any = None
@staticmethod
def _is_empty(response):
"""An empty string is a failure too, not a valid answer.
DeepSeek occasionally returns empty content on long generations. That
is NOT an exception, so a try/except alone never catches it: CrewAI
just raises "Invalid response from LLM call - None or empty" a few
steps later and the whole crew dies. Treat it as a failure here.
"""
return response is None or (isinstance(response, str) and not response.strip())
def call(self, *args, **kwargs):
try:
response = self.primary.call(*args, **kwargs)
if not self._is_empty(response):
return response
reason = "returned an EMPTY response"
except Exception as exc:
reason = f"raised {type(exc).__name__}: {str(exc)[:150]}"
if self.secondary is None:
raise
if self.secondary is None:
raise RuntimeError(f"Primary LLM {reason} and no fallback is configured")
print(f"\n[llm] PRIMARY {reason}\n[llm] falling back to secondary provider...\n")
response = self.secondary.call(*args, **kwargs)
if self._is_empty(response):
raise RuntimeError("Both primary and fallback LLMs returned empty responses")
return response
# delegate capability checks to the primary
def supports_function_calling(self):
return self.primary.supports_function_calling()
def supports_stop_words(self):
return self.primary.supports_stop_words()
def get_context_window_size(self):
return self.primary.get_context_window_size()
def make_groq_llm(max_tokens):
return LLM(
model=f"openai/{os.getenv('GROQ_MODEL')}",
api_key=os.getenv("GROQ_API_KEY"),
base_url=os.getenv("GROQ_BASE_URL"),
max_tokens=max_tokens,
temperature=0.3, # triage should be consistent, not creative
)
def make_deepseek_llm(max_tokens=DEEPSEEK_MAX_TOKENS):
return LLM(
model=f"deepseek/{os.getenv('DEEPSEEK_MODEL', 'deepseek-v4-flash')}",
api_key=os.getenv("DEEPSEEK_API_KEY"),
base_url=os.getenv("DEEPSEEK_BASE_URL", "https://api.deepseek.com"),
max_tokens=max_tokens,
temperature=0.3,
)
# Which provider leads. Groq is fast, free and reliable here, so it leads. Its
# catch is the 8000-token request cap above, which is why every task also
# carries an explicit length limit: an agent that writes past its budget gets
# cut off mid-sentence. DeepSeek has no such cap but intermittently returns
# empty or short responses under long prompts, so it rides shotgun instead.
# Set LLM_PROVIDER=deepseek in .env to flip the order.
PRIMARY_PROVIDER = os.getenv("LLM_PROVIDER", "groq").lower()
def make_llm(groq_tokens, deepseek_tokens):
"""Build one LLM per agent, with the other provider wired in as a parachute.
Each provider gets its own budget because their limits differ: Groq must
stay under GROQ_TPM_LIMIT (prompt + output), DeepSeek can breathe.
"""
has_groq = bool(os.getenv("GROQ_API_KEY"))
has_deepseek = bool(os.getenv("DEEPSEEK_API_KEY"))
groq = make_groq_llm(groq_tokens) if has_groq else None
deepseek = make_deepseek_llm(deepseek_tokens) if has_deepseek else None
if PRIMARY_PROVIDER == "groq":
primary, secondary = groq, deepseek
else:
primary, secondary = deepseek, groq
primary = primary or secondary
if primary is None:
raise RuntimeError("Set GROQ_API_KEY and/or DEEPSEEK_API_KEY in .env")
if secondary is primary:
secondary = None
if secondary is None:
return primary
return FallbackLLM(
model=f"{PRIMARY_PROVIDER}-with-fallback", primary=primary, secondary=secondary
)
# Budgets: (groq, deepseek). The Groq numbers keep prompt + output under
# GROQ_TPM_LIMIT; the DeepSeek numbers are what these reports actually need.
triage_llm = make_llm(4000, 4000) # smallest prompt, runs first
rca_llm = make_llm(4000, 5000) # prompt also carries task 1's output
sdet_llm = make_llm(3200, 6000) # prompt also carries task 1 + task 2 output
01_test_analyst_Agent.py
# Test Ananlyst Agent# # a senior QA with 15 years (JIRA MD)# of experience. Based on the feature, # it will just analyze the requirement# and suggest a 5-10 testcases(p0 testcases).from crewai import Agent,Task, Crew
from crewai import LLM
from dotenv import load_dotenv
import os
# By Default crew AI actually the brain which# OpenAI - GROQ API Key# Step 0 - Set up the Brain# Step 1. - Define the Agent (identity)# Step 2. - Give the Task to the Agent# Step 3. Add them to the Crew# Step 4. Kick Off Agent.# Prompt vs Skill vs AI Agent# We need use the GROQ gpt-oss-120b model# Step 0 - Set up the Brain (Groq LLM)load_dotenv() # reads the .env file in this folder# Groq exposes an OpenAI-compatible API, so we use the "openai/" provider# prefix with a custom base_url pointing at Groq.# GROQ_MODEL in .env = openai/gpt-oss-120b (exact ID from groq.com console)
groq_llm = LLM(
model=f"openai/{os.getenv('GROQ_MODEL')}",
api_key=os.getenv("GROQ_API_KEY"),
base_url=os.getenv("GROQ_BASE_URL"),
)
# Step 1. - Define the Agent (identity)
qa_agent = Agent(
role="QA Enginner",
goal="Analyse the feature or the requirements, and create 5-10 test cases out of it.",
backstory="You are a senior QA engineer with 15 years of experience in test planning and testcases creation",
llm = groq_llm,
verbose=True
)
# Step 2 - Give the Task to the Agent
test_case_task = Task(
description="Create 5-10 test cases",
expected_output="A numbered list of 5-10 test cases with brief descriptions for a app.vwo.com Login page with the username, password and submit button with remember me functionality",
agent=qa_agent
)
# Step 3. Add them to the Crew
crew = Crew(
agents=[qa_agent],
tasks=[test_case_task],
verbose=True
)
# Step 4. kickOffif __name__ == "__main__":
result = crew.kickoff()
print(result)
04_Build_QABugTriageCrew_Prod.py (lines 262-351)
bug_analyst = Agent(
role="Senior Bug Triage Analyst and Defect Classification Specialist",
goal=(
"Read the raw bug report and return one defensible triage verdict for it: "
"exactly one SEVERITY, exactly one PRIORITY, and exactly one CATEGORY. "
"Every one of the three must be justified by quoting the specific line of "
"the bug report that drove the decision. Never guess numbers that are not "
"in the report, and always call out what information is missing that would "
"change your verdict."
),
backstory="""You are a veteran QA engineer with 15+ years of experience. You have
personally triaged well over 20,000 defects across e-commerce checkouts, payment
gateways, and B2B SaaS dashboards, and you have run the daily bug triage meeting
for teams of 30+ engineers. Triage is not paperwork to you. A wrong severity either
wakes an on-call engineer at 2 AM for a typo, or lets a money-losing bug sit in the
backlog for three sprints. You take it seriously.
## THE ONE RULE PEOPLE ALWAYS GET WRONG
Severity and Priority are NOT the same thing, and you never collapse them:
- SEVERITY = technical impact on the system. How badly is it broken? Set by QA.
This is objective and does not care about the release calendar.
- PRIORITY = business urgency. How soon must it be fixed relative to everything
else? Driven by users affected, money at risk, and whether a workaround exists.
A typo in the company name on the homepage is LOW severity but HIGH priority.
A crash in an admin tool used by two internal people is HIGH severity but LOW
priority. You explain this distinction whenever the two ratings differ.
## SEVERITY SCALE (technical impact)
- S0 Blocker : System down, data loss or data corruption, security breach,
money computed wrong, complete checkout/login failure. Nothing
can proceed.
- S1 Critical : Major feature completely broken with NO workaround. Core user
journey blocked for a large segment.
- S2 Major : Feature impaired or partially wrong, but a reasonable workaround
exists. Journey is painful, not blocked.
- S3 Minor : Cosmetic, layout, wording, or minor inconvenience. Function is
intact.
- S4 Trivial : Typo, alignment nitpick, enhancement request, "would be nice".
## PRIORITY SCALE (business urgency and fix order)
- P0 : Fix now, hotfix today, stop other work. Production + revenue or security.
- P1 : Fix in the current sprint, before the next release ships.
- P2 : Schedule into the next sprint. Normal backlog flow.
- P3 : Fix when the area is touched next. Opportunistic.
- P4 : Backlog / icebox. May legitimately never be fixed.
## CATEGORY TAXONOMY (pick exactly one, the closest root fit)
Functional Logic, Data and Calculation, UI/UX and Layout, Performance,
Security, API and Integration, Compatibility (browser/OS/device),
Configuration and Deployment, Regression, Usability and Content.
## HOW YOU ACTUALLY DECIDE (your triage checklist)
1. Environment first. Production outweighs staging outweighs local dev.
2. Blast radius. All users, one segment, or one account? If the report does not
say, you say so instead of inventing a percentage.
3. Money and data path. Anything touching price, total, discount, tax, payment,
or stored customer data starts at S0/S1 by default, even when the symptom
looks cosmetic. A wrong total is never "just a display issue".
4. Workaround. If a real workaround exists, severity drops one level. If it does
not, priority rises one level.
5. Reproducibility. Consistent and 100% reproducible beats intermittent. An
intermittent bug is NOT automatically lower severity, it is often harder and
you say so.
6. Regression signal. "It worked before the last deploy" is a strong escalator.
A freshly shipped regression is more urgent than an old known defect.
7. Layer isolation. If the API response is correct but the UI shows something
else, the defect is frontend/presentation, not backend. If the UI input is
correct but the stored value is wrong, it is backend/data. Say which layer.
8. Security and compliance. Any auth bypass, data exposure, PII leak, or
injection vector is S0/P0 regardless of how few users hit it.
## RULES OF YOUR DESK
- You NEVER inflate severity to get attention, and you never deflate it to
protect a release date.
- The reporter's severity is an input, not an instruction. If the evidence does
not support it, you override it and state plainly why you disagreed.
- You never invent facts. If user impact, error logs, or affected version are
missing, you list them under "Missing Information" and state the assumption
you triaged under.
- You state a confidence level (High / Medium / Low) on your verdict. Low
confidence with a clear list of what you need beats a confident guess.
- You write for two audiences at once: an engineer who needs the technical
signal, and a product manager who needs to know if it can wait.""",
llm=triage_llm,
verbose=True,
allow_delegation=False, # This agent handles its own work
)
A backstory is where QA judgement lives: severity versus priority, the money-path rule, layer isolation and "never invent facts". The model applies these rules; it does not replace the person who wrote them.