Chapter 12
Chapter 12 . Agent frameworks . CrewAI

CrewAI: a QA team in code

CrewAI writes the person behind a prompt as code: an Agent has a role, a goal and a backstory, a Task says what to deliver, and a Crew runs the tasks in order. The chapter goes from one test-analyst agent to a three-agent crew that triages a Jira bug, finds the likely root cause and plans the tests, and shows how to keep it inside a free Groq key's 8,000-token request limit.

5
Scripts
01 to 05. 03 is the in-class skeleton; 04 is its finished version.
3
Specialists in the triage crew
Triage, root cause and test strategy, chained with context.
8,000
Groq tokens per request
Free tier: prompt plus max_tokens together.
3,200
Groq budget for agent 3
8,000 minus its ~4.8k prompt. The last agent has the least room.

01What CrewAI gives a tester

Up to now a prompt was something you pasted into a chat. CrewAI lets you write the expert behind the prompt once, as code, and reuse it on every ticket, in a pipeline or on a schedule. Five names cover everything this chapter uses:

ConceptIn CrewAIIn the triage crew (04)
LLMLLM(model, api_key, base_url, max_tokens)One per agent, each with its own token budget
AgentAgent(role, goal, backstory, llm): who does the workA senior triage analyst whose backstory carries the S0-S4 and P0-P4 rulebook
TaskTask(description, expected_output, agent, context): what to deliver, and in what shape"Triage this bug report", with a 400-word HARD LIMIT
CrewCrew(agents, tasks, process): the team and the orderThree specialists, Process.sequential
kickoff()Runs the crew and returns a result objectresult.tasks_output holds all three reports
LLM: the brain model, api_key, base_url, max_tokens Process.sequential task 1, then 2, then 3 outputs flow on as context crew.kickoff() result.raw: last task result.tasks_output: all Agent: who role goal backstory llm Task: what description expected_output agent context=[earlier tasks] Crew: the team agents=[...] tasks=[...] process=... verbose=True llm= agent= tasks= process= kickoff()
The four objects every crew in this chapter uses. The agent is the who, the task is the what, and the crew decides the order.
Why a triage crew? The docstring of 04_Build_QABugTriageCrew_Prod.py makes the business case: a daily triage meeting of 30 minutes for 30 people is about 300 hours a month. A crew that pre-triages every ticket before the meeting buys most of that back, and because each verdict is asked to quote the line of the bug report that drove it, a person can check it quickly.

02Your first agent in four steps

01_test_analyst_Agent.py is the smallest useful crew: set up the brain, define the agent, give it a task, add both to a crew and kick it off. The comments in the file name the steps the same way.

01_test_analyst_Agent.py (lines 29-67)
# Step 0 - Set up the Brain (Groq LLM)
load_dotenv()  # reads the .env file in this folder

# Groq exposes an OpenAI-compatible API, so we use the "openai/" provider
# prefix with a custom base_url pointing at Groq.
# GROQ_MODEL in .env = openai/gpt-oss-120b (exact ID from groq.com console)
groq_llm = LLM(
    model=f"openai/{os.getenv('GROQ_MODEL')}",
    api_key=os.getenv("GROQ_API_KEY"),
    base_url=os.getenv("GROQ_BASE_URL"),
)

# Step 1. - Define the Agent (identity)
qa_agent = Agent(
    role="QA Enginner",
    goal="Analyse the feature or the requirements, and create 5-10 test cases out of it.",
    backstory="You are a senior QA engineer with 15 years of experience in test planning and testcases creation",
    llm = groq_llm,
    verbose=True
)

# Step 2 - Give the Task to the Agent
test_case_task = Task(
    description="Create 5-10 test cases",
    expected_output="A numbered list of 5-10 test cases with brief descriptions for a app.vwo.com Login page with the username, password and submit button with remember me functionality",
    agent=qa_agent
)

# Step 3. Add them to the Crew
crew = Crew(
    agents=[qa_agent],
    tasks=[test_case_task],
    verbose=True
)

# Step 4. kickOff
if __name__ == "__main__":
    result = crew.kickoff()
    print(result)
  • The brain is Groq. CrewAI defaults to OpenAI. Groq speaks the same API, so the code sets base_url to Groq and prefixes the model with openai/.
  • The agent is the who. role, goal and backstory become the persona the model is told to play. One agent can take many tasks.
  • The task is the what. Here description is only "Create 5-10 test cases" and the page details sit in expected_output. Drill 1 moves them where they belong.
flowchart LR
  ENV[".env: GROQ_API_KEY, GROQ_MODEL, GROQ_BASE_URL"] --> LLM["LLM: openai/ + model, base_url = Groq"]
  LLM --> AG["Agent: role, goal, backstory"]
  AG --> TASK["Task: description, expected_output"]
  TASK --> CREW["Crew: agents, tasks"]
  CREW -->|"kickoff()"| OUT["5-10 numbered test cases"]
Script 01 from configuration to output. The root README lists what a run covers: valid login, empty fields, password masking, remember-me persistence, injection protection and accessibility.

03Live demo: kick off a crew

Pick a crew and press Kick off, or use Run next task to go one task at a time. Each task card shows the agent and the exact context wiring from the script, and the panel on the right shows what the next task receives, including the earlier outputs. For the 04 crew, change a Groq budget or a key and watch the FallbackLLM rules from the script decide what happens.

Crew runner: sequential tasks, context and the token squeezeNo API key needed
Agents and tasks, as defined in the script
    data-testid=crew-pickdata-testid=crew-kickoffdata-testid=crew-stepdata-testid=crew-resetdata-testid=crew-task-1data-testid=crew-statusdata-testid=crew-log
    Event log
      LLM settings for the 04 crew (make_llm)
      Groq max_tokens per agent
      data-testid=crew-providerdata-testid=crew-groq-keydata-testid=crew-ds-keydata-testid=crew-429data-testid=crew-ds-emptydata-testid=crew-max-3
      Groq request size: prompt + max_tokens against the 8,000 cap
      Inputs to the current task
      
            Task outputs (illustrative text, not a recorded run)
            
      What is real here: the agents, tasks, context lists, budgets, the 8,000 cap, the approximate prompt sizes (from the repo's learnings note) and the fallback logic. What is not: the task outputs. The repo keeps no recorded run, so they are illustrative text written for this page. A real run words, rates and scores things differently every time.

      04Sequential crews and context

      With Process.sequential the tasks run in list order. What each task receives from the earlier ones depends on context:

      • No context set (script 02): the sequential process passes the earlier output along, so the writer sees the research without any wiring.
      • context=[...] set (script 04): the task receives exactly the outputs of the tasks you list. The root cause task gets the triage verdict; the test strategy task gets both.
      04_Build_QABugTriageCrew_Prod.py (lines 595-637)
      test_task = Task(
          description=f"""Using the triage verdict and the root cause analysis from the
      previous tasks, design the test strategy for this bug:
      
      {bug_report}
      
      Deliver:
      1. THE MISSING TEST: name the single test that should have existed and would have
         caught this before production, and explain why the current suite missed it.
      2. VERIFICATION TEST: the one test that proves the fix works. State its level.
      3. REGRESSION SET: 3-5 cases that stop this specific bug returning, each pinned to
         the cheapest level that can catch it (unit / API / E2E) with a one-line reason.
      4. BOUNDARY AND EDGE CASES: apply boundary value analysis to the failing threshold
         in this report, plus the negative and state-transition cases.
      5. PLAYWRIGHT TYPESCRIPT: runnable code for the E2E tests worth automating,
         following your standards (web-first assertions, no hard waits, role/testid
         locators, API-seeded state, route interception where it isolates the layer).
      6. WHAT TO KEEP MANUAL and why.
      7. EXIT CRITERIA: what must be green before this ticket is closed.""",
      
          expected_output="""A test strategy in markdown: the Missing Test, a table of tests
          (ID, Level, Type, What it proves, Priority), runnable Playwright TypeScript code
          blocks for the E2E cases, a Manual-only section with justification, and Exit
          Criteria.
      
          HARD LIMIT: 800 words including the code. You are the last agent and have
          the least room, so budget it: keep the tables narrow, write ONE complete
          Playwright spec file rather than several half-finished ones, and make sure
          you reach Exit Criteria. A complete short plan beats a truncated long one.""",
          agent=test_recommender,
          context=[triage_task, root_cause_task],  # Uses both outputs
      )
      
      
      # ---------------------------------------------------------------------------
      # Step 3 + 4 - Assemble the Crew and kick it off
      # ---------------------------------------------------------------------------
      crew = Crew(
          agents=[bug_analyst, root_cause_agent, test_recommender],
          tasks=[triage_task, root_cause_task, test_task],
          process=Process.sequential,
          verbose=True,
      )
      flowchart TD
        ID["BUG_ID, default VWO-48"] --> J["requests.get: /rest/api/3/issue/BUG_ID"]
        J -->|"200"| ADF["_adf_to_text(): flat bug text"]
        J -->|"any error"| C["bug_cache/VWO-48.txt"]
        ADF --> T1
        C --> T1
        T1["Task 1: triage_task<br/>Bug Triage Analyst"] -->|"context"| T2["Task 2: root_cause_task<br/>Root Cause Investigator"]
        T1 -->|"context"| T3["Task 3: test_task<br/>Test Strategy Advisor"]
        T2 -->|"context"| T3
        T3 --> R["result.tasks_output: three reports"]
      The 04 crew. Jira first, the local cache when Jira fails, then three specialists, each reading the verdicts before it.
      Read every agent, not just the last. Printing crew.kickoff() shows only the final task's output. 04 loops over result.tasks_output to print all three reports, labelled by agent.

      05The token squeeze: budget every agent

      On Groq's free tier one request may not exceed 8,000 tokens, and that counts the prompt and max_tokens together. In a sequential crew the prompt grows at every step, because each task carries the earlier outputs. So the agent with the most to write has the least room to write it. The repo's learnings note measured the shape:

      Agent (LLM)Prompt size, approx.Room under 8,000Groq max_tokensDeepSeek max_tokensWord limit
      Bug Triage Analyst (triage_llm)~2.3k~5.7k40004000400
      Root Cause Investigator (rca_llm)~3.0k~5.0k40005000650
      Test Strategy Advisor (sdet_llm)~4.8k~3.2k32006000800
      04_Build_QABugTriageCrew_Prod.py (lines 137-171)
      def make_llm(groq_tokens, deepseek_tokens):
          """Build one LLM per agent, with the other provider wired in as a parachute.
      
          Each provider gets its own budget because their limits differ: Groq must
          stay under GROQ_TPM_LIMIT (prompt + output), DeepSeek can breathe.
          """
          has_groq = bool(os.getenv("GROQ_API_KEY"))
          has_deepseek = bool(os.getenv("DEEPSEEK_API_KEY"))
      
          groq = make_groq_llm(groq_tokens) if has_groq else None
          deepseek = make_deepseek_llm(deepseek_tokens) if has_deepseek else None
      
          if PRIMARY_PROVIDER == "groq":
              primary, secondary = groq, deepseek
          else:
              primary, secondary = deepseek, groq
      
          primary = primary or secondary
          if primary is None:
              raise RuntimeError("Set GROQ_API_KEY and/or DEEPSEEK_API_KEY in .env")
          if secondary is primary:
              secondary = None
      
          if secondary is None:
              return primary
          return FallbackLLM(
              model=f"{PRIMARY_PROVIDER}-with-fallback", primary=primary, secondary=secondary
          )
      
      
      # Budgets: (groq, deepseek). The Groq numbers keep prompt + output under
      # GROQ_TPM_LIMIT; the DeepSeek numbers are what these reports actually need.
      triage_llm = make_llm(4000, 4000)  # smallest prompt, runs first
      rca_llm = make_llm(4000, 5000)     # prompt also carries task 1's output
      sdet_llm = make_llm(3200, 6000)    # prompt also carries task 1 + task 2 output

      Two fixes work together. Each agent gets its own LLM with a budget that fits its position in the chain (make_llm(groq_tokens, deepseek_tokens)), and every expected_output ends with a HARD LIMIT word count and the reason for it ("every word you spend costs them room"), because a wordy first agent steals room from the last one.

      How the author found it: raising max_tokens to 8000 turned silent mid-sentence truncation into a loud error, Limit 8000, Requested 10421. The note also records that every model on the key had the same 8,000 limit, so switching models was not an escape: the budget had to be engineered.

      06A fallback that also catches empty answers

      A free key will fail mid-demo: a 429 rate limit, a 413 request too large, an expired key, an outage. FallbackLLM wraps two providers and re-issues the failed call on the second one, so the crew keeps going. The detail that matters is the first branch: an empty string counts as a failure too.

      04_Build_QABugTriageCrew_Prod.py (lines 77-95)
          def call(self, *args, **kwargs):
              try:
                  response = self.primary.call(*args, **kwargs)
                  if not self._is_empty(response):
                      return response
                  reason = "returned an EMPTY response"
              except Exception as exc:
                  reason = f"raised {type(exc).__name__}: {str(exc)[:150]}"
                  if self.secondary is None:
                      raise
      
              if self.secondary is None:
                  raise RuntimeError(f"Primary LLM {reason} and no fallback is configured")
      
              print(f"\n[llm] PRIMARY {reason}\n[llm] falling back to secondary provider...\n")
              response = self.secondary.call(*args, **kwargs)
              if self._is_empty(response):
                  raise RuntimeError("Both primary and fallback LLMs returned empty responses")
              return response
      flowchart TD
        A["agent needs a completion"] --> P["primary.call()"]
        P -->|"text"| OK["return the answer"]
        P -->|"empty string"| R["reason: returned an EMPTY response"]
        P -->|"exception"| E["reason: raised Type: message"]
        R --> S{"secondary configured?"}
        E --> S
        S -->|"no"| X["raise"]
        S -->|"yes"| F["print [llm] PRIMARY reason, call the secondary"]
        F -->|"text"| OK
        F -->|"empty"| B["RuntimeError: both returned empty"]
      FallbackLLM.call(). Without the empty check, DeepSeek's empty answer is not an exception, so the wrapper never fires and CrewAI fails a few steps later.
      • Why wrap, not subclass: the docstring explains that CrewAI's LLM is a factory that returns a provider-specific object, so the code wraps BaseLLM and delegates.
      • Groq leads, DeepSeek catches. DeepSeek has no 8,000 cap but sometimes returns empty or short answers on long prompts; Groq is tight but predictable. LLM_PROVIDER=deepseek flips the order.
      • The cost: with the wrapper in place, CrewAI's usage counter sees the wrapper, so 04 reports that token usage is not tracked.

      07The scripts, and how to run them

      Five scripts, a build prompt and a cached bug. Python 3.10 or later; the root README installs only crewai and python-dotenv for this chapter.

      FileWhat it doesHow to run it
      01_test_analyst_Agent.pyOne agent, one task: 5-10 test cases for a login page with remember me, on Groq.python 01_test_analyst_Agent.py
      02_Research_Write_AI_Agent.pyTwo agents in sequence: a researcher lists bug categories, a writer turns them into a pre-PR checklist. Runs as soon as the file is imported (no main guard).python 02_Research_Write_AI_Agent.py
      03_Build_QABugTriageCrew.pyIn-class skeleton: the business case, a fetch_jira_ticket stub and empty Agent() / Task() calls to fill in. Not runnable as written.Exercise; 04 is the answer
      04_Build_QABugTriageCrew_Prod.pyThe finished triage crew: Jira REST fetch with a local cache fallback, three specialists chained with context, per-agent budgets and FallbackLLM.python 04_Build_QABugTriageCrew_Prod.py
      05_QAPipeline_TP_TC_TW_PW.pyFour agents: a Jira analyst using the mcp-atlassian MCP server, then test plan, test cases and Playwright writers that save Markdown files to output/.Defines run_crew(ticket_id) but never calls it: add the call yourself
      Final_QA_Pipeline.mdThe RICE-POT build prompt that specifies the chapter 13 Streamlit app.Read it before chapter 13
      bug_cache/VWO-48.txtA cached copy of bug VWO-48, used by 04 when Jira cannot be reached.Read by 04 automatically
      terminal
      cd chapter_12_CrewAI
      python3 -m pip install crewai python-dotenv
      cp .env.sample .env          # add GROQ_API_KEY; DEEPSEEK_API_KEY and the Jira values are optional
      python 01_test_analyst_Agent.py
      python 02_Research_Write_AI_Agent.py
      BUG_ID=VWO-48 python 04_Build_QABugTriageCrew_Prod.py
      LLM_PROVIDER=deepseek python 04_Build_QABugTriageCrew_Prod.py

      Environment variables (names only): GROQ_API_KEY, GROQ_MODEL, GROQ_BASE_URL, DEEPSEEK_API_KEY, DEEPSEEK_BASE_URL, DEEPSEEK_MODEL, JIRA_EMAIL, JIRA_API_TOKEN, JIRA_BASE_URL (all in .env.sample), plus LLM_PROVIDER and BUG_ID (read by 04) and JIRA_URL / JIRA_USERNAME (read by 05).

      05 hands the Jira analyst tools from an MCP server. It imports MCPServerAdapter from crewai_tools and launches uvx mcp-atlassian, so it also needs the crewai-tools package and uv. It keeps only two of the server's tools, because every tool schema is added to the prompt:

      05_QAPipeline_TP_TC_TW_PW.py (lines 368-386)
              ALLOWED = {"jira_get_issue", "jira_search"}
              filtered_tools = [t for t in mcp_tools if t.name in ALLOWED]
              print(f"🔧 Using {len(filtered_tools)} tool(s) to stay under TPM limit: "
                    f"{[t.name for t in filtered_tools]}\n")
      
              # ── Create Agents (inject MCP tools into Agent 1) ─────────
              agents = create_agents(filtered_tools, ticket_id)
      
              # ── Create Tasks (wired sequentially) ─────────────────────
              tasks = create_tasks(agents, ticket_id)
      
              # ── Assemble the Crew ─────────────────────────────────────
              crew = Crew(
                  agents=list(agents),
                  tasks=list(tasks),
                  process=Process.sequential,
                  verbose=True,
                  max_rpm=4,  # Rate limit for Groq free tier
              )

      When 04 cannot reach Jira, it reads the cached bug instead. This is the text the crew triages:

      bug_cache/VWO-48.txt (lines 10-26)
      Environment: Production, Chrome 120, Windows 11
      Severity (Reporter): High
      
      Steps to Reproduce:
      1. Add 3+ items to shopping cart (total > $50)
      2. Apply discount code "SAVE20" (20% off)
      3. Observe the cart total
      
      Actual Result: Cart total shows $0.00 instead of discounted price
      Expected Result: Cart total should show original price minus 20%
      
      Additional Info:
      - Happens only when cart has 3+ items
      - Works fine with 1-2 items
      - Started after last Friday's deployment (v2.4.1)
      - No errors in browser console
      - API response shows correct discounted amount

      08Gotchas

      • Double model prefix. GROQ_MODEL is openai/gpt-oss-120b and the code adds another openai/, so CrewAI receives openai/openai/gpt-oss-120b. The outer prefix picks the OpenAI-compatible client; base_url points it at Groq. Chapter 13's .env.example shows the same id.
      • 02 runs on import. It has no if __name__ == "__main__": guard, so importing it from another file starts the crew.
      • 03 is not runnable. Agent() and Task() with no arguments fail CrewAI's required fields. It is the exercise; 04 is the worked answer.
      • 05 never runs its crew. run_crew() is defined but not called, and the docstring's "fallback to API" is not implemented (requests is imported but unused). Chapter 13 builds that fallback properly.
      • The 8,000 limit is per request. A higher max_tokens does not buy room; it moves you from truncated output to a 413.
      • Point Jira at your own site. The Jira URL defaults in 04 and 05 are the course's own workspace. Set JIRA_BASE_URL (04) or JIRA_URL (05) to something like https://your-site.atlassian.net.
      • The first role string has a typo. 01's role is "QA Enginner". Harmless, but it is part of the prompt the model reads.
      • Model output is a draft. A triage verdict is a suggestion with quoted evidence, not a decision. The crew is useful because a person can check each claim against the bug report quickly.