The Testing Academy · Class Notes Sunday, 23 August (IST)
Live class · study guide

Finishing Python, and building your first CrewAI agent

The last of Python (collections, the main guard, files, .env secrets and pytest) and then the real thing: what CrewAI is, the anatomy of an agent, and a first QA-analyst agent built in five steps. Now carries the full pytest and CrewAI reference too: fixtures, parametrize, markers and flags, every agent, task and crew field worth setting, the model-provider table, and the fix for the provider error that took up half the live build.

By Pramod Dutta, The Testing Academy. Study notes from the live AI Tester Blueprint 3x class, rebuilt from the session recording, with the pytest and CrewAI reference material folded in. Every snippet was executed before publishing: Python and pytest 9.1 on CPython 3.12, and the CrewAI code against a real install of crewai 1.15.17, where every field listed was confirmed on the actual class and only the paid model call was left unmade. Two things the class fought with live are resolved here with the verified fix.

01

Closing out Python

The last Python session before the agents work. Everything here exists because CrewAI and the evaluation framework need it, not for its own sake.

The collections module is the standard library's set of upgraded containers. The honest framing given in class: these are a plus version of the built-in list, tuple, set and dictionary, and most of them you will rarely use in automation. Recognise them, do not memorise them.

Python
from collections import namedtuple, Counter, defaultdict

# a tuple whose positions have names
Info = namedtuple("Info", ["name", "age", "is_married", "grade"])
t = Info(name="Pramod", age=34, is_married=True, grade=9.8)
print(t)         # Info(name='Pramod', age=34, is_married=True, grade=9.8)
print(t.name)    # 'Pramod'   by name
print(t[0])      # 'Pramod'   by index, both still work

The problem it solves is real. A plain tuple ("Pramod", 34, True, 9.8) gives you no way to know what True means or whether 9.8 is a grade or a version. A named tuple keeps tuple behaviour and adds labels.

Python
c = Counter("aaaaabbbbccc")
print(c.most_common(3))   # [('a', 5), ('b', 4), ('c', 3)]
print(c.total())          # 12

d = defaultdict(int)
for w in "test test build".split():
    d[w] += 1
print(dict(d))            # {'test': 2, 'build': 1}
print(d["never-seen"])    # 0, where a plain dict raises KeyError

Counter is the interview one. Counting occurrences in a string is a common coding-round question, and the standard library already does it in one line. defaultdict earns its place by removing the "check if the key exists first" dance.

02

The main guard

Python does have an entry point, and it looks like nothing else:

Python
def f1(): print("f1")
def f2(): print("f2")

if __name__ == "__main__":
    print("main guard ran")
    f1()
    f2()

Run the file directly and the block executes. Import the same file from somewhere else and it does not. That is the whole point: it lets a file be both a runnable script and an importable module, which is exactly what you want when a file holds an agent you sometimes run and sometimes import.

The name is not decoration and it is not arbitrary either. Python sets __name__ to "__main__" in whichever file you ran, and to the module's own name in anything imported. So the condition is literally asking "am I the file that was run?" Verified both ways: run directly and the guard fires, import it and it stays silent.

03

Files, paths, and the thing that goes wrong

The class hit the same problem three separate times, so it is worth naming: a relative path resolves against where you ran the command, not against the file that contains the code.

Python
import os

file_path = os.path.join(os.getcwd(), "testdata.txt")

with open(file_path, "r") as f:
    print(f.read())

with open(...) and the older f = open(...) then f.close() do the same thing. The with form closes the file for you even if something throws in the middle, which is why it is the one to use.

Wrap it when the file might genuinely be missing:

Python
try:
    with open(file_path, "r") as f:
        print(f.read())
except FileNotFoundError:
    print("file not found")

CSV works the same way through the csv module, and pandas does it in fewer lines if you already have it.

04

Secrets, and keeping them out of git

An API key does not belong in your source. It goes in a .env file:

Code
GROQ_API_KEY=your_key_here
DB_USER=admin
DB_PASSWORD=super_secret
Python
from dotenv import load_dotenv
import os

load_dotenv()                       # reads .env into the environment
print(os.getenv("DB_USER"))         # 'admin'
print(os.getenv("NOT_SET"))         # None, not an exception

Install it with pip install python-dotenv. Note the mismatch that catches people: the package is python-dotenv, the import is dotenv.

.env goes in .gitignore, without exception. This is the single most common way a working key ends up public. os.getenv returning None rather than raising is worth knowing too: a missing key does not fail loudly, it fails later and more confusingly, so check for None if the value matters.

05

pytest, and how a test is written

pytest is here because the evaluation framework later in the course is built on it. If you know TestNG, you know this: a test is a function whose name starts with test_, and the assertion is a plain assert.

Terminal
pip install pytest
Terminal
pytest                            # everything it can find
pytest test_login.py              # one file
pytest test_login.py::test_valid  # one test
pytest -q                         # quiet, one character per test
pytest -v                         # verbose, one line per test with the result

Discovery is convention, not configuration. pytest collects files named test_*.py or *_test.py, and inside them functions whose names start with test_. A function named anything else is never run, which is the most common reason a test "does not exist".

Assertions are just assert

No assertEquals, no matcher library. pytest rewrites the plain assert statement so that a failure still shows you both sides.

Python
def test_equality():
    assert 3 == 3

def test_membership():
    assert "ok" in "not okay"

When one fails you get the expression and the values, not just "false":

Code
>   def test_fails():        assert 1 == 2
                             ^^^^^^^^^^^^^
E   assert 1 == 2

For floating point, comparing directly is a trap, because 0.1 + 0.2 is not 0.3:

Python
import pytest

def test_almost():
    assert 0.1 + 0.2 == pytest.approx(0.3)

To assert that something does fail:

Python
def test_raises():
    with pytest.raises(ZeroDivisionError):
        1 / 0

def test_raises_match():
    with pytest.raises(ValueError, match="invalid literal"):
        int("abc")

match takes a regular expression and is checked against the message, which stops the test passing on the right exception type raised for the wrong reason.

Fixtures

A fixture is setup that a test asks for by naming it as a parameter.

Python
import pytest

@pytest.fixture
def user():
    return {"name": "admin", "role": "qa"}

def test_fixture(user):
    assert user["name"] == "admin"

Use yield when there is teardown. Everything before the yield is setup, everything after runs once the test is done, pass or fail.

Python
@pytest.fixture
def resource():
    made = ["open"]
    yield made              # the test runs here
    made.append("closed")   # teardown, runs even if the test failed

Scope controls how often it is built. The default is function, meaning fresh for every test.

Scope Built once per
function test (the default)
class test class
module file
package package
session whole run

Verified: two tests sharing a session-scoped fixture see the same object, so a mutation in the first is visible in the second. That is the point of session scope and also its main hazard.

autouse=True applies a fixture to every test without it being requested, which is handy for logging and dangerous for anything else.

Put shared fixtures in conftest.py. Tests in that directory and below can use them with no import at all.

Parametrize

One test, many cases, each reported separately.

Python
@pytest.mark.parametrize("a,b,expected", [
    (1, 1, 2),
    (2, 3, 5),
    (10, -10, 0),
])
def test_add(a, b, expected):
    assert a + b == expected

Three tests appear in the report, not one. When case two fails you see exactly which inputs did it, which is the whole advantage over a loop inside a single test.

06

pytest markers, skips, and the flags worth knowing

Markers are pytest's tags, the same idea as TestNG groups or Cucumber tags.

Python
import pytest

@pytest.mark.smoke
def test_addition():
    assert 3 == 3

@pytest.mark.regression
def test_subtraction():
    assert 1 - 1 == 0

Declare your markers in pytest.ini to avoid warnings about unknown ones:

Code
[pytest]
markers =
    smoke: quick sanity tests
    regression: fuller suite
Terminal
pytest -m smoke                 # only tests marked smoke
pytest -m "smoke and not slow"  # boolean expressions work
pytest -m "smoke or slow"

The class ran the marked tests with -k smoke, and it worked, so this is a sharpening rather than a correction. -m is the flag for markers; -k matches keywords, which includes both marker names and test names. That difference bites, and two cases were measured on pytest 9.1 rather than asserted. Given a test called test_alpha marked smoke and a second called test_beta_smoke with no marker at all, pytest -m smoke selected only the marked one, while pytest -k smoke selected both. Given tests named test_a, test_b and test_after, the filter pytest -k "test_a or test_b" selected three, because test_after contains test_a as a substring. Neither surprise is rare in a real suite where names share prefixes. Use -m when you mean the tag, and keep -k for when you genuinely want to match on names.

Skips and expected failures

Python
import sys
import pytest

@pytest.mark.skip(reason="not implemented yet")
def test_skipped(): ...

@pytest.mark.skipif(sys.version_info < (3, 11), reason="needs 3.11+")
def test_conditional(): ...

@pytest.mark.xfail(reason="known bug, ticket QA-123")
def test_known_bug():
    assert False

The distinction matters when you read a summary line:

  • skipped: did not run.
  • xfailed: ran, failed, and that was expected. Not a failure.
  • xpassed: ran, and unexpectedly passed. Usually means the bug is fixed and the marker should come off.

A real run of every pytest example on this page reported 12 passed, 2 skipped, 1 xfailed, 1 xpassed.

Flags worth memorising

Flag Does
-q / -v quieter / more verbose output
-x stop at the first failure
--lf rerun only the tests that failed last time
--ff run last failures first, then the rest
-k EXPR select by keyword, name or marker substring
-m EXPR select by marker
--collect-only list what would run, without running it
-s do not capture output, so print shows
--maxfail=N stop after N failures
-n 4 run on 4 processes, needs pytest-xdist

--lf is the one that changes how a day feels. Fix, rerun only what broke, repeat, and run the suite once at the end.

You will meet pytest twice: once as a Python test framework, and again underneath DeepEval, the LLM evaluation framework later in the course, which is built on it. Evaluating a model output is the same shape as any other test: an expected result, an actual result, and an assertion between them. The markers, fixtures and flags above all carry over unchanged.

07

What CrewAI is

CrewAI is a Python framework for orchestrating role-based LLM agents. Open source, and free unless you want to run it on their infrastructure.

It offers two shapes:

  • Crew. A team of agents with autonomy, collaborating and passing results to each other.
  • Flow. Event-driven. When something happens, do something.
Crew: a team Agent 1Agent 2Agent 3 each has one role, and passes its result on Flow: event driven EventHandlerAct when this happens, do that A crew you would actually build Run build 1 and build 2 through Jenkins, compare results.json against results2.json, find the tests that changed answer between runs, and post the flaky list to Slack. Nobody watching.
The flaky-test example from the session: one role, done unattended, forever.

The example used to make it concrete: you have a results.json from a Playwright run, 100 tests with 3 failures. An agent whose single role is finding flaky tests reruns the build, compares the two result files, works out which tests changed their answer, and messages you the list. It runs whether or not your laptop is on.

08

Three levels of building an agent

L0 Vibe coding it works, you cannot say why L1 n8n, Langflow visual, wired by hand L2 CrewAI, LangChain code, deployable, runs unattended Go left to right. Each level teaches you what the next one is doing under the surface. The dividing line is autonomy, not difficulty. A coding assistant waits for you to ask. A deployed agent does not.
Work through the levels rather than jumping. Level two is where an agent stops needing you.

The distinction that answers "why not just use a coding assistant for this": a coding assistant is triggered, and an agent is autonomous. Deploy the agent on a server and it runs on its trigger, at any hour, without you. The point was made with a real figure: a production team running more than thirty-five independent agents for flaky-test detection, root-cause analysis, bug triage and downtime tracking, none of which anyone wakes up.

09

The anatomy of an agent

AGENT rolegoalbackstorytoolsllm who it iswhat it must achievewhy it is crediblewhat it can reachits brain TASK descriptionexpected_outputcontext what to dowhat good looks likewhat it may draw on CREW holds both kickoff() output Everything except the kickoff line is configuration. That is the whole framework.
Five fields, three fields, one container, one call.

Agent accepts more than sixty fields and Task more than thirty. These are the ones worth knowing, all confirmed present on the real classes in crewai 1.15.17.

Agent field What it does
role The job title. Shapes how the model answers.
goal What this agent is trying to achieve, in one sentence.
backstory Why it is credible. This is prompt context, not decoration.
llm The model. Omit it and CrewAI defaults to OpenAI.
tools A list of tools it may call.
verbose Print the reasoning as it works. Invaluable while learning.
memory Let it remember across executions.
allow_delegation Whether it may hand work to another agent.
max_iter Cap on reasoning loops, so a confused agent stops.
max_rpm Requests per minute, for staying inside a rate limit.
cache Reuse identical tool results.

role, goal and backstory are prompt engineering, not metadata. They are assembled into the system prompt, so vague ones produce vague output. "QA Engineer" beats "Assistant", and a backstory naming the domain beats a generic one.

Task field What it does
description What to do.
expected_output What good looks like. Skipping this is the most common cause of a disappointing result.
agent Which agent owns it.
context A list of other tasks whose output feeds this one.
tools Override the agent's tools for this task.
async_execution Run without blocking the next task.
output_file Write the result to a path.
output_json / output_pydantic Force structured output instead of prose.
callback A function to run when the task finishes.

context is how agents chain. It takes task objects, not strings:

Python
analyse = Task(description="Analyse the ticket", expected_output="a summary", agent=analyst)
write   = Task(description="Write test cases",  expected_output="a list",    agent=writer,
               context=[analyse])          # write now receives analyse's output
10

Building the first agent

Install it, with either package manager. CrewAI requires Python 3.10 or newer (and, as of this version, below 3.14):

Terminal
pip install crewai
Terminal
uv tool install crewai

Tools are a separate package, which is what the missing-module error is telling you when an import fails:

Terminal
pip install crewai-tools

The core install pulls in a lot, including chromadb and openai, so expect it to take a while and to want its own virtual environment.

Then five steps, in order. This is the whole file.

0BrainLLM(...) 1Agentwho it is 2Taskwhat to do 3Crewholds both 4kickoff()the only line that runs everything inside the dashed box is configuration Step 0 is the one that will cost you time, so read the next section before you write it.
Four boxes of setup and one call. Step 0 is where the live build lost half an hour.
Python
import os
from dotenv import load_dotenv
from crewai import Agent, Task, Crew, LLM

load_dotenv()

# 0. the brain
llm = LLM(
    model="openai/gpt-oss-120b",
    base_url="https://api.groq.com/openai/v1",
    api_key=os.getenv("GROQ_API_KEY"),
)

# 1. the agent identity
qa_agent = Agent(
    role="QA Engineer",
    goal="Analyse the feature or requirement and create 5 to 10 test cases",
    backstory="You are a senior QA engineer with 15 years of experience "
              "in test planning and test case design.",
    llm=llm,
    verbose=True,
)

# 2. the task
task = Task(
    description="Create 5 to 10 test cases for the login page of the practice app.",
    expected_output="A numbered list of test cases, each with a one-line description.",
    agent=qa_agent,
)

# 3. the crew
crew = Crew(agents=[qa_agent], tasks=[task], verbose=True)

# 4. kick off
if __name__ == "__main__":
    result = crew.kickoff()
    print(result)

Run it with python test_analyst_agent.py. That is it.

verbose=True prints everything the agent does as it does it, which is how you see the crew start, the task get assigned, and the final answer come back.

Notice how little of that file is code. The agent, the task and the crew are configuration; the only line that does anything is crew.kickoff(). That is a useful lens on the whole framework, and it is why "can a coding assistant not do this?" misses the point. The difference is not the output, it is that this file can be deployed and triggered without you.

11

The brain, and the error that ate half the build

This is the part that fought back live, and it is worth getting right because you will hit it within five minutes of starting.

CrewAI defaults to OpenAI. If you do not have an OpenAI key, you have to point it somewhere else, and the obvious guess does not work.

What everyone tries first model="groq/openai/gpt-oss-120b" ImportError: did not match any supported native provider groq is not in CrewAI's native provider list. Two routes work: A. Use the openai prefix, no extra install model="openai/gpt-oss-120b" base_url="https://api.groq.com/openai/v1" Groq speaks the OpenAI protocol, so this just works B. Install LiteLLM, then groq/ works pip install litellm model="groq/openai/gpt-oss-120b" the error message itself tells you this Natively supported prefixes include openai, anthropic, google, bedrock, openrouter, deepseek, ollama, cerebras
The failure and both fixes, checked against a real install.

Route A is the one used in the code above, and it needs nothing extra. Groq exposes an OpenAI-compatible endpoint, so the openai/ prefix plus Groq's base_url is all it takes.

This is why the live build kept returning "model not found" and why the assistant's fixes kept missing. The error was never the model name; groq/ simply is not a provider CrewAI resolves natively in this version. Verified against crewai 1.15.17: the natively supported prefixes are openai, anthropic, claude, azure, azure_openai, google, gemini, bedrock, aws, openrouter, deepseek, ollama, ollama_chat, hosted_vllm, cerebras, dashscope and snowflake. Anything else needs LiteLLM. OpenRouter is on that list, which matters if you are using it for another model.

The general lesson is worth more than the fix: the error message named the solution ("install LiteLLM for broad model support") and it went unread while a fix was guessed at instead. Read the whole error before asking anything to repair it.

The same pattern covers any OpenAI-compatible endpoint, which is most of them.

Provider What to write
OpenAI model="openai/gpt-4o"
Groq model="openai/<model>" plus base_url="https://api.groq.com/openai/v1"
OpenRouter model="openrouter/<vendor>/<model>", natively supported
Ollama, local model="ollama/llama3"
Anthropic model="anthropic/<model>"
Anything else install LiteLLM, then use its prefix
12

Crew, and the two processes

The crew is the container. It holds the agents and the tasks, and kickoff() is the only line that runs anything.

Python
from crewai import Crew, Process

crew = Crew(
    agents=[analyst, writer],
    tasks=[analyse, write],
    process=Process.sequential,     # the default
    verbose=True,
)
  • Process.sequential runs tasks in list order, each receiving what came before. Start here.
  • Process.hierarchical adds a manager that decides who does what and in which order.

A hierarchical crew requires manager_llm. Leaving it out raises a ValidationError at construction, before anything runs, which is at least a fast failure. Verified on crewai 1.15.17.

Python
Crew(agents=[...], tasks=[...], process=Process.hierarchical)                  # ValidationError
Crew(agents=[...], tasks=[...], process=Process.hierarchical, manager_llm=llm) # fine

Other crew fields worth knowing: memory, cache, max_rpm, and planning for a planning pass before execution.

13

Common errors, and what each one means

Every error here was produced against a real install rather than recalled, so the symptom column is the text the library actually prints.

Symptom Cause Fix
did not match any supported native provider The prefix is not native Use openai/ with a base_url, or install LiteLLM
Asks for an OpenAI key you never set llm was omitted, so it defaulted to OpenAI Pass llm= explicitly on the agent
ValidationError on a hierarchical crew No manager_llm Add one, or use Process.sequential
No module named 'crewai_tools' Tools are a separate package pip install crewai-tools
Vague or rambling output No expected_output Describe the shape you want
Agent loops without finishing No iteration cap Set max_iter
14

When AI helps too much

The most useful thing in the session was not the agent. It was watching an assistant asked to add a small pytest wrapper, and instead rewrite the file, add a main function, generate a config, and start creating files in an unrelated folder. The session's response was to stop it, delete the extra work, and just run the file:

Terminal
python test_analyst_agent.py

It worked first time, and had worked all along.

The line to keep: if you do not understand what you are doing, AI will make a fool out of you. Not because the model is bad, but because you cannot review a change you do not understand, and a change you cannot review is a change you have accepted. The defence is the same one from the async class: read every edit before accepting it, and when a one-line command will do, run the one-line command.

15

What comes next

The agent built here is entirely hardcoded, which was stated plainly rather than glossed over. The task is a fixed string and the output goes to the terminal. Both become dynamic next: agents and tasks can be parameterised, and the task description can come from a Jira ticket instead of a literal.

The build trailed for the coming sessions: a multi-agent Jira triage system, where one agent fetches the ticket, another does root-cause analysis, a third recommends test cases, and the result is posted to Slack. Guardrails and hallucination checks come after that.

There is a second shape to know about before that. Crew is a team of agents collaborating on tasks. Flow is event-driven: when something happens, run this. Most QA work starts as a crew, and a flow is what you reach for when the trigger matters as much as the work, such as a build finishing or a ticket changing state.

The realistic first project, if you want one before the triage system: an agent whose only role is finding flaky tests. Rerun a build, compare the two result files, work out which tests changed their answer between runs, and post the list to Slack. One role, done unattended, on a schedule nobody has to remember.

16

Tasks and announcements

  • Today's task: build your first CrewAI agent, run it, and post a screenshot of the run.
  • Install CrewAI with pip or uv, and put a GROQ_API_KEY in a .env file that is listed in .gitignore.
  • Practise the Python exercises. The running count is past 180.
  • Ask your questions in the community thread, which was noticeably empty.
  • Tuesday and Thursday evening: the AI Fluency certification sessions. Do not miss these.
  • Coming next: the rest of CrewAI, a research agent, triage agents, and the multi-agent Jira system.

Check yourself: what does if __name__ == "__main__" actually test, and when is it false? Which pytest flag selects by marker, and what else does the other one match? Name the five fields of an agent and the three of a task. Why does groq/ fail as a model prefix, and what are the two ways to fix it? In the whole agent file, which single line actually does anything?