Projects . AI Tester Blueprint
Learn AI for testing through 27 hands-on projects
The course starts with local AI tools and prompt engineering, moves into apps and automation, then expands into LangFlow, RAG, MCP, CrewAI agents and custom MCP servers. Every project links to its files in the course repository.
27
Projects
From a first local LLM to an autonomous QA agent.
6
Phases
Local AI, apps and agents, retrieval, MCP and crews, evaluation, agentic QA.
75
Repo and chapter links
README, code and chapter pages, one click from each card.
0 to 26
Project numbers
Repo folders Project_00 to Project_23 today; the agentic chapters ship from their chapter pages.
How you move through the program
Every later phase gets easier once the earlier one is clear.
1 Start with control Projects 0 to 4 show how to control prompt behaviour, model behaviour and input structure before adding more moving parts.
2 Move to useful systems Projects 5 to 10 turn AI into visible workflow value: apps, agents, Jira integration, and content or bug workflows.
3 Learn retrieval and flow engineering Projects 11 to 15 go from LangFlow fundamentals to retrieval theory, flow engineering, product code and embeddings.
4 Extend into MCP and multi-agent systems Projects 16 to 19 connect AI to MCP, Python foundations, CrewAI crews and custom MCP servers built from scratch.
5 Score what the model says Projects 20 to 22 replace assertEquals with scored evaluation: golden datasets, DeepEval metrics in pytest, then a full evaluation framework.
6 Go agentic, end to end Projects 23 to 26 move from a LangChain agent to a LangGraph Jira-to-report graph and an autonomous QA crew with stage gates.
Phase 1: Local AI and prompting
Private local LLM workflows, prompt frameworks and reusable prompt assets. Projects 0 to 4.
Project 00 Phase 1
LLM Basics for QA and SDET
A practical LLM foundation for testers: the core concepts you need to test LLM systems, the Transformer paper explained in QA terms, and a glossary of every term the course uses.
Focus LLM fundamentals, what to test in an LLM app Stack Markdown guide, HTML tutorial, glossary Key idea Attention Is All You Need, read as a tester
Why LLM apps need probabilistic testing The Transformer paper explained for QA and SDETs A keyword glossary for the whole course
Project 01 Phase 1
Local Test Case Generator
Generate structured test cases from user stories with a local LLM while keeping everything private on the machine.
Focus Test case generation, prompt engineering Stack Python, FastAPI, Vanilla JS, Ollama Key idea B.L.A.S.T. protocol for agentic AI
Local execution with Llama 3.2 through Ollama Structured JSON-style output for QA use Clear backend, tool, and UI separation
Project 02 Phase 1
Selenium to Playwright Converter
Convert Selenium Java tests into Playwright TypeScript with a privacy-first local AI workflow.
Focus Code conversion, legacy migration Stack React, Vite, Node.js, TailwindCSS, Ollama Key idea Local secure migration assistant
CodeLlama-powered local conversion Express proxy over Ollama Modern UI with code editing and quick testing
Project 03 Phase 1
RICE-POT Prompt Framework
Show how a disciplined prompting framework produces enterprise-style automation output instead of vague code.
Focus Prompt engineering, framework generation Stack Java, Selenium, Maven, TestNG Key idea RICE-POT prompt structure
Role, instruction, context, example, parameters, output, tone Enterprise automation constraints baked into the prompt Salesforce login used as the target app
Project 04 Phase 1
Local LLM Prompt Templates
Create reusable prompt templates that turn PRDs and context files into actionable QA output.
Focus Prompt templates, PRD analysis Stack Markdown templates, Playwright TypeScript, context files Key idea Context-constrained prompting
Separates PRD source, constraints, and output template Useful for repeatable requirement-to-test workflows Turns one-off prompting into reusable assets
Phase 2: Apps, agents and automation
Product-like UIs, Jira agents, no-code automation and AI-assisted workflows. Projects 5 to 10.
Project 05 Phase 2
Job Board Assistant
Build a usable Kanban-style product with AI assistance while keeping data local in the browser.
Focus AI-assisted full-stack development Stack React 19, TypeScript, Vite, Tailwind CSS Key idea Production-feeling app built with AI support
Kanban board with six lifecycle columns Search, filters, import/export, and stats LocalStorage-based persistence
Project 06 Phase 2
AI Resume Fix for LinkedIn
Use AI to reposition QA resumes for stronger role targeting and clearer professional storytelling.
Focus Resume optimization, career tools Stack AI prompts, DOCX, PDF Key idea Career artifact enhancement with AI
Turns prompting into a real career deliverable Generates role-targeted resume assets Extends the course beyond pure code generation
Project 07 Phase 2
TestPlan AI Agent + JIRA Integration
Generate structured QA artifacts from JIRA tickets through a full-stack AI agent workflow.
Focus AI agents, JIRA integration, app development Stack Node.js, Express, React, TypeScript, Tailwind CSS Key idea A.N.T. 3-layer architecture
Combines settings, JIRA fetch, template handling, and generation routes Supports Groq and local model modes Connects QA planning directly to ticket data
Project 08 Phase 2
n8n AI Workflow Automation
Teach testers how to automate multi-step AI workflows without writing every integration by hand.
Focus No-code AI agents, workflow automation Stack n8n, Groq API, Jira API, Google Docs, Google Sheets Key idea Visual workflow automation for testers
Introduces no-code orchestration for QA work Connects docs, tickets, sheets, and models Builds comfort with multi-step workflow thinking
Project 09 Phase 2
Content Creation Agent (n8n)
Plan a scheduled AI pipeline that can generate and publish content automatically.
Focus Content automation, scheduling Stack n8n, API workflows, automation prompts Key idea Repeatable publishing automation
Workflow thinking for daily content production Mixes topic discovery, generation, imagery, and publishing Good example of AI automation outside testing execution
Project 10 Phase 2
BugSnap
Improve bug reporting workflows by adding AI, workflow automation, and vector-style enrichment ideas.
Focus Bug reporting tools Stack Workflow JSON, markdown, web tooling Key idea Bug reporting as an AI-enhanced workflow
Links reporting to retrieval and vector operations Makes bug workflows part of the AI curriculum Useful transition toward semantic systems
Phase 3: LangFlow and retrieval systems
LangFlow basics, RAG theory, visual flow engineering, modular RAG apps and embeddings. Projects 11 to 15.
Project 11 Phase 3
LangFlow Fundamentals
Teach LangFlow basics and starter QA agents before students start building retrieval systems.
Focus LangFlow basics, starter agents Stack LangFlow, Groq, prompt templates, API request nodes Key idea Visual AI building blocks before RAG
Simple chatbot and prompt-based QA assistant RICE-POT test case generator flow JIRA and PDF starter agent flows
Project 12 Phase 3
RAG Basics
Explain what RAG is, why retrieval matters, and how the main architecture families differ.
Focus RAG theory, architectures, evaluation Stack Python, LangChain, markdown Key idea Grounding model answers in private documents
Covers the major RAG patterns clearly Includes evaluation and testing guidance Acts as the theory layer before LangFlow or app implementation
Project 13 Phase 3
RAG with LangFlow
Translate retrieval patterns into importable LangFlow flows that students can inspect and test visually.
Focus Low-code AI, visual node programming Stack LangFlow, Chroma, Groq, AstraDB, n8n Key idea Drag-and-drop RAG pipeline engineering
Naive, advanced, modular, graph, and other RAG flows Import instructions for LangFlow and n8n Makes prompts, retrieval, and routing inspectable on a canvas
Project 14 Phase 3
RAG VIBE Coding App
Show how retrieval theory and visual flow ideas become a working modular application.
Focus Full-stack RAG app, modular retrieval Stack FastAPI, ChromaDB, Python, static HTML Key idea Domain-routed ingestion and answer generation
Upload route for API, UI, and performance documents Router decides the correct vector store per query Static dashboard for ingestion and chat
Project 15 Phase 3
Vector Embeddings Visualizer
Make chunking, embeddings, similarity, and vector projection easy to explain in workshops.
Focus Embeddings, chunking, similarity, RAG foundations Stack FastAPI, Vanilla HTML/CSS/JS, Ollama, OpenAI, Mistral Key idea Visible teaching model for vector search
Chunk cards, vector previews, and heatmaps Demo mode plus real embedding providers Useful bridge between theory and retrieval systems
Phase 4: MCP, CrewAI and custom agents
MCP workflows, Python for AI, CrewAI multi-agent crews and custom MCP servers. Projects 16 to 19.
Project 16 Phase 4
MCP Basics
Introduce MCP-driven QA workflows using Playwright orchestration, execution evidence, and JIRA-linked reporting ideas.
Focus MCP, browser orchestration, evidence capture Stack Playwright MCP, JIRA workflow concepts, HTML reporting Key idea Tool-connected QA execution with AI assistance
Playwright-driven test execution through MCP Failure evidence and screenshot-oriented reporting Bridges AI prompting with browser and issue-tracking actions
Project 17 Phase 4
Python for AI Testers
Build a solid Python foundation covering basics through advanced concepts needed for AI-powered testing tools.
Focus Python fundamentals, data structures, modules Stack Python, JSON, OS module, Lambda functions Key idea Python fluency as a prerequisite for AI tooling
21 progressive Python exercises from Hello World to CrewAI intro Covers variables, strings, lists, loops, dicts, tuples, functions Includes JSON handling, OS module, lambda functions, and module imports
Project 18 Phase 4
CrewAI Multi-Agent Systems
Build multi-agent AI crews for QA tasks including bug triage, JIRA test plan generation, and custom tool creation.
Focus Multi-agent AI, CrewAI framework, QA automation Stack CrewAI, Python, JIRA API, custom tools, memory Key idea Collaborative AI agents for QA workflows
7 progressive CrewAI examples from hello world to MCP integration Bug triage crew with HTML report generation JIRA-connected test plan agent with memory support
Project 19 Phase 4
MCP Server Creation
Learn to build custom MCP servers from scratch, progressing from simple calculators to real QA dashboard tools.
Focus Custom MCP servers, tool creation, QA dashboards Stack Python, MCP SDK, FastAPI, test data Key idea Building your own AI-connected tool servers
Hello World calculator MCP server Weather MCP server with API integration QA Dashboard MCP server with real test data
Phase 5: LLM evaluation with DeepEval
Scoring model output instead of asserting it, from first metrics to a 25-metric framework. Projects 20 to 22.
Project 20 Phase 5
LLM Evaluation Basics
Why assertEquals fails on LLM output and what replaces it: ground truth, golden datasets, evaluator families, thresholds and a live scoring demo.
Focus Scoring LLM output instead of asserting it Stack Python, Groq or local models, golden datasets, scoring scripts Key idea A score with a threshold is the new assertion
Ground truth and golden datasets for QA prompts Three evaluator families: exact match, similarity and LLM-as-judge A live scoring demo you can rerun on your own prompts
Project 21 Phase 5
DeepEval Basics
DeepEval as a pytest plugin: LLMTestCase fields, relevancy and hallucination metrics, thresholds, and a Groq judge wired up with deepeval set-local-model.
Focus Scored LLM tests inside pytest Stack DeepEval, pytest, Groq judge, LLMTestCase Key idea LLM checks run in the same suite as your other tests
LLMTestCase: input, actual output, expected output and context Answer relevancy and hallucination metrics with thresholds A Groq model as the judge via deepeval set-local-model
Project 22 Phase 5
DeepEval Framework Creation
A full LLM evaluation suite: 25 DeepEval metrics and 289 pytest cases grading a Groq chatbot and a RAG app, with an attack library and a recorded run.
Focus Framework design for LLM evaluation Stack DeepEval, pytest, Groq, a RAG app under test, HTML reports Key idea Evaluation as a maintained framework, not a one-off script
25 metrics organised into reusable suites 289 cases across a chatbot and a RAG app An attack library for jailbreak and prompt-injection checks
Phase 6: Agentic QA with LangChain and LangGraph
From a first LangChain agent to a Jira-to-report graph and a fully autonomous QA crew. Projects 23 to 26.
Project 23 Phase 6
LangChain for Testers
LangChain 1.x for QA in 13 scripts: create_agent, streaming, tools, typed output, and the first browser agent.
Focus Agents, tools and typed output Stack LangChain 1.x, Python, Groq or OpenAI-compatible models, Playwright Key idea An agent is a model plus tools plus a loop you control
create_agent, streaming and tool calling, step by step Typed, structured output you can assert on A first Playwright browser agent driven from LangChain
Project 24 Phase 6
LangChain + Playwright: Full Agentic QA Project
The complete agentic QA build: a Jira ticket becomes a browser test plan, LangChain tools drive Playwright, and the results flow back as a report.
Focus End-to-end agentic UI testing Stack LangChain, Playwright, Jira API, Python Key idea The agent plans, acts in the browser and reports, with you reviewing each gate
Jira-to-browser test pipeline from the LangChain chapter Playwright actions exposed to the agent as tools Run logs and a result report you can attach back to the ticket
Project 25 Phase 6
LangGraph: Jira to Execution to Results
LangGraph for testers in eleven lessons: state, routing, retries, parallel suites, checkpoints, human approval, a ReAct loop and a Jira-to-report capstone.
Focus Stateful agent graphs with approvals Stack LangGraph, LangChain, Jira API, Playwright, Python Key idea Graphs make long QA workflows resumable and reviewable
State, routing and retries as graph nodes Checkpoints and human approval before execution Capstone: a Jira ticket in, an executed suite and a sent report out
Project 26 Phase 6
Full Autonomous QA Agent
The capstone: an autonomous QA crew that turns Jira tickets into a test plan, test cases and Playwright code, with validation gates, an MCP-to-REST fallback and coverage reporting.
Focus Autonomous, gated QA pipeline Stack CrewAI, Streamlit, MCP and REST, Playwright, Jira Key idea Autonomy with stage gates beats autonomy without them
Jira tickets to plan, cases and Playwright code Validation gates between every stage MCP first with a REST fallback, plus coverage output
Get the code
Clone the repository once; each card above links straight to its folder and files.
terminal Copy
git clone https://github.com/PramodDutta/AITesterBlueprint.git
cd AITesterBlueprint