The Testing Academy · Class Notes Saturday, 19 September (IST)
Live class · study guide

The end-to-end QA co-pilot, and what an agentic run costs

A LangChain recap, then one pipeline: a Jira ticket in, a local RAG for existing test cases, a planner, an executor driving a real browser, a reporter posting to Slack. Run live end to end. Plus the two numbers usually left out of a demo: what one agentic run costs, and the suite size at which this stops being maintainable.

By Pramod Dutta, The Testing Academy. Study notes from the live AI Tester Blueprint 3x class, rebuilt from the session recording and the batch repository, which received the whole pipeline at the end of the session (commit "feat: add the full Jira to browser QA pipeline", pushed as the class closed). The Eraser deck was not reachable while this page was written. The pushed code is more complete than what was shown on screen, in two places that matter, and both differences are flagged below. Cost figures are the instructor's own from a live run and from consulting work, reported as stated.

01

What this class covered

  • A recap of the LangChain work: models, streaming, chains, tools, structured tools
  • Recording a flow once with Playwright MCP, then replaying it through LangChain
  • What one full agentic run costs, priced live from real token usage
  • Where fully agentic execution stops working, and the suite size at which it does
  • Choosing between a vibe-coded script, n8n or Langflow, and Crew.ai or LangChain
  • What LangSmith adds: traceability and observability
  • The six-stage QA co-pilot pipeline, with three agents and two human gates
  • The local RAG behind it, in about a hundred lines
  • What is still missing, and why LangGraph comes next
02

Where this sits

The batch is near the end. LangChain is covered, the exercises run, and today was the recap that turns them into one thing, before LangGraph arrives next week as the production form of it.

The recap list, as run through in class: invoking a model, Gemini, streaming, parallel and sequential chains, tools, multiple tools, structured tools, and then Playwright as an orchestrator driven by DeepSeek.

A small live hiccup worth keeping. The first rerun of the orchestrator did nothing, because the API key had been deleted between sessions. Worth knowing that the failure mode looks like a silent hang rather than a clear authentication error, so check the key before you start debugging your chain.

03

Record once with MCP, replay forever without it

This is the pattern from the recap that is worth taking to work on Monday, and it is not obvious.

You have two ways to produce an end-to-end flow for the orchestrator. Write it by hand, or let the agent explore the app and write it for you. The second uses Playwright MCP, and only once:

PAY THE EXPLORATION COST ONCE First run: Playwright MCP explores the app, finds the login page, verifies every step Writes the confirmed path into the task file, as the steps that are known to work Every later run LangChain plus Playwright tools, no MCP, no re-exploring Why it matters Exploration is the expensive part of an agentic run, in both seconds and tokens. Doing it once and keeping the result is the difference between a demo and something you can afford to run nightly.
The first pass is slow and thorough. Every pass after it follows a path that has already been proven against the real application.

What you give up, and what you do not. You still are not writing locators. The orchestrator works them out from the page through the Playwright tools, on every run. What gets cached is the route, the sequence of steps that is known to work, not the selectors.

Honest about the failure modes, as shown live: the model occasionally stalls or takes a wrong step, more often when it is rushing. The stated fix is better instructions rather than retries.

04

What one run actually costs

The part almost no demo includes, and the reason it was included here: if you cannot answer this, you cannot propose the thing at work.

The full end-to-end run, on DeepSeek v4.1 flash, was priced live by asking the agent to fetch current pricing and compute from its own token usage.

Measure Figure
One full end-to-end run about $0.69, under one rupee
Run time about 16 seconds
5,000 test cases, typical load roughly 15,000 to 24,000 rupees a month
5,000 test cases, 80% load up to about 100,000 rupees a month, described as the worst case

The comparison offered: a commercial grid such as BrowserStack or LambdaTest, or a person doing the same work. And a real number from consulting: one project running 3,000 test cases pays $3,000 to $4,000 a month to a third-party vendor.

Why the run happened outside peak hours matters. DeepSeek prices differ by time of day, and Indian working hours fall in the cheaper window for it. That is part of why the number is this low, and it is worth stating if you quote the figure to your management.

05

Where fully agentic stops working

The most useful five minutes of the session, and it came from the room pushing back rather than from the slides.

Two objections were raised: what happens when tests fail because they are flaky rather than broken, and how you debug when many tests fail at once. Both were accepted as correct.

The honest answer, stated plainly in class. "Maintenance is the real challenge in orchestrated test cases." And the reason this has not become the industry default: at scale it does not hold.

  • Small organisation, around 1,000 test cases and two QA engineers: this works well, and costs a few thousand rupees a month.
  • Large organisation, around 12,000 test cases: going fully agentic is described as not even possible. "Maintenance will kill us."

The framing to take away: agentic execution is a real tool with a real ceiling. Knowing where the ceiling is, and saying so, is what separates an engineer from a demo.

06

Which framework, when

Asked directly, and answered as three tiers rather than a winner.

Problem size Use
Quick A vibe-coded solution. Just write it.
Medium n8n or Langflow. Visual, fast, good enough for a proof of concept.
Production Crew.ai or LangChain.

And within production:

  • Crew.ai when you want multiple crews collaborating and handing work to each other, each with a role.
  • LangChain when you want a defined sequence, and when you want the rest of the Lang family, particularly LangSmith.

LangSmith, in one line. Traceability plus observability: seeing what your agent actually did, step by step, and what it replied. If you expect to need that (and in production you will), that is the argument for staying inside the LangChain family rather than a technical superiority claim.

On the comparison itself. Both support RAG, and the class was explicit that neither is wrong. The phrase used was that choosing between them is like choosing between two pizza chains. The recommendation that came back when the question was put to a model, with this specific QA co-pilot described: LangGraph plus LangChain, with Crew.ai called a valid but second choice. That matches where the course is heading next week.

07

The pipeline, as pushed

The session built up to the full thing and ran it live. The code pushed at the end of class is more complete than the version on screen, so what follows is the repository.

Six stages, three of them agents:

Text

1. FETCH     the Jira ticket            (REST v3, offline fixture fallback)
2. RETRIEVE  similar existing test cases from a local RAG      <- human gate
3. PLAN      a plan grounded in BOTH the ticket and the RAG hits   (Groq)
4. EXECUTE   the automatable cases in Chromium                 <- human gate
5. REPORT    a reporter agent posts the verdict to a dummy Slack MCP
Terminal
python 013_FULL_E2E_Fetch_JIRA_Local_RAG_QA_Orch.py VWO-49

Three flags, and the third is the interesting one:

Flag Effect
--yes skip both human gates, for CI
--dry-run stop after stage 3, never open a browser
--no-rag skip retrieval entirely, to see what grounding is actually worth
SIX STAGES, THREE AGENTS, TWO GATES 1. FETCH Jira REST v3, with an offline fixture fallback 2. RETRIEVE local RAG over your existing test cases human gate review the hits 3. PLAN agent, Groq. Grounded in ticket AND hits human gate before the browser 4. EXECUTE agent, DeepSeek. Chromium via the Playwright tools 5. REPORT agent. Posts the verdict to a dummy Slack MCP The flags --yes skip both gates, for CI --dry-run stop after stage 3 --no-rag skips stage 2 entirely. Run it once with and once without on the same ticket: the difference in the plan is exactly what grounding in your own test library is worth, measured rather than asserted.
The gates are the part that makes it deployable. An agent that can open a browser against your app without anyone confirming the plan is a demo, not a pipeline.

Two models on purpose, which is worth copying. Groq writes the plan, DeepSeek drives the browser. The stated reason is that DeepSeek is markedly better at long multi-step tool discipline, which is what a browser run needs, while planning is a single well-shaped generation. Both are cheap. Picking a different model per stage is a real optimisation, not a fussy one.

Two differences between the class and the pushed code, both worth knowing. First: in class the ticket VWO-49 was described as containing negative login cases for TTACart. The fixture in the repository is a different ticket: "Add Passkey and SSO Login Options to VWO Login Page", a story with six acceptance criteria covering UI visibility, SSO flow, passkey flow, error handling, compatibility and security. Use the repository version. Second: the session closed by saying the pipeline was "too much agentic" and that human-in-the-loop and state management were missing. The pushed code already has two human gates, one after retrieval and one before the browser opens. The state-management gap is real and stands, and it is the actual argument for LangGraph.

08

The local RAG, which is smaller than you expect

The retrieval step does not use a vector database server. It is about 100 lines over a JSON file, and reading it is the fastest way to understand what RAG actually is.

Piece Choice
Corpus 20 TTACart test cases in a JSON file
Embeddings fastembed, model BAAI/bge-small-en-v1.5, 384 dimensions, about 130MB
Store a numpy array
Similarity vectors normalised to unit length, so a dot product is cosine similarity
Cache the matrix saved to disk, keyed by a hash of the corpus, so it rebuilds only when the corpus changes
Fallback if fastembed is not installed, it degrades to keyword overlap so the demo still runs offline
Python
rag = TestCaseRAG()
hits = rag.search("SSO and passkey login buttons", k=3)
# [(0.71, {...}), (0.66, {...}), (0.61, {...})]  best first

The line in the module worth quoting. "At 20 documents a vector database is a numpy array, and anything heavier is ceremony." Swap the body of search() for a Qdrant call when the corpus outgrows memory; the interface does not change. If you are reading the RAG concepts notes from the 4x session on the same morning, this is that whole diagram in one file you can run.

What gets embedded is the title, area, type, tags, steps and expected result of each case. What comes back into the prompt is each hit with its id, type, priority, area, status, last run result and similarity score, so the planner can cite an existing case by id instead of reinventing it.

09

What is still missing, and what comes next

Not production-ready out of the box, and the list is specific. The pipeline is ready to adapt, not to deploy. Before it goes anywhere real: point it at your Jira, point the RAG at your test case library, then verify that the plans it writes are correct, that coverage is what you expect, and that the token cost holds at your volume. "If you have not tested it from your end with multiple tickets, what is the point?"

LangGraph arrives next week as the production form of this: explicit branching, investigation loops, proper state, and human decisions as first-class parts of the graph rather than prompts bolted on. It runs on LangChain underneath. LangSmith comes as extra sessions rather than being dropped.

10

Tasks and announcements

  • No class the following day. The week is deliberately cleared so the batch can catch up.
  • Finish every pending recording and exercise before the next class: RAG, MCP, LangChain and Crew.ai. The stated minimum before LangGraph starts is 12 to 13 LangChain exercises.
  • The reason for the insistence: LangGraph assumes all three. Arriving without RAG, MCP and LangChain means not following it.
  • Course completion certificates come next week, after LangGraph.
  • Hackathon certificates are being released now. Prize money has already been transferred.
  • The next hackathon is planned for around October, and will include the AI 4x batch, so there will be competitors from outside this cohort.
  • A resume template for describing agentic experience was requested and will be shared.
  • Recordings stay available for life, confirmed in response to a question about post-batch access.