Lesson 10 of 11
Chapter 18 . LangGraph . Lesson 10 of 11

Flaky test analyzer

Chapter 5's flaky test finder, rebuilt with a decision, an approval pause and an optional LLM explanation.

01Real-world example

Illustration: Compare two CI runs of the same suite: a test that passed in one run and failed in the other is flaky.
Compare two CI runs of the same suite: a test that passed in one run and failed in the other is flaky.

02The graph

Flaky test analyzer as a graph
Clean runs go straight to the report. Flaky runs ask you first.

03How the code does it

  • load_runs reads result1.json and result2.json, two Playwright JSON reports, with playwright_results.load_report.
  • compare: a test is flaky if its verdict flipped between runs or it was retried inside a run; tests that failed in both runs are real bugs.
  • after_compare skips the approval when nothing is flaky; after_approve runs explain only when a key is set.
  • The count never depends on the model: the LLM only adds notes after a human approves.

Key points

  • The count is plain Python, never the AI
  • Pause before changing anything
  • Uses every idea from lessons 1 to 9

04Run it

terminal
cd chapter_18_LangGraph/src/chapters
python 010_Flaky_Analyzer_Graph.py --yes
output
Quarantine these flaky tests? (yes/no)
  - loginTests/auth.spec.ts > @P0 Login > redirects to dashboard after successful login

FLAKY TEST COUNT: 1
Compared 50 tests present in both runs
Same count as the chapter 5 LangFlow flow. With a GROQ key it also adds LLM notes.
Needs a key. Copy chapter_18_LangGraph/.env.sample to .env and add a free Groq key from console.groq.com. Without one the script still runs and skips the LLM notes.

05Then and now

Chapter 5 flow compared with the chapter 18 graph
Same answer as the chapter 5 LangFlow flow. Now it can choose a path, wait for you, and explain.