Chapter 5
Chapter 5 . Agents and workflows . LangFlow agents

LangFlow: every flow is an API

LangFlow lets you wire components on a canvas and then call the whole flow over HTTP. This chapter builds a flaky test analyzer twice: once with a model reading two Playwright reports, once as plain Python whose count can gate CI. It adds a Jira bug triage flow, an API contract check, and scripts that run LangFlow in Docker.

1
Flaky test in the bundled runs
FLAKY TEST COUNT: 1 and 2 consistent failures across 50 tests; chapter 18 gets the same answer.
8
Cases in test_flow.py
Six fixtures with known answers plus two error paths, run against a freshly imported copy of the flow.
22,737
Tokens for the LLM version
Recorded run on deepseek/deepseek-v4-pro: 40.9 s for the same verdict.
1.12.3
LangFlow image
Pinned in langflow-up.sh; most flows in the chapter were exported from 1.10.0.

01Why a tester cares

A red build is either a real bug or a flaky test, and the two need opposite actions: a bug goes to engineering, a flaky test goes to quarantine and a rerun. Telling them apart by hand means opening two reports and comparing test by test. This chapter turns that comparison into a LangFlow flow and then calls it like any other HTTP API, from a React UI, from a script or from CI.

LangFlow is a visual builder: you drop components (models, prompts, file readers, parsers, your own Python) on a canvas and connect their ports. The point for testers is what happens next: every saved flow gets a REST endpoint, POST /api/v1/run/{flowId}, so the flow you prototyped is already the service your tests call.

Flow or fileComponents (as wired in the JSON)What it does
Project/AI3X_001_HelloWorld.jsonChat Input, Groq (llama-3.1-8b-instant, key from the global variable GROQ_API_KEY), Chat OutputProves the canvas and your Groq key work.
Project/Hello_AIAgent.jsonChat Input, Ollama (qwen3.5:4b at http://localhost:11434), Chat OutputThe same check against a local model, no key.
Project/AI3X_002_Flaky_Test_AIAgent.jsonTwo Read File nodes, a Prompt Template with {file1} and {file2}, OpenRouter (deepseek/deepseek-v4-pro), Chat OutputThe LLM flaky analyzer. The React UI in ui/ uploads two reports and renders its Markdown answer.
flaky_test_analyzer_ai_Agent/flow/Flaky_Test_Analyzer.jsonChat Input, a custom Python component FlakyTestAnalyzer, Chat OutputThe same job with no model: a folder path in, a report with FLAKY TEST COUNT: N out. test_flow.py checks it.
Project/AI3X_003_Bug_Triage_AI_Agent.jsonAPI Request (GET one Jira issue), Parser, Prompt Template, OpenRouter (deepseek/deepseek-v4-flash), Chat OutputAsks for severity, priority, impact areas, a root-cause hypothesis and a justification for one ticket.
Project/AI3X_004_API_Contract_Validator.mdSpec only: a GET request, a sample response and a JSON Schema. There is no flow file.Build it yourself: call the API, then ask a model whether the response still matches the schema.
langflow-up.sh, langflow-down.shDocker Desktop + the langflowai/langflow:1.12.3 containerStart LangFlow on port 7860 with persistent storage, and stop it again.
flowchart LR
  subgraph LLM["AI3X_002: LLM agent flow"]
    FA["Read File (File-daKW7)"] --> PT["Prompt Template: file1, file2"]
    FB["Read File (File-IKmcY)"] --> PT
    PT --> OR["OpenRouter deepseek-v4-pro"] --> CO1["Chat Output"]
  end
  subgraph DET["Flaky_Test_Analyzer.json: deterministic flow"]
    CI["Chat Input: folder path"] --> AN["Flaky Test Analyzer, custom Python"] --> CO2["Chat Output"]
  end
The flaky analyzer, built twice. The LLM flow reasons over both files; the deterministic flow computes the count with Python and never calls a model.

02Flaky or broken: the rule the analyzer applies

The deterministic component works in three steps, and the order matters.

  1. Collapse retries. A test can have several results in one run (Playwright retries). They become one verdict per run: passed, failed, skipped, or flaky-in-run when one run saw both a pass and a failure.
  2. Compare the two runs. A test is flaky if its verdict differs between the runs (in either direction), or if Playwright already caught it flaky on retry. It is a consistent failure if it failed in both runs: a real, reproducible bug.
  3. Report. A test present in only one run is not comparable and never counted as flaky.
flaky_test_analyzer_ai_Agent/flow/flaky_analyzer_component.py
    def _verdict(self, statuses):
        """Collapse Playwright retries into one verdict per test."""
        if not statuses:
            return SKIPPED
        bad = {"failed", "timedOut", "interrupted"}
        had_pass = any(s == "passed" for s in statuses)
        had_fail = any(s in bad for s in statuses)
        if had_pass and had_fail:
            return FLAKY_IN_RUN  # Playwright's own retry already proved instability
        if had_pass:
            return PASSED
        if had_fail:
            return FAILED
        return SKIPPED
flaky_test_analyzer_ai_Agent/flow/flaky_analyzer_component.py
        shared = set(run_a) & set(run_b)
        only_a = sorted(set(run_a) - set(run_b))
        only_b = sorted(set(run_b) - set(run_a))

        # Flaky = same test, different verdict across the two runs,
        # plus anything Playwright already flagged flaky via retries.
        flipped = sorted(t for t in shared if run_a[t] != run_b[t])
        retry_flaky = sorted(
            t for t in shared
            if FLAKY_IN_RUN in (run_a[t], run_b[t]) and t not in flipped
        )
        flaky = flipped + retry_flaky

        consistent_fail = sorted(
            t for t in shared if run_a[t] == FAILED and run_b[t] == FAILED
        )

On the chapter's two real reports (50 tests each, Playwright 1.60.0, 8 workers) run 1 has 3 failures and run 2 has 2. The two failures are shared and exactly one test recovered, so the count is 1, not 3 and not 2:

report: result1.json vs result2.json
FLAKY TEST COUNT: 1

Run A: result1.json  (50 tests)
Run B: result2.json  (50 tests)
Compared: 50 tests present in both runs

## FLAKY_TESTS (1)
- [failed -> passed] loginTests/auth.spec.ts > loginTests/auth.spec.ts > @P0 Login > redirects to dashboard after successful login

## CONSISTENT_FAILURES (2)
- dashboardTests/dashboard.spec.ts > dashboardTests/dashboard.spec.ts > @P0 Dashboard > renders revenue chart with correct totals
- loginTests/auth.spec.ts > loginTests/auth.spec.ts > @P0 Login > rejects login with expired session token

## RERUN_RECOMMENDATION
Quarantine and rerun the 1 flaky test(s) above; they changed verdict between identical runs, so they are unreliable signals.
The 2 consistent failure(s) are real bugs, not flakiness. Fix them.

## RAW_STATS
- result1.json: {"startTime": "2026-06-18T02:44:27.764Z", "duration": 41666.884, "expected": 47, "skipped": 0, "unexpected": 3, "flaky": 0}
- result2.json: {"startTime": "2026-06-18T02:44:27.764Z", "duration": 39509.619, "expected": 48, "skipped": 0, "unexpected": 2, "flaky": 0}

The output of the component's analyze() on the bundled files, run offline for this page. Each key repeats the file name because the top-level suite title is the file and the spec's file field is added again; the flow README shows a shorter form. Chapter 18 rebuilds the same analyzer as a LangGraph graph and gets the same count: lesson 10, flaky test analyzer.

03Live demo: flaky or broken?

The table holds the test titles and statuses of the chapter's real result files and fixture folders; the report on the right is the component's exact output format, recomputed in your browser. Click any status to change it (passed, failed, failed then passed on retry, skipped, not in the run) and watch the count. Below the table you can see the REST call that would send the same folder to the published flow.

Flaky or broken? Compare two Playwright runsNo API key needed
data-testid=lf-presetdata-testid=lf-resetdata-testid=lf-tabledata-testid=lf-only-changeddata-testid=lf-flaky-countdata-testid=lf-consistent-countdata-testid=lf-nc-countdata-testid=lf-reportdata-testid=lf-api-detdata-testid=lf-api-llm
0flaky
0consistent failures
0not comparable
0compared
Tests (click a status to change it)
#TestRun ARun BVerdict
Report from the Flaky Test Analyzer component

    
Call the published flow (shown, not sent)
Chat Input Flaky Test Analyzer (Python) Chat Output
Request (the folder as seen inside the container)

      Response (built from the analyzer report)
      

    
Try these. In stable, flip one result in Run B: one changed verdict is enough to count as flaky. In retry_flaky, the test that failed then passed inside run 1 shows as [flaky-in-run -> passed]. In ui/samples, a test that was skipped in run 2 counts as flaky too: decide whether your team agrees with that rule.

04Call a published flow over REST

Every flow answers POST /api/v1/run/{flowId}?stream=false. The body carries input_value (what a Chat Input receives), input_type, output_type, an optional session_id and tweaks: overrides for any component field, keyed by component id. The report text comes back at outputs[0].outputs[0].results.message.text.

The LLM flow takes files, so the React UI makes two kinds of call: it uploads each report to get a server file_path, then runs the flow with both paths as tweaks on the two Read File components.

sequenceDiagram
  participant B as React UI (port 5173)
  participant V as Vite proxy
  participant L as LangFlow
  B->>V: POST /api/v1/files/upload/flowId with result1.json
  V->>L: same request, same origin for the browser
  L-->>B: file_path A
  B->>V: POST /api/v1/files/upload/flowId with result2.json
  V->>L: forward
  L-->>B: file_path B
  B->>V: POST /api/v1/run/flowId?stream=false, paths as tweaks
  V->>L: forward
  L-->>B: outputs, report text inside
Two uploads, one run. The browser only ever talks to its own origin; Vite forwards to LangFlow.
flaky_test_analyzer_ai_Agent/ui/src/lib/api.js
// Uploads one file and returns the server-relative path to feed into tweaks.
export async function uploadFile({ apiBase, apiKey, flowId }, file) {
  const form = new FormData()
  form.append('file', file)
  const res = await fetch(`${trimBase(apiBase)}/api/v1/files/upload/${flowId}`, {
    method: 'POST',
    headers: { 'x-api-key': apiKey },
    body: form,
  })
  if (!res.ok) throw new Error(`Upload failed for "${file.name}": ${await readError(res)}`)
  const data = await res.json()
  if (!data?.file_path) throw new Error(`Upload of "${file.name}" returned no file_path`)
  return data.file_path
}

// Runs the flow with the two uploaded paths and a prompt. Returns the raw response.
export async function runFlow(cfg, { pathA, pathB, prompt, sessionId }) {
  const { apiBase, apiKey, flowId, fileIdA, fileIdB } = cfg
  const res = await fetch(`${trimBase(apiBase)}/api/v1/run/${flowId}?stream=false`, {
    method: 'POST',
    headers: { 'Content-Type': 'application/json', 'x-api-key': apiKey },
    body: JSON.stringify({
      output_type: 'chat',
      input_type: 'text',
      input_value: prompt,
      session_id: sessionId,
      tweaks: {
        [fileIdA]: { path: [pathA] },
        [fileIdB]: { path: [pathB] },
      },
    }),
  })
  if (!res.ok) throw new Error(`Analysis failed: ${await readError(res)}`)
  return res.json()
}

Why the proxy. LangFlow's upload endpoint does not answer the browser's CORS preflight (OPTIONS returns 422), so a direct upload from the page fails with "Failed to fetch". The UI keeps apiBase blank and lets Vite forward /api:

flaky_test_analyzer_ai_Agent/ui/vite.config.js
import { defineConfig } from 'vite'
import react from '@vitejs/plugin-react'

// LangFlow's file-upload endpoint does NOT answer the browser's CORS preflight
// (OPTIONS -> 422), so calling it cross-origin from the browser fails with
// "Failed to fetch". We sidestep CORS entirely by proxying same-origin /api
// requests through Vite to the LangFlow server. Override the target with
// LANGFLOW_URL if LangFlow runs elsewhere.
const LANGFLOW_URL = process.env.LANGFLOW_URL || 'http://localhost:7861'

export default defineConfig({
  plugins: [react()],
  server: {
    port: 5173,
    strictPort: false,
    open: false,
    proxy: {
      '/api': {
        target: LANGFLOW_URL,
        changeOrigin: true,
      },
    },
  },
})

Auth. Since LangFlow 1.5 the /run endpoint needs an x-api-key even with auto-login. Create one under Settings, API Keys, or do what test_flow.py does:

flaky_test_analyzer_ai_Agent/flow/test_flow.py
def bootstrap():
    """Auto-login, then mint an API key (run endpoints require one since v1.5)."""
    st, out = _call("/api/v1/auto_login")
    if st != 200:
        sys.exit(f"Langflow not reachable at {BASE} (auto_login -> {st})")
    bearer = {"Authorization": "Bearer " + out["access_token"]}
    st, out = _call("/api/v1/api_key/", {"name": "flaky-flow-test"}, headers=bearer)
    if st not in (200, 201):
        sys.exit(f"could not create api key: {st} {out}")
    return bearer, (out.get("api_key") or out.get("key"))


def run_flow(flow_id, key, folder):
    st, out = _call(
        f"/api/v1/run/{flow_id}?stream=false",
        {"output_type": "chat", "input_type": "chat", "input_value": folder},
        headers={"x-api-key": key},
    )
    if st != 200:
        return st, str(out)
    try:
        return 200, out["outputs"][0]["outputs"][0]["results"]["message"]["text"]
    except (KeyError, IndexError, TypeError):
        return 200, json.dumps(out)

The same call from a terminal, from the flow README:

terminal
curl -s -X POST "http://localhost:7860/api/v1/run/<FLOW_ID>?stream=false" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $LANGFLOW_API_KEY" \
  -d '{"output_type":"chat","input_type":"chat",
       "input_value":"/test-results/flaky_test_analyzer_ai_Agent"}'

05Run LangFlow locally

Two scripts run LangFlow in Docker on macOS with Docker Desktop. langflow-up.sh starts Docker, reuses or creates a container named langflow from the pinned image, and polls /health every 3 seconds until it answers 200.

langflow-up.sh
# Create container if missing (first run / after prune), else just start it.
if "$DOCKER" ps -a --format '{{.Names}}' | grep -qx "$NAME"; then
  echo "==> Starting existing '$NAME' container..."
  "$DOCKER" start "$NAME" >/dev/null
else
  echo "==> Container '$NAME' not found. Creating with persistent volume..."
  "$DOCKER" run -d --name "$NAME" \
    -p 7860:7860 \
    -v "$DATA":/app/langflow-data \
    -v "$CHAPTER":/test-results:ro \
    -e LANGFLOW_CONFIG_DIR=/app/langflow-data \
    -e LANGFLOW_SAVE_DB_IN_CONFIG_DIR=true \
    -e LANGFLOW_AUTO_LOGIN=true \
    "$IMAGE" >/dev/null
fi

echo "==> Waiting for Langflow to be ready..."
for i in $(seq 1 60); do
  code=$(curl -s -o /dev/null -w '%{http_code}' "$URL/health" 2>/dev/null || echo 000)
  if [ "$code" = "200" ]; then echo "    READY (~$((i*3))s)"; break; fi
  sleep 3
done

echo "==> Langflow: $URL"
"$DOCKER" ps --filter "name=$NAME" --format '    {{.Names}} | {{.Status}} | {{.Ports}}'
langflow-down.sh
#!/usr/bin/env bash
# Stop Langflow container and quit Docker Desktop.
set -euo pipefail

DOCKER="/Applications/Docker.app/Contents/Resources/bin/docker"
NAME="langflow"

if "$DOCKER" info >/dev/null 2>&1; then
  echo "==> Stopping '$NAME'..."
  "$DOCKER" stop "$NAME" >/dev/null 2>&1 || echo "    (not running)"
  echo "==> Quitting Docker Desktop..."
  osascript -e 'quit app "Docker Desktop"' >/dev/null 2>&1 || true
  sleep 3
fi

if pgrep -f "Docker Desktop" >/dev/null; then
  echo "    Docker still running (force-quit if needed)."
else
  echo "==> Docker Desktop quit. All stopped."
fi
Edit two paths first. DATA and CHAPTER at the top of langflow-up.sh are absolute paths on the author's machine. Point them at your own checkout: DATA holds the LangFlow database, CHAPTER is mounted read-only at /test-results so flows can read the result files.

What the container settings are for, from the chapter's learning notes:

  • Persistence. A host bind mount plus LANGFLOW_CONFIG_DIR and LANGFLOW_SAVE_DB_IN_CONFIG_DIR=true. Without the last one the SQLite database stays inside the image and your flows vanish with the container.
  • Auto-login. LangFlow 1.12 switched the auto-login default off; the script sets LANGFLOW_AUTO_LOGIN=true explicitly.
  • A pinned image. langflowai/langflow:1.12.3, not latest: a new version migrates the database in place on boot, so back up the data folder before you change the tag.
terminal
./chapter_05_AI_Agents_LangFlow/langflow-up.sh        # prints http://localhost:7860 when ready

# the deterministic flow: import flow/Flaky_Test_Analyzer.json in the UI, or test it end to end
cd chapter_05_AI_Agents_LangFlow/flaky_test_analyzer_ai_Agent/flow
python3 test_flow.py                                   # needs LangFlow on :7860

# the React UI for the LLM flow (Node.js 20+); the proxy defaults to port 7861
cd ../ui && npm install
LANGFLOW_URL=http://localhost:7860 npm run dev          # http://localhost:5173

Keys by name only: the model nodes read the LangFlow global variables GROQ_API_KEY and OPENROUTER_API_KEY; the UI reads its LangFlow key from the Connection panel or VITE_API_KEY; the curl example uses LANGFLOW_API_KEY.

06Bug triage and the API contract check

Bug triage flow

AI3X_003_Bug_Triage_AI_Agent.json is a straight line: an API Request component GETs one Jira issue over REST, a Parser turns the response into text (pattern {result}), a Prompt Template drops it into {issue}, and OpenRouter (deepseek/deepseek-v4-flash, temperature 0.7) answers in a Chat Output. The prompt asks a senior triage engineer for five things: severity (Blocker to Trivial), priority (P0 to P4), impact areas, a root-cause hypothesis and a one- or two-sentence justification, based only on the issue and without inventing logs or stack traces.

Two notes before you reuse it. The flow's own note says it uses Groq and returns strict JSON; the wired model is OpenRouter and the prompt never asks for JSON, so expect labelled prose. The issue key is fixed in the request URL: feed it from a Chat Input or a tweak instead, and read the Jira token from a LangFlow global variable the way the model nodes read theirs, because an exported flow carries every field value with it.

API contract validator (spec only)

Project/AI3X_004_API_Contract_Validator.md describes a flow you build yourself: an API Request component calls GET https://gorest.co.in/public/v2/users, and an OpenRouter model (DeepSeek V4 Flash) compares the live response with a JSON Schema and reports drift: missing fields, wrong types, extra keys. The README's expected verdict is a PASS: all 10 objects have an integer id and string name, email, gender and status.

Project/AI3X_004_API_Contract_Validator.md
{
  "$schema": "http://json-schema.org/draft-04/schema#",
  "type": "array",
  "items": [
    {
      "type": "object",
      "properties": {
        "id": {
          "type": "integer"
        },
        "name": {
          "type": "string"
        },
        "email": {
          "type": "string"
        },
        "gender": {
          "type": "string"
        },
        "status": {
          "type": "string"
        }
      },
      "required": [
        "id",
        "name",
        "email",
        "gender",
        "status"
      ]
    },

The start of the spec's schema: the first of ten identical item schemas.

A schema gotcha worth a test. The spec writes "items": [ ... ], the tuple form: one schema per array position, so it only describes the first 10 users and says nothing about an eleventh. The README shows "items": { ... }, one schema for every item. Use the README form for a list endpoint.

07LangFlow, LangGraph or LangSmith?

LangFlow vs LangGraph vs LangSmith.md compares the three tools on fourteen dimensions. In short: LangFlow builds it visually, LangGraph builds it in code with full control, LangSmith watches, debugs and grades whatever you built. They stack rather than compete.

LangFlow (this chapter)LangGraph (chapter 18)LangSmith
What it isA visual builder; every flow is also a REST APIA code framework for stateful graphsTracing and evaluation for any LLM app
You work inA canvas of components and edgesPython or JavaScriptA web dashboard plus an SDK
Best atBuilding and changing a flow quicklyLoops, branches, retries, human approvalKnowing why an app behaved the way it did
Weak atDeep custom logic, version controlQuick prototypesBuilding anything: it only observes
For a testerReproduce a flow visually before writing the evalDeterministic seams to test againstEval datasets are the test suite
The chapter 5 LangFlow flow (Chat Input, Flaky Analyzer, Chat Output) above the chapter 18 LangGraph version, which adds a clean or flaky branch, a human approval and an LLM explanation
Same answer, two tools. Chapter 18 rebuilds this analyzer as a graph that can branch, wait for a person and explain.

Next step: LangGraph lesson 10 keeps the count in plain Python, pauses for approval before quarantining, and only then lets a model add notes.

08What to watch

  • 7860 or 7861. langflow-up.sh publishes port 7860, but the UI's proxy defaults to 7861. Start the UI with LANGFLOW_URL=http://localhost:7860.
  • The UI's instruction reaches no component. The exported LLM flow has no Chat Input, so the input_value the UI sends is not wired anywhere; the Prompt Template's fixed text drives the analysis.
  • Three sections or four. PROMPTS.md says the agent returns three sections; the template asks for four, adding SUMMARY.
  • The container sees only what you mount. Paths must be under /test-results; anything else fails with Folder not found inside the Langflow container. The component remaps one host path to that mount.
  • Match on the title, not the key. Report keys repeat the file name (loginTests/auth.spec.ts > loginTests/auth.spec.ts > ...), unlike the shorter form in the flow README.
  • A skip is a flip. Any verdict change counts, so [passed -> skipped] is flaky under this rule. Decide whether that is what your team wants before the number gates CI.
  • RAW_STATS is Playwright's count. The component echoes each file's stats block; it does not recompute it.
  • Test the exported file. test_flow.py imports the JSON fresh, runs eight cases and deletes the copy, so it proves the file you ship, not whatever is loaded in the UI.