The Testing Academy · Class Notes Sunday, 9 August (IST)
Live class · study guide

Local LLMs for QA, and building a Jira test case generator

Why open-weight models exist, the one QA use case that makes a local LLM pay for itself, how to pick a model your laptop can actually run, and building an independent Jira test case generator with Ollama and a Groq fallback.

By Pramod Dutta, The Testing Academy. Study notes from the live AI Tester Blueprint 4x class, rebuilt from the session recording. Commands and configuration are reproduced as they were shown on screen.

01

Why open-weight models exist

Closed models never tell you what data they were trained on or what weights they ended up with. Open models publish four things: the model weights, the source code needed to run them, the training data, and a license.

That transparency is the point, and the second-order effect matters just as much: a market with credible free models keeps the price of the paid ones down. Concentrating every capable model in one or two vendors is the failure mode to avoid.

The families named in class:

Family From
Llama Meta
Gemma Google
Qwen Alibaba
DeepSeek DeepSeek
Mistral a European company, not a Chinese one
Kimi Moonshot
Nemotron Nvidia

Two clarifications from the floor. Ollama is a tool for running models, not a model family. And Perplexity is a curation layer that sends your question to other providers' models, so it does not belong on a list of model families either.

The class also flagged that Sarvam is not yet an independent large language model of its own, and was blunt about India not being in the frontier-model race today, which is the argument for learning to use these models well.

02

The use case that pays for a local model

This is the section to remember, because it is the one you will actually deploy at work.

An AI chatbot under test does not return a fixed string. Ask it for the status of order 123 twice and you get two differently worded answers. There is nothing to assert against.

The fix is a second, much smaller model used purely as an extractor:

Text
AI chatbot answer (varies)  ->  small LLM  ->  {"order_number": "123", "status": "delivered"}
                                              structured, fixed, assertable

The chatbot keeps its expensive model. The extractor does one narrow job, so it can run on something tiny. Your test then asserts on JSON fields instead of on prose.

This is what the class called using an LLM as an extractor, and it generalises: any time an AI feature returns free text that a test needs to check, put a cheap model between the answer and the assertion.

03

The cost math that forces the decision

The extraction step runs on every test, every run, which is where a premium model stops being affordable. The figures below are the class's own back-of-envelope illustration, not a price list:

Scenario Rough monthly cost
12,000 test cases, 5 model calls each, on a premium model around 1,000 US dollars per day
5,000 test cases a day for 30 days, frontier model around 180,000 US dollars
Same volume, a 20-billion-parameter open model a few thousand rupees
Same volume, a small hosted open model around 150 US dollars
Self-hosted on your own machine hardware only, paid once

The conclusion the class drew: generation can use a premium model, execution cannot. Generating a test suite once with a frontier model is fine. Calling one on every assertion of every run is not, at any realistic scale.

04

Local vs cloud, honestly

Local open model Cloud model
Where it runs your laptop or a company server the provider's servers
Internet not required required
Data stays in your environment leaves your environment
Setup install plus a model download an API key
Ongoing cost hardware and electricity per token
Control full, including fine-tuning limited
Speed and quality slower, smaller faster, stronger

The disadvantages are real and were stated plainly: local models are slow, they want serious RAM, and their knowledge is narrower and updated less often.

Fine-tuning came up as the customisation route: teaching a small model one skill it needs for your project, rather than reaching for a bigger general model.

05

Setting up Ollama, and sizing the model to your laptop

Ollama runs open models locally, and it also lists hosted models in the same dropdown. Entries marked cloud are not running on your machine and need a paid plan, so ignore them for this exercise. LM Studio does the same job with an easier interface and is the fallback if Ollama misbehaves. llama.cpp also exists.

On a work laptop, get approval before installing any of this. That was stated twice and it is not optional.

Sizing is the part people get wrong. Stay at or under roughly 1 to 3 billion parameters on a normal laptop:

Model Verdict on 8 or 16 GB RAM
Llama 3.2, 1B (about 1.3 GB) yes, the safest starting point
Gemma 3, 1B yes, and noticeably fast
DeepSeek R1, 1.5B yes
Code Llama, small variants yes
Qwen, 4B works, slower
Anything 7B, 14B, 20B, 26B no, the machine will hang

Install from the Ollama app, or from the terminal:

Terminal
ollama run llama3.2:1b   # first run downloads the weights, then chats locally

Turn the wifi off and ask it something. It still answers, which is the whole point of local.

Same prompt, small local model versus a frontier cloud model, produced visibly different quality. The class's framing: you are comparing a young child with someone holding a PhD. Keep the task small and the small model does fine.

06

Cloud-hosted open models: Groq and OpenRouter

If your machine cannot run anything locally, the middle path is a provider that hosts open models cheaply:

  • Groq (groq.com, not Elon Musk's Grok) serves open models and gives roughly 1 million free tokens a day, which is plenty for classwork.
  • OpenRouter does the same across many providers, also with free models available.

Both are reached with an API key from a free account, and both are pay-as-you-go once the free allowance runs out.

07

Project: a Jira test case generator

The build: a local application that takes a Jira ID, fetches the requirement, and generates test cases from it, using a local model with a cloud fallback.

Prerequisites: a requirement to work from and Jira access, which means three values: the Jira URL (yourname.atlassian.net), the email on the account, and an API token created from the Atlassian API tokens page. A free Jira trial is enough. Azure DevOps works too, since it also exposes an API.

The build order mattered more than the code:

  1. Sketch the app first. A rough drawing of the two screens: a main screen with an input and a send button, and a settings screen holding the Jira email, token and URL, the Ollama URL, and the Groq token, with a save button. The sketch was saved into the repository as an image.
  2. Write the prompt, then improve the prompt. The first draft went into a file, then was handed back to a chat model with the instruction to rewrite it in RICEPOT form. The refined version carried a role, context, parameters and tonality, and that was the version handed to the coding agent.
  3. Plan before building. The refined prompt plus the sketch went to a coding agent in plan mode, which produced a plan.md. Only after the plan looked right was the agent switched to agent mode to write files.
  4. Let it build, then verify the connections. The agent was explicitly asked to confirm that both the Ollama and the Groq connections worked, not just that the code existed.

The resulting stack: Python with Streamlit for the interface (Streamlit is a Python front end, so no separate web stack), a Jira client hitting the REST API, and an LLM client that tries Ollama first and falls back to Groq.

Text
.env
  JIRA_URL=https://yourname.atlassian.net
  JIRA_EMAIL=you@example.com
  JIRA_API_TOKEN=...
  OLLAMA_URL=http://localhost:11434
  OLLAMA_MODEL=gemma3:1b
  GROQ_TOKEN=...

The .env file holds every credential and is never committed. The Ollama URL is plain HTTP because it is local.

Running it: the settings screen has test buttons for each connection, green meaning good. Enter a Jira ID on the main screen and the app fetches that issue and returns test cases. Switching the brain from Ollama to Groq is a settings change, nothing else.

Formatting of the generated output was poor on the first pass. The fix was to ask the agent to improve it, which it did, testing its own output as it went. Cosmetic problems are worth deferring until the pipeline works end to end.

Jira is reached here through its REST API rather than through an MCP server. MCP is coming in a later session and will replace this plumbing.

Why build it at all when Atlassian's own assistant can draft test cases inside Jira? Because that assistant only works inside Jira and bills you for it. An independent agent is not locked to a vendor, and the next agents on the syllabus, flaky-test analysis and root cause analysis, cannot be built inside a single tool's walls.

08

The gap this build leaves, and BLAST

The honest closing question of the class: the app works, but do you know how it was built? Which file does what, how the front end reaches the back end, what the architecture is. If the answer is no, you are a blind vibe-coder shipping AI slop.

That is the motivation for the BLAST framework, which is not a prompting framework but a way to keep control of an application you build with AI:

Letter What it forces you to know
B Blueprint: what the application actually is
L Linking: how the front end and back end connect, and what information flows
A Architecture: the core structure
S Style: the UI you chose
T Trigger and deploy: how it runs and ships

BLAST itself, plus deployment (a VPS, or Vercel for free hosting), is the next session. It was named and deliberately deferred here.

09

Tasks and announcements

Four tasks, in order of urgency:

  1. Finish the overdue RICEPOT task by end of day: a test plan, test cases and a bug report, pushed to your GitHub repository.
  2. Keep your prompt templates in a templates folder: STLC, Playwright, Selenium and API. This is a hard prerequisite, because next week's skill files are built from these templates. If the templates do not exist, the skill work cannot start.
  3. Install Ollama (or LM Studio) with a model of 3 billion parameters or fewer, and generate at least 25 test cases from the shared requirement using it.
  4. Replicate the Jira test case generator in your own repository. You were not expected to build it live, only to follow, then rebuild it afterwards.

Schedule:

  • Tuesday, 8:00 PM IST: Claude 101 part two, with a doubt session alongside.
  • Skills masterclass: Wednesday 12 August, 8:00 PM IST (moved earlier from 9:00 PM), mandatory for the 4x batch: going from prompt to skill file. Announced in this class for Thursday, then moved to the Wednesday.
  • Friday: no class. 15 August is a holiday.
  • Templates and today's code are being pushed to the repository.