Why open-weight models exist
Closed models never tell you what data they were trained on or what weights they ended up with. Open models publish four things: the model weights, the source code needed to run them, the training data, and a license.
That transparency is the point, and the second-order effect matters just as much: a market with credible free models keeps the price of the paid ones down. Concentrating every capable model in one or two vendors is the failure mode to avoid.
The families named in class:
| Family | From |
|---|---|
| Llama | Meta |
| Gemma | |
| Qwen | Alibaba |
| DeepSeek | DeepSeek |
| Mistral | a European company, not a Chinese one |
| Kimi | Moonshot |
| Nemotron | Nvidia |
Two clarifications from the floor. Ollama is a tool for running models, not a model family. And Perplexity is a curation layer that sends your question to other providers' models, so it does not belong on a list of model families either.
The class also flagged that Sarvam is not yet an independent large language model of its own, and was blunt about India not being in the frontier-model race today, which is the argument for learning to use these models well.
The use case that pays for a local model
This is the section to remember, because it is the one you will actually deploy at work.
An AI chatbot under test does not return a fixed string. Ask it for the status of order 123 twice and you get two differently worded answers. There is nothing to assert against.
The fix is a second, much smaller model used purely as an extractor:
AI chatbot answer (varies) -> small LLM -> {"order_number": "123", "status": "delivered"}
structured, fixed, assertable
The chatbot keeps its expensive model. The extractor does one narrow job, so it can run on something tiny. Your test then asserts on JSON fields instead of on prose.
This is what the class called using an LLM as an extractor, and it generalises: any time an AI feature returns free text that a test needs to check, put a cheap model between the answer and the assertion.
The cost math that forces the decision
The extraction step runs on every test, every run, which is where a premium model stops being affordable. The figures below are the class's own back-of-envelope illustration, not a price list:
| Scenario | Rough monthly cost |
|---|---|
| 12,000 test cases, 5 model calls each, on a premium model | around 1,000 US dollars per day |
| 5,000 test cases a day for 30 days, frontier model | around 180,000 US dollars |
| Same volume, a 20-billion-parameter open model | a few thousand rupees |
| Same volume, a small hosted open model | around 150 US dollars |
| Self-hosted on your own machine | hardware only, paid once |
The conclusion the class drew: generation can use a premium model, execution cannot. Generating a test suite once with a frontier model is fine. Calling one on every assertion of every run is not, at any realistic scale.
Local vs cloud, honestly
| Local open model | Cloud model | |
|---|---|---|
| Where it runs | your laptop or a company server | the provider's servers |
| Internet | not required | required |
| Data | stays in your environment | leaves your environment |
| Setup | install plus a model download | an API key |
| Ongoing cost | hardware and electricity | per token |
| Control | full, including fine-tuning | limited |
| Speed and quality | slower, smaller | faster, stronger |
The disadvantages are real and were stated plainly: local models are slow, they want serious RAM, and their knowledge is narrower and updated less often.
Fine-tuning came up as the customisation route: teaching a small model one skill it needs for your project, rather than reaching for a bigger general model.
Setting up Ollama, and sizing the model to your laptop
Ollama runs open models locally, and it also lists hosted models in the same dropdown. Entries marked cloud are not running on your machine and need a paid plan, so ignore them for this exercise. LM Studio does the same job with an easier interface and is the fallback if Ollama misbehaves. llama.cpp also exists.
On a work laptop, get approval before installing any of this. That was stated twice and it is not optional.
Sizing is the part people get wrong. Stay at or under roughly 1 to 3 billion parameters on a normal laptop:
| Model | Verdict on 8 or 16 GB RAM |
|---|---|
| Llama 3.2, 1B (about 1.3 GB) | yes, the safest starting point |
| Gemma 3, 1B | yes, and noticeably fast |
| DeepSeek R1, 1.5B | yes |
| Code Llama, small variants | yes |
| Qwen, 4B | works, slower |
| Anything 7B, 14B, 20B, 26B | no, the machine will hang |
Install from the Ollama app, or from the terminal:
ollama run llama3.2:1b # first run downloads the weights, then chats locally
Turn the wifi off and ask it something. It still answers, which is the whole point of local.
Same prompt, small local model versus a frontier cloud model, produced visibly different quality. The class's framing: you are comparing a young child with someone holding a PhD. Keep the task small and the small model does fine.
Cloud-hosted open models: Groq and OpenRouter
If your machine cannot run anything locally, the middle path is a provider that hosts open models cheaply:
- Groq (groq.com, not Elon Musk's Grok) serves open models and gives roughly 1 million free tokens a day, which is plenty for classwork.
- OpenRouter does the same across many providers, also with free models available.
Both are reached with an API key from a free account, and both are pay-as-you-go once the free allowance runs out.
Project: a Jira test case generator
The build: a local application that takes a Jira ID, fetches the requirement, and generates test cases from it, using a local model with a cloud fallback.
Prerequisites: a requirement to work from and Jira access, which means three values: the Jira URL (yourname.atlassian.net), the email on the account, and an API token created from the Atlassian API tokens page. A free Jira trial is enough. Azure DevOps works too, since it also exposes an API.
The build order mattered more than the code:
- Sketch the app first. A rough drawing of the two screens: a main screen with an input and a send button, and a settings screen holding the Jira email, token and URL, the Ollama URL, and the Groq token, with a save button. The sketch was saved into the repository as an image.
- Write the prompt, then improve the prompt. The first draft went into a file, then was handed back to a chat model with the instruction to rewrite it in RICEPOT form. The refined version carried a role, context, parameters and tonality, and that was the version handed to the coding agent.
- Plan before building. The refined prompt plus the sketch went to a coding agent in plan mode, which produced a
plan.md. Only after the plan looked right was the agent switched to agent mode to write files. - Let it build, then verify the connections. The agent was explicitly asked to confirm that both the Ollama and the Groq connections worked, not just that the code existed.
The resulting stack: Python with Streamlit for the interface (Streamlit is a Python front end, so no separate web stack), a Jira client hitting the REST API, and an LLM client that tries Ollama first and falls back to Groq.
.env
JIRA_URL=https://yourname.atlassian.net
JIRA_EMAIL=you@example.com
JIRA_API_TOKEN=...
OLLAMA_URL=http://localhost:11434
OLLAMA_MODEL=gemma3:1b
GROQ_TOKEN=...
The .env file holds every credential and is never committed. The Ollama URL is plain HTTP because it is local.
Running it: the settings screen has test buttons for each connection, green meaning good. Enter a Jira ID on the main screen and the app fetches that issue and returns test cases. Switching the brain from Ollama to Groq is a settings change, nothing else.
Formatting of the generated output was poor on the first pass. The fix was to ask the agent to improve it, which it did, testing its own output as it went. Cosmetic problems are worth deferring until the pipeline works end to end.
Jira is reached here through its REST API rather than through an MCP server. MCP is coming in a later session and will replace this plumbing.
Why build it at all when Atlassian's own assistant can draft test cases inside Jira? Because that assistant only works inside Jira and bills you for it. An independent agent is not locked to a vendor, and the next agents on the syllabus, flaky-test analysis and root cause analysis, cannot be built inside a single tool's walls.
The gap this build leaves, and BLAST
The honest closing question of the class: the app works, but do you know how it was built? Which file does what, how the front end reaches the back end, what the architecture is. If the answer is no, you are a blind vibe-coder shipping AI slop.
That is the motivation for the BLAST framework, which is not a prompting framework but a way to keep control of an application you build with AI:
| Letter | What it forces you to know |
|---|---|
| B | Blueprint: what the application actually is |
| L | Linking: how the front end and back end connect, and what information flows |
| A | Architecture: the core structure |
| S | Style: the UI you chose |
| T | Trigger and deploy: how it runs and ships |
BLAST itself, plus deployment (a VPS, or Vercel for free hosting), is the next session. It was named and deliberately deferred here.
Tasks and announcements
Four tasks, in order of urgency:
- Finish the overdue RICEPOT task by end of day: a test plan, test cases and a bug report, pushed to your GitHub repository.
- Keep your prompt templates in a templates folder: STLC, Playwright, Selenium and API. This is a hard prerequisite, because next week's skill files are built from these templates. If the templates do not exist, the skill work cannot start.
- Install Ollama (or LM Studio) with a model of 3 billion parameters or fewer, and generate at least 25 test cases from the shared requirement using it.
- Replicate the Jira test case generator in your own repository. You were not expected to build it live, only to follow, then rebuild it afterwards.
Schedule:
- Tuesday, 8:00 PM IST: Claude 101 part two, with a doubt session alongside.
- Skills masterclass: Wednesday 12 August, 8:00 PM IST (moved earlier from 9:00 PM), mandatory for the 4x batch: going from prompt to skill file. Announced in this class for Thursday, then moved to the Wednesday.
- Friday: no class. 15 August is a holiday.
- Templates and today's code are being pushed to the repository.