The Testing Academy · Class Notes Saturday, 8 August (IST)
Live class · study guide

Python sets, dictionaries, map and filter, and the start of OOP

Why a set silently drops your duplicates, the difference between map and filter, the dictionary trick that solves half of all string interview questions, and the class-and-object model the AI frameworks are written in.

By Pramod Dutta, The Testing Academy. Study notes from the live Python class, rebuilt from the session recording. Code is reproduced as it was written on screen, with the output the class saw.

01

From the class: Tarun's offer

The session opened with news from Tarun. After almost a year without a job, despite 18 years of experience, he accepted an offer and joined this week.

His message, read out in class: the prompting skills, framework creation, MCP, RAG and LangChain covered in the course are what let him stand out and speak confidently in the interviews.

Tarun has agreed to share his resume with the batch, with personal details removed, so everyone can see what a profile that gets selected actually looks like.

Worth holding on to while you work through the rest of this page. The Python below is not the destination, it is what the DeepEval, LangChain and CrewAI work is built on.

02

Sets: unique, unordered, unindexed

A set is an iterable, mutable collection that refuses duplicates. Three properties decide everything else about it:

  • No duplicates. Add the same value twice and the set keeps one.
  • Unordered. Elements come back in whatever order the set feels like, and that order can change as the set grows.
  • Unindexed. There is no s[0]. Lists and tuples have positions, sets do not.
Python
my_set = {1, 2, 3, 3, 4, 5}
print(my_set)          # {1, 2, 3, 4, 5}   the duplicate 3 is gone
print(len(my_set))     # 5
print(type(my_set))    # <class 'set'>

Brackets matter and the class spelled them out: () for a tuple, [] for a list, {} for a set.

Sets can be built from other collections, which is the fastest way to deduplicate anything:

Python
set([1, 1, 2, 3])                                 # {1, 2, 3}
set(("The Testing Academy", "The Testing Academy"))  # one element, not two

Mixed types are allowed. One result surprises people:

Python
print({1, True, "qa", 2.5})   # True vanishes

True equals 1 in Python, so a set that already holds 1 treats True as the same element. False and 0 collide the same way. This turns up as an interview question specifically because it looks wrong.

An empty set is written set(), not {}. Empty curly braces create an empty dictionary, which is a different type. It is the one place the brace notation does not do what you expect.

A frozenset is a set that cannot be changed after it is created. Same behaviour, no add or remove.

Set comprehensions work like list comprehensions:

Python
squares = {x**2 for x in range(5)}
print(squares)   # {0, 1, 4, 9, 16}

Read it as a loop turned inside out: run x from 0 to 4, square each one, collect the results into a set. The class read x**2 aloud as "x into 2", but the values on screen (0, 1, 4, 9, 16) are squares, so the operator is the power operator.

03

Set operations: union, intersection, difference

The reason sets exist. Given two sets:

Python
a = {1, 2, 3}
b = {3, 4, 5}

a | b    # {1, 2, 3, 4, 5}   union, everything, 3 counted once
a & b    # {3}               intersection, only what both have
a - b    # {1, 2}            difference, a without anything in b
b - a    # {4, 5}            order matters here

The named methods a.union(b), a.intersection(b) and a.difference(b) do exactly the same thing, and turn up less often in real code than the operators.

Operator Name Answers the question
\| union what appears in either?
& intersection what appears in both?
- difference what is in the first and not the second?
04

Interview drill: first non-repeating character

Problem: given a string from the user, return the first character that appears exactly once. For "swiss", the answer is w.

Python
def first_non_repeating(text):
    seen = set()
    for ch in text:
        if text.count(ch) == 1:
            seen.add(ch)
            return ch
    return None

print(first_non_repeating("swiss"))   # w

The walkthrough, one character per pass:

ch text.count(ch) Action
s 3 repeats, skip
w 1 appears once, return w
i 1 never reached, the function already returned
s 3 never reached
s 3 never reached

count() scans the whole string for each character, and return exits the moment the first match is found, which is why the later characters are never examined.

Today's task: extend this to return all non-repeating characters, not just the first. The set starts earning its keep there, because you collect instead of returning, and the loop condition has to change.

05

filter and map

Both take a function and a collection. The difference is what comes back.

filter keeps the elements where the function returns True, so the result is smaller or equal:

Python
numbers = [1, 2, 3, 4, 5, 6]

def is_even(x):
    return x % 2 == 0

print(list(filter(is_even, numbers)))   # [2, 4, 6]

Element by element: 1 gives False and is dropped, 2 gives True and is kept, 3 dropped, 4 kept, 5 dropped, 6 kept.

A lambda is a one-line function, and it fits filter perfectly:

Python
results = ["pass", "fail", "pass", "skip"]
print(list(filter(lambda r: r == "pass", results)))   # ['pass', 'pass']

names = ["QA", "", "Automation Tester", ""]
print(list(filter(lambda n: n != "", names)))   # ['QA', 'Automation Tester']

map applies the function to every element and returns the same number of items, transformed:

Python
print(list(map(lambda x: x**2, [1, 2, 3, 4, 5])))     # [1, 4, 9, 16, 25]
print(list(map(str.upper, ["qa", "sdet"])))           # ['QA', 'SDET']

response_times = [1200, 1500, 1800]
print(list(map(lambda ms: ms / 1000, response_times)))   # [1.2, 1.5, 1.8]
filter map
What the function returns True or False a new value
Size of the result reduced unchanged
Use it to refine a list transform a list

These are the same ideas as Java 8 streams and the JavaScript array methods. If you know stream().filter() or array.map(), you already know this.

06

Dictionaries

A dictionary is the key-value structure, the same shape as JSON and the same job as a Java Map.

Python
person = {"name": "Pramod", "age": 34, "role": "SDET"}

person["age"]          # 34
person["city"] = "Delhi"     # add
del person["role"]           # delete
"name" in person             # True

for key, value in person.items():
    print(key, value)

Three behaviours that come up in interviews:

Python
# 1. duplicate keys: the last one wins, silently, no error
d = {"name": "Pramod", "age": 65, "age": 67}
print(d["age"])   # 67

# 2. order does not affect equality
{"a": 1, "b": 2} == {"b": 2, "a": 1}   # True

# 3. zip pairs two lists, ignoring anything unmatched on either side
keys = ["name", "role", "experience"]
values = ["Pramod", "SDET", 3, 90]
print(dict(zip(keys, values)))   # {'name': 'Pramod', 'role': 'SDET', 'experience': 3}

Merging two dictionaries uses the same pipe as set union:

Python
merged = dict1 | dict2

Dictionaries nest freely, which is how API responses arrive:

Python
students[0]["address"]["office"]   # reach into a dict inside a list
07

Interview drills: character frequency and vowel count

Count how often each character appears. The whole trick is dict.get(key, default), which returns the stored value if the key exists and the default if it does not:

Python
text = "automation"
char_count = {}

for ch in text:
    char_count[ch] = char_count.get(ch, 0) + 1

print(char_count)
# {'a': 2, 'u': 1, 't': 2, 'o': 2, 'm': 1, 'i': 1, 'n': 1}

The first few passes:

ch Already a key? get(ch, 0) New value
a no 0 1
u no 0 1
t no 0 1
o no 0 1
m no 0 1
a yes 1 2
t yes 1 2

Count the vowels, and collect them:

Python
vowels = "aeiou"
count = 0
found = []

for ch in "hello world":
    if ch in vowels:
        count += 1
        found.append(ch)

print(count)   # 3
print(found)   # ['e', 'o', 'o']

Duplicates are counted, because the question asks how many vowels appear, not how many distinct ones.

When a loop stops making sense, draw the table. Write one row per pass with the variable values after that pass, the same expression-and-result table used in the Playwright batch. It converts a confusing loop into arithmetic you can check.

08

OOP: classes and objects

The reason this matters right now: CrewAI, LangChain and the DeepEval framework are all written in Python OOP. You do not need to write those libraries, but reading roughly half of their source is what separates using them from guessing at them.

Before object orientation there was procedural programming, the C style where a program is a pile of functions and the data lives apart from them. It emphasised doing things, and it did not scale.

Object orientation puts data and behaviour together:

  • A class is a blueprint. It has attributes (also called properties or data members) and behaviour (methods).
  • An object is a real entity built from that blueprint, an instance of the class.
Python
class Person:
    name = None
    age = None

    def eat(self):
        print("eating")

geeta = Person()    # object
amit = Person()     # a different object from the same class

The analogies from class, in order: a building blueprint versus the actual buildings put up from it; a blueprint for a person versus Omkar, Sindhuja and Amit, who all differ; a Dog class versus a Mastiff, a Maltese and a Chow Chow. And the closest one to home: AI Tester Blueprint is the class, every student in the batch is an object.

Two Python specifics:

  • No curly braces. Indentation defines what belongs to the class. A function written back at the left margin is outside the class, even if it sits directly underneath it.
  • An object is created by calling the class, Person(), with no new keyword.

Constructors, instance variables, encapsulation, inheritance, polymorphism, abstraction, static members, super, overriding and overloading, and whether Python has interfaces, were all named as the rest of the OOP series and were not covered here. They start in the next class.

09

Tasks and announcements

  • Today's task: extend the non-repeating-character program to find all non-repeating characters.
  • Python fundamentals test: submit your screenshots. It can be retaken.
  • 15 August: holiday. 16 August: hackathon, running like the tests do. DeepEval is not required for it.
  • Remaining syllabus: finish Python, then the DeepEval framework, LangChain and CrewAI, plus MCP creation in Python.
  • Extra classes: Claude 101 part two on Tuesday evening, which completes the certification, then Cursor and Codex masterclasses.
  • Today's code is pushed. Questions go in the doubt thread.