This notebook assembles the everyday machinery of Python: the built-in containers (list, tuple, dict), the statements that drive decisions and repetition, and the functions that keep code organised and reusable.
Subchapter 1.1 handled one measurement at a time. Here we keep a small catalogue of weather stations, each with an elevation and a few daily mean temperatures, carry it through the whole notebook, and close with a generated-code bug that is silent until the function is called a second time.

Figure 1:Sphinx Observatory, Jungfraujoch — the station behind this book’s running example. Photo by Julius Silver, via Wikimedia Commons, licensed under CC BY-SA 4.0.
1.2.1 Data Structures¶
Lists¶
A list is an ordered, mutable sequence — the default container for a series of measurements. We can create lists by separating different items with commas in square brackets: [Item1, Item2, Item3]. Lists come with built-in methods to modify their contents, such as .append() to add an item to the end, .insert() to add at a specific position, and .remove() to delete a specific value.
# a list: ordered, mutable sequence of daily mean temperatures (°C)
readings_celsius = [-2.3, -1.1, 0.4, 1.2, -0.8]
readings_celsius.append(-3.1) # add to the end
readings_celsius.insert(0, -1.9) # add at a specific position
readings_celsius.remove(0.4) # remove a specific value
print(readings_celsius)[-1.9, -2.3, -1.1, 1.2, -0.8, -3.1]
# indexing and slicing: access individual items or a range of items
print(readings_celsius[0], readings_celsius[-1])
print(readings_celsius[1:3]) # a slice: items at index 1 and 2-1.9 -3.1
[-2.3, -1.1]
# get the length of a list
print(len(readings_celsius))6
Tuples¶
A tuple is an immutable sequence. You need a tuple instead of a list when the data represents a fixed entity (like a coordinate pair or a date) where accidental modification of a single element would render the entire record physically invalid.
# a tuple: immutable record, here a (lat, lon) coordinate in degrees
station_coords = (46.5475, 7.9853)
lat, lon = station_coords # unpacking into two names
print(lat, lon)46.5475 7.9853
Dictionaries¶
A dictionary maps unique, immutable keys to values. It is written with braces, as
{key: value} pairs separated by commas:
{
"key_1": value_1,
"key_2": value_2,
}This structure provides very fast lookup by name or id, rather than by sequential position.
station = {
"name": "Jungfraujoch",
"elevation_m": 3571,
"coordinates": (46.5475, 7.9853),
"readings_celsius": [-2.3, -1.1, 0.4, 1.2, -0.8],
}You can look up an element using square brackets, add new key-value pairs by assignment, and remove elements using the del keyword.
# look up an element
print(station["name"])
print(station.get("name", "Warning: 'name' key not found!")) # get with default valueJungfraujoch
Jungfraujoch
# lookup an element that does not exist, with a default value
print(station.get("dates", "Warning: 'dates' key not found!"))Warning: 'dates' key not found!
# add a new value
station["country"] = "Switzerland"
print(station){'name': 'Jungfraujoch', 'elevation_m': 3571, 'coordinates': (46.5475, 7.9853), 'readings_celsius': [-2.3, -1.1, 0.4, 1.2, -0.8], 'country': 'Switzerland'}
# remove an element
del station["coordinates"]
print(station){'name': 'Jungfraujoch', 'elevation_m': 3571, 'readings_celsius': [-2.3, -1.1, 0.4, 1.2, -0.8], 'country': 'Switzerland'}
# .items() yields the key-value pairs
print(station.items())dict_items([('name', 'Jungfraujoch'), ('elevation_m', 3571), ('readings_celsius', [-2.3, -1.1, 0.4, 1.2, -0.8]), ('country', 'Switzerland')])
Going deeper: nested structures and safe lookups
Real datasets are usually nested. Here is a dictionary of station codes, each mapping to a dictionary of attributes.
stations = {
"JFJ": {"name": "Jungfraujoch", "elevation_m": 3571, "readings_celsius": [-2.3, -1.1, 0.4, 1.2, -0.8]},
"BAS": {"name": "Basel-Binningen", "elevation_m": 316, "readings_celsius": [6.1, 7.4, 5.9, 8.2, 6.8]}
}
# keyed lookup
print(stations["JFJ"]["name"])
# safe lookup with a specifiable default using .get()
print(stations.get("ZRH", {}).get("name", "unknown")) # 'ZRH' is absent1.2.2 Control Flow¶
Code normally runs linearly from top to bottom. Control flow allows a program to make decisions based on data values (conditional statements), repeat operations for multiple elements (loop statements), and interrupt or skip operations (control statements).
Decisions: if, elif, else¶
Here, we use conditional statements to combine relational operators and logical operators so that a program can have different information flow according to some conditions. In other words, some code snippets are executed only if some conditions are satisfied.
temp_celsius = 6.1
if temp_celsius < 0.0:
category = "freezing"
elif temp_celsius < 10.0:
category = "cold"
elif temp_celsius < 20.0:
category = "mild"
else:
category = "warm"
print(category)cold
Repetition: for, while¶
We use loop statements to execute code multiple times, for instance, applying an algorithm to every item in a list. In Python, for loops iterate over a sequence, while while loops repeat until a condition is no longer met.
For loops¶
# a second station's data: Basel-Binningen's daily mean temperatures (°C)
basel_readings_celsius = [6.1, 7.4, 5.9, 8.2, 6.8]
for temp in basel_readings_celsius:
print(temp)6.1
7.4
5.9
8.2
6.8
# loop over indices in list
for i in range(len(basel_readings_celsius)):
print(i, basel_readings_celsius[i])0 6.1
1 7.4
2 5.9
3 8.2
4 6.8
# better way: enumerate gives index + value
for i, temp in enumerate(basel_readings_celsius):
print(i, temp)0 6.1
1 7.4
2 5.9
3 8.2
4 6.8
dates = ["2024-01-01", "2024-01-02", "2024-01-03", "2024-01-04", "2024-01-05"]
# zip pairs two sequences
for date, temp in zip(dates, basel_readings_celsius):
print(date, temp)2024-01-01 6.1
2024-01-02 7.4
2024-01-03 5.9
2024-01-04 8.2
2024-01-05 6.8
# .items() yields (key, value) pairs and can be used in a for loop
for key, value in station.items():
print(key, value)name Jungfraujoch
elevation_m 3571
readings_celsius [-2.3, -1.1, 0.4, 1.2, -0.8]
country Switzerland
Going deeper: comprehensions
A comprehension builds a list or dict in a single expression, optionally with a filter. It is a compact replacement for a short loop that only transforms or selects.
readings_celsius = [6.1, 7.4, 5.9, 8.2, 6.8]
# list comprehension: convert each value to kelvin
readings_kelvin = [t + 273.15 for t in readings_celsius]
print(readings_kelvin) # [279.25, 280.55, 279.05, 281.35, 279.95]
# list comprehension with a filter: keep the days above 7 °C
warm = [t for t in readings_celsius if t > 7.0]
print(warm) # [7.4, 8.2]
# dict comprehension: station code -> number of readings, reusing the `stations`
# dict from the box above
n_readings = {code: len(v["readings_celsius"]) for code, v in stations.items()}Read a comprehension as “the expression, for each item, (optionally) if a condition”. Exercise 13 is optional practice with both kinds of comprehension.
While loops¶
countdown = 3 # index variable for the while loop
while countdown > 0:
print("Countdown:", countdown)
countdown = countdown - 1 # decrement the counterCountdown: 3
Countdown: 2
Countdown: 1
# while loop using the length of a list, with a condition inside
day = 0
while day < len(basel_readings_celsius): # loop until the index reaches the length
if basel_readings_celsius[day] > 7.0:
print("warm day:", dates[day], basel_readings_celsius[day])
day += 1warm day: 2024-01-02 7.4
warm day: 2024-01-04 8.2
Break and continue¶
Sometimes you need to interrupt a loop. The break statement exits a loop entirely, whereas continue skips the remainder of the current iteration and jumps to the next one.
# break: stop the loop entirely
hot_temp = None # define it first, in case no hot day is found
day = 0
while day < len(basel_readings_celsius):
temp = basel_readings_celsius[day]
if temp > 8.0:
print("Found a hot day, stopping!")
hot_temp = temp
break
print("Processing:", temp)
day += 1
print("Hot day temperature:", hot_temp)Processing: 6.1
Processing: 7.4
Processing: 5.9
Found a hot day, stopping!
Hot day temperature: 8.2
# continue: skip missing values represented as None
mixed = [6.1, None, 5.9, None, 6.8]
total, n = 0.0, 0
for value in mixed:
if value is None:
print("Found None value in list.")
continue
total += value
n += 1
print("Mean of present values:", round(total / n, 2))Found None value in list.
Found None value in list.
Mean of present values: 6.27
Skipping missing values by hand like this is the idea behind the NaN-aware operations you will meet in the next subchapter: numpy marks a gap as nan and provides reductions such as np.nanmean that ignore it for you.
| Control flow | Syntax | What it does |
|---|---|---|
if / elif / else | if cond: … elif cond: … else: | Runs the first branch whose condition is True; else runs when none match. |
for | for item in sequence: | Repeats the body once for each item in a sequence. |
while | while condition: | Repeats the body as long as the condition stays True. |
break | break | Leaves the loop immediately. |
continue | continue | Skips the rest of this iteration and moves to the next. |
enumerate() | for i, x in enumerate(seq): | Pairs each item with its index, counting from 0. |
zip() | for a, b in zip(A, B): | Walks two or more sequences together, stopping at the shortest. |
1.2.3 Functions¶
A function names a reusable block of code. You define it with def, give it arguments, and return a value. A pure function depends only on its arguments and has no side effects, which makes it easy to test and reason about. Default arguments supply sensible fallbacks that callers can override by keyword.
def compute_sum(values):
"""Return the sum of a list of numbers."""
total = 0.0
for v in values:
total += v
return totalprint(compute_sum([1.0, 2.0, 3.0]))6.0
Going deeper: docstrings
A docstring is a string literal on the first line of a function body. Unlike a # comment
it is kept at runtime, so tools and the built-in help() can read it. Write it as a single
sentence in the imperative: what the function returns, not how it works.
def compute_sum(values):
"""Return the sum of a list of numbers."""
total = 0.0
for v in values:
total += v
return total
print(compute_sum.__doc__) # Return the sum of a list of numbers.
help(compute_sum) # shows the signature and the docstringLonger functions use a multi-line docstring that also describes the arguments, the return value, and the units involved — a habit worth forming early, since a docstring is where the unit of a quantity should be recorded when the name alone cannot carry it.
A function can return more than one value by separating them with commas. Python packs them into a tuple, which the caller can unpack into separate names — the same unpacking you saw for coordinates earlier.
def summarise(values):
"""Return the total and the mean of a list of numbers."""
total = 0.0
for v in values:
total += v
return total, total / len(values) # two values, packed into a tuple# unpack the two returned values into two names
total_celsius, mean_celsius = summarise(basel_readings_celsius)
print(total_celsius, round(mean_celsius, 2)) # 34.4 6.88
# or keep them together as a tuple
result = summarise(basel_readings_celsius)
print(type(result), result) # <class 'tuple'> (34.4, 6.88)34.4 6.88
<class 'tuple'> (34.4, 6.88)
# function with a default argument, callers may override it by keyword
def classify_temperature(temp_celsius, mild_max=20.0):
if temp_celsius < 0.0:
return "freezing"
elif temp_celsius < 10.0:
return "cold"
elif temp_celsius < mild_max:
return "mild"
else:
return "warm"print(classify_temperature(15.0)) # default mild_max = 20 -> mild
print(classify_temperature(15.0, mild_max=12.0)) # override -> warmmild
warm
A function can also accept a variable number of positional arguments by prefixing a parameter name with *. Python packs everything the caller passes into a tuple under that name — useful when the number of values is not known in advance. The same * unpacks an existing sequence back into separate positional arguments at the call site.
def mean_of(*values):
"""Return the mean of any number of positional arguments."""
total = 0.0
for v in values:
total += v
return total / len(values)
print(round(mean_of(6.1, 7.4, 5.9), 2)) # three readings, no list needed
print(round(mean_of(*basel_readings_celsius), 2)) # unpack an existing list into positional arguments6.47
6.88
Going deeper: error handling with raise and try/except
Use raise to reject input that breaks a precondition, and try/except to handle an expected failure without crashing the program.
def to_kelvin(temp_celsius):
if temp_celsius < -273.15:
raise ValueError("temperature is below absolute zero")
return temp_celsius + 273.15
for value in [20.0, -300.0]:
try:
print(round(to_kelvin(value), 2)) # 293.15, then the error is caught
except ValueError as err:
print("skipped:", err)Going deeper: lambda and scope (LEGB)
A lambda is a small anonymous function, handy as a sort key. Reusing the stations dict from above:
# station codes sorted by elevation, highest first
by_elevation = sorted(stations, key=lambda code: stations[code]["elevation_m"], reverse=True)names resolve by the LEGB rule — Local, then Enclosing, then Global, then Built-in. A name assigned inside a function is local unless declared global or nonlocal, which is one more reason pure functions that avoid touching globals are easier to reason about.
Going deeper: type hints
Annotations document the types a function expects and returns. They are not enforced at runtime, but they make intent explicit and can be checked by tools such as mypy or ty.
def to_kelvin(temp_celsius: float) -> float:
return temp_celsius + 273.15temp_celsius: float says the argument is a float; -> float says the function returns one. A
list of floats is written list[float].
Going deeper: pure and impure functions
A function is pure when it depends only on its arguments and changes nothing outside itself. A function that modifies one of its arguments is impure, and the surprise is that the caller’s own data changes:
def drop_last_impure(readings):
readings.pop() # modifies the caller's list
return readings
original = [6.1, 7.4, 5.9]
drop_last_impure(original)
print(original) # [6.1, 7.4] — the original is now shorterThe pure version copies first, leaving the input untouched:
def drop_last_pure(readings):
trimmed = readings.copy() # work on a copy
trimmed.pop()
return trimmed
original = [6.1, 7.4, 5.9]
shorter = drop_last_pure(original)
print(original, shorter) # [6.1, 7.4, 5.9] [6.1, 7.4]Lists, dicts, and other mutable objects are bound to the same object as the caller’s variable, not a copy, in a potentially risky way: if the object is changed, even inside a function, the caller’s original data changes too. The generated-code bug at the end of this subchapter — a mutable default argument — is related: the same list object gets reused and mutated across every call.
When generated code lies: the shared default list¶
Ask an assistant for a function that collects the readings above a threshold and the version below is a common answer: a list parameter, defaulting to an empty list, filled in and returned. Called on one station it gives the right answer. Called on a second station it returns that station’s warm days with the first station’s still in front of them, and reports no error.
def collect_warm_days(readings_celsius, threshold=0.0, out=[]):
# append days warmer than the threshold (as an assistant returned it)
for r in readings_celsius:
if r > threshold:
out.append(r)
return out
print("Station 1 warm days:", collect_warm_days([-2.3, -1.1, 0.4, 1.2, -0.8]))
print("Station 2 warm days:", collect_warm_days([6.1, 7.4, 5.9, 8.2, 6.8]))Station 1 warm days: [0.4, 1.2]
Station 2 warm days: [0.4, 1.2, 6.1, 7.4, 5.9, 8.2, 6.8]
def collect_warm_days(readings_celsius, threshold=0.0):
out = [] # a fresh list is built on every call
for r in readings_celsius:
if r > threshold:
out.append(r)
return out
# assign each call's result to its own variable
jfj_warm = collect_warm_days([-2.3, -1.1, 0.4, 1.2, -0.8])
bas_warm = collect_warm_days([6.1, 7.4, 5.9, 8.2, 6.8])
print("Station 1 warm days:", jfj_warm)
print("Station 2 warm days:", bas_warm)Station 1 warm days: [0.4, 1.2]
Station 2 warm days: [6.1, 7.4, 5.9, 8.2, 6.8]
Summary¶
| Concept | Rule to remember |
|---|---|
| Lists | Ordered and mutable; index from 0, and .append grows them. |
| Tuples | Immutable records — safe to pass around, and they cannot be edited in place. |
| Dicts | Map keys to values; read safely with .get(key, default), walk pairs with .items(). |
| Branching | if/elif/else runs at most one branch, and exactly one when including else. |
| Loops | for walks a sequence, while repeats on a condition; break and continue steer them. |
| Pairing | enumerate adds an index, zip walks two sequences together. |
| Functions | def bundles logic with arguments, defaults, and a return value; *args collects extra positional arguments. |
| Default arguments | Never make a default mutable (out=[]): it is created once and shared across every call. |
Resources¶
Automate the Boring Stuff with Python — free online; the chapters on flow control, functions, lists, and dictionaries map directly onto this subchapter.
Software Carpentry — Programming with Python — a research-oriented introduction to lists, loops, conditionals, and functions.