Exercise 1: Name your measurements¶
Create scientifically named variables for a single observation at a river gauge: a discharge of 4.7 m³ s⁻¹, a station name of your choice, an integer count of 24 hourly samples, and a boolean flag for whether discharge exceeds 5 m³ s⁻¹. Print each value together with its type.
# Your solution hereExercise 2: Unit conversions¶
A temperature is recorded as 18.5 °C. Compute it in kelvin and in degrees Fahrenheit. Separately, a discharge of 4.7 m³ s⁻¹ drains a catchment of area 250 km². Express the specific discharge in mm per day (use: 1 m³ s⁻¹ spread over 1 km² for one day equals 86.4 mm). Print each result to two decimals with an f-string.
# Your solution hereExercise 3: Operator families and a quality flag¶
Three daily readings arrive from a station: reading_1_celsius = 18.2, reading_2_celsius = -300.0, and reading_3_celsius = None (the sensor reported nothing).
Build a boolean
is_plausiblethat isTruewhen the first reading is above absolute zero and below 60 °C.Build
is_implausiblefor the second reading: apply the same test as in step 1, and negate it withnot.Use
isto check whether the third reading is missing.Starting from
n_valid = 0, use+=twice to count how many of the first two readings pass the step 1 test. Note thatint()castsTrueto 1 andFalseto 0.The sensor is rated for -40 °C to 60 °C, both bounds included. Use
>=and<=to test whether the first reading lies in that range, and!=to check that the first two readings differ.
# Your solution hereExercise 4: Parse a record¶
A logger emits records like "2024-01-15 , JUNGFRAUJOCH , -2.3 , degC". Extract the date string, the station name in title case, the numeric value as a float, and the unit as written, then print them. Then print the station name in lower case as well, the form you would use in a file name.
# Your solution hereExercise 5: Indexing and slicing a record¶
Weather records often arrive in a fixed shape, where each field always occupies the same positions. Given
record = "ALO01.03.2022"the first three characters are the station code and the rest is a date as dd.mm.yyyy.
Using only indexing and slicing (no
.split()), extract the station code, the day, the month, and the four-digit year, and print them.Now extract the month a second way, using
.split(".")on the date part, and use==to confirm you get the same answer.In a comment, say which of the two approaches you would trust to extract the month from a record where the station code can be two or four characters long, and why.
# Your solution hereExercise 6: Formatting π to the precision you choose¶
math.pi holds π to about 16 significant digits, and an f-string format specifier prints it to as many decimals as you ask for.
Create the
_filesfolder if it does not exist, as in the lecture, then write the sentence"This is my first I/O exercise."to_files/pi_exercise.txtwith mode"w".Reopen the file with mode
"a"and append three more lines: one built with an f-string that reportsmath.pito four decimal places, and two lines of your own choosing.Read the file back with
.readlines()and print its full contents.Change the format specifier to two decimals instead of four, rerun, and confirm the printed value now rounds to
3.14.
# Your solution hereExercise 7: Round-trip through a file, with an append¶
Write the Jungfraujoch week from the lecture — the daily maximum temperatures -8.4, -8.9, -9.2,
-9.0, -9.5 (°C) — to _files/jfj_week.txt, one value per line, using pathlib and a with
block. Create the _files folder first if it does not exist; do not rely on Exercise 6 having
made it.
Write the five values with mode
"w", one.write()call per value.Reopen the file with mode
"a"and append the last two days of the week,-7.8and-6.1.Read the file back with
.readlines(), cast each of the seven lines tofloatby name, and compute the mean by summing the seven values by hand with+and dividing bylen(). Print it to two decimals.
# Your solution hereGoing deeper (optional)¶
Exercise 8: Same seven values, same mean?¶
Adding the same floats in a different order does not have to give exactly the same result,
because each + rounds its result to the nearest float. This exercise uses math.isclose from
the going-deeper box on floating-point precision.
Compute the mean of the Jungfraujoch week — -8.4, -8.9, -9.2, -9.0, -9.5, -7.8, -6.1 — by adding the values in that order and dividing by 7.
Compute it again, adding the same seven values in reverse order.
Compare the two means with
==, and print the result.Compare the same two means with
math.isclose, and print the result.In a comment, say which of the two comparisons you would put in a script that checks whether two calculations of the same quantity agree, and why.
# Your solution hereExercise 9: A year of station data¶
The exercises so far used values typed by hand. This one uses a real record: daily maximum temperatures for 2022 from the automated weather station at Independence Municipal Airport, Iowa, US, station code IIB. The file has one row for each of the 365 days, but 16 of them, from 25 October to 9 November, have no temperature value. Work through the steps in order, since each step reuses variables from the one before. The same file returns in 1.2, where you turn the whole year into a monthly summary.
The data has three columns: the station code, the day as dd.mm.yy, and the daily maximum
temperature in degrees Fahrenheit. Converting to SI units is part of the job.

Figure 1:Image by Tobias Hämmer from Pixabay
# Pre-supplied: download the data file and cache it locally.
# You do not need to understand this cell yet — fetching data is covered in the
# reproducible-data-pipelines bonus subchapter.
import pooch
data_file = pooch.retrieve(
url="https://raw.githubusercontent.com/gse-unil/2026_MLEES_book/main/data/part-I/station_iib_daily_max_temp_2022.csv",
known_hash="sha256:034755fb289b4e157e5d029995481786a359eb25a80df62c09ee83d063013144",
fname="station_iib_daily_max_temp_2022.csv",
path=pooch.os_cache("mlees"),
)
print("data cached at:", data_file)data cached at: /home/runner/.cache/mlees/station_iib_daily_max_temp_2022.csv
Step 1. Open the file with a with block in read mode and read all its lines with
.readlines().
Print how many lines the file has, then print the first three lines exactly as they come off disk.
# Your solution here
# Hint: data_file is a plain string; after from pathlib import Path, Path(data_file) is a
# path you can .open()
# Hint: with Path(data_file).open("r", encoding="utf-8") as fhandler: ...
# Hint: .readlines() gives you a list of strings, one per line
# Hint: lines[0] is the header row, not an observation
# Hint: print lines[0], lines[1] and lines[2] one at a timeStep 2. Look at the first record (the second line of the file). Split it on the separator and print the resulting list. Then print the type of each of the three fields.
Answer in a comment: are the temperatures numbers at this point?
# Your solution here
# Hint: .strip() first, to remove the trailing newline
# Hint: .split(",") turns one line into a list of fields
# Hint: type() will tell you what you are really holdingStep 3. Take that same first record and turn it into three properly typed, well-named variables: the station code as text, the day as text, and the maximum temperature as a float in degrees Celsius. Print them in one line with an f-string, showing the temperature to one decimal with its unit.
Recall that °C = (°F − 32) × 5/9.
# Your solution here
# Hint: name each variable for the quantity and unit it holds, e.g. day_1 and
# max_temp_celsius_1; step 4 needs a second set of names for the second day
# Hint: f"{value:.1f} °C" formats to one decimalStep 4. Now do the same for the second record (the third line of the file), so you have two days side by side.
Ask which of the two days was warmer by comparing the two raw temperature fields, before any casting.
Ask the same question again, this time comparing the two values after casting to float.
The two answers disagree. In a comment, explain which one is right and why the other one is wrong.
# Your solution here
# Hint: you already have the first record's fields; do the same for lines[2]
# Hint: compare the raw strings first, then compare float() of each
# Hint: give the second day its own names, e.g. day_2 and max_temp_celsius_2Step 5. Write your own summary file, in SI units this time.
Build a path to
_files/station_iib_summary.csvwithpathlib.Path, creating the folder if it does not exist.In one
withblock in mode"w", write a header linestation,day,max_temp_celsiusand then one line for each of the two days you converted, with the temperature formatted to one decimal.Reopen the file in mode
"r", read it back with.read(), and print it.
# Your solution here
# Hint: summary_path = Path("_files/station_iib_summary.csv")
# Hint: summary_path.parent.mkdir(exist_ok=True)
# Hint: build each line with an f-string, and remember the trailing \n
# Hint: f"{station_code},{day_1},{max_temp_celsius_1:.1f}\n"