This notebook builds the smallest complete toolkit for handling environmental data in Python: variables, the core scalar types, arithmetic, casting, string formatting, and reading and writing text files.
We work throughout with one running example, a week of daily maximum air temperatures from the Jungfraujoch research station in the Swiss Alps, and finish by showing how a single type confusion can make generated code return the wrong answer without raising an error.

Figure 1:Image by haim charbit from Pixabay
1.1.1 Variables and the Core Scalar Types¶
A variable binds a name to a value, and Python infers the value’s type, which determines what operations are allowed. Five scalar types cover almost everything a single measurement needs: int (whole counts) , float (real-valued measurements), str (text), bool (true/false flags), and None (a deliberate “no value”, useful for missing or not-yet-assigned information).
Our running example is the Jungfraujoch research station in the Swiss Alps (3571 m), one of the highest year-round weather stations in Europe. Let’s assume we have one week of daily maximum air temperatures from it. The number of readings is a whole count; each temperature is a decimal-point number.
n_readings = 7 # Integer (int): number of measurements
temp_celsius = -8.4 # Floating-point (float): degrees Celsius, day 1 of the week
print(n_readings, type(n_readings)) # print() displays a value, type() its type
print(temp_celsius, type(temp_celsius))7 <class 'int'>
-8.4 <class 'float'>
The station also has a name, each reading can carry a flag that is either true or false, and a quality field may have no value assigned to it yet.
station_name = "Jungfraujoch" # String (str): name of the station
is_subzero = True # Boolean (bool): True if the temperature is below zero
quality_flag = None # NoneType (None): quality information not yet assigned
print(station_name, type(station_name))
print(is_subzero, type(is_subzero))
print(quality_flag, type(quality_flag))Jungfraujoch <class 'str'>
True <class 'bool'>
None <class 'NoneType'>
| Data Type | Python Identifier | Operational Scope |
|---|---|---|
| Integer | int | Whole numerical counts |
| Floating-Point | float | Real-valued measurements |
| String | str | Textual data |
| Boolean | bool | True/False logical flags |
| NoneType | None | Deliberate absence of value |
1.1.2 Arithmetic, Comparison, and Casting¶
Arithmetic (+ - * / // % **) and comparison (< <= > >= == !=) operators behave as in mathematics, with two things to watch:
/always returns afloateven if the operands divide evenly;Operators are type-dependent:
+adds numbers but concatenates strings.
# Arithmetic: unit conversions, with units encoded in the names
temp_kelvin = temp_celsius + 273.15
temp_fahrenheit = 9 / 5 * temp_celsius + 32
print("Temperature in Kelvin:", temp_kelvin, "K")
print("Temperature in Fahrenheit:", temp_fahrenheit, "F")Temperature in Kelvin: 264.75 K
Temperature in Fahrenheit: 16.88 F
Two of the seven days in our week fell strictly below -9.0 °C. / gives the fraction of the week that was that cold, and always returns a float:
# Division: / is always float, even when the numbers divide evenly
n_below_minus9 = 2
fraction_below = n_below_minus9 / n_readings
print("Fraction of days below -9 C:", fraction_below)Fraction of days below -9 C: 0.2857142857142857
// and % split a count into whole groups and a remainder. Over a longer record — the station’s full January, 31 days — they answer “how many complete weeks, and how many days left over?”:
n_days_january = 31
print("Complete weeks:", n_days_january // 7) # floor division: 4
print("Days left over:", n_days_january % 7) # remainder (modulus): 3Complete weeks: 4
Days left over: 3
# Comparison operators return bool
print(temp_kelvin > 270.0)False
Four operator families appear throughout this book, including the comparison operators you have just met.
| Family | Operators | What they do |
|---|---|---|
| Comparison | < <= > >= == != | Compare two values, return a bool |
| Assignment | = += -= *= /= | Assign, or update a variable in place |
| Logical | and or not | Combine or negate bool values |
| Identity | is is not | Test whether two names refer to the same object |
count += 1 is shorthand for count = count + 1. Use is only to compare with None;
for values, use ==.
# Assignment operators update a variable in place
count = 0
count += 3 # same as count = count + 3
count -= 1
count *= 2
print(count) # 44
# Logical operators: `and` needs both, `or` needs at least one
is_freezing = temp_celsius < 0.0
is_valid = temp_celsius > -273.15
print(is_freezing and is_valid) # True: both conditions hold
print(is_freezing or False) # True: at least one holdsTrue
True
# `not` negates a bool; `is` tests identity, which is how you check for None
print(not is_freezing) # False
print(quality_flag is None) # True
print(temp_celsius is not None) # TrueFalse
True
True
Use the Exercises section to explore all the different comparison operators.
Casting converts between types with int(), float(), and str(). The station also logs air pressure, and at 3571 m it is far below sea level pressure — around 651 hPa, or 65120 Pa. Here it arrives as text:
# Values read from text arrive as str and must be cast before arithmetic
raw = "65120.0" # station air pressure, Pa, as text
print(raw, type(raw))65120.0 <class 'str'>
# Execution would halt here due to TypeError:
# Uncomment this cell (with e.g., Ctrl+/) to see
# print(raw + 1.0)pressure_pa = float(raw)
print(pressure_pa, type(pressure_pa))
print(pressure_pa + 100.0) # now arithmetic works
count_str = str(n_readings) # casting the other way, to build text
print(count_str + " readings") # str + str concatenates65120.0 <class 'float'>
65220.0
7 readings
print(int(3.99)) # truncates toward zero -> 3, not 43
1.1.3 Rounding Numbers¶
Use round() for nearest-integer behaviour. A second argument says how many decimal places to keep.
print(round(1.666667, 2)) # round to 2 decimal places1.67
1.1.4 Formatting Numbers with f-strings¶
f-strings (f"...") interpolate variables and apply a format specifier after a colon: :.2f fixes two decimals, :.3e uses scientific notation, :.1% formats a fraction as a percentage. This keeps reported precision honest and units explicit.
# f-strings interpolate values and apply format specifiers after ':'
print(f"{station_name}: {temp_celsius:.1f} °C ({temp_kelvin:.2f} K)")
print(f"pressure = {pressure_pa:.3e} Pa")
print(f"days below -9 °C = {fraction_below:.1%}")Jungfraujoch: -8.4 °C (264.75 K)
pressure = 6.512e+04 Pa
days below -9 °C = 28.6%
1.1.5 Working with Text¶
Measurements often arrive embedded in text — station records, log lines, CSV fields.
A string is a sequence of characters, so you can pick out single characters by index and ranges of them by slice. Indexing starts at 0; negative indices count from the end.
# indexing: single characters
print(station_name[0], station_name[-1]) # J hJ h
A slice [start:stop] includes start and excludes stop. Either end can be left out, and it defaults to the beginning or the end of the string.
# slicing: [start:stop] includes start, excludes stop
print(station_name[0:5]) # Jungf
print(station_name[:5]) # Jungf — start defaults to 0
print(station_name[5:]) # raujoch — stop defaults to the end
print(station_name[-4:]) # joch — the last four charactersJungf
Jungf
raujoch
joch
Station records are often fixed width: every field always occupies the same positions, so slicing alone can take them apart. Here the first three characters are the station code and the rest is the date.
record = "JFJ2022-01-03"
print(record[:3]) # JFJ — the station code
print(record[3:]) # 2022-01-03 — the dateJFJ
2022-01-03
The same slicing syntax will return on lists in the next subchapter and on arrays in 1.3.
Methods¶
Not every record is that tidy. String methods clean them up: .strip() removes surrounding whitespace, .split(sep) breaks a string into parts, and .lower()/.upper()/.title() normalise case. The parts are still text and must be cast before arithmetic.
All of these are methods: functions that belong to a type and are called on a value of that type with a dot. The functions used so far, such as print(), type(), float(), and round(), take the value to work on inside their parentheses. A method takes it from in front of the dot instead, and the parentheses hold only any further arguments: record.split(",") below splits the string record, and "," says where to cut. The parentheses are required even when there is nothing to put in them, as in .strip(). Each type has its own methods: every str has .split(), but a float does not, so temp_celsius.split(",") raises an AttributeError. A string method never changes the string it is called on. It returns a new value, which you keep by assigning it to a name, and which you can call a further method on straight away: in fields[0].strip().title(), .strip() returns a trimmed string and .title() is then called on that.
# the same station, logged messily: extra spaces, mixed case, comma-separated
record = " JUNGFRAUJOCH , -9.2 , degC "
fields = record.split(",")
print(fields)[' JUNGFRAUJOCH ', ' -9.2 ', ' degC ']
name = fields[0].strip().title()
print(name)
name_lowercase = name.lower()
print(name_lowercase)Jungfraujoch
jungfraujoch
value_celsius = float(fields[1].strip())
print(value_celsius)-9.2
unit = fields[2].strip()
print(unit)degC
Not everything after a dot is a method. An attribute is a value stored on an object, and it is read with the dot but without parentheses, since there is nothing to run. Paths in 1.1.7 have both: data_path.parent is an attribute holding the folder the file sits in, and data_path.exists() is a method that checks the disk. Python looks up a method the same way, as an attribute whose value is a function, which is why temp_celsius.split(",") fails with an AttributeError: a float has no attribute called split. Adding parentheses to a plain attribute, as in data_path.parent(), raises a TypeError. Leaving them off a method raises nothing: record.strip gives back the method itself instead of a trimmed string, and any error appears only later, where that result is used as text.
1.1.6 Importing Libraries¶
Python keeps extra tools in libraries (also called modules), and you bring them into your environment with import. The standard library ships with many tools; math is a good example, and the exercises for this subchapter use it.
import math
print(math.sqrt(16)) # 4.0
print(math.pi) # 3.1415926535897934.0
3.141592653589793
Going deeper: rounding down and up with math.floor() and math.ceil()
math.floor() and math.ceil()int() truncates toward zero and round() goes to the nearest integer. The math module adds the two remaining directions: math.floor() always rounds down, toward negative infinity, and math.ceil() always rounds up, toward positive infinity. Both return an int.
print(math.floor(3.99), math.ceil(3.01)) # 3 4
print(math.floor(-8.4), math.ceil(-8.4)) # -9 -8
print(int(-8.4)) # -8, the same as ceil for a negative valueFor positive values math.floor() gives the same result as int(); for negative values, such as the sub-zero Jungfraujoch readings, it does not. math.ceil() is the one to use when a count has to cover everything: storing 31 days of readings in 7-day files takes math.ceil(31 / 7), which is 5 files, whereas the floor division 31 // 7 gives 4 and leaves three days out.
The dot in math.floor(x) does not make it a method: it picks the function floor out of the math module, and the value goes inside the parentheses, as with round().
1.1.7 Files and Paths with pathlib¶
Standard .txt and .csv files are universally readable plain text. pathlib.Path builds filesystem paths that work on any operating system: join the parts with the / operator, Path("_files") / "station_daily.csv", and Python uses the right separator for your system. Never type backslashes (e.g., \t inside a string is a tab), and do not start a path with / unless you mean the root of the disk. We open paths with .open(), specifying a mode:
| Mode | Meaning |
|---|---|
"r" | Read an existing file (the default); fails if it does not exist |
"w" | Write, creating the file or overwriting everything in it |
"a" | Append, adding to the end and keeping the existing contents |
An open file must be closed again, or the last writes may never reach disk. Doing that by
hand is easy to forget, so the standard form is a with block, which closes the file for you
as soon as the block ends — even if an error occurs inside it.
from pathlib import Path
data_path = Path("_files/station_daily.csv")
data_path.parent.mkdir(exist_ok=True) # create the folder if needed
print(f"Does the file exist? {data_path.exists()}")Does the file exist? False
Open the path in mode "w" to create the file and write the header row plus the first day of the week. The with block closes it again as soon as the indented lines are done.
# "w" creates the file (or overwrites it); the with block closes it automatically
with data_path.open("w", encoding="utf-8") as fhandler:
fhandler.write("station,day,max_temp_celsius\n")
fhandler.write("JFJ,2022-01-01,-8.4\n")
print(f"Does the file exist now? {data_path.exists()}")Does the file exist now? True
In Python, indentation is part of the syntax. A line that opens a block ends in a colon, like with ... as fhandler: above, and every line indented beneath it belongs to that block, up to the first line back at the previous level. The two write calls are indented, so they run while the file is open; the print is not, so it runs after the with block has closed the file. Many languages mark blocks with braces and ignore indentation, but in Python the indentation is the only marker, so moving a line in or out of a block changes when it runs, and Python raises no error as long as the result is still valid code. Use four spaces per level, the convention in Python code; notebook editors indent the line after a colon for you. A colon with no indented line after it, or an indented line where no block was opened, stops the cell with an IndentationError. The bodies of the if statements, loops, and functions in the next subchapter follow the same rule.
Mode "a" reopens the same file and adds to the end instead of overwriting it. That gets the remaining six days of the week in.
with data_path.open("a", encoding="utf-8") as fhandler:
fhandler.write("JFJ,2022-01-02,-8.9\n")
fhandler.write("JFJ,2022-01-03,-9.2\n")
fhandler.write("JFJ,2022-01-04,-9.0\n")
fhandler.write("JFJ,2022-01-05,-9.5\n")
fhandler.write("JFJ,2022-01-06,-7.8\n")
fhandler.write("JFJ,2022-01-07,-6.1\n")# "r" reads it back; .read() returns the whole file as one string
with data_path.open("r", encoding="utf-8") as fhandler:
raw_text = fhandler.read()
print(raw_text)station,day,max_temp_celsius
JFJ,2022-01-01,-8.4
JFJ,2022-01-02,-8.9
JFJ,2022-01-03,-9.2
JFJ,2022-01-04,-9.0
JFJ,2022-01-05,-9.5
JFJ,2022-01-06,-7.8
JFJ,2022-01-07,-6.1
Going deeper: the manual form, and the built-in open()
open()You will also see code that manages files by hand, calling .open() and then .close() yourself
instead of using a with block:
fhandler = data_path.open("a", encoding="utf-8")
fhandler.write("JFJ,2022-01-08,-5.4\n")
fhandler.close()It behaves the same as long as the close() call is never skipped — but an exception raised
between open() and close() leaves the file unclosed, and it is easy to forget the call
entirely. The with block avoids both risks, which is why it is the pattern used throughout the
rest of this book. Recognise the manual form in code you did not write; we do not recommend adopting it.
You will also meet the built-in open() used directly on a path that is already a plain string,
which is common when a path comes from somewhere else:
with open("_files/station_daily.csv", "r", encoding="utf-8") as fhandler:
raw_text = fhandler.read()Path.open() and the built-in open() do the same job; use whichever matches what you are
holding.
# .readlines() gives a list of lines instead of one long string
with data_path.open("r", encoding="utf-8") as fhandler:
all_lines = fhandler.readlines()
print("number of lines:", len(all_lines)) # 8: one header plus seven daysnumber of lines: 8
first_record = all_lines[1].strip() # .strip() removes the trailing newline
fields = first_record.split(",")
print(fields) # ['JFJ', '2022-01-01', '-8.4']
temp_text = fields[2]
print(temp_text, type(temp_text)) # -8.4 <class 'str'> — still text!['JFJ', '2022-01-01', '-8.4']
-8.4 <class 'str'>
# cast before doing arithmetic
temp_celsius_read = float(temp_text)
print(temp_celsius_read, type(temp_celsius_read)) # -8.4 <class 'float'>
print(temp_celsius_read + 273.15) # 264.75-8.4 <class 'float'>
264.75
When generated code lies: a hidden type bug¶
AI assistants write code that runs and looks right, but is sometimes silently wrong. A classic trap: values read from a text file are strings, and comparing strings does not behave like comparing numbers.
# day 1 and day 3 of the week, straight off the file — so they are strings
reading_day1 = "-8.4"
reading_day3 = "-9.2"
# "was day 1 warmer than day 3?" — comparing the strings directly
print("is", reading_day1, "warmer than", reading_day3, "?", reading_day1 > reading_day3)is -8.4 warmer than -9.2 ? False
# cast to float first, then compare as numbers
print("is", reading_day1, "warmer than", reading_day3, "?", float(reading_day1) > float(reading_day3))is -8.4 warmer than -9.2 ? True
Going deeper: floating-point precision
Floats are binary approximations of real numbers, so familiar decimals are not stored exactly:
print(0.1 + 0.2) # 0.30000000000000004
print(0.1 + 0.2 == 0.3) # False
import math
print(math.isclose(0.1 + 0.2, 0.3)) # TrueNever test floats for equality with ==; use math.isclose (or numpy.isclose for arrays) with an explicit tolerance. This matters whenever you compare or threshold measured quantities.
Summary¶
| Concept | Rule to remember |
|---|---|
| Types | Every value has a type; type() shows it (int, float, str, bool, None). |
| Casting | Text from a file is str; use float() or int() before doing maths. |
| Operators | + adds numbers but joins strings; / always gives a float. |
| Comparison | Strings compare character by character, not by numeric value. |
| Naming | Encode the quantity and unit: temp_celsius, pressure_pa. |
| Formatting | f"{value:.2f}" fixes the number of decimals shown. |
| Files | Use pathlib.Path with a with path.open(mode) block; "w" overwrites, "a" appends. |
Resources¶
Project Pythia Foundations — geoscience-flavoured tutorials on the core scientific-Python stack; the natural next step after this chapter.
The Python tutorial — the authoritative reference for the built-in types and syntax used above.