This is the first chapter of the book; it covers the following subchapters. One further subchapter, held back as optional self-study material, lives in the Bonus material section after this chapter.
Variables, scalar types, casting, string formatting, and reading and writing text files.
Lists, tuples, and dicts, plus the control flow and functions that tie them together.
The numpy array: creation, indexing, vectorised math, broadcasting, and reductions.
Building figures with matplotlib, and labelled, multi-dimensional arrays with xarray.
Series and DataFrames: selecting, resampling, grouping, and handling missing data.
Points, lines, and polygons with geopandas: projections, joins, and measuring correctly.
Bundling state and behaviour into classes, then reading and trusting code you didn’t write — assertions, exceptions, logging, and tests.
From descriptive statistics to a first, honestly evaluated regression, classification, and clustering model.
Resources¶
Every subchapter above ends with its own short reading list, pointing at the specific tool or dataset just used. This section collects all of them in one place, and adds the resources worth knowing about for learning Python itself — at whatever level you are starting from.
Learning Python: books and courses¶
Getting started, no prior programming assumed
Automate the Boring Stuff with Python — Al Sweigart; free online. Practical, task-driven, and the resource already pointed to from subchapter 1.2.
Think Python — Allen Downey; free (the current edition is a set of runnable Jupyter notebooks). A slower, more conceptual first course than Automate the Boring Stuff, closer to how a first computer-science class is taught.
Python Crash Course — Eric Matthes (No Starch Press). The most widely used printed introduction; fast-paced and project-based, ending in small data-analysis, web, and game projects.
The Python Tutorial — the official documentation’s own introduction. Terser than a book, but authoritative and always current; already this book’s own reference in subchapter 1.1.
Software Carpentry — Programming with Python — a research-oriented introduction built around lists, loops, conditionals, and functions; already used throughout Part I.
Comfortable with the basics, building fluency
A Whirlwind Tour of Python — Jake VanderPlas; free. A fast second pass through the language for someone who already codes in something else, written explicitly as the on-ramp to the Data Science Handbook below.
Python Data Science Handbook — Jake VanderPlas; free online. Covers numpy, pandas, matplotlib, and scikit-learn in real depth — arguably the single closest match to this part’s own scope, and already this book’s numpy resource in subchapter 1.3.
Effective Python — Brett Slatkin. Specific, well-explained practices for writing idiomatic Python once the basics are solid; the 3rd edition covers 125 of them.
Going deeper
Fluent Python — Luciano Ramalho (O’Reilly). How Python’s data model actually works underneath the syntax — for once the language feels comfortable and the question becomes why, not just how.
The scientific Python stack¶
Python Data Science Handbook — Introduction to NumPy — arrays, broadcasting, masking, and ufuncs. (1.3)
Scientific Python Lectures — NumPy — a concise, research-oriented tour of the same material. (1.3)
Pythia Foundations — Matplotlib Basics — the figure/axes model and the core plot types. (1.4)
Pythia Foundations — Introduction to Xarray — DataArray/Dataset, label-based selection, built-in plotting. (1.4)
An Introduction to Earth and Environmental Data Science — Abernathey and Key; numpy, matplotlib, and xarray on real geoscience data. (1.4)
Python for Data Analysis, 3rd ed. — Getting Started with pandas — McKinney; Series/DataFrame mechanics and data cleaning. (1.5)
pandas — Getting started — the official task-oriented tutorials. (1.5)
Geospatial data¶
GeoPandas — Managing projections — setting and transforming coordinate reference systems. (1.6)
GeoPandas basics: maps, projections, and spatial joins — a worked introduction to GeoDataFrames and spatial joins. (1.6)
Object-oriented and defensive programming¶
Object-Oriented Programming (OOP) in Python — classes, instances, attributes, methods, inheritance. (1.7)
Python Classes: The Power of Object-Oriented Programming — instance vs. class attributes, dataclasses, and when not to use a class. (1.7)
pytest documentation — writing tests, fixtures, parametrisation. (1.7)
Machine learning and statistical graphics¶
scikit-learn — Getting started — the estimator interface, train/test evaluation, pipelines. (1.8)
seaborn — statistical visualisation on dataframes. (1.8)
scikit-learn — Clustering and Decomposing signals (PCA) — k-means and PCA reference documentation. (1.8)
palmerpenguins — Horst, Hill & Gorman (2020); the real dataset behind 1.8’s clustering and PCA content. (1.8)
Reproducibility and tooling¶
Project Pythia Foundations — geoscience-flavoured tutorials across the whole scientific-python stack; a natural next step after this part. (1.1)
uv documentation — project creation, the src layout, dependency management. (Bonus A)
Ruff documentation — the linter and formatter used in Bonus A’s going-deeper boxes. (Bonus A)
pooch documentation — retrieving single files, registries, hashing, DOI downloads. (Bonus A)
Earth and Environmental Data Science — All About Data — Abernathey and Key on fetching remote data, Zenodo DOIs, FAIR practice. (Bonus A)