Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Open In Kaggle

A result is only as trustworthy as the environment and the data behind it. This subchapter has two halves: the first packages reusable code into an installable, version-pinned project with uv, so an environment can be rebuilt identically on any machine; the second surveys the file formats you will meet, fetches a remote file once and verifies it, and builds the idempotent cache → hash → skip pattern that makes a data pipeline repeatable. At the end, the generated code trusts that a file on disk is the file you wanted, without checking.

Packaging with uv: The src Layout

For code you will reuse or share, promote it from a notebook to an installable package. uv init --lib scaffolds a src/ layout with a pyproject.toml, and uv.lock records the exact dependency versions.

uv init --lib mypackage       # create a src/ layout with pyproject.toml
cd mypackage
uv add numpy                  # add a dependency, updating uv.lock
uv run pytest                 # run the tests in the locked environment

The generated pyproject.toml records the package name, version, and the dependencies uv add pins:

[project]
name = "mypackage"
version = "0.1.0"
dependencies = [
    "numpy>=2.0",
]

Placing code under src/mypackage/ means tests import the installed package, not stray files in the working directory, so packaging mistakes surface immediately. uv.lock pins the full dependency graph, making the environment reproducible on any machine.

From Reproducible Code to Reproducible Data

Packaging pins how your code runs: the same dependencies, the same versions, on any machine. The rest of this subchapter pins what it runs on — making sure a data file is fetched, verified, and cached the same way every time, so a result doesn’t quietly drift when a remote file changes underneath you.

A Tour of Formats

Match the format to the data. CSV is universal plain text, human-readable but untyped and bulky — fine for small tables and interchange. Unlike CSV, parquet is a columnar binary format: typed, compressed, and fast for large tables. For labelled n-dimensional arrays with metadata, netCDF is the standard single-file container. zarr stores the same array model as a chunked directory (or cloud object store), built for parallel and out-of-core access. When the data is a raster grid rather than a table or array, GeoTIFF holds it together with its coordinate reference system.

csv bytes:     294209
parquet bytes: 116098
parquet keeps dtypes: {'time': dtype('<M8[us]'), 'temp_celsius': dtype('float64'), 'station': <StringDtype(na_value=nan)>}
netcdf is a file:     True
zarr is a directory:  True
reopened netcdf shape: (12, 4, 6)

Fetching Data over HTTP

requests performs the HTTP GET underneath any download. The essential pattern checks the status and writes the bytes:

import requests
r = requests.get(url, timeout=30)
r.raise_for_status()                 # turn a 404/500 into an exception
Path("data.csv").write_bytes(r.content)

This works, but it re-downloads on every run and verifies nothing about which file arrived. Pinning and caching are the next step.

pooch: Download Once, Verify Always

pooch.retrieve wraps the request, caches the file, and checks its hash against a value you pin. Calling it again returns the cached path with no download.

import pooch
path = pooch.retrieve(
    url="https://example.org/data/temperature.csv",
    known_hash="sha256:1a2b3c...",      # pin the exact file
    path=pooch.os_cache("mlees"),        # OS-appropriate cache folder
)
data = pd.read_csv(path)

The pinned hash is what makes the pipeline reproducible: if the remote file ever changes, the check fails loudly instead of feeding you different data in silence. The call is shown rather than run so this book builds without network access; the cell below implements the same cache → hash → skip logic on a local file, using pooch’s real hashing.

The Idempotent Fetch: Cache, Hash, Skip

A robust fetch is idempotent — calling it repeatedly does the least work and always yields the same verified file. The logic is: if a cached copy exists and its hash matches, use it; otherwise (re)download and verify. Here a local file stands in for the remote source so the logic runs offline.

known sha256: 428eeaa65c3c10f6 ...
cache miss: download and verify
cache hit:  hash matches, skip download
PosixPath('_files/cache/temperature.csv')
shape: (5, 2) | mean temp: 17.92
<xarray.Dataset> Size: 80B
Dimensions:       (date: 5)
Coordinates:
  * date          (date) datetime64[us] 40B 2024-06-01 2024-06-02 ... 2024-06-05
Data variables:
    temp_celsius  (date) float64 40B 18.2 17.5 19.1 16.8 18.0

When generated code lies: trusting a file that is merely present

Asked to “download the data if it is not already there”, an assistant checks only whether the file exists. But a file can exist and still be wrong — truncated by an interrupted download, or left over from an older version.

rows returned: 1 (should be 5)
rows returned: 5 (now correct)

Summary

ConceptRule to remember
PackagingPromote reusable code to a src/ package with uv init --lib; pin dependencies with uv.lock.
FormatsCSV for small interchange, parquet for large tables, netCDF for labelled arrays, zarr for chunked or cloud arrays, GeoTIFF for rasters.
FetchingA bare requests.get re-downloads every time and verifies nothing.
PinningA hash says what a file is, not just what it is called.
The idempotent fetchcache → hash → skip; re-fetch on a miss or a hash mismatch.
Existence is not integrityA truncated or stale file passes an exists() check and corrupts everything downstream.

Resources