Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Open In Kaggle

Exercise 1: CSV versus parquet

Build a 1000-row DataFrame with a time column (pd.date_range, hourly) and a numeric column. Write it to CSV and to parquet, print both file sizes, and confirm that parquet preserves the datetime dtype on read-back.

Exercise 2: netCDF versus zarr

Build a small xarray Dataset with a monthly time coordinate. Write it to netCDF and to zarr, then confirm that the netCDF output is a file and the zarr output is a directory.

Exercise 3: Hashing for integrity

Write a short text file, copy it to a second path, and use pooch.file_hash to confirm that identical content produces identical hashes.

Exercise 4: An idempotent fetch

Using a local file as the “source”, implement fetch(expected_hash) that copies the source into a cache path on a miss and skips on a hash match. Call it twice and show that the first call downloads and the second is a cache hit.

Exercise 5: Fix the existence-only check

A stale, truncated file sits in the cache. Write fetch_verified(expected_hash) that re-fetches from the source whenever the cached file is missing or its hash does not match, then returns the correct content.

source = Path("_files/src.csv"); source.write_text("x\n1\n2\n3\n")
cache = Path("_files/c5/src.csv"); cache.parent.mkdir(exist_ok=True, parents=True)
cache.write_text("x\n1\n")   # stale / truncated

Exercise 6: A real pooch.retrieve call

Write a pooch.retrieve call that fetches a CSV from a URL, pins it with a SHA256 known_hash, caches it in an OS-appropriate folder, and reads it with pandas.

Use a file this book already ships, so the call actually runs:

https://raw.githubusercontent.com/gse-unil/2026_MLEES_book/main/data/part-I/forest_fires.csv
sha256:0d6586a1fa52f55bef48578aef14eb97273f1e9330e1a53423df497a77065253

Then run the cell a second time and watch what it does not print. The first run downloads; the second finds the file in the cache, checks its hash, and returns the path without touching the network.