Chapter 4: Unsupervised Learning (Clustering, Dimensionality Reduction) and Environmental Complexity
This chapter turns to unsupervised learning — dimensionality reduction with PCA and clustering with K-means, GMMs, and DBSCAN — closing with a real application to identifying dynamical regimes in ocean circulation data.
Supervised vs. unsupervised vs. semi-supervised learning, PCA and other dimensionality-reduction methods, and clustering with K-means, Gaussian mixture models, and DBSCAN.
Does PCA always speed up training and improve performance? Comparing a random forest and a logistic regression classifier on MNIST, with and without PCA.
Choosing the number of clusters for K-means on a subsample of MNIST digits, using the silhouette score and inertia, with and without PCA to speed up training.
Clustering reduced-dimensionality ECCO ocean-model fields with xarray and K-means to identify dynamical regimes in the North Atlantic, following Sonnewald et al.'s THOR method.
Resources¶
Textbooks and papers this chapter’s exercises adapt¶
Hands-On Machine Learning with Scikit-Learn — Aurélien Géron; chapters 8 (dimensionality reduction) and 9 (unsupervised learning) are the source of the classifier exercises. (4.2, 4.3). The current edition, Hands-On Machine Learning with Scikit-Learn and PyTorch, covers the same material as chapters 7 and 8.
Sonnewald, M., Wunsch, C., & Heimbach, P. (2019). Unsupervised learning reveals geography of global ocean dynamical regions. Earth and Space Science — the THOR method and the source of the ocean-regimes exercise. (4.4)
Sonnewald, M., & Lguensat, R. (2021). Revealing the Impact of Global Heating on North Atlantic Circulation Using Transparent Machine Learning. Journal of Advances in Modeling Earth Systems — a follow-up applying THOR to circulation change under global heating. (4.4)
Python scripts adapted from Maike Sonnewald’s own research code. (4.4)
Scikit-learn¶
Decomposition: PCA — principal component analysis and its variants. (4.1, 4.2)
Clustering — K-means, Gaussian mixture models, and DBSCAN. (4.1, 4.3, 4.4)
Clustering performance evaluation — the silhouette score and inertia used to choose the number of clusters. (4.3)
Xarray¶
Data structures —
Dataset,DataArray, and the coordinate/dimension model used to reformat the ocean-model data. (4.4)Reshaping and reorganizing data —
stack/unstack, used to flatten the gridded fields for clustering and rebuild them for plotting. (4.4)
Datasets¶
MNIST — handwritten digits, used throughout the dimensionality-reduction and clustering exercises. (4.2, 4.3)
ECCO (Estimating the Circulation and Climate of the Ocean) — the realistic ocean-model output behind the ocean-regimes dataset. (4.4)
- Sonnewald, M., Wunsch, C., & Heimbach, P. (2019). Unsupervised Learning Reveals Geography of Global Ocean Dynamical Regions. Earth and Space Science, 6(5), 784–794. 10.1029/2018ea000519
- Sonnewald, M., & Lguensat, R. (2021). Revealing the Impact of Global Heating on North Atlantic Circulation Using Transparent Machine Learning. Journal of Advances in Modeling Earth Systems, 13(8). 10.1029/2021ms002496