This chapter introduces three classical classifiers — support vector machines, decision trees, and random forests/ensembles — then applies them to a real environmental risk-mapping problem: wildfire susceptibility in the Liguria region of Italy.
Support vector machines, decision trees, and random forests/ensemble modeling: how each works, their regularization hyperparameters, and where they show up in environmental science.
Comparing LinearSVC, SVC, and SGDClassifier on the iris dataset, then training an SVM
regressor on California housing prices.
Training and fine-tuning a decision tree on the moons dataset with grid search, then building a random forest from scratch by growing and combining many trees.
Comparing individual classifiers against voting and stacking ensembles on MNIST digits.
Training classifiers on real topography, land-cover, and climate data to map wildfire susceptibility in the Liguria region of Italy.
Resources¶
Textbooks and papers this chapter’s exercises adapt¶
Hands-On Machine Learning with Scikit-Learn — Aurélien Géron; chapters 5 (SVMs), 6 (decision trees), and 7 (ensembles/random forests) are the source of the classifier exercises. (3.2, 3.3, 3.4). The current edition, Hands-On Machine Learning with Scikit-Learn and PyTorch, covers decision trees and random forests as chapters 5 and 6 but no longer has a standalone SVM chapter.
Tonini, M., et al. “A Machine Learning-Based Approach for Wildfire Susceptibility Mapping. The Case Study of the Liguria Region in Italy.” Geosciences 10.3 (2020): 105 — the source of the wildfire-mapping exercise and its Liguria dataset. (3.5)
Trucchia, A., et al. “Defining Wildfire Susceptibility Maps in Italy for Understanding Seasonal Wildfire Regimes at the National Level.” Fire 5.1 (2022): 30 — generalizes the Liguria case study to all of Italy. (3.5)
Scikit-learn¶
Support vector machines —
LinearSVC,SVC, kernels, and regularization. (3.1, 3.2)Decision trees — the CART algorithm and its regularization hyperparameters. (3.1, 3.3)
Ensemble methods — voting, bagging, random forests, and stacking. (3.1, 3.3, 3.4)
Model selection: GridSearchCV — hyperparameter search with cross-validation. (3.3)
Datasets¶
The iris dataset and the California housing dataset — bundled with scikit-learn. (3.2)
MNIST — handwritten digits, used for comparing individual and ensemble classifiers. (3.4)
Liguria wildfire susceptibility data (topography, land cover, climate, and historical wildfire occurrence) — Trucchia, Meschi, and Tonini. (3.5)