Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Open In Kaggle


Credits

This online tutorial would not be possible without invaluable contributions from Andrea Trucchia (reduced data, methods), Giorgio Meschi (code, methods), and Marj Tonini (presentation, methods). The methodology builds upon the following article:

Trucchia, A.; Meschi, G.; Fiorucci, P.; Gollini, A.; Negro, D., Defining Wildfire Susceptibility Maps in Italy for Understanding Seasonal Wildfire Regimes at the National Level, Fire, (2022)

which generalizes the study below from the Liguria region (our case study) to all of Italy:

Tonini, Marj, et al. “A machine learning-based approach for wildfire susceptibility mapping. The case study of the Liguria region in Italy.” Geosciences 10.3 (2020): 105.


In week 3’s final notebook, we will train classifiers on real wildfire data to map the fire risk in different regions of Italy. To keep the data size manageable, we will focus on the coastal Liguria region that experiences a lot of wildfires, especially during the winter.

Machine Learning for Environmental Risk Analysis

For environmental sciences practioners, one of the scenarios where machine learning can be particularly useful is risk analysis. Environmental risk analysis involves predicting where potential hazards may occur; it also involves understanding why some regions are more vulnerable to hazards than others. Machine learning models can be useful for these tasks because they can analyze large datasets containing different environmental predictors (e.g., weather conditions, soil conditions etc.) in an effective manner. By learning the hidden links between predictors and hazard risks, we may also gain new insights on what predictors or patterns are useful for creating early hazard warning systems.

In this exercise, we ask you to use the machine learning classifiers we learned in this chapter to recreate wildfire susceptibility maps for the Liguria region of Italy. The basic idea is to use ML classifiers to analyze a dataset with observations in weather conditions, vegetation cover, and topography information. The goal will be to understand how different factors enhance or reduce the probability of firefire in Liguria, which can help authorities and decision-makers to apply resources to critical areas for hazard prevention.

6409e1d442e86ebe5987bd9902a95f44.jpg

Caption: A wildfire in Italy. Can we predict which locations are most susceptible to wildfires using simple classifiers? 🔥

Source: ANSA

Let’s start by downloading and loading the datasets into memory using the pooch and GeoPandas libraries:

/home/runner/work/2026_MLEES_book/2026_MLEES_book/.venv/bin/python: No module named pip
Note: you may need to restart the kernel to use updated packages.

Part I: Pre-Processing the Dataset for Classification

Q1) After analyzing the topography and land cover data provided in variables, create your input dataset inputs from variables to predict the occurence of wildfires (wildfires). Keep at least one categorical variable (veg, bioclim, or phytoclim).

Hint 1: Refer to the documentation at this link to know what the different keys of variables refer to.

Hint 2: You may refer to Table 1 of Tonini et al., copied below, to choose your input variables, although we recommend starting with less inputs at first to build a simpler model and avoid overfitting.

Marj_Table1.PNG

Here are some pandas commands you could use to explore your data. .head() .columns() .describe()

There are 25 columns (variables) in the DataFrame. It is probably best to start simple and just use a few variables to make the wildfire prediction.

Can we use a very simple model to predict wildfires ❓

A simple model might contain dem, slope, veg and bioclim. We can use it as a baseline to evaluate model performance when you use other combinations to train the model.

However, we cannot tell you which combination would perform the best as we have not done an extensive search while preparing the notebook.

Q2) To avoid making inaccurate assumptions about which types of vegetation and non-flammable area are most similar, convert your categorical inputs into one-hot vectors.

Hint 1: You may use the fit_transform method of scikit-learn’s OneHotEncoder class to convert categorical inputs into one-hot vectors.

Hint 2: Don’t forget to remove the categorical variables from your input dataset, e.g. using drop if you are still using a GeoDataFrame, or del/pop if you are working with a Python dictionary.

Hint 3: There are numerous ways to change the categorical data into one-hot vectors. It is quite easy to do in pandas, but scikit-learn also provides some transformers that could be useful, including .ColumnTransformer() and Pipeline().

In the guided reading, you have seen how these functions are used. Try experiment with them and see if you prefer using these scikit transformers or pandas.

Now that we built our inputs dataset, we are ready to build our outputs dataset!

Q3) Using the point_index column of wildfires and variables, create your outputs dataset, containing 1 when there was a wildfire and 0 otherwise.

Hint: Check that inputs and outputs have the same number of cases by looking at their .shape[0] attribute.

Q4) Separate your inputs and outputs datasets into a training and a test set. Keep at least 20% of the dataset for testing.

Hint 1: You may use scikit-learn’s train_test_split function.

Hint 2: If you are considering optimizing the hyperparameters of your classifier, form a validation dataset as well.

Hint 3: We recommend performing the split on the indices so that it is easier to track what points are in which dataset after splitting. You will have an easier time when plotting the susceptibility map.

Congratulations, you have created a viable wildfire dataset to train a machine learning classifier! 😃 Now let’s get started 🔥

Part II: Training and Benchmarking the Machine Learning Classifiers

Q5) Now comes the machine learning fun! 🤖 Train multiple classifiers on your newly-formed training set, and make sure that at least one has the predict_proba method once trained.

Hint: You may train a RandomForestClassifier or an ExtraTreesClassifier, but we encourage you to be creative and include additional classifiers you find promising! 💻

Q6) Compare the performance and confusion matrices of your classifiers on the test set. Which classifier performs best in your case?

Hint 1: You may use the accuracy_score to quantify your classifier’s performance, but don’t forget there are many other performance metrics to benchmark binary classifiers.

Hint 2: You can directly calculate the confusion matrix using scikit-learn’s confusion_matrix function.

For comparison, below is the confusion matrix obtained by the paper’s authors:

download (2).png

Part III: Making the Susceptibility Map

Q7) Using all the classifiers you trained that have a predict_proba method, predict the probability of a wildfire over the entire dataset.

Hint: predict_proba will give you the probability of both the presence and absence of a wildfire, so you will have to select the right probability.

Q8) Make the susceptibility map 🔥

Hint 1: The x and y coordinates for the map can be extracted from the variables dataset.

Hint 2: You can simply scatter x versus y, and color the dots according to their probabilities (c=probability of a wildfire) to get the susceptibility map.

You should get a susceptibility map that looks like the one below. Does your susceptibility map depend on the classifier & the inputs you chose? Which map would you trust most?

download (3).png

It seems like our model was too simple to generate a useful map 😞 . Your TA actually experimented training a RandomForest model with 10 variables and got a 91% accuracy!

So you should definitely try combinations of different variables to get a map that is better than what you just got.

Bonus Exercise 4: Exploring the Susceptibility Map’s Sensitivity to Seasonality and Input Selection

josh-hild-N3e9vYJGZ1w-unsplash (1).jpg

Caption: The Liguria region (Cinque Terre), after you save it from raging wildfires using machine learning ✌

Part I: Seasonality

Q1) Using the season column of wildfires, separate your data into two seasonal datasets (1=Winter, 2=Summer).

Hint: When splitting your inputs into two seasonal datasets, keep in mind that temp_1 and prec_1 are the climatological mean temperatures and precipitation during winter, while temp_2 and prec_2 are the climatological mean temperature and precipitation during summer.

Q2) Use these two seasonal datasets to make the Liguria winter and summer susceptibility maps using your best classifier(s). What do you notice?

Hint: Feel free to recycle as much code as you can from the previous exercise. For instance, you may build a library of functions that directly train the classifier(s) and output susceptibility maps!

Part II: Input Selection

The details of the susceptibility map may strongly depend on the inputs you chose from the variables dataset. Here, we explore two different ways of selecting inputs to make our susceptibility maps as robust as possible.

Q3) Using your best classifier, identify the inputs contributing the most to your model’s performance using permutation feature importance.

Hint: You may use scikit-learn’s permutation_importance function using your best classifier as your estimator.

Q4) Retrain the same type of classifier only using the inputs you identified as most important, and display the new susceptibility map.

Hint: Feel free to recycle as much code as you can from the previous exercise. For instance, you may build a library of functions that directly train the classifier(s) and output susceptibility maps!

Can you explain the differences in susceptibility maps based on the inputs’ spatial distribution?

If the susceptibility map changed a lot, our best classifier may initially have learned spurious correlations. This would have affected our permutation feature importance analysis, and motivates re-selecting our inputs from scratch! 🔨

Q5) Use the SequentialFeatureSelector to select the most important inputs. Select as few as possible!

Hint: Track how the score improves as you add more and more inputs via n_features_to_select, and stop when it’s “good enough”.

Which inputs have you identified as the most important? Are they the same as the ones you selected using permutation feature importance?

Q6) Retrain the same type of classifier using as little inputs as possible, and display the new susceptibility map.

Hint: Feel free to recycle as much code as you can from the previous exercise. For instance, you may build a library of functions that directly train the classifier(s) and output susceptibility maps!

How does it compare to the authors’ susceptibility map below?

download (3).png