Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Open In Kaggle

image.png

“A Black Box Model is a system that does not reveal its internal mechanisms. In machine learning, “black box” describes models that cannot be understood by looking at their parameters (e.g. a neural network). The opposite of a black box is sometimes referred to as White Box, and is referred to in this book as interpretable model. Model-agnostic methods for interpretability treat machine learning models as black boxes, even if they are not.”

Can you enhance the trustworthiness of machine learning models by employing eXplainable Artificial Intelligence methods to challenge their black-box nature?

Image/Quote source: Christopher Molnar’s “Interpretable Machine Learning” book

For this chapter’s first tutorial, our focus will be on implementing fundamental techniques that shed light on the process through which machine learning models generate predictions. The ability to elucidate the how behind these predictions is pivotal for enhancing the trustworthiness of the models. Moreover, this exploration may unveil opportunities for new scientific discoveries along the journey!

Upon completing this tutorial, you will gain proficiency in:

  1. Implementing permutation feature importance and partial independence plots.

  2. Applying these elucidation tools to linear models, tree models, and neural networks.

  3. Acquiring a foundational understanding of Shapely values.

  4. Mastering the utilization of the SHAP package to explicate various machine learning models.

Throughout this tutorial, we will apply eXplainable Artificial Intelligence (XAI) plots, such as Partial Dependence Plots (PDPs), on straightforward ML models trained with sample datasets. Subsequently, in the following tutorial, we will put our acquired knowledge into practice by employing XAI tools on a real-world environmental science dataset.

I) Partial Dependence Plots and Permutation Feature Importance

We will first load a simple tabular datafile containing data about the Titanic. 🛥

Here's a data sample. You can copy the row header text from here if you need it later 😉
Loading...

The data table contains records of passengers onboard the Titantic cruise ship. Included in the table are the personal details of different passengers, the fare and cabin class they are in, and whether or not they managed to survive the disaster.

Based on this dataset, we will train a simple classification model to determine if a given passenger can survive the incident, based on available information on his/her age, ticket fare, etc.

This exercise is based on a Kaggle Competition Notebook series (https://www.kaggle.com/code/dansbecker/partial-dependence-plots)

Q1: Train a simple machine learning model on the Titanic Dataset and measure its accuracy

Our problem can be phrased as a binary classification problem that requires a ML model to give binary outputs on whether a passenger can survive the disaster or not.

Remember that the code snippets below are suggestions and that we encourage you to be creative in how you solve this question 🎨

Hints:

  1. We will use the ‘PassengerID’, ‘Age’, and ‘Fare’ columns in the titanic_df Pandas DataFrame to predict the ‘Survived’ column.

  2. We will need to standardize our input data before training our models. We recommend using the SimpleImputer module from sklearn.impute for this task. (Reference link)

  3. Use train_test_split to split the standardized inputs and outputs into training set, validation set, and test set.

  4. Use either GradientBoostingClassifier or RandomForestClassifier to train accurate binary classifiers. You may use a baseline LogisticRegression.

The accuracy score for the classifier may not be terribly good (we got 0.66 with a GradientBoostingClassifier).

We encourage you to test different ways to tune the model hyperparameters to get a better classification performance than the model we just created, for instance using GridSearchCV 🤖

The first tool of ML explainability is Partial Dependence Plot (PDPs). Simply put, partial dependence plots summarizes all possible ML model outputs when only one feature in the input is perturbed.

The PDP function for regression tasks can be defined as (Molnar Ch. 8.1):

f^S(xS)=EXC[f(xS,XC)]=∫f^(xS,XC) dP(XC)\hat{f}_S(x_S) = \mathbb{E}_{X_C} \left[ f(x_S, X_C) \right] = \int \hat{f}(x_S, X_C) \, dP(X_C)

where xSx_{S} is the input feature we would like to perturb, XCX_{C} include the reminding features in the input dataset, and f^()\hat{f}() representing the trained machine learning model.

image.png

From Molnar’s interpretable ML book: PDPs for the bicycle count prediction model and temperature, humidity and wind speed. The largest differences can be seen in the temperature. The hotter, the more bikes are rented. This trend goes up to 20 degrees Celsius, then flattens and drops slightly at 30. Marks on the x-axis indicate the data distribution.

Luckily, there are existing packages to calculate PDPs so we do not need to start from scratch!

In the second question, we ask you to produce a simple PDP to show how the three features in the Titanic Dataset affects the ML model predictions.

Q2: Produce a Partial Dependence Plot with the trained classification model

Complete the get_pdp_values_shap function, get the PassengerId, Age and Fare PDPs, and plot them in a 1x3 matplotlib panel plot.

Hints:

(1) The SHAP package have a nice function shap.partial_dependence_plot to calculate PDPs. However, it does not work well with matplotlib subplots. \ (2) Here, we ask you to complete a function to get data from the PDP plots. \ (3) The function should return six arrays as outputs. These arrays correspond to the x-coordinates and y-coordinates of the PDP curve, as well as the coordinates of the median values of the data and the coordinates of the prediction based on the median data.

<Figure size 1900x400 with 3 Axes>
image.png

For reference, above are the partial dependence plots we got when using a Gradient Boosting Classifier.

What can you conclude from the three PDPs that you just created?

What variable has a positive impact on the survival chance of a particular passenger? What variable has little to no impact?

Permutation Feature Importance is another type of model-agnostic method that can be used to explain different machine learning models.

Both PDPs and Permutation Feature Importance involves summarizing all model predictions when we perturb the input feature, with the goal of isolating the effect of the perturbed feature on the predictions.

The interested feature in the permutation importance is perturbed differently to that used to create PDPs. Permutation Feature Importance is defined as the deterioration in model performance when the interested feature in shuffled randomly.

Q3: Calculate the mean permutation feature importance for all three input features

Hint:

  1. We will import the permutation_importance module from sklearn.inspection. (Reference link)

  2. We repeat the calculation 30 times to make sure the values we get are valid. This is determined with the n_repeats parameter in the permutation_importance function.

  3. The mean and standard deviation of the permutation importance values can be accessed from a permutation_importance object with two attributes: .importances_mean and .importances_std.

II) Introduction to SHAP

image.png

In this section, we will introduce another popular library for ML Explainability: SHAP. The usefulness of the SHAP package relies on using it to calculate and analyze Shapely values.

Shapely value is a useful concept from comparatively game theory that is applied to machine learning context. In order to link machine learning models and game theory, we need to treat the prediction task for each instance of the dataset as a game. The Shapely value for an individual input feature - anologous to players in the game - quantifies the average affect of adding that feature to the model has on the game outcome, i.e., the difference between a prediction on a sample made with the interested input and one without.

For a simple multiple linear regression model,

f^=β0+β1x1+β2x2+...+βixi\hat{f} = \beta_0+\beta_1 x_1+\beta_2 x_2 + ... + \beta_i x_i

the contribution ϕj\phi_j for the j-th feature in the input to the game outcome is:

ϕj(f^)=βjxj−βjE(Xj)\phi_j (\hat{f}) = \beta_j x_j - \beta_j E(X_j)

where βjE(Xj)\beta_j E(X_j) is the mean effect estimate of the feature, i.e., the prediction when the feature value is not known. For a linear model, we can directly read the shapely value from the PDPs.

In the SHAP package, the shapely value for sample 50 in the dataset is simply the difference between the values for that sample in the PDP and the ML prediction made with the median values of the MedInc feature.

One of the more attractive properties of shapely value is its additive properties.

If we sum the effect of all the features in a linear model, the result for a particular sample is simply its prediction values minus the mean prediction over the entire dataset.

∑j=1pϕj(f^)=f^(x)−E(f^(X))\sum_{j=1}^p \phi_j (\hat{f}) = \hat{f}(x) - E(\hat{f}(X))

We can use a waterfall plot to summarize all the contributions ϕj\phi_j from different input features.

The above plot summarizes how each input feature in the housing dataset impact the prediction for the fiftieth sample in the dataset. From the above plot, we also see that only a subsample of the input features yield non-negligible effect on the ML prediction.

Q4: Train an XGBoost Classifier to predict wine quality

Now let us try to put what we have learned in practice and create some of these beautiful SHAP plots!

First, we will train a tree-based XGBoost classifier to predict wine quality from different wine properties.

image.png

Can you predict the quality of a wine based on its chemical properties?

Image source: eyetronic

Here's a data sample. You can copy the row header text from here if you need it later 😉
Loading...

You will notice that the Quality column contains various integer values. To simplify our problem and convert it into a straightforward binary classification task, we’ll transform this column as follows:

  • If (Quality<5)\left( \text{Quality} < 5 \right), we’ll label it as 0, indicating “bad wine.”

  • If (Quality≥5)\left( \text{Quality} \geq 5 \right), we’ll label it as 1, signifying “good wine.”

This transformation creates a binary classification problem, where the goal is to classify wines as either “bad” or “not bad” based on the quality scores.

Q4a: Data Processing

We recommend following the steps below:

(1) Extract quality column as the output (y) dataset. In the mean time, remove the quality column from the input (X) dataset.

(2) Convert the quality column into 0 and 1 with the above criteria.

(3) Use train_test_split to split the data into training and test set. We will use a 70%-30% split here.

Q4b: Initiatize and train an XGBClassifier

Hints:

(1) Refer the XGBClassifier tutorial link if you want to learn more about training a XGBoost model.

(2) We will set the XGBClassifier objective as binary:logistic.

(3) Evaluate model performance with the test set. Report two performance scores: Accuracy score and f1_score.

Reference for these performance scores:

(a) Accuracy Score

(b) F1 Score

Q5: Train a TreeExplainer on the XGBoost Classifier and calculate the SHAP values for the X_test dataset

(1) Train a TreeExplainer (Reference) on the trained XGBoost Classifier model.

(2) Apply the Explainer to the test dataset, and save the calculated SHAP values.

Q6: Produce two different plots to summarize the mean absolute SHAP values for all input features

For this problem, we will try to create some plots that explain the entire dataset, i.e., global explanation methods.

In the previous exercise, we introduced PDPs and Permutation Importance, which are global methods. Several more advanced plots are available in the SHAP package to above us to summarize more information in pretty plots!

We recommend experimenting with bar plots and beeswarm plots. You can refer to the following tutorial if you have doubts.

Bar plots

Beewsarm plots

Q7: Use two different types of plots - waterfall plots and force plots - to explain how the prediction for the second sample is made

Now, we turn our focus to more detailed, sample-based explainations. We briefly introduced the waterfall plots in the tutorial. We will now try to create such plots and one more type of plots - the force plots - to explain the contributions of different inputs to the prediction.

As an example, we will try to explain the prediction for the second sample.

Reference for the two plots:

Waterfall plots

Force Plots

Hint: The Force plots require the javascripts library to be displayed. We need to use the following code: shap.initjs() to initialize javescript in colab.

Q8: What do the numbers in the force plots represent?

In the force plot and waterfall plot we’ve just examined, we gained insights into how various factors collaborate to steer the model’s prediction. For the second element (index 1), we go from an average expectation of 0.103 to a final output of -4.37.

Now, considering our model is designed as a binary classifier, exclusively capable of yielding outputs of 0 or 1, it prompts a natural question:

Why does SHAP provide explanations for an output like -4.37 for the test set’s second element?

Let’s delve into this apparent paradox to enhance our understanding 🕵

(1) Utilize the .predict_proba() method for the test set’s element you investigated above

(2) Implement the Logistic function (Reference link) to transform one of the values obtained in Step (1).

What do you see?

III) SHAP DeepExplainer for Neural Networks

In this last part, our goal is to train a CNN on the MNIST dataset, and apply SHAP’s DeepExplainer on the trained model to understand the pixels it uses to make accurate predictions 🔢

image.png

Which part of the image does SHAP use to accurately classify digit pictures? Source

Q9: Design a CNN model on the MNIST dataset

Here, your task is to design a CNN model named Net() on the MNIST dataset.

There is flexibility in how the model architecture should be like. You can directly apply some CNN architectures on MNIST that we played with in previous exercise or some architectures available online here.

The only instruction here is that the model should end with a linear layer that output 10 values that are converted into probability after passing through a Softmax layer for multi-class classification.

We have already prepared some ready made functions to train the model and report the performance skill of the trained model on the test set.

Q10: Train the CNN model

(1) Use the train() function to train your CNN model.

(2) You can experiment with the optimizer to get better skills. Especially the learning rate lr=, and if you are using SGD, the momentum momentum=.

(3) Make sure the model is accurate enough to have learned reliable strategies to classify digits.

Q11: Train a DeepExplainer on your CNN model

(1) Train a SHAP DeepExplainer with the trained CNN model. We will use the first 100 images in the test set as the training data.

(2) Use the DeepExplainer to explain the predictions for the test images with the following index: 100, 101, 102.

(3) Use Image plots to visualize the explanation.

Below are examples of attribution maps we obtained for 0 and 1 classifications.

image.png