Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Open In Kaggle

One of the best ways to improve skills of machine learning models is to combine multiple machine learning models. Since each model in the ensemble will make slightly different predictions, the models could be more robust and generalizable to unseen data. We can also characterize the uncertainty of ML predictions with an ensemble approach.

Here we will create multiple individual classifiers on the MNIST data, a dataset with hand-written digit images, and perform skill evaluation. The skills of individual models will be compared with skills of ensemble models.

MNIST_Examples.png

The goal is to train and compare individual classifiers on MNIST data, before combining them into an ensemble model. Will the power of teamwork shine through? 🔢

Let’s start by loading the MNIST database!

(70000, 784)

Q1) Split the MNIST dataset into a training, a validation, and a test sets

Hint 1: The documentation for scikit-learn’s train_test_split function is at this link.

Hint 2: You may use 50k instances for training, 10k instances for validation, and 10k instances for testing.

Q2) Train various classifiers on the training set and compare them on the validation set

Hint: You may compare a RandomForestClassifier, an ExtraTreesClassifier, and a SVC, but we encourage you to be creative and include additional classifiers you find promising! The more the merrier 😀

Note from TA: The SVC can be slow to train. Test RandomForest and ExtraTrees first for quick results.

Now it’s time to make the individual classifiers vote to form an ensemble model

Q3) Combine the classifiers into an ensemble that outperforms them all on the validation set, using a soft or hard voting classifier.

Hint: The documentation for scikit-learn’s VotingClassifier class can be found at this link. Note that its argument voting can be changed from hard to soft.

Hint: If your ensemble does significantly worse than individual classifiers, consider deleting the individual classifiers negatively affecting the performance of your ensemble using del Voting_Classifier.estimators_[index_of_model_to_delete], where the estimators_ attribute of your Voting_Classifier’s lists the individual classifiers that were trained as part of the ensemble.

Q4) Does your ensemble clearly outperform your individual classifiers on the test set

Your voting classifier may only slightly beat the best model. Maybe voting isn’t the best way to get the best prediction!

Let’s try the brute-force approach: Training a classifier on the individual model’s predictions to beat the voting approach.

Bonus Exercise 3: From Individual Classifiers to Ensemble Stacking via Blenders

Blender.jpg

Let’s learn how to best blend the individual classifiers’ predictions!

Q1) Run the individual classifiers from the previous exercise to make predictions on the validation set, and create a new training set with the resulting predictions

Hint: The target stays the same, but now each training instance is a vector containing the set of predictions from all your individual classifiers. You may group all these vectors into a feature array X_val_predictions that should have the shape: (Number_of_validation_instances,Number_of_individual_classifiers).

Q2) Train a classifier on this new training set

Hint 1: You may train a RandomForestClassifier.

Hint 2: You could fine-tune this blender or try other types of blenders (e.g., a LogisticRegression or an MLPClassifier), then select the best one using cross-validation.

Congratulations! 😃

You have just trained a blender, and together with classifiers they form a stacking ensemble. Now let’s evaluate the ensemble on the test set.

Q3) Evaluate the blender on the test set and compare it to the voting classifier you trained earlier

Hint 1: You will have to first calculate the predictions of your individual classifiers on the test set, similar to what you did in Question 1.

Hint 2: Make sure you use the same score (e.g., the accuracy_score) to compare both ensemble models.

Is the blender worth the effort?