Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Open In Kaggle

Caption: Denise diagnoses an overheated CPU at our data center in The Dalles, Oregon.
For more than a decade, we have built some of the world's most efficient servers.


Photo from the Google Data Center gallery

Our world is increasingly filled with data from all sorts of sources, including environmental data. Can we reduce the data to a reduced, meaningful space to save on computation time and increase explainability?

This notebook will be used in the lab session for week 4 of the course, covers Chapter 8 of the old edition of Géron’s book (Chapter 7, “Dimensionality Reduction”, in the current Hands-On Machine Learning with Scikit-Learn and PyTorch edition), and builds on the notebooks made available on Github.

Need a reminder of last week’s labs? Click here to go to notebook for week 3 of the course.

Notebook Setup

First, let’s import a few common modules, ensure MatplotLib plots figures inline and prepare a function to save the figures. We also check that Python 3.5 or later is installed (although Python 2.x may work, it is deprecated so we strongly recommend you use Python 3 instead), as well as Scikit-Learn ≥0.20.

Dimensionality Reduction using PCA

This week we’ll be looking at how to reduce the dimensionality of a large dataset in order to improve our classifying algorithm’s performance! With that in mind, let’s being the exercise by loading the MNIST dataset.

Q1) Load the input features and truth variable into X and y, then split the data into a training and test dataset using scikit’s train_test_split method. Use test_size=0.15, and remember to set the random state to rnd_seed!

Hint 1: The 'data' and 'target' keys for mnist will return X and y.

Hint 2: Here’s the documentation for train/test split.

We now once again have a training and testing dataset with which to work with. Let’s try training a random forest tree classifier on it. You’ve had experience with them before, so let’s have you import the RandomForestClassifier from sklearn and instantiate it.

Q2) Import the RandomForestClassifier model from sklearn. Then, instantiate it with 100 estimators and set the random state to rnd_seed!

Hint 1: Here’s the documentation for RandomForestClassifier

Hint 2: Here’s the documentation for train/test split.

Hint 3: If you’re still confused about instantiation, there’s a blurb on wikipedia describing it in the context of computer science.

We’re now going to measure how quickly the algorithm is fitted to the mnist dataset! To do this, we’ll have to import the time library. With it, we’ll be able to get a timestamp immediately before and after we fit the algorithm, and we’ll get the time by calculating the difference.

Q3) Import the time library and calculate how long it takes to fit the RandomForestClassifier model.

Hint 1: Here’s the documentation to the function used for getting timestamps

Hint 2: Here’s the documentation for the fitting method used in RandomForestClassifier.

We care about more than just how long we took to trian the model, however! Let’s get an accuracy score for our model.

Q4) Get an accuracy score for the predictions from the RandomForestClassifier

Hint 1: Here is the documentation for the accuracy_score metric in sklearn.

Hint 2: Here is the documentation for the predict method in RandomForestClassifier

Let’s try doing the same with with a logistic regression algorithm to see how it compares.

Q5) Repeat Q2-4 with a logistic regression algorithm using sklearn’s LogisticRegression class. Hyperparameters: solver='lbfgs' (multinomial is the default behavior for multiclass problems with this solver)

*Hint 1: Here is the documentation for the LogisticRegression class.

Up to now, everything that we’ve done are things we’ve done in previous labs - but now we’ll get to try out some algorithms useful for reducing dimensionality! Let’s use principal component analysis. Here, we’ll reduce the space using enough axes to explain over 95% of the variability in the data...

Q6) Import scikit’s implementation of PCA and fit it to the training dataset so that 95% of the variability is explained.

Hint 1: Here is the documentation for scikit’s PCA class.

Hint 2: Here is the documentation for scikit’s .fit_transform() method.

Q7) Repeat Q3 & Q4 using the reduced X_train dataset instead of X_train.

Q8) Repeat Q5 using the reduced X_train dataset instead of X_train.

You can now compare how well the random forest classifier and logistic regression classifier performed on both the full dataset and the reduced dataset. What were you able to observe?

Write your comments on the performance of the algorithms in this box, if you’d like 😀 (Double click to activate editing mode)