Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Open In Kaggle

Photo Credits: Three Assorted-Color Garbage Cans by Hamza Javaid licensed under the Unsplash License

Sorting takes a lot of effort - is there a way that we can get computers to do it automatically for us when we don’t even know where to begin?

This notebook will be used in the lab session for week 4 of the course, covers Chapter 9 of the old edition of Géron’s book (Chapter 8, “Unsupervised Learning Techniques”, in the current Hands-On Machine Learning with Scikit-Learn and PyTorch edition), and builds on the notebooks made available on Github.

Need a reminder of last week’s labs? Click here to go to notebook for week 3 of the course.

Notebook Setup

First, let’s import a few common modules, ensure MatplotLib plots figures inline and prepare a function to save the figures. We also check that Python 3.5 or later is installed (although Python 2.x may work, it is deprecated so we strongly recommend you use Python 3 instead), as well as Scikit-Learn ≥0.20.

Data Setup

We need to load the MNIST dataset from OpenML - we won’t be loading it as a Pandas dataframe, but will instead use the Dictionary / ndrray representation.

We’re going to subsample the digits in the dataset, choosing a random set of digit classes - we won’t even know the number of different digits we will chose; it will be somewhere between 4 and 8!

We will learn on a total of 8000 digits, evenly distributed amongst the randomly selected digits in the dataset. Let’s store the number of samples to be taken from each class.

With that out of the way, let’s generate the dataset!

We now have a set of 8000 random samples that we know belong to somewhere between 4 and 8 clusters 😃
Can we divide them into these groups without knowing the labels beforehand?

Warning: Don’t expect near perfect results this time.

Clustering with KMeans

The first thing we need to do is to import the KMeans model from scikit learn. Let’s go ahead and do so.

Q1) Import the KMeans model from scikit learn.

Hint: Here is the documentation for the Kmeans implementation in sklearn.

We don’t know how many clusters we should use to split the data. Our first instinct would be to use as many clusters as the number of digits we have, but even that is not necessarily optimal. Why dont we try all K’s between 4 and 40?

To do so, we’ll need to begin by training a Kmeans algorithm for each value of K we’re interested in.

How long does it take to train a single KMeans model with 10 initial centroid settings? This will give us an idea as to whether it may be a good idea to apply a dimensionality reduction algorithm before fitting our models.

Q2) Import python’s time library and measure how long it takes to train a single KMeans model with 3 clusters on the raw data subset.

Hint 1: Here is the documetation for the function used to get timestamps

This should seem like a bit longer than we want. (3.42 seconds during testing of the notebook). Doing this many times means there’s a possibility we’ll be sitting around doing nothing, possiblity for several minutes. Who has time for this?

Let’s reduce the dataset using PCA, capturing 95% of the variability in the data.

Q3) Import PCA from scikit and reduce the dimensionality of our input data. 95% of the variance in the data should be captured.

Hint 1: Here is the documentation for PCA.

Hint 2: .fit_transform() will be very useful

Let’s try training a KMeans model on the reduced dataset and see if our training time improved...

Q5) Repeat Q2 using the reduced dataset

Reminder: We’re still splitting the data into 3 clusters

Q6) Train a KMeans model for   2≤k≤20\; 2 \le k \le 20

Hint 1: Set up a range using python’s range function or numpy’s arange

Hint 2: You can store each trained model by appending it to a list as you iterate

You should hopefully remember the silhouette score and inertia metrics from the reading. Let’s go ahead and make a kk vs silhouettesilhouette plot and a kk vs inertiainertia plot.

Q7) Import the silhouette score metric and generate a silhouette score value for each model we trained

Hint 1: Here is the documentation for the Silhouette score implementation in scikit learn

Hint 2: The silhouette score needs the model input data and the model labels as arguments

Hint 3: The model labels are stored as an attribute in each model. Check the list of available attributes in the sklearn KMeans documentation.

Q8) Plot comparing kk vs Silhouette Score. Highlight the maximum score

Hint 1: You’ll need to find the position of the best score in the silhouette score list. Here is the documentation to a numpy function that would be very useful for this.

Hint 2: matplotlib’s pyplot has been imported as plt. Here is the documentation to the subplots() method.

Hint 3: Here is the documentation to the plot() method in matplotlib’s pyplot. Note that it is also implemented as a method in the axes objects created with plt.subplots()

We got this figure when we ran the code. Does your figure look the same?

s3_clustering1.png

Now we’ll plot the value of kk vs inertia

Q9) Plot comparing kk vs Inertia. Highlight the inertia for the model with the highest silhouette score

Hint 1: If you followed the previous step as it was written, you have the index for the best model stored inbest_index.

Hint 2: The KMeans documentation details the attribute in which the model’s intertia is stored.

You should see a figure similar to the one below.

s3_clustering1.png

If you ran the notebook with the default random seed, you may be surprised to see that the best performing KMeans model is the one that breaks it off into two clusters! The next best two will be those associated with k=4k=4 and k=5k=5. Let’s get some plots to try to make sense of the results, since we have the actual labels.

There aren’t any more questions to answer from here on out, but you will have to change the code if you started out from a different random seed! I’ll try to be good about pointing out what code you’ll need to change.

Let’s begin by using scikit’s TSNE implementation to reduce the dimensionality of the dataset for plotting. This will take a bit of computation time!

Let’s continue by making a list of the three best models.

And now, let’s get a set of predictions for each model and store it in a list!

Pandas will make producing a nice plot a lot simpler. Let’s import it and make a dataframe with the reduced input components, the truth labels, and the predicted cluster labels.

Note that the predicted labels don’t correspond to the digit labels since this is an unsupervised model!

And now we’ll make a nice, big 4x4 plot that allows us to see the true answers and how our algorithm clustered our data!

Our figure looks like this.

s3_clustering1.png

Assuming you started out from the intended random seed, you’ll be able to see that the “Best” model is trying to separate the 0 digits from the non-zero digits. The 2nd2^{nd} best model is able to separate the digits pretty well, but lumps 4s and 9s into a single cluster!

The 3rd3^{rd} best model begins to group digits into clusters that may not have much significance to us at first glance, but whose metrics seem to indicate worse performance and whose results would need further analysis to try to understand.