Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Open In Kaggle

Now that you have learned the basics of Support Vector Machines and tuning regularization parameters can make the trained SVMs more generalizable, it is time to learn how to create and train a simple SVM.

We will start with a sample dataset including measurements of different physical characteristics of flowers. We would like to train a support vector machine to automatically differentiate two different types of flowers. After training our first SVM model, we will make additional experiments to see how the decision boundaries of regularized versions of trained SVMs differ from less regularized ones.

Goal: Building similar models based on different types of Support Vector Machines (SVMs) to classify linearly separable classes, here Iris Setosa and Iris Versicolor from the Iris dataset.

christina-brinza-TXmV4YYrzxg-unsplash.jpg

Caption: Iris flowers in the evening light. Are they Irises Setosa or Irises Versicolor?

Source: Photo by Christina Brinza on Unsplash

First, let’s load the Iris dataset! 💐

Now we have our pre-processed dataset 💐:

Our features are (petal length, petal width) in X.

Our target is (Iris species) in y.

Q1) Train a Linear Support Vector Classification model on the pre-processed dataset

Hint: The documentation for LinearSVC is at this link

Q2) Plot the decision boundary of this classifier

Hint: According to the documentation, given a SVC object svc:

  • Weights: W = svc.coef_[0], and

  • Intercept: I = svc.intercept_

the decision boundary is the line:

yboundary=−W[0]x+I[0]W[1]y_{boundary} = -\frac{W\left[0\right] x + I\left[0\right]}{W\left[1\right]}

⚠ If you normalized your inputs before feeding it to the SVM in the previous question (e.g., via the StandardScaler), the equation above is only valid in “normalized” coordinates.

Now show the decision boundary that you just got in a scatter plot. Can it cleanly separate different flowers?

Hint: (1) We will use plt.scatter() to plot the flower data. Check documentation for details. (2) We will need to initiate a X array to plot the decision boundary. There are many ways to create such an array, but let’s use np.linspace() for now. Documentation

Q3) Train a SVC and a SGDClassifier for the same task and compare these two models to the LinearSVC. Use kernel = 'Linear' when instantiating the SVC!

Hint: Here is the documentation for the SVC class and the SGDClassifier class.

Q4) Create more regularized versions of each model and compare these new models to the previous ones

Hint: Vary the hyperparameter C for the LinearSVC and SVC models, and vary the hyperparameter alpha for the SGDClassifier model. Consult the documentation to know whether to increase or decrease the regularization parameters.

How does regularization affect the decision boundary in this simple case?

Bonus Exercise 1: Training a SVM Regressor on the California Housing Dataset

CA_house_beach (2).jpg

Can we use SVMs to predict the price of a house in California (in 1990) based on its characteristics (longitude, latitude, total_rooms, etc.)?

The dataset was originally used in Pace, R. Kelley, and Ronald Barry. “Sparse spatial autoregressions.” Statistics & Probability Letters 33.3 (1997): 291-297.

Let’s first load the dataset using Scikit-Learn’s fetch_california_housing() function:

Let’s split the data into a training set and a test set:

Q1) Normalize the features X before training the regressor

Hint: You may use the StandardScaler to normalize X using its z-score.

Q2) Start by training a simple Linear Support Vector Regression and assess its performance on the training and test sets

Hint 1: Here’s the documentation for scikit-learn’s LinearSVR

Hint 2: You may assess the regressor’s performance using the any scikit-learn regression metric you find interpretable.

Q3) How large is the model error in $?

Hint: The unit of y in the dataset is $10,000.

Q4) Try to beat your linear model using more complex SVM regressors.

Hint 1: The performance of a model should be assessed using the test set.

Hint 2: Géron’s model uses a SVR for which the hyperparameters gamma and C were optimized using a Randomized Search, and gets the root mean-squared error down to approximately 6,000 dollars on the test dataset.