Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Open In Kaggle

This week’s notebook is based off of the exercises in Chapter 4 of Géron’s book.

Notebook Setup

Let’s begin like in the last notebook: importing a few common modules, ensuring MatplotLib plots figures inline and preparing a function to save the figures. We also check that Python 3.5 or later is installed (although Python 2.x may work, it is deprecated so once again we strongly recommend you use Python 3 instead), as well as Scikit-Learn ≥0.20.

You don’t need to worry about understanding everything that is written in this section.

/home/runner/work/2026_MLEES_book/2026_MLEES_book/.venv/bin/python: No module named pip
Note: you may need to restart the kernel to use updated packages.

Data Setup

In this notebook we will be working with the Palmer Penguins dataset. Each entry in the dataset includes the penguin’s species, island, sex, flipper length, body mass, bill length, bill depth, and the year the study was carried out. Let’s take a moment and observe our subjects!

🐧


In order: Adélie (Pygoscelis adeliae), Chinstrap (Pygoscelis antarcticus), and Gentoo (Pygoscelis papua) penguins


As you can imagine, this dataset is normally used to train multiclass/multinomial classification algorithms and not binary classification algorithms, since there are more than 2 classes.

“Three classes, even!” - an observant TA

For this exercise, however, we will implement the binary classification algorithm referred to as the logistic regression algorithm (also called logit regression).

Like with the Titanic dataset in the previous notebook, the data here is loaded as a Pandas DataFrame. Feel free to play around with it in the cell below!

As we mentioned before, there are three species of penguin in the dataset. However, today we’ll be implementing a binary classification algorithm, which means we need to have exactly two target classes! Let’s go ahead and filter the data so that we keep the Adelie and Gentoo species.

We now have a dataframe with all the information that we need. Let’s go ahead and extract the bill length and depth to use as input data, storing it in xx. Then we’ll store the labels (i.e., the targets) in yy.

Q1) Extract the bill length and bill depth to use as the input vector xx, and store the label (i.e., the target data) in yy

We have our input data, but we need a target to predict. We previously filtered the data to only include Adélie and Gentoo penguins, but we still have them as strings! Let’s convert them to a binary representation (i.e., 0 or 1). Make sure you have the same penguins as in your input!

Q2) Convert the species label to a binary classification, and filter the target data to match the input data.

We now have a set of binary classification data we can use to train an algorithm.

As we saw during our reading, we need to define three things in order to train our algorithm:

⋅\cdot the type of algorithm we will train, \ ⋅\cdot the cost function (which will tell us how close our prediction is to the truth), and \ ⋅\cdot a method for updating the parameters in our model according to the value of the cost function (e.g., the gradient descent method).

Let’s begin by defining the type of algorithm we will use. We will train a logistic regression model to differentiate between two classes. A reminder of how the logistic regression algorithm works is given below.


The logistic regression algorithm will thus take an input tt that is a linear combination of the features:

$t_{\small{n}} = \beta_{\small{0}} + \beta_{\small{1}} \cdot X_{1,n} + \beta_{\small{2}} \cdot X_{2,n}$

where

  • nn is the ID of the sample

  • X0X_{\small{0}} represents the bill length

  • X1X_{\small{1}} represents the bill width

This input is then fed into the logistic function, σ\sigma:

σ:t↦11+e−t\begin{align} \sigma: t\mapsto \dfrac{1}{1+e^ {-t}} \end{align}

Let’s define the logistic function for later use.

Q3) Define the logistic function

Now that the logistic function has been defined, we can plot it (this will help us remember what it looks like!) Run the code below - you won’t have to fill anything in for this one 😀 But feel free to show the code and read through it - some of the functions used can be helpful to you down the line!

With the logistic function, we define inputs resulting in σ≥0.5\sigma\geq0.5 as belonging to the one class, and any value below that is considered to belong to the zero class.

We now have a function which lets us map the value of the bill length and width to the class to which the observation belongs (i.e., whether the length and width correspond to Adélie or Gentoo penguins). However, there is a parameter vector θ\theta with a number of parameters that we do not have a value for:
θ=[β0,β1\theta = [ \beta_{\small{0}}, \beta_{\small{1}}, β2]\beta_{\small{2}} ]

Q4) Set up an array of random numbers between 0 and 1 representing the θ\theta vector.

In order to determine whether a set of β\beta values is better than the other, we need to quantify well the values are able to predict the class. This is where the cost function comes in.

The cost function, cc, will return a value close to zero when the prediction, p^\hat{p}, is correct and a large value when it is wrong. In a binary classification problem, we can use the log loss function. For a single prediction and truth value, it is given by:

c(p^,y)={−log⁡(p^)if  y=1−log⁡(1−p^)if  y=0\begin{align} \text{c}(\hat{p},y) = \left\{ \begin{array}{cl} -\log(\hat{p})& \text{if}\; y=1\\ -\log(1-\hat{p}) & \text{if}\; y=0 \end{array} \right. \end{align}

However, we want to apply the cost function to an n-dimensional set of predictions and truth values. Thankfully, we can find the average value of the log loss function JJ for an an-dimensional set of y^\hat{y} & yy as follows:

J(p^,y)=−1n∑i=1n[yi⋅log⁡(p^i)]+[(1−yi)⋅log⁡(1−p^i)]\begin{align} \text{J}(\mathbf{\hat{p}},y) = - \dfrac{1}{n} \sum_{i=1}^{n} \left[ y_i\cdot \log\left( \hat{p}_i \right) \right] + \left[ \left( 1 - y_i \right) \cdot \log\left( 1-\hat{p}_i \right) \right] \end{align}

We now have a formula that can be used to calculate the average cost over the training set of data.

Now let’s code 💻

Q5) Define a log_loss function that takes in an arbitrarily large set of prediction and truths

Hint 1: You need to encode the function JJ above, for which Numpy’s functions may be quite convenient (e.g., log, mean, etc.)

Hint 2: Asserting the dimensions of the vector is a good way to check that your function is working correctly. Here’s a tutorial on how to use assert. For instance, to assert that two vectors X and y have the same dimension, you may use:

assert X.shape==y.shape

We now have a way of quantifying how good our predictions are. The final thing needed for us to train our algorithm is figuring out a way to update the parameters in a way that improves the average quality of our predictions.



Warning: we’ll go into a bit of math below

Let’s look at the change in a single parameter within θ\theta: β1\beta_1 (given X1,i=X1X_{1,i} = X_1,   p^i=p^\;\hat{p}_{i} = \hat{p},   yi=y\;y_{i} = y). If we want to know what the effect of changing the value of β1\beta_1 will have on the log loss function we can find this with the partial derivative:

$ \dfrac{\partial J}{\partial \beta_1} $

This may not seem very helpful by itself - after all, β1\beta_1 isn’t even in the expression of JJ. But if we use the chain rule, we can rewrite the expression as:

$\dfrac{\partial J}{\partial \hat{p}} \cdot \dfrac{\partial \hat{p}}{\partial \theta} \cdot \dfrac{\partial \theta}{\partial \beta_1}$

We’ll spare you the math (feel free to verify it youself, however!):

$\dfrac{\partial J}{\partial \hat{p}} = \dfrac{\hat{p} - y}{\hat{p}(1-\hat{p})}, \quad \dfrac{\partial \hat{p}}{\partial \theta} = \hat{p} (1-\hat{p}), \quad \dfrac{\partial \theta}{\partial \beta_1} = X_1 $

and thus

$ \dfrac{\partial J}{\partial \beta_1} = (\hat{p} - y) \cdot X_1 $

We can calculate the partial derivative for each parameter in θ\theta which, as you may have realized, is simply the θ\theta gradient of JJ: ∇θ(J)\nabla_{\theta}(J)

With all of this information, we can now write ∇θJ\nabla_{\theta} J in terms of the error, the feature vector, and the number of samples we’re training on!

$\nabla_{\mathbf{\theta}^{(k)}} \, J(\mathbf{\theta^{(k)}}) = \dfrac{1}{n} \sum\limits_{i=1}^{n}{ \left ( \hat{p}^{(k)}_{i} - y_{i} \right ) \mathbf{X}_{i}}$

Note that here kk represents the iteration of the parameters we are currently on.

We now have a gradient we can calculate and use in the batch gradient descent method! The updated parameters will thus be:

θ(k+1)=θ(k)−η ∇θ(k)J(θ(k))\begin{align} {\mathbf{\theta}^{(k+1)}} = {\mathbf{\theta}^{(k)}} - \eta\,\nabla_{\theta^{(k)}}J(\theta^{(k)}) \end{align}

Where η\eta is the learning rate parameter. It’s also worth pointing out that   p^i(k)=σ(θ(k),Xi)\;\hat{p}^{(k)}_i = \sigma\left(\theta^{(k)}, X_i\right)

In order to easily calculate the input to the logistic regression, we’ll multiply the θ\theta vector with the X data, and as we have a non-zero bias β0\beta_0 we’d like to have an X matrix whose first column is filled with ones.

Xwith bias=(1X1,0X2,01X1,1X2,1...1X1,nX2,n)\begin{align} X_{\small{with\ bias}} = \begin{pmatrix} 1 & X_{1,0} & X_{2,0}\\ 1 & X_{1,1} & X_{2,1}\\ &...&\\ 1 & X_{1,n} & X_{2,n} \end{pmatrix} \end{align}

Q6) Prepare the X_with_bias matrix.

Our X_with_bias matrix looks like this: \ [[ 1. \quad -0.69346042 \quad 0.92572752] \ [ 1. \quad -0.6164717 \quad 0.28005659] \ [ 1. \quad -0.46249427 \quad 0.57805856] \ [ 1. \quad -1.15539273 \quad 1.22372949] \ [ 1. \quad -0.65496606 \quad 1.86940041] \ [ 1. \quad -0.73195478 \quad 0.47872457] \ [ 1. \quad -0.67421324 \quad 1.37273047] \ [ 1. \quad -1.65581939 \quad 0.62772555] \ [ 1. \quad -0.13529222 \quad 1.67073244] \ [ 1. \quad -0.94367375 \quad 0.13105561]]

Q7) Write a function called predict that takes in the parameter vector θ\theta and the X_with_bias matrix and evaluates the logistic function for each of the samples.

If everything is set up correctly and you didn’t change the debug data and theta, the output for your predict function should be:

[0.35434369 0.46118934 0.57172409 0.67553632 0.76454801]

Q8) Now that you have a predict function, write a gradient_calc function that calculates the gradient for the logistic function.

Hint: You’ll have to feed theta, X, and y to the gradient_calc function.

Hint: You can use this equation to calculate the gradient of the cost function.

If you kept the same dummy data we included by default in the notebook, you should get [ 0.16546829 -0.19307376 -0.15630302] as the output of your gradient calculator! 💻

We can now write a function that will train a logistic regression algorithm!

Your logistic_regression function needs to:

  • Take in a set of training input/output data, validation input/output data, a number of iterations to train for, a set of initial parameters θ\theta, and a learning rate η\eta

  • At each iteration:

  • Generate a set of predictions on the training data. Hint: You may use your function predict on inputs X_train from the training set.

  • Calculate and store the loss function for the training data at each iteration. Hint: You may use your function log_loss on inputs X_train and outputs y_train from the training set.

  • Calculate the gradient. Hint: You may use your function grad_calc.

  • Update the θ\theta parameters. Hint: You need to implement this equation.

  • Generate a set of predictions on the validation data using the updated parameters. Hint: You may use your function predict on inputs X_valid from the validation set.

  • Calculate and store the loss function for the validation data. Hint: You may use your function log_loss on inputs X_valid and outputs y_valid from the validation set.

  • Bonus: Calculate and store the accuracy of the model on the training and validation data as a metric!

  • Return the final set of parameters θ\theta & the stored training/validation loss function values (and the accuracy, if you did the bonus)

Q9) Write the logistic_regression function

¡¡¡Important Note!!!

The notebook assumes that you will return

  1. a Losses list, where Losses[0] is the training loss and Losses[1] is the validation loss

  2. a tuple with the 3 final coefficients (β0\beta_0, β1\beta_1, β2\beta_2)


Now that we have our logistic regression function, we’re all set to train our algorithm! Or are we?

There’s an important data step that we’ve neglected up to this point - we need to split the data into the train, validation, and test datasets.

train ✂️ validation ✂️ test

Now we’re ready!

Q10) Train your logistic regression algorithm. We recommend you use 500 iterations, η\eta=0.1

Hint: It’s time to use the logistic_regression function you defined in Q5.

Let’s see how our model did while learning!

Congratulations on training a logistic regression algorithm from scratch!

Your loss function graph should look something similar to this...

And the accuracies we got during development of the notebook are:

Training Accuracy:99.4%
Validation Accuracy:100.0%
Test Accuracy:100.0%

Once you’re done with the upcoming environmental science applications notebook, feel free to come back to take a look at the challenges 😀

Challenges

  • C1) Add more features to try to improve our accuracies!

  • C2) Add early stopping to the training algorithm! (e.g., stop training when the accuracy is greater than a target accuracy)