Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Open In Colab Open In Kaggle

EuroSAT overview image

This notebook is converted from the source: https://www.coursera.org/learn/getting-started-with-tensor-flow2

Exercise Instruction

By the end of this notebook, you’ll be able to 😃😃😃

  1. Construct CNNs that classifies EuroSAT images into one of its 10 classes;

  2. Save and load trained models;

  3. Explore ways to improve the model performance.


Land Cover Classification aims to automatically provide labels describing the represented physical land type or how a land area is used (e.g., residential, industrial).

Convolutional Neural Networks (CNNs), the state of-the-art image classification method in computer vision and machine learning, have been reported to be suitable for the classification of remotely sensed images.

However, the classification of remotely sensed images is a challenging task, particularly due to the lack of reliably labeled ground truth datasets.

The EuroSAT dataset provides large quantity of training data for this purpose. It consists of 27000 labelled Sentinel-2 satellite images of different land uses: residential, industrial, highway, river, forest, pasture, herbaceous vegetation, annual crop, permanent crop and sea/lake.

For a reference, see the following papers:

  • Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. Patrick Helber, Benjamin Bischke, Andreas Dengel, Damian Borth. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2019.

  • Introducing EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification. Patrick Helber, Benjamin Bischke, Andreas Dengel. 2018 IEEE International Geoscience and Remote Sensing Symposium, 2018.


⚡⚡⚡ You can create your own code by following each question or complete the code with blanks.


Data Setup

Using the EuroSAT dataset which consists of 27000 images and labels might use too much memory, thus we use a smaller subset of the original dataset - 4000 training images and 1000 testing images with roughly equal numbers of each class.

Since PyTorch has no .compile()/.fit()/.evaluate(), we’ll define three small reusable helpers before building any models — one to train a model for a number of epochs (optionally with callbacks and a validation set), one to evaluate a model on a dataset, and a History class that mimics Keras’s history.history dict-of-lists so the rest of this notebook can stay close to the original.

Q1 Build a CNN model1 to classify Eurosat data.

Let’s construct a CNN called model1 using nn.Sequential, according to the following specifications:

  • The first layer should be a Conv2d layer with 16 filters, a 3x3 kernel size, and ‘same’ padding. Name this layer ‘conv_1’.

  • Follow it with a ReLU activation, named ‘relu_1’.

  • The third layer should also be a Conv2d layer with 8 filters, a 3x3 kernel size, and ‘same’ padding. Name this layer ‘conv_2’.

  • Follow it with a ReLU activation, named ‘relu_2’.

  • The fifth layer should be a MaxPool2d layer with a pooling window size of 8x8. Name this layer ‘pool_1’.

  • The sixth layer should be a Flatten layer, named ‘flatten’.

  • The seventh layer should be a Linear layer with 32 units, followed by a ReLU activation. Name these layers ‘dense_1’ and ‘relu_3’.

  • The final layer should be a Linear layer with 10 units and no activation (unlike Keras, nn.CrossEntropyLoss expects raw logits, not softmax probabilities). Name this layer ‘dense_2’.

Hint 1: Unlike keras.models.Sequential, nn.Sequential doesn’t have an .add() method — but if you pass it an OrderedDict of (name, layer) pairs, you get the same named-layer access (model1.conv_1, etc.) that Keras’s name= argument gives you.

Hint 2: nn.Conv2d needs an explicit number of input channels — 3 for the first layer (RGB), and conv_2’s input channels equal conv_1’s output channels.

Hint 3: As in the artificial-neural-networks and deep-computer-vision exercises, nn.LazyLinear can save you from computing the flattened size by hand.

❓❓❓ Do you have a model of the following structure?

(This screenshot is from the original Keras version of this exercise — print(model1) will show a differently formatted, but structurally equivalent, layer list.)

Q2 Set up model1’s loss function, optimizer, and an evaluation metric.

  • Use the Adam optimiser, cross entropy loss function, and a single accuracy metric.

Q3 Evaluate the initial accuracy of model1: Is the initial accuracy of the model as you have expected?

❓❓❓ Does your model have a similar initial accuracy & why?

Q4 Train model1 with 15 epochs, store the result in variable ‘history’;

  • Store the fitting result in a variable history;

❓❓❓ Do you have similar print output?

(Again, expect a differently formatted but comparable progress report — there’s no Keras-style progress bar here.)

Q5 Evaluate the accuracy of fitted model1, plot training, validation set loss and accuracy, and print the model’s structure

  • Evaluate the fitted model

    • What’s the test accuracy after training the model? Does it improve from the initialization?

  • Print the model’s structure

❓❓❓ Do you have similar test accuracy?

❓❓❓ Do you have a similar plot?

❓❓❓ Do you have a similar structure?

(This reference was generated with Keras’s plot_model utility, which has no PyTorch equivalent — print(model) gives you the same layer-by-layer information as text instead of a diagram.)

Q6 Create two callbacks to save the model weights at each epoch and the best validation accuracy epoch

  1. checkpoint_every_epoch: a callback that saves the model weights every epoch during training;

  2. checkpoint_best_only: a callback that saves only the weights with the highest validation accuracy.

Hint 1: Since train_model calls each callback’s on_epoch_end(epoch, model, logs) method after every epoch, a callback here is just a small class with that one method — logs is the same dict of "loss"/"accuracy"/"val_loss"/"val_accuracy" values printed for that epoch.

Hint 2: Use torch.save(model.state_dict(), path) to save weights — there’s no PyTorch equivalent of Keras’s ModelCheckpoint, so both callbacks save with plain torch.save().

Q7 Build a CNN model with the same initial structure as model1 and train for 15 epochs using the callbacks from Q6

Now, you will train the model using the two callbacks you created. If you created the callbacks correctly, two things should happen:

  • At the end of every epoch, the model weights are saved into a directory called checkpoints_every_epoch

  • At the end of every epoch, the model weights are saved into a directory called checkpoints_best_only only if those weights lead to the highest validation accuracy

You should then have two directories:

  • A directory called checkpoints_every_epoch containing filenames that include checkpoint_000.pt, checkpoint_001.pt, etc, with the numbers corresponding to the epoch

  • A directory called checkpoints_best_only containing a single checkpoint.pt file, containing only the weights leading to the highest validation accuracy

❓❓❓ Do you have similar print output?

Q8 Create new models model_last_epoch and model_best_epoch with model1’s initial structure; load weights from the latest saved epoch and the saved epoch with the highest validation accuracy respectively

Now you will use the weights you just saved in a fresh model. You should load into two freshly instantiated model instances:

  • model_last_epoch should contain the weights from the latest saved epoch

  • model_best_epoch should contain the weights from the saved epoch with the highest validation accuracy

Hint: glob.glob("checkpoints_every_epoch/*.pt") gives you every saved checkpoint path; sorting that list gives you the latest one at the end (since the filenames are zero-padded).

❓❓❓ Are the saved models’ validation accuracy as expected?

Q9 Explore to improve the model performance by trying to reduce the bias: design and train a model2

  • For example, train a CNN called model2, which has the same structure as model1, but with more convolution units: conv_1 has 32 units, conv_2 has 64 units, dense_1 with 256 units

❓❓❓

  • Does the model performance improve?

  • What are other potential strategies to improve model performance by reducing the bias?

❓❓❓ Are your printed output similar to the following screenshot?

❓❓❓ Does your model’s test accuracy improve and why?

Q10 Explore to improve the model performance by trying to reduce the variance: design and train a model3

  • For example, train a CNN called model3, which has the same structure as model2, add dropout layers one after the max pooling layer, one before the final dense layer with dropout rate of 0.2.

❓❓❓

  • Does the model performance improve?

  • What are other potential strategies to improve model performance by reducing the variance?

❓❓❓ Are your printed output similar to the following screenshot?

❓❓❓ Does your model’s test accuracy improve and why?

Q11 Explore to improve the model performance using transfer learning from a pretrained model

  • Use a pretrained VGG16 model from TorchVision.

    • Does the model performance improve after the above explorations?

    • How would you further improve the model performance?

Hint 1: torchvision.models.vgg16(weights=torchvision.models.VGG16_Weights.DEFAULT).features gives you just VGG16’s convolutional layers (no classifier head) — the PyTorch equivalent of Keras’s include_top=False.

Hint 2: Freeze every parameter in the pretrained feature extractor with param.requires_grad = False, so training only updates the new layers you add on top.

❓❓❓ Are your printed output similar to the following screenshot? It takes a really long time to train ...

✌✌✌ Congratulations! You have completed this exercise. Now you know how to train CNNs to classify remote sensing images.

Still, the accuracy reported from the Eurosat dataset paper - Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification is 98.57%. ⚡⚡⚡Have you found strategies to improve the performance to that level?