This chapter summarizes Chapters 9 and 10 of “Hands-On Machine Learning with Scikit-Learn and PyTorch” by Aurélien Géron, which introduce the fundamentals of artificial neural networks (ANNs) and how to build them with PyTorch.
5.1.1 Understanding Neural Networks:¶
Artificial Neural Networks (ANNs) are computational models inspired by the human brain. They consist of interconnected nodes (neurons) organized in layers. Information flows through these nodes, undergoing weighted computations and transformations, allowing ANNs to learn patterns and relationships from data, used widely in machine learning tasks.
Neurons: Neurons are the basic building blocks of ANNs. They receive inputs, perform computations, and produce outputs.
Layers: ANNs consist of an input layer, one or more hidden layers, and an output layer.
Activation Functions: Activation functions introduce non-linearity to neurons, enabling the network to model complex relationships.
Feed-forward Neural Networks (FNNs): FNNs pass information in one direction, from input to output. They are used for tasks like classification and regression.
Example: Building an FNN to classify species based on environmental data.
Backpropagation: Backpropagation is the training algorithm for ANNs. It adjusts the network’s weights and biases to minimize the error.
Example: Training an ANN to predict air quality based on historical data.
Hyperparameter Tuning: ANNs have numerous hyperparameters that influence their performance, including the number of layers, neurons per layer, and learning rate.
Example: Experimenting with different network architectures to optimize accuracy in predicting forest fire occurrences.
Overfitting and Regularization: ANNs are susceptible to overfitting, where they memorize training data but perform poorly on new data.
Techniques like dropout and early stopping mitigate overfitting. Example: Implementing dropout layers to improve the generalization of an ANN for predicting species distribution.
Environmental Applications 🌄: ANNs are used in environmental sciences for various tasks, such as weather forecasting, climate modeling, and remote sensing image analysis.
Example: Using ANNs to predict rainfall patterns based on climate data for flood risk assessment.
5.1.2 Perceptron¶
A perceptron is one of the simplest ANN architectures, used for binary classification tasks. In its simplest form, composed of one input layer and one output layer, a perceptron takes multiple binary inputs, applies weights to these inputs, sums them up, adds a bias term, and then passes the result through an activation function (e.g. step function) to produce an output.
Here’s a breakdown of the components and the function of a perceptron:
Inputs (X1, X2, X3, ...): These are the features or inputs to the perceptron. Each input is assigned a weight (W1, W2, W3, ...) which determines its importance in the computation. Weights (W1, W2, W3, ...): Weights are associated with each input and represent the strength of the connection between the input and the perceptron. Summation (Σ): The weighted sum of the inputs is calculated as follows:
Sum = (X1 * W1) + (X2 * W2) + (X3 * W3) + ...
Bias (B): A bias term is added to the weighted sum. The bias allows the perceptron to shift its decision boundary. Sum with Bias = Sum + B
Activation Function (e.g., Step function): The result of the summation is passed through an activation function. The most common activation function for a perceptron is the step function. If the result is above a certain threshold, the perceptron outputs a “1” (or “True”); otherwise, it outputs a “0” (or “False”). Output = 1 if (Sum with Bias > Threshold), else 0
Here is an environmental science 🌄 example where a perceptron can be used:
Perceptron for Species Classification: Imagine you’re working on a project to identify whether a particular bird species is present in a given area based on environmental features like temperature, humidity, and vegetation density. You have data on the presence (1) or absence (0) of the bird species in different locations.
In this scenario, you can use a perceptron to build a binary classifier. The inputs would be the environmental features (e.g., temperature, humidity, vegetation density), each with its corresponding weight representing its importance in determining bird presence. The perceptron computes the weighted sum of these inputs, adds a bias term, and passes the result through a step function. If the output is 1, it predicts the presence of the bird species; otherwise, it predicts absence.
A single perceptron may have limitations in capturing complex relationships in environmental data, which calls for more complex neural network architectures, such as multi-layer perceptrons (MLPs).
Multilayer Perceptron (MLP)
The stack of multiple perceptrons, composed of one input layer, one or more hidden layers, and one final output layer.
5.1.3 Create a Neural Network¶
Building a neural network in practice involves several steps, from data preparation to model evaluation. Let’s go through these steps using an example related to environmental science 🌄: predicting air quality based on meteorological data.
Step 1: Data Collection and Preprocessing
Data Collection: Gather historical data containing meteorological variables (e.g., temperature, humidity, wind speed) and corresponding air quality measurements (e.g., PM2.5 levels).
Data Preprocessing: Clean the data by handling missing values, outliers, and scaling features to a similar range (e.g., using Min-Max scaling). Split the data into training and testing sets.
Step 2: Model Selection and Architecture
Choose a Neural Network Architecture: Based on the problem, we select an FNN. Determine the number of input features (based on meteorological data) and the number of output neurons (1 for air quality prediction). Define the Model: Using a deep learning framework like
PyTorch, define the neural network’s architecture, including the number of hidden layers and neurons per layer, as well as the activation functions.
For simplicity, let’s create a model with one hidden layer.
import torch
import torch.nn as nn# Synthetic data standing in for meteorological features (temperature, humidity,
# wind speed) predicting an air quality measurement (e.g. PM2.5)
torch.manual_seed(42)
num_features = 3
n_samples = 200
X = torch.randn(n_samples, num_features)
y = (2.0 * X[:, 0] - 1.5 * X[:, 1] + 0.5 * X[:, 2]).unsqueeze(1) + torch.randn(n_samples, 1) * 0.1
train_size = int(0.8 * n_samples)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]model = nn.Sequential(
nn.Linear(num_features, 32),
nn.ReLU(),
nn.Linear(32, 1) # Output layer for air quality prediction
)Step 3: Loss Function and Optimizer
Choose a Loss and an Optimizer: PyTorch has no single
compilestep — instead, create the loss function (e.g. mean squared error for regression tasks) and the optimizer (e.g. Adam or SGD) as separate objects.
loss_fn = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters())Step 4: Training
Train the Model: PyTorch has no built-in
.fit()— write the training loop explicitly. For each epoch, compute predictions on the training data, calculate the loss, backpropagate the gradients, and update the weights. Mini-batching (splitting the data into batches rather than using it all at once) is typically handled with aDataLoader, omitted here for simplicity.
for epoch in range(50):
optimizer.zero_grad()
predictions = model(X_train)
loss = loss_fn(predictions, y_train)
loss.backward()
optimizer.step()Step 5: Evaluation
Evaluate the Model: Use the testing data to assess the model’s performance. Evaluate metrics like Mean Absolute Error (MAE) or Root Mean Squared Error (RMSE) to measure prediction accuracy.
model.eval()
with torch.no_grad():
predictions = model(X_test)
test_loss = loss_fn(predictions, y_test)
test_mae = torch.mean(torch.abs(predictions - y_test))
print(f"Test MAE: {test_mae:.2f}")Test MAE: 1.71
Step 6: Hyperparameter Tuning
Hyperparameters defining the network architecture:
Number of Layers: Defines the depth of the network.
Number of Neurons per Layer: Determines the width of each layer.
Activation Functions: Specifies the functions used to introduce non-linearity.
Architecture-Specific Parameters: Parameters unique to specific network types. Hyperparameters controlling the training process:
Learning Rate: Governs the step size in updating weights during training.
Batch Size: Specifies the number of samples processed before updating model parameters.
Optimizer: Algorithms adjusting weights during training (e.g., SGD, Adam, RMSprop).
Regularization Techniques: Methods to prevent overfitting (e.g., dropout rates, L1/L2 regularization).
Initialization Schemes: Initial values for weights and biases.
Early Stopping: Criterion to stop training based on validation performance.
Training Epochs: Number of iterations through the entire dataset during training.
Step 7: Deployment
Deploy the Model: Once satisfied with the model’s performance, deploy it in a production environment. It can be used to make real-time air quality predictions based on incoming meteorological data.
In this example, we’ve built a simple neural network for air quality prediction. In practice, you can enhance the model by incorporating more complex architectures (e.g., convolutional neural networks for image-based environmental data) and larger datasets. Neural networks in environmental science can be applied to various tasks, such as predicting pollutants, forecasting weather patterns, or modeling climate phenomena.