Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

1.8) Statistical Foundations and the Machine-Learning Stepping-Stone

Open In Colab Open In Kaggle

This closing subchapter connects descriptive statistics to a first supervised model. Using an environmental lapse rate — air temperature as a function of elevation — it moves from exploratory plots, histograms, and correlation, through the scikit-learn estimator interface (instantiate, fit, predict), to an honest evaluation on held-out data with RMSE and the coefficient of determination. It then turns to unsupervised learning on a real dataset — k-means clustering and PCA on Palmer Archipelago penguin measurements — before the generated-code bug. The estimator-as-object pattern is the one introduced in the object-oriented subchapter. At the end, the generated code judges a model by how well it fits the data it was trained on: a near-perfect score on the training set, and a badly wrong prediction on new points.

1.8.1 Exploratory Data Analysis with seaborn

seaborn draws statistical graphics directly from a dataframe. We generate a synthetic set of stations whose temperature falls with elevation at roughly the environmental lapse rate of −6.5 °C km⁻¹, plus noise, and add a loosely elevation-dependent count of snow days.

seaborn is not a replacement for matplotlib but a layer on top of it: pass column names instead of extracted arrays, and get a sensible default style, a legend, and — for regplot below — a fitted trend line, all in one call. ax= still accepts a matplotlib axes, so the two compose rather than compete; reach for matplotlib directly whenever a plot needs more control than seaborn’s defaults give you.

<Figure size 500x350 with 1 Axes>

1.8.2 The Palmer Penguins Dataset

The second half of this subchapter works on a real dataset: the Palmer Penguins, collected by Kristen Gorman at Palmer Station, Antarctica. It holds 344 records of the physical attributes of penguins in the Palmer Archipelago, belonging to three species — two of them have no body measurements at all and are dropped below, leaving 342.

Illustrations of the three penguin species in the dataset — Chinstrap, Gentoo and Adelie

Figure 2:The three species in the dataset. Artwork by Allison Horst (@allison_horst), from the palmerpenguins package. Data by K. Gorman and the Palmer Station Antarctica LTER, released under CC-0.

1.8.3 Histograms, KDE, and Density Plots

A histogram counts observations into bins; a kernel density estimate (KDE) smooths that count into a continuous curve. sns.histplot and sns.kdeplot take the same data=/x= arguments as the rest of seaborn, and hue= splits either one by a categorical column in a single call — no manual loop over groups needed. The examples below use real measurements of three penguin species from the Palmer Archipelago, Antarctica (Horst, Hill & Gorman, 2020; palmerpenguins, doi:10.5281/zenodo.3960218; data collected by the Palmer Station Antarctica LTER and K. Gorman).

<Figure size 600x350 with 1 Axes>

kdeplot also works on two variables at once, showing where two measurements co-vary rather than just their individual spreads; jointplot adds the two 1D marginal densities for free.

<seaborn.axisgrid.JointGrid at 0x7fd84f687440>
<Figure size 600x600 with 3 Axes>

Two of those four measurements describe the bill. The raw data calls them culmen length and depth, after the ridge along the top of the bill; the column names here use the plainer word.

A diagram of a penguin head showing which dimension is bill length and which is bill depth

Figure 3:What bill_length_mm and bill_depth_mm measure. Artwork by Allison Horst (@allison_horst), from the palmerpenguins package.

1.8.4 Descriptive Statistics and Correlation

describe summarises each column; corr gives the pairwise Pearson correlation. Temperature is strongly anti-correlated with elevation — the lapse rate — while snow days rise with it.

       elevation_m  temp_celsius  snow_days
count        80.00         80.00      80.00
mean       1894.42          2.67      56.92
std         978.78          6.68      21.61
min         209.04         -9.87       6.52
25%        1161.85         -2.89      40.06
50%        1960.25          2.51      57.56
75%        2655.25          8.02      72.30
max        3490.79         16.17      99.78
<Figure size 400x300 with 2 Axes>

1.8.5 The scikit-learn Estimator Interface

scikit-learn covers three pillars of statistical modelling, and this subchapter touches all three: regression, which learns the relationship between continuous inputs and a continuous output; classification, which learns which of a set of classes an observation belongs to; and clustering, which groups observations that resemble each other without being told what the groups are.

Every scikit-learn model is an object used the same way: instantiate it, fit it to a feature matrix X and target vector y, then predict. By convention X is two-dimensional with shape (n_samples, n_features) and y is one-dimensional. The fitted parameters are stored on attributes ending in an underscore.

For a straight line y^=wx+b\hat{y} = wx + b, fit chooses the slope ww and intercept bb that minimise the sum of squared errors

J(w,b)=∑i=1N(yi−y^i)2,J(w, b) = \sum_{i=1}^{N}\left(y_i - \hat{y}_i\right)^2,

solved directly rather than by trial and error. Least squares is what makes the fitted line the one that passes as close as possible to every point at once.

recovered lapse rate: -6.67 °C km^-1
intercept (sea-level temp): 15.31 °C
prediction at 2000 m: 1.96 °C

The recovered lapse rate is close to the −6.5 °C km⁻¹ built into the data, but not exact: 80 noisy stations only pin it down so far. More data narrows the gap.

   80 stations -> lapse rate  -6.35 °C km^-1, intercept 14.61 °C
  800 stations -> lapse rate  -6.54 °C km^-1, intercept 15.04 °C
 8000 stations -> lapse rate  -6.51 °C km^-1, intercept 15.02 °C

1.8.6 Train/Test Split, RMSE, and R^2

A model must be judged on data it has not seen. train_test_split holds out a test set; the root-mean-square error (RMSE) reports the typical error in the target’s units, and R^2 the fraction of variance explained.

test RMSE: 1.595 °C
test R^2:  0.947
<Figure size 500x350 with 1 Axes>

1.8.7 Scaling and Pipelines

Many estimators (regularised regressions, distance-based methods) require features on comparable scales; StandardScaler centres and scales them. A Pipeline chains preprocessing and model into one estimator, so the scaler is fitted on the training data only — which is exactly what prevents information from the test set leaking into training.

pipeline test R^2: 0.947

1.8.8 Linear Classification

Not every target is continuous. When y is a category — here, whether a station’s reading fell below freezing — a classifier follows the identical fit/predict interface. LogisticRegression is the classification analogue of linear regression: it fits a linear boundary in feature space and reports class probabilities rather than a continuous value.

test accuracy: 0.958
predicted frost at 3000 m: True

Separating two species by bill shape

That target was manufactured from a threshold, so the classifier had an easy job. A real one: tell an Adelie from a Chinstrap using nothing but the two bill measurements. Gentoo is left out — it is the easy case, separable on almost any pair of features — leaving the two species that overlap.

Look at the two measurements one at a time first. Bill length separates the two species almost completely; bill depth barely separates them at all.

<Figure size 1000x320 with 2 Axes>

Bill length runs from roughly 32 to 58 mm and bill depth from 15 to 21, so the two features differ in scale by about a factor of three. Left alone, the fitted weights would reflect that difference as much as the information in the features, which is exactly what StandardScaler inside a pipeline prevents.

test accuracy: 0.97

The fitted model is a straight line through the standardised bill plane: one weight per feature plus an intercept. Reading the weights says which measurement the decision actually rests on.

bill_length_mm    3.79
bill_depth_mm    -1.19
dtype: float64
intercept: -1.79
<Figure size 500x240 with 1 Axes>

1.8.9 Unsupervised Learning: Clustering and Dimensionality Reduction

Every model so far had a known target. Two more scikit-learn estimators follow the same fit/predict shape with no target at all: KMeans groups similar observations, and PCA finds the directions of greatest variance in correlated features. Both are demonstrated on the penguin measurements loaded above.

k-means clustering

Four panels showing successive k-means iterations, the cluster centres moving and the point colours settling as the algorithm converges

Figure 5:Four successive iterations of k-means on the same points. Each panel assigns every point to its nearest centre and then moves each centre to the mean of the points assigned to it; by the fourth panel neither step changes anything and the algorithm has converged. Figure from Jake VanderPlas, Python Data Science Handbook (code MIT-licensed).

The algorithm alternates two steps until neither changes anything, and it never sees a label.

KMeans partitions points into n_clusters groups by repeatedly assigning each point to the nearest centroid and moving each centroid to the mean of its group. It is easiest to see on data built to have clusters: make_blobs draws points around a chosen number of centres, and the fitted centroids land on them.

<Figure size 900x350 with 2 Axes>

On a real dataset the clusters are rarely that obliging. The penguin measurements below have no species column as far as KMeans is concerned — it sees two physical measurements and nothing else.

<Figure size 900x350 with 2 Axes>
cluster     0    1   2
species               
Adelie      0  116  35
Chinstrap   0   53  15
Gentoo     68    1  54

Two features are not quite enough to cleanly separate three species by shape alone: Gentoo (heavier, longer-flippered) forms its own cluster, but Adelie and Chinstrap overlap and mix into the other two clusters — the crosstab shows exactly where. Clustering finds structure in the features you give it; it cannot recover a label the features do not distinguish.

That is a statement about the features, not about the algorithm. Gentoo separates because it is heavier and longer-flippered than the other two; Adelie and Chinstrap overlap because in this plane they genuinely are alike. Swap body mass for bill length, and the same three species come apart — Chinstrap flippers resemble Adelie’s, but Chinstrap bills resemble Gentoo’s.

<Figure size 1000x380 with 2 Axes>

Choosing k: the elbow method

We fixed k=3 because we already know there are three species — in general the right number of clusters is not known in advance. The elbow method refits KMeans across a range of k and plots the inertia (each point’s squared distance to its cluster’s centroid, summed over every point): inertia always falls as k grows, but the rate of improvement drops sharply once extra clusters stop capturing real structure — a bend, or “elbow”, in the curve.

<Figure size 500x350 with 1 Axes>

Silhouette analysis

The elbow is not always sharp: here the curve bends somewhere between two and four clusters without saying which. The silhouette score puts a number on it. For one observation,

S=b−amax⁡(a,b),S = \frac{b - a}{\max(a, b)},

where aa is its mean distance to the other points in its own cluster and bb its mean distance to the points of the nearest other cluster. A value near 1 means the observation sits comfortably inside its own cluster; near 0 means it sits on the boundary between two. Averaged over every observation, the score can be compared across k, and the largest average wins.

<Figure size 500x350 with 1 Axes>
k=2: 0.630
k=3: 0.580
k=4: 0.555
k=5: 0.540
k=6: 0.523

Two clusters score highest — and two clusters is not three species. Both diagnostics are answering the question they were asked, which is how well-separated the groups in these two features are, not how many species the archipelago holds. A clustering that scores well can still be the wrong grouping for the question you care about.

Dimensionality reduction with PCA

PCA finds new axes — linear combinations of the original features — ordered by how much variance they capture, so a handful of components can summarise many correlated measurements. Scale first, exactly as for the pipeline above: PCA is sensitive to each feature’s variance, not just its values.

explained variance ratio: [0.688 0.193]
<Figure size 550x400 with 1 Axes>

Unlike k-means, PCA needs no target and no assumed number of groups — here the first two components already separate the three species almost as cleanly as the full four-feature space, using all four correlated measurements at once instead of picking two by hand.

When generated code lies: scoring on the training data

Asked to “evaluate the model”, an assistant fits a flexible model and reports R^2 on the same data it trained on. On a small, noisy sample a high-degree polynomial fits almost perfectly — and predicts new points disastrously.

degree-8 R^2 on TRAINING data: 0.971
degree-8 R^2 on TEST data: -7.1
linear    R^2 on TEST data: 0.984
<seaborn.axisgrid.FacetGrid at 0x7fd842d273e0>
<Figure size 900x300 with 3 Axes>
<Figure size 500x350 with 1 Axes>
<Figure size 500x350 with 1 Axes>

Summary

ConceptRule to remember
Explore firstseaborn plots, describe, and corr reveal structure before any model is fitted.
Distributionshistplot/kdeplot show one variable (split it with hue=); jointplot adds the marginals to a 2D view.
Estimator interfaceEvery scikit-learn estimator is instantiate → fit(X, y) → predict.
ShapesX is two-dimensional, y is one-dimensional.
EvaluatingFit on the training set, score on a held-out test set; report RMSE in target units and R^2.
OverfittingA high training score with a low test score; a training score alone is never performance.
LeakagePut StandardScaler in a Pipeline so preprocessing is fitted on training data only.
ClassificationLogisticRegression brings the same fit/predict interface to a categorical target.
UnsupervisedKMeans groups by proximity, PCA summarises by variance; the elbow and silhouette scores help choose k.

Resources