Learn ML with transit dataLondon Underground, Central line

Interactive machine learning course · real TfL data

Machine learning, explained with the Central line

How much of the Central line's timetable will actually run? Using 109 periods of real TfL data, we build models of growing power, from a plain average to a neural network, and see at every step how each one works and why it eventually stops being enough.

Every model on this page is trained live in your browser. Chapters 1 to 3 use the real published service operated (the share of the scheduled service that actually ran) of the Central line; chapters 4 to 9 use small simulated datasets so that each model's behaviour is easy to see.

88.5%average service operated, 2018-19 P1 to 2026-27 P6
27.3%worst period: 2020-21 P1
97.2%best period: 2018-19 P10
84.0%latest: 2026-27 P6

The Central line crosses London from west to east. Transport for London reports Tube performance in four-week periods, 13 per financial year; the measure used here is service operated, the share of the scheduled service that actually ran. Source: Underground services performance, service operated (Transport for London). Powered by TfL Open Data.

Course recap · vocabulary

  • Observation: one row of the table. Here, one period of the Central line.
  • Features, written x: what we know before we predict.
  • Target, written y: what we want to predict. A number gives a regression; a category (on time or late) gives a classification.
  • Training: tuning the model's parameters on examples whose answer we know.
  • Test: measuring the error on examples the model never saw. It is the only score that counts.

Chapter 1 · Regression

The average, the model to beat

RecapA model turns features x into a prediction ŷ, and we judge it on test data.

The chart shows 109 periods of service operated (the share of the scheduled service that actually ran) for the Central line, from 2018-19 P1 to 2026-27 P6. The simplest possible model ignores all context and always predicts the same value c. Which c should we pick? Move the slider. Each vertical stroke is an error, called a residual.

The mean squared error (MSE) averages the squares of those residuals. The small chart below shows the MSE for every value of c: a parabola whose lowest point sits exactly on the average, 88.5%. An optimist who always predicts 100% does much worse.

The worst period in the series, 2020-21 P1 at 27.3%, drags the mean down more than the median, because its huge residual gets squared. That period coincides with the first COVID-19 lockdown in spring 2020, when the Tube ran a reduced service. If we measure the absolute error (MAE) instead, the median (90.0%) wins. The choice of error function changes the "best" answer.

This naive model is called the baseline. Any more complex model has to beat it, or it is useless.

Course recap

  • Residual: e = y − ŷ
  • MSE = (1/n) Σ (yi − ŷi)² and RMSE = √MSE, here in percentage points.
  • MAE = (1/n) Σ |yi − ŷi|
  • The mean minimises the MSE; the median minimises the MAE.
  • Squaring punishes big misses hard; the MAE is more robust to extreme values.

Quick quiz

Why does one terrible period move the MSE so much?

Because the error is squared: missing by 40 points costs 1,600, as much as 100 periods missed by 4 points.

Your model gets an RMSE of 3.9 points and the baseline 4.0. Good news?

Not really: the gain is tiny and could come from the luck of the split. Compare across several splits (cross-validation) before concluding.

Why it is not enough

The average predicts 88.5% for every period, even right after a terrible one. We need a model that uses information known in advance. Chapter 2: linear regression.

Lab · service operated, the Central line

Chapter 2 · Regression

The straight line and gradient descent

RecapThe best constant, in the MSE sense, is the mean. It is the baseline to beat.

We add one piece of information known in advance: last period's figure. Each dot pairs one period (vertical axis) with the period before it (horizontal axis). The model becomes a line ŷ = a·x + b. On this data, the best line has a slope of about 0.70: a bad period tends to be followed by another below-average one, but less extreme.

Tune a and b by hand to bring the MSE down. The map on the right shows the MSE for every pair (a, b): the cost function, a valley whose lowest point is marked with a star. Click anywhere on the map to move the model there.

Gradient descent does the job on its own. At each step it measures the slope of the valley at the current point and steps downhill. The step size is the learning rate η. Run it as is: the values are around 80, so the valley is a long, thin canyon and the descent crawls along it. Push η one notch higher and it diverges.

Now tick "Standardise x". Rescaling x to mean 0 and standard deviation 1 turns the canyon into a round bowl, and the descent reaches the bottom in a few steps, even with a large η. This is why we almost always normalise features.

Course recap

  • Model: ŷ = a·x + b
  • Cost: J(a, b) = (1/n) Σ (a·xi + b − yi)²
  • Gradient, with ei = ŷi − yi: ∂J/∂a = (2/n) Σ ei·xi and ∂J/∂b = (2/n) Σ ei
  • Update: a ← a − η·∂J/∂a, same for b.
  • Standardise: x' = (x − mean) / std
  • Linear regression has an exact formula (least squares). Gradient descent matters where none exists: logistic regression, neural networks.

Quick quiz

The best model finds a ≈ 0.70. How do you read it?

Each extra point of service operated last period adds about 0.70 point this period. Good and bad spells persist, but they fade.

The MSE grows at every step and finally explodes. What do you do?

Lower the learning rate η, and standardise the features so that the valley is better shaped.

Why it is not enough

Last period only tells part of the story. Over the years, service operated drifts up and down in waves that no straight line follows. Chapter 3: polynomials.

Lab · this period vs the previous one

Chapter 3 · Regression

Polynomials, bias and overfitting

RecapGradient descent follows the slope of the cost; too large a step makes it diverge, and standardising features helps a lot.

Now x is time and y is the service operated of each period. To follow curves, we give the model powers of x: ŷ = w0 + w1x + w2x² + … + wd·xd. It is still a linear model… in its parameters w. The degree d sets how flexible it is.

The model only sees a random sample of periods (solid dots); the others (hollow) are the test set. Raise the degree. At first, training and test errors fall together. Then the curve starts threading through every training dot: the training error keeps falling while the test error climbs back up. The model is learning the noise by heart. That is overfitting.

Click "New sample" at a high degree: the curve changes completely from one sample to the next. That is variance. At degree 1 it barely moves but misses every wave: that is bias. Drag "Training periods" up: with more data, the same high degree behaves much better.

Finally, tick "Forecast the last two years". The model trains on the past only and must extend its curve into the future. Even a curve that looks fine on the past shoots off the chart: polynomials are terrible forecasters.

Course recap

  • Underfitting (high bias): the model is too rigid, the error is high everywhere.
  • Overfitting (high variance): low training error, high test error.
  • test error ≈ bias² + variance + noise, and the noise cannot be reduced.
  • Ridge: minimise MSE + λ Σ wj². The larger λ, the more constrained the model.
  • More data reduces variance.
  • Pick d and λ on a validation set and keep the test set for the final score. Here we peek at the test set for the demonstration.

Quick quiz

Training error 1.2, test error 95. Diagnosis?

Overfitting. Lower the degree, raise λ or add data.

Why not choose the degree directly on the test set?

The test set would become disguised training data, and the reported error would be too optimistic.

Why it is not enough

So far we predicted a percentage. A rider mostly wants a yes or no answer: will my train be late? That is a classification. Chapter 4: logistic regression.

Lab · service operated over time

Chapter 4 · Classification

Logistic regression

RecapToo rigid means bias; too flexible means variance. Regularisation looks for the compromise.

From here on the data is simulated, so that each model's behaviour is easy to see. Two features: a train's delay at departure, and the share of late trains on the line. Each dot is a trip, blue if it arrived on time, orange if it arrived 5 or more minutes late.

Logistic regression computes a linear score z = w0 + w1x + w2y and turns it into a probability with the sigmoid. The boundary where p = 0.5 is a straight line. Click "Watch it learn" to see gradient descent move the line into place.

The threshold decides from which probability we announce "late". Lower it and you miss fewer delays (recall goes up) but cry wolf more often (precision goes down). Watch the confusion matrix.

Now switch to "Circles" or "XOR": no straight line separates the classes and accuracy stalls. Tick "Add x², y², x·y": with these new features the boundary can bend. Crafting the right features by hand is called feature engineering.

Course recap

  • Sigmoid: σ(z) = 1 / (1 + e−z), always between 0 and 1.
  • Cost (log-loss): J = −(1/n) Σ [yi·log pi + (1 − yi)·log(1 − pi)]
  • Precision = TP / (TP + FP): when I announce a delay, am I right?
  • Recall = TP / (TP + FN): of all real delays, how many do I catch?
  • Accuracy = (TP + TN) / n, misleading when one class is rare.

Quick quiz

Only 10% of trains are late. What accuracy does a model that always says "on time" get?

90%, with a recall of 0. That is why we look at precision and recall.

Why does logistic regression fail on XOR?

Its boundary is a straight line in the space of input features; XOR needs at least two lines.

Why it is not enough

Adding x², y², x·y works for circles, but you have to guess the right shape in advance. On the spiral, which features would you invent? We want models that find the shape on their own. Chapter 5: k-nearest neighbours.

Lab · will it arrive on time?

A click on the chart adds a point

Chapter 5 · Classification

k-nearest neighbours

RecapLogistic regression draws a straight boundary; to bend it, you have to invent features by hand.

This model looks for no formula at all. For a new trip, it finds the k closest trips in the training data and lets them vote. Hover over the chart: the lines connect the hovered point to the neighbours consulted.

With k = 1, every training point paints its own colour around itself: 100% training accuracy, islands everywhere, and a test score that suffers. As k grows, the boundary smooths out. k plays the same role as the polynomial degree: it sets the flexibility, only in reverse.

Course recap

  • Euclidean distance: d = √((x1 − x2)² + (y1 − y2)²)
  • Prediction: p = share of the k neighbours in class 1
  • Small k: high variance. Large k: high bias.
  • Always normalise features, or a feature in minutes will crush a feature in percent.
  • No training at all, but every prediction scans every point. In high dimension, all points become "far": the curse of dimensionality.

Quick quiz

Why is training accuracy 100% with k = 1?

Each training point is its own nearest neighbour.

What happens if k equals the number of training points?

The model always predicts the majority class: we are back to the baseline.

Why it is not enough

k-NN explains nothing: there is no readable rule. It also gets slow and unreliable as features pile up. We want a flexible model that produces rules. Chapter 6: the decision tree.

Lab · the neighbours vote

A click on the chart adds a point

Chapter 6 · Classification

The decision tree

Recapk-NN lets the k closest neighbours vote; k sets the flexibility, and features must be normalised.

A tree asks one question after another, such as "delay at departure ≤ 6 min?". At each node it picks the question that best separates the classes, the one that makes the two resulting groups as pure as possible. Its boundaries are therefore made of horizontal and vertical segments.

The maximum depth sets the flexibility. At depth 1, a single question. At depth 10, the tree carves little boxes around every isolated point. The learned rules are printed under the chart.

Click "Resample" several times: the tree is rebuilt on a draw with replacement of the same data. The rules change a lot for a small change in the data. A deep tree is unstable.

Course recap

  • Gini impurity of a group: G = 2p(1 − p), zero when the group is pure.
  • Pick the split that minimises the impurity of the two children, weighted by their size.
  • Stop at a maximum depth or a minimum number of points per leaf.
  • Strengths: readable, no normalisation needed.
  • Weakness: high variance.

Quick quiz

Do you need to standardise features before a tree?

No: a split "x ≤ threshold" does not depend on the scale of x.

Why does the boundary look like a staircase?

Each question looks at a single feature at a time.

Why it is not enough

A single tree is readable but unstable. The idea: grow many of them and let them vote. Chapter 7: forests and boosting.

Lab · one question at a time

A click on the chart adds a point

    

Chapter 7 · Classification

Random forests and boosting

RecapA deep tree follows the data closely but changes a lot from one sample to another.

A random forest trains dozens of trees, each on a different resample and with some randomness in the choice of features. Their mistakes differ, so averaging them cancels much of the error. Add trees: the boundary smooths out and the test error falls.

Boosting works the other way round. The trees are small and built one after another, each one correcting the errors left by the previous ones. It mostly reduces bias. This is the family of XGBoost, LightGBM and CatBoost, often the best choice on tabular data.

Compare: a forest of 1 tree is the unstable tree of chapter 6; with 100 trees it becomes stable. Boosting with many deep trees ends up learning the noise.

Course recap

  • Bagging (forest): average of deep, diverse trees. Reduces variance.
  • Boosting: sum of weak trees, each fitted to the remaining errors. Reduces bias.
  • Fm(x) = Fm−1(x) + η·hm(x), with hm fitted to yi − pi.
  • Adding trees to a forest does not cause overfitting; adding boosting rounds does. Hence early stopping.
  • The price: you lose the readability of a single tree.

Quick quiz

Why does an average of trees make fewer mistakes?

Their errors are partly independent: the average of noisy predictions is less noisy than each of them.

Boosting with 100 trees of depth 8: what is the risk?

Overfitting: boosting keeps chasing the noise. Prefer shallow trees and a modest η.

Why it is not enough

Tree ensembles shine on tables. For images, sound, text or very twisted shapes, we need a model that builds its own features. Chapter 8: the neural network.

Lab · many trees

A click on the chart adds a point

Chapter 8 · Classification

The neural network

RecapBagging reduces variance, boosting reduces bias; both combine many trees.

A neuron is a logistic regression: a weighted sum followed by an activation function. A network stacks layers of neurons. The diagram shows what each hidden neuron "sees": those of the first layer each draw a straight boundary, those of the next layer combine them into curved shapes.

The network therefore learns its own features, where chapter 4 made you invent them by hand. On the spiral, try 2 neurons, then 8: with too few neurons the shape is impossible to draw.

Training reuses the gradient descent of chapter 2. Computing the gradient layer by layer, from the output back to the input, is called backpropagation. In the diagram, orange links are positive weights and blue ones negative; their thickness follows their strength.

Course recap

  • Neuron: a = f(w · x + b), with f = tanh or ReLU.
  • Output: p = σ(z), with the same log-loss as chapter 4.
  • Backpropagation: the chain rule, applied layer by layer.
  • Adam: gradient descent that adapts the step of each parameter.
  • Many parameters: it needs lots of data and careful tuning, and it is hard to interpret.

Quick quiz

A 10-layer network with no activation, f(z) = z: what is it worth?

No more than a linear model: a composition of linear functions is linear.

Training loss keeps falling but test accuracy drops. What is happening?

Overfitting: shrink the network, regularise it, or stop training earlier.

Why it is not enough

More powerful does not mean better everywhere. Let us compare every model on the same data. Chapter 9: the showdown.

Lab · a small network

Chapter 9 · Wrap-up

The showdown

RecapA neural network learns its own features, at the cost of many parameters.

Eight models, trained on the same points and scored on the same test points. Change the dataset. On "Trains (realistic)", almost everyone ties and logistic regression, simple and readable, is enough. On the spiral, only flexible models cope. The green frame marks the best test score.

ModelBoundaryReadabilityData neededWatch out for
Mean, majority classnonetotalvery littleonly useful as a yardstick
Linear or logistic regressionstraight linereadable coefficientslittlebias when the relation is curved
Polynomial, crafted featureshand-picked curvemediumlittle to mediumoverfitting at high degree, wild forecasts
k-nearest neighboursfreelowmediumnormalisation, speed, high dimension
Decision treestaircasereadable rulesmediuminstability
Forest, boostingsmoothed staircaselow (feature importance)medium to largetuning, compute time
Neural networkfreevery lowlargetuning, overfitting, cost

Back to the Central line

On the Central line, the plain average misses by 9.2 points in a typical period (RMSE). Simply using last period's figure, as in chapter 2, brings that down to 6.5 points. A serious forecast would add what TfL also publishes, such as lost customer hours and status messages, and would likely end with a gradient-boosted model. The rule of this course holds: start simple, measure, and only add complexity when the test error really drops.

The same course runs on other networks, each with its own real data:

The golden rule

  • Start with the baseline.
  • Add complexity one step at a time.
  • Keep a more complex model only if the validation error really drops.
  • Save the test set for the very end.

Terminus

Cheat sheet

RecapEverything worth remembering from the journey, on one page.

Evaluate

  • MSE = (1/n) Σ (yi − ŷi)²
  • MAE = (1/n) Σ |yi − ŷi|
  • Precision, recall, accuracy, confusion matrix.
  • Training to learn, validation to tune, test to score.

Optimise

  • A cost J to minimise.
  • θ ← θ − η·∇J(θ)
  • η too large: divergence. Too small: crawling.
  • Standardised features speed up the descent.

Control complexity

  • Bias: too rigid. Variance: too flexible.
  • Levers: degree, k, depth, number of neurons.
  • Regularisation: + λ Σ wj²
  • More data reduces variance.

Models in one line

  • Logistic: p = σ(w · x + b)
  • k-NN: the k neighbours vote.
  • Tree: questions that minimise Gini.
  • Forest: average of trees. Boosting: corrective sum.
  • Network: layers of neurons, backpropagation.