Free interactive course · real open data
Learn machine learning with real transit data
Nine short chapters take you from a plain average to a neural network. Every model is trained live in your browser, on real reliability figures published by the MTA, Transport for London and SNCF. Each chapter shows how a model works, then why you need the next one.
Pick a network. The course is the same; the data, the numbers and the story change.
New York
The NYC subway, A train
Will the A train be on time?
on-time performance 69.4% on averagelatest 77.9% (2026-08)140 months, 2015-01 to 2026-08Start the course →LondonThe London Underground, Central line
How much of the Central line's timetable will actually run?
service operated 88.5% on averagelatest 84.0% (2026-27 P6)109 periods, 2018-19 P1 to 2026-27 P6Start the course →ParisThe Paris RER B
Will the RER B get me there on time?
punctuality 85.2% on averagelatest 90.1% (2026-08)159 months, 2013-01 to 2026-08Start the course →What you will learn
- The baseline: why the average is the model to beat, MSE versus MAE.
- Linear regression and gradient descent, with a live cost map and learning rate.
- Polynomials, bias and variance, overfitting, regularisation and why forecasting is hard.
- Logistic regression, the sigmoid, the threshold, precision and recall.
- k-nearest neighbours and the curse of dimensionality.
- Decision trees, Gini impurity and instability.
- Random forests and gradient boosting.
- Neural networks, activations and backpropagation.
- A showdown of eight models on the same data.
How it works
- No install, no account: everything runs in the page.
- Chapters 1 to 3 use each network's real published data.
- Chapters 4 to 9 use small simulated datasets, so each model's behaviour is easy to see.
- Each chapter ends with a recap card and a quick quiz.