latetrains.pkLate, on record44 live
Guides/Where our train delay data comes from

Where our train delay data comes from

Late Trainsabout 6 min readmethodologyaboutdata

Published 5 September 2026. Not revised since.

If a website tells you a train is typically forty minutes late, the reasonable next question is how it knows. Most of the time there is no answer, or the answer is that a number was copied from somewhere that copied it from somewhere else.

This page is the answer for this site. It is longer than a marketing sentence because the details are where the trust either is or is not.

We keep what we were told, before we interpret it

The first rule, and the one everything else rests on: every payload we fetch is written to an archive before anything is parsed. The archive is append only. Nothing is edited and nothing is deleted.

That sounds like housekeeping. It is the foundation of the whole thing, for one reason: it means every figure we publish can be recomputed. If we later find our interpretation was wrong, we can fix the code and rebuild the history from what was actually received at the time, rather than patching a number and hoping.

It also means our published figures are falsifiable in principle. A number that cannot be rederived from a record of what was observed is not a measurement, it is an assertion.

Two ways an arrival gets recorded

An arrival reaches our database by one of two paths, and every stored arrival carries a label saying which. That label matters, so the site distinguishes them rather than blending them.

The first path is the operator's own report. When the source reports that a train has passed a station, along with how late it was, we store that. It is the operator's measurement, not ours, and it is the better one when it exists.

The second path is our own observation. We collect position reports continuously while a train is running, which gives a dense trail of where it was and when. From that trail we can determine when the train reached each station, including stations the operator did not report on.

The second path fills the gaps in the first. It never overwrites it. When both exist for the same stop, the operator's own figure stands.

When we refuse to record a time

This is the part most worth reading, because a data source is defined at least as much by what it declines to record as by what it collects.

A GPS trail is not continuous in practice. Transponders go quiet. A train can vanish for hours and reappear well down the line, and when it reappears the naive reading is that it arrived at every station in between at the moment it came back, which would produce arrival times that imply impossible speeds.

So we hold a rule: a station arrival gets a clock only when we were genuinely watching. If the previous position report is too old, or the reappearing report is too far from the station to witness anything about it, we record that the train passed and how late the operator said it was, and we leave the arrival time blank.

A blank is not a gap in the data in the sense of something missing that should be there. It is the honest output. We knew the train got there. We did not see when. Recording a confident time would have manufactured a fact.

We hold a similar floor on how early a train can be recorded as arriving. A recorded arrival an hour before the booked time is not an unusually punctual train, it is a mismatch between our copy of the timetable and the run we matched it against. That is our error, and letting it through would put our error into a figure about them, where it would read as a service running early.

Why some pages are empty

You will find pages here with no figures on them, and pages that say a service has too little recorded history to summarise.

That is deliberate and it is not a bug to be fixed by lowering the bar. A punctuality figure drawn from two runs is not a summary of anything. We hold a minimum amount of recorded history before a train gets a published figure, and until then the page says so.

The same applies when a data source is down. If we cannot confirm something, the page shows nothing rather than the last thing we knew, presented as though it were current. Every live figure on the site is stamped with the moment it was true and ages itself as you look at it, so a stale number cannot quietly pass for a fresh one.

Where the timetable comes from, and why that matters

Delay is a difference between an actual time and a booked time, so it is only as good as our copy of the booked time.

Timetables are versioned here. When a service is retimed, the new version is stored alongside the old one rather than replacing it, and a run is always measured against the version that was in force on the day it ran. Measuring a journey from last month against a timetable introduced last week would produce a delay figure that describes nothing.

Source timetables also contain errors, and a wrong booked time at one station shifts every delay we calculate there by the same amount permanently. We check for shapes that cannot be right, such as a booked gap between two stops that no train could plausibly take, and we refuse them rather than ingesting them. Where a bad cell has already been ingested and repaired, the delay figures derived from it are withdrawn rather than left standing on a foundation that has been removed.

What we do not do

We do not estimate a position we did not receive. We do not fill a missing arrival with an interpolation between the two nearest known ones. We do not publish a distance we did not measure. We do not carry a figure forward from an earlier day and present it as today's.

Each of those would make the site look more complete. Each would put a number in front of you that nothing behind it supports, and you would have no way of telling it apart from the rest.

The site is more blank than it could be. That is the trade, and it is deliberate.

See it in the data

  • The archive — what we collect, and the dataset itself