How often is it
actually right?
A forecast that never publishes its error rate is not a forecast, it is decoration. Here is the measured record: the hits, the average miss, and the failures.
Measured on real returns
How the test is run
The numbers above are not a self-assessment on training data. They come from a walk-forward test: for each real return, the model is rebuilt using only the information that existed before that car appeared, and its published window is then compared against what actually happened. Nothing from the future leaks in, including the population-wide statistics the model leans on for thinly-tracked cars.
Only returns that happened after the site went public count. Anything earlier is training history, and grading ourselves on it would be circular.
Reading the two numbers honestly
The two hit rates measure different promises. The calibrated range is the one designed to be reliable: it contains the return in the large majority of cases, and it is the number to trust if you want to know whether a car is realistically due. The likely period is deliberately much narrower, and its hit rate is close to a coin flip.
That is not a defect being hidden. It is the honest cost of a tight band. A narrow window that is right half the time is useful precisely because we tell you it is right half the time. If we widened it until it was almost always correct, it would stop being actionable.
The average error is the third number worth keeping in mind: roughly two weeks between the central estimate and the real return. That single figure is the reason this site stopped publishing precise dates.
A worked audit of the last five returns
Aggregate percentages can hide a lot, so here is every one of the five most recent returns at the time of writing, each scored against the window that was live before the car came back.
| Car | Central estimate off by | Likely period | Calibrated range |
|---|---|---|---|
| Mazda RX500 | 1 day | hit | hit |
| Jaguar D-type | 4 days | hit | hit |
| Toyota GT-One (TS020) | 5 days | missed by 2d | hit |
| McLaren MP4/4 | 21 days | missed | hit |
| Lamborghini Countach LP400 | 46 days | missed | missed |
Four of the five landed inside the calibrated range and two inside the tight likely period, almost exactly what the aggregate rates predict, which is itself a good sign that the sample is representative rather than lucky.
The Countach is the interesting failure. It returned after 33 days when 22 recorded appearances pointed to roughly 79, so the forecast had it months away. No amount of recalibration would have caught that. The information needed, that this particular car was coming back early, simply does not exist in its rotation history. It is a clean example of a blind spot built into the model, not a bug.
Within these five the errors run both ways, three early and two late. Across the full run of live predictions they do not: the mean signed error is −5.8 days, meaning cars come back almost a week sooner than the centre of the estimate on average. That is a real lean, not noise, and it is the model's clearest known weakness. A five-car sample was too small to show it, and an earlier version of this page wrongly concluded from those five that no such lean existed. Correcting it is the next piece of work on the forecast.
Where the data itself is weak
One limitation is worth stating plainly because it affects what can honestly be claimed. Appearance history that predates this site was imported in bulk, and that import carries a uniform assumed dealer stay rather than measured ones. Statistics about how long cars remain purchasable are therefore computed only from the period we have actually observed day by day, a much smaller sample, and one that shows real variation rather than the neat constant the imported history appears to suggest.
Return intervals, which drive every forecast on the site, are not affected: those come from appearance dates, which are reliable throughout.
This page will change
These figures are recomputed from live data, so they move as new returns are recorded. They may get worse. If a change to the model improves the hit rate, this is where you will see it; if it makes things worse, that will show up here too.
The full method behind these numbers is documented on the methodology page.