How to Validate Horse-Racing Predictions Before Trusting Them
10 Sep 2026 · 7 min read · True Overlay team
A racing forecast can look impressive and still be poorly calibrated. Five winners from ten selections might be luck, a favourable sample, or a model that only identifies favourites. Validation asks a harder question: when a system says 30%, do events in that group happen about 30% of the time, and does the system add information beyond a sensible baseline?
That distinction matters because a single result is noisy. A horse either wins or it does not; a probability forecast can still be good when the favourite loses. The useful object to audit is the locked probability, its information set, and the result observed later—not a screenshot of a pick after the race.
Start with a proper baseline
The first comparison should be a market baseline built from the prices that were actually available when the forecast was made. A model that beats an unrealistic baseline tells you very little. Record the quote time, remove the bookmaker margin when comparing probabilities, and keep the model's forecast frozen so later price moves cannot leak into the test.
A second baseline can be simpler still: the historical win rate for the relevant field or race class. The point is not to prove that a model is useful because it beats a straw man. It is to ask whether its extra inputs improve a fair comparison on races it did not use for development.
Use calibration, not just winners
Group forecasts into probability bands such as 10–20%, 20–30%, and 30–40%. In a well-calibrated sample, the observed win rate in each sufficiently large band should be close to the midpoint. A reliability diagram makes that gap visible. Small samples can swing wildly, so show the number of runners and an uncertainty interval instead of turning one band into a claim.
The Brier score is the mean squared error of a probability forecast against the eventual outcome: (forecast probability − outcome)². Lower is better. Allan H. Murphy's original 1973 decomposition separates reliability, resolution, and uncertainty (https://doi.org/10.1175/1520-0450(1973)012<0595:ANVPOT>2.0.CO;2). The scikit-learn calibration guide explains why a Brier score should not be treated as a calibration-only measure (https://scikit-learn.org/stable/modules/calibration.html).
Protect the test from hindsight
Use chronological splits: earlier races for design decisions and later races for the holdout. Do not tune weights, thresholds, or feature choices on the same races used to report the final score. If a forecast is regenerated after a going change or non-runner, keep the revision history and evaluate the contract that was actually live at the decision time.
Report enough context to make the result interpretable: race count, runner count, date range, region, missing-data rules, market timestamp, model version, and the exact outcome definition. A result with no denominator, no time window, or no frozen forecast is marketing copy rather than evidence.
What a useful public record looks like
A useful ledger shows the forecast before the result, keeps losing records, and labels what is measured versus what is still unconfigured. It can include calibration bands, a matched market comparison, and drawdown context, but it should not imply that past scores guarantee future profit. Responsible use also means staking only money you can afford to lose and never changing a plan to chase the last result.
True Overlay's public methodology and results pages are designed around those boundaries: inspect the arithmetic in the worked example, then check what the live platform has actually settled. If the underlying feed, evaluation window, or sample is unavailable, the honest status is unavailable—not a made-up estimate.
See the idea in a worked example
Compare an estimated chance with a market price, understand the uncertainty, and see what the product is designed to show. No account required.
Explore the example
