Skip to main content
TRUE OVERLAYAI racing intelligence

Research standard

From licensed racecard to settled record.

The final probability is a transparent blend of de-vigged market consensus, stored ratings, and a bounded AI component. Application code controls the blend, evidence, arithmetic, immutability, and settlement.

The analysis contract

Ranked field

Every active runner appears exactly once in model order.

Win probability

A deterministic blend of de-vigged market consensus, stored ratings, and bounded AI is normalized to 100%.

Fair decimal price

Calculated deterministically from the normalized probability.

Finish bands

Paid top-two and top-three estimates are deterministically derived from the locked win vector under a disclosed Harville assumption and prospectively scored in the public platform ledger.

Market comparison

Median bookmaker prices are converted to a de-vigged field baseline; best price is kept separately for expected-value arithmetic.

Pace map and shape

AI classifies early run style only from supplied form and comments; code then derives a coverage-gated whole-field pressure diagnosis without changing the forecast.

Field-level evidence

Evidence labels resolve to non-empty source fields for that runner.

Uncertainty

Missing inputs and data completeness remain visible beside the estimate.

Audit identity

Model ID, prompt version, input hash, source time, and generation time are retained.

Settlement

Official results close the record without deleting unsuccessful predictions.

Calibration

Settled runner probabilities feed a public multiclass Brier score, matched market-skill benchmark and forecast-band reliability table.

Input policy

The AI payload is limited to the normalized race, active runner identities, form, conditions, and comments. A runner name is treated only as an identifier; jockey or trainer reputation is not inferred from a name. Every active runner must have form or a source comment. The projected early pace map classifies run style only from those two qualitative runner fields, must cite one of them, and uses an explicit unknown state when the text cannot support an inference. It is not a sectional-timing feed. Bookmaker prices and numerical ratings are withheld from the AI: market consensus and official, performance, and speed ratings enter only through their separate deterministic components, preventing either channel from being counted twice. Prediction generation also requires a source update no more than 30 minutes old and at least 50% completeness across active runners' form and rating fields. Provider credentials and unrelated account data never enter the prompt.

Field-relative pace shape

The application converts the locked runner-level styles into a whole-field setup diagnosis without another model call. At least three supported classifications, 60% field coverage, and two medium/high-confidence styles are required; otherwise the result is withheld. Only medium- or high-confidence leader and prominent labels determine whether early pressure is limited, balanced, or strong. Setup notes identify conditions such as a lone control candidate, a contested lead, pace to aim at, or pace dependency, but do not upgrade or downgrade a runner's win probability. The calculation does not use sectional timing, course, distance, draw, going, or historical pace-bias evidence, and planned tactics or the start can invalidate it.

Prospective regime evidence

The trust map uses five fixed cuts—region, surface, race type, field size, and locked input confidence—on the current forecast contract. For every segment, model and de-vigged-market multiclass Brier errors are measured on the same complete settled races. The paired advantage equals market error minus model error. A 95% t interval is calculated over the per-race differences; no directional label is allowed below 30 matched races, and a larger sample remains inconclusive when its interval crosses zero. Multiple segment views can reveal operating weakness, but they are descriptive diagnostics rather than independently validated or optimized strategies.

Validation sequence

  1. 1Schedule the shared forecast inside the final 30 minutes before the off; confirm the race is upcoming, active, has at least two runners, and passes the published freshness and completeness floors.
  2. 2Require every active runner and rank exactly once; reject unknown identities.
  3. 3Require the advisory AI distribution within a narrow total tolerance, then normalize it.
  4. 4Require a bounded pace style, confidence, and non-empty form-or-comment evidence field for every runner; replace all evidence text with the exact stored value.
  5. 5Reject stale quotes, filter prices more than 35% from the runner median, require at least two accepted books, and withhold the market component when accepted prices remain too dispersed.
  6. 6Build a median-book, de-vigged market distribution from the accepted quotes and keep the best accepted price separate.
  7. 7Transform stored official, performance, and speed ratings into a field-relative distribution.
  8. 8Blend market, ratings, and AI with explicit weights; then rank the final probabilities in code.
  9. 9Calculate fair price, expected value, and probability disagreement in application code.
  10. 10Withhold a value signal when the market is unreliable or data completeness is too low, and remove it from the live board after its quote is 30 minutes old.
  11. 11Retry malformed structured output once; otherwise return analysis unavailable.
  12. 12Store the completed prediction with its versioned input hash and generation identity.

Forecast revisions

Eligible races are normally checked for a refreshed shared forecast at most once every ten minutes inside the final 30-minute window. Every current-contract forecast locks normalized going, surface, distance, active-runner count, and a one-way active-field fingerprint. A field or condition mismatch bypasses the ordinary cooldown inside that window; until a compatible snapshot is stored, current repricing and signal interpretation fail closed while the old record remains visible for audit. An otherwise unchanged normalized input reuses the existing record, while a changed input creates a new immutable snapshot. If only market prices or numerical ratings changed, the system reuses the validated qualitative generation and deterministically recomputes the blend with zero additional model tokens, retaining a lineage pointer to the source AI record. A model, contract, race-condition, active-field, form, or comment change requires fresh structured generation. Overlay members can compare the two newest same-contract snapshots, see whether qualitative AI was reused or regenerated, and inspect market, ratings, and AI component movements—including movements that offset before reaching the final probability. These are arithmetic comparisons of stored values, not causal attributions, and the view does not prove that newer is more accurate.

Live decision monitoring

The paid racecard keeps every stored model probability fixed while recomputing the filtered, de-vigged market from the newest timestamped quotes. It compares each current signal, observed price, and expected value with the locked record and labels the decision as new, strengthened, unchanged, weakened, lost, or unavailable. EV-only WATCH and VALUE prices answer what decimal price would meet the respective arithmetic threshold, rounded upward to two decimals; data-confidence and probability-edge gates still apply separately. A scratch or other active-field change makes the locked distribution incompatible, while stale, sparse, or dispersed quotes withhold the current decision. This is market-monitoring intelligence, not a new forecast, a prediction edit, or evidence that the displayed price was executable.

Prospective research receipts

A paid member can freeze a shortlist, monitor, or pass intent for a runner before the scheduled off. The browser submits only identities, intent, and an optional rationale; the server reloads the race and current-contract prediction, rechecks field and condition compatibility, rebuilds the full accepted market, requires 100% reliable active-field coverage, and stores the selected quote and model evidence atomically under a one-per-member, revision, and runner identity. Receipts cannot be edited into a better-looking hindsight record. Confirmed card drift produces a deduplicated in-app warning while preserving the original snapshot. Per-race result ingestion uses one repeatable-read transaction to upsert official outcomes, settle forecasts, and create a deduplicated member receipt summary for the provider result revision. Receipt lifecycle notices do not enter the email queue because capture is not email consent. The personal dashboard calculates observed win rate minus average locked probability as a calibration gap and mean squared binary probability error as a Brier score, alongside quote-versus-SP summaries overall and by intent. It uses at most the newest 1,000 receipts and discloses truncation. Because receipts are self-selected and several runners can share one race outcome, the sample is neither random nor independent. The metrics are descriptive process evidence, not proof of execution, profit, ROI, closing-line efficiency, calibration, or model skill.

Finish-band projection

The paid top-two and top-three lens applies Harville's sequential-ranking formula to the complete locked win distribution. It asks what the finish bands would be if the remaining runners' relative strengths stayed proportional after each earlier finisher is removed. The calculation is deterministic and its derived fair odds contain no bookmaker margin. Top-two and top-three binary Brier scores and calibration bands accumulate prospectively in the public platform-controlled ledger only when every forecast runner has an official result; non-finishers are negative outcomes and dead heats can produce more positives than the nominal band. This is not independent validation, does not independently model finish-position behavior, and must not be interpreted as bookmaker place terms or a recommendation. Malformed, duplicate, incomplete, or non-normalized runner vectors produce no projection.

Scenario Lab

The paid Scenario Lab is a sensitivity calculation over the three derived probability components already stored with a forecast. Members can rebalance market, ratings, and qualitative AI weights; the application renormalizes the resulting field to 100% and recalculates scenario ranks, fair prices, expected value at the locked observed price, and disagreement with the locked market. An automated stress test also evaluates the locked weights and a disclosed finite grid of balanced and channel-led mixes, reporting probability, rank, expected-value, and probability-edge ranges plus how often the locked leader remains first. An edge is labelled robust only when both EV and probability disagreement remain strictly positive in every disclosed mix. Duplicate effective mixes are removed and an unavailable market channel receives zero weight. These ranges measure weight sensitivity, not statistical confidence or total outcome uncertainty; the grid cannot prove that an edge is real or complete. The lab makes no additional model request, stores no scenario, cannot alter the immutable forecast or its signals, and never enters public performance reporting.

Performance scoring

After official settlement, the public ledger calculates a multiclass Brier score across every runner, with lower values indicating less probability error. It also calculates Brier skill against the locked, filtered and de-vigged market only where both complete probability ledgers exist for the same race: 100 × (1 − model Brier ÷ market Brier). Positive skill means the model scored better on that matched sample; negative means it scored worse. Dead heats use equal fractional outcomes. Region and locked data-confidence slices expose operating regimes that an aggregate can hide. Sample counts and evidence tiers remain visible because a short run is descriptive, not proof of a durable edge.

Interpretation limits

A fair price is a model estimate, not the unknowable true probability. Market comparison does not account for an individual's commission, limits, liquidity, tax position, staking, or jurisdiction. Performance conclusions should wait for a sufficiently large, prospectively recorded sample.