math · Updated 21 Jul 2026
Calibration: The Only Promise a Prediction Can Keep
"65%" is not a pick — it's a promise that over many such calls, about 65 in 100 land. Calibration is the audit of that promise.
A single match proves nothing
When a 65% favourite loses, nothing went wrong — the forecast itself said that happens 35% of the time. This is the uncomfortable core of probabilistic prediction: no single result can validate or refute it.
What can be audited is the collection. Gather every call where the model said 60–70%, and count. If about two-thirds of them won, the model is calibrated in that range: its numbers mean what they say. If only half won, the model is overconfident — its "65%" was really a 50%, dressed up.
The two ways to fail
- Overconfidence — promising 84%, delivering 62%. The common failure of sports models (and of humans). Every number reads stronger than it is; users slowly learn not to trust it.
- Underconfidence — promising 55%, delivering 70%. Rarer, and it wastes information: the model knows more than it dares to say.
A calibrated model can still be beaten by a sharper one (see the Brier score, which grades precision and calibration together) — but an uncalibrated model is worse than useless, because its numbers actively mislead.
How this shapes what you see on Tofiko
Our confidence chips are probability bands, and each carries a checkable claim:
- ▲ FAV — top outcome at ≥65%. "Favourites like this have won ~3 of 4."
- ▲ CLR — 55–65%. ~ LEAN — 45–55%. · OPEN — below 45%.
Those sentences are only honest if calibration holds — so we test it weekly, in public. On the Performance page, every league group's drilldown shows "promised vs delivered": the model's average stated probability in each band next to the actual win rate, on genuine pre-kickoff forecasts only (backfilled history is excluded — it can prove anything and therefore proves nothing).
One thing goal-based models systematically underrate is the draw — across thousands of matches our models' average draw probability ran ~2.5 points below the actual draw frequency. Until that's recalibrated at the source, we don't trust the model's own draw number for warnings; the ✕ DRAW risk chip fires off the market's fair draw probability instead, which tests better. Honesty includes knowing which of your own numbers not to use.
The takeaway for reading any forecast
Ask two questions of any prediction service — including us:
- Do they publish probabilities, or just picks? A pick without a number can never be audited.
- Do they show promised-vs-delivered? If a site claims "80% accuracy" but won't show the calibration table behind it, the number is marketing, not measurement.
Related
Frequently asked questions
What does it mean for a prediction model to be calibrated?
When it says 65%, that outcome happens about 65% of the time over many matches. Calibration compares promised probabilities with delivered frequencies, bucket by bucket.
Can a single wrong prediction prove a model bad?
No. A 65% favourite losing is expected 35% of the time — that's the forecast working, not failing. Only a pattern across many matches (65% calls winning far less than 65%) is evidence of a problem.
What is overconfidence in a prediction model?
Promising more certainty than reality delivers — e.g. calls labelled 84% that win only 62% of the time. It's the most common failure mode of sports models and the main thing calibration testing exists to catch.
How does Tofiko check its own calibration?
Weekly, on genuine pre-kickoff forecasts only. The Performance page shows 'promised vs delivered' for every confidence band in every league group — the model's average stated probability next to the actual win rate.