modelsUpdated 22 Jul 20264 min read

The Track record page

The track record page is the receipt. It shows the run of forecasts the models actually made — what was picked, and how those picks resolved — as a history you can scroll rather than a score you have to trust.

Two pages, two jobs

Track record = the run of real forecasts. Performance = how good the probabilities behind them were. A track record can look decent while the underlying probabilities are mediocre, and the other way round. Read both.

The Tofiko Track record page listing the run of pre-kickoff forecasts and how they resolved
Every entry here was recorded before kick-off. The list is short on purpose — nothing backfilled is counted.

How it differs from Performance#

The Performance page is a grading exercise. It asks whether the probabilities were well made: Brier score against the market, calibration, whether strong favourites won as often as claimed. It reports things like model Brier 0.610 against the market's 0.594 — the market ahead, stated plainly.

Track record asks a simpler question: what did the models say, and what happened next? It is the raw material Performance is calculated from. If a headline number ever looks surprising, the track record is where you go to see the individual forecasts behind it.

Neither page is a substitute for the other:

  • A track record with a good stretch of winners proves nothing if the probabilities were poorly calibrated — you got lucky on prices.
  • Good calibration with almost no forecasts behind it proves nothing either — there is not enough evidence yet.

Why backfilled history is excluded#

The honest counts on this site — 915 genuine forecasts — include only predictions recorded before kick-off.

It would be trivial to generate a much larger history by running the current model over past seasons. Sites do this all the time, and it always produces a flattering record. The reason is not fraud, it is leakage: a model built and tuned on data that includes those results already "knows" too much about them. Team strength ratings, league baselines and calibration curves all absorb information that was not available at the time.

So a backfilled record measures how well the model fits the past, which is a question nobody betting next weekend needs answered. Excluding it costs a lot of impressive-looking rows and keeps the number meaningful.

Why sample size decides whether any of this matters#

A run of picks needs to be long before it separates skill from luck, and "long" is much longer than intuition suggests.

  • Short runs are dominated by variance. Twenty winning bets in a row is a thing that happens to people with no edge at all.
  • The size of the edge sets the size of the sample. A small genuine edge needs tens of thousands of bets before profit alone proves it exists — see how many bets you actually need for the arithmetic.
  • This is precisely why Tofiko's verified-signal gate requires at least 200 real picks, four consecutive weeks and closing-line value, rather than a profit figure. At present 0 of 17 league groups clear it.
The trap to avoid

Do not scroll the track record looking for a hot streak to follow. A streak is the easiest pattern to find in random data and the hardest to profit from. Judge the record by its length and its Brier score, not by its best stretch.

How to use it#

  1. Check the count before anything else. How many pre-kickoff forecasts are there for the leagues you care about? If it is thin, stop there.
  2. Cross-read with Performance. Take the by-league-group table on the Performance page and see whether the model or the market scored the better Brier for that group.
  3. Treat the whole thing as a second opinion. On 1X2 the market is currently ahead. The models are useful for framing a price, not for outsourcing a decision.
  4. Keep your own record separately. Your results depend on the prices you got and the stakes you set, neither of which the model track record knows about — log them in My Picks.
  5. Test rules properly rather than eyeballing history. The Labs exist for that, with the usual warnings about overfitting.

Related

Frequently asked questions

What's the difference between Track record and Performance?

Track record shows the run of actual forecasts — what was picked and how it resolved. Performance grades the quality of the probabilities behind those picks using Brier score and calibration. One is the history, the other is the scoring.

Why is the track record so short?

Because only forecasts recorded before kick-off count. There are 915 of those. Backfilled predictions — generated after results were known — are excluded, even though including them would make the record look far longer and far better.

Does a good run on the track record mean the model has an edge?

No. A run of winners over a small sample is the single most common way people fool themselves. Tofiko's own gate requires at least 200 real picks, four consecutive weeks and closing-line value before a league group counts as verified, and 0 of 17 groups currently pass.

Should I judge the track record on profit?

No. Profit over a short run is dominated by luck and by which prices you happened to take. Closing-line value and calibration move much faster towards the truth, which is why Tofiko grades on those instead.