modelsUpdated 22 Jul 20264 min read
Labs: Strategy Lab, Backtest and Martingale Lab
The Labs group in the sidebar holds three tools for testing ideas before they cost you money: Strategy Lab, Backtest and Martingale Lab. They are simulation tools. Nothing you run here places a bet, and nothing you run here proves an edge.
Labs is for falsifying ideas, not confirming them. The useful outcome of a session here is usually finding out that a rule you liked does not survive contact with history — that is a saving, not a failure.

What each tool does#
Strategy Lab lets you define a betting rule — the kind of selection you would take, and the conditions under which you would take it — and run it over historical fixtures to see what it would have produced.
Backtest runs a defined approach across a body of past matches and reports the outcome. It is the measuring instrument behind any strategy you build.
Martingale Lab simulates the doubling-after-a-loss staking system so you can watch what happens to a bankroll under it. See why martingale fails for the mechanics.
All three sit apart from the model grading on the Performance page. Performance reports what the models actually did before kick-off; Labs reports what a rule would have done in hindsight. Those are very different claims.
The traps, in order of how often they bite#
1. Backtests overfit. Every historical sample contains real signal and a great deal of noise. A rule with enough adjustable parts will happily describe the noise. The tell is a strategy that only works with one specific set of thresholds, and collapses if you nudge any of them.
2. Tuning until it looks profitable is not research. If you try forty variations and keep the best one, you have not found an edge — you have run a search for the luckiest configuration in your sample. Somebody doing this with random numbers would also find a "winner".
3. A strategy tested on the data that inspired it proves nothing. This is the subtle one. If you noticed a pattern in last season's away underdogs and then backtested away underdogs on last season, the test cannot fail. You already know the answer; the backtest is just repeating it back to you.
4. Simulated prices are not the prices you would have got. Historical odds are a snapshot. Limits, movement and availability all reduce what a rule captures in practice.
5. Small samples flatter everything. Read how many bets you actually need before you trust any result with a few dozen selections behind it.
Tofiko's own models do not currently clear the bar the Labs are testing against. On 1X2 the market's Brier score is better than the model's (0.594 against 0.610), and 0 of 17 league groups pass the profitability gate. Any backtest that suddenly shows a large, easy edge deserves more suspicion, not less.
In-sample versus walk-forward, in plain words#
In-sample means you designed the rule and tested it on the same matches. The result is a description of that period. It is close to worthless as a prediction, because the rule was shaped by the very outcomes it is being scored against.
Walk-forward means you build the rule on an early stretch of history, then test it on the next stretch it has never seen — and repeat, rolling forward. Each test is a genuine question rather than a rehearsal.
A simple discipline, even without dedicated tooling:
- Split your history in two before you start. Design on the first half only.
- Write down the rule and its thresholds. Do not change them.
- Run it on the second half, once.
- If it falls apart there, believe the second half. That is the honest result.
- If you go back and tweak, you have spent your out-of-sample data. Any further test is in-sample again.
The Martingale Lab exists to demonstrate ruin#
Martingale doubles the stake after every loss so that one win recovers everything. It produces a long, smooth run of small wins followed by a single catastrophic sequence, and the mathematics guarantee that sequence arrives.
The Lab is there so you can see it rather than take it on faith: run it, extend the sample, and watch the stake requirement explode. Eight losses in a row from a one-unit start needs 256 units on the ninth bet.
It is not a recommendation. If you want staking that survives a losing run, read staking plans and work out the strike rate a price actually requires with the break-even calculator.
Taking an idea out of the Lab#
If a rule survives an out-of-sample test and you still want to use it, treat live results as the real experiment: log every selection in My Picks at the price you actually got, decide the sample size in advance, and judge it on closing line value rather than on how the first fortnight went.
Related
- How to read the Performance page
- The Track record page
- My Picks: your betting workspace
- The Martingale System: Why Doubling Up Always Ends the Same Way
- Staking Plans Compared: Flat, Kelly and the Progressions That Ruin You
- Sample Size: How Many Bets Before You Know You Have an Edge
- Bankroll Management: Why Staking Beats Picking
Frequently asked questions
If my backtest is profitable, does the strategy work?
Not on its own. A rule adjusted until it looks good on history is fitted to that history, not to football. The test that matters is whether it holds up on data you did not use while designing it.
What is overfitting, in one sentence?
Overfitting is when a rule describes the noise in your sample rather than anything that repeats — it looks excellent on the data it was built from and falls apart everywhere else.
Why does Tofiko include a Martingale Lab if martingale doesn't work?
To show you what it does. Doubling after a loss looks unbeatable until the losing run arrives, and the fastest way to understand that is to watch a simulation hit the table limit or the end of the bankroll.
How many bets does a strategy need before I believe it?
More than feels reasonable. Tofiko's own verified-signal gate demands at least 200 real pre-kickoff picks, four consecutive weeks and closing-line value — and 0 of 17 league groups currently pass it.