StrategyUpdated 21 Jul 20266 min read
Sample Size
Almost everyone dramatically underestimates how many bets it takes to tell a real edge from luck. The honest answer, if profit is your only evidence, is tens of thousands.
That number sounds like an exaggeration until you work it out. So let's work it out.
Why a season of bets tells you almost nothing#
The problem is a ratio. The signal you're looking for — your edge — is tiny. The noise around it — the variance of individual bets — is enormous by comparison.
Take a flat-staked even-money bet. Each one returns +1 unit or -1 unit. The standard deviation of a single bet is therefore about 1 unit, which is fifty times larger than the 2% edge you're hoping to detect. Over n bets, expected profit grows with n but the noise only grows with the square root of n. That's the whole story: you win the race eventually, but slowly.
Run the numbers on 500 bets — a busy season — with a 2% edge:
Your entire expected profit for the year is less than half of one standard deviation. Finishing the season up 30 units or down 30 units are both completely ordinary outcomes for the same method. A season of results does not distinguish a 2% edge from no edge, from a 2% disadvantage.
The formula behind the big number#
Put that in general terms. With per-bet edge e and per-bet standard deviation of about 1 unit at even money:
- Total expected profit after
nbets isn × e - Standard deviation of that total is the square root of
n - Signal in standard errors is therefore
e × √n
Set that equal to 2 — the usual "two standard errors from zero" bar — and rearrange: n = 4 ÷ e².
| Per-bet edge | Bets to reach 2 standard errors | Realistic? |
|---|---|---|
| 10% | 400 | Fantasy |
| 5% | 1,600 | Very unlikely |
| 2% | 10,000 | A good edge |
| 1% | 40,000 | Plausible and unprovable |
Two caveats, both of which make it worse. First, clearing two standard errors is the bar for this particular sample looking convincing; designing a test that would actually find a 2% edge most of the time it exists pushes the requirement to around 20,000 bets. Second, betting at longer prices increases the per-bet variance roughly in proportion to the odds, so backing 3.00 shots doesn't rescue you — it costs you.
At 10,000 bets a week, this is a research problem. At ten bets a week it's twenty years.
One percentage point is the whole game#
The reason the numbers explode is that the difference between a winning and losing method is microscopic. At even money, break-even is a 50% strike rate. 51% is a 2% edge and a career; 49% is a 2% loss and a slow bleed. Over 500 bets that is 255 wins against 245 — the whole distance between a career and a bleed is ten results out of five hundred.
Set the price to 2.00 in the break-even calculator below and drag the strike-rate slider by a single point in each direction. Then try it at 1.50 and at 4.00 to see how the required rate moves with the price.
At 2.50 you only need to be right 40.0% of the time. Winning 45% clears that, so every unit staked returns 12.5% on average — regardless of how often that feels like losing.
This is the whole reason strike rate is a bad scoreboard: 40% at 3.00 makes money and 70% at 1.30 loses it. The price sets the bar; your job is to clear it.
Five coin flips separating a good method from a bad one is why 500 bets can't tell them apart. It isn't a flaw in the statistics; it's the size of the thing being measured.
What a losing run actually proves#
Nothing, usually. Consider a method winning 55% of even-money bets — a 10% edge, wildly beyond anything the football markets realistically offer. The expected number of losing runs of length r or more in n bets is about n × p × q^r, where q is the loss probability. Over a 500-bet season with q = 0.45:
| Losing run | Chance of seeing it in 500 bets |
|---|---|
| 7 or more in a row | about 64% |
| 8 or more in a row | about 37% |
| 10 or more in a row | about 9% |
So a seven-bet losing run is the expected experience of an exceptional method. It is not evidence that anything is broken, and abandoning a method because of one is how people convert a good process into a bad one. This is also the entire reason staking plans and bankroll management exist: not to make money, but to guarantee that a routine run doesn't end you before the arithmetic gets a chance.
The reverse holds too. A ten-bet winning streak is not evidence you've found something.
Why ROI-based tipster claims are unfalsifiable#
Now apply the same maths to a sales pitch. Over 200 even-money bets, the standard error of ROI is 1 ÷ √200 = 7.1 percentage points. A tipster with exactly zero edge produces a 200-bet record somewhere between about -14% and +14% ROI, ninety-five times out of a hundred, by luck alone.
So a +12% ROI over 200 bets is not a claim you can test. It is inside the range that nothing produces.
It gets worse with selection. Start a thousand tipsters with no skill whatsoever, let each post 200 bets, and roughly twenty-five of them finish above +14% ROI. Those twenty-five now have a genuine record and a website. Nobody hears from the other 975. Every published ROI you see has already passed through that filter, which is why an ROI number with a sample size in the hundreds carries almost no information — and why "verified profit" verifies the arithmetic, not the edge.
What to measure instead#
If profit is too slow, measure the things that record a number on every bet regardless of outcome.
- Closing line value — did you take a better price than the market's final one, adjusted for the book's margin? Answers in hundreds of bets rather than tens of thousands, because it doesn't wait for the ball to go in.
- Calibration — when the model says 65%, does it happen 65% of the time? Every forecast contributes, winners and losers alike.
- Brier score — a single number for the accuracy of the probabilities themselves, comparable against the market's.
None of these prove profitability. They fail much faster than profit does, which is exactly what makes them useful: a method that's broken shows up in calibration and CLV within a season, while its ROI is still telling you a comfortable story.
Tofiko publishes calibration, Brier score and CLV instead of a headline return, because a return over any sample we could realistically collect wouldn't mean anything. On the CLV measure we currently show no demonstrated edge over closing prices, and we publish that rather than a profit figure from a sample too small to falsify.
Related
- Closing Line Value: The Only Scoreboard That Answers in a Season
- Calibration: The Only Promise a Prediction Can Keep
- The Brier Score: How We Grade Our Own Predictions
- Bankroll Management: Why Staking Beats Picking
- Break-Even Calculator — The Strike Rate Every Price Demands
- Closing Line Value Calculator — Grade a Bet Without Waiting for the Result
Frequently asked questions
How many bets do you need to prove a betting edge?
Far more than most people assume. At even money, reaching two standard errors above zero requires roughly 4 divided by the square of your edge. A 2% edge therefore needs about 10,000 bets to clear that bar, and closer to 20,000 for a test that would reliably find the edge if it were there.
Is 200 bets enough to judge a tipster?
No. Over 200 even-money bets the standard error of ROI is about 7 percentage points, so a tipster with no edge at all produces a record anywhere between roughly -14% and +14% ROI purely by chance. A +12% return over that sample is not evidence of anything.
How long can a losing run last with a genuine edge?
Long enough to break most people's nerve. A method winning 55% of even-money bets — an enormous edge by real standards — has a losing run of seven or more in a 500-bet season about two seasons in three, and a run of ten or more roughly one season in eleven.
What should I measure instead of profit?
Closing line value, calibration and Brier score. All three record something on every bet rather than waiting for outcomes to accumulate, so they produce an interpretable answer in hundreds of bets rather than tens of thousands.