# Sample Size in Betting Tests: Why There Is No Magic Number

The direct answer

There is no fixed number of bets that proves a strategy works. The sample required depends on the expected edge, odds range, strike rate, variance, dependence between bets and the precision needed. A thousand bets can be inadequate for a tiny claimed edge, while a much smaller sample can reveal a gross calculation error.

The honest question is not, “Have I reached 500 bets?” It is, “How uncertain is the estimate, and could ordinary variation still explain what I see?”

This guide uses binomial uncertainty and reproducible simulations to show why. The calculations are illustrations, not proof that any betting strategy is profitable.

Why a fixed threshold fails

Imagine two systems. One backs selections near evens and claims a 55 per cent strike rate. The other backs 20/1 outsiders and claims a small return on investment. Their result distributions are completely different. Five hundred bets from each do not provide equal evidence.

Even within one market, repeated selections may not be independent. Bets from the same match, the same model version or the same weather pattern can share errors. Counting every leg as a fresh observation exaggerates the effective sample.

Changing the rules during the test creates another problem. A strategy tried with ten filters until one historical version looks profitable has consumed more evidence than its final bet count suggests. The displayed sample ignores the selection process.

A simple uncertainty calculation

For independent binary outcomes with probability `p` and sample size `n`, the standard error of the observed strike rate is:

`square root of p x (1 - p) / n`

A rough 95 per cent range is the observed rate plus or minus 1.96 standard errors. Near a 50 per cent strike rate, this produces the following illustrative widths.

Number of betsApproximate 95% uncertainty around a 50% rate
100plus or minus 9.8 percentage points
400plus or minus 4.9 points
1,000plus or minus 3.1 points
2,500plus or minus 2.0 points
10,000plus or minus 1.0 point

After 100 bets, a 55 per cent observed strike rate is not sharply distinguishable from 50 per cent. Even after 1,000, the rough range remains wide enough to matter when the claimed edge is small.

For small samples or probabilities far from 50 per cent, use an appropriate binomial interval such as Wilson's rather than relying blindly on the rough normal approximation.

Odds and returns add more variance

Strike rate alone does not describe a betting return. A win at decimal odds of 2.00 and one at 15.00 have different effects. Variable stakes, each-way terms, dead heats, voids and commission add further variation.

Return on investment can therefore remain unstable long after the win-rate estimate begins to look respectable. A few long-priced winners may dominate the full history. Publish the distribution of odds and contribution of the largest wins rather than one headline percentage.

The sports betting variance calculator shows simulated ending balances and drawdowns under declared assumptions. It is a way to understand possible spread, not a forecast of future profit.

A reproducible simulation

Consider a deliberately simple model: 1,000 independent bets, £1 flat stake, true win probability 52 per cent and decimal odds of 2.00. The theoretical expected profit is £40 because each bet has expected profit of four pence.

To test the result distribution, set a published pseudo-random seed, generate 1,000 sequences under those fixed assumptions and record the profit from each sequence. Anyone using the same generator, seed and code should reproduce the same figures.

A typical simulation will show many profitable runs, but also some losing runs despite the positive assumption. Changing the seed changes individual paths, not the theoretical expectation. The responsible report includes:

1. generator and seed; 2. number of simulated runs; 3. bets per run; 4. true win probability; 5. odds and staking rule; 6. median result and selected percentiles; 7. proportion of losing runs.

Never present one attractive simulated path. That is illustration shopping.

Power depends on the edge you want to detect

Large effects are easier to detect than small ones. Distinguishing a genuine 60 per cent win probability from 50 per cent at even money needs fewer observations than distinguishing 51 per cent from 50 per cent.

That creates a practical problem in efficient betting markets. Plausible sustainable edges are often small, so the sample needed for a confident statistical distinction can be very large. Market conditions, limits and the model itself may change before that evidence accumulates.

The answer is not to lower the standard after seeing a profitable run. Use prior reasoning, out-of-sample testing, closing-price evidence, calibration and transparent uncertainty together. No single statistic carries the entire claim.

Backtests need a larger honesty budget

A backtest can process thousands of historical events and still mislead. Look-ahead bias, survivor bias, unavailable prices and repeated filter testing can make the apparent sample much stronger than it is.

Ask whether the information existed at the recorded betting time, whether the odds were genuinely obtainable, whether limits and commission were represented and how many alternative rules were tried. Preserve failed versions. The guide to evaluating a betting system covers these checks in more detail.

A final rule chosen after inspecting all history should be tested on later data that played no part in selection. Calling the same data both training and proof is not an independent test.

Dependence reduces effective sample size

Ten player bets from one football match are not necessarily ten independent observations. A red card, tactical change or faulty team assumption can affect all of them. Likewise, several horse-racing selections exposed to the same going forecast may share one error.

Cluster results by event, day, competition or model decision where appropriate. A conservative analysis can resample whole clusters rather than individual bets. At minimum, disclose that the independence assumption is doubtful.

This is especially important for Bet Builders. Their legs are deliberately connected. The guide to Bet Builder correlation explains why multiplying individual probabilities is unreliable.

Stop rules can bias the story

If a test stops as soon as it reaches £1,000 profit but continues indefinitely when losing, the reported winners have a built-in advantage. The same issue occurs when results are checked daily and announced only when a significance threshold first appears.

Choose the test horizon, decision rules and primary metric before the results are known. If monitoring must be continuous, use a method designed for sequential analysis and report every look. Do not quietly turn a fixed test into repeated attempts.

A practical evidence standard

1. State the strategy before testing it. 2. Define the target market, odds range, staking and exclusions. 3. Record every qualifying bet, not only those placed. 4. Preserve offered prices and timestamps. 5. Report sample count, turnover, return, odds distribution and uncertainty. 6. Separate model development from later evaluation. 7. Test sensitivity to a slightly worse price and plausible rule changes. 8. Show drawdown and the contribution of the largest winners. 9. Check probability calibration where forecasts exist. 10. Avoid a profitability claim stronger than the evidence.

The expected value guide provides the companion process for assessing whether a forecast and its available price justify a bet. Together, expected-value and sample-size analysis can expose confident stories built on weak records.

Frequently asked questions

Are 100 bets enough to judge a strategy?

Usually not for a modest claimed edge. Near a 50 per cent strike rate, the rough 95 per cent uncertainty is almost ten percentage points after 100 independent bets.

Is 1,000 bets enough?

It is stronger evidence than 100, but not a universal proof. A small edge, long odds, correlated bets, rule changes or backtest selection can still leave substantial uncertainty.

Should void bets count in the sample?

Record them, but define how they enter the primary metric. A void normally adds no win or loss and no effective turnover, yet a high void rate may reveal an operational issue.

Can closing-line value replace a large sample?

No. It can provide useful supporting evidence about price quality, but it depends on the market and reference price. It does not prove realised profitability.

Does a profitable backtest show that future bets will win?

No. It shows what the chosen rules produced on the tested historical record, subject to data and design quality. Future performance can differ materially.

Free newsletter

Join the BetOwl newsletter

Receive regular free racing tips, betting articles, and updates.