When a strategy tested on historical data produces a good result, the feeling it produces is confidence. That feeling is usually larger than the result and almost never matched going forward.
The cause is rarely that the strategy is bad. It’s a misreading of what the backtest measured.
A backtest is a measurement, not a prediction
A backtest answers exactly one question: if these rules had been run over this slice of data, what would have happened?
That’s a valuable answer. It is not the answer to “what happens next”, and several conditions have to hold before it even approximates one. When they don’t, what you measured isn’t the strategy’s performance. It’s how well you fitted the strategy to that data.
Three mechanisms that inflate the result
Overfitting. The more parameters a strategy has, the easier it becomes to match past data. Give it enough of them and you can produce a near-perfect result for any historical window, with no bearing on the future whatsoever. The practical tell: if moving a threshold by one percent changes the result substantially, what you found is a coincidence, not a rule.
Look-ahead. The test using information that couldn’t have been known at that moment. The most common form is trading inside a candle at its closing price: a rule that assumes you knew the close cannot be executed in real time. The results table doesn’t flag this as an error. It just looks good.
Survivorship bias. Testing history against the assets that are listed today. On the crypto side this effect is large: taking today’s traded list back three years excludes everything that disappeared during those three years. Looking backwards, the universe is better than the one you would actually have faced.
Where friction sits in the result
For a backtest to approach reality, three items have to be inside it: trading fees, spread and slippage.
Their effect looks small per trade and turns decisive in aggregate. In a strategy that trades once a day the difference is limited; in one that trades ten times a day the same percentage applies ten times over. That is the most common reason a high-frequency rule is profitable in a backtest and unprofitable in reality, and it has nothing to do with the strategy’s logic.
Slippage in particular gets skipped. Assuming the order fills at the price available rather than the price you saw makes the test harder, and closer to what happens.
How to make the test harder
The way to make a backtest more trustworthy isn’t to search for a better result. It’s to try to break the one you have.
Split the period in two. Develop the rules on the first half only and never open the second. When the parameters are settled, run the second half once. If the result is close to the first half, you have something. If it’s markedly worse, what you had was excess fit.
Look at market regimes separately. A total return collapses a strategy that gains in a trend and bleeds in a range into one favourable number. Separating the periods reveals which assumption the strategy actually depends on.
Look at the worst run. Not the total, but the longest string of consecutive losses and the deepest drawdown. What you’ll live through going forward isn’t the average. It’s that run.
What comes after the backtest
A good backtest isn’t the decision to go live; it’s the precondition for making that decision. Two things the test can’t measure remain: whether orders actually fill, and what you do during a losing run.
Both only show up with a small amount of money under real conditions. In Finbula, running a bot in Demo mode and taking it to Real mode with a small amount happen from the same screen, and control of the bot always stays with you. But knowing which stage you’re in isn’t something a screen can tell you.



