The curve was beautiful until real fills got involved
A backtest that works and a live account that loses are usually two different strategies wearing the same name. Four gaps cause almost all of it. Each one leaves evidence you can find in a few evenings.
Four gaps explain nearly every backtest that dies live
The test was fitted to its own sample. The test skipped the costs you actually pay. You took trades the rules never described. And real money changed how you pressed the buttons.
Those four cover the great majority of cases, and they fail in different ways. Cost gaps make every trade slightly worse in a uniform way. Overfitting makes the live results look random rather than merely smaller than the test. Drift shows up as trades in your statement that have no matching signal in the test.
Start by measuring rather than guessing. Line the two records up trade by trade. The shape of the damage points at one of the four. Traders usually assume it is psychology, and the arithmetic usually finds costs first.
Price the costs on a real contract and watch the edge shrink
Most tests fill at the mid, instantly, at every price they asked for. Live, you cross the spread going in and again coming out.
Run the numbers on the E-mini S&P, where one tick is 12.50 dollars. Give up a single tick each way and that is 25 dollars gone. Say your commission and fees come to 4 dollars round turn. Every trade now starts 29 dollars behind where the test started it.
If the backtest showed an average of 45 dollars a trade over 400 trades, that is 18,000 dollars gross. Take off 29 dollars times 400, which is 11,600. You are left with 6,400 and one bad month wipes it.
Any system whose average trade is a handful of ticks is mainly a costs experiment. Re-run the test using your worst realistic fill.
A curve tuned until it looked good has memorised the sample
Every time you nudge a parameter and keep the better result, you fit the test harder to the noise. That noise belongs to one stretch of data and nothing else.
Two checks catch it. First, hold back data you never looked at while building, then run the final rules on that block exactly once. A second pass turns it into more of the same sample. Second, move each parameter a few notches either side of your chosen value.
That second check matters more than people expect. A real edge sits on a plateau: a 12-period setting and a 14-period setting both work, roughly as well. A fitted result sits on a spike, where 13 prints a lovely curve and 12 and 14 both lose. If your settings are a spike, live trading will find that out for you.
Forward test at small size and log every deviation
Print the rules on one page before you start. Entry condition, stop, target, size, and the trades you are forbidden to take.
Then run it at the smallest size your broker allows. Fix the number of trades before the first order. Small size is the point. It keeps the money quiet enough for you to follow the rules while still making the fills real.
Keep a deviation log next to the trade list. One line per exception: signal skipped, entry chased, stop moved, or a trade taken with no signal behind it. That log is the measurement of the third and fourth gaps, and no amount of reflection replaces it.
At the end of the block, count deviations first. If a fifth of your trades were off-plan, the strategy has not been tested yet.
Compare live to the test per setup and per hour, never in total
A single blended difference tells you the account lost. It never tells you where.
Match the two records on the same axes. Count the signals the test produced while you were at the desk. Then count the ones you took. A gap there is drift, and the missing trades are usually the uncomfortable ones, which are often the good ones.
Then split both records by hour. Say the tested edge lived in the first thirty minutes. Check whether your fills in that window are two ticks worse than the rest of the day. If they are, the problem is cost, with a time of day attached.
One more comparison worth running: average winner and average loser in R, live against test. Same win rate with a smaller average winner puts the blame on your exits.
The forward test setup, step by step
Ten steps. The first six happen before you place a single live order.
- Write the entry, stop, target and size rules on one page.
- Re-run the backtest with your worst realistic slippage and real commissions.
- Test your parameters a few notches either side and keep only plateau settings.
- Run the final rules once on data you never used while building.
- Fix the number of forward-test trades and the size before you start.
- Create a deviation log with one row per off-plan action.
- Tag every live trade with the setup name and the hour it was taken.
- Record realised R next to the R the rules predicted.
- At the end of the block, count deviations before you judge the strategy.
- Compare live and tested results per setup and per hour, then change one thing.
Frequently asked questions
Start with at least one tick each way plus your real commission. After fifty live trades, compare that assumption against your own fills and correct it.
It catches rule errors and misses the money. Queue position and hesitation only appear with real fills. Treat paper as step one and small live size as step two.
Split the backtest by hour and compare like with like. If the tested edge was concentrated at the open, wider spreads and faster fills there are the first suspects.
Put the live half in TraderLog
Your real fills come in from Schwab, IBKR or a CSV. The stats page gives win rate, average winner, average loser, profit factor and expectancy on what actually filled. That is the number to hold your simulated results against.
Every trading morning at 8:40 ET our model draws the SPY, QQQ and IWM zones it expects to matter, on a TradingView chart, before the open. Where price reaches one, it has turned 74 percent of the time. Free, no account. See today's map