All articles
Trading·2 July 2026·7 min read

I Trust My 6-Month Backtest More Than a 10-Year One

丂卩ㄖㄖҜㄚ丂卩ㄖㄖҜㄚResearch notes

I Trust My 6-Month Backtest More Than a 10-Year One

More history isn't always better — but not for the reason the hot-take crowd thinks. The recent window measures the size of your edge; the long window tells you when it's dying.

The "more history is always better" rule is wrong — but not for the reason the hot-take crowd thinks. Every backtesting guide tells you the same thing: more data is better. A ten-year backtest with three thousand trades sounds far more trustworthy than a six-month one with eighty. Bigger sample, tighter statistics, more regimes covered. Who could argue?

Me.

But I'm going to do something most wouldn't, which is argue against myself halfway through. The honest version of this is more useful than the clean version. So here's the claim, and then the full working.

Why a long backtest feels safe

Start by being fair to the orthodoxy, because it's half right.

A longer backtest gives you more trades. More trades means your measured win rate, your profit factor, your average winner are all estimated with more confidence. The error bars shrink. Eighty trades might show a profit factor of 2.0 by luck; three thousand trades showing 1.4 is far harder to fake. Statistical significance is real, and a six-month sample genuinely is a noisier estimate than a ten-year one.

If markets were stationary, this argument would be the end of it, and you should stop reading and go run the longest backtest you can.

The one assumption holding it all up

But that entire edifice rests on a single load-bearing assumption that nobody says out loud — that the market which generated 2016's data is the same market generating today's. That the data is all drawn from one stable distribution. Statisticians call it stationarity, and a ten-year backtest is only meaningful if it holds.

Right now, it doesn't. We are in a non-stationary stretch, and you can see it directly. This year does not behave like the eighteen months before it. When the data-generating process changes underneath you, old data isn't just less relevant. It's actively misleading, because it's describing a market that no longer exists.

And here is the bit that the "more history" crowd misses.

A ten-year backtest doesn't give you a precise measurement of today's edge. It gives you a precise measurement of the average edge across ten years of different markets. You get a beautifully tight confidence interval around a number that may describe no market that is actually in front of you. Precision about an average of regimes is not the same as accuracy about the current one. A long backtest buys down variance by buying in bias, and nobody tells you that's the trade you're making.

The trade I'd have binned if I trusted the long number

Concrete example, from my own accounts.

I run an IB Fade in the New York session. Over the most recent six months, in the current high-volatility regime, it runs at a profit factor around 2.0. Over the full twelve months, that same strategy drops to roughly 1.0 to 1.17. Marginal. Borderline discard.

So which number is true? If I'd deferred to the longer window the way every backtesting guide tells me to, I'd have looked at a profit factor barely above breakeven and binned a strategy that is strongly profitable in the market I'm actually trading today. The twelve-month average wasn't more honest. It was more diluted — it blended the current regime, where the edge is excellent, with an earlier regime where it wasn't, and handed me the mush in the middle. The longer window made a live edge look dead.

That's the case for recency. Now the case against myself.

Why "more informative" is not "more reliable"

Here is the trap.

Recency being more informative is not the same as recent being more reliable. Those are two different axes, and the entire risk lives in the gap between them. My six-month New York sample is more relevant than the twelve-month, yes. It's also half the trades, which makes it a noisier estimate of whatever the current edge truly is. The recency argument and the sample-size argument point in opposite directions, and both are valid at the same time.

So the honest position was never "trust the short backtest." It's this: the short window tells you the right thing, but with wider error bars. So you weight it more heavily, and you hold it more loosely. You don't confuse "this is the relevant regime" with "this number is precise." It is the relevant regime, measured imprecisely. Treating a noisy but relevant number as if it were exact is just a different way to blow up.

The time the long window saved me

And to prove I'm not just cheerleading for recent data, here's a case where the long backtest earned its keep, on my own data.

I had an M2K (Micro E-mini Russell 2000) configuration that looked excellent over six months. Strong profit factor, clean curve. The six-month window alone said deploy. But the twelve-month run revealed a weak stretch through the previous autumn and winter that the recent half simply couldn't show me. That older data wasn't valuable as a truth to regress toward. It was valuable because it was a long enough window to answer a different question — is this recent strength durable, or is it fading?

That reframes what long data is actually for. Its job is not to validate your edge. Its job is to tell you whether your recent edge is holding or decaying. You literally need the longer window to run the recency test, because you cannot see a taper with only the recent half. The two windows answer two different questions, and the error, made by both camps, is asking one window to do the other's job.

The actual rule

So I don't trust the six-month over the ten-year because short is better. I trust it because the two windows have different jobs, and most people give them the wrong ones.

The recent window measures the size of your edge in the regime you're trading now. The long window measures its trajectory — whether that edge is strengthening, stable, or quietly dying. The "trust the ten-year average" crowd uses the long window for sizing, and ends up trading a diluted number from a market that's gone. The naive "just use recent data" crowd throws the long window away, and loses the only tool that can warn them their edge is fading.

The read I actually use is simple. If the recent window is stronger than the full period, the edge is holding: healthy. If the recent window is weaker than the full period, that's a taper: the warning sign. My deployable M2K config runs better over six months than over twelve. That's not me cherry-picking the flattering number. That's the specific signal that the edge is still alive.

Rolling profit factor versus the twelve-month average — the recent regime versus the earlier one

The obligation nobody mentions

There is one more thing worth adding — recency weighting comes with a duty that the people who advocate it usually skip.

If you accept that recent data is more relevant because the regime can change, you have also accepted that your current edge has an expiry date you cannot predict. The same logic cuts both ways. So weighting recency is not permission to find a hot recent strategy and ride it forever. It obliges you to keep watching the long window for the taper, and to have a rule for switching the strategy off when the regime that feeds it turns. Recency without a kill rule isn't an edge. It's just being early to your own blow-up.

That's the part that separates this from recklessness. The contrarian-sounding headline is true, but only if you carry the obligation that comes with it.

The sacred cow, gently slaughtered

More history isn't better. More history is more precise, about a market that may no longer exist. In a stationary world, precision and accuracy are the same thing and the long backtest wins.

In the non-stationary world we're actually in, they come apart, and I would rather have a noisier measurement of the market I'm trading today than a flawless measurement of one that's already gone.

Use the long window. Just use it for the right job: not to tell you whether your edge is real, but to tell you when it's about to die.

Trade the truth behind the lesson

Every base rate in this piece lives in the Hit Rates library, free to read. Or connect your broker and see which of them your own trading actually survives.

Browse the library Start free
Keep reading
[ tradestar ]Evidence for every serious trader. Broker sync, prop tracking, and edge stats in one workspace.
© 2026 Tradestar. Trading involves risk of loss.Historical stats, not predictions. Trade your own plan.