DEV Community

Pedro Groppo
Pedro Groppo

Posted on

Why Your Backtest Is Probably Lying to You (And How ATR Fixes Half the Problem)

If you've ever backtested a trading strategy in Python, there's a good chance you've made one of two mistakes without realizing it. Both are easy to make, both make your results look better than they are, and both have simple fixes.

Mistake #1: Fixed-percentage stop losses

The most common way to size a stop loss is a fixed distance from entry: "always exit 2% below entry." It's simple, but it ignores something important — how volatile the asset actually is right now.

  • On a calm asset, a 2% stop is often wider than necessary. You're risking more capital than the trade needs.
  • On a volatile asset, that same 2% can be too tight. Price hits it on ordinary noise, not because your thesis was wrong.

The fix: size your stop off the asset's own recent volatility instead of an arbitrary universal number. The standard tool for this is ATR (Average True Range) — the average range the asset moves, bar by bar, over a lookback period.

def calculate_atr(df, period=14):
    high, low, close = df["high"], df["low"], df["close"]
    prev_close = close.shift(1)

    true_range = pd.concat([
        high - low,
        (high - prev_close).abs(),
        (low - prev_close).abs(),
    ], axis=1).max(axis=1)

    return true_range.rolling(window=period).mean()
Enter fullscreen mode Exit fullscreen mode

Once you have ATR, your stop becomes a multiple of it (e.g., entry_price - 2 * atr), and your position size is calculated so that hitting the stop costs exactly the % of your account you decided to risk — no more, no surprise:

def calculate_position_size(account_size, risk_pct, entry_price, stop_loss):
    dollar_risk = account_size * risk_pct
    risk_per_unit = abs(entry_price - stop_loss)
    return dollar_risk / risk_per_unit
Enter fullscreen mode Exit fullscreen mode

Now a calm asset gets a tighter stop and a bigger position; a volatile one gets a wider stop and a smaller position — for the same dollar risk either way.

Mistake #2: Backtesting on your entire dataset

This one is sneakier. Say you're testing a moving-average crossover strategy. You try a 20-day MA, get a mediocre Sharpe ratio. You try 35 days, better. You try 42, even better. You lock in 42 days and report that Sharpe as "the strategy's performance."

The problem: you just spent your validation budget finding the parameters that fit the noise in that specific dataset. You didn't discover an edge — you found the specific numbers that happened to work on data you've now already seen. This is a classic form of overfitting, sometimes called data-mining bias, and it's the single most common reason DIY backtests look great and live strategies don't.

The fix is a discipline, not a formula: split your data chronologically before you tune anything.

def split_data(df, split_date, date_column="date"):
    df = df.copy()
    df[date_column] = pd.to_datetime(df[date_column])
    split = pd.to_datetime(split_date)

    in_sample = df[df[date_column] < split].reset_index(drop=True)
    out_of_sample = df[df[date_column] >= split].reset_index(drop=True)

    return in_sample, out_of_sample
Enter fullscreen mode Exit fullscreen mode
  • In-sample (older data): this is the only data you're allowed to look at while tuning parameters.
  • Out-of-sample (newer data): run the strategy here with parameters already frozen. Never adjust anything after seeing this result.

The question that actually matters: does the out-of-sample performance look roughly like the in-sample performance, or does it collapse? A collapse is the most reliable signal you have that the strategy was overfit — not that it "stopped working."

Putting it together

Neither fix is complicated in isolation. The hard part is discipline: actually doing the ATR-based sizing instead of a round number, and actually holding out real data instead of peeking at it "just this once." Most DIY backtesting mistakes come from skipping these two habits under time pressure, not from not knowing about them.

If you want a reference implementation, I open-sourced the ATR sizing piece here: github.com/pedrogroppo2-cell/atr-position-sizing — MIT licensed, no dependencies beyond pandas.

I also packaged the full flow (ATR sizing + the out-of-sample backtesting engine + a notebook walking through both) into a small kit, in case skipping the setup time is useful: Risk Management + Backtesting Kit.

Happy to answer questions about either piece in the comments.

Top comments (2)

Collapse
 
arhancanli profile image
Arhan Canli •

Mistake #2 is the one that bites hardest, and your 20 / 35 / 42 example is exactly the shape of it. The holdout is the right fix. It's also possible to put a number on how much the search inflated the in-sample Sharpe before you ever touch the holdout, which tells you whether there's anything worth testing.

The deflated Sharpe ratio (Bailey and López de Prado) does that from four inputs: the best Sharpe, how many variants you tried, how far their Sharpes spread, and how long the backtest is. Say the best window gives a Sharpe of 1.2 over two years of daily data, and the windows you tried had Sharpes spreading by about 0.5. Taken on its own, 1.2 looks like a 95.5% chance the edge is real. But if it was the best of 30 windows, pure luck would be expected to produce a best of about 1.04, and the probability falls to 59%. Best of 10 windows: 72%. So "I tried 30 and kept the best" eats most of the evidence on its own, before costs or regime change.

One more holdout trap: a holdout is only clean the first time you look at it. If the out-of-sample result disappoints and you go back to tune, it has quietly become in-sample, so it's worth writing down how many times you've checked it.

If useful, there's a free browser calculator for this at canlicapital.com/tools/deflated-sharpe (I built it; it's also available as an MCP server for Claude and Cursor).

Some comments may only be visible to logged-in visitors. Sign in to view all comments.