DEV Community

Cover image for I Asked Claude to Invest $10,000. Then I Changed the Starting Date
Kevin Meneses González
Kevin Meneses González

Posted on Originally published at Medium

I Asked Claude to Invest $10,000. Then I Changed the Starting Date

The portfolio was never the biggest risk.

The date was.

Three weeks ago I asked Claude to build a $10,000 allocation for a friend who'd had cash sitting in a savings account for eight months. Five positions, each one backed by real numbers pulled from a live data feed. The piece did well. The replies were mostly the same question, worded fifty different ways:

"OK, but would it have worked?"

My friend asked a sharper version. She runs financial models for a living, so she went straight to the weak spot:

"What if I'd done this in January 2022, right before everything fell?"

Fair question. So I ran a market timing backtest. Same portfolio, same $10,000, but instead of one start date I tested 105 of them: every month from January 2015 to September 2023, each one held for three years.

Then I compared every single run against the S&P 500, a classic 60/40 fund, and the portfolio most people actually want to own right now: the Magnificent 7.

If you're:

  • sitting on cash and waiting for "a better moment" to invest,
  • building a backtesting tool or an AI agent that reasons over market data,
  • or reading performance charts that conveniently start at the bottom,

this is worth your ten minutes.

Every backtest you've seen has one start date

And the person who wrote it picked that date.

That's the quiet problem with most "I invested $X in Y" content. One entry point. One exit point. One number at the end that looks like evidence.

It isn't. It's an anecdote with a chart attached.

Shift the start by six months and the same strategy can look brilliant or broken. Most authors don't do it on purpose. They pick a round date like January 1st, run the numbers once, and publish whatever comes out.

Here's the uncomfortable part: when you run the first article's portfolio from January 2019, $10,000 becomes $19,135 in three years. Start it in January 2022 and the exact same allocation ends at $12,382.

Same stocks. Same weights. Same rebalancing rule.

$6,753 of difference, and none of it came from the portfolio.

A single start date is an anecdote. A distribution is evidence.

The fix is called rolling-window backtesting.

Instead of asking "what did this portfolio return?", you ask "what did this portfolio return for every possible entry date?" You get a distribution instead of a number. And a distribution tells you three things a single backtest can't:

  • how bad your luck can realistically get,
  • how often the strategy beats a simple benchmark,
  • whether the edge is in the returns or somewhere else.

That third point turned out to be the whole story.

The setup: rolling 3-year windows on real adjusted prices

Four portfolios, $10,000 each, rebalanced back to target weights every 12 months:

Portfolio Holdings
Claude's portfolio (from the first article) 35% VOO, 20% MSFT, 20% JNJ, 15% GLD, 10% SHY
Magnificent 7 (popular stocks) Equal weight AAPL, MSFT, NVDA, AMZN, GOOGL, META, TSLA
S&P 500 benchmark 100% VOO
60/40 benchmark 100% AOR (iShares Core 60/40 allocation ETF)

Start dates: the first trading day of every month from January 2015 to September 2023. That's 105 entry points, and every window closes by September 2026.

The data comes from EODHD's end-of-day API, using the adjusted_close field. That matters more than it sounds. Adjusted prices fold in dividends and splits, so JNJ's 2–3% yield and NVDA's and TSLA's stock splits are reflected correctly. Using raw closing prices would quietly understate the dividend payers and wreck the split stocks.

One API key covered all twelve tickers back to 2014, on the same endpoint and in the same JSON shape.

→ If you want to reproduce this, EODHD's free tier is enough to start.

The code: a rolling-window backtest in Python

About 50 lines. One function downloads prices, one simulates a portfolio from any start date, and a loop runs all 105 windows.

pip install requests pandas
Enter fullscreen mode Exit fullscreen mode
import os
import requests
import pandas as pd

API_TOKEN = os.environ["EODHD_API_KEY"]

PORTFOLIOS = {
    "Claude": {"VOO": .35, "MSFT": .20, "JNJ": .20, "GLD": .15, "SHY": .10},
    "Mag7":   {t: 1/7 for t in ["AAPL", "MSFT", "NVDA", "AMZN", "GOOGL", "META", "TSLA"]},
    "VOO":    {"VOO": 1.0},
    "60/40":  {"AOR": 1.0},
}
YEARS = 3

def get_prices(ticker):
    url = f"https://eodhd.com/api/eod/{ticker}.US"
    params = {"api_token": API_TOKEN, "fmt": "json", "from": "2014-01-01"}
    df = pd.DataFrame(requests.get(url, params=params).json())
    return df.set_index(pd.to_datetime(df["date"]))["adjusted_close"]

tickers = {t for w in PORTFOLIOS.values() for t in w}
prices = pd.DataFrame({t: get_prices(t) for t in tickers}).dropna()

def simulate(weights, start, capital=10_000):
    """Buy on `start`, rebalance every 12 months, hold YEARS years."""
    window = prices.loc[start:start + pd.DateOffset(years=YEARS), list(weights)]
    cuts = [window.index.searchsorted(start + pd.DateOffset(years=y)) for y in range(YEARS)]
    cuts.append(len(window) - 1)
    value, path = capital, []
    for a, b in zip(cuts, cuts[1:]):
        seg = window.iloc[a:b + 1]
        shares = pd.Series(weights) * value / seg.iloc[0]
        curve = (seg * shares).sum(axis=1)
        path.append(curve if not path else curve.iloc[1:])
        value = curve.iloc[-1]
    return pd.concat(path)

rows = []
for start in pd.date_range("2015-01-01", "2023-09-01", freq="MS"):
    for name, weights in PORTFOLIOS.items():
        curve = simulate(weights, start)
        rows.append({
            "start": start.date(),
            "portfolio": name,
            "final": round(curve.iloc[-1]),
            "max_drawdown": round((curve / curve.cummax() - 1).min() * 100, 1),
        })

results = pd.DataFrame(rows)
final = results.pivot(index="start", columns="portfolio", values="final")
print(final.describe(percentiles=[.1, .5, .9]).T.round(0))
print("Claude beats VOO in", f"{(final['Claude'] > final['VOO']).mean():.0%}", "of start dates")
Enter fullscreen mode Exit fullscreen mode

Output:

           count     mean     std      min      10%      50%      90%      max
portfolio
60/40      105.0  12326.0  1261.0  10387.0  10779.0  12187.0  14190.0  15312.0
Claude     105.0  15055.0  1662.0  12382.0  12936.0  14879.0  17207.0  19135.0
Mag7       105.0  27870.0  8738.0  16639.0  18374.0  25574.0  39156.0  59977.0
VOO        105.0  14920.0  1764.0  11098.0  13109.0  14473.0  17526.0  20004.0
Claude beats VOO in 54% of start dates
Enter fullscreen mode Exit fullscreen mode

One detail in simulate() is easy to get wrong. If the anniversary falls on a weekend, the rebalance has to happen on the next trading day, and the portfolio value has to be measured on that same day. Measure value on Friday and buy at Monday's prices, and you silently drop a day of returns at every rebalance. My first version had exactly that bug. It moved some results by more than $1,000.

From here you can build:

  • a "what if I'd started earlier" widget for any watchlist,
  • a risk report that shows the worst 10% of entry dates instead of the average,
  • an agent that answers "should I wait to invest?" with a distribution instead of an opinion.

Five start dates with names

Before the full distribution, here are the dates people actually ask about. Value of $10,000 after three years, with the worst peak-to-trough drop along the way:

Entry date Claude's portfolio S&P 500 (VOO) 60/40 (AOR) Magnificent 7
Jan 2018 (before the Q4 selloff) $15,743 / −22.8% $14,736 / −34.0% $12,422 / −22.9% $38,183 / −34.9%
Feb 2020 (pre-COVID peak) $12,926 / −22.9% $13,308 / −34.0% $11,250 / −22.9% $19,987 / −50.7%
Apr 2020 (after the COVID crash) $15,371 / −18.3% $17,448 / −24.5% $13,112 / −21.7% $29,128 / −49.8%
Jan 2022 (market peak) $12,382 / −17.2% $12,839 / −24.5% $10,789 / −21.5% $18,006 / −49.0%
Oct 2022 (near the bear market low) $17,839 / −10.3% $19,043 / −18.7% $15,312 / −9.8% $36,297 / −31.2%

Read the Claude column top to bottom. The same portfolio ranges from $12,382 to $17,839 depending only on the month you clicked "buy."

Now read across any row. Claude vs. the S&P 500 is usually a few hundred dollars apart.

The month mattered far more than the portfolio did.

What 105 start dates actually show

1. Timing luck is about 10× bigger than portfolio choice

Across all 105 windows, the median gap between Claude's portfolio and the S&P 500 was $637.

The gap between Claude's best and worst start date was $6,753.

That ratio is the headline. If you spend three weeks optimizing what to buy and zero minutes thinking about how you'll handle a bad entry, you've put your effort in the wrong place.

2. On returns alone, Claude's portfolio was a coin flip against the S&P 500

It finished ahead of VOO in 54% of start dates. Median ending value: $14,879 vs. $14,473.

That's not an edge. That's noise.

If the only goal was "beat the S&P 500," the honest conclusion would be: just buy VOO.

3. The edge was in the drawdowns

This is where the two portfolios stopped looking alike.

Across all 105 start dates Claude's portfolio S&P 500 (VOO) 60/40 (AOR) Magnificent 7
Median ending value $14,879 $14,473 $12,187 $25,574
Worst ending value $12,382 $11,098 $10,387 $16,639
Median max drawdown −17.7% −24.5% −21.7% −36.0%
Worst max drawdown −23.5% −34.0% −22.9% −50.7%
Avg. months below $10,000 2.8 3.9 6.3 3.6
Beats VOO (share of start dates) 54% n/a 0% 100%

Claude's portfolio delivered roughly the same return as the S&P 500 with drawdowns about a third smaller. Its worst entry still ended higher than VOO's worst entry.

That's what the gold, JNJ and short-term treasuries were doing. They didn't add return. They made the bad start dates less bad.

And that matters more than it looks on a spreadsheet. A −34% drop is the point where real people sell. A −23% drop is the point where they complain and hold.

The 60/40 fund made an interesting control. In this specific decade, it never beat the S&P 500 over a 3-year window, and its typical drawdown (−21.7%) was deeper than Claude's (−17.7%). The 2022 bond crash hit it from both sides.

The popular-stocks portfolio: huge returns, huge hindsight

The Magnificent 7 portfolio beat the S&P 500 in 100% of start dates. Median ending value: $25,574. Best case, starting January 2019: $59,977.

Before you sell everything and buy seven tickers, look at two numbers.

The range. Same seven stocks, same weights. Starting in January 2019 gave you $59,977. Starting in February 2021 gave you $16,639. That's a $43,339 difference driven only by the entry month. The more concentrated the portfolio, the more the calendar decides your outcome.

The drawdowns. The median window included a −36% drop. The worst included −50.7%. Half your money, gone on paper, at least once. Most people who say they'd hold through that haven't had to.

Then there's the problem no backtest fixes on its own: look-ahead bias.

Nobody called these seven companies "the Magnificent 7" in 2015. That name dates from 2023. I picked them in 2026 because they won. Running them back to 2015 is like betting on a race after reading the results.

The same caveat applies, more mildly, to Claude's portfolio. MSFT and gold were picked with today's information too. A rolling backtest removes start-date bias. It does not remove selection bias.

Lump sum vs. dollar-cost averaging: the obvious fix that mostly isn't

If the entry date is the risk, the natural reaction is to spread it out. Instead of investing $10,000 on day one, invest $1,667 a month for six months.

I ran that too, on Claude's portfolio, with uninvested cash earning zero:

Lump sum 6-month DCA
Median ending value $14,879 $14,697
Worst ending value $12,382 $12,030
Best ending value $19,135 $17,803
Wins (share of 105 start dates) 78% 22%
Feb 2020 start (pre-COVID) $12,926 $13,338
Jan 2022 start (market peak) $12,382 $12,763

DCA won in only 22% of start dates. It helped when you started right at a peak: about $400 better in February 2020 and in January 2022. Averaged over the 10 worst lump-sum start dates, though, the two came out essentially tied ($12,555 vs. $12,564). And it cost real money in the windows where markets simply went up.

This matches what Vanguard and others have published for decades: lump sum wins about two-thirds of the time because markets rise more often than they fall. In this decade it won even more often.

DCA isn't a return strategy. It's a regret strategy. If investing $10,000 the day before a crash would make you sell everything, the six-month version is worth its cost. If it wouldn't, it isn't.

The honest caveats

  • This was a good decade. Not one of the 420 windows (105 dates × 4 portfolios) lost money over three years. Run the same test starting in 2000 or 2007 and that changes fast. EODHD has 30+ years of history if you want to stress-test it.
  • No taxes, no trading costs. Annual rebalancing in a taxable account has a cost this backtest ignores.
  • Overlapping windows. 105 windows over ~11 years aren't 105 independent experiments. Neighboring months share most of their data.
  • Three years is short. It's the horizon where timing hurts most. Over 10–15 years, entry dates matter less. They don't stop mattering.

Key takeaways

  • Your entry date moved the result about 10× more than your portfolio choice did. $6,753 between best and worst start, vs. a $637 median gap to the S&P 500.
  • A diversified portfolio's real job is shrinking the worst case, not beating the index. Claude's allocation matched VOO on returns and cut the median drawdown from −24.5% to −17.7%.
  • Concentration amplifies timing. The Magnificent 7 turned a $10,000 timing gap into a $43,000 one, and its 100% win rate rests on hindsight.

FAQs

Does market timing matter for long-term investing?

Yes, more than most people expect over 3–5 year horizons. In a rolling backtest of 105 start dates between 2015 and 2023 using EODHD price data, the same $10,000 portfolio ended anywhere from $12,382 to $19,135 after three years, depending only on the entry month. The effect shrinks over longer horizons, but it doesn't disappear.

How much does the start date change a portfolio's returns?

In this test, entry timing changed 3-year results by up to $6,753 on a $10,000 diversified portfolio and up to $43,339 on a concentrated Magnificent 7 portfolio. The difference between the diversified portfolio and the S&P 500 was much smaller: a median of $637.

Is lump sum investing better than dollar-cost averaging?

Usually, yes. In this backtest, investing $10,000 at once beat a 6-month dollar-cost averaging plan in 78% of start dates. DCA only helped when the lump sum landed right before a crash, such as February 2020 or January 2022. It reduces regret more than it improves returns.

Did the Magnificent 7 beat the S&P 500?

Yes. From 2015 to 2026, an equal-weight Magnificent 7 portfolio beat the S&P 500 in every 3-year window tested, with a median ending value of $25,574 vs. $14,473 per $10,000. But the result has strong look-ahead bias: the group was selected for having already won, and it suffered drawdowns of up to −50.7%.

Is a 60/40 portfolio still worth it?

Between 2015 and 2026, a 60/40 fund (AOR) never beat the S&P 500 over a 3-year window, and the 2022 bond selloff left its median drawdown (−21.7%) deeper than that of a stock-heavy portfolio with gold and short-term treasuries (−17.7%). Its value is in lower volatility in most periods, not higher returns.

How do I backtest a portfolio with different start dates in Python?

Download adjusted daily prices (for example from EODHD's /api/eod/ endpoint), write a function that simulates the portfolio from a given start date, then loop it over every monthly start date and collect the ending value and maximum drawdown. Use adjusted prices so dividends and splits are included, and rebalance on the first trading day on or after each anniversary.

Can Claude backtest an investment portfolio?

Claude can write and run the backtesting code, and with a market data connection such as EODHD's MCP server or REST API, it can work with real historical prices instead of remembered ones. The output is only as reliable as the data and assumptions behind it, so ask it to show the exact numbers and start dates it used.

The real question was never what to buy

My friend read the results and said something I hadn't expected:

"So I've been waiting eight months for the right moment, and the moment matters more than anything I was researching."

Yes. Which is exactly why waiting doesn't solve it. Nobody knows which month will be the January 2019 and which will be the January 2022.

What you can choose is a portfolio whose worst start date you could live with.

Run the backtest yourself with EODHD's API. The free tier covers every endpoint used here, including adjusted end-of-day prices with 30+ years of history.

This is not financial advice, and I'm not a financial advisor. Past performance, across 105 start dates or one, doesn't guarantee future results.

This article ran 420 backtests before writing a single claim. That's how I write for fintech and API companies: real calls, real data, numbers your readers can reproduce.

→ See more of my work at kevinmeneses.com


Looking for technical content for your company? I can help — LinkedIn · kevinmenesesgonzalez@gmail.com

Top comments (0)