Daily Picks
Klement on Investing
Klement on InvestingJoachim Klement

The performance decay of LLM trading strategies

TL;DR

LLM trading strategies excel in backtests within their training data window but suffer sharp performance drops out-of-sample, exposing a serious overfitting problem.

In an experiment at the San Francisco Fed, ChatGPT's inflation forecasts were miserable. Once the predictions went beyond the training window, the model broke down. Training data works like a leaking pipe, sneaking answers into the test.

Change the scene, and 30% becomes 9%

Chinese researchers tested five LLM-based trading agents, all running on GPT-4o with a known training cutoff of October 2023. Backtested within the training window (Q2–Q3 2021), the models earned 30%–44%, comfortably beating the Nasdaq 100.

Then they asked the same models to trade out-of-sample (Q3–Q4 2024). The market was similar (the index returned about 13.5% in both periods), but the trading returns collapsed to 9%–22%. Deprived of their data crutch, the AI lost its edge.

Monte Carlo simulation to wean AI off data

The researchers' fix: let the LLM devise its own strategy, then test it not only in a backtest but also against counterfactual scenarios. In plain terms, it is good old Monte Carlo simulation with a few extra bells and whistles, creating a simulated environment the model has never seen.

Tested on five stocks and Bitcoin, it worked surprisingly well.

In a word: AI trading strategies ace the exam by memorizing the answer key, but flunk when the exam changes. Monte Carlo simulation is their 'experiential learning' course.

Read the original →
Share to

You might also read

5 articles worth reading every day

Curated from high-quality sources, with concise summaries and key takeaways.

Daily Picks

5 articles worth reading every day

Curated from high-quality sources, with concise summaries and key takeaways.

All posts from that day2026-08-27 · 16 in total