Do AIs Make Good Traders, and Do They Make Good Traders Better?

nassim nicholas taleb

Two years ago, we put that conjecture to the test. We called it “The Crystal Ball Challenge.” We staked 120 finance-trained adults with $50 each and handed them the front page of the Wall Street Journal – one day before publication, with any mention of market moves blacked out. In effect, we gave them what every trader dreams of: a working crystal ball.

They could go long or short the S&P 500 and 30-year Treasury bonds, with leverage if desired. For example, shown Wednesday’s front page – reporting on Tuesday’s events – they placed their trades at Monday’s close and were closed out at Tuesday’s close, once the news had played out in the market. Each player got 15 trading opportunities, one front page per year from 2008 to 2022.

The result: Our paid players didn’t do too well! On average, they broke even and about 1/6 went bust. They weren’t great at inferring market direction from the crystal ball, but they were particularly bad at position-sizing. We also invited six very successful macro-traders to play the game and they did pretty well – more on that shortly.1

The Crystal Ball game has been hugely popular since we made it freely available to play on our website. About 60,000 people have given it a whirl, and they fared substantially worse on average than our paid players – not surprising, given they were staked with $1 million in play money rather than real greenbacks.

Many of our readers have asked: Why should humans have all the fun? So we invited four leading LLMs – Claude, ChatGPT, Gemini, and Grok – to try their hand at our game. We even went a step further and updated the game so you can compete against the AIs mano-a-machina, no tears. Everyone trades simultaneously – sealed bid – and the AIs see the same headlines you do, reasoning from first principles without knowing the market returns. There’s also a leaderboard where players can see their own and the AIs performance. Want to give it a try before reading on? Play the Crystal Ball Game.

How’d the AIs Do?

Below are the results for the four AIs over ten rounds of play, with a starting wealth of $1 million. In terms of average ending wealth, Claude does best, followed by ChatGPT. Grok just about breaks even, while Gemini on average loses a considerable amount of starting wealth.

chart

The AIs Took Way Too Much Risk

Whether the AIs made money, as Claude and ChatGPT did, isn’t the end of the story. In order to assess how well they played, we need to assess their performance relative to the objective we provided before they played the game. Our instructions to the AIs were:

Please play the game as if you are a typical middle-aged, wealthy investor in the United States. The starting bankroll represents 100% of your financial wealth, you have pension income covering your subsistence needs, and there are no taxes.2

A high-functioning LLM could reasonably have inferred from this that their objective should not be to maximize expected gain, but rather their risk-adjusted profit, assuming a typical investor’s degree of risk-aversion.

We believe all the AIs were too aggressive in their position-sizing. One perspective on this is to note that the US stock market has moved by over 5% on 23 days and by over 9% on seven days since the year 2000. Given average position sizing in stocks of 7x to 12x across the AIs, we think they were taking too much risk of a catastrophic loss of capital, given none of them had (or could reasonably expect to have) super high hit ratios.