Claude beat the S&P 500 index, but maybe it's just luck

27 million dollars are already copying the model's trades, and investors are sure they are on the right track. An AI expert explains why a 14-stock portfolio and a five-month period are not proof of genius. And why AI is not yet a substitute for a human advisor.

N12Author: Efrat Nomberg-Junger
Source
Claude beat the S&P 500 index, but maybe it's just luck
Photo: N12 / אילוסטרציה: traffic_analyzer, getty images

The artificial intelligence model Claude was given a simple task: to manage a $50,000 investment portfolio and beat the S&P 500 index. After about five months, the experiment's operators reported that the portfolio yielded a return of about 19%, compared to about 12% for the index. The result attracted attention online, and investors who followed the experiment began to copy the model's trades. According to reports, the volume of investments copying its actions has already reached about $27 million.

Language models like Claude do not understand money. They scan vast amounts of text, thousands of financial reports, news headlines, and financial analyses in seconds, trying to find correlations between textual information and stock performance. This is how they identify trends that an average investor might miss due to data overload. In the next stage, the model translates its statistical forecasts into buy and sell decisions to maximize a goal defined for it in advance. In this experiment, the goal was one: to beat the S&P 500 index.

Dr. Mike Ehrlichson is not quick to be impressed. Ehrlichson, head of AI at a tech startup and manager of a professional LinkedIn community of over 60,000 members, says that five months is not enough to know if someone is an excellent investor or just lucky.

The portfolio included only 14 stocks, and about a quarter of it was held in bonds, he explains. The portion invested in stocks was much more concentrated and risky than the S&P 500 index. When you take more risk, sometimes you earn more, and sometimes you lose much more.

Ehrlichson also points to survivorship bias. There are quite a few similar experiments with AI models running online today, but the one that beat the index is the one that gets the headlines, he says. The explanations the model gives in hindsight are also not necessarily impressive. A language model knows how to write a convincing explanation for almost any decision, even if it stemmed from luck.

What the machine doesn't know

In fact, the model and human advice have quite a bit in common: both analyze companies, read reports, and try to predict the market's direction, and both build a diversified investment mix across different assets, such as stocks and bonds, to manage risk.

But the differences are sharp. Artificial intelligence has no psychology: it doesn't get stressed when the market falls, doesn't act out of fear or euphoria, and can wait patiently for opportunities. These are exactly the things that cause many investors to make mistakes. A human advisor, no matter how experienced, is exposed to emotional biases and pressure from clients who have panicked.

On the other hand, a human advisor sits in front of the client and understands the full picture: when they will need the money, whether they are saving for retirement or an apartment, what their personal risk tolerance is. The model in the experiment functioned as a machine that sees only numbers and returns, without context for the investor's life. Ehrlichson adds that a human advisor also has a basic understanding of how the world works, while language models are built to generate text and therefore might provide convincing explanations for wrong decisions, or those that simply stemmed from luck.

And yet, Ehrlichson does not rule out the idea. "I would let the model manage a small and pre-defined amount, one that if it disappears, my life won't change," he says. But only with clear rules: no concentrating everything in one stock, no leverage, and always with a person who oversees and approves exceptional decisions.

What does bother him is the scale. "Not the fact that they gave Claude a $50,000 portfolio to manage," he says. "What bothers me is that already today there are tens of millions of dollars that are copying its trades almost automatically." Here, people might think that artificial intelligence is a substitute for judgment.

The public is still not convinced

The general public is also showing caution. A Gallup poll found that only about 18% of Americans have used artificial intelligence for financial decision-making, and only a minority express high trust in it. Another survey by the investment firm Janus Henderson yielded a similar finding: investors use AI for research and information gathering, but their comfort level drops when artificial intelligence actually recommends how to invest. Most respondents noted that they would like to see a human in the picture even as technology improves.

"People might get confused and think that artificial intelligence is a substitute for judgment," says Ehrlichson. "In practice, at least in the coming years, it can serve as an excellent researcher and fast data analyst, and perhaps also a tool that maintains discipline. But not a substitute for the person who makes the decisions."

Related News