The Next Big Leap for AI Is Into the Real World
Engineers and investors are flocking to "world models," also known as "large action models," in the hope of sparking a revolution in robotics similar to the one ChatGPT triggered in writing and programming.

Large language models of artificial intelligence today may excel at office work, but according to people in the field, to perform tasks in the physical world, large action models, also known as "world models," are required.
According to the method, which is intended to train artificial intelligence to operate robots, the models are trained using video games and simulations, rather than literature, code, and images. Recently, the technology has reached a point where the models are capable of navigating in three-dimensional space and even operating objects independently through a simple series of instructions.
The technology is still in its infancy, according to Moritz Baier-Lentz, who invests in several companies in the field. "If you compare it to large language models, we are roughly at the GPT-2 stage," he said (the model that OpenAI launched in 2019).
Although world models are still in a relatively early stage, some of the most prominent names in the technology industry are already rushing into the field. The bet is that a fundamentally new architecture will be able to bring artificial intelligence to places that systems based on today's large language models are unable to reach. And the pioneers of the field hope that one day it will be possible to merge the two approaches.
How "world models" work
The goal
To embed artificial intelligence in robots that operate in the real world, while reducing errors and the possibility of harming people.
Training
With the help of video games. The model records every frame in the game alongside the actions performed by the users. Thus, the artificial intelligence can link actions to their results.
The limitation
The robot must be four-legged, a wheeled vehicle, or a drone, and it must be controllable via a game controller, mouse, or keyboard.
From words to worlds
For about half a billion years, animals have gradually developed brains capable of creating a model of the world around them and using this mental map to plan their next action. Language, on the other hand, has existed for only about a hundred thousand years, and it is a kind of shortcut that humans created to process and convey ideas, both large and small.
"Text is just a partial representation of the real world, where information is lost," said Kent Rollins, Chief Product Officer at General Intuition and former head of the Fortnite ecosystem at Epic Games. "The world existed long before we had text, so when we use text to describe it, we necessarily miss essential aspects of it."
Robotics researchers have known this for a long time. The control systems of more sophisticated robots today rely on simulations based on the laws of physics to simulate the physical world. These are also world models, but they are carefully written in code and adapted for very specific purposes. A system that works for one type of robot does not necessarily fit another.
Startups currently developing world models aim to create a control system that will be as diverse in operating robots as today's large language models are diverse in creating text, from sonnets to software.
Training based on video games
A New York startup called General Intuition, which develops world models, is currently completing a funding round at a valuation of more than $6 billion, which will make it the most expensive AI lab of its kind. The basis for training its model is video games: the model records every frame in the game alongside the actions performed by the users, from every button press to the movement of the joystick. Unlike robots trained using video only, in this method, the artificial intelligence can link user actions to their results in virtual worlds, and thus a large action model is created.
After a slight adjustment, the model can operate a robot in the real world as if it were a character in a video game. However, there is a limitation: the robot must be four-legged, a wheeled vehicle, or a drone, of the type that usually transmits video to the user in real-time and can be controlled via a game controller, mouse, or keyboard. These features are quite common among widely used robots, but are not suitable for most humanoid robots that walk on two legs.
The General Intuition model is trained on millions of hours during which humans played video games. The data was collected from Medal.tv, a sister service of the company and a platform for recording and sharing game clips.
In the offices of General Intuition in New York, Geneva, and their surroundings, the company showcases a robotic dog. While regular artificial intelligence for operating such a robot might require training on a massive scale, this model only needs a few minutes of adjustment. The system needs to know where it is, what type of body it is operating in, and what its goal is, and from there it is ready to go, according to Baier-Lentz, who invests in General Intuition.
Previous demonstrations of world models, such as the Genie models presented by Google DeepMind, focused on creating new worlds in which robots could be trained. The new generation is already capable of perceiving the world around it and deciding what to do next. Jack Parker-Holder, who previously led Google's activities in the field of world models, was one of the founders of the London-based Emulate. The young lab is in talks to raise more than $500 million from investors, and its technical team includes six former Google employees, according to documents reviewed by The Wall Street Journal.
Language + Action = ?
Other companies developing world models rarely reveal details about the AI systems they are building, but their acquisitions and the publications of their engineers provide some clues.
At the head of the startup World Labs is Prof. Fei-Fei Li of Stanford University, who previously worked at Google and is known as the "godmother of AI." The company recently acquired the robotics company Sanix. At the head of AMI Labs is Yann LeCun, former Chief AI Scientist at Meta. It seems that the company is moving in a broader direction and is examining several architectures that go beyond traditional large language models.
While General Intuition and other startups emphasize their capabilities in the field of robotics, investors and potential business partners raise another question: how can world models improve the capabilities of existing language-based models?
For large language models to be able to process images and sound, they must be trained directly on this content and convert the input into tokens, as they do with words. The attempt to process a three-dimensional environment in this way proved to be extremely inefficient, and the result is models that are too slow to operate a robot in most situations.
In the real world "you can't just miss something roughly"
But world models also have their own problems. It was relatively easy to advance ChatGPT from one generation to the next precisely because at the beginning of its journey it dealt only with text. As large language models grow and are fed more and more data, they become larger and smarter, but they continue to make mistakes, according to George Konidaris, a professor of robotics at Brown University and co-founder of Realtime Robotics, which develops systems for industrial robots.
When it comes to robots, there are far fewer situations where we would be willing to accept hallucinations and other errors.
"The real world is very complex, and it has a lot of edge cases and hard boundaries, and you can't just miss something roughly," said Konidaris. "If your robot hits something while it's moving, everything changes."
Some believe that improving the efficiency of large language models might be enough to turn them into the artificial intelligence that will operate robots and other 3D applications. But there is also the possibility of combining the two approaches: large language models could be assisted by world models, just as they are currently assisted by other software to perform tasks.
Meanwhile, it is easy to notice that neither General Intuition nor Emulate are located in the San Francisco Bay Area, which was fascinated by the possibility that AI based on large language models would lead to "superintelligence."
When I asked the CEO of General Intuition, Pim de Witte, if superintelligence is the goal of his company, he was reserved. "I am in New York because I want to stay away from all this cult-like behavior and just focus on a scientific evaluation of the capabilities of models in robots," he said. "There is no need to make it something bigger than it is."





