Artificial intelligence developers are increasingly pivoting away from large language models (LLMs) toward “world models,” a new class of AI designed to understand the laws of physics and spatial relationships rather than just text. This shift represents a move toward physical AI, where systems are trained to interact with the environment, navigate space, and predict the consequences of actions in the real world, according to researchers in the field.
While generative AI chatbots like ChatGPT and Claude have dominated the industry, some computer scientists argue that LLM research has reached a point of diminishing returns. By focusing on how objects respond to force or how light interacts with surfaces, proponents believe world models could provide the “brain” necessary for the next generation of robotics and autonomous systems. This transition is backed by significant private investment from firms like Kindred Ventures, which is currently funding startups focused on these specialized architectures.
Moving Beyond Text-Based Intelligence
The core limitation of current language models is their reliance on the statistical structure of text, which does not translate to physical competence. Martial Hebert, dean of the School of Computer Science at Carnegie Mellon University, notes that while LLMs excel at predicting the next word in a sequence, they lack an understanding of the physical world required to perform basic tasks, such as picking up a coffee mug. According to Hebert, true intelligence requires an awareness of geometry, dynamics, and the physical interaction between objects, which he describes as far more complex than linguistic prediction.

This pursuit of “physical AI” is viewed by many as the natural evolution of traditional robotics. Rather than programming robots for specific, rigid tasks, researchers aim to build systems with a general model of movement and balance. This allows the machine to adapt to its environment in real time, similar to how a human nervous system adjusts to physical changes like an injury or a shift in terrain. The goal is to create AI that functions with an intuitive sense of how the world works, rather than relying on pre-programmed instructions for every possible scenario.
The Taxonomy of World Models
The term “world model” has become a central, yet often ambiguous, concept in the AI sector. Fei-Fei Li, a pioneer in the field and founder of the San Francisco-based startup World Labs, has attempted to categorize these systems to clarify their disparate functions. In a recent essay, Li outlined three distinct categories for these models, noting that they are often conflated despite having different objectives:

- Renderers: These models prioritize visual fidelity, creating highly realistic virtual environments. While visually impressive, they often lack the physical grounding necessary to teach robots how to interact with the real world.
- Simulators: These systems create virtual training grounds that mirror the physical structure of reality. They serve as environments where AI can learn the laws of physics safely before being deployed to hardware.
- Planners: These are designed to predict the outcomes of actions in unstructured environments. Li emphasizes that a robot capable of planning is a robot capable of effective work, making this the most sought-after capability in the industry.
Yann LeCun, a prominent AI researcher and former chief AI scientist at Meta, echoes the importance of predictive capability. On a recent episode of the Unsupervised Learning podcast, LeCun explained that he defines a world model as a system that enables an AI agent to predict the consequences of its own actions, allowing it to navigate and operate within an environment independently.
Investment and Future Applications
The push toward these models is attracting substantial capital, even as the broader tech industry continues to pour trillions of dollars into traditional LLM development. Steve Jang, co-founder and managing partner at Kindred Ventures, suggests that the future of the field will not be defined by a single, massive model, but rather by a diverse ecosystem of architectures tailored to specific physical and spatial challenges. His firm is currently backing companies like Overworld, which is building adaptable video game environments, and Causal Labs, which focuses on AI models for weather prediction.
Louis Castricato, who left his doctoral studies at Brown University to found Overworld, argues that the industry has largely exhausted the potential for fundamental LLM research. His company aims to develop AI that can navigate complex, detailed environments where objects respond dynamically to user interaction. By optimizing for interaction, these startups hope to create virtual and physical spaces that feel responsive and grounded in reality, rather than static or pre-rendered.
As the industry races to develop the first truly capable planning agents, the focus remains on closing the gap between digital intelligence and physical execution. While LLMs continue to transform office-based work and creative industries, the next phase of AI development appears centered on the ability to “read the room” and interact with the physical world in a meaningful, autonomous way. Stakeholders are expected to watch for further academic publications and startup product launches as researchers work to refine these models and move them from simulation into real-world applications.
Worth a look