General
World Models vs. Large Language Models: Two Visions of Machine Intelligence
Published March 19, 2026
There is a quiet civil war in AI research, and the stakes are not just technical. On one side: large language models, the architecture behind ChatGPT, Claude, and Gemini, which have consumed hundreds of billions of dollars in compute and crowned a handful of companies as the gatekeepers of machine intelligence. On the other: world models, a fundamentally different approach that aims to learn how reality works — not by predicting the next word, but by building internal simulations of the world.
If world models succeed, they could break the equation that currently governs AI: more money equals better models. That would reshape who controls AI, who profits from it, and who gets displaced by it. If they fail, the existing oligopoly hardens further. Either way, this architectural divide will determine the distribution of power in the next era of artificial intelligence.
LLMs: A Brief Recap
Large language models are, at their core, autocomplete engines. Given a sequence of tokens — fragments of text — they predict the next one. Then the next. Then the next. Everything these systems produce, from legal briefs to poetry to code, emerges from this single operation repeated billions of times (for a fuller treatment, see How Language Models Actually Work).
The mechanism is powerful but carries specific constraints. LLMs learn from text. They scale by training on more text with more parameters on more GPUs. Their intelligence, such as it is, comes from statistical patterns extracted from written language — the compressed residue of human expression. They can produce a beautiful essay about a ball rolling down a hill. They cannot predict the ball's trajectory.
This matters because the scaling paradigm — more data, more compute, better performance — has made LLMs enormously expensive to build. As documented in The Economics of Training a Frontier Model, training runs now cost upward of $1 billion. That cost structure creates a natural oligopoly: only a handful of companies can afford to play (The Frontier Lab Oligopoly). The architecture, in other words, is not just a technical choice. It is an economic structure.
What World Models Are
A world model is an AI system that learns a compressed, abstract representation of how an environment works. Instead of predicting the next word in a sequence, it predicts the next state of a system — not pixel by pixel or token by token, but in a learned abstract space that captures the underlying dynamics.
Think of it this way. If you see a cup teetering on the edge of a table, you don't need to simulate every photon of light or every molecule of ceramic to predict what happens next. You have an internal model — built from years of interacting with the physical world — that tells you the cup will fall and break. World models aim to give machines the same kind of compressed, efficient understanding.
The most prominent proposal is Yann LeCun's Joint Embedding Predictive Architecture, or JEPA. LeCun, Meta's chief AI scientist and a Turing Award laureate, laid out the vision in his 2022 paper "A Path Towards Autonomous Machine Intelligence." The core argument: current AI systems, including LLMs, lack the ability to plan, to reason about cause and effect, or to learn efficiently from limited experience. JEPA addresses this by learning to predict representations of future states in an abstract embedding space, rather than reconstructing raw sensory data.
The distinction is subtle but consequential. A generative model trying to predict video, for instance, must account for every pixel in every frame — an enormous computational burden that includes vast amounts of irrelevant detail (the precise pattern of shadows, the exact shade of the sky). A JEPA-style model instead learns to predict high-level features: the object moved left, the ball accelerated, the door opened. It discards the noise and keeps the structure.
LeCun's framework also introduces the idea of a "world model" as a core module in an intelligent system, sitting alongside a perception module, an actor module, and a cost function — a modular architecture that looks more like cognitive science than the monolithic scaling of today's LLMs. As the cognitive scientist Kenneth Craik argued as early as 1943, internal models of reality are what allow organisms to anticipate events and act adaptively. LeCun is, in a sense, proposing to engineer what evolution stumbled into.
The Technical Difference That Matters
The split between LLMs and world models is not a matter of degree. It is a difference in kind.
LLMs generate sequences. They produce one token after another, with each token conditioned on what came before. Their architecture — the transformer — is optimized for this sequential generation across enormous text corpora. The quality of their outputs depends on the breadth and quality of their training data and the scale of the model.
World models simulate. They build internal representations of environments and predict how those environments evolve. They can learn from video, from sensor data, from physical interaction — not just text. Where LLMs require vast written corpora, world models can, in principle, learn from the kind of unstructured, continuous data that the physical world produces in abundance.
LLMs scale by getting bigger. The scaling laws that have governed AI progress since 2017 suggest that performance improves predictably with more parameters, more data, and more compute (Kaplan et al., 2020). World models aim for a different kind of efficiency — intelligence through abstraction rather than brute-force statistical coverage. A world model that truly captures the dynamics of rigid-body physics shouldn't need to see a billion examples of objects falling. It should learn the principle.
This is the technical claim. Whether it holds at scale is the open question.
Where Each Excels
Today, the division of labor is relatively clear.
LLMs dominate language. Chatbots, code generation, writing assistance, legal document drafting, customer service, translation, summarization — any task that can be framed as "produce the right sequence of words given a context" is LLM territory. These systems have reshaped white-collar knowledge work with startling speed (The White-Collar Displacement Wave).
World models target the physical world. Robotics, autonomous vehicles, industrial automation, drug discovery, climate modeling, surgical assistance — domains where understanding physics, spatial relationships, and cause-and-effect matters more than generating text. These are tasks where you need a system that can predict what will happen when a robotic arm applies force to a deformable object, or what trajectory a vehicle should follow through an intersection, or how a protein will fold under specific conditions.
The gap between these capabilities is revealing. An LLM can write a compelling explanation of orbital mechanics. It cannot simulate an orbit. It can describe in detail how a surgical procedure should unfold. It cannot guide a surgical robot through an unexpected complication. The difference is between knowing about the world through language and understanding the world through dynamics.
The Redistribution Angle
This is where the architectural debate stops being academic and starts being political.
Compute Economics
If world models achieve capable intelligence through efficient abstraction rather than scaling brute-force compute, the central equation of the current AI era breaks. The "more compute = better AI" paradigm is what justifies billion-dollar training runs, what makes NVIDIA the most valuable company on Earth, and what creates The Hardware Bottleneck that concentrates power in a handful of actors.
World models propose an alternative path: better representations instead of bigger models. If that path works, the compute barrier to building capable AI systems drops. Organizations that cannot afford $1 billion training runs might build effective AI with a fraction of the resources. The moat around the frontier labs gets shallower.
Who Controls What
LLM dominance concentrates power in companies that can afford massive text corpora, massive GPU clusters, and the engineering talent to orchestrate both. The bottleneck is capital.
World models could shift the bottleneck. But "shift" does not necessarily mean "democratize." The new bottleneck could be proprietary sensor data — the millions of hours of driving footage that Waymo has collected, the industrial process data locked inside manufacturing companies, the medical imaging archives controlled by hospital networks. It could be robotics platforms, where a handful of companies control the hardware through which world models interact with physical reality.
Redistribution is not guaranteed. The question is whether world models create new access points or new chokepoints. The answer will depend partly on whether the field develops open data-sharing norms — projects like Common Crawl democratized text data; no equivalent exists yet for robotics and sensor data.
Which Workers Are Affected
LLMs displace knowledge workers — writers, coders, analysts, customer service representatives, paralegals. These are predominantly white-collar, often college-educated, and concentrated in service economies.
World models, if they succeed, displace physical-world workers. Drivers, warehouse operators, manufacturing line workers, agricultural laborers, logistics coordinators. These are often blue-collar, less likely to hold degrees, and concentrated in different geographic and demographic populations.
The political dynamics are entirely different. White-collar displacement happens quietly — a team gets smaller, an entry-level role disappears, a freelancer finds fewer clients. Physical-world displacement is visceral and visible: a factory installs robots, a fleet goes autonomous, a warehouse stops hiring. The workers affected have different unions (or none), different political representatives, different capacity to retrain or relocate.
As Daron Acemoglu has argued, the distributional consequences of automation depend not on the technology itself but on the institutional structures that mediate its adoption — labor regulations, collective bargaining power, social safety nets, and the political choices societies make about who bears the cost (Acemoglu & Restrepo, 2019). A shift from LLM-driven to world-model-driven displacement would test an entirely different set of those structures.
AMI Labs as Test Case
In March 2025, Yann LeCun co-founded AMI Labs with a $1.03 billion funding round — the largest-ever seed investment for an AI company (AMI Labs: LeCun's Billion-Dollar Bet Against LLMs). The startup's explicit mission is to build world models that achieve machine intelligence through LeCun's JEPA-based approach rather than the autoregressive token prediction of LLMs.
AMI Labs is the highest-profile institutional bet that the world-model paradigm is not merely a research curiosity but a viable commercial alternative to the LLM stack. Whether it delivers will shape market structure: if world models prove capable at scale, they validate an alternative path to AI development that is potentially less capital-intensive. If they fail, the LLM paradigm hardens into orthodoxy, and the companies that dominate today — OpenAI, Google DeepMind, Anthropic, Meta — maintain their structural advantage.
The Skeptic's Case
It would be negligent not to say this plainly: world models are largely theoretical. JEPA has shown promising results in research settings — learning useful representations from video, outperforming some contrastive learning methods on specific benchmarks (Assran et al., 2023). But no world-model system has demonstrated the kind of broad, general capability that LLMs have achieved. The gap between a research prototype and a deployed system that works reliably in the real world is vast, and littered with the remains of ideas that were elegant in principle and intractable in practice.
Meanwhile, LLMs keep getting better. They are absorbing new modalities — vision, audio, video. The frontier labs are investing heavily in multimodal capabilities and in their own versions of world-model research. Google DeepMind's Genie and Veo projects, OpenAI's Sora video model, and Meta's own internal research all incorporate elements of world modeling within the broader transformer paradigm. The incumbents may simply absorb the world-model paradigm rather than be disrupted by it.
The scaling laws that have powered LLM progress may be plateauing (The Scaling Laws Plateau), but that does not automatically validate the alternative. It is possible that LLMs hit diminishing returns and world models fail to scale — leaving the field in a productive crisis rather than a clean paradigm shift.
What to Watch
The resolution of this debate will not come from theoretical arguments. It will come from engineering results: can world models learn general-purpose representations efficiently enough to compete with LLMs in real-world tasks? Can they do so at a cost that changes the economics of AI?
Watch AMI Labs for early signals. Watch whether autonomous driving and robotics companies adopt JEPA-style architectures or stick with transformer variants. Watch whether the compute cost of capable AI systems starts falling — not because hardware gets cheaper, but because the architecture gets smarter.
And watch who benefits. If world models work, the winners will not necessarily be the people who need redistribution most. New technologies have a persistent habit of enriching their creators before reaching the public. Policy choices about public investment, openness norms, and worker bargaining power could change that pattern — but so far, nothing in the world models research agenda suggests its proponents are asking the redistribution question. The question, as always, is not just whether the technology works — but who it works for.
Related
- How Language Models Actually Work
- Economics of Training a Frontier Model
- The Frontier Lab Oligopoly
- The Hardware Bottleneck
- The White-Collar Displacement Wave
- The Scaling Laws Plateau
- AMI Labs
- The Training Pipeline
- Physical AI: When Machines Learn to Touch the Real World
- Multimodal AI
Sources
- LeCun, Yann. "A Path Towards Autonomous Machine Intelligence." Meta AI, June 2022. OpenReview
- Assran, Mahmoud et al. "Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture." CVPR, 2023. arXiv
- Kaplan, Jared et al. "Scaling Laws for Neural Language Models." arXiv, 2020. arXiv
- Acemoglu, Daron and Pascual Restrepo. "Automation and New Tasks: How Technology Displaces and Reinstates Labor." Journal of Economic Perspectives 33, no. 2 (2019): 3–30. AEA
- Craik, Kenneth. The Nature of Explanation. Cambridge University Press, 1943.
- Ha, David and Jürgen Schmidhuber. "World Models." arXiv, 2018. arXiv
- Cottier, Ben et al. "The Rising Costs of Training Frontier AI Models." Epoch AI, 2024. Epoch AI