What Are World Models - Background

What Are World Models

How AI Learns How the World Works

by Csaba Fekszi

You run a world model every day, like predicting where a ball will land, showing how these models directly impact your operations and decision-making, which should inspire confidence in their practical value for your business.

A transformer model trained on New York taxi rides learned to give turn-by-turn directions through Manhattan with near-perfect accuracy. Its researchers at MIT, Harvard, and Cornell then reconstructed the street map the model had formed, revealing imaginary streets, flyovers, and roads with impossible orientations. As MIT News reports, closing just 1% of the streets reduced the model’s accuracy from nearly 100% to 67%, showing how small changes can significantly affect AI’s understanding of real environments.

The model knew the routes and misread the city. Closing that gap is the aim of world models: AI systems that learn how an environment behaves and what happens when someone acts in it.

Google DeepMind, Meta, and NVIDIA have released such models, demonstrating their growing importance. Two start-ups founded by leading researchers, Yann LeCun’s AMI Labs and Fei-Fei Li’s World Labs, raised more than $2 billion between them in February and March 2026, according to TechCrunch. This article explains how world models learn, where they already pay off in real-world scenarios, and what your company can prepare now to leverage this technology.

What a world model does

You run a world model every day. Before you catch a ball, you predict where it will land. Meta’s AI researchers describe this inner model as a simulator: it lets you try an action in your head and pick the one most likely to reach your goal, illustrating how these models mirror human intuition and decision-making in daily operations.

An AI world model does the same job in software. In Google DeepMind’s definition, it simulates aspects of the world so an agent can foresee how its surroundings will develop and how its actions will change them.

The difference from a large language model lies in what gets predicted. A language model predicts the next word. A world model predicts the next state of a scene, a machine, or a road, often in response to an action.

If you want to see what happens inside the language-model side of that comparison, we explain it in The Secret Life of LLMs: How AI Actually Works →

Accurate predictions, wrong picture of the world

Good predictions and correct understanding are two separate measurements. The taxi experiment by Keyon Vafa and colleagues, presented at NeurIPS 2024, showed it for maps. A follow-up study from the same group, presented at ICML 2025, showed it for physics.

What Are World Models - Ábra 1 (EN)
Figure 1. The model learned the routes and misread the city.

A model trained on planetary orbits predicted trajectories accurately, even for solar systems outside its training data. Asked for the forces behind those orbits, it produced a nonsensical law of gravitation, and a different one depending on the slice of data it worked from.

The authors draw a historical parallel. Kepler could predict where the planets would be; Newton explained why, and his laws carried over to new problems. The models reached Kepler’s level and stopped there.

Video models show a similar gap. On IntPhys 2, a benchmark Meta released alongside V-JEPA 2, a system must spot which of two videos breaks the laws of physics. Humans are close to perfect, and current video models score at or near chance.

A system can pass every test built from typical days and still fail on the first unusual one.

Three ways machines learn how the world works

Research labs explore three broad routes (imagine, watch, generate) to teach machines about the world, presenting opportunities your company can explore to stay ahead in AI innovation.

Learning by imagining

The agent builds a model of its environment, then practices inside it. The approach goes back to a 2018 paper in which David Ha and Jürgen Schmidhuber trained an agent entirely inside a dream generated by its own world model.

DreamerV3, published in Nature, turned the idea into a general-purpose algorithm. With a single configuration, it beat specialized methods across more than 150 tasks, and it became the first algorithm to collect diamonds in Minecraft from scratch, learning purely from its own experience.

Its successor, Dreamer 4, achieved the same feat from recorded gameplay alone, using a hundredth of the data an earlier OpenAI agent needed. Its authors point to robotics, where trial and error on real equipment is slow and unsafe.

Learning by watching

Video is a rich and plentiful record of how objects move and interact. Meta trained V-JEPA 2, a world model with 1.2 billion parameters, on more than 1 million hours of video and 1 million images. It predicts in a compressed representation of the scene and leaves out pixel-level detail, an approach Yann LeCun proposed in 2022 and now pursues at his own start-up, AMI Labs.

After this pre-training, the model needed only 62 hours of robot data to plan. In environments that were new to the model, robot arms picked up and placed unfamiliar objects with success rates of 65% to 80%.

Learning by generating

The third route produces the world itself. From a text prompt, Google DeepMind’s Genie 3 generates an interactive environment you can move through in real time, at 24 frames per second and 720p resolution. Consistency is the hard part. Each frame builds on the last, so small errors accumulate, and Genie 3 worlds stay largely consistent for a few minutes.

NVIDIA’s Cosmos models, trained on 20 million hours of video covering human activity, industrial settings, robotics, and driving, generate synthetic footage for training robots and autonomous vehicles.

All three routes learn from raw observations linked to actions, encouraging your company to start with internal data to develop effective world models for operational improvements.

Where world models already earn their keep

Autonomous driving is the clearest case today. The Waymo Driver has covered nearly 200 million fully autonomous miles on public roads and billions more in simulation. In February 2026, Waymo introduced a simulator built on Genie 3 that generates rare events its fleet can hardly record at scale, from a tornado to an elephant on the road.

Waymo notes that most driving simulators are trained only on their own fleet’s road data, which limits them to situations the fleet has already encountered.

Robotics and industrial vision follow the same pattern. NVIDIA’s Cosmos 3 release of May 2026 added tools for generating images of product defects, and NVIDIA claims the model cuts training and evaluation cycles from months to days. That is a vendor claim, worth testing on your own data.

The nearest reference point for Hungarian readers sits in Debrecen. BMW describes its plant there as the group’s first factory planned and validated completely virtually, in a digital twin built on NVIDIA Omniverse. A digital twin is engineered from design data and known physical rules; a world model learns its rules from observation. The two are starting to meet, since Cosmos can generate synthetic video from 3D scenes built in Omniverse.

Beyond these niches, the market is young. At a world models panel in September 2026, reported by TechCrunch, AMI Labs said it was still in a research and building phase and kept its product plans private. World Labs’ Marble, probably the most developed product in the field, mostly demonstrates media creation, game environments, and visual effects.

Gartner’s Hype Cycle for Physical AI, 2026 concludes that hype and risk in the field currently outweigh business outcomes. AMI Labs’ own chief executive, Alexandre LeBrun, anticipated the buzz when the company announced its funding round:

“In six months, every company will call itself a world model to raise funding.”

Alexandre LeBrun, CEO of AMI Labs, speaking to TechCrunch in March 2026

What your company can prepare now

For most mid-sized companies, world models will likely arrive inside products: a robot, an inspection camera, a planning tool, or a simulation service. The preparation that pays off sits in your own data. Gartner names limited real-world data as a core constraint on physical AI and advises AI leaders to streamline their sensor and OT/IoT platforms.

That preparation starts with knowing what data you have, who owns it, and whether it can be trusted across systems — the same foundation we examine in Why Do AI Projects Fail Without Data Governance →

Every approach above learns from the same three things: what the situation looked like, what was done, and what happened next. Picture sensor readings in one database, work orders in the maintenance system, and quality results in a spreadsheet, each with its own clock. Before any model can connect an intervention to its effect, those clocks have to agree.

What Are World Models - Ábra 2 (EN)
Figure 2. A world model can learn only what your data records. Start by making your own operation readable.

Five checks before you rely on an AI’s picture of your operation

1. The detour test
Test every AI system on changed conditions, such as a new product variant, a closed line, or a new supplier, alongside typical days.

2. State, action, outcome
Check whether your data links what the situation looked like, what was done, and what happened next, with timestamps that match across systems.

3. Rare events.
List the failures that cost the most and happen the least. That is where simulation and synthetic data earn their cost, and where a model trained on ordinary days is weakest.

4. The vendor’s world
Ask what data a model learned from, how long its simulations stay consistent, and how it performed outside its training conditions.

5. The human checkpoint
Decide where a person approves a plan before it reaches real equipment.

Clear answers to these five checks prepare you for world-model-based tools as they mature. The same data work also strengthens the AI you can deploy today, from visual quality inspection to maintenance forecasting.

From routes to the whole map

The taxi model learned the routes and misread the city. World models aim to learn the city itself: its streets, its rules, and what happens when a road closes. The research is moving quickly, and the commercial market is still taking shape.

For your company, the most useful step is closer to home. A world model can learn only what your data records. So start by making your own operation readable: link what happened, what was done, and what followed, and test every AI system on the detours as well as the daily routes.

The map of your own operation is the part you can start drawing today.

Where does this technology fit in your operation?

Understanding the technology is a useful first step. The question that matters for your company is where it can create business value, and the answer depends on how your own operation runs. In a free 30-minute consultation, we help you clarify whether the technology described here belongs in your operation and, if so, which first step makes sense.

Sources

  • BMW Group. (2023, March 21). BMW Group at NVIDIA GTC: Virtual production under way in future plant Debrecen. Source of the example of BMW’s Debrecen plant being planned and validated completely virtually before physical production began. Read article →

  • Gartner. (2026, March 19). Leverage Data and World Models to Advance Physical AI. Source of findings on limited real-world data as a constraint for physical AI and the role of sensor and OT/IoT platforms. Read article →

  • Gartner. (2026, July 16). Hype Cycle for Physical AI, 2026. Source of Gartner’s assessment that hype and risk around physical AI currently outweigh demonstrated business outcomes. Read article →

  • Google DeepMind. (2025, August 5). Genie 3: A New Frontier for World Models. Source of the world-model definition, Genie 3’s 24 fps 720p generation, short-term consistency, and accumulating-error limitations. Read article →

  • Ha, D., & Schmidhuber, J. (2018). World Models. arXiv. Source of the early world-model concept in which an agent can be trained entirely inside a simulated “dream” generated by its own learned model. Read article →

  • Hafner, D., Pasukonis, J., Ba, J., & Lillicrap, T. (2025, April). Mastering Diverse Control Tasks Through World Models. Nature. Source of DreamerV3 results across more than 150 tasks using one configuration, including obtaining diamonds in Minecraft without human data. Read article →

  • Hafner, D., Yan, W., & Lillicrap, T. (2025, September). Training Agents Inside of Scalable World Models. arXiv. Source of Dreamer 4, its use of offline data, its reported data efficiency relative to OpenAI’s VPT agent, and the robotics rationale for scalable world models. Read article →

  • Meta AI. (2025, June 11). Introducing the V-JEPA 2 World Model and New Benchmarks for Physical Reasoning. Source of the internal-simulator framing, V-JEPA 2’s scale and training data, robotic pick-and-place results, and IntPhys 2 performance. Read article →

  • MIT News. (2024, November 5). Despite Its Impressive Output, Generative AI Doesn’t Have a Coherent Understanding of the World. Source of the navigation experiment showing near-perfect route performance despite incoherent internal street maps, including the drop in accuracy when only 1% of streets were closed. Read article →

  • NVIDIA. (2025, January 6). NVIDIA Makes Cosmos World Foundation Models Openly Available to Physical AI Developer Community. Source of details on Cosmos training data, including 20 million hours of video and synthetic video generated from Omniverse scenes. Read article →

  • NVIDIA. (2026, May 31). NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI. Source of Cosmos 3 use cases such as defect-image generation and NVIDIA’s claim of reducing development timelines from months to days. Read article →

  • TechCrunch. (2026, March 9). Yann LeCun’s AMI Labs Raises $1.03B to Build World Models. Source of funding information on AMI Labs and World Labs and comments on the emerging world-model market. Read article →

  • TechCrunch. (2026, September 18). World Model Companies Are Keeping a Lot of Secrets. Source of the current development status of major world-model companies, including AMI Labs’ research phase and Marble as one of the most developed products. Read article →

  • Vafa, K., Chen, J., Rambachan, A., Kleinberg, J., & Mullainathan, S. (2024). Evaluating the World Model Implicit in a Generative Model. NeurIPS 2024. Source of the New York taxi-route experiments, including impossible street orientations and flyover-like internal representations. Read article →

  • Vafa, K., Chang, J., Rambachan, A., & Mullainathan, S. (2025). What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models. ICML 2025. Source of the orbital-trajectory experiments, the emergence of a nonsensical force law, and the comparison with Keplerian and Newtonian structure. Read article →

  • Waymo. (2026, February 6). The Waymo World Model: A New Frontier for Autonomous Driving Simulation. Source of Waymo’s world-model simulator, its use of Genie 3, rare-event generation, nearly 200 million autonomous miles of experience, and the limitations of relying only on fleet-collected data. Read article →

Picture of Csaba Fekszi

Csaba Fekszi

Csaba Fekszi is an IT expert with more than two decades of experience in data engineering, system architecture, and AI-driven process optimization. His work focuses on designing scalable solutions that deliver measurable business value.

Related posts

Artificial Intelligence Explained - Background
AI in Business
Why ChatGPT Is Not the Same as AI
What Is an AI Agent - Background
AI Building Blocks
And How Is It More Than a Chatbot
Common Pitfalls to Avoid in an AI Pilot - Background
AI in Business
Why So Many AI Projects Stall — and How to Finally Move Beyond the Pilot Phase​
From Video to Insight - Background
AI Business Use Case
Where Video Analysis Creates Operational Value
The Key Steps to a Successful AI Implementation - Background
AI in Business
Turning Ambition into Real, Scalable Results

Are you sure AI is the right next step?

We help uncover the real opportunities, limitations, and realistic next steps.

Comments are closed.