AI Is Learning How the Physical World Works

AI-powered humanoid robot navigating obstacles in a robotics testing environment

New “world models” are being designed to predict motion, understand space and anticipate what happens next, capabilities that could move artificial intelligence from screens into robots, vehicles and other physical systems.

THE UNIVERSAL RECORD

Sourced reporting. No opinions.

Brad Socha | August 31, 2026 | 5:35 AM EST

Artificial intelligence has become remarkably capable at working with language, images and software. A harder challenge is now moving to the centre of AI research: teaching machines enough about the physical world to predict what will happen when objects move, collide, fall or are manipulated, and eventually to act safely within those environments.

The approach is commonly associated with “world models,” systems designed to build internal representations of an environment and predict how it may change. Rather than merely identifying a cup on a table, for example, a useful physical model might help determine where the cup is in three-dimensional space, how it can be grasped and what is likely to happen if a robot pushes instead of lifts it.

Progress from Google DeepMind, Meta and Nvidia suggests this is becoming an important frontier of AI development. But demonstrations of physical reasoning should not be mistaken for machines possessing a complete understanding of physics. Independent research continues to expose significant weaknesses when models encounter unfamiliar situations or must make accurate predictions over longer periods. 

AI Moves Beyond Words

One route to physical understanding is video. Unlike individual photographs, video contains information about movement, cause and effect, object permanence and how environments change over time.

Meta’s V-JEPA 2, introduced in 2025, was pretrained using more than one million hours of video and images. Instead of attempting to reproduce every pixel of a future frame, its predictive architecture learns abstract representations of what is happening and tries to anticipate how those representations will change. Meta subsequently trained the system with 62 hours of robot interaction data and demonstrated it planning actions including reaching, grasping and moving objects in environments it had not encountered during robot training. 

Google DeepMind is pursuing related goals through both simulated worlds and robotics. Its Genie 3 world model generates interactive environments from text descriptions and can maintain navigable scenes in real time at 720p for several minutes. DeepMind describes world simulation as a way for agents to learn how environments evolve and how their actions affect them. The company also acknowledges limitations: generated environments are not perfectly accurate representations of real locations, and continuous interaction remains limited in duration. 

The connection between prediction and physical action becomes clearer in robotics.

In July 2026, DeepMind introduced Gemini Robotics 2, a family of models intended to combine perception and reasoning with robot control. The company’s demonstrations include humanoid robots bending, reaching and manipulating objects, robots performing delicate tasks with hands and grippers, and multiple machines coordinating on shared tasks. Its embodied-reasoning model can interpret visual surroundings and plan sequences of actions lasting several minutes. 

Those results remain controlled demonstrations and company evaluations rather than evidence that robots can reliably handle the enormous variety of situations encountered in everyday human environments. DeepMind itself continues to describe the technology as progress toward general-purpose physical AI, with some models remaining in private preview or restricted testing. 

Building Models of What Happens Next

Nvidia is taking another approach through Cosmos, a family of foundation models aimed specifically at what the company calls “physical AI.”

Cosmos models can use combinations of video, images, text and other information to reason about scenes, generate simulated environments and predict possible actions. In May 2026, Nvidia introduced Cosmos 3, combining visual reasoning, world generation and action prediction in a single architecture. The company is positioning the technology for robotics, autonomous vehicles and systems that must anticipate changing physical conditions. 

Simulation could address one of robotics’ fundamental problems: obtaining enough training experience.

Language models can learn from enormous quantities of digitized text. Robots require information about forces, movement, geometry, object interactions and consequences. Collecting millions of hours of physical robot experience is expensive and slow, while some dangerous or unusual situations cannot safely be reproduced repeatedly.

A sufficiently accurate world model could generate synthetic scenarios in which an AI system practices before encountering comparable situations in reality. An autonomous machine could theoretically evaluate several possible actions internally, similar in principle to considering what might happen before acting.

But generating a convincing video and accurately predicting reality are different achievements.

WorldBench, a research benchmark developed to test physical prediction, found that evaluated world models became less accurate as their prediction horizon increased. The researchers also reported substantial difficulty with object permanence and found that vision-language models performed poorly on physics-reasoning tests. The benchmark illustrates why photorealistic output alone cannot establish that a system has learned reliable physical laws. 

That distinction becomes critical when an error can have physical consequences. A chatbot producing an incorrect sentence is inconvenient; a robot incorrectly predicting the trajectory of an object, vehicle or person could create a safety hazard. Developers therefore need conventional engineering safeguards, sensors and control systems alongside learned AI models.

The potential nevertheless extends far beyond humanoid robots. Better physical reasoning could support warehouse automation, industrial equipment, autonomous transportation, assistive devices and machines operating in environments too hazardous for people. Models capable of learning partly through observation could also reduce the need to program every possible situation individually.

The deeper significance is a change in what researchers are asking AI to learn. Large language models demonstrated that machines could extract powerful patterns from enormous collections of human-generated information. World models are testing whether similar learning methods can capture enough of space, motion, time and cause-and-effect to make useful predictions about reality.

Current systems provide evidence of progress, not proof that the problem has been solved. The physical world is considerably less forgiving than a digital one. For AI to move reliably from generating answers to taking actions, predicting what happens next may prove just as important as understanding what humans say.

Sources:

Google DeepMind — Gemini Robotics 2
https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/ 

Google DeepMind — Gemini Robotics 2
https://deepmind.google/models/gemini-robotics/ 

Google DeepMind — Gemini Robotics ER 2 Model Card
https://deepmind.google/models/model-cards/gemini-robotics-er-2/ 

Google DeepMind — Genie 3
https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/ 

Meta AI — Introducing V-JEPA 2
https://ai.meta.com/blog/v-jepa-2-world-model-benchmarks/ 

Meta AI Research — V-JEPA 2
https://ai.meta.com/research/publications/v-jepa-2-self-supervised-video-models-enable-understanding-prediction-and-planning/ 

NVIDIA — Cosmos 3
https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-3-the-open-frontier-foundation-model-for-physical-ai 

WorldBench — How Close Are World Models to the Physical World?
https://world-bench.github.io/ 


About the Author
Brad Socha is the founder of The Universal Record, focused on sourced, factual global reporting. Coverage includes international news, geopolitics, technology, and major developments.


Discover more from The Universal Record

Subscribe to get the latest posts sent to your email.

Get the Universal Record App

Read verified global news anywhere.
Free on iPhone and Android.

Official Apple App Store badge displaying the Apple logo and the text “Download on the App Store” on a black background.
Official Google Play badge displaying the Google Play logo and the text “Get It on Google Play” on a black background.

Discover more from The Universal Record

Subscribe now to keep reading and get access to the full archive.

Continue reading