Chatbots needed large language models to learn from text at scale. Robots, self-driving cars and industrial machines need something different. Nvidia just released the most advanced version of it yet.
Nvidia launched Cosmos 3, an open world foundation model for physical AI trained on 20 trillion tokens of multimodal data — including nearly a billion images, 400 million real and synthetic videos, audio and action data from humans and robots.
The model targets machines that need to understand the physical world before acting in it. At the launch, Nvidia founder and CEO Jensen Huang said that “the big bang of physical AI is just around the corner thanks to breakthroughs in multimodal reasoning language, vision and world models.”
The distinction matters for anyone building or deploying physical AI. An LLM learns from text. A world foundation model learns from physical environments: how objects move, collide, fall and interact over time. What makes Cosmos different from a video generator, Axios noted, is the action data: Cosmos 3 doesn’t just generate realistic scenes, it predicts what a robot or vehicle should do next within them.
The Data Problem Physical AI Has to Solve
Training a chatbot on internet-scale text is expensive but tractable. Training a robot or autonomous vehicle on real-world physical experience is neither. A robot learning to handle objects needs millions of interaction examples. An autonomous vehicle needs exposure to rare and dangerous scenarios, fog, pedestrian edge cases and unexpected road conditions, that can’t be safely or cheaply collected at scale on public roads.
World foundation models solve this by generating synthetic training data that reflects real physics. Instead of driving a test fleet for years, an autonomous vehicle developer can run millions of simulated scenarios in days.
Cosmos 3 compresses what previously required separate models into one. MarkTechPost noted that earlier Cosmos releases split physical reasoning, world generation and action generation across separate systems. Cosmos 3 unifies them in a single open model, reducing training cycles from months to days according to Nvidia.
Who Is Building on It
The Cosmos platform already has a working ecosystem. Agile Robots, Doosan Robotics, LG Electronics and Samsung are building robotics applications on it. Li Auto is using it for autonomous vehicle development. Waabi, Wayve and Foretellix use Nvidia’s Cosmos models to simulate traffic scenarios, weather conditions and pedestrian behaviors without physical trials, AI Multiple reported.
In January, Mercedes-Benz launched its first premium robotaxi service on the Uber network using Nvidia’s physical AI stack built on Cosmos, 4D Pipeline noted.
Nvidia also launched the Cosmos Coalition alongside Cosmos 3, a global collaboration including Agile Robots, Black Forest Labs, Runway and Skild AI to advance next-generation open world foundation models.
AMI Labs, founded by Yann LeCun and seeking a $3 billion valuation, and World Labs, founded by Fei-Fei Li and in talks at a $5 billion valuation, are among the frontrunners building competing world model platforms. Google DeepMind’s Genie 3 is positioned as a rival, with differentiation already emerging: Genie 3 excels at generating novel environments from text, while Cosmos maintains strict physical consistency for industrial applications.
For all PYMNTS AI coverage, subscribe to the daily AI Newsletter.