NVIDIA Calls Cosmos 3 the World’s First Fully Open Omnimodel, Giving Robots and Autonomous Vehicles a Powerful Physics-Grounded Brain

TL;DR AI
2 min readKey summary
NVIDIA unveiled Cosmos 3 at GTC Taipei, calling it an open omnimodel for understanding and generating text, images, video, ambient sound, and actions.
The system pairs reasoning and generation transformers to better model physical interactions and produce more grounded outputs.
NVIDIA says Cosmos 3 can function as a vision-language model, a simulated world model, and a foundation for building other world models.
Cosmos 3 Super and Cosmos 3 Nano are available now, while Cosmos 3 Edge is planned for real-time edge inference.
The goal is to improve robotics, autonomous vehicles, and other vision systems with stronger physical-world understanding for reasoning and action.



