AgiBot WITA-Omni Full-Modal Model Tops DailyOmni Global Leaderboard: Beating Google Gemini, ByteDance Doubao, and Alibaba Qwen at Embodied Cross-Modal Understanding

TL;DR AI
2 min readKey summary
AgiBot’s WITA-Omni Preview scored 85.21 on the DailyOmni benchmark, ranking first overall.
It led six of eight metrics, including audio-visual alignment, temporal reasoning, and long-video understanding.
The model uses a Thinker-Talker-Actor design that runs reasoning, speech, and action generation in parallel to better synchronize a robot’s words, gestures, and expressions in real time.
The result underscores how embodied AI tuned for real-world interaction can outperform general-purpose multimodal models on robotics-relevant tasks.
