Google DeepMind announces multimodal generative model “Gemini Omni,” enabling video generation and editing through natural-language dialogue and reasoning

TL;DR AI
2 min readKey summary
Google DeepMind unveiled Gemini Omni, a new multimodal model family for generating and editing video from diverse inputs.
The model supports reference-based editing, natural language control, and AI avatars that can use a user's voice.
Google also said generated videos will include SynthID watermarking for transparency.
The update aims to make consistent, text-driven video creation more practical for creators and everyday users.



