Google’s Gemini Omni turns images, audio, and text into video — and that’s just the start

TL;DR AI
2 min readKey summary
Google unveiled Gemini Omni, a multimodal model family that can generate and edit video, images, audio, and text from mixed inputs.
The first release, Gemini Omni Flash, can make 10-second videos in the Gemini app, YouTube Shorts, and Flow.
Google also added text-based photo editing, digital-avatar video creation with safeguards, and SynthID watermarking for synthetic media.
Longer videos and broader media-generation features are planned next, as Google pushes deeper into cross-modal AI.



