tech
Google's Gemini Omni turns images, audio, and text into video — and that's just the start
Google's Gemini Omni is a new multimodal model that reasons across text, images, audio, and video to generate and edit videos through simple conversation — starting with Omni Flash.

TL;DR
- Gemini Omni is Google's new family of multimodal AI models designed to reason across text, images, audio, and video.
- The model can generate high-quality videos from combined inputs and allows text-based photo editing.
- The first iteration, Gemini Omni Flash, is available now, focusing on consumer use cases like creating personalized videos and editing vacation footage.
- A more advanced Omni Pro model is planned for later release.
- To prevent deepfakes, avatar creation requires user verification, and all generated videos will have Google's SynthID digital watermark.
- Gemini Omni aims to simplify video creation for consumers and will also be available via API for enterprise and creative professionals.