tech

Google's Gemini Omni turns images, audio, and text into video — and that's just the start

Google's Gemini Omni is a new multimodal model that reasons across text, images, audio, and video to generate and edit videos through simple conversation — starting with Omni Flash.

Google's Gemini Omni turns images, audio, and text into video — and that's just the start

TL;DR

  • Gemini Omni is Google's new family of multimodal AI models designed to reason across text, images, audio, and video.
  • The model can generate high-quality videos from combined inputs and allows text-based photo editing.
  • The first iteration, Gemini Omni Flash, is available now, focusing on consumer use cases like creating personalized videos and editing vacation footage.
  • A more advanced Omni Pro model is planned for later release.
  • To prevent deepfakes, avatar creation requires user verification, and all generated videos will have Google's SynthID digital watermark.
  • Gemini Omni aims to simplify video creation for consumers and will also be available via API for enterprise and creative professionals.