Switch language한국어
Back to the list

Google DeepMind announces multimodal generative model “Gemini Omni,” enabling video generation and editing through natural-language dialogue and reasoning

TL;DR AI

Key summary

2 min read
  1. Google DeepMind unveiled Gemini Omni, a new multimodal model family for generating and editing video from diverse inputs.

  2. The model supports reference-based editing, natural language control, and AI avatars that can use a user's voice.

  3. Google also said generated videos will include SynthID watermarking for transparency.

  4. The update aims to make consistent, text-driven video creation more practical for creators and everyday users.

Read the original