Google unveils 'Gemini Omni'... generates videos from text, images, and audio

TL;DR AI
2 min readKey summary
Google unveiled Gemini Omni, a multimodal AI that can understand and generate video from text, images, audio, and video together.
The initial Gemini Omni Flash model focuses on creating 10-second videos and supports avatar-style generation plus SynthID watermarking.
Google is rolling it out first in the Gemini app, YouTube Shorts, and creator tools.
The company also plans to offer an API soon and is preparing higher-end models for advertising and video production.


