Skip to main content

Gemini Omni: DeepMind explains Google's new video model — from conversational edits to SynthID

Ilustrační obrázek
Google DeepMind's Gemini Omni is past the demo stage. The native multimodal model — which generates and edits video from any combination of text, image, audio, and video inputs — is now rolling out through Google's paid AI subscriptions, YouTube Shorts, YouTube Create, and a developer API. This week, DeepMind's engineers detailed how it works. And every single output carries Google's invisible SynthID watermark — a detail with real teeth under the EU AI Act.

When Gemini Omni was first shown at Google I/O in May, it was easy to file under "impressive demo." Three months later, the model is in production surfaces, and the engineering team — including Google DeepMind's Mohammad Babaeizadeh, Anish Nangia, and Sarah Xu — has published a technical breakdown of the architecture behind it, as The Tech Buzz reports. This is the point where a product stops being a promise and starts being something developers and creators have to make actual decisions about.

What Gemini Omni actually is

Gemini Omni is a native multimodal model, not a Frankenstein pipeline of separate image, audio, and video networks. The initial release, Gemini Omni Flash, takes text, images, audio, or video as input and produces video as output. It handles 9:16 vertical video natively, which sounds like a small detail until you remember that most AI video tools default to 16:9 and force you to crop everything for mobile feeds.

The bigger conceptual shift is conversational editing. Instead of generating a clip, watching it, and typing an entirely new prompt from scratch, you keep the same video and request changes conversationally — altering a motion, a character's action, or the scene itself. That is a fundamentally different workflow from the "type it, render it, pray" generation style, and it changes how much leverage a creator actually has over the result.

What DeepMind's engineers actually said

Babaeizadeh, Nangia, and Xu walked through three technical pillars. First, conversational editing: the model treats the video as a persistent object you can iterate on with natural language, rather than regenerating everything from zero. Second, physical dynamics: the model reasons about how objects move, collide, and interact — the exact area where most video generators still produce nonsense. Third, unified multimodal reasoning: one model handles all modalities internally, which is why the audio, the lip movements, and the visuals stay in sync instead of being stitched together from separate components.

The benchmark claim: multi-character scenes and Punjabi dialogue

Google reports benchmark results with superior text-to-video fidelity compared to competing models, including accurate multi-character scene logic. One showcased example — synchronized Punjabi spoken dialogue — deserves a pause. Generating visuals and lip-synced speech in a non-English language in a single pass is genuinely hard, and most video models do not even attempt it. Punjabi is not a European language, but the implication matters here: the model is not obviously English-centric. With 24 official EU languages, the question European users should ask is whether that multilingual dialogue holds up in Czech, Polish, or Finnish. Google has not shown that yet, and we won't take it on faith.

Where you can actually use it

Access is already wider than the usual staged rollout. Gemini Omni is available to paying subscribers across Google AI Plus, Pro, and Ultra; free users get it inside YouTube Shorts and the YouTube Create app; and external developers get API access for building their own pipelines. For European users, there is no separate EU add-on — it sits inside the same Google AI subscription tiers that already carry models like Gemini 3.7 Flash.

The YouTube Shorts integration is the distribution move to watch. Google is putting AI video generation directly into an app millions of Europeans already use daily, which is exactly how a feature becomes boring and ubiquitous. It also has a workflow logic to it: you do not export a vertical video from a web tool and re-import it — you edit in the same surface where Shorts are produced.

SynthID: the watermark EU compliance teams should like

Every video generated by Gemini Omni automatically receives Google's invisible SynthID watermark — built for AI attribution, not for plastering a visible badge across the frame. Here is the European angle: the EU AI Act's transparency obligations for AI-generated content are no longer a voluntary exercise. As of August 2026, national authorities and the EU AI Office are enforcing binding rules on transparency, governance, and risk mitigation under the AI Act and the 2026 AI Omnibus framework. A hidden watermark baked in at generation time is a practical compliance layer — much more robust than post-hoc detection tools that struggle with re-encoded video. For EU companies processing video through the API, that is one less thing to build. GDPR still applies to any personal data you feed into prompts, but the synthetic-content labeling problem is handled at the source.

Our take from the AI Arena side

Gemini Omni is a cloud-only model with closed weights — it is not something we will benchmark locally on our RTX 5060 Ti 16 GB rig in AI Arena. That is not a criticism: a model that renders lip-synced multilingual dialogue needs data-center scale, and local video generation remains a completely different performance class. But it does mean the "Google says" caveat applies until independent developers hammer the API. The benchmark claims are Google-reported, and our editorial instinct is the same as always: impressive demo — now let us see it in production. That production test starts now, in Shorts, Create, and the API.

What to do with it

For creators: try it inside YouTube Create if you are already producing Shorts — native vertical output removes the crop-and-reframe step entirely.

For developers: the API is the more significant unlock. Conversational editing over a persistent video object means you can build real iteration loops — generate, request a change, render only the changed pass — into pipelines for ads, product videos, or localization experiments.

For EU businesses: the built-in SynthID watermark aligns with AI Act transparency duties, but keep GDPR in mind for anything containing personal data. And set your non-English dialogue expectations carefully: the architecture is promising, but Google has yet to prove the Punjabi result generalizes to smaller European languages.

Is Gemini Omni available in Europe?

Yes — Google is rolling it out through the paid Google AI tiers (Plus, Pro, Ultra) and the free surfaces YouTube Shorts and YouTube Create, with API access for developers. No US-only restriction was announced; EU subscribers use the same plans they already have for other Gemini models.

What is conversational video editing?

It means you keep the same generated video and ask for changes in natural language — modifying a motion, a detail, or a character's action — instead of writing a new prompt and regenerating the whole clip. DeepMind's engineers list it as one of the three core capabilities of the model.

Can the SynthID watermark be removed or detected?

SynthID is invisible and is added automatically to every Gemini Omni output. It is designed for attribution — Google's technical answer to making AI-generated video identifiable, which is directly relevant to EU AI Act transparency rules.

Discussion

No comments yet — be the first to share your thoughts.
X

Don't miss out!

Subscribe for the latest news and updates.