Gemini Omni 1.1 Flash Shipped Today. The News Is Editing, Not Generation
Google made gemini-omni-1.1-flash generally available today. Read the feature list and one thing stands out: almost none of it is about making a better eight-second clip.
Scene extension that reads ten seconds of context instead of one. The ability to hand the model a first frame and a last frame and have it build the middle. Reference clips that carry a character across shots. These are continuity features. They are the things you need when a clip is not the deliverable but a piece of one, and shipping them together is a clearer statement of direction than any quality claim would have been.

What is actually new
The predecessor here is Gemini Omni Flash — gemini-omni-flash-preview, in API preview since 30 June 2026. Note the naming, because the internet is already getting it wrong: there is no product called "Omni Flash 1.0". Google's own string is "Gemini Omni 1.1 Flash", and the thing it replaces is simply Gemini Omni Flash.
| Capability | Gemini Omni Flash (preview) | Gemini Omni 1.1 Flash |
|---|---|---|
| Scene extension context | Final second of the clip | Up to 10 seconds of prior footage |
| Total extended length | — | 40 seconds, built in 3–10s continuations |
| First-and-last frame | — | Supply both, model generates between them |
| Reference clips | — | Up to 3 clips, 3 seconds each |
| Resolutions | 720p | 360p, 720p, 1080p and 4K (the top two upscaled) |
| Per-generation duration | 3–10 seconds | 3–10 seconds |
The extension change is the one to read twice. A model that sees only the last frame of what it is continuing has no idea what the camera was doing, so extensions drift — the light shifts, a jacket changes colour, the dolly that was moving left stops. Ten seconds of context is enough to carry a camera move and a lighting setup across the join. Google describes the continuations as 3–10 second segments up to a 40-second total, so this is not a fixed block size; you are stitching in variable pieces.
There is a 360p mode as well, which Google says runs up to 60% faster at a third of the cost of 720p. That figure is from the announcement blog rather than the API reference, so treat it as a vendor claim rather than a measured number — but the shape of it is obviously right, and a cheap draft tier is exactly what iterative video work has been missing.
What it costs from Google directly
Video output bills at $17.50 per million tokens, and Google's pricing footnote converts that for you: 5,792 tokens per second of 720p video, "an effective price of approximately $0.10 per second." Input is $1.50 per million. There is no free tier for video output.
The other resolutions are published too, just nowhere near that footnote, which is why almost nobody quotes them. Google's launch announcement carries a price table as an image, and its Cloud pricing page carries the exact version as a sentence: Omni bills 1,931 tokens per second at 360p, 5,792 at 720p, 8,688 at 1080p and 17,376 at 4K. At $17.50 per million that is $0.0338, $0.1014, $0.1520 and $0.3041 a second. Note the shape rather than the digits — against 720p those token counts are exactly a third, one, one and a half, and three. The ladder triples from 720p to 4K, which is worth remembering when you are tempted to render a first pass at full resolution.
Ten cents a second is worth sitting with for a moment. A forty-second sequence, assembled from extensions at 720p, is about four dollars of output before you have generated a single alternative take. Video pricing is not image pricing, and the iteration habits that work at a fraction of a cent per image will empty an account here.
Where you can call it today
- Gemini API (
generativelanguage.googleapis.com) — GA - Google AI Studio — GA
- Gemini Enterprise Agent Platform — for enterprise API access
- Google Flow — for AI Plus, Pro and Ultra subscribers
- Gemini app — scene extension for Plus, Pro and Ultra
Not Vertex AI. Its release notes make no mention of Omni at all, and Google's announcement points enterprise users at the Agent Platform instead. If you have read otherwise, that page is guessing.
One dated item sits underneath all of this: the preview endpoint gemini-omni-flash-preview is scheduled for deprecation on 30 September 2026. Anything you built against the preview during the summer has about a month of runway.

The specifications, in one place
Per-generation output runs 3 to 10 seconds at 24 FPS, in 16:9 or 9:16 — those are the only two aspect ratios, so square and vertical-cinema framings are out. Audio is generated natively alongside the picture, and you cannot upload an audio reference to steer it. Any video you supply for editing or extension has to be 10 seconds or shorter.
Every frame carries SynthID, Google's invisible provenance watermark. It is undetectable to viewers and cannot be switched off by any provider, which matters only if you have a contract that forbids watermarked deliverables.
Google is also candid in the model card about what does not work yet:
"Maintaining complete consistency throughout edits, generating scenes with complex motion, or rendering perfectly accurate text remains a challenge."
Read that next to the feature list and it is a useful tension. The headline additions are consistency features; the model card says consistency is still the hard part. Both things are true, and a launch post that only quoted the first would be selling you something.
One capability is deliberately held back. The model can alter recorded speech, and Google has chosen not to ship it:
"Gemini Omni Flash is capable of changing people's speech. For now, we are restricting this capability and working to better understand how to safely and responsibly bring it to our users."
There are regional limits too: in the EEA, Switzerland and the UK you cannot upload or edit images containing minors or certain recognisable people.
What this means on E2X
We said we would not put a date on this one. It landed the next day. Gemini Omni 1.1 Flash went live on E2X on 28 August 2026, with all four of its capabilities callable through the same endpoint shape as everything else in the catalog.
One price covers all four: $0.09 per second at 720p for text-to-video, image-to-video, start-end-frame-to-video and reference-to-video, checked 28 August 2026. Durations of 4, 6, 8 and 10 seconds, at 360p, 720p, 1080p or 4K — $0.0297 a second on the draft tier, $0.18 at 4K. The middle two are new here: the previous generation only ever shipped text-to-video and reference-to-video on this platform, so animating a still and interpolating between a first and last frame are both first-time capabilities for this family.
Set that against Google's own per-second rates and the difference is really a difference of multipliers. Google runs $0.0338, $0.1014, $0.1520 and $0.3041 as you climb; we run $0.0297, $0.09, $0.135 and $0.18. The two ladders start in the same shape — a third of the 720p rate down at 360p, half again as much at 1080p — and then part company at the top, where Google triples and we double. That leaves us under Google at every step: about twelve percent at 360p, eleven at 720p and 1080p, and forty-one at 4K. What the price also buys is one API surface and one set of credentials across every model in our catalog, which is worth something without being a claim to be cheapest in every configuration.
The genuinely cheap option sits next to it, and most readers should look at it first. Veo 3.1 Fast text-to-video is $0.01 per second — a ninth of Omni 1.1 Flash. It is also the one figure on this page that is not a close call: Google's own table prices Veo 3.1 Fast at $0.10 a second at 720p, so the same model costs a tenth here of what it costs there. For a straight text-to-video shot with no editing, no extension and no character continuity, the nine-times premium over it buys you very little. We would rather you spent a dollar than nine and came back.
Where Omni earns the difference is exactly where the 1.1 features point: continuation, first-and-last frame control, and carrying a subject across shots. If your work is one clip at a time, Veo 3.1 Fast is the better call, and it also covers image-to-video, reference-to-video and start-and-end-frame jobs.
Prices and product terms can change. Check each provider's current pricing before making a purchasing decision. This is a scoped comparison, not a claim that E2X is the cheapest option in every configuration.
What to do this week
If you were on the preview endpoint, put the 30 September date in the calendar and move. If you are choosing a video model for the first time, start on Veo 3.1 Fast at a cent a second and only move up when a job actually needs continuity. And if you want to compare schemas rather than read prose, every model publishes a machine-readable spec — the Omni 1.1 Flash one lists every parameter, constraint and current price.
Browse the rest by job from text-to-video or reference-to-video.
Frequently asked questions
What is Gemini Omni 1.1 Flash?
Gemini Omni 1.1 Flash is Google's video generation and editing model, made generally available on 27 August 2026 under the API id gemini-omni-1.1-flash. It generates 3 to 10 second clips at 24 FPS with native audio, and adds editing features its predecessor lacked: scene extension to 40 seconds, first-and-last frame interpolation, and reference clips for character consistency.
When was Gemini Omni 1.1 Flash released?
27 August 2026, as a generally available model rather than a preview. Google's API changelog records it as "Gemini Omni Flash generally available (GA): Released gemini-omni-1.1-flash". The earlier preview model, gemini-omni-flash-preview, had been available since 30 June 2026 and is scheduled for deprecation on 30 September 2026.
Is there an Omni Flash 1.0?
No. Google has never shipped a product called Omni Flash 1.0. The model that came before Gemini Omni 1.1 Flash is called Gemini Omni Flash, with the API id gemini-omni-flash-preview, and it was a preview release. Pages describing an "Omni Flash 1.0" are inventing a version number that does not exist.
How much does Gemini Omni 1.1 Flash cost?
Calling Google directly, video output bills at $17.50 per million tokens, with input at $1.50 per million and no free tier for video. Google's Cloud pricing page gives the token cost of one second at each resolution — 1,931 at 360p, 5,792 at 720p, 8,688 at 1080p and 17,376 at 4K — which works out at $0.0338, $0.1014, $0.1520 and $0.3041 a second. On E2X the same tiers are $0.0297, $0.09, $0.135 and $0.18, which puts the default eight-second 720p clip at $0.72.
How long can a Gemini Omni 1.1 Flash video be?
A single generation produces 3 to 10 seconds. Using scene extension you can build up to 40 seconds total, adding 3 to 10 second continuations to an existing clip. Any video you supply as input for editing or extension must itself be 10 seconds or shorter.
Can I use Gemini Omni 1.1 Flash on Vertex AI?
Not as of 27 August 2026. Vertex AI's release notes make no mention of the Omni family. Google's announcement directs enterprise API users to the Gemini Enterprise Agent Platform, and the model is also available through the Gemini API, AI Studio, Flow and the Gemini app.
Is Gemini Omni 1.1 Flash available on E2X?
Yes, since 28 August 2026, across all four capabilities: text-to-video, image-to-video, start-end-frame-to-video and reference-to-video. It is $0.09 per second at 720p, with $0.0297 at 360p, $0.135 at 1080p and $0.18 at 4K, and durations of 4, 6, 8 and 10 seconds. That sits under Google's own per-second rate at every tier — about twelve percent at 360p, eleven percent at 720p and 1080p, and forty-one percent at 4K. Audio is generated natively and cannot be switched off. For single clips with no editing or continuity requirement, Veo 3.1 Fast at $0.01 per second is still usually the better value.