Back to Blog

Five Places Sell Gemini Omni 1.1 Flash. The Price Moves 44%

E2X Team··15 min read
gemini-omnigooglepricingtext-to-videovideo-generation

Nobody comparison-shops model weights. You pick a model name, land on whichever provider page Google surfaced, and pay what the page says.

On 28 August 2026 we went looking for every place publishing a callable price for one model — gemini-omni-1.1-flash, shipped by Google the day before. Five do. Same weights, same native audio, same SynthID watermark in every frame. The number at the bottom of the page moved by forty-four percent.

What follows is that audit — including the part where we tell you to buy something nine times cheaper than the model this post is about.

A row of identical unlabelled tin canisters on a worn zinc counter, each with a blank paper tag hanging at a different height

The five prices, read on the same afternoon

Everything below is 720p, per second of finished video, read from each provider's live page on 28 August 2026. Audio is included everywhere: the model has no audio switch, so sound is generated with the picture and nobody bills it separately.

Where you call itPer second, 720pEight-second clip
E2X$0.09$0.72
fal.ai v1.1/edit$0.10$0.80
Runware$0.10$0.80
Google, direct$0.1014 effective$0.81
fal.ai base text-to-video$0.125$1.00
fal.ai base reference / image-to-video$0.13$1.04
WaveSpeed$0.13$1.04

Bottom to top, $0.09 against $0.13. That is the forty-four percent, and it buys nothing — no better second of video, no longer clip, no higher ceiling. Same file, same datacentre.

Three of those rows belong to one company, which is the next section.

Google's row needs a footnote, because Google's pricing page bills in tokens rather than seconds. Video output is $17.50 per million tokens, and 720p consumes 5,792 tokens a second; multiply and you get $0.1014. That conversion is Google's own arithmetic, not ours — the page prints it as "an effective price of approximately $0.10 per second."

One company, one model family, two pricing logics

The strangest finding is not a competitor. It is fal.ai disagreeing with itself.

Open fal.ai/models/google/gemini-omni-flash/v1.1/edit and you get $0.10 a second on a ladder identical to Google's rounded list, rung for rung, no margin at all. That is not an independent price but a passthrough — treat it as a second data point and you are counting Google twice. Open the base gemini-omni-flash pages on the same domain and the model is priced in tokens instead, at $21.875 per million — about $0.125 for text-to-video, $0.13 for reference and image jobs. Same company, same family, one listing at cost and one marked up.

That markup is exactly twenty-five percent over Google's token rate. Nothing wrong with it: an aggregator handling retries and giving you one auth surface is entitled to charge. But it means "the price of Omni 1.1 Flash" has no single referent even inside one storefront. So check the version string on the page you are billing against — a page reading gemini-omni-flash with no version may be selling you neither 1.1 nor its price.

A wall of identical antique brass post office boxes, one small door standing slightly ajar to show a dark empty interior

One provider is missing a number entirely. Replicate lists the model and publishes no per-second rate for it — we looked, and there is nothing to quote. That is a finding, not an omission on our part: a video model whose cost you cannot read before you run it is a model you cannot budget.

Google's resolution ladder is published twice, in two different shapes

Everything above is one rung. Omni 1.1 Flash sells four, and finding the other three took far longer than it should have.

Google publishes it twice, in neither of the places you would look. The Cloud pricing footnote gives token counts per second, per tier — the billing truth, since the invoice is cut on tokens. The launch announcement carries a rounded dollar table, but as a rendered image, so scrapers and the articles built from them skip past it. The Gemini API pricing page, the obvious stop, has only the 720p cell.

We got far enough down that road to conclude the four-rung ladder was a reseller's invention, and drafted a section saying so, before going back and reading the picture. Worth admitting, because it is the exact failure mode this post is about: a number that is published, sourced and correct can still be invisible to the way most people check.

The ladder from the token counts, which is what you actually pay:

ResolutionGoogle tokens/secGoogle $/secE2XDifference
360p1,931$0.0338$0.0297E2X 12% cheaper
720p5,792$0.1014$0.09E2X 11% cheaper
1080p8,688$0.1520$0.135E2X 11% cheaper
4K17,376$0.3041$0.18E2X 41% cheaper

Read the top row for the thing that actually changes how you work. A 360p second costs exactly a third of a 720p second on our ladder — $0.0297 against $0.09 — so three drafts cost what one full-resolution pass costs. That is the whole argument for a draft tier in one sentence, and it is why the cheap rung is worth having even when the quality is obviously worse.

Now the bottom row, where working from token counts pays off: the multiplier argument stops being rhetoric and becomes Google's own data. Take 720p as 1×. Google's tiers run 0.33 · 1 · 1.5 · 3. Ours run 0.33 · 1 · 1.5 · 2. Identical at the bottom, identical in the middle, and then Google triples into 4K where we double — which is why ten seconds of 4K is $3.04 calling Google and $1.80 calling us.

Four tall glass measuring cylinders on pale limestone, each filled with clear water to a progressively higher level, soft window light from the right

All four capabilities bill on that same ladder here — reference-to-video costs what start-end-frame-to-video costs, which is what image-to-video and plain text-to-video cost. Not universal: on fal's base pages, reference and image jobs cost more than prompt-only ones.

Runware's cheap rung sits above an expensive one

At 720p Runware matches the field at $0.10. Climb one rung and it charges $0.16 at 1080p and $0.32 at 4K — against Google's own $0.1520 and $0.3041. It is more expensive than the source it resells at exactly the two resolutions where a second costs the most. Only about five percent over Google; but compare providers by the first number you see, pick it for a 4K job, and you pay nearly double the $0.18 that rung costs here.

The lesson is worth more than the figure: compare at the tier you actually render at. A headline rate is almost always a provider's cheapest useful rung, and the ratio between rungs is set by the reseller.

What calling Google directly does not include

The instinct says cut the middle layer and buy from the source. Here the source is neither the cheapest option nor the least restrictive.

  • No free tier for video output. Text and image work on Gemini's free tier; video does not. Nothing to test against.
  • No Batch discount. Other Gemini 3.x models take roughly fifty percent off for batch submission. Video is excluded, so the lever that halves a large job is gone.
  • Data handling. Even on the paid tier, Google's terms mark content submitted to this model as usable for product improvement.
  • A price you have to hunt for. The docs show one video cell; the ladder is an image in a blog post.

None of that makes calling Google wrong. It makes "buy direct, skip the margin" a weaker heuristic than it sounds — on this model there is barely a margin to skip, since the reseller mirroring the list price hands it to you at cost.

The clip below cost $0.72

We generated this with Omni 1.1 Flash itself, eight seconds at 720p, at the price this post is arguing about. Eight times nine cents.

A night aerial over midtown Manhattan pushes into Times Square until one digital board fills the shot: OMNI FLASH 1.1, then AVAILABLE NOW ON E2X.AI, resolving to the E2X logo over wet asphalt.

The prompt, verbatim:

Cinematic 8-second commercial set in New York City at night. Opens on a wide aerial shot of Manhattan, then pushes rapidly into Times Square. Neon billboards, yellow taxis, pedestrians, wet asphalt reflecting colored light. The camera flies toward one massive digital billboard dominating the square. On the billboard, clearly readable text appears: "OMNI FLASH 1.1". Then: "THE LOWEST PRICE." The billboard transitions to: "AVAILABLE NOW ON E2X.AI". Final text: "START USING OMNI FLASH 1.1 TODAY". End with a clean bold E2X.AI logo filling the billboard. Fast energetic camera movement, realistic New York atmosphere, cinematic lighting, photorealistic detail, premium tech commercial look, high contrast, shallow depth of field, realistic reflections, subtle lens flares, smooth transitions. No narration, no one speaks. The message is carried entirely by the giant Times Square billboard. Sound design: city ambience, distant traffic, a rising synth swell, no voices, no speech.

Run as text-to-video, eight seconds, 720p, 16:9. It returned in 94 seconds against our own pipeline's estimate of about 240 — that estimate is ours; Google publishes no latency figure. One sample is not a benchmark, but if you have been budgeting four minutes per generation, budget less.

Now read the prompt against the picture, because this became an accidental test of the model's weakest skill. The prompt names billboard text four times, in quotation marks, and the hero board delivers — both lines clean, logo landed. Every other billboard in frame, the ones the prompt gestured at with "neon billboards", comes back as convincing gibberish. So the rule is not "this model cannot do text" but something narrower: text you specify tends to render; text the model invents to dress the set does not. Shaky text rendering has been reported on this family since launch — implicator.ai flagged it, and we could not find it stated that way on Google's model card — so take the clip as our evidence, not a vendor admission.

Before you pay nine cents, look at one cent

Here is where the audit argues against its own subject.

Veo 3.1 Fast costs $0.01 per second on the same API, same key — one ninth of Omni 1.1 Flash. The clip above would have been eight cents there instead of seventy-two, and for a standalone shot with nothing depending on it, very likely good enough. On Google's own table Veo 3.1 Fast lists at $0.10 a second, so there we are a tenth of the source rather than eleven percent under it.

So the honest ranking is not "Omni is cheapest here." It is that most people reading a post about the cheapest way to call Omni should be calling something else. Reach for it when the job needs what only it does — a specified first and last frame, a character who survives across shots, extension to forty seconds. A dolly shot is not one of those.

One comparison would be dishonest to make backwards. Omni Flash, the June preview, bills $0.0825 a second for text-to-video and $0.075 for reference jobs here — so 1.1 is nine to twenty percent more expensive than the model it replaces, not cheaper. Worth being precise about where that comes from: all three sit on the same $0.15 list rate and differ only in the discount we apply, 40 percent against 45 and 50. The increase is a smaller discount, not a dearer product. What it buys is image-to-video and start-and-end-frame control, a 360p rung the older enum has no entry for, and a ten-second ceiling — a fair trade, but an increase, and anyone telling you the new model is cheaper has the direction backwards.

Running this audit yourself

Five habits that will outlive these particular numbers:

  1. Read the version string, not the model name. Two pages on one domain sold us the same family at $0.10 and $0.125 on one afternoon.
  2. Check whether the source is text. Google's dollar ladder is an image and its real one a token footnote, so scrapers quote only the number sitting in the docs as HTML.
  3. Price the tier you render at, not the one the provider leads with. Runware is level at 720p and the dearest listing at 4K.
  4. Convert tokens to seconds before comparing. $17.50 per million means nothing until you multiply by the tokens your tier burns.
  5. Check whether a per-second price exists at all. Replicate carries the model without publishing one.

Prices and product terms can change. Check each provider's current pricing before making a purchasing decision. Google Batch or Flex pricing is not directly equivalent to a standard on-demand API request because scheduling, availability, and processing conditions differ; it is therefore excluded here. This is a scoped comparison across the five providers publishing a callable per-second price, not a claim that E2X is the cheapest option in every configuration. The cheapest way to make most clips is not this model at all, as the section above says.

Then draft at the cheapest rung you can tolerate and render the keeper at 720p. Parameters and current rates per capability are in the machine-readable spec; the lineup is in our catalog. We covered what shipped in 1.1, what the leaderboards measured, how to prompt it and why a 40-second scene arrives as four separate bills separately.

Frequently asked questions

What is the cheapest way to call Gemini Omni 1.1 Flash?

Across the five providers publishing a callable price on 28 August 2026, E2X came out lowest on all four rungs — $0.0297 for a 360p second, nine cents at 720p, $0.135 one tier up, $0.18 at 4K. The 720p field runs $0.10 at fal.ai's v1.1 page and Runware, $0.1014 from Google, $0.13 at WaveSpeed and on fal's base reference and image endpoints. Our margin is thinnest at 360p, about one percent against fal, and widest at 4K.

What does Gemini Omni 1.1 Flash cost if I call Google directly?

Google bills in tokens, not seconds: $17.50 per million for video output. Its Cloud pricing footnote lists the tokens each resolution burns per second, and multiplying through gives a shade over three cents for a 360p second, a shade over ten for 720p, fifteen and a fifth for 1080p, and just past thirty at 4K. A rounded version of the same ladder appears in the launch announcement, as an image rather than text. No free tier, no Batch discount.

Why do different providers charge different prices for the same model?

They price on different bases. Some mirror Google's rounded list, which is why fal.ai's v1.1 page and Runware both land on $0.10. Others price in tokens with margin added: fal's base pages use $21.875 per million against Google's $17.50, a 25% uplift reaching $0.125 to $0.13. WaveSpeed states a flat $0.13. The weights and the output are identical throughout.

Does Replicate publish a price for Gemini Omni 1.1 Flash?

Replicate lists the model but publishes no per-second rate for it, as of 28 August 2026. That makes it impossible to include in a like-for-like comparison and hard to budget, since video billing scales with duration and resolution rather than request count.

Is Gemini Omni 1.1 Flash cheaper than the model it replaces?

No. The previous Gemini Omni Flash preview bills $0.0825 per second for text-to-video and $0.075 for reference-to-video on E2X, against $0.09 for Omni 1.1 Flash — nine to twenty percent more. All three run off the same $0.15 list rate and differ only in discount, so the rise is a smaller discount rather than a dearer model. The extra buys image-to-video, start-and-end-frame control, a 360p draft rung and ten-second clips.

Which resolution should I generate at to keep costs down?

Draft at 360p and render only the keeper at 720p. On E2X a 360p second is exactly a third of a 720p one, so four seconds costs $0.1188 against $0.36 — five drafts plus one full render comes to $0.954, still under three straight 720p attempts at $1.08. 4K rarely earns its keep for web delivery: ten seconds is $1.80 here against $3.04 from Google.

Is there a cheaper video model than Gemini Omni 1.1 Flash?

Yes, and for most single-clip work it is the better call. Veo 3.1 Fast is a cent a second on E2X — one ninth of Omni 1.1 Flash, and a tenth of Google's own $0.10 listing for it — covering prompt-only jobs, animated stills, reference clips and first-and-last-frame work. Omni earns its premium only on scene extension, ten-second generations, the 360p tier, or reference consistency across shots.