Gemini Omni 1.1 Flash
Generate video from a text prompt with Gemini Omni 1.1 Flash. Four output resolutions up to 4K and durations from 4 to 10 seconds.
- video
- high-fidelity
Faster, lower-cost Veo 3.1 tier covering text-to-video, image-to-video, reference-to-video and start-end frame transitions.
Sign in to run models
No output yet
Run the model to see generated results.
import requests
result = requests.post(
'https://api.e2x.ai/v1/jobs/submit',
headers={
'Authorization': f'Bearer {API_KEY}',
'Content-Type': 'application/json'
},
json={
'model': 'google/veo-3.1-fast/text-to-video',
'input': {}
}
)import time
job_id = result.json()['jobId']
while True:
response = requests.get(
f'https://api.e2x.ai/v1/jobs/{'{job_id}'}',
headers={'Authorization': f'Bearer {'{API_KEY}'}'}
)
data = response.json()['data']
if data['status'] == 'completed':
print('Done!', data['outputs'][0]['url'])
break
elif data['status'] == 'failed':
raise Exception(f"Job failed: {'{'}data['error']['message']{'}'}")
time.sleep(2)Estimated cost per generation by duration and resolution. You only pay for what you run.
Discover more AI models for Text to Video
Veo 3.1 Fast is Google's speed-tuned Veo 3.1: the same resolutions, the same 4/6/8-second durations, the same always-on native audio, turned around faster and priced lower. We expose it as google/veo-3.1-fast/text-to-video and we bill per second of output — $0.01 per second, ×1.5 when the audio track is on. An 8-second 720p clip with sound comes to $0.12. Nothing to upload. A prompt is the whole input.
Got a still image you want brought to life instead? Then you want Veo 3.1 Fast image-to-video — same model, one starting frame. Need the same face and the same jacket to survive across a batch of shots? That is Veo 3.1 Fast reference-to-video.
Every serious provider of this model bills by the second, us included. So the only honest comparison is a fixed configuration — here it is at 720p with audio on, checked on August 26, 2026:
| API provider | Billing | Rate (720p, audio on) | 8-second clip | vs. us |
|---|---|---|---|---|
| E2X (current) | per second | $0.01/s ×1.5 audio | $0.12 | — |
| Google Gemini API | per second | $0.10/s | $0.80 | 6.7× our price |
| fal.ai | per second | $0.15/s | $1.20 | 10× our price |
| Replicate | per second | $0.15/s | $1.20 | 10× our price |
| E2X list (pre-discount) | per second | $0.10/s ×1.5 | $1.20 | our undiscounted rate |
Read the last row against fal.ai and Replicate. Our list price for an 8-second clip is $1.20 — exactly what they charge. What you pay today is a tenth of that, and still 6.7× below Google's own API for the identical request.
Budget around one nuance: 720p → 1080p is free on fal.ai and Replicate, while Google charges more per second for it. Resolution is a multiplier on our side too, so pull the live figure from this model's llms.txt before committing to a 1080p pipeline.
Prices and product terms can change. Check each provider's current pricing before making a purchasing decision. Google Batch or Flex pricing is not directly equivalent to a standard on-demand API request because scheduling, availability, and processing conditions differ; it is therefore excluded from this comparison. This is a scoped comparison, not a claim that E2X is the world's cheapest option in every configuration.
Most text-to-video models hand you a silent MP4 and leave the sound to you. Veo 3.1 Fast doesn't. Google's API generates audio natively and you cannot switch it off there — synchronised dialogue, foley and room tone, produced alongside the picture and locked to the frames. Google's own documented example prompt contains spoken lines.
That changes what one call is worth. Write "a fishmonger slams a crate down and shouts fresh in, five minutes ago" and you get the slam and the shout on the right frames — not a mime you have to score later. For social cuts and animatics, an entire post step disappears.
What you do get to pick:
has_sound, defaulted to true, driving the ×1.5 multiplier.Eight seconds is the ceiling for one clip. Sequences are multiple calls.
Make a key in the dashboard, keep it server-side, post the job. Only prompt is required.
curl -X POST https://api.e2x.ai/v1/jobs/submit \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google/veo-3.1-fast/text-to-video",
"input": {
"prompt": "Handheld shot down a rain-slick Osaka alley at night. A ramen vendor lifts the noren, steam pours out, and he calls to a passing cyclist: \"Last bowl of the night!\" Neon reflections shiver in the puddles.",
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": "8",
"has_sound": "true"
}
}'
We return a job ID. Poll it, or don't:
const headers = {
Authorization: `Bearer ${process.env.E2X_API_KEY}`,
"Content-Type": "application/json",
};
const submitted = await fetch("https://api.e2x.ai/v1/jobs/submit", {
method: "POST",
headers,
body: JSON.stringify({
model: "google/veo-3.1-fast/text-to-video",
input: {
prompt:
"Slow push-in on a lighthouse keeper writing in a logbook. Wind rattles the glass. She murmurs, \"Third night, still no signal.\" Lamp sweeps past her face every four seconds.",
aspect_ratio: "9:16",
resolution: "720p",
duration: "6",
has_sound: "true",
seed: 41205,
},
}),
}).then((r) => r.json());
const jobId = submitted.data.jobId;
while (true) {
const job = await fetch(`https://api.e2x.ai/v1/jobs/${jobId}`, { headers })
.then((r) => r.json());
if (job.data.status === "completed") {
console.log(job.data.outputs[0].url);
break;
}
if (job.data.status === "failed") {
throw new Error(job.data.error?.message || "Generation failed");
}
await new Promise((r) => setTimeout(r, 5000));
}
Statuses run pending → processing → completed, or land on failed / cancelled. Send a webhookUrl instead and we'll call you. Budget about four minutes per job; Google publishes a floor of 11 seconds and a peak-hours ceiling of 6 minutes, so build for the ceiling.
One schema trap: duration and resolution are strings. "8", not 8.
Same weights, same per-second price, three different front doors. The input you already have decides it:
| Endpoint | 8s clip, 720p + audio | Pick it when |
|---|---|---|
| Veo 3.1 Fast text-to-video | $0.12 | You have words and nothing else — and you want 4s or 6s |
| Veo 3.1 Fast image-to-video | $0.12 | One still exists and should become the first frame |
| Veo 3.1 Fast reference-to-video | $0.12 | A person, product or place must stay identical across shots |
Price won't break the tie, so follow the brief. If a client already signed off on a key visual, sending it as a starting frame beats re-describing it in prose. If six clips have to look like one campaign, load up to three references — just know reference-to-video forces 8 seconds, no exceptions. This endpoint keeps the shorter durations, which makes it the cheapest way to fill a storyboard with rough motion. The rest of what we run sits in the text-to-video category and the full model catalog.
Treat the prompt as a shot list with a sound mix attached. Four things move the needle more than adjectives:
Name the camera move. "Locked-off wide", "slow dolly in", "handheld follow" — the model respects these. Vague framing gives you drifting, uncommitted motion.
Write dialogue in quotes. Short lines, one or two per clip. Eight seconds is not long. Paraphrased intent ("he says something encouraging") produces mumbled filler.
Ask for the ambience. Rain on tin, a fridge hum, distant traffic. The model invents a bed if you don't name one, and it may not match your edit.
Say what the light is doing. Overcast noon, sodium streetlight, one practical lamp. Lighting decides whether a Fast-tier clip reads as deliberate or as a render.
English is fully supported; Google hasn't evaluated other languages here. Write the dialogue in whatever language you want spoken and expect more variance outside English.
Download within 48 hours. Google keeps the file for two days — their wording is that you must download your video within 2 days of generation. Copy every output to your own bucket in the same job that reads outputs[0].url. A CDN link is not an archive.
Every frame carries SynthID. Google's imperceptible watermark is applied to all Veo output and is checkable on their verification platform. Nobody can turn it off, us included.
Region rules on people. In the EU, UK, Switzerland and MENA, personGeneration is restricted to allow_adult.
We bill per second: $0.01 per second of output, multiplied by 1.5 when audio is enabled. An 8-second 720p clip with sound is $0.12, a 6-second one is $0.09, and 4 seconds is $0.06. Resolution above 720p carries its own multiplier.
At the configuration and date we checked — 8 seconds, 720p, audio on, August 26, 2026 — yes: $0.12 with us against $1.20 there. Our own undiscounted list price is that same $1.20. The gap is a discount, not a different product.
Our schema exposes has_sound and we default it to true. Google's own Gemini API does not offer an off switch at all, so audio is best treated as part of what this model is — the same holds on our image-to-video endpoint. It's also the reason for the ×1.5 price multiplier.
Eight seconds, in a single call. You can request 4 or 6 on this endpoint, but 1080p and 4k require the full 8. Longer sequences mean stitching several jobs together yourself, or holding a look steady with reference-to-video.
16:9 and 9:16 only, at 24fps, in 720p, 1080p or 4k. We default to 16:9 and 720p. There is no square or 4:5 option, so crop in post if you need one.
Yes — every Veo output is watermarked with SynthID, Google's imperceptible provenance marker, and it can be verified through their platform. No provider exposes a parameter to disable it.
We budget around four minutes per job. Google publishes a minimum of 11 seconds and a maximum of 6 minutes at peak, so use a webhook rather than a tight polling loop for bulk work.