Gemini Omni 1.1 Flash
Animate a still image into video with Gemini Omni 1.1 Flash. The source image becomes the first frame and drives the generated motion.
- video
- high-fidelity
Faster, lower-cost Veo 3.1 tier covering text-to-video, image-to-video, reference-to-video and start-end frame transitions.
Sign in to run models
No output yet
Run the model to see generated results.
import requests
result = requests.post(
'https://api.e2x.ai/v1/jobs/submit',
headers={
'Authorization': f'Bearer {API_KEY}',
'Content-Type': 'application/json'
},
json={
'model': 'google/veo-3.1-fast/image-to-video',
'input': {}
}
)import time
job_id = result.json()['jobId']
while True:
response = requests.get(
f'https://api.e2x.ai/v1/jobs/{'{job_id}'}',
headers={'Authorization': f'Bearer {'{API_KEY}'}'}
)
data = response.json()['data']
if data['status'] == 'completed':
print('Done!', data['outputs'][0]['url'])
break
elif data['status'] == 'failed':
raise Exception(f"Job failed: {'{'}data['error']['message']{'}'}")
time.sleep(2)Estimated cost per generation by duration and resolution. You only pay for what you run.
Discover more AI models for Image to Video
Hand this endpoint one still and it becomes the first frame of a moving, scored shot. We run Google's Veo 3.1 Fast as google/veo-3.1-fast/image-to-video, and we bill per second of output: $0.01 per second, ×1.5 with the audio track on. So an 8-second 720p clip with sound is $0.12 — one image plus one prompt, no reference wrangling.
No starting image to work from? Then Veo 3.1 Fast text-to-video is your endpoint — same model, same price, prompt only, and it lets you drop to 4 or 6 seconds. If instead you have several images and the point is keeping one character consistent across shots, go to Veo 3.1 Fast reference-to-video.
Nobody sells this model per video. Google, fal.ai, Replicate and we all meter output second by second, so a fair comparison has to fix the configuration. Here is 8 seconds, 720p, audio on, when we last checked on August 26, 2026:
| API provider | Billing | Rate (720p, audio on) | 8-second clip | vs. us |
|---|---|---|---|---|
| E2X (current) | per second | $0.01/s ×1.5 audio | $0.12 | — |
| Google Gemini API | per second | $0.10/s | $0.80 | 6.7× our price |
| fal.ai | per second | $0.15/s | $1.20 | 10× our price |
| Replicate | per second | $0.15/s | $1.20 | 10× our price |
| E2X list (pre-discount) | per second | $0.10/s ×1.5 | $1.20 | our undiscounted rate |
Sit with that bottom row. Our undiscounted list price is $1.20 for an 8-second clip — the identical figure fal.ai and Replicate charge. The $0.12 you pay now is a tenth of it, and still 6.7× below Google's own API.
Resolution hits the bill differently depending on where you buy. Stepping 720p up to 1080p costs nothing extra on fal.ai or Replicate; Google charges a higher per-second rate. We treat resolution as a multiplier too, so read the current figure off this model's llms.txt before planning a 1080p run.
Prices and product terms can change. Check each provider's current pricing before making a purchasing decision. Google Batch or Flex pricing is not directly equivalent to a standard on-demand API request because scheduling, availability, and processing conditions differ; it is therefore excluded from this comparison. This is a scoped comparison, not a claim that E2X is the world's cheapest option in every configuration.
Your image anchors frame one. Everything after it is generated — hold that mental model, because it explains both the strength and the failure mode.
The strength: composition, palette, wardrobe and set are already settled. You aren't gambling on prose to reproduce a product shot the client already approved. Feed the still, describe the motion, get a clip that starts exactly where the still did.
The failure mode: the further the clip travels from that frame, the more it improvises. Ask for a 180° orbit in eight seconds and the far side of the subject is invention. Ask for a slow push-in and a head turn, and it holds.
The audio comes along for the ride. Google generates it natively and their API has no switch to disable it — dialogue in sync, footsteps on the footfall, room tone underneath. A still photograph goes in; something with a soundtrack comes out.
Google's constraints on the source image, plainly:
image_url when the job runs, not when you submit it.That last one causes more failures than the rest combined: short-lived signed links expire mid-queue.
Two required fields here: prompt and image_url.
curl -X POST https://api.e2x.ai/v1/jobs/submit \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google/veo-3.1-fast/image-to-video",
"input": {
"prompt": "The barista slides the cup forward and steam curls up. Slow push-in on the crema. Espresso machine hiss, low café chatter behind.",
"image_url": "https://example.com/counter-shot.jpg",
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": "8",
"has_sound": "true"
}
}'
You get a job ID back. Poll it, or hand us a webhook:
const headers = {
Authorization: `Bearer ${process.env.E2X_API_KEY}`,
"Content-Type": "application/json",
};
const submitted = await fetch("https://api.e2x.ai/v1/jobs/submit", {
method: "POST",
headers,
body: JSON.stringify({
model: "google/veo-3.1-fast/image-to-video",
input: {
prompt:
"She lowers the sunglasses and glances off-frame left. Hem lifts in the wind. Locked-off camera, gulls and surf underneath.",
image_url: "https://example.com/lookbook-03.jpg",
aspect_ratio: "9:16",
resolution: "720p",
duration: "8",
has_sound: "true",
seed: 88117,
},
}),
}).then((r) => r.json());
const jobId = submitted.data.jobId;
while (true) {
const job = await fetch(`https://api.e2x.ai/v1/jobs/${jobId}`, { headers })
.then((r) => r.json());
if (job.data.status === "completed") {
console.log(job.data.outputs[0].url);
break;
}
if (job.data.status === "failed") {
throw new Error(job.data.error?.message || "Animation failed");
}
await new Promise((r) => setTimeout(r, 5000));
}
Jobs move pending → processing → completed, or stop at failed / cancelled. Post a webhookUrl if you'd rather not poll. We plan for roughly four minutes; Google's published range is 11 seconds at best, 6 minutes at peak.
Watch the types: duration, resolution and has_sound are strings. "true", not true.
Three doors into identical weights, at an identical per-second price. Your input decides which one:
| Endpoint | 8s clip, 720p + audio | Pick it when |
|---|---|---|
| Veo 3.1 Fast image-to-video | $0.12 | One approved still should become frame one |
| Veo 3.1 Fast text-to-video | $0.12 | Nothing exists yet — and you want a 4s or 6s option |
| Veo 3.1 Fast reference-to-video | $0.12 | Up to three images must define a recurring face, product or place |
The honest split: use this endpoint when the still is the opening frame. Use reference-to-video when your images are guidance instead — the same actor across five scenes — and accept that it locks you to 8 seconds. Only exploring? Don't upload anything: text-to-video at 4 seconds costs $0.06 and burns through a storyboard fast. Browse the rest in the image-to-video category or the whole model catalog.
The common mistake is re-describing the image you just uploaded. The model can see it. Every word spent on "a woman in a red coat on a pier" is a word not spent on what happens next — and it invites reinterpretation of a frame you had already locked.
Write the change instead. Three habits carry most of the quality:
Give one primary motion. A head turn, a camera push, a door opening. Stack four actions into eight seconds and none of them read.
Say where the camera sits and whether it moves. "Locked off", "slow dolly in", "handheld drift right". Silence on this point produces aimless float.
Score it out loud. Name the ambience and any line of dialogue in quotes. The audio is generated whether you plan it or not, so plan it.
Prompts in English are fully supported. Google has not evaluated other languages for this model, so keep instructions in English even when the spoken line is in another one.
Two days. That's your download window. Google keeps the generated video for two days — their wording: you must download your video within 2 days of generation. Pull outputs[0].url into your own storage in the same code path that reads it. The returned link is temporary.
SynthID is on every frame. Google watermarks all Veo output with its imperceptible provenance marker, verifiable through their platform. No provider can disable it, us included.
People in regulated regions. Across the EU, UK, Switzerland and MENA, personGeneration is capped at allow_adult.
Match aspect ratio to your source. A 16:9 still with 9:16 requested forces the model to invent the edges.
We charge $0.01 per second of generated video, multiplied by 1.5 when audio is on. That makes an 8-second 720p clip with sound $0.12. Higher resolutions apply their own multiplier, so check the live model page before budgeting a 1080p or 4k batch.
Under 8 MB, at least 720p, and either 16:9 or 9:16. We fetch it from the image_url you provide, so the link must still resolve when the job runs — expiring signed URLs are the most common failure here.
Not on this endpoint — it takes exactly one image_url. If several images need to guide a generation, use reference-to-video, which accepts an array of up to three.
Yes, and it's a main reason to reach for this model. Google produces synchronised dialogue, effects and ambience natively, with no off switch in their own API. Our schema exposes has_sound, defaulted to true, carrying the ×1.5 multiplier.
Up to 8 seconds. This endpoint accepts 4, 6 or 8, but 1080p and 4k output require the full 8. Nothing in the Veo 3.1 Fast family exceeds 8 seconds in a single call.
Every Veo video carries SynthID, Google's imperceptible AI-provenance watermark, and it can be checked on their verification platform. Nobody offers a way to switch it off.
Plan for about four minutes per job. Google's official figures put the floor at 11 seconds and the peak-hour ceiling at 6 minutes, so a webhook beats a polling loop once you're running any real volume.