Gemini Omni 1.1 Flash
Generate video from up to seven reference images with Gemini Omni 1.1 Flash. The reference images carry character and style consistency into the result.
- video
- high-fidelity
Faster, lower-cost Veo 3.1 tier covering text-to-video, image-to-video, reference-to-video and start-end frame transitions.
Sign in to run models
No output yet
Run the model to see generated results.
import requests
result = requests.post(
'https://api.e2x.ai/v1/jobs/submit',
headers={
'Authorization': f'Bearer {API_KEY}',
'Content-Type': 'application/json'
},
json={
'model': 'google/veo-3.1-fast/reference-to-video',
'input': {}
}
)import time
job_id = result.json()['jobId']
while True:
response = requests.get(
f'https://api.e2x.ai/v1/jobs/{'{job_id}'}',
headers={'Authorization': f'Bearer {'{API_KEY}'}'}
)
data = response.json()['data']
if data['status'] == 'completed':
print('Done!', data['outputs'][0]['url'])
break
elif data['status'] == 'failed':
raise Exception(f"Job failed: {'{'}data['error']['message']{'}'}")
time.sleep(2)Estimated cost per generation by duration and resolution. You only pay for what you run.
Discover more AI models for Reference to Video
This is the endpoint you use when the same face, the same jacket or the same shopfront has to survive across a whole batch of clips. We run Google's Veo 3.1 Fast as google/veo-3.1-fast/reference-to-video, it takes an image_urls array of up to three reference images, and we bill per second: $0.01 per second, ×1.5 with audio on. Every clip here is 8 seconds, so the price is $0.12 per clip at 720p with sound.
One image enough, and a shorter clip would help? Use Veo 3.1 Fast image-to-video — it animates a single starting frame and will give you 4 or 6 seconds. Nothing to upload at all? Veo 3.1 Fast text-to-video runs on a prompt alone.
Per-second metering is universal for this model — Google, fal.ai and Replicate all do it, and so do we. Since reference-to-video is fixed at 8 seconds, the table below is the only configuration that matters. Figures checked on August 26, 2026, at 720p with audio:
| API provider | Billing | Rate (720p, audio on) | 8-second clip | vs. us |
|---|---|---|---|---|
| E2X (current) | per second | $0.01/s ×1.5 audio | $0.12 | — |
| Google Gemini API | per second | $0.10/s | $0.80 | 6.7× our price |
| fal.ai | per second | $0.15/s | $1.20 | 10× our price |
| Replicate | per second | $0.15/s | $1.20 | 10× our price |
| E2X list (pre-discount) | per second | $0.10/s ×1.5 | $1.20 | our undiscounted rate |
Read the last row against the two marketplaces. Our list price is $1.20 for an 8-second clip — precisely what fal.ai and Replicate bill. You pay a tenth of that today, and it still lands 6.7× under Google's own API. That gap is a live discount on an identical model, not a stripped-down tier.
Price 1080p deliberately: fal.ai and Replicate include the 720p → 1080p step at no extra cost, while Google charges a higher per-second rate. Resolution is a multiplier on our side too — the current number lives in this model's llms.txt.
Prices and product terms can change. Check each provider's current pricing before making a purchasing decision. Google Batch or Flex pricing is not directly equivalent to a standard on-demand API request because scheduling, availability, and processing conditions differ; it is therefore excluded from this comparison. This is a scoped comparison, not a claim that E2X is the world's cheapest option in every configuration.
Google's limit is explicit: "Use up to three reference images to guide the content", described in their schema as "Up to three images to be used as style and content references." Three. A fourth is outside what the model accepts.
That budget forces a decision, and it's the interesting part of this endpoint. You spend three slots on whatever must not drift. A campaign shoot usually goes: one clean portrait of the presenter, one full-length shot showing the outfit, one photo of the location. Every clip in the set then comes back with the same person, same clothes, same place — you change only the prompt.
Product work splits differently: two angles of the object, one frame of the brand's colour treatment. Same principle. References carry identity, the prompt carries action.
This is not image-to-video's job. There, your image is frame one. Here the references are guidance and the opening frame is generated from them. Need a clip that starts on an exact approved still? That's the other endpoint.
Google's rule reads: duration "Must be '8' when using extension, reference images or with 1080p and 4k resolutions." Reference images are on that list, so the 4- and 6-second options the other two endpoints offer don't exist here. Our schema still shows the duration enum; for this capability the only valid value is "8".
Budget accordingly. A six-clip sequence is 48 seconds of output — $0.72 with us, $7.20 on fal.ai or Replicate.
Required fields: prompt and image_urls.
curl -X POST https://api.e2x.ai/v1/jobs/submit \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google/veo-3.1-fast/reference-to-video",
"input": {
"prompt": "The woman from image 1, wearing the coat from image 2, walks into the lobby from image 3 and shakes off the rain. She looks up and says: \"You waited.\" Warm tungsten light, marble echo.",
"image_urls": [
"https://example.com/talent-portrait.jpg",
"https://example.com/wardrobe-coat.jpg",
"https://example.com/lobby-set.jpg"
],
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": "8",
"has_sound": "true"
}
}'
Submit returns a job ID; poll it or register a webhook:
const headers = {
Authorization: `Bearer ${process.env.E2X_API_KEY}`,
"Content-Type": "application/json",
};
const submitted = await fetch("https://api.e2x.ai/v1/jobs/submit", {
method: "POST",
headers,
body: JSON.stringify({
model: "google/veo-3.1-fast/reference-to-video",
input: {
prompt:
"The mascot from image 1 hops onto the counter from image 2, knocks over a cup and freezes. Squeaky footsteps, a ceramic clink, then silence.",
image_urls: [
"https://example.com/mascot-turnaround.png",
"https://example.com/kitchen-set.jpg",
],
aspect_ratio: "9:16",
resolution: "720p",
duration: "8",
has_sound: "true",
seed: 60411,
},
}),
}).then((r) => r.json());
const jobId = submitted.data.jobId;
while (true) {
const job = await fetch(`https://api.e2x.ai/v1/jobs/${jobId}`, { headers })
.then((r) => r.json());
if (job.data.status === "completed") {
console.log(job.data.outputs[0].url);
break;
}
if (job.data.status === "failed") {
throw new Error(job.data.error?.message || "Generation failed");
}
await new Promise((r) => setTimeout(r, 5000));
}
Status goes pending → processing → completed, or ends at failed / cancelled. Plan on about four minutes per job — Google's official window is an 11-second floor and 6 minutes at peak. Reuse one seed across a set if you want the generations to rhyme.
Identical weights, identical per-second price. The shape of your input picks the endpoint:
| Endpoint | Duration options | 8s clip, 720p + audio | Pick it when |
|---|---|---|---|
| Veo 3.1 Fast reference-to-video | 8s only | $0.12 | Up to 3 images must fix a face, outfit or place across clips |
| Veo 3.1 Fast image-to-video | 4 / 6 / 8s | $0.12 | One still should be the literal first frame |
| Veo 3.1 Fast text-to-video | 4 / 6 / 8s | $0.12 | Prompt only, nothing to upload |
| Omni Flash reference-to-video | — | $0.80 (default config) | A different family, when Veo's look isn't the one you want |
We'd rather you spent less. If one image does the job, image-to-video is the same price and unlocks 4- and 6-second clips — a 4-second version costs $0.06, half of the fixed 8 seconds you pay here. Come to this endpoint when consistency across multiple clips is the actual requirement, because that is the one thing the others can't do. The rest of the category sits under reference-to-video; everything else is in the model catalog.
The model receives an unlabelled array. Nothing in the payload says slot two was wardrobe rather than a second character — so say it in the prompt.
Indexing works: "the woman from image 1", "the coat from image 2", "the lobby from image 3". A few words, and most of the ambiguity that produces mixed-up outputs disappears.
Three more habits:
Give each reference one job. A busy photo holding a person, a product and a room asks the model to guess what you cared about. Crop tight.
Repeat the identity in words. "Same short silver hair, same round glasses" reinforces what the reference shows. Cheap insurance.
Write the dialogue out. Google generates audio natively with no off switch in their API, so sound is happening either way — speech, effects, ambience. Quote the line and it lands on the right frames.
Google supports English fully and has evaluated nothing else for this model, so keep the instruction text English — the quoted dialogue can be whatever language the scene needs.
Two days to download. Google's retention on generated video is two days — "you must download your video within 2 days of generation." For a campaign of twenty clips built over a week, that's an architecture decision, not a footnote. Copy outputs[0].url into your own storage the moment a job completes.
SynthID on every output. Google's imperceptible watermark is applied to all Veo video and is verifiable through their platform. No provider can disable it.
Person generation is region-limited. In the EU, UK, Switzerland and MENA, personGeneration is restricted to allow_adult.
Version your reference set. Consistency depends on sending byte-identical references every time. A re-exported portrait swapped in halfway through a batch will show.
We bill $0.01 per second of output, ×1.5 with audio on. This endpoint is fixed at 8 seconds, so a 720p clip with sound is $0.12. Higher resolutions carry their own multiplier — check the live model page before planning a 1080p batch.
Three at most. Google's documentation says you may use up to three reference images to guide the content, and their schema describes them as style and content references. Fewer is fine; two is common.
Google requires 8 seconds whenever reference images are in play — the same rule that covers 1080p and 4k. Need 4 or 6 seconds? Use the image-to-video or text-to-video endpoint instead: same model, same per-second price.
Image-to-video makes one image the literal opening frame. Reference-to-video takes up to three images as guidance for identity, wardrobe, style or location, then generates an opening frame consistent with them. Use references when consistency across several clips matters more than a specific first frame.
Yes. Google generates dialogue, sound effects and ambience natively, with no way to disable it in their own API. We expose has_sound with a true default, and that setting applies the ×1.5 price multiplier.
All Veo output carries SynthID, Google's imperceptible marker for AI-generated content, verifiable on their platform. No provider — us included — can turn it off.
We budget roughly four minutes. Google publishes an 11-second minimum and a 6-minute maximum at peak, so for a full set of clips use a webhookUrl rather than holding polling loops open.