Back to Models

Veo 3.1 Fast

Faster, lower-cost Veo 3.1 tier covering text-to-video, image-to-video, reference-to-video and start-end frame transitions.

Pricingstarting $0.01/request
Latency~240 seconds average
Resolution720P/1080P/4K
Best forvideo, google, high-fidelity

Parameters

Estimated Cost

You save 90% on this model
Base cost per second$0.10
Discount-90%
Your Total$0.01
You save $0.09 per request

Sign in to run models

Output Preview

Ready

No output yet

Run the model to see generated results.

API Example— Current Parameters

generate.py
import requests

result = requests.post(
    'https://api.e2x.ai/v1/jobs/submit',
    headers={
        'Authorization': f'Bearer {API_KEY}',
        'Content-Type': 'application/json'
    },
    json={
  'model': 'google/veo-3.1-fast/reference-to-video',
  'input': {}
}
)

Get Job— Poll for result

get_job.py
import time

job_id = result.json()['jobId']

while True:
    response = requests.get(
        f'https://api.e2x.ai/v1/jobs/{'{job_id}'}',
        headers={'Authorization': f'Bearer {'{API_KEY}'}'}
    )
    data = response.json()['data']

    if data['status'] == 'completed':
        print('Done!', data['outputs'][0]['url'])
        break
    elif data['status'] == 'failed':
        raise Exception(f"Job failed: {'{'}data['error']['message']{'}'}")

    time.sleep(2)

Pricing Breakdown

Estimated cost per generation by duration and resolution. You only pay for what you run.

Duration
720P
1080P
4K
4
$0.40
$0.40
$0.60
6
$0.60
$0.60
$0.90
8
$0.80
$0.80
$1.20
Explore Alternatives

Related Models

Discover more AI models for Reference to Video

View all models

Veo 3.1 Fast Reference-to-Video API on E2X

This is the endpoint you use when the same face, the same jacket or the same shopfront has to survive across a whole batch of clips. We run Google's Veo 3.1 Fast as google/veo-3.1-fast/reference-to-video, it takes an image_urls array of up to three reference images, and we bill per second: $0.01 per second, ×1.5 with audio on. Every clip here is 8 seconds, so the price is $0.12 per clip at 720p with sound.

One image enough, and a shorter clip would help? Use Veo 3.1 Fast image-to-video — it animates a single starting frame and will give you 4 or 6 seconds. Nothing to upload at all? Veo 3.1 Fast text-to-video runs on a prompt alone.

What consistency costs per clip

Per-second metering is universal for this model — Google, fal.ai and Replicate all do it, and so do we. Since reference-to-video is fixed at 8 seconds, the table below is the only configuration that matters. Figures checked on August 26, 2026, at 720p with audio:

API providerBillingRate (720p, audio on)8-second clipvs. us
E2X (current)per second$0.01/s ×1.5 audio$0.12—
Google Gemini APIper second$0.10/s$0.806.7× our price
fal.aiper second$0.15/s$1.2010× our price
Replicateper second$0.15/s$1.2010× our price
E2X list (pre-discount)per second$0.10/s ×1.5$1.20our undiscounted rate

Read the last row against the two marketplaces. Our list price is $1.20 for an 8-second clip — precisely what fal.ai and Replicate bill. You pay a tenth of that today, and it still lands 6.7× under Google's own API. That gap is a live discount on an identical model, not a stripped-down tier.

Price 1080p deliberately: fal.ai and Replicate include the 720p → 1080p step at no extra cost, while Google charges a higher per-second rate. Resolution is a multiplier on our side too — the current number lives in this model's llms.txt.

Prices and product terms can change. Check each provider's current pricing before making a purchasing decision. Google Batch or Flex pricing is not directly equivalent to a standard on-demand API request because scheduling, availability, and processing conditions differ; it is therefore excluded from this comparison. This is a scoped comparison, not a claim that E2X is the world's cheapest option in every configuration.

Three images, and not a fourth

Google's limit is explicit: "Use up to three reference images to guide the content", described in their schema as "Up to three images to be used as style and content references." Three. A fourth is outside what the model accepts.

That budget forces a decision, and it's the interesting part of this endpoint. You spend three slots on whatever must not drift. A campaign shoot usually goes: one clean portrait of the presenter, one full-length shot showing the outfit, one photo of the location. Every clip in the set then comes back with the same person, same clothes, same place — you change only the prompt.

Product work splits differently: two angles of the object, one frame of the brand's colour treatment. Same principle. References carry identity, the prompt carries action.

This is not image-to-video's job. There, your image is frame one. Here the references are guidance and the opening frame is generated from them. Need a clip that starts on an exact approved still? That's the other endpoint.

Eight seconds, no negotiation

Google's rule reads: duration "Must be '8' when using extension, reference images or with 1080p and 4k resolutions." Reference images are on that list, so the 4- and 6-second options the other two endpoints offer don't exist here. Our schema still shows the duration enum; for this capability the only valid value is "8".

Budget accordingly. A six-clip sequence is 48 seconds of output — $0.72 with us, $7.20 on fal.ai or Replicate.

API request: an array of references

Required fields: prompt and image_urls.

curl -X POST https://api.e2x.ai/v1/jobs/submit \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/veo-3.1-fast/reference-to-video",
    "input": {
      "prompt": "The woman from image 1, wearing the coat from image 2, walks into the lobby from image 3 and shakes off the rain. She looks up and says: \"You waited.\" Warm tungsten light, marble echo.",
      "image_urls": [
        "https://example.com/talent-portrait.jpg",
        "https://example.com/wardrobe-coat.jpg",
        "https://example.com/lobby-set.jpg"
      ],
      "aspect_ratio": "16:9",
      "resolution": "720p",
      "duration": "8",
      "has_sound": "true"
    }
  }'

Submit returns a job ID; poll it or register a webhook:

const headers = {
  Authorization: `Bearer ${process.env.E2X_API_KEY}`,
  "Content-Type": "application/json",
};

const submitted = await fetch("https://api.e2x.ai/v1/jobs/submit", {
  method: "POST",
  headers,
  body: JSON.stringify({
    model: "google/veo-3.1-fast/reference-to-video",
    input: {
      prompt:
        "The mascot from image 1 hops onto the counter from image 2, knocks over a cup and freezes. Squeaky footsteps, a ceramic clink, then silence.",
      image_urls: [
        "https://example.com/mascot-turnaround.png",
        "https://example.com/kitchen-set.jpg",
      ],
      aspect_ratio: "9:16",
      resolution: "720p",
      duration: "8",
      has_sound: "true",
      seed: 60411,
    },
  }),
}).then((r) => r.json());

const jobId = submitted.data.jobId;

while (true) {
  const job = await fetch(`https://api.e2x.ai/v1/jobs/${jobId}`, { headers })
    .then((r) => r.json());

  if (job.data.status === "completed") {
    console.log(job.data.outputs[0].url);
    break;
  }
  if (job.data.status === "failed") {
    throw new Error(job.data.error?.message || "Generation failed");
  }
  await new Promise((r) => setTimeout(r, 5000));
}

Status goes pending → processing → completed, or ends at failed / cancelled. Plan on about four minutes per job — Google's official window is an 11-second floor and 6 minutes at peak. Reuse one seed across a set if you want the generations to rhyme.

Which Veo 3.1 Fast door to use

Identical weights, identical per-second price. The shape of your input picks the endpoint:

EndpointDuration options8s clip, 720p + audioPick it when
Veo 3.1 Fast reference-to-video8s only$0.12Up to 3 images must fix a face, outfit or place across clips
Veo 3.1 Fast image-to-video4 / 6 / 8s$0.12One still should be the literal first frame
Veo 3.1 Fast text-to-video4 / 6 / 8s$0.12Prompt only, nothing to upload
Omni Flash reference-to-video—$0.80 (default config)A different family, when Veo's look isn't the one you want

We'd rather you spent less. If one image does the job, image-to-video is the same price and unlocks 4- and 6-second clips — a 4-second version costs $0.06, half of the fixed 8 seconds you pay here. Come to this endpoint when consistency across multiple clips is the actual requirement, because that is the one thing the others can't do. The rest of the category sits under reference-to-video; everything else is in the model catalog.

Label your references

The model receives an unlabelled array. Nothing in the payload says slot two was wardrobe rather than a second character — so say it in the prompt.

Indexing works: "the woman from image 1", "the coat from image 2", "the lobby from image 3". A few words, and most of the ambiguity that produces mixed-up outputs disappears.

Three more habits:

Give each reference one job. A busy photo holding a person, a product and a room asks the model to guess what you cared about. Crop tight.

Repeat the identity in words. "Same short silver hair, same round glasses" reinforces what the reference shows. Cheap insurance.

Write the dialogue out. Google generates audio natively with no off switch in their API, so sound is happening either way — speech, effects, ambience. Quote the line and it lands on the right frames.

Google supports English fully and has evaluated nothing else for this model, so keep the instruction text English — the quoted dialogue can be whatever language the scene needs.

Ship checklist

Two days to download. Google's retention on generated video is two days — "you must download your video within 2 days of generation." For a campaign of twenty clips built over a week, that's an architecture decision, not a footnote. Copy outputs[0].url into your own storage the moment a job completes.

SynthID on every output. Google's imperceptible watermark is applied to all Veo video and is verifiable through their platform. No provider can disable it.

Person generation is region-limited. In the EU, UK, Switzerland and MENA, personGeneration is restricted to allow_adult.

Version your reference set. Consistency depends on sending byte-identical references every time. A re-exported portrait swapped in halfway through a batch will show.

Common questions

How much does Veo 3.1 Fast reference-to-video cost on E2X?

We bill $0.01 per second of output, ×1.5 with audio on. This endpoint is fixed at 8 seconds, so a 720p clip with sound is $0.12. Higher resolutions carry their own multiplier — check the live model page before planning a 1080p batch.

How many reference images can I send?

Three at most. Google's documentation says you may use up to three reference images to guide the content, and their schema describes them as style and content references. Fewer is fine; two is common.

Why can't I generate a 4-second clip here?

Google requires 8 seconds whenever reference images are in play — the same rule that covers 1080p and 4k. Need 4 or 6 seconds? Use the image-to-video or text-to-video endpoint instead: same model, same per-second price.

What is the difference between this and image-to-video?

Image-to-video makes one image the literal opening frame. Reference-to-video takes up to three images as guidance for identity, wardrobe, style or location, then generates an opening frame consistent with them. Use references when consistency across several clips matters more than a specific first frame.

Does the output include synchronised audio?

Yes. Google generates dialogue, sound effects and ambience natively, with no way to disable it in their own API. We expose has_sound with a true default, and that setting applies the ×1.5 price multiplier.

Are the generated videos watermarked?

All Veo output carries SynthID, Google's imperceptible marker for AI-generated content, verifiable on their platform. No provider — us included — can turn it off.

How long does one job take to finish?

We budget roughly four minutes. Google publishes an 11-second minimum and a 6-minute maximum at peak, so for a full set of clips use a webhookUrl rather than holding polling loops open.