Back to Models

Veo 3.1 Fast

Faster, lower-cost Veo 3.1 tier covering text-to-video, image-to-video, reference-to-video and start-end frame transitions.

Pricingstarting $0.01/request
Latency~240 seconds average
Resolution720P/1080P/4K
Best forvideo, google, high-fidelity

Parameters

Estimated Cost

You save 90% on this model
Base cost per second$0.10
Discount-90%
Your Total$0.01
You save $0.09 per request

Sign in to run models

Output Preview

Ready

No output yet

Run the model to see generated results.

API Example— Current Parameters

generate.py
import requests

result = requests.post(
    'https://api.e2x.ai/v1/jobs/submit',
    headers={
        'Authorization': f'Bearer {API_KEY}',
        'Content-Type': 'application/json'
    },
    json={
  'model': 'google/veo-3.1-fast/image-to-video',
  'input': {}
}
)

Get Job— Poll for result

get_job.py
import time

job_id = result.json()['jobId']

while True:
    response = requests.get(
        f'https://api.e2x.ai/v1/jobs/{'{job_id}'}',
        headers={'Authorization': f'Bearer {'{API_KEY}'}'}
    )
    data = response.json()['data']

    if data['status'] == 'completed':
        print('Done!', data['outputs'][0]['url'])
        break
    elif data['status'] == 'failed':
        raise Exception(f"Job failed: {'{'}data['error']['message']{'}'}")

    time.sleep(2)

Pricing Breakdown

Estimated cost per generation by duration and resolution. You only pay for what you run.

Duration
720P
1080P
4K
4
$0.40
$0.40
$0.60
6
$0.60
$0.60
$0.90
8
$0.80
$0.80
$1.20
Explore Alternatives

Related Models

Discover more AI models for Image to Video

View all models

Veo 3.1 Fast Image-to-Video API on E2X

Hand this endpoint one still and it becomes the first frame of a moving, scored shot. We run Google's Veo 3.1 Fast as google/veo-3.1-fast/image-to-video, and we bill per second of output: $0.01 per second, ×1.5 with the audio track on. So an 8-second 720p clip with sound is $0.12 — one image plus one prompt, no reference wrangling.

No starting image to work from? Then Veo 3.1 Fast text-to-video is your endpoint — same model, same price, prompt only, and it lets you drop to 4 or 6 seconds. If instead you have several images and the point is keeping one character consistent across shots, go to Veo 3.1 Fast reference-to-video.

Per-second billing, and what that means at checkout

Nobody sells this model per video. Google, fal.ai, Replicate and we all meter output second by second, so a fair comparison has to fix the configuration. Here is 8 seconds, 720p, audio on, when we last checked on August 26, 2026:

API providerBillingRate (720p, audio on)8-second clipvs. us
E2X (current)per second$0.01/s ×1.5 audio$0.12—
Google Gemini APIper second$0.10/s$0.806.7× our price
fal.aiper second$0.15/s$1.2010× our price
Replicateper second$0.15/s$1.2010× our price
E2X list (pre-discount)per second$0.10/s ×1.5$1.20our undiscounted rate

Sit with that bottom row. Our undiscounted list price is $1.20 for an 8-second clip — the identical figure fal.ai and Replicate charge. The $0.12 you pay now is a tenth of it, and still 6.7× below Google's own API.

Resolution hits the bill differently depending on where you buy. Stepping 720p up to 1080p costs nothing extra on fal.ai or Replicate; Google charges a higher per-second rate. We treat resolution as a multiplier too, so read the current figure off this model's llms.txt before planning a 1080p run.

Prices and product terms can change. Check each provider's current pricing before making a purchasing decision. Google Batch or Flex pricing is not directly equivalent to a standard on-demand API request because scheduling, availability, and processing conditions differ; it is therefore excluded from this comparison. This is a scoped comparison, not a claim that E2X is the world's cheapest option in every configuration.

What the model does with your frame

Your image anchors frame one. Everything after it is generated — hold that mental model, because it explains both the strength and the failure mode.

The strength: composition, palette, wardrobe and set are already settled. You aren't gambling on prose to reproduce a product shot the client already approved. Feed the still, describe the motion, get a clip that starts exactly where the still did.

The failure mode: the further the clip travels from that frame, the more it improvises. Ask for a 180° orbit in eight seconds and the far side of the subject is invention. Ask for a slow push-in and a head turn, and it holds.

The audio comes along for the ride. Google generates it natively and their API has no switch to disable it — dialogue in sync, footsteps on the footfall, room tone underneath. A still photograph goes in; something with a soundtrack comes out.

Input requirements we enforce

Google's constraints on the source image, plainly:

  • Under 8 MB. Larger files are rejected before generation starts.
  • 720p or higher. Upscale a small asset before sending it, not after.
  • 16:9 or 9:16. Square and 4:5 won't do — re-crop first.
  • Reachable over HTTPS. We fetch image_url when the job runs, not when you submit it.

That last one causes more failures than the rest combined: short-lived signed links expire mid-queue.

API request: one image, one prompt

Two required fields here: prompt and image_url.

curl -X POST https://api.e2x.ai/v1/jobs/submit \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/veo-3.1-fast/image-to-video",
    "input": {
      "prompt": "The barista slides the cup forward and steam curls up. Slow push-in on the crema. Espresso machine hiss, low café chatter behind.",
      "image_url": "https://example.com/counter-shot.jpg",
      "aspect_ratio": "16:9",
      "resolution": "720p",
      "duration": "8",
      "has_sound": "true"
    }
  }'

You get a job ID back. Poll it, or hand us a webhook:

const headers = {
  Authorization: `Bearer ${process.env.E2X_API_KEY}`,
  "Content-Type": "application/json",
};

const submitted = await fetch("https://api.e2x.ai/v1/jobs/submit", {
  method: "POST",
  headers,
  body: JSON.stringify({
    model: "google/veo-3.1-fast/image-to-video",
    input: {
      prompt:
        "She lowers the sunglasses and glances off-frame left. Hem lifts in the wind. Locked-off camera, gulls and surf underneath.",
      image_url: "https://example.com/lookbook-03.jpg",
      aspect_ratio: "9:16",
      resolution: "720p",
      duration: "8",
      has_sound: "true",
      seed: 88117,
    },
  }),
}).then((r) => r.json());

const jobId = submitted.data.jobId;

while (true) {
  const job = await fetch(`https://api.e2x.ai/v1/jobs/${jobId}`, { headers })
    .then((r) => r.json());

  if (job.data.status === "completed") {
    console.log(job.data.outputs[0].url);
    break;
  }
  if (job.data.status === "failed") {
    throw new Error(job.data.error?.message || "Animation failed");
  }
  await new Promise((r) => setTimeout(r, 5000));
}

Jobs move pending → processing → completed, or stop at failed / cancelled. Post a webhookUrl if you'd rather not poll. We plan for roughly four minutes; Google's published range is 11 seconds at best, 6 minutes at peak.

Watch the types: duration, resolution and has_sound are strings. "true", not true.

Picking the right Veo 3.1 Fast endpoint

Three doors into identical weights, at an identical per-second price. Your input decides which one:

Endpoint8s clip, 720p + audioPick it when
Veo 3.1 Fast image-to-video$0.12One approved still should become frame one
Veo 3.1 Fast text-to-video$0.12Nothing exists yet — and you want a 4s or 6s option
Veo 3.1 Fast reference-to-video$0.12Up to three images must define a recurring face, product or place

The honest split: use this endpoint when the still is the opening frame. Use reference-to-video when your images are guidance instead — the same actor across five scenes — and accept that it locks you to 8 seconds. Only exploring? Don't upload anything: text-to-video at 4 seconds costs $0.06 and burns through a storyboard fast. Browse the rest in the image-to-video category or the whole model catalog.

Describe the motion, not the picture

The common mistake is re-describing the image you just uploaded. The model can see it. Every word spent on "a woman in a red coat on a pier" is a word not spent on what happens next — and it invites reinterpretation of a frame you had already locked.

Write the change instead. Three habits carry most of the quality:

Give one primary motion. A head turn, a camera push, a door opening. Stack four actions into eight seconds and none of them read.

Say where the camera sits and whether it moves. "Locked off", "slow dolly in", "handheld drift right". Silence on this point produces aimless float.

Score it out loud. Name the ambience and any line of dialogue in quotes. The audio is generated whether you plan it or not, so plan it.

Prompts in English are fully supported. Google has not evaluated other languages for this model, so keep instructions in English even when the spoken line is in another one.

Before you ship

Two days. That's your download window. Google keeps the generated video for two days — their wording: you must download your video within 2 days of generation. Pull outputs[0].url into your own storage in the same code path that reads it. The returned link is temporary.

SynthID is on every frame. Google watermarks all Veo output with its imperceptible provenance marker, verifiable through their platform. No provider can disable it, us included.

People in regulated regions. Across the EU, UK, Switzerland and MENA, personGeneration is capped at allow_adult.

Match aspect ratio to your source. A 16:9 still with 9:16 requested forces the model to invent the edges.

Questions we get asked

How much does Veo 3.1 Fast image-to-video cost on E2X?

We charge $0.01 per second of generated video, multiplied by 1.5 when audio is on. That makes an 8-second 720p clip with sound $0.12. Higher resolutions apply their own multiplier, so check the live model page before budgeting a 1080p or 4k batch.

What are the requirements for the input image?

Under 8 MB, at least 720p, and either 16:9 or 9:16. We fetch it from the image_url you provide, so the link must still resolve when the job runs — expiring signed URLs are the most common failure here.

Can I animate more than one image in a single call?

Not on this endpoint — it takes exactly one image_url. If several images need to guide a generation, use reference-to-video, which accepts an array of up to three.

Does the generated video include sound?

Yes, and it's a main reason to reach for this model. Google produces synchronised dialogue, effects and ambience natively, with no off switch in their own API. Our schema exposes has_sound, defaulted to true, carrying the ×1.5 multiplier.

How long is the generated clip?

Up to 8 seconds. This endpoint accepts 4, 6 or 8, but 1080p and 4k output require the full 8. Nothing in the Veo 3.1 Fast family exceeds 8 seconds in a single call.

Is there a watermark on the output?

Every Veo video carries SynthID, Google's imperceptible AI-provenance watermark, and it can be checked on their verification platform. Nobody offers a way to switch it off.

How fast is a generation?

Plan for about four minutes per job. Google's official figures put the floor at 11 seconds and the peak-hour ceiling at 6 minutes, so a webhook beats a polling loop once you're running any real volume.