Back to Models

Veo 3.1 Fast

Faster, lower-cost Veo 3.1 tier covering text-to-video, image-to-video, reference-to-video and start-end frame transitions.

Pricingstarting $0.01/request
Latency~240 seconds average
Resolution720P/1080P/4K
Best forvideo, google, high-fidelity

Parameters

Estimated Cost

You save 90% on this model
Base cost per second$0.10
Discount-90%
Your Total$0.01
You save $0.09 per request

Sign in to run models

Output Preview

Ready

No output yet

Run the model to see generated results.

Sample Output

API Example— Current Parameters

generate.py
import requests

result = requests.post(
    'https://api.e2x.ai/v1/jobs/submit',
    headers={
        'Authorization': f'Bearer {API_KEY}',
        'Content-Type': 'application/json'
    },
    json={
  'model': 'google/veo-3.1-fast/text-to-video',
  'input': {}
}
)

Get Job— Poll for result

get_job.py
import time

job_id = result.json()['jobId']

while True:
    response = requests.get(
        f'https://api.e2x.ai/v1/jobs/{'{job_id}'}',
        headers={'Authorization': f'Bearer {'{API_KEY}'}'}
    )
    data = response.json()['data']

    if data['status'] == 'completed':
        print('Done!', data['outputs'][0]['url'])
        break
    elif data['status'] == 'failed':
        raise Exception(f"Job failed: {'{'}data['error']['message']{'}'}")

    time.sleep(2)

Pricing Breakdown

Estimated cost per generation by duration and resolution. You only pay for what you run.

Duration
720P
1080P
4K
4
$0.40
$0.40
$0.60
6
$0.60
$0.60
$0.90
8
$0.80
$0.80
$1.20
Explore Alternatives

Related Models

Discover more AI models for Text to Video

View all models

Veo 3.1 Fast Text-to-Video API on E2X

Veo 3.1 Fast is Google's speed-tuned Veo 3.1: the same resolutions, the same 4/6/8-second durations, the same always-on native audio, turned around faster and priced lower. We expose it as google/veo-3.1-fast/text-to-video and we bill per second of output — $0.01 per second, ×1.5 when the audio track is on. An 8-second 720p clip with sound comes to $0.12. Nothing to upload. A prompt is the whole input.

Got a still image you want brought to life instead? Then you want Veo 3.1 Fast image-to-video — same model, one starting frame. Need the same face and the same jacket to survive across a batch of shots? That is Veo 3.1 Fast reference-to-video.

What an 8-second clip costs, everywhere

Every serious provider of this model bills by the second, us included. So the only honest comparison is a fixed configuration — here it is at 720p with audio on, checked on August 26, 2026:

API providerBillingRate (720p, audio on)8-second clipvs. us
E2X (current)per second$0.01/s ×1.5 audio$0.12—
Google Gemini APIper second$0.10/s$0.806.7× our price
fal.aiper second$0.15/s$1.2010× our price
Replicateper second$0.15/s$1.2010× our price
E2X list (pre-discount)per second$0.10/s ×1.5$1.20our undiscounted rate

Read the last row against fal.ai and Replicate. Our list price for an 8-second clip is $1.20 — exactly what they charge. What you pay today is a tenth of that, and still 6.7× below Google's own API for the identical request.

Budget around one nuance: 720p → 1080p is free on fal.ai and Replicate, while Google charges more per second for it. Resolution is a multiplier on our side too, so pull the live figure from this model's llms.txt before committing to a 1080p pipeline.

Prices and product terms can change. Check each provider's current pricing before making a purchasing decision. Google Batch or Flex pricing is not directly equivalent to a standard on-demand API request because scheduling, availability, and processing conditions differ; it is therefore excluded from this comparison. This is a scoped comparison, not a claim that E2X is the world's cheapest option in every configuration.

The audio track is the headline

Most text-to-video models hand you a silent MP4 and leave the sound to you. Veo 3.1 Fast doesn't. Google's API generates audio natively and you cannot switch it off there — synchronised dialogue, foley and room tone, produced alongside the picture and locked to the frames. Google's own documented example prompt contains spoken lines.

That changes what one call is worth. Write "a fishmonger slams a crate down and shouts fresh in, five minutes ago" and you get the slam and the shout on the right frames — not a mime you have to score later. For social cuts and animatics, an entire post step disappears.

What you do get to pick:

  • Duration: 4, 6 or 8 seconds. This is the only one of our three Veo 3.1 Fast endpoints where that choice is genuinely open. A 4-second 720p clip with audio runs $0.06; 6 seconds is $0.09.
  • Resolution: 720p (our default), 1080p or 4k, at 24fps. 1080p and 4k are 8-second-only — Google's constraint, not ours.
  • Aspect ratio: 16:9 (default) or 9:16. No square, no 4:5.
  • has_sound, defaulted to true, driving the ×1.5 multiplier.

Eight seconds is the ceiling for one clip. Sequences are multiple calls.

API request: prompt in, MP4 out

Make a key in the dashboard, keep it server-side, post the job. Only prompt is required.

curl -X POST https://api.e2x.ai/v1/jobs/submit \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/veo-3.1-fast/text-to-video",
    "input": {
      "prompt": "Handheld shot down a rain-slick Osaka alley at night. A ramen vendor lifts the noren, steam pours out, and he calls to a passing cyclist: \"Last bowl of the night!\" Neon reflections shiver in the puddles.",
      "aspect_ratio": "16:9",
      "resolution": "720p",
      "duration": "8",
      "has_sound": "true"
    }
  }'

We return a job ID. Poll it, or don't:

const headers = {
  Authorization: `Bearer ${process.env.E2X_API_KEY}`,
  "Content-Type": "application/json",
};

const submitted = await fetch("https://api.e2x.ai/v1/jobs/submit", {
  method: "POST",
  headers,
  body: JSON.stringify({
    model: "google/veo-3.1-fast/text-to-video",
    input: {
      prompt:
        "Slow push-in on a lighthouse keeper writing in a logbook. Wind rattles the glass. She murmurs, \"Third night, still no signal.\" Lamp sweeps past her face every four seconds.",
      aspect_ratio: "9:16",
      resolution: "720p",
      duration: "6",
      has_sound: "true",
      seed: 41205,
    },
  }),
}).then((r) => r.json());

const jobId = submitted.data.jobId;

while (true) {
  const job = await fetch(`https://api.e2x.ai/v1/jobs/${jobId}`, { headers })
    .then((r) => r.json());

  if (job.data.status === "completed") {
    console.log(job.data.outputs[0].url);
    break;
  }
  if (job.data.status === "failed") {
    throw new Error(job.data.error?.message || "Generation failed");
  }
  await new Promise((r) => setTimeout(r, 5000));
}

Statuses run pending → processing → completed, or land on failed / cancelled. Send a webhookUrl instead and we'll call you. Budget about four minutes per job; Google publishes a floor of 11 seconds and a peak-hours ceiling of 6 minutes, so build for the ceiling.

One schema trap: duration and resolution are strings. "8", not 8.

Choosing between our Veo 3.1 Fast endpoints

Same weights, same per-second price, three different front doors. The input you already have decides it:

Endpoint8s clip, 720p + audioPick it when
Veo 3.1 Fast text-to-video$0.12You have words and nothing else — and you want 4s or 6s
Veo 3.1 Fast image-to-video$0.12One still exists and should become the first frame
Veo 3.1 Fast reference-to-video$0.12A person, product or place must stay identical across shots

Price won't break the tie, so follow the brief. If a client already signed off on a key visual, sending it as a starting frame beats re-describing it in prose. If six clips have to look like one campaign, load up to three references — just know reference-to-video forces 8 seconds, no exceptions. This endpoint keeps the shorter durations, which makes it the cheapest way to fill a storyboard with rough motion. The rest of what we run sits in the text-to-video category and the full model catalog.

Prompts that earn the sound

Treat the prompt as a shot list with a sound mix attached. Four things move the needle more than adjectives:

Name the camera move. "Locked-off wide", "slow dolly in", "handheld follow" — the model respects these. Vague framing gives you drifting, uncommitted motion.

Write dialogue in quotes. Short lines, one or two per clip. Eight seconds is not long. Paraphrased intent ("he says something encouraging") produces mumbled filler.

Ask for the ambience. Rain on tin, a fridge hum, distant traffic. The model invents a bed if you don't name one, and it may not match your edit.

Say what the light is doing. Overcast noon, sodium streetlight, one practical lamp. Lighting decides whether a Fast-tier clip reads as deliberate or as a render.

English is fully supported; Google hasn't evaluated other languages here. Write the dialogue in whatever language you want spoken and expect more variance outside English.

What to check before you ship

Download within 48 hours. Google keeps the file for two days — their wording is that you must download your video within 2 days of generation. Copy every output to your own bucket in the same job that reads outputs[0].url. A CDN link is not an archive.

Every frame carries SynthID. Google's imperceptible watermark is applied to all Veo output and is checkable on their verification platform. Nobody can turn it off, us included.

Region rules on people. In the EU, UK, Switzerland and MENA, personGeneration is restricted to allow_adult.

Frequently asked questions

How much does Veo 3.1 Fast text-to-video cost on E2X?

We bill per second: $0.01 per second of output, multiplied by 1.5 when audio is enabled. An 8-second 720p clip with sound is $0.12, a 6-second one is $0.09, and 4 seconds is $0.06. Resolution above 720p carries its own multiplier.

Is E2X really ten times cheaper than fal.ai for this model?

At the configuration and date we checked — 8 seconds, 720p, audio on, August 26, 2026 — yes: $0.12 with us against $1.20 there. Our own undiscounted list price is that same $1.20. The gap is a discount, not a different product.

Can I turn the audio off?

Our schema exposes has_sound and we default it to true. Google's own Gemini API does not offer an off switch at all, so audio is best treated as part of what this model is — the same holds on our image-to-video endpoint. It's also the reason for the ×1.5 price multiplier.

How long can one generated clip be?

Eight seconds, in a single call. You can request 4 or 6 on this endpoint, but 1080p and 4k require the full 8. Longer sequences mean stitching several jobs together yourself, or holding a look steady with reference-to-video.

Which aspect ratios and resolutions are supported?

16:9 and 9:16 only, at 24fps, in 720p, 1080p or 4k. We default to 16:9 and 720p. There is no square or 4:5 option, so crop in post if you need one.

Do the videos have a watermark?

Yes — every Veo output is watermarked with SynthID, Google's imperceptible provenance marker, and it can be verified through their platform. No provider exposes a parameter to disable it.

How long does a generation take?

We budget around four minutes per job. Google publishes a minimum of 11 seconds and a maximum of 6 minutes at peak, so use a webhook rather than a tight polling loop for bulk work.