Loading...
Loading...
Text to video turns a written prompt into a short moving clip. Nothing but the description goes in, so the model has to invent the subject, the camera behaviour and the motion over time. That makes it the most open-ended capability in the catalog and the one where the exact wording of a prompt changes the result the most.
Useful prompts describe the shot the way a director would: subject, action, camera move, setting and lighting. Clips are short by design, so plan a single beat per generation and cut several together rather than expecting one prompt to carry a whole scene. Video jobs also run noticeably longer than image jobs and cost more per request, which is a good reason to lock the look in with a cheap image model first and then bring the winning description across.
E2X currently serves this category with two Google models, Veo 3.1 Fast and Omni Flash. Both are called like every other model here: submit a model slug and an input object to the jobs endpoint, take the job id that comes back, and poll it until the clip is ready. Billing is per request from a prepaid balance, with the price shown on each model's catalog entry before you call it.
Short. These models produce single-shot clips rather than full scenes, so longer sequences are built by generating several clips and editing them together.
Video generation renders many frames that have to stay consistent with each other, so the job simply runs for longer. Submit it, then poll the job id instead of holding a request open.
Yes, but through a different category. Image to video animates a single frame, and start and end frame to video interpolates between two of them.
One request per clip, charged against your prepaid balance at the per-request price listed on the model you pick.