Two ways to generate video
Text-to-Video (T2V) - describe a scene, get a video. No reference needed. Best for completely original ideas where you have no visual yet.
Image-to-Video (I2V) - start from a still image, animate it. The image controls the look; your prompt controls the motion. Best when you already have album art or a character design you want to bring to life.
Grok Imagine 1.5 is a separate model in the picker, listed alongside Grok Imagine. Same price as 1.0 and the cheapest video on the platform. The difference that matters is how each treats reference images: 1.0 turns your first reference into the opening frame, while 1.5 uses your references as guidance and composes a new shot. If you want twenty scenes that do not all open the same way, use 1.5.
Important - Grok I2V: when using Grok for video with a reference image, a text prompt is NOT required. Grok will animate the uploaded image on its own. You can still add a prompt to steer the motion if you want more control, but it is optional in that flow. Other models (Kling, Wan, Seedance) still benefit from a motion prompt even in I2V mode.
For album promo, music videos, and most creative work, I2V gives you far more control - generate the image first in Image Studio, then animate it here.
Your first video
- Open Video Studio.
- Choose T2V or I2V at the top. If I2V, upload or select a reference image from your gallery.
- Pick a model - see the cheatsheet below.
- Write a motion prompt. For I2V this should describe what moves, not what is in the image:
"camera slowly dollies forward, hair ripples in the wind, neon sign flickers". - Set duration (most models offer 4-10 seconds).
- Click Generate. Videos take longer than images - usually 60-180 seconds.
Model cheatsheet
| Model | Mode | Best for | Duration |
|---|---|---|---|
| Runway Gen-4 Turbo | T2V / I2V | Cinematic AI video with consistent characters across scenes and plain-language camera control | 5-10s |
| VEO 3.1 Quality | T2V / I2V (first and last frame) | Highest quality, cinematic results | 4, 6 or 8s |
| VEO 3.1 Fast | T2V / I2V (first and last frame) | Same family, quicker and cheaper | 4, 6 or 8s |
| VEO 3.1 Lite | T2V / I2V (first and last frame) | The cheapest way onto Veo, built for volume | 4, 6 or 8s |
| Kling 3.0 Standard | T2V / I2V + audio | Newest Kling, 720p with native audio | 3-15s |
| Kling 3.0 Pro | T2V / I2V + audio | Newest Kling, 1080p with native audio | 3-15s |
| Kling 3.0 4K | T2V / I2V + audio | Maximum fidelity Kling tier | 3-15s |
| Kling 2.6 | T2V / I2V + audio | Strong motion, native audio | 5-10s |
| Wan 3.0 | T2V / I2V + audio | Newest Wan, native audio, and the only model that bills honestly down to 2 seconds | 2-30s |
| Wan 3.0 Pro | T2V / I2V + audio | The 2K and 4K tier, and cheaper per second than standard Wan 3.0 at 1080p | 2-30s |
| Wan 2.7 | T2V / I2V | Cost-effective, decent quality | 4-5s |
| Wan 2.6 Flash | I2V / Reference | Cheapest video on the platform, optional soundtrack | 2-10s |
| Happy Horse 1.1 | T2V / I2V / Reference | Up to 9 reference images, always arrives with a soundtrack | 3-15s |
| Seedance 2 | T2V / I2V + audio | Stylized motion, artistic output, and native audio that costs no more than silence | 4-8s |
| Seedance 2.5 | T2V / I2V + audio | The premium Seedance tier, native audio, and the only one that runs to 30s | 4-30s |
| Seedance 1.5 Pro | T2V / I2V + audio | Budget tier - cinematic at a fraction of Seedance 2 cost, optional native audio, up to 1080p | 4-14s |
Live pricing in the model picker. Videos are significantly more expensive than images - budget accordingly.
Short clips are not rounded up on Wan 3.0. Most video models bill a minimum length, so asking for 2 seconds quietly costs you the same as 5. We measured Wan 3.0 and it bills by the second all the way down to its 2 second floor, which makes it the model to reach for when you need a quick cutaway, a loop, or a single beat of motion rather than a whole scene.
VEO 3.1 now offers 4, 6 or 8 seconds, and all three cost the same. Veo is priced per video rather than per second, so a 4 second clip costs exactly what an 8 second one does. We measured this rather than assuming it. That makes the shorter options useful for getting a usable take quickly, or for a cutaway where 8 seconds would be too long, but there is no money saved by picking 4 over 8. If cost per second is what matters to you, Wan 3.0 bills by the second and is the better tool.
If you want 1080p, pick Wan 3.0 Pro rather than standard Wan 3.0. Pro is the higher tier in every other respect, and at 1080p it also happens to cost less per second. There is no case for choosing standard Wan 3.0 at 1080p.
Reference video: put several subjects in one shot. Happy Horse 1.1 and Wan 2.6 Flash both have a Reference mode that is different from image-to-video. In I2V your image becomes the literal first frame of the clip. In Reference mode your images are subjects to feature - a person, a product, a prop - and the model composes a scene containing them. How literally each model does this varies, and we measured it. Wan 2.6 Flash genuinely recomposes: it keeps your subjects and re-lights the whole scene. Happy Horse 1.1 stays much closer to the source, and its first frame often looks like a redrawn version of your reference rather than a new shot, so reusing one image across several scenes produces near-duplicates. Happy Horse takes up to 9 references and Wan 2.6 Flash takes up to 5, counting the main image slot. Attach anything in the reference strip and the model switches to Reference mode automatically.
Which models let you choose the aspect ratio. This differs per model and it is not guesswork - we test it:
- Grok I2V honours it and overrides your image. Ask for 16:9 with a portrait photo and you get a landscape video.
- Happy Horse I2V ignores it - the clip always takes the shape of your source image, so the dropdown is hidden in that mode on purpose. Use its Reference mode instead if you need to pick a shape.
- Happy Horse Reference and text-to-video both honour it, including vertical.
If the aspect dropdown disappears when you attach an image, that is deliberate - it means the model would have ignored the setting.
Happy Horse 1.1 always produces a soundtrack. There is no silent option; the model scores every clip itself. If you are cutting it into an edit that already has audio, mute the track in your editor. Wan 2.6 Flash is the opposite - audio is off by default, and turning it on doubles the per-second price, so leave it off unless you want it.
Audio is free on the Seedance 2.x family, and it is the reason we render talking characters there. Across Seedance 2.0 Mini, 2.0 Fast, 2.0 and 2.5 the "with audio" rate is identical to the silent rate, so a speaking clip costs exactly what a mute one does. Almost nothing else prices this way - Wan 2.6 Flash and Seedance 1.5 Pro both roughly double for speech. Rendering silent and adding a voice afterwards means paying for a second pass, so when the character needs to talk, ask the engine to speak in the first place.
Happy Horse and Wan 2.6 Flash run on Alibaba rather than our main supplier. That is deliberate: when the usual provider has an outage, these keep working. Wan 2.6 Flash is the cheapest of the two, though Grok Imagine is cheaper still - Wan has no 480p tier, so its floor is 720p.
Extend a Runway clip. A finished Runway Gen-4 Turbo 720p clip can be continued. After it generates, click Extend, then either keep extending in 720p to chain it further, or finalize in 1080p (a 1080p extension is terminal and cannot be extended again). 1080p source clips cannot be extended, so generate at 720p if you want to keep building a chain. If an extend occasionally returns a temporary upstream error, just try again - you are not charged for a failed attempt.
What makes a good motion prompt
For I2V:
- Describe the camera motion:
"slow dolly in","crane shot rising","static camera, handheld shake" - Describe subject motion:
"she turns her head and smiles","waves crash against the rocks","neon flickers" - Keep it short - 1 to 2 sentences is plenty
For T2V:
- Full scene description + motion:
"A biker rides through a rainy Tokyo street at night, neon reflections, slow tracking shot from behind"
Bad prompts:
- Describing the image details the model already sees (I2V)
- Trying to cram multiple camera moves in one shot
- Asking for edits or cuts - these models generate single shots
Image-to-video quality rules
The input image is the single biggest quality factor for I2V:
- High resolution (at least 1024x1024) - low-res inputs produce blurry video
- Clean subject - dense backgrounds confuse motion models
- One clear focus - crowd scenes animate poorly
- Readable lighting - harsh shadows can cause weird motion artifacts
If your generated image is going to become video, generate it with that end goal in mind - keep the composition simpler than you would for a stand-alone image.
Outputs and storage
Videos save to your gallery automatically. As soon as a generation finishes it is copied to Cloudflare R2, our permanent CDN, and a gallery entry is created for you. You do not have to click anything, and closing the tab mid-render does not lose the video - it will be waiting in your gallery. There is still a Save to Gallery button on the result for when you want to be sure.
This matters because the provider's own link expires within about 24 hours. The copy on R2 does not expire, and that is the one your gallery points at.
Note that images work differently - those are not saved automatically, and you do need to click Save on each one before leaving the Image Studio.
From the Gallery you can:
- Download (blob download, no session navigation)
- Share to the community feed
- Use as a share video for one of your tracks
Known issues
- Infinitalk lip sync is currently 100% offline. Hidden from the model picker until the upstream provider restores service. This matters less than it used to: a talking character is normally a native-audio render now (Grok, or the Seedance 2.x tiers in UGC Studio), where the engine speaks the line itself and charges no more than a silent clip. Lip sync is for putting a specific recording on a face.
- Some image editing models (Kontext Pro/Flex, Seedream V4 Edit, Qwen Edit) are hidden from I2V because results are unreliable.
- Grok Imagine has intermittent upstream outages - the frontend retries automatically, but if you hit a hard failure, try VEO 3.1 or Kling.
For deeper troubleshooting see generation-failures.