Reference images are not one feature
Every video model here has a slot for reference images and the button looks the same everywhere. What happens next does not, and in two cases what happens is that your images are ignored.
Start with the one that costs you money
Grok Imagine in text-to-video mode accepts reference images and discards them. The uploader fills, the counter reads "2/7", the generation runs, you are charged, and a perfectly plausible video comes back. The images were never sent. We only found it by pulling the provider's request log, where the payload had no image field at all.
If you want Grok to use your references, you must be in image-to-video mode. Same model, same pictures, completely different outcome.
This is the reason to check the mode indicator before you spend anything. A video that looks fine is not evidence your inputs were used.
The number of images matters more than the model
This is the part almost nobody knows, and we published the wrong version of it ourselves until we tested it properly.
With Grok Imagine, in either version:
| images attached | what happens |
|---|---|
| one | that image becomes the opening frame, and the video takes its shape. Your aspect ratio setting is ignored. |
| two or more | no opening frame is locked. All images act as references, Grok composes a new shot, and your aspect ratio is honoured. |
So a single square image with 16:9 selected gives you a square video. Add a second reference and the same request gives you widescreen.
That surprises people, and it is not a fault. One image reads as "animate this picture", so the model preserves it. Several images read as "compose something from these", so nothing is preserved and the frame is yours to shape.
If your aspect ratio is being ignored, add a second reference image. That is the whole fix.
Measured 2026-08-22 across nine generations on both Grok 1.0 and 1.5, using square and portrait sources against landscape requests. The two versions behave identically. An earlier version of this article said they differed; they do not, and the difference we saw then was the image count, not the model.
The three behaviours
Conditioning. The model reads your images for palette, subject and mood, then composes a new shot. Your image never appears directly. This is what you want for a sequence of scenes that share a look without repeating each other.
First-frame anchoring. Your image becomes frame one and is animated forward. Useful when you have an exact opening shot in mind. Awkward across a sequence, because every scene built from that image opens identically. On Grok this happens only with a single image.
Subject featuring. The model keeps the things in your image but rebuilds the scene around them, often relit. You get your buildings back under someone else's sky.
What each model does
| Model | Reference behaviour | Notes |
|---|---|---|
| Grok Imagine 1.0 and 1.5 | One image anchors the frame, two or more condition | Identical to each other. Measured 2026-08-22 |
| Grok Imagine, text mode | Silently discards references | Switch to image-to-video |
| Seedance family | Conditioning, dedicated reference parameter | No first frame at all, by design |
| Wan 2.6 Reference | Subject featuring, relights the scene | |
| Happy Horse 1.1 | Close reconstruction of the reference | In image mode it derives the aspect from your frame and ignores the ratio you pick. Measured 2026-08-08 |
| Kling family | Single image, treated as a frame |
How many images each model takes
| Model | Limit |
|---|---|
| Grok Imagine 1.0 and 1.5 | 7 |
| Seedance family | 9 |
| Happy Horse 1.1 Reference | 9 |
| Wan 2.6 Reference | 5 |
| Kling family | 1 |
These capacities were measured on 2026-08-18 and have not been re-tested since. The Grok behaviour above was re-measured on 2026-08-22. We are dating them separately on purpose, because provider behaviour changes and an undated number is the thing that misleads people later.
Choosing
A music video or any sequence of scenes? Use a conditioning model with two or more references. Grok Imagine 1.5 is the cheapest option and varies most between runs, which is what a twenty scene edit needs.
An exact opening image you want on screen? Give a model a single image. On Grok that locks the frame, and remember it will also take that image's shape.
Your product or character in a scene you did not shoot? Wan 2.6 Reference features subjects rather than copying compositions.
One caution on near-duplicates. A model that reconstructs your reference approximately is harder to work with than one that copies it exactly. Exact repeats read as a deliberate motif. Near-misses read as a mistake.
Why this article changed
An earlier version stated that Grok 1.0 anchored the first frame while Grok 1.5 composed a new shot, as though the versions differed. They do not. Both do both, depending on how many images you attach.
We repeat that here rather than quietly editing it, because the mistake is instructive: the true statement was never "Grok ignores the aspect ratio", it was "Grok ignores the aspect ratio when only one image is provided". Drop the condition and a correct claim becomes a wrong one. That has now caught us three separate times on this same behaviour.
Everything above is measured on the models we host, with the date attached. Where we have not re-tested something, it says so.
See the comparison
Every reference behaviour described here is playable side by side, with the prompt and both reference images available so you can run the same test yourself: The Arena.