Discovery
AI video models that take reference images
A reference image is the answer to the hardest problem in AI video: keeping the same face, the same product or the same costume across more than one shot. You give the model a picture of the subject and describe the scene, and the subject survives into it.
That is not the same as animating a still. Image-to-video starts from your picture and moves on from it; a reference is carried through a scene the model generates around it. These 9 on Musevate accept one, starting at 14 credits for five seconds.
9 models, read from the live registry when this page loaded.
| # | Model | 5 seconds | Longest | Resolution | Example |
|---|---|---|---|---|---|
| 1 | MiniMax H3 Max H3 retrained for closer prompt adherence. | 19 cr | 10s | 1080p, 768p, 480p | Watch |
| 2 | Gemini Omni Flash 1.1 Faster, cheaper, and now takes reference images. | 25 cr | 10s | 4k, 1080p, 720p, 360p | Watch |
| 3 | Wan 3.0 Reference Wan 3.0 driven by reference images rather than one frame. | 30 cr | 10s | 1080p, 720p, 480p | Watch |
| 4 | MiniMax H3 The most cost efficient way to animate an image. | 14 cr | 10s | 4k, 1440p, 768p, 480p | Not yet |
| 5 | Gemini Omni Flash Google's Gemini Omni Flash, built for speed. | 35 cr | 10s | 720p | Not yet |
| 6 | 35 cr | 10s | 1080p, 720p | Not yet | |
| 7 | Kling O3 Pro The Kling O3 professional tier. | 40 cr | 15s | 720p | Not yet |
| 8 | Seedance 2.0 Mini The smallest Seedance 2.0 tier. | 45 cr | 15s | 720p, 480p | Not yet |
| 9 | Seedance 2.5 Reference Seedance 2.5 driven by reference material. | 125 cr | 30s | 1080p, 720p, 480p | Not yet |
Questions
What is a reference image in AI video?
A picture you supply that the model has to keep: a person's face, a product, a costume, a style. It is not the first frame — that is image-to-video, where the clip starts from your picture and moves on. A reference is carried through the whole generation, so the same character can appear in shot after shot and still be the same character.
Which models accept reference images?
9 of the 88 on Musevate. It is a newer capability than text-to-video or image-to-video and the catalogue reflects that; a model is listed here only when the endpoint that takes references is one the router has verified and will actually send to.
How is this different from image-to-video?
Image-to-video takes one still and animates it: the video begins as your picture. A reference is used differently — the model generates a new scene and keeps the subject you gave it inside that scene. If you want your photograph to move, that is image-to-video. If you want your subject to appear in a shot you are describing, that is a reference.
Can I use more than one reference?
On most of these, yes — a character and a product, or a subject and a style. How many each accepts comes from its own schema and the studio will not offer more than the model takes, because a reference that is dropped silently is worse than one that was never offered.
Try any of them on one account
Musevate routes each generation to the strongest compatible model, and every model above is available without a separate subscription.
Create your first video
