NewNew models in the catalogue
Musevate

Discovery

AI video models that take reference images

A reference image is the answer to the hardest problem in AI video: keeping the same face, the same product or the same costume across more than one shot. You give the model a picture of the subject and describe the scene, and the subject survives into it.

That is not the same as animating a still. Image-to-video starts from your picture and moves on from it; a reference is carried through a scene the model generates around it. These 9 on Musevate accept one, starting at 14 credits for five seconds.

9 models, read from the live registry when this page loaded.

AI video models that accept reference images, cheapest first, models with an example clip listed before those without
#Model5 secondsLongestResolutionExample
1
MiniMax H3 Max

H3 retrained for closer prompt adherence.

19 cr10s1080p, 768p, 480pWatch
2
Gemini Omni Flash 1.1

Faster, cheaper, and now takes reference images.

25 cr10s4k, 1080p, 720p, 360pWatch
3
Wan 3.0 Reference

Wan 3.0 driven by reference images rather than one frame.

30 cr10s1080p, 720p, 480pWatch
4
MiniMax H3

The most cost efficient way to animate an image.

14 cr10s4k, 1440p, 768p, 480pNot yet
5
Gemini Omni Flash

Google's Gemini Omni Flash, built for speed.

35 cr10s720pNot yet
635 cr10s1080p, 720pNot yet
7
Kling O3 Pro

The Kling O3 professional tier.

40 cr15s720pNot yet
8
Seedance 2.0 Mini

The smallest Seedance 2.0 tier.

45 cr15s720p, 480pNot yet
9
Seedance 2.5 Reference

Seedance 2.5 driven by reference material.

125 cr30s1080p, 720p, 480pNot yet

Questions

What is a reference image in AI video?

A picture you supply that the model has to keep: a person's face, a product, a costume, a style. It is not the first frame — that is image-to-video, where the clip starts from your picture and moves on. A reference is carried through the whole generation, so the same character can appear in shot after shot and still be the same character.

Which models accept reference images?

9 of the 88 on Musevate. It is a newer capability than text-to-video or image-to-video and the catalogue reflects that; a model is listed here only when the endpoint that takes references is one the router has verified and will actually send to.

How is this different from image-to-video?

Image-to-video takes one still and animates it: the video begins as your picture. A reference is used differently — the model generates a new scene and keeps the subject you gave it inside that scene. If you want your photograph to move, that is image-to-video. If you want your subject to appear in a shot you are describing, that is a reference.

Can I use more than one reference?

On most of these, yes — a character and a product, or a subject and a style. How many each accepts comes from its own schema and the studio will not offer more than the model takes, because a reference that is dropped silently is worse than one that was never offered.

Try any of them on one account

Musevate routes each generation to the strongest compatible model, and every model above is available without a separate subscription.

Create your first video
AI Video Models That Take Reference Images · Musevate