Documentation
Everything Musevate does, and how to get it right
The whole platform in one page: what each control does, how credits are spent, and the handful of habits that separate a clip you keep from one you regenerate. Read the last part even if you skip the rest — it is the difference between two attempts and ten.
What Musevate is#
One place to generate video with many different AI models, without having to know which one to use.
There are dozens of video models and they are good at different things. One is cheap and fast and cannot hold a face steady. Another is slow and expensive and does hands properly. Choosing between them normally means a separate account for each, a different interface for each, and enough experience to know which is which.
Musevate runs them all behind one account, measures what each one costs and what it can actually do, and picks the right one for the request you describe. You buy credits, a generation spends them, and the difference between raw compute and a model worth using is ours to carry.
You describe the shot; we choose the model
The field moves weekly. Models are retired, replaced and repriced, and work built around one name has to be rebuilt when that name goes. Say what you want on screen and leave the rest to us, which is the part we exist to do.
How a generation works#
Five steps, and knowing them explains almost every question support gets asked.
- You describe a shot and set the length, shape and quality.
- We quote a price in credits, before you commit. The quote reflects the model that will actually run.
- Credits are reserved the moment you press Generate, so two tabs cannot spend the same balance twice.
- The model runs. Anywhere from about twenty seconds to several minutes, depending on length and quality.
- The video lands in your gallery. If the model fails instead, the credits come back automatically.
You can close the tab
The generation runs on our servers, not in your browser. Closing the tab, losing signal or reloading the page does not cancel it and does not lose it. It will be in your generations when you come back.
The four ways to make a video#
Each starts from something different. Choosing the right one matters more than the prompt does.
| Mode | You provide | Best for |
|---|---|---|
| Text | A written description | A shot that does not exist yet — an idea, a mood, a scene you are inventing. |
| Image | One image, and optionally a closing frame | Animating something specific: a product, a photograph, artwork you already have. |
| References | Several images the prompt can refer to | Keeping a character, a product or a style consistent across shots. |
| Lip-sync | A photo of a person and a voice recording | Making a still portrait speak the words in the audio. |
Text to video
The most freedom and the least control. The model invents everything, so two runs of the same prompt give two different videos. Use it to explore, then move to Image once you have a frame you like.
Image to video
The uploaded image becomes the first frame and the prompt describes what happens next. This is the highest-control mode and the one most people should be using: you have already decided the composition, the subject and the colour, so the model only has to supply motion.
Where the chosen model supports it you can also give a closing frame — the shot then travels from the first image to the second. It is the difference between “pan right” and “end up here”.
The first frame decides almost everything
A soft, badly lit or low-resolution starting image produces a soft, badly lit video, every time, at any price. Fixing the image costs nothing. Regenerating around a bad one costs credits repeatedly and does not work.
References
Several images at once, and the prompt can point at them by position — @Image1, @Image2 and so on. “The woman from @Image1 holding the bottle from @Image2, walking through @Image3” is a shot a model can assemble from parts you control.
This is how you keep a face or a product recognisable between clips. Not every model supports it, and the studio only offers the mode when the selected model does.
Lip-sync
A portrait and an audio file, and the mouth follows the words. The face should be front-on, evenly lit and large in the frame. The audio should be clean speech — music underneath it confuses the timing.
You can upload a recording or have a voice generated for you. Both count as the same input.
The studio controls#
What each one changes, and which of them change the price.
| Control | Options | What it does to the result and the cost |
|---|---|---|
| Duration | From about 2 seconds upward, depending on the model | The biggest cost driver: most models bill by the second, so doubling the length roughly doubles the price. It does not make the shot better. |
| Aspect ratio | 16:9, 9:16, 1:1 | Landscape, vertical and square. Free to change. Decide it before generating rather than cropping after, which throws away half your resolution. |
| Resolution | 720p up to 4K, where the model offers it | Higher costs more and takes longer. It does not fix a weak composition; it renders the weak composition more sharply. |
| Quality | Fast, Quality, Cinema | Which class of model gets used. Fast is for finding out whether an idea works. Cinema is for the version you publish. |
| Sound | On, off, or unavailable | Some models generate audio, some never do, some always do. The toggle only appears where the model actually has one. |
A control that is not there is telling you something
The studio hides settings the selected model cannot honour rather than accepting them and quietly ignoring them. If an option disappears when you change the length or the mode, it is because the model that fits your request does not support it.
Credits, and what they buy#
One currency across every model, so the price of a shot does not depend on which one we route it to.
A generation’s price is quoted before you commit, and is worked out from the model that will run, the length, the resolution and whether sound is included. The same settings can cost a different number of credits on different days, because the cost of running each model moves and we re-measure it continuously rather than trusting the figure from the day a model was added.
What happens when a generation fails
Credits come back automatically. A model that errors, times out or refuses the request refunds in full, without you asking and usually within seconds. If a job hangs with no result at all, a sweep closes it and refunds it within half an hour.
What does not come back
A generation that ran successfully and produced a video you do not like is not refunded. The work was done and the compute was spent, the same way a print shop cannot un-print a page you decide you do not want. This is the most important sentence on this page, and it is why the rest of it exists: the way to not waste credits is to spend a few of them cheaply first.
Expiry and rollover
- Credits bought in a pack last a year from purchase.
- Trial credits are shorter-lived — the offer states how long.
- On a subscription, unused credits roll into the next month up to a ceiling that depends on the plan. Above that ceiling they stop accumulating, so a balance that keeps growing means the plan is bigger than the need.
- When you spend, the credits closest to expiring go first.
Pricing has the current packs, plans and ceilings.
Avatars#
A face you keep, so the same person appears across different videos.
An avatar is a saved image and a name — either uploaded or generated from a description. Once saved it can be used as the starting image or as a reference, which is how a character stays recognisable from one clip to the next.
Only faces you have the right to use
Do not upload a photograph of someone who has not agreed to it. This applies to public figures as much as to people you know. Accounts used to make a real person appear to say or do something they did not are closed, and Trust & safety sets out the rest.
Getting a video worth keeping#
Almost all wasted credits come from the same five mistakes. None of them are about the model.
1. One action, one camera move, one shot
Five seconds is about one thing happening. A prompt asking a character to stand, turn, walk to a window and look out will produce a smear of all four. Ask for the turn. Generate the walk separately.
Weaker
A woman stands up from her desk, walks across the office, opens the door and greets a colleague, cinematic, 4k, beautiful lighting, masterpiece
Stronger
A woman turns from her desk toward the window, slow dolly in, late afternoon light through venetian blinds
The second one is a shot. The first is a scene, and asking for a scene in five seconds is how you get four half-actions and no usable frame.
2. Say what is in the frame before you say how it looks
Models weight the start of a prompt most heavily. Lead with the subject and what it is doing, then where, then the treatment. A prompt that opens with cinematic, moody, anamorphic, 8k has spent its strongest position on adjectives.
3. Describe the camera as a movement
“Slow dolly in”, “handheld push”, “static wide”, “crane up” — these describe motion over time, which is what a video model produces. “Shot on ARRI Alexa” describes a look, and mostly gets interpreted as “slightly more filmic”.
One move per clip. Two competing moves usually produce neither.
4. Test at Fast, commit at Cinema
A short, fast, low-resolution generation tells you whether the idea works — whether the composition holds, whether the action reads, whether the model understood you. It costs a fraction of the final version. Get the prompt right there, then run it once at the settings you actually want.
Most people who spend a lot of credits without getting a usable clip did the opposite: went straight to the most expensive settings and repeated them.
5. Start from an image whenever you can
If the composition matters at all — a product, a person, a place — make or find the first frame and use Image mode. You remove every variable except motion, and motion is the part the model is actually good at.
The same prompt will not give the same video twice
This is how the models work, not a fault. If you get something close to right, generate it again before changing the words — the next run may simply be a better take of the same idea. Rewriting a prompt that was nearly working is how people talk themselves away from a result they were about to get.
What models are still bad at#
Said plainly, because finding out by spending credits is an expensive way to learn it.
| This | Why | What to do instead |
|---|---|---|
| Readable text | Letterforms are drawn, not typed. Signs, logos and captions come out as convincing-looking nonsense. | Generate the shot clean and add text in an editor. |
| Hands doing something precise | Fingers are the classic failure. Holding is usually fine; manipulating, counting or gesturing is not. | Frame above the wrists, or keep hands still. |
| Faces far from the camera | A face across a room is a handful of pixels, and the model invents the rest between frames. | Come closer, or accept that the face will drift. |
| Several people interacting | Identities swap, limbs merge, eyelines miss. Two is hard, three is a lottery. | One subject per shot. Cut between them. |
| Exact counts | “Five birds” is a suggestion, not an instruction. | Ask for “a few”, or crop so the number is not visible. |
| Long unbroken action | Coherence degrades with time. The last second of a long clip is the least reliable part of it. | Shorter clips, joined in an edit. |
Before you spend a lot#
Thirty seconds of checking, against a generation that cannot be undone.
- Is this one action, or have I described a scene?
- Does the prompt start with the subject rather than the adjectives?
- Is there exactly one camera move?
- Have I run it cheaply once, at Fast and short, to see whether the idea holds?
- If composition matters, am I starting from an image instead of from text?
- Is the aspect ratio the one I will actually publish in?
- Am I asking for something on the list above that models cannot do yet?
When something goes wrong#
What each situation means, and whether it costs you anything.
| What you see | What it means | Credits |
|---|---|---|
| Failed, with a message | The model refused or errored. Usually the prompt hit a content limit, or an input was not something that model accepts. | Refunded automatically. |
| Stuck on processing | Long generations genuinely take minutes. If nothing has happened after half an hour, a sweep closes it. | Refunded when it closes. |
| Not enough credits | The quote exceeded your balance. Nothing was started. | Nothing spent. |
| A setting is missing | The model that fits your request does not support it. | Nothing spent. |
| The video is fine but not what you wanted | The generation succeeded. This is the case the rest of this page exists to prevent. | Not refunded. |
Anything that does not fit those rows, write to us with the generation and we will look at it the same day.
Account, plans and billing#
What you can change yourself, and where.
- Balance and history — every credit in and out is itemised in Billing.
- Plans can be changed or cancelled at any time, from the same page. A cancelled plan runs to the end of the period you have paid for.
- Packs are one-off purchases with no subscription, and can be bought alongside a plan.
- Invoices are emailed, and are also in the billing portal.
- Deleting your account removes your generations and uploads. It cannot be undone, and it does not refund unused credits.
Your videos are yours to use commercially. The details are in the Terms, and what we do with your files is in the Privacy policy.
Where to go next#
Writing prompts that work→
The craft in more depth: subject, camera and light, and the phrasing models actually respond to.
Help centre→
The questions support gets asked most, answered directly.
The model catalogue→
What each model is good at, with an example of its output.
Examples→
Real generations, with the prompt and settings that made them.
Pricing→
Current packs, plans, rollover ceilings and what a credit buys.
Talk to a person→
Anything this page did not answer. Same-day reply.

