We tested 15 AI video models on the same face. Here is what each one is for
We ran fifteen image-to-video models through the same test: one start image, one motion prompt, one scene. Only the model changed. What follows is what came out, including the three that turned out to be unusable for the job, and what each one is actually good for.
Every number here is measured, not quoted from a landing page. Prices are the provider's own, converted at one credit per cent.
The two things that decided it
Almost every model produces something watchable at thumbnail size. The differences only appear when you look closely, and they appear in two places.
Skin texture survives, or it does not
A generated face carries pores, fine hair and small blemishes. The video model re-renders that face frame by frame, and most of them quietly smooth it away. At thumbnail size you will not see it. At full screen on a phone, the viewer does, and reads it as fake without being able to say why.
The mouth opens when it should not
If your video has no voiceover, a character whose lips move is the single most obvious tell there is. Several models add speech on their own, because a person filmed head-on usually talks. Telling them not to in the prompt does not reliably work.
The ranking, and what each model is for
There is no single best model, there is a best model per job. Here is the one we reach for in each case.
Kling 3 Pro, for silent UGC with a face
Our default. It keeps the grain of the source image, and it was the only model in the test that kept the mouth closed for the whole shot. It outputs 1176x1764, the highest resolution of the group, and follows the aspect ratio of your start image instead of reframing on its own.
Weakness: it generates an audio track by default, which you have to strip if your soundtrack comes from the ad platform. Its predecessor, Kling 2.5, is a different animal entirely and smooths skin badly, so do not judge one by the other.
Veo 3.1 Fast, when the character must speak
The same quality of skin as Kling 3 Pro, and the model that got us the cleanest hands in the whole test, which is rare enough to note. It also generates native audio, and it opens the mouth even when you ask it not to.
That makes it wrong for a silent ad and exactly right for a talking object, a spoken tip or any format where the line is heard rather than read. Veo 3.1, the premium version, gave us no visible gain for two and a half times the price.
Kling 3 Standard, for animating a screenshot
When the subject is a user interface rather than a person, the job is not to invent motion but to preserve what is there. Push in slowly, keep the interface sharp, add nothing. Kling 3 Standard does that for a third less than the Pro, and skin fidelity is irrelevant on a screen recording.
Wan 3 Prime, the serious second option
Equal on resolution and close on skin. It drifts on identity, giving our subject a noticeably different hairstyle from her own start image, and costs two and a half times Kling 3 Pro. Worth trying when a scene does not work on Kling.
MiniMax H3 Max, for volume
Softer skin, and it outputs 768 pixels wide, which throws away detail you paid for upstream. At eight credits a second it is the sensible pick when you are producing many clips and no single one carries the campaign.
Watch the price: it currently advertises a promotional rate that expires, so the number you plan with should be the standard one.
Three models we had to rule out
These are not bad models. They are structurally wrong for vertical UGC with a human face, and it costs nothing to know that before you spend on them.
| Model | What happened | Cost of finding out |
|---|---|---|
| seedance-2.5 | Refused the request outright: its content policy blocks images that may contain likenesses of real people. The most expensive model in the catalogue cannot be used on a photorealistic face. | Nothing. Refused before generating. |
| flux-3 | Output 800x1088, and its default duration is auto, which gave us 15 seconds instead of 5. | A 435 credit clip we could not use. |
| gemini-omni-flash | Output 1280x720, in landscape, ignoring the vertical orientation of the start image. | One wasted run. |
The rule that matters more than the model
The most useful thing we learned is not in the ranking. Four separate observations pointed at the same principle: a video model animates what is already there, it does not create.
| What we saw | Why |
|---|---|
| Skin texture survives the animation | It is on the start image. |
| A hand raised to the mouth came out waxy, smooth and paler than the face | It was not on the start image. The model invented it. |
| A smile asked for in the motion prompt never appeared | The start image had a neutral face, and the expression never left it. |
| Asking for hands to stay out of frame did nothing | An absence cannot be requested. Describe what you want, not what you do not. |
The practical consequence: put everything that has to look real into the image prompt, and let the motion prompt only extend it. Once we moved the smile into the start image, it held for the whole shot on the first try.
What a five second clip costs
| Model | Credits per second | 5 s clip |
|---|---|---|
| kling-3-standard | 8.4 | 42 |
| kling-3-pro | 11.2 | 56 |
| veo-3.1-fast | 12.5 | 50 (billed 4 s) |
| minimax-h3-max | 8 | 40 |
| wan-3-prime | 28 | 140 |
| veo-3.1 | 30 | 120 (billed 4 s) |
| seedance-2.5 | 47.3 | 237 |
Two traps hide in that table. Veo only accepts durations of 4, 6 or 8 seconds, so asking for 5 gives you 4. And several of these models charge two to four times more at 1080p than at 720p, which is why the estimate in the app is computed from the exact parameters that will be sent rather than from a headline rate.
How we tested
One start image, generated once and reused for every model. One motion prompt, identical across all of them. Frames extracted at the same timestamp, cropped to the same region of the face, compared at the same magnification. Anything else compares two variables at once and proves nothing.
The honest limit: this is one scene, an indoor evening shot under a warm lamp. A ranking established on one scene does not generalise without checking, and we would not be surprised to see the order shift on daylight or on a non-human subject.
If you want the shorter version: our templates already carry the model that fits what each of them does, and the documentation explains how credits map to what a provider charges.