← All posts

Why AI UGC ads look fake, and what fixes it

The four tells that give AI UGC ads away, in the order viewers notice them, plus the prompt habits and the order of operations that remove each one.

Most AI UGC ads fail in the same four ways, and none of them are the model's fault. People reach for an AI UGC video generator, get something that is almost right, and cannot say why it feels wrong. Viewers can. They just do not articulate it, they scroll.

I build SimpliGen, which makes these clips, so I read the support threads where they go wrong. Here is what actually gives them away, in the order people notice.

What gives an AI UGC ad away?

Four things, and they arrive in a fixed order: the face, then the motion, then the face again, then the packaging. Each has a different cause, and only one of them is a model limitation you cannot prompt your way out of.

What the viewer noticesWhy it happensWhat to do
Skin too smooth, eyes slightly too largeThe model is flattering the faceJudge the still frame before animating, and push back with negative prompts
Motion that drifts towards slow motionNo phone video moves like thatKeep the clip short and the action to one gesture
The face at second eight is not the face at second oneIdentity drift in current image to video modelsLock the character, keep clips short, do not fix it in the prompt
Label text morphing into near-miss gibberishText is the hardest thing in the frameCheck every frame, and prefer shots that do not linger on the label

Why does the face change halfway through?

Because current image to video models do not hold an identity perfectly across a clip, and it gets worse the longer you run. Nothing tips a viewer off faster than a person who becomes their own cousin by the end of five seconds.

This is the one you cannot prompt away, so you design around it. Start from a locked character rather than a fresh face each time: Character Studio locks in a base portrait and every image and video after that reuses the same face, which removes the drift between clips even though it cannot fully remove drift inside one. Keep clips short. Wan's own repository ships a dedicated character animation variant precisely because general video generation does not solve this by itself.

Why does "hyperrealistic, 8k, cinematic" make it worse?

Because those words push the model towards CGI, which is the opposite of what a UGC clip needs. The words people add to make it look real are the words that make it look fake.

This is documented rather than folklore: our own quality guide says plainly do not stack quality keywords such as "hyperrealistic, 8k, ultra detailed, unreal engine", because they "often push output toward CGI looks rather than realism". The same page suggests going the other way and putting "cartoon, anime, 3d render, cgi, plastic skin, doll" into your negative prompt. It also notes that very long prompts dilute, since models latch onto the first concrete instructions.

A UGC ad is trying to look like a phone video shot in a kitchen. Every word that reaches for production value moves it away from that.

What should the prompt say instead?

Name the person, point at the reference, describe one thing they do. The first mistake almost everyone makes is writing the prompt like a search query: "woman using face cream" gives you something that looks like a stock photo which learned to move.

The specific version is boring to write and works far better. Say who the subject is and which reference picture they are, then keep the action to one natural gesture and put any spoken line in quotation marks. Our community arrived at this on its own and now trades structured prompt templates that open by defining the subject and pointing at the reference image, because the loose version does not survive contact with a video model.

Why does the product look wrong, especially the label?

Because text on packaging is the hardest thing in the frame, and it fails independently of everything else. Labels morph, brand names come out as near-miss gibberish, and it happens even when the lighting, the hands and the face are all convincing.

So treat the product as its own problem rather than a detail of the shot. Build it once in Product Studio so the same product carries across clips, then check every frame of any shot where the label is legible. A shot that lingers on your packaging is a harder shot than one where the product simply sits in someone's hand, and it is worth choosing the easier composition when the label is not the point.

What separates a clip you ship from one you bin?

Not the prompt. The order of operations, and it is the most reliable habit we have.

Get the still frame right first. Generate the image, iterate until the framing, the product and the face are what you want, and only then animate that frame. Judging a video by watching videos is slow and expensive; judging it as a photograph first is neither.

Then draft cheaply. Run low-step versions, pick the one whose motion you like, and re-render that exact seed at full steps: the same prompt, settings and seed produce the same result, so the draft you liked is reproducible rather than a lucky accident you cannot find again. Generating several at once with different seeds is the fastest way to find a keeper.

How many attempts does a usable clip take?

I do not know, and I would rather say so than publish a number I cannot stand behind. It depends entirely on the shot, and anyone quoting you a single figure is quoting an average of work that looks nothing like yours.

The more useful question is what each attempt costs. Generating locally on your own GPU, an attempt costs a few minutes and some electricity, so throwing most of them away is simply how the work goes. That is a different creative process from one where every attempt is metered, and it changes what you are willing to try. On cloud credits the arithmetic returns, since resolution, duration and steps drive what a generation costs, which is exactly why the cheap-drafts habit above is worth having in both modes. SimpliGen itself is a one-time licence, not a subscription.

For what a local attempt actually costs in minutes on real hardware, the timing numbers are in our post on local video.

When should you still hire a real creator?

When the ad depends on a real person's credibility rather than a face saying words. If the claim is "I used this for three months and here is my skin", a generated person cannot make that claim honestly, and the fix is not a better prompt.

The same line applies to using someone's likeness. Real people's faces are not raw material, and our content policy is explicit about likeness and deepfakes. Generated presenters are for the shots where the presenter is a device, not a witness. Knowing which of those you are making is most of the judgement in this work.

What should you do next?

Take your worst clip and check it against the four tells above, in order. Most of the time it is the second and fourth, the motion and the packaging, and both are fixable without touching the model.

Then change the order you work in: still frame first, one gesture, cheap drafts, then the full render of the seed you liked. If you want the underlying setup, UGC Studio uses the character and product you have already built so the same face and the same product carry across every clip.

Try it on your own PC

SimpliGen is a one-time purchase for Windows. Generate locally on your GPU or on our cloud, without touching a node graph.

See pricing