How to Make AI UGC Look Like Real iPhone Footage

Seedance produces beautiful, cinematic, polished output by default.
that polish is what triggers ad detection in viewers.
the pattern recognition system has been trained on years of polished advertising, and the moment a reel reads as produced, the brain files it as an ad and everything after plays against that categorization.
real phone footage looks different.
the framing is slightly off, the lighting isn't balanced, the skin has texture, the camera drifts, the environment has clutter.
every one of those signals "real person filming with their phone" rather than "production team filming an ad."
the whole job is pulling Seedance away from its polished defaults.
Skin Texture in Every Prompt
this one has more impact on viewer perception of authenticity than anything else.
every prompt with a human face needs this, no exceptions.
the exact language:
> "realistic skin texture, visible pores around nose and cheeks, natural slight unevenness, no filter quality"
without it, Seedance defaults to smooth, poreless, waxy skin that viewers identify as AI inside the first half-second.
the smooth rendering is a model bias toward what its training data labeled as attractive, and that skin doesn't exist under a real phone camera.
this single line forces the model to render the imperfections real skin actually has: pores, slight unevenness, natural shine in some areas and matte texture in others.
tuning it:
if the output still looks plastic, move this instruction earlier in the prompt and strip any competing beauty language.
if it overcorrects into leathery, soften to "subtle natural texture" and drop the unevenness mention.
Anti-Polish Language
beyond skin, the whole aesthetic needs pulling away from the cinematic defaults.
telling the model what the output should NOT be carries more weight than describing what it should.
the language:
> "handheld phone camera feel, casual slightly unsteady framing, filmed in a real environment, not a professional set, organic not studio quality, slight imperfections, authentic phone footage aesthetic"
without this, Seedance gives you:
- locked-off camera framing with no movement
- evenly balanced studio lighting
- composed scenes where everything is centered
- clean backgrounds with nothing accidentally in shot
every one of those reads as a production team to a viewer's pattern recognition.
this instruction stops the model from producing all 4 at once.
Lighting Gets Its Own Section
lighting carries more mood and authenticity information than almost any other variable.
burying it in a list of scene descriptors produces weaker output than giving it its own dedicated section.
the base structure:
> "soft warm window light from camera-left, casting natural shadows across the face, no harsh highlights, skin properly illuminated without overexposure, golden hour quality with realistic falloff"
every lighting description needs 4 things: direction, quality, shadow behavior, and skin illumination.
scenario variations:
morning bedroom:
> "soft cool morning light filtering through closed blinds, subtle blue cast in the shadows, warm skin tones contrasting against the cool ambient, no harsh edges"
evening kitchen:
> "warm yellow practical lighting from overhead, slight shadows under the eyes from the angle, no overexposure on the cheek highlights, realistic interior lighting falloff"
outdoor walking:
> "diffused overcast daylight, soft shadows wrapped around the face, no harsh sun angles, natural skin illumination without specular highlights"
pick one lighting condition per creator or series and name it identically every time.
inconsistent wording between generations produces visible drift across a series.
The Environment Needs Real Clutter
clean, organized backgrounds read as set design.
real spaces have things in them.
a charger cable on the nightstand, a glass with a smudge on it, a t-shirt over the back of a chair, mail stacked on the counter, these signal "real person in their real space" in ways that clean backgrounds can't.
specify it by scene:
bedroom:
> "messy unmade bed visible behind her, clothes scattered on the chair, charger cable dangling from the nightstand, water glass on the side table, makeup products scattered on the dresser"
kitchen:
> "dish rack visible behind her, wooden cabinets with slight wear, skincare items near the sink, coffee mug with ring stains on the counter, scattered mail on the side"
car:
> "slightly cluttered front seat, AirPods case on the dashboard, charging cable plugged into the dash, water bottle in the cup holder, slight dust on the steering wheel"
the specificity matters more than the mess.
"cluttered room" generates generic clutter.
naming the charger cable and the water glass generates a specific room that belongs to a specific person.
Subtle Handheld Motion
real phone footage has subtle motion even when someone is trying to hold still.
perfectly stable footage is what production cameras produce, which is exactly the aesthetic to avoid.
the motion description:
> "subtle natural handheld motion mimicking a propped-up phone, slight micro-shake from breathing and small movements, no static locked-off framing, slight autofocus breathing on the face, organic camera drift"
for content meant to look propped-up, like a vanity mirror selfie or kitchen counter content, keep the motion very subtle.
for content meant to look handheld, like walking-and-talking, make it more pronounced.
matching the motion to the implied filming scenario is what makes the output feel internally consistent.
Emotional Arc
without explicit direction, Seedance defaults to flat neutral expression throughout the whole clip.
flat neutral expression triggers the uncanny valley because real people shift constantly during natural speech.
the base arc:
> "warm and conversational at the opening, slight surprise registering at the midpoint, settling into measured authenticity at the close, natural micro-expression shifts throughout, occasional blinks, subtle head movements while speaking"
for more dramatic arcs:
> "opens with visible frustration in the brow and mouth, transitions to growing curiosity as the discovery happens, lands on relieved approval with a slight smile in the eyes, each transition reads as genuine internal response, not performed"
specificity is what produces emotional realism.
vague emotional direction produces flat output.
this matters more on Seedance 2.5 than it did on 2.0.
a 30 second single pass means the arc has to carry the whole clip, so you're directing a performance now, not a 5 second beat.
Keeping the Character Consistent
for multi-shot reels assembled from several generations, character consistency is non-negotiable.
any drift in the person's appearance between shots breaks the reel at the cut points.
viewers might not consciously spot what feels off, but the inconsistency registers as "something is wrong" and produces the scroll-away behavior that kills hold rate.
the anchoring instruction:
> "maintain exact appearance from @image1, no drift, no deformation, avoid jitter, face stable, natural smooth movement, stable picture throughout, no flickering or ghosting"
this forces the model to lock the character's appearance to the reference image rather than letting it drift.
how 2.5 changes this:
the 30 second single pass holds character far better within one generation.
you're anchoring across fewer seams than before, but any assembly from more than 1 render still needs this.
What Not to Add
Seedance adds things you didn't ask for if you don't explicitly say not to.
background music, ambient sound, text overlays, random objects appearing in the scene, the model fills in details it thinks the clip needs.
the exclusion list:
> "no background music, no ambient noise, no text overlay, no captions, no platform UI elements, no logos, no on-screen text of any kind"
this matters more on 2.5.
native audio generation is stronger now, so the audio exclusions carry more weight than they did on 2.0.
captions get added in CapCut, audio gets mixed separately, logos get composited later.
the generation should produce visuals and dialogue only.
The Technical Closer
every prompt ends the same way.
> "4K, ultra HD, rich detail, sharp clarity, cinematic texture, natural color grading, soft lighting balance, no blur, no ghosting, no flickering, stable picture throughout"
these don't dramatically change the output.
they prevent the small failure modes, subtle ghosting, occasional flickering, soft focus, that otherwise force a regeneration.
at volume, that prevention is worth the 2 lines it costs.
What a Full Prompt Looks Like
putting it all together:
> "@image1 + @image2. the woman from @image1 holds @image2 exactly like the reference selfie, fingers wrapped naturally around the product. she looks directly into the front-facing phone camera with a relaxed half-smirk and soft eyes.""[0-3s] warm and conversational, slight smile at the corners of her mouth. realistic skin texture, visible pores around the nose and cheeks, natural slight unevenness, no filter quality. soft warm window light from camera-right, casting uneven shadows across her cheek and hair, no harsh highlights. messy kitchen behind her, dish rack visible, wooden cabinets, skincare items near the sink, everyday clutter.""she says, '[line 1].' subtle natural handheld motion mimicking a propped-up phone, slight autofocus breathing on the face.""[3-8s] quick natural cut. she demonstrates the product, camera stays close and casual. skin texture and lighting continue from the previous beat. her expression shifts to genuine interest. she says, '[line 2].'""[8-13s] cut back to selfie framing, hands lightly pressed to her cheeks, leaning toward the lens with candid disbelief. tiny natural blink and jaw movement. handheld breathing motion stays subtle. she says, '[line 3].'""[13-15s] she lifts the product toward the camera until it fills the frame, face softly visible behind it. quiet natural laugh. slight edge distortion from the close lens. casual creator energy, not posed. she says, '[CTA line].'""maintain exact appearance from @image1, no drift, no deformation, face stable, stable picture throughout.""4K, ultra HD, rich detail, sharp clarity, natural colors, soft lighting, no blur, no ghosting, no flickering.""no background music, no ambient noise, no text overlay."
this prompt structure consistently passes the realism test when the avatar appears on screen in the first 2 seconds with no further context.
The Post-Production Pass
the production principles get you 90% there.
the remaining 10% comes from CapCut.
the 5 things that close the gap:
- extra subtle camera shake layer, since real phone footage has more handheld motion than even well-prompted Seedance generates
- film grain overlay at low intensity, adding the subtle digital noise real phone cameras pick up
- warm iPhone colour grade, matching the specific colour signature phone footage has that Seedance's neutral output doesn't
- trending audio match for TikTok specifically, since the algorithm weights audio fit as a distribution signal
- platform-native captions, TikTok caption style for TikTok, Stories text for Instagram, longer blocks for Facebook
the whole pass takes 2 to 3 minutes per finished reel.
those 2 to 3 minutes separate AI UGC that converts from AI UGC that gets filed as an ad in the first second.
Running This at Volume
the principles aren't complicated.
every one of them is a specific line or description that goes into the prompt.
the work is standardising them across every production cycle rather than applying them when you remember.
at real volume, 30 to 50 reels a day across a portfolio of accounts, these have to be templated into the workflow.
a master prompt document per creator, loaded into every generation session, means the full production setup rides into every prompt automatically.
the guys who template it produce content that consistently passes the realism test.
the guys who don't produce technically capable output that viewers identify as AI inside the first half-second and scroll past before it has a chance to do anything.
the gap between those 2 groups is a production discipline, not a model choice.
Related articles

How to CONSISTENTLY Get High Views on TikTok With AI UGC
*"realistic skin texture, visible pores, natural slight unevenness, no filter quality"*

The economics of an AI UGC phone farm
So you're thinking of marketing your app on social media with organic UGC. You have a few options to pick from:

How to run your first AI UGC campaign (step-by-step guide)
Restrict the model's exploration space. Force it to define every detail it would otherwise fill in with default in JSON format