i make $70M/year with these 3 AI UGC formats...

these 3 formats carry most of my ad spend.
podcast ads, pixar style, and street interviews.
each one solves a different selling problem.
this article covers what each is for, how to write it, how to build it, and when to run which...
format 1: the podcast ad:
2 people on a podcast set talking about the product like it came up mid conversation.
half the feed is already podcast clips, so it blends in as content and gets watched through before it gets identified as an ad.
why it converts.
one person plays the skeptic and asks what the viewer is thinking.
while the other has the answer and walks through it.
when the skeptic is won over on screen, the viewer follows the same path.
a single person making the claim to camera has to be believed on their own.
2 people arriving at it together REMOVES that step.
what it suits.
products that need explaining.
supplements with a mechanism, anything with a process behind it, and any offer where the objection is complexity rather than price.
writing it.
- open on the skeptic raising the problem or doubting something
- the other person reacts and introduces the angle
- the product enters through the conversation
- 1 of them handles the obvious objection out loud
- soft close pointing to where to buy
give the skeptic the objection your landing page gets asked most.
when they raise it and the other person answers, the viewer's own objection is handled before they finish forming it.
write both characters differently.
the skeptic is quicker and more casual, the other slower and more considered.
identical delivery gives the format away in 10 seconds.
leave the imperfections in.
interruptions, half sentences, one of them laughing mid answer, etc
length.
2 to 3 minutes long, which is where most AI UGC softwares fail because they stop a single output at 8 to 15 seconds.
building it.

infinite ugc has a podcast feature that builds the 2 person scene from the script directly.
(they're running a sale right now where you only pay $1 for your first 3 days btw)
this is what i use.
- generate both avatars first from tiktok screenshots so neither reads as a stock render
- upload them as the character references
- upload the product image and type the product name to match the script
- set 9:16
- paste the script and pick a voice for each character
- check the preview showing length, image count, and clip count before you spend
give each character a distinct voice.
2 similar reads make the conversation feel like one person talking to themselves.
check the scene proposal for 4 things:
- both characters stay the same people across the cut
- the set stays consistent, meaning the same mics, table, and background
- the camera alternates between them rather than sitting static
- the product appears when it enters the dialogue
testing.
swap which character carries the product first,
since the same script converts differently depending on whether the skeptic comes around or the expert leads with it.
then swap the objection the skeptic raises, since that line does most of the persuading.
format 2: pixar style:

a fully animated story ad with a 3d character.
why it converts.
a 3d character reads as warm and safe, so the viewer watches it as a story before processing it as an ad.
it also stages what a camera cannot film, meaning anything happening inside the body.
what it suits.
emotional benefits like sleep, confidence, energy, and comfort. products aimed at parents.
invisible mechanisms where the mechanism is the entire sell.
the tradeoff.
pixar carries less literal proof than a person holding your product.
so it opens the relationship and a demo or a talking head closes it on retarget.
build both in the same batch so the message stays consistent across the 2 touches.
writing it.
story rules rather than ad rules:
- open on a character in a situation the viewer recognizes
- give them a clear want in the first few seconds
- put an obstacle in the way, which is your problem
- the product is the turn
- close on the character in a better state
the style renders facial expression well, so write the character reacting rather than explaining.
a face changing across 3 scenes carries more than 3 lines of dialogue would.
give the character one exaggerated physical trait, since the style handles distinctive features well and it makes them recognizable across a batch.
length.
40 to 60 seconds long with 5 to 7 scenes.
keep the scene count low, since every transition is a chance for the character to drift.
building it.

runs out of the animation tab in infinite ugc on the pixar setting.
- select pixar style
- upload the product image and type the product name to match the script
- set the aspect ratio
- paste the script and pick the voice
- check the preview before you spend
it builds the character sheet first, meaning every character gets a colour palette and a set of expressions designed before any frame exists.
that is what holds one character across every scene, and it is the single thing that breaks most on other AI UGC tools.
check the proposal for:
- the character staying the same across every scene
- the emotional state visibly changing between the problem and the resolution
- the product entering at the turn
- backgrounds staying simple so the character carries the frame
critique specifically when a scene misses.
if the voice comes back flat, hit regen audio instead of rebuilding the video.
format 3: the street interview:
a person with a mic stops someone on the street and the answer becomes the ad.
why it converts.
3 mechanics at once.
an answer from a stranger reads as an opinion instead of a claim.
people watch this format for entertainment, so it gets watched before it gets identified as an ad.
the question does the hooking, so your first 3 seconds are someone asking something the viewer also wants answered.
what it suits.
anything with a price, a comparison, or a common frustration attached. products where social proof moves the buyer more than a mechanism would.
writing it.
the product has to come from the interviewee.
the interviewer asking about a product reads as an ad while the person answering with a product reads as a recommendation.
- the interviewer asks, straight in, no setup
- the first answer sets up the problem
- a follow-up question deepens it
- the product enters through the answer
- the close is the person's own recommendation
on the question, number questions pull hardest.
"what are you spending on skincare a month" beats "do you like skincare",
since a number gives the viewer something to measure themselves against in the first 2 seconds.
the other types that work: a confession question, a comparison question, and an opinion question on something people already argue about.
match the setting to the product.
a gym entrance for supplements, a shopping district for beauty, a campus for anything aimed younger.
length.
20 to 45 seconds, longer for a multi person cut.
building it.
- screenshot a real street interview on tiktok and json prompt both avatars from it
- upload the interviewer and the interviewee as character references
- upload the product image and type the product name to match
- set 9:16
- paste the script and pick voices that sound different from each other
on a multi person cut, generate each answer as its own scene and let infinite ugc stitch them.
order the answers so the weakest comes first and the one carrying your product comes last, since the sequence builds instead of peaking early.
check the proposal for:
- the mic staying in frame, since it is the visual signal carrying the format
- the setting staying consistent across the cut
- the camera reading as handheld rather than mounted
- the background having movement in it, since an empty street reads as staged
style the captions like the organic version, meaning large centred text appearing word by word.
the software and tools used across these 3 formats:

- Claude — scripts, dialogue for 2 character formats, and hook variations
- Infinite UGC — this is what is used to actually build and generate the AI UGC videos. does literally everything for you except scripting
- ChatGPT — json prompts for the avatars in the podcast and street interview builds
- Nano Banana — generating the avatar images from those prompts
- Meta Ad Library — finding which of the 3 formats competitors are running longest
- TikTok Creative Center — checking format saturation in your category
the tool list is shorter than it looks because 3 of those steps happen inside one place.
the podcast scene, the pixar render, the multi scene street cut, and the captions all sit in infinite ugc,
and the length cap that stops other generators at 8 to 15 seconds does not exist there.
that matters most on the podcast format, since a 3 minute conversation on a capped tool is 12 separate renders and a manual assembly job.
how i run all 3 together:
the order:
- start with a talking head to prove the angle and the script
- rebuild the winner as a podcast ad if the product needs explaining
- rebuild it as pixar if the benefit is emotional or the mechanism is invisible
- rebuild it as a street interview if the buyer moves on social proof
one proven script becomes 4 ads without writing anything new.
keep the core claim identical across all 4 so the data tells you about the format rather than the message.
testing.
around $20 a day per creative, read ctr first, then hold rate, and cut under a 1% ctr by day 3.
weight hold rate on pixar and podcast ads, since both are built around staying to the end.
weight ctr on street interviews, since the question is doing the stopping.
the diagnostics differ by format:
- on a podcast ad, a drop before the halfway mark means the skeptic dynamic is not working, so rewrite the first exchange
- on pixar, a strong hold rate with weak conversion means the product entered too late in the story
- on a street interview, weak ctr means the question failed and the rest of the ad never got a chance
they also fatigue at different speeds,
so when one stops picking up spend, rotate the same script into the next.
then build the library as you go.
every winning script, every avatar that performed, and the questions that pulled.
a script that works in one format is a candidate for the other 2, so the library compounds faster than the testing does.
- CEO
Related articles

How to Make AI UGC Look Like Real iPhone Footage
*"realistic skin texture, visible pores around nose and cheeks, natural slight unevenness, no filter quality"*

How to CONSISTENTLY Get High Views on TikTok With AI UGC
*"realistic skin texture, visible pores, natural slight unevenness, no filter quality"*

The economics of an AI UGC phone farm
So you're thinking of marketing your app on social media with organic UGC. You have a few options to pick from: