Blog

How to make street-interview ads without a film crew

20 September 2026

A street interview is the cheapest-looking ad format that consistently outperforms expensive ones. Someone holds a mic, stops a stranger, asks a question with a sharp edge to it, and the stranger answers in a way nobody could have scripted. It works because the viewer believes the exchange happened.

Producing one the traditional way is not cheap. Below is what the format actually requires, what each part costs in a real shoot, and how to get the same thing without any of it.

What a street interview is made of

Strip a good street poll down and there are five components, and only five.

Everything else — the color grade, the music, the lower thirds — is decoration.

What it costs to shoot

A modest street-interview shoot in a US city runs an interviewer at a day rate, a camera operator, a sound operator or at least a competent lav setup, a permit if you are shooting on a sidewalk with a crew and a tripod, and half a day of editing. Call it a day of five people's time, plus the pre-production of writing questions and finding a corner that does not have a construction site on it.

The part people forget is the release form. Every stranger whose face appears has to sign. Some will, some will not, and the best reaction of the day is frequently from someone who then declines. You shoot for six hours to get four usable exchanges, and your edit is constrained by what you happened to capture, not by what the ad needed.

Then the ad fatigues in three weeks and you do it again.

The generated version

The same five components can be produced without any of the logistics. The interviewer becomes a character record — a set of reference images, a voice, and a name — that you reuse in every video. The street becomes a scene you describe once and keep. The respondents are generated people, so there are no releases to chase and no one to decline after the fact.

What changes is where the difficulty sits. You are no longer solving for weather, permits and availability. You are solving for believability, and believability in generated footage comes from specific things:

Working in long takes

The biggest practical difference from a shoot is that you should generate the exchange as one continuous take rather than as separate shots you assemble afterwards. A long take keeps the mic in the same hand, the interviewer in the same position, and the respondent's posture continuous. Cutting between separately generated shots reintroduces exactly the discontinuities that make AI video look like AI video — the jacket that changes, the hand that swaps sides, the background car that teleports.

Generate thirty or forty seconds of the exchange, then mark the two or three seconds that actually land and cut to them. This is the same thing an editor does with real footage, and for the same reason: the usable material is a fraction of what was captured.

Writing the question

Question quality determines everything downstream, and it is the part most people rush. A good street-poll question has a correct answer the respondent does not have, or an honest answer that is mildly embarrassing. "How long do you think you spend on your phone a day?" works because every answer is wrong and the respondent knows it is wrong as they say it.

Bad questions ask for opinions. "What do you think about sleep?" produces nothing, because there is no gap between what the respondent says and what they mean. You want the gap. The gap is the ad.

Once you have the question, the product enters late — after the reaction, not before it. A poll that opens with the product is a commercial. A poll that opens with an embarrassing admission and arrives at the product twenty seconds later is a piece of content that happens to sell something.

The series effect

The reason to build a recurring interviewer rather than a new face each time is that street polls compound. The fifth video with the same interviewer on a different street performs better than the first, because returning viewers recognize the format before the first word. That only works if the face, the voice and the mic are identical across all of them — which is straightforward when the interviewer is a record you reuse, and nearly impossible when it depends on booking the same person for another shoot day.

Keep the question style consistent too. If your first three are about phone use, sleep and reading, the fourth should be recognizably from the same family. You are building a format, not a set of one-offs.

What to do first

Write one question with a trap in it. Decide who is asking it and what they look like. Generate the exchange as a long take, mark the seconds that land, cut, caption, post. Then change one variable — the street, the respondent, the question — and run it again. The second one takes a fraction of the time of the first, which is the actual advantage of not having a film crew: not that the shoot is cheaper, but that the tenth one costs the same as the second.

Start free
© 2026 Klovera · Internet Mastery
BlogDone for you TermsPrivacyContact