Video ad testing: a complete how-to guide
Learn how to run video ad testing step by step, from sample size to survey questions. Start your test with a free template today.
Summary:
Video ad testing is the practice of showing a video ad, or several versions of one, to a sample of your target audience before it goes live, then measuring how they react.
Instead of guessing which cut or message will land, you get direct feedback on recall, clarity, emotional response, and purchase intent while there is still time to make changes.
This guide walks through the full process: setting an objective, choosing between monadic and sequential monadic testing, calculating a sample size, writing survey questions that predict performance, fielding the study, and turning results into a creative decision.
You will also find ready-to-use question wording and a list of mistakes that quietly wreck test results.
Video ads are expensive to get wrong. Test yours with a real audience before you spend on production.
A video ad test only produces useful answers when the setup is deliberate. Treat the steps below as a sequence, not a menu.
Before you write a single question, decide what decision this test needs to inform. Common objectives include:
Your objective determines which metrics matter most. A launch-decision test should weight purchase intent and overall appeal heavily. A diagnostic test should lean on open-ended feedback and message clarity questions instead.
Monadic testing shows each respondent only one video ad. Nobody compares options side by side, so you avoid comparison bias and get a cleaner read on how an ad performs on its own merits. The tradeoff: you need a separate, sufficiently large group of respondents for every ad you test, which raises the total sample size and cost.
Sequential monadic testing shows each respondent two or more ads, one after another, and asks the same questions after each. This lets you test more concepts with a smaller total sample and get a direct, within-person comparison. The tradeoff is order bias, since the first ad someone sees can color how they judge the next one, so rotate ad order across respondents to cancel it out.
Choose monadic for a small number of well-developed concepts where you want the most rigorous, unbiased read on each. Choose sequential monadic when you are screening several early-stage concepts on a tighter budget.
Sample size drives whether your results are noise or signal. For monadic video ad testing, 200 to 300 respondents per ad variant is a commonly cited industry range for detecting meaningful differences in attributes like appeal or purchase intent. Sequential monadic designs can often work with a smaller total sample, since every respondent evaluates multiple ads.
A few factors push the number up or down:
If you are unsure where to start, a sample size calculator can translate your target confidence level and margin of error into a concrete number.
Structure your survey in this order:
Keep the survey focused: once a questionnaire creeps past 30 questions, completion rates drop and answer quality suffers, so trim anything that does not map back to your objective.
Your results are only as good as the people answering.
Recruit respondents who resemble your actual target market, whether that means current customers, a lookalike audience, or a panel filtered by the same age, income, or interest criteria you use in media buying.
A market research audience panel can help you reach a specific demographic quickly if your own contact list is too small.
Launch the survey and monitor completion rates as responses come in. Watch for:
Fielding windows for a targeted panel typically run from a few hours to a couple of days, depending on how niche your audience criteria are.
Once responses are in, compare ads, or compare an ad against your own prior benchmarks, across each metric you set out to measure.
A widely used method is the Top 2 Box score, which combines the two most positive answer choices for a question into a single percentage.
If 35 percent of respondents rated an ad "extremely likely to buy" and 25 percent rated it "very likely," the Top 2 Box purchase intent score is 60 percent.
Layer in a statistical significance check before declaring a winner. A gap of a few points between two ads might sit within the margin of error, especially at smaller sample sizes, so do not treat every numeric difference as a real one. Read the open-ended responses too: quantitative scores tell you what happened, and the verbatim comments tell you why.
Translate results into a decision:
Keep a record of scores across campaigns so you build your own internal benchmarks over time, rather than treating each test as a one-off.
Below are ready-to-use questions covering recall, message clarity, emotional response, and purchase intent. Adapt the wording to fit your brand voice, but keep the scale structure consistent across every ad you test so comparisons stay fair.
Round out the survey with a screener question up front and demographic questions at the end, such as age, gender, or income.
Global ad spend runs into the hundreds of billions of dollars a year, and video commands a growing share of that budget across streaming, social, and connected TV. Every dollar spent on a weak concept is a dollar that could have gone toward a stronger one, which is why pre-testing creative before it reaches paid media has become standard practice at disciplined marketing teams.
Testing before launch gives you real advantages over launching untested and hoping for the best:
None of this replaces good creative judgment; it gives that judgment something concrete to stand on before impressions are on the line.
Even a well-intentioned test can produce misleading results if it falls into one of these traps.
Running a test with too few respondents per ad variant is the most common mistake. A small sample can produce a score that looks decisive but is really just statistical noise. If budget is tight, consider a sequential monadic design instead of stretching a monadic design too thin.
If two versions differ in music, voiceover, pacing, and call-to-action all at once, a winning score tells you almost nothing about which change drove the result. Isolate one or two variables per test.
A small gap between two ads might sit well within the margin of error for your sample size. Always check significance before declaring a winner, and treat close results as a tie rather than a victory.
A blurry upload or a slow-loading link can tank scores for reasons that have nothing to do with the creative itself.
Testing a niche product with a general population panel, or skipping screener questions, produces feedback from people who were never going to buy it.
Numeric scores tell you what happened; skipping verbatim feedback means missing why, often the more useful part of the test.
For monadic testing, 200 to 300 respondents per ad variant is a commonly cited industry range for detecting meaningful differences in metrics like appeal or purchase intent. Larger samples improve your confidence interval and let you reliably analyze subgroups. Sequential monadic designs can often work with fewer total respondents, since each person evaluates multiple ads.
Monadic testing shows each respondent a single video ad, which avoids comparison bias but requires a full sample for every ad you test. Sequential monadic testing shows each respondent multiple ads in sequence and asks the same questions after each, which is more cost-effective and lets you test more concepts, at the cost of some order bias that rotation can help offset.
Video ad testing surveys your target audience before launch and asks direct questions about clarity, emotion, and intent. A live A/B test runs two or more versions in an actual ad platform and measures real behavior, such as click-through or conversion rate, from a slice of live traffic. Pre-launch testing is lower-risk since no media budget is spent yet; a live A/B test shows how an ad performs in the real environment. Many teams use both, running a survey-based test to narrow the field before an in-platform A/B test on the finalists.
That depends on your method. Monadic testing works best with a handful of well-developed concepts, since each one needs its own full sample. Sequential monadic testing can handle more concepts in a single study, since every respondent evaluates several, though a long list of videos will strain attention and survey length.
Most tests should cover recall, message clarity, emotional response, believability, relevance, and purchase intent at a minimum. A launch-decision test should weight purchase intent and overall appeal heavily, while a diagnostic test on an underperforming ad should lean more on clarity and open-ended feedback.
You do not need a big research team to test a video ad properly. You need a clear objective, the right method for your number of concepts, a sample size that can detect a real difference, and questions that map back to the decision you are trying to make.
If you are ready to put this process into practice, start from a video ad testing template to build your questionnaire in minutes, or try the video ad testing feature if you want a guided setup with built-in scorecards and an integrated respondent panel.

SurveyMonkey can help you do your job better. Discover how to make a bigger impact with winning strategies, products, experiences, and more.

A diary study is a qualitative research method where people log experiences over time. Learn when to use one and see real examples.

Learn how to run a win-loss analysis with a repeatable framework, real interview questions and a free template. No CI vendor required.

Learn how to run a market assessment: a TAM, SAM, SOM walkthrough, sample questions, and a go/no-go checklist. Try SurveyMonkey free.