How to design a creative testing survey
Learn how to design a creative testing survey, from exposure order and scale anchors to cell sizing and fielding QA that keeps results readable.
A creative testing survey measures how a target audience responds to specific ad concepts before those concepts go live. The survey is the instrument: the screener, the exposure order, the scales, and the open ends that turn a set of headlines, images, or video cuts into comparable numbers.
Creative testing is the wider practice around that instrument. It covers what you measure, how much budget to set aside, how long a round takes, and how results feed the media plan. If you need that broader picture first, start with the guide to creative testing, then come back here to build the questionnaire. This page is about the survey itself: what you ask, in what order, with which answer options, and how you check the data before you trust it.
That distinction matters because most creative tests fail at the questionnaire, not the strategy. Teams pick the right concepts, field to the right audience, and then compare two scores that were never comparable in the first place.
A creative test produces one number per concept per metric. Everything that number means comes from how the question was built. Small design mistakes do not add noise evenly, they bend results in one direction, which is worse: you get a confident answer that points the wrong way.
Here are the flaws that show up most often in review, what each one corrupts, and how to fix it.
| Design flaw | What it corrupts | The fix |
| Unlabeled slider or 0 to 100 scale | Cross-concept comparability. Respondents anchor differently, so a 62 in one cell is not a 62 in another. | Use a fully labeled five-point or seven-point scale with the same anchors in every cell. |
| Showing every concept to every respondent | Independent scores. By concept three, ratings reflect comparison and fatigue, not response to the ad. | Use a monadic design so each respondent sees one concept, or cap sequential monadic at three concepts. |
| Purchase intent asked only after exposure | Attribution. You cannot tell whether the concept moved intent or the audience already had it. | Ask a baseline intent item before exposure, then re-ask the same item after. |
| Diagnostic battery before recall | Unaided recall. Once you name the brand or the message in a question, recall is contaminated. | Always ask unaided recall first, then aided takeaway, then diagnostics. |
| Too many concepts in one survey | Data quality across the whole survey. Length drives dropoff and straightlining. | Keep the concept count low and the question count per concept lower. |
| Uneven cell sizes | Statistical readability. A cell with half the sample has roughly a wider margin of error. | Set a fixed target per cell and quota to it. |
| Leading question wording | Every score downstream. "How much did you love this ad?" is not a measurement. | Use neutral stems and a balanced scale with an equal number of positive and negative points. |
| No attention or speeder checks | Effect size. Inattentive completes flatten differences between concepts. | Add a trap item and a minimum time threshold before the exposure block. |
Every creative testing survey is assembled from four decisions. Get these right and the rest of the questionnaire mostly writes itself.
Exposure design decides who sees what. Monadic gives each respondent a single concept, which keeps every score independent and makes it the default for creative testing across the industry, as YouGov notes in its guidance on the method. Sequential monadic shows each respondent several concepts in randomized order, which costs less sample but introduces order effects.
The practical limits are well documented. Drive Research recommends testing no more than three concepts in a sequential monadic design, and points out that pushing 20 concepts into one survey produces roughly 100 questions and about a 60-minute interview. Attest advises no more than five or six ideas per survey, ideally two or three, and no more than five questions per asset.
Comparative cells, where you deliberately place two concepts side by side and ask the respondent to choose, answer a different question. They tell you which concept wins a head-to-head, not how either performs on its own. Run them as a final tiebreaker block after the monadic scores, never as a substitute.
Order is not cosmetic. Each block in a creative testing survey has to come before the block it would otherwise contaminate. Here is a question sequence you can copy directly into your survey and adapt:
The pre-post pairing in steps three and eight is the piece most teams skip.
Asking the same intent item before and after exposure gives you a within-respondent shift rather than a single post-exposure number, and it separates concepts that raised intent from concepts that were simply shown to a high intent audience.
Keep the wording, scale, and anchor labels identical in both instances. If you change one word, you have measured two different things.
This is where most creative testing surveys quietly break, and it is the part almost nobody publishes. Write the anchors out.
Five-point purchase intent. Use this when you want a familiar, low effort item and you plan to report Top 2 Box:
Seven-point agreement. This is a standard Likert scale format, and it gives you more room to detect differences between similar concepts:
Three rules govern the choice.
On scoring: Top 2 Box, the share of respondents choosing the top two options, is the standard readout for creative testing and it is what the ad testing guide uses.
It is easy to explain and it maps to a decision. Means are more sensitive to small differences but they hide distribution, so a concept that polarizes can post the same mean as a concept nobody cares about.
Report Top 2 Box as the headline and keep the mean and the full distribution alongside it. Pick one metric as your decision rule before you field, not after you see the data.
Plan sample as cells, not as completes. The working target for creative testing is about 200 completes per cell, which gives you enough base to read differences between concepts without overspending. If you want to check the margin of error that base buys you, run the number through a sample size calculator before you field.
The math flows from your exposure design. A monadic test of two concepts at 200 per cell needs 400 completes. A sequential monadic test of four concepts at 200 per cell needs 800, because sequential designs multiply cells by rotation position rather than collapsing them.
Build the budget as a table before you write a single question:
| Input | Example A | Example B |
| Concepts | 2 | 4 |
| Design | Monadic | Sequential monadic |
| Completes per cell | 200 | 200 |
| Total completes needed | 400 | 800 |
| Incidence rate in target audience | 50% | 25% |
| Screener starts required | 800 | 3,200 |
Incidence rate is the input teams forget. If only a quarter of the general population qualifies for your screener, you need four times as many starts as completes, and that is what actually sets the cost and the field time.
Estimate incidence from a prior round or from category penetration data before you commit to a concept count, and check the targeting options on the market research panel to see how tightly you can define the audience without gutting your incidence.
Length is the other constraint on cell count. The monadic design guide frames survey length as concepts multiplied by metrics, and holds the total under 30 questions. Run that multiplication first. If it breaks the ceiling, cut concepts rather than cutting the diagnostics that tell you why a concept lost.
You do not have to write a creative testing questionnaire from a blank page. A few starting points cover most rounds:
A creative testing survey follows a fixed block order: screener, pre-exposure baseline, exposure, unaided recall, aided takeaway, diagnostics, post-exposure intent, open end, and demographics. Write the exposure block first, then work outward, because every other block is positioned relative to it.
A creative testing survey asks one open-ended recall question, one aided message takeaway question, four to six labeled diagnostic scale items, and a purchase intent item repeated before and after exposure. Keep the count per concept low, because question volume drives dropoff faster than survey topic does.
A creative test needs roughly 200 completes per cell to read differences between concepts reliably. Multiply that by your cell count, then divide by your screener incidence rate to get the number of starts you actually have to buy.
Monadic testing shows each respondent one concept, while sequential monadic shows each respondent several concepts in randomized order. Monadic keeps scores independent and costs more sample, and sequential monadic saves sample but introduces order effects, which is why it should stay under three concepts.
The questionnaire is the part of creative testing you fully control.
Fix the exposure order, label every scale point, repeat the intent item before and after exposure, and size your cells against a real incidence rate, and the numbers you get back will hold up in the room where the media decision gets made.
Everything else in a creative test is downstream of those four choices.
Explore the product to see how automated ad testing fields a monadic creative test in as little as an hour.

SurveyMonkey can help you do your job better. Discover how to make a bigger impact with winning strategies, products, experiences, and more.

A logo testing survey only works with the right methodology. Compare monadic vs. comparative design, sample size, and scoring.

Explore 30+ logo testing questions organized by funnel stage, design element, and methodology, then build your survey from a template.

See real concept testing examples across logos, packaging, names, and ads, plus how to turn each into a survey today.