What questions should you ask in a concept test?
A copyable question set grouped by purpose, the monadic versus sequential monadic choice, and how to read the scores.
A concept test should ask the same core questions about every concept: overall appeal, personal relevance, uniqueness, believability, purchase intent and value for money, each on a fixed 5-point scale, plus open questions on likes and dislikes. Ask them right after people see the concept, and read the scores against concepts tested the same way rather than as sales forecasts.
Which questions belong in a concept test, grouped by purpose?
A usable concept test asks one screener up front, then the same six scored questions and three open questions after each concept, in the order below. Replace the square brackets, keep the wording and scales identical for every concept, and keep the same order every time.
| # | Purpose | Question | Scale | Read it as |
|---|---|---|---|---|
| 1 | Screening (before any concept) | Which of the following have you [bought or used] in the past [3 months]? | Category list with "None of these" last | Admit only people in the target category |
| 2 | Appeal | Overall, how appealing is this [product] to you? | Extremely, very, somewhat, not very, not at all appealing | Top-two-box share |
| 3 | Relevance | To what extent do you agree or disagree: "This [product] is relevant to me." | 5-point agreement | Top-two-box share |
| 4 | Uniqueness | How different is this [product] from others you can buy today? | Very different to not at all different (5 points) | Top-two-box share |
| 5 | Believability | How believable is what this [product] says it does? | Completely to not at all believable (5 points) | Top-two-box share |
| 6 | Purchase intent | If [product] were available at [price], how likely would you be to buy it? | Definitely, probably, might or might not, probably not, definitely not | Top-two-box share, as a ranking |
| 7 | Price and value | Considering its price of [price], how would you rate it as value for money? | Excellent to very poor value (5 points) | Top-two-box share |
| 8 | Diagnostics: likes | What, if anything, do you like about this [product]? | Open text | Themes, counted |
| 9 | Diagnostics: dislikes | What, if anything, do you dislike or find unclear? | Open text | Themes, counted |
| 10 | Diagnostics: change | What one thing would you change about it? | Open text | Themes, counted |
The screener, purchase intent and agreement wordings are the ones in our free survey question bank, which gives the source and the bias to avoid for each. The appeal, uniqueness, believability and value wordings follow common industry practice; there is no single published standard for them. If you add a clarity check, use one statement per idea, such as "The description is easy to understand", and never ask about two things at once.
Should you use a monadic or a sequential monadic design?
Use a monadic design when you need clean, comparable scores for each concept; use sequential monadic when budget limits the sample and you mainly need a ranking, and then rotate the order.
In a monadic test each respondent sees one concept. In a sequential monadic test each respondent sees several, one at a time, and answers the same questions after each. Sequential designs cost less but the order changes the scores. In commercial tests where each respondent rated three concepts, concepts were rated higher when shown first, and a strong concept pulled down the rating of the next one while a weak one lifted it (Friedman and Schillewaert, 2013). An Ipsos analysis of 22 two-product tests (2013) found a second-position effect and notes the standard fix: rotate so each concept appears equally often in each position.
| Design | Four concepts, 150 ratings per concept | Strength | Risk |
|---|---|---|---|
| Monadic | 600 completes, each sees one concept | No order or comparison effects | Four times the sample |
| Sequential monadic | 150 completes, each sees all four, rotated | A quarter of the sample; within-person comparison | First position flatters; only about 38 people see each concept first |
A useful habit: in a sequential monadic test, analyse the first concept each person saw on its own as a small monadic read, and check that it ranks the concepts the same way as the full data.
How many respondents does each concept need?
Enough that the gap you care about is bigger than the noise. On 150 people, a top-two-box score near 50% has a margin of error of about ±8 points at 95% confidence, and two concepts with 150 ratings each need to differ by about 11 points before the gap is statistically significant. For ±5 points you need 385 ratings per concept. The sample size calculator does this arithmetic for your own settings, and the survey significance calculator tests a gap between two concepts.
How should you read concept test results?
Read concepts against each other and against your own past concepts tested the same way, and use the open answers to explain the scores.
- Report top-two-box shares and the full distribution. Two concepts with the same top-two-box can differ in how many people chose the strongest answer.
- Treat purchase intent as a ranking, not a forecast. Stated intentions "often contain systematic biases" and may not predict purchases (Sun and Morwitz, 2010).
- Read the measures together. The pattern tells you what to fix:
| Pattern | Likely meaning | Next step |
|---|---|---|
| High appeal, low believability | People want it but doubt the claim | Add proof, then retest the claim |
| High uniqueness, low relevance | Different, but for someone else | Check the audience, or narrow the target |
| High relevance, low uniqueness | Wanted, but seen as like existing options | Sharpen the difference before launch |
| High intent, low value for money | Likes it, resists the price | Run a pricing method on the price |
| Middling everything, few open answers | The concept is unclear | Rewrite the description and test again |
- Split by segment. A concept that wins overall can lose with the buyers who matter, so report the target segment separately and size the sample for it.
Our free survey question bank holds the screener, purchase-intent and agreement questions above with their sources, and the sample size calculator sizes each cell. In PulseLake, a concept test runs as traditional survey research, with its objectives, method, evidence and decisions kept in one study context.
Frequently asked questions
How many concepts can one respondent evaluate?
There is no fixed limit, but every extra concept lengthens the survey and adds order effects. The order-effect evidence cited on this page comes from commercial tests where each respondent saw three concepts, so test more than that only if you rotate the order and check that later positions still behave.
Should the concept show a price?
Show the price if you ask purchase intent at a price, as the standard wording does. If the price is still open, test the concept without one and run a pricing method such as Van Westendorp or Gabor-Granger afterwards.
What is a good top-two-box purchase intent score?
There is no universal benchmark. Published norms are category-specific and usually proprietary, so the useful comparison is with your own earlier concepts, tested with the same wording, scale, audience and design.
Can you test concepts in interviews instead of a survey?
Yes, for early concepts whose wording still changes: interviews show what people understand and why they react. A handful of interviews will not give scores you can compare between concepts, so use them to refine a concept and a survey to choose between concepts.
Sources
Sources for the facts on this page, last checked October 9, 2026.
- Friedman and Schillewaert, Order and quality effects in sequential monadic concept testing (EMAC 2013), UCLouvain record checked October 9, 2026
- Ipsos, Understanding the Order Effects in Sequential Monadic Product Tests (2013) checked October 9, 2026
- Sun and Morwitz, Stated intentions and purchase behavior: A unified model, International Journal of Research in Marketing 27 (2010) checked October 9, 2026
Related guides
- Survey question bank, with exact wording and sources
- Survey sample size calculator
- Gabor-Granger vs Van Westendorp: which pricing method should you use?Van Westendorp finds a range of acceptable prices; Gabor-Granger finds the price that earns most within a set you choose. Worked examples of both, with numbers.
- MaxDiff vs conjoint analysis: when should you use each?MaxDiff ranks a single list of items; conjoint measures how attributes and levels, often including price, trade off inside whole products. When to use each.
Run research end to end. Keep the knowledge working.
One AI-native operating system for market research and insight professionals — from study design and evidence generation to agents, institutional knowledge, delivery and action.
