Survey significance calculator
Check whether the gap between two survey results is bigger than sampling noise would explain. Enter the percentage and the base for each group, such as two segments or two waves, and get the difference, the z statistic, the two-sided p-value and a plain-English reading.
A two-proportion z-test compares two independent percentages. For 48% of 400 against 40% of 350, the gap is 8 points, z = 2.20 and p = 0.028, which is significant at the 5% level for a single comparison. That says the gap is unlikely to be sampling noise, not that it is large or important.
Difference between A and B
+8.0 percentage points (A − B)
Statistically significant at the 5% level (p = 0.028)
- Group A
- 48.0% of 400
- Group B
- 40.0% of 350
- Difference (A − B)
- +8.0 percentage points
- Pooled percentage
- 44.27%
- Standard error (pooled)
- 3.635 percentage points
- z statistic
- 2.201
- p-value (two-sided)
- 0.0278
- 95% confidence interval for the difference
- +0.9 to +15.1 percentage points
- Significance threshold
- 0.05
- Statistically significant?
- Yes
Workingpooled p = (192 + 140) ÷ (400 + 350) = 0.4427SE = √(p × (1 − p) × (1 ÷ 400 + 1 ÷ 350)) = 0.03635z = (0.4800 − 0.4000) ÷ 0.03635 = 2.201p-value = 2 × (1 − Φ(|z|)) = 0.0278
How to read this
Group A (48.0%) is 8.0 percentage points higher than group B (40.0%).
The p-value of 0.028 is below the threshold of 0.05. If the two populations really had the same percentage, random samples of these sizes would produce a gap this large or larger about 2.8 times in 100. So sampling noise alone is an unlikely explanation. The 95% confidence interval for the true difference runs from +0.9 to +15.1 percentage points.
If this is one of many comparisons (several segments, questions or waves), some will look significant by chance. Raise the number of comparisons above to see the Bonferroni-adjusted threshold.
Statistical significance is not practical significance. With large samples, tiny differences become significant, and with small samples large differences may not. Judge the size of the difference, and its interval, against what would change a decision.
How does the survey significance calculator work?
It runs a two-sided, two-proportion z-test on two independent groups. Under the null hypothesis that both groups have the same true percentage, the best estimate of that percentage pools the two samples:
pooled p = (x₁ + x₂) ÷ (n₁ + n₂)SE = √(pooled p × (1 − pooled p) × (1 ÷ n₁ + 1 ÷ n₂))z = (p₁ − p₂) ÷ SEp-value = 2 × (1 − Φ(|z|))
Here p₁ and p₂ are the two percentages as fractions, x₁ and x₂ the numbers answering, n₁ and n₂ the bases, and Φ the standard normal cumulative distribution function. The p-value is the probability, if the two populations really had the same percentage, of seeing a gap at least this large in either direction from random samples of these sizes.
The confidence interval for the difference uses the unpooled standard error, √(p₁(1 − p₁) ÷ n₁ + p₂(1 − p₂) ÷ n₂), and equals (p₁ − p₂) ± z* × that value. The test uses the pooled error and the interval the unpooled one, so in borderline cases the two can differ slightly.
The test assumes two independent random samples and enough respondents in all four cells (answering and not answering, in each group). A common rule of thumb is at least 10 in each; the calculator warns when a cell has fewer.
What is a worked example?
Segment A is 48% of 400 respondents (192 people) and segment B is 40% of 350 (140 people).
pooled p = (192 + 140) ÷ (400 + 350) = 0.4427SE = √(0.4427 × 0.5573 × (1 ÷ 400 + 1 ÷ 350)) = 0.03635z = (0.48 − 0.40) ÷ 0.03635 = 2.20p-value = 2 × (1 − Φ(2.20)) = 0.028
The 95% confidence interval for the difference runs from 0.9 to 15.1 points. The p-value is below 0.05, so the gap is statistically significant for a single planned comparison.
The line between significant and not is thin. A 45% against 38% gap with the same bases (7 points) gives p = 0.052, just over 0.05. And the same 8-point gap from ten times the sample (4,000 against 3,500) gives z = 7.0. Treat p-values as a measure of evidence, not as a switch.
How should you read the result?
- p-value. The probability of a gap at least this large if there were truly no difference. A small value means the data are hard to square with no difference. It is not the probability that the difference is real or that the result will replicate.
- Confidence interval. The range of true differences consistent with the data. It shows both the direction and the likely size, so it is more informative than the verdict alone.
- Not significant. This means the evidence is not strong enough, not that the groups are equal. A wide interval that includes both zero and large differences calls for more data.
- Significant. This means sampling noise is an unlikely explanation. Bias, weighting or how groups were formed can still produce a real gap that is not caused by what you think.
Is a significant difference an important one?
Not necessarily. Statistical significance depends on both the size of the gap and the size of the samples. With large samples, very small gaps become significant: with enough respondents even a 1-point gap passes the test. With small samples, a large gap may fail to reach significance.
Judge practical significance separately. Ask what size of difference would change a decision, then compare that with the confidence interval. If the whole interval lies below the size that matters, the result is significant but unimportant. If it spans sizes that matter and sizes that do not, collect more data before acting.
What about multiple comparisons?
Each test at the 5% level has a 5% chance of a false positive when there is no real difference. Run many tests, across segments, questions or waves, and some will look significant by chance. For k independent comparisons, the chance of at least one false positive is 1 − (1 − α)ᵏ.
| Comparisons (k) | Chance of at least one false positive | Bonferroni threshold (0.05 ÷ k) |
|---|---|---|
| 1 | 5.0% | 0.0500 |
| 3 | 14.3% | 0.0167 |
| 5 | 22.6% | 0.0100 |
| 10 | 40.1% | 0.0050 |
| 20 | 64.2% | 0.0025 |
The Bonferroni adjustment divides the significance level by the number of comparisons. It is simple and safe but conservative. Holm's step-down method is never less powerful and equally simple, and the Benjamini-Hochberg procedure controls the false discovery rate when many tests are run. Enter the number of comparisons in the calculator to apply Bonferroni to the verdict and the interval.
Decide which comparisons you will make before looking at the data. Testing every segment and reporting only those that look significant makes the adjustment meaningless.
When should you not use this calculator?
- The same people are in both groups. Before-and-after results from the same respondents, or two questions answered by the same people, are paired. Use McNemar's test.
- The groups overlap. Comparing a segment with the total, which includes it, breaks independence. Compare the segment with everyone else.
- The data are weighted or clustered. The test treats the bases as simple random samples. Weighting and clustering reduce precision, so use effective bases (base divided by the design effect) or the survey software's own test.
- The sample is not random. P-values assume random sampling. For opt-in panels and convenience samples they describe the sample, not the population, and should be read with caution.
- Counts are small. If any of the four cells has fewer than about 10 people, use Fisher's exact test.
- The measure is not a percentage. Use a t-test for means such as average scores, ANOVA for means across more than two groups, and a chi-square test for percentages across more than two groups.
Frequently asked questions
What does a p-value mean?
It is the probability of seeing a difference at least as large as the one observed if the two populations really had the same percentage. It does not give the probability that the difference is real, and it says nothing about the size of the difference.
What p-value counts as significant?
The convention is below 0.05, which matches 95% confidence, but the threshold is a choice rather than a law. A p-value of 0.049 and one of 0.051 carry almost the same evidence. Report the p-value and the confidence interval, not only the verdict.
If two margins of error overlap, is the difference not significant?
Not necessarily. Two 95% intervals can overlap and the difference between them can still be significant, because the standard error of the difference is smaller than the sum of the two standard errors. Non-overlapping intervals do imply a significant difference. Use the test itself.
How do I handle many comparisons?
Fix the list of comparisons in advance and adjust the threshold. Bonferroni divides the significance level by the number of tests; Holm's method and the Benjamini-Hochberg procedure are less conservative alternatives. The calculator applies Bonferroni when you enter the number of comparisons.
Can I use it for A/B tests?
Yes, for conversion-style percentages from two independent groups, if you decided the sample size in advance. Checking the result repeatedly and stopping when it first looks significant raises the false positive rate well above 5%.
Why can the test and the interval disagree slightly?
The test uses a pooled standard error, which assumes no difference, while the confidence interval uses an unpooled one. When the result is very close to the threshold, one can show significance while the other just misses.
Cite or link this tool
You are welcome to link to this tool or cite it in a report, article, course or blog post. Copy the HTML to link to it with attribution, or use the plain-text citation.
Link with attribution (HTML)
Plain-text citation
Embed this calculator
Put this calculator on your own website, course page or blog post. It is free to embed, on one condition: keep the credit link under the calculator. The credit link is the licence for free use; if you remove it, please remove the embed as well.
Embed code
The embed is a compact version of this calculator that works inside an iframe. It runs in the visitor's browser, and what they type is never sent to PulseLake or anyone else. Change the height if your layout needs it.
Related tools
- Survey sample size calculatorHow many completed responses do you need? Set the confidence level, margin of error, expected proportion and, if you know it, the population.
- Margin of error calculatorWhat margin of error does your sample give you? Enter the sample size, confidence level and proportion, and optionally the population.
- Van Westendorp price sensitivity calculatorPaste or load answers to the four Van Westendorp price questions and get the four price points, an acceptable range and a chart.
- Gabor-Granger price calculatorEnter the share of respondents who would buy at each price and get the demand curve, the revenue-maximising price and, with a unit cost, the profit-maximising price.
- Survey question bankStandard, widely used survey questions grouped by purpose, with exact wording, scale, when to use each, a common bias to avoid and its source.
- Research brief templateTurn a research request into a clear brief: objective, decision, audience, method, timeline and deliverables, ready to copy.
- All free survey and research calculatorsThe full list, with the formulas each one uses.
Formulas and references
The standard sources behind the formulas on this page:
- Agresti, A. (2013). Categorical Data Analysis (3rd ed.). Hoboken, NJ: John Wiley and Sons. Two-proportion z-test, confidence interval for a difference of proportions and the chi-square equivalence.
- Bonferroni, C. E. (1936). Teoria statistica delle classi e calcolo delle probabilità. Pubblicazioni del R. Istituto Superiore di Scienze Economiche e Commerciali di Firenze, 8, 3-62.
- Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65-70.
- Benjamini, Y. and Hochberg, Y. (1995). Controlling the false discovery rate. Journal of the Royal Statistical Society B, 57(1), 289-300.
- Wasserstein, R. L. and Lazar, N. A. (2016). The ASA statement on p-values: context, process, and purpose. The American Statistician, 70(2), 129-133.
- Abramowitz, M. and Stegun, I. A. (1964). Handbook of Mathematical Functions. Normal distribution functions (series 26.2.10 and the continued fraction 26.2.14, used to compute the normal distribution).
- Acklam, P. J. (2003). An algorithm for computing the inverse normal cumulative distribution function, with one step of Halley's method (used to find the critical value z).
Run research end to end. Keep the knowledge working.
One AI-native operating system for market research and insight professionals — from study design and evidence generation to agents, institutional knowledge, delivery and action.
