The Multiple Comparisons Problem: A Reader’s Guide
Learn why multiple testing inflates false positives, when Bonferroni, Holm, or FDR control fits, and how to audit multiplicity in a research paper.

The multiple comparisons problem appears when a study has many opportunities to find a statistically significant result but interprets each result as though it came from a single planned test. Outcomes, time points, subgroups, doses, models, and interim looks can all multiply those opportunities.
The remedy is not “apply Bonferroni to every p-value.” A careful reader first identifies the family of claims, the role of each analysis, and the error rate the study intended to control. Only then can the reader judge whether the method and conclusion align.
Why many tests create false alarms
Suppose a study runs 20 independent tests, every null hypothesis is true, and each test uses α = 0.05. The probability that no test rejects is:
(1 − 0.05)²⁰ = 0.95²⁰ ≈ 0.358
So the probability of at least one false rejection is:
1 − 0.95²⁰ ≈ 0.642, or about 64%.
With 40 independent null tests, that probability rises to about 87%. These examples show the mechanism, not a universal calculation. Real outcomes and test statistics are often correlated, some null hypotheses may be false, and the precise family-wise risk then differs. The core lesson remains: more analytic opportunities create more chances for a chance result to look special.
The FDA guidance on multiple endpoints treats this as a risk of false conclusions when several endpoints support confirmatory claims.
Define the family before choosing a correction
A multiple-testing procedure controls error across a family of hypotheses. That family should follow the scientific decision, not be chosen after seeing results.
Examples include:
- several primary endpoints where success on any one could establish efficacy;
- secondary endpoints intended to support additional confirmatory claims;
- comparisons of several doses with one control;
- many subgroup claims supporting tailored treatment effects;
- thousands of features screened to generate candidates for follow-up.
Not every p-value in a paper necessarily belongs to one family. A diagnostic model check, a primary efficacy claim, and a descriptive baseline summary serve different purposes. Conversely, splitting closely related confirmatory outcomes into separate “families” after analysis can hide the true number of chances to claim success.
Ask: Which set of rejections would lead the authors to one scientific or regulatory conclusion? That is often the most useful starting point.
A worked example: one positive subgroup among many
Imagine a randomized trial with a null overall treatment effect. The authors then examine treatment effects across sex, three age bands, four disease-severity groups, smoking status, and several biomarkers. One subgroup produces p = 0.03.
That p-value does not answer the most important question: whether treatment effects truly differ across subgroups. The relevant evidence includes:
- whether the subgroup and direction were prespecified;
- how many subgroup hypotheses were examined;
- the interaction test between treatment and subgroup;
- the subgroup effect estimate and confidence interval;
- consistency with biological rationale and prior evidence;
- independent replication.
Comparing “significant in subgroup A” with “not significant in subgroup B” is not a valid interaction test. This is the same contrast error described in the guide to statistical flaws in AI-assisted review.
Where multiplicity hides in research papers
Multiple outcomes and time points
A trial may specify one primary outcome but measure dozens of secondary outcomes at several visits. If a favorable week-4 result is highlighted after week 8 and week 12 are null, time has become another analysis dimension.
Subgroups and treatment arms
Age, sex, baseline severity, genotype, site, dose, and comparator can produce a large grid of tests. Small subgroup samples also create unstable estimates, making apparently dramatic effects more likely.
Alternative outcome definitions and models
Continuous versus responder outcomes, adjusted versus unadjusted models, different covariate sets, outlier rules, and missing-data methods can all change results. If choices are made after inspecting the data, the nominal p-value does not reflect the full search process.
Interim analyses and repeated looks
Testing accumulating data repeatedly and stopping when significance appears inflates false-positive risk unless the design uses an appropriate sequential or alpha-spending procedure.
Selective reporting
Multiplicity becomes especially misleading when readers see only the successful tests. Compare the paper with its registry, protocol, statistical analysis plan, and supplement to reveal the planned analysis space and omitted outcomes.
Correction methods and what they control
| Approach | Controls or structures | Typical role | Important limitation |
|---|---|---|---|
| Bonferroni | Family-wise error rate by testing each hypothesis at α/m | Small, defined confirmatory families | Can be conservative and ignores helpful dependence structure |
| Holm | Family-wise error rate with a step-down procedure | Confirmatory families needing strong control | Still requires the family to be defined in advance |
| Prespecified hierarchy | Ordered testing; later claims unlock only after earlier success | Primary and key secondary endpoints | Claims outside the sequence remain exploratory unless separately controlled |
| Benjamini–Hochberg | False discovery rate under stated dependence conditions | High-dimensional discovery with follow-up | Does not provide the same “no false positive in the family” protection as FWER control |
| Multilevel or model-based strategy | Models related outcomes or groups jointly | Structured correlated hypotheses | Validity depends on model choice, assumptions, and prespecification |
Bonferroni and Holm
Bonferroni divides the family-wise significance level by the number of tests. With 10 hypotheses and a desired family-wise level of 0.05, each test uses 0.005. It is transparent and valid under broad dependence conditions, but it can sacrifice power.
Holm’s step-down procedure orders p-values and compares them with progressively less stringent thresholds. It controls the family-wise error rate and is never less powerful than the corresponding plain Bonferroni procedure.
Hierarchical testing
A prespecified hierarchy may test the primary endpoint first and proceed to key secondary endpoints only if the primary test succeeds. This can preserve strong error control while reflecting scientific priorities. A failed gate means later nominally significant results do not automatically become confirmatory claims.
False discovery rate
The Benjamini–Hochberg procedure targets the expected proportion of false rejections among rejections rather than the probability of any false rejection. This trade-off is often useful when screening many hypotheses and following discoveries with independent validation.
“FDR controlled at 5%” does not promise that exactly 5% of the significant results in one realized study are false. It describes a long-run expectation under the procedure’s assumptions.
When no blanket correction is needed
Absence of a named correction is not automatically a flaw.
- A single prespecified primary hypothesis tested once has no multiplicity within that primary family.
- Co-primary endpoints may require success on all components; that rule can reduce rather than inflate false efficacy conclusions, though power and interpretation still need attention.
- A valid prespecified gatekeeping or sequential design may control error through its structure.
- Descriptive estimates may be presented without formal hypothesis claims.
- Exploratory analyses may remain unadjusted if the paper clearly treats them as hypothesis-generating and does not convert isolated p-values into confirmation.
The test is whether the strength of the claim matches the error control and prespecification, not whether the methods contain the word “Bonferroni.”
How to audit multiple comparisons in a paper
1. Identify the claims
List the conclusions the abstract and discussion treat as established. Separate primary, secondary, safety, subgroup, and exploratory claims.
2. Reconstruct the analysis space
Count outcomes, time points, arms, doses, subgroup variables, model variants, and interim looks. You do not need an exact product of every dimension; you need enough visibility to understand how isolated the highlighted result really was.
3. Compare with prespecified documents
Use the registry record history, protocol, and analysis plan. Check the outcome definition, family, hierarchy, direction, time point, and correction before and after amendments.
4. Find the error-control method
Search for “multiplicity,” “family-wise,” “adjusted p-value,” “Holm,” “Bonferroni,” “gatekeeping,” “alpha spending,” “false discovery rate,” and “Benjamini–Hochberg.” Then confirm that the stated method actually covers the claims being made.
5. Read estimates, not only adjusted significance labels
Adjustment addresses false-positive risk; it does not establish clinical importance, eliminate bias, or make a poor model appropriate. Inspect effect sizes, confidence intervals, outcome definitions, and missing-data sensitivity.
6. Calibrate the conclusion
A prespecified result that survives an appropriate confirmatory procedure can support a stronger claim. An unplanned isolated finding among many tests is better framed as a hypothesis requiring confirmation.
Multiple testing, p-hacking, and forking paths
These ideas overlap but are not identical.
- Multiple testing is a statistical structure that can be planned and handled transparently.
- Selective reporting hides some analyses or outcomes from the reader.
- P-hacking repeatedly changes analyses or reporting decisions to obtain a favorable threshold crossing.
- Researcher degrees of freedom or forking paths describe the many defensible choices whose combined selection can make nominal inference too optimistic even when no explicit grid of tests is shown.
A correction applied only to the final visible tests cannot fully repair an undisclosed, outcome-driven analysis process. Prespecification, transparent reporting, robustness checks, and replication address that broader problem.
Common interpretation mistakes
- Treating every p-value in an article as one family without considering its scientific role.
- Assuming any missing Bonferroni correction invalidates the paper.
- Treating FDR control as if it guarantees the truth of each selected finding.
- Declaring adjusted statistical significance clinically important.
- Comparing significance in two subgroups instead of testing their interaction.
- Ignoring the many analyses described only in supplements or registry history.
- Applying a correction after selectively choosing which tests to include.
Using AI to map multiplicity
An AI-assisted review can inventory outcomes, time points, groups, and analysis labels across a long document. It can also compare the abstract with the methods and flag a secondary result presented as if it were primary.
The reviewer still has to define the scientifically coherent family, verify the procedure, inspect dependence and model assumptions, and judge the claim. Ask the tool to cite source passages and expose uncertainty rather than deliver a binary “correct/incorrect” verdict.
Related guides
- Apply multiplicity checks within the full clinical trial critical-appraisal checklist.
- Use the peer-review workflow to turn a concern into a specific, proportionate reviewer comment.
- Preserve outcome definitions in an accurate research paper summary.
- Compare adjusted and unadjusted findings across studies with a structured evidence matrix.
- Interpret effect ranges before and after adjustment with the confidence-interval reader’s guide.
Frequently asked questions
What is the multiple comparisons problem?
It is the increased opportunity for false-positive findings when a study tests many hypotheses and interprets each result as if it were the only test. The relevant risk depends on the family of claims, dependence among tests, and the error rate the study intends to control.
When should researchers use a Bonferroni correction?
Bonferroni is a simple option when the goal is strong family-wise error control across a defined set of tests. It can be conservative, especially with many correlated tests, so Holm procedures, prespecified hierarchies, or other methods may be more suitable.
What is the difference between family-wise error rate and false discovery rate?
Family-wise error rate controls the probability of making at least one false rejection in a family of hypotheses. False discovery rate controls the expected proportion of false rejections among all rejections, using a convention for the case with no rejections.
Do exploratory analyses need multiple-testing correction?
Exploratory work still needs an error-aware interpretation. FDR control may fit high-dimensional discovery, while clearly labeling analyses as hypothesis-generating and requiring independent confirmation may be more informative than treating isolated unadjusted p-values as established findings.
Related posts
Continue exploring the methods and concepts used in this guide.

Statistics
Confidence Intervals Explained for Research Readers
Interpret confidence intervals using effect size, precision, null values, and practical thresholds—without mistaking non-significance for no effect or certainty.
Read guide →
Peer Review
How to Peer Review a Research Paper
A peer-review checklist for assessing a manuscript’s question, methods, statistics, results, reporting, conclusions, and writing useful comments.
Read guide →
Critical Appraisal
How to Read a Clinical Trial Critically
Use eight practical checks to evaluate a randomized clinical trial’s protocol, randomization, missing data, outcomes, analysis, effects, harms, and applicability.
Read guide →
SinaPilot
Check a paper’s multiplicity in context
Upload a readable paper and use SinaPilot’s Review workflow to map outcomes, subgroups, time points, and statistical claims for your own verification.