How to Read a Clinical Trial Critically
Use eight practical checks to evaluate a randomized clinical trial’s protocol, randomization, missing data, outcomes, analysis, effects, harms, and applicability.

To read a clinical trial critically, reconstruct the question the trial was designed to answer, then follow the evidence from assignment through outcomes to the reported treatment effect. Do not start with the discussion. Start with the protocol, participant flow, methods, and results.
The eight checks below are designed for randomized trials. They complement rather than replace design-specific expertise. The CONSORT 2025 statement defines the minimum information a randomized-trial report should include, while Cochrane’s RoB 2 framework organizes bias assessment around a specific result.
Before reading: assemble the trial record
Open the paper, supplement, trial-registry record, protocol, and statistical analysis plan when available. Capture the registration date and use the registry’s record history to distinguish prespecified decisions from later amendments.
At minimum, write down:
- population and setting;
- intervention and comparator;
- primary outcome, metric, and time point;
- planned sample size;
- analysis population;
- primary treatment-effect measure.
The PICO framework is a useful starting structure, but a complete trial question also needs a time point and an explicit estimand: exactly which treatment effect is being estimated under which assumptions.
1. Was the primary question prespecified?
Compare the publication with the earliest accessible registry entry, protocol, and analysis plan. Check whether the primary outcome has the same variable, metric, aggregation method, and time point in each document.
Look for:
- registration before participant enrollment or before outcomes were known;
- a clearly labeled primary outcome;
- a prespecified analysis population and model;
- documented amendments with dates and reasons;
- agreement between the registered outcome and the paper’s abstract.
A changed outcome is not automatically improper. The concern is an unexplained change made after results could influence the choice, especially when the new outcome is more favorable. ClinicalTrials.gov exposes both outcome definitions and the record history needed to inspect changes.
2. Did randomization protect the comparison?
“Randomized” is a label, not a complete method. Ask how the sequence was generated and whether allocation was concealed from the people enrolling participants. Predictable assignment can allow selection decisions to distort otherwise random groups.
Then inspect baseline characteristics, but avoid performing a battery of significance tests on Table 1. Chance imbalances can occur after valid randomization. The important questions are whether an imbalance is large enough to affect interpretation, whether it was anticipated in the analysis plan, and whether the adjusted analysis was prespecified.
For cluster-randomized, crossover, factorial, or adaptive trials, confirm that the design features appear in both sample-size planning and analysis. A patient-level model applied to cluster-assigned data, for example, can understate uncertainty.
3. Were the comparator, masking, and deviations handled fairly?
Identify what the control group actually received: placebo, usual care, active treatment, waitlist, or something else. The comparator determines the question. “Better than no additional treatment” is different from “better than the current standard.”
Masking is most important where knowledge of assignment can influence behavior, co-interventions, outcome assessment, or decisions to remain in the study. If participants or staff could not be masked, look for protections such as blinded outcome adjudication and objective outcome definitions.
Record major deviations from assigned treatment, crossovers, prohibited co-interventions, and adherence. Then check whether the paper estimates the effect of assignment to treatment, the effect of receiving treatment as planned, or another prespecified estimand. These are different questions and may require different assumptions.
4. What happened to every randomized participant?
Use the flow diagram and outcome tables to reconcile the numbers from randomization to analysis. For each group, record:
- number randomized;
- number receiving the assigned intervention;
- withdrawals and exclusions with reasons;
- number with outcome data at the primary time point;
- number included in the primary analysis.
There is no universal dropout threshold below which bias disappears or above which a trial becomes unusable. Missing data are more concerning when the reason is related to the unobserved outcome, differs between groups, or could plausibly change the result.
Do not judge the issue only by the name of the method. “Multiple imputation” is not a guarantee, and “complete-case analysis” is not always invalid. Examine the assumptions and whether sensitivity analyses test plausible departures from them. Cochrane’s RoB 2 guidance explicitly evaluates missing outcome data through the likely consequences for the result.
The dedicated guide to assessing missing data in a research paper explains missingness mechanisms, analysis choices, sensitivity analyses, and practical audit questions in more detail.
5. Were outcomes measured and reported appropriately?
For the primary outcome, identify:
- the exact variable measured;
- who measured it and whether they were masked;
- the instrument or event definition;
- the analysis metric, such as final value, change, response, or time to event;
- the aggregation method;
- the primary time point.
Subjective outcomes are not automatically weak, but they are more vulnerable when participants or assessors know treatment assignment. Surrogate outcomes may be measured precisely yet remain an incomplete proxy for how patients feel, function, or survive.
Compare every prespecified primary and secondary outcome with the publication and supplement. A missing outcome, altered scale, or newly favored time point needs an explanation. Also inspect harms: how they were collected, for how long, with what denominators, and whether withdrawals due to harms are visible.
6. Did the analysis match the design and question?
Do not try to approve a paper by matching each data type to one “correct” named test. Statistical validity depends on the estimand, design, distribution, dependency structure, model assumptions, and missing-data strategy.
High-value checks include:
| Trial feature | Question for the analysis |
|---|---|
| Repeated measurements | Is within-participant correlation modeled, and is the between-group contrast reported? |
| Cluster assignment | Is clustering reflected in sample size, standard errors, and degrees of freedom? |
| Time-to-event outcome | Are censoring, proportional-hazards assumptions, and absolute risks addressed where relevant? |
| Multiple endpoints or looks | Is the confirmatory error rate controlled through adjustment or a prespecified hierarchy? |
| Non-inferiority or equivalence | Is the margin justified and are the planned analysis populations reported? |
Be alert to a particularly common mistake: a significant before-after change in the treatment group and a non-significant change in the control group do not prove that groups differ. The paper must estimate the contrast between them.
7. How large and precise is the treatment effect?
A p-value is calculated under a statistical model and, assuming the null hypothesis and other assumptions, summarizes how incompatible the observed data are with that model. It is not the probability that the null hypothesis is true, the probability the result occurred “by chance,” or a measure of clinical importance.
Read the effect estimate and confidence interval first. The confidence-interval guide explains how to read both limits against the null and a practical threshold without equating non-significance with no effect:
- For binary outcomes, inspect absolute risks, risk difference, relative effect, and number needed to treat when appropriate.
- For continuous outcomes, interpret the difference on the original scale and compare it with a justified clinically important threshold.
- For time-to-event outcomes, pair relative measures such as a hazard ratio with absolute event risks over a meaningful period when available.
- For null results, ask which important effects the interval still permits.
Use the effect-size interpretation guide when you need to distinguish relative from absolute effects or translate standardized measures back into practical meaning.
Statistical significance can coexist with a trivial effect in a large trial. A non-significant result can coexist with substantial uncertainty in a small one. Neither conclusion should be reduced to whether p < 0.05.
8. Do benefits, harms, and applicability support the conclusion?
The abstract’s efficacy result is only one part of the decision. Compare benefits with serious and common harms, treatment burden, follow-up duration, and withdrawals. Confirm that denominators and observation periods are comparable.
Then ask whether the participants, intervention, comparator, setting, and follow-up match the decision you face. Restrictive eligibility, specialist centers, run-in periods, or unusually intensive monitoring can limit transportability even when internal validity is strong.
Read funding and competing-interest disclosures, but do not use sponsorship as a shortcut for judging a result. Instead, inspect design choices, access to data, the sponsor’s role, publication control, and whether conclusions match the estimates.
A compact clinical trial appraisal table
Use one row per central result:
| Field | What to record |
|---|---|
| Question | Population, intervention, comparator, outcome, time point |
| Prespecification | Registry, protocol, analysis-plan location and dates |
| Result | Analysis population, denominator, estimate, confidence interval |
| Bias concerns | Randomization, deviations, missing data, measurement, selective reporting |
| Clinical meaning | Absolute effect, important threshold, harms, follow-up |
| Applicability | Who and which settings the result reasonably applies to |
| Bottom line | What the result supports, what remains uncertain |
This structure also makes it easier to compare the trial with related papers without flattening important design differences.
Common critical-appraisal mistakes
- Reading only the abstract and discussion.
- Treating “randomized” or “double-blind” as sufficient proof of low bias.
- Applying a universal acceptable-dropout percentage.
- Comparing significance within groups instead of effects between groups.
- Treating non-significance as equivalence or no harm.
- Ignoring absolute effects and confidence intervals.
- Checking benefits carefully while accepting a thin harms section.
- Rejecting an entire trial because one domain raises concern instead of judging a specific result.
Using AI as a second reader
An AI review can help reconcile outcome definitions, trace claims to tables, identify missing reporting, and produce a repeatable checklist. It cannot recover unavailable data, verify model assumptions from prose alone, or decide clinical importance for you.
Ask for source locations behind every concern. Verify each one against the paper and protocol. If a result depends on a specialized model, seek statistical expertise rather than converting an automated question into a definitive criticism.
For claim-level checks, Paper Q&A can help locate a reported endpoint, effect estimate, table, or disclosure in the available full text. Use AI Peer Review for the broader structured critique, then verify both outputs against the trial and protocol.
Related research workflows
- Establish a neutral record with the research paper summary template.
- Turn appraisal findings into constructive comments with the peer-review checklist.
- Apply a design- and result-specific risk-of-bias assessment.
- Diagnose multiplicity with the multiple comparisons reader’s guide.
- Place a trial inside a reproducible systematic literature review workflow.
Frequently asked questions
What should I check first in a clinical trial paper?
Start with the prespecified research question and primary outcome. Compare the paper with its trial registry, protocol, and statistical analysis plan before interpreting the headline result.
Does randomization automatically make a clinical trial reliable?
No. Randomization can reduce confounding, but validity also depends on allocation concealment, deviations from assigned treatment, missing outcome data, outcome measurement, selective reporting, and an analysis that matches the design.
How much missing data is acceptable in a randomized trial?
There is no universal percentage that makes missing data safe or fatal. Risk depends on why data are missing, whether missingness differs between groups, how much the result could change under plausible assumptions, and whether sensitivity analyses support the conclusion.
What matters more, a p-value or a confidence interval?
Neither should be interpreted alone, but a confidence interval is usually more informative because it shows the estimated effect and its precision. Interpret it with the outcome scale, clinical threshold, study design, model assumptions, harms, and prior evidence.
Related posts
Continue exploring the methods and concepts used in this guide.

Critical Appraisal
How to Assess Missing Data in a Research Paper
Assess missing outcome data by reconciling denominators, examining reasons and timing, checking assumptions, and reading sensitivity analyses.
Read guide →
Critical Appraisal
Risk of Bias Assessment: A Practical Guide by Study Design
Assess risk of bias with the right tool, at the right result level, using transparent domain judgments for trials, observational studies, diagnostic studies, and reviews.
Read guide →
Statistics
Effect Sizes Explained: A Practical Guide for Research Readers
Interpret effect sizes using the outcome scale, absolute and relative effects, uncertainty, baseline risk, and practical importance.
Read guide →
SinaPilot
Critically appraise the trial in front of you
Upload a readable trial report and use SinaPilot’s Review workflow to surface study-design, statistical, reporting, and conflict-of-interest questions for verification.