11 min readStatistics

Effect Sizes Explained: A Practical Guide for Research Readers

Interpret effect sizes using the outcome scale, absolute and relative effects, uncertainty, baseline risk, and practical importance.

Paired scientific distributions with different separations being calibrated onto a common effect-size plane

An effect size describes how large a difference, association, or change is. To interpret one, identify the outcome and comparison, confirm the effect measure and its null value, translate the estimate into meaningful units, inspect its confidence interval, and compare the plausible range with a practical threshold.

Do not ask whether an effect size is “good” before asking what it measures. A risk difference of 2 percentage points, a risk ratio of 0.80, and a standardized mean difference of −0.30 can describe different views of an effect—or entirely different outcomes.

Effect measure, effect estimate, and uncertainty

These terms play different roles:

  • Effect measure: the statistical scale chosen for the comparison, such as mean difference or risk ratio.
  • Effect estimate: the value calculated from the observed data on that scale.
  • Uncertainty interval: the range, often a confidence interval, describing statistical precision under the model.
  • Practical threshold: the smallest difference that would matter for a decision, ideally defined before results are known.

For example, “risk ratio” is the measure, 0.75 is the estimate, and 95% CI 0.58 to 0.96 expresses uncertainty. None of those values establishes whether the effect matters without the outcome, baseline risk, follow-up, bias, and decision context.

The Cochrane Handbook chapter on effect measures distinguishes difference measures from ratio measures and shows how the appropriate choice changes with the outcome data.

Choose the measure from the outcome

Outcome typeCommon effect measuresNull valueEssential context
Binary eventRisk ratio, odds ratio, risk difference1 for ratios; 0 for differencesEvent definition, baseline risk, follow-up
Continuous measurementMean difference, standardized mean difference0Instrument, direction, variability, meaningful change
Time to eventHazard ratio, survival difference, difference in restricted mean survival time1 for hazard ratio; 0 for differencesTime horizon, censoring, proportional-hazards assumption
Count or rateRate ratio, rate difference1 for ratios; 0 for differencesPerson-time, recurrent events, exposure duration
AssociationCorrelation or regression coefficientUsually 0Coding, adjustment set, design, functional form

The same dataset can support several measures. Authors should explain which one answers the research question and report enough group-level information for readers to interpret it.

Interpret binary outcomes in absolute and relative terms

Suppose a harmful event occurs in 20 of 100 participants in the comparator group and 15 of 100 in the intervention group.

  • Comparator risk: 20%.
  • Intervention risk: 15%.
  • Risk ratio: 0.15 / 0.20 = 0.75.
  • Relative risk reduction: 1 − 0.75 = 0.25, or 25%.
  • Risk difference: 0.15 − 0.20 = −0.05, or 5 percentage points fewer.
  • Number needed to treat: approximately 1 / 0.05 = 20, if the assumptions and time horizon support this transformation.

“25% lower risk” sounds larger than “5 percentage points fewer,” but both arise from the same hypothetical data. The relative measure describes proportional change. The absolute measure describes the change from this baseline risk over this follow-up period.

Baseline risk changes absolute benefit

Apply the same risk ratio of 0.75 to two populations:

  • At a 20% comparator risk, expected intervention risk is 15%: a 5-point difference.
  • At a 4% comparator risk, expected intervention risk is 3%: a 1-point difference.

Relative stability does not make absolute benefit constant. Treatment burden, harms, cost, and patient preferences may differ even when the relative effect is similar.

Risk ratio and odds ratio are not interchangeable

Risk is the probability of an event. Odds are risk / (1 − risk). When events are common, an odds ratio can be farther from 1 than the corresponding risk ratio and should not be described as if it were a relative risk.

Odds ratios are natural outputs of some study designs and models, including logistic regression and many case-control analyses. Interpret the measure that was actually estimated, and convert only when the required baseline information and assumptions are available.

Read mean differences in original units

A mean difference preserves the measurement scale. If a blood-pressure trial reports a mean difference of −4.5 mmHg, the estimate is directly interpretable in mmHg.

Check:

  • whether the outcome is an endpoint or change from baseline;
  • whether lower or higher values are better;
  • which groups were subtracted;
  • whether the scale is linear enough for the difference to be meaningful;
  • whether means represent a skewed distribution adequately;
  • how the estimate compares with a clinically or practically important difference.

Original units are usually easier to use than a standardized value. Do not standardize merely because the resulting number looks more portable.

Understand standardized mean differences

A standardized mean difference divides a mean contrast by a standard deviation so studies using different instruments for the same underlying construct can be compared or combined.

Common variants include:

  • Cohen’s d: a mean difference divided by a standard-deviation estimate;
  • Hedges’ g: a standardized mean difference with a correction that reduces small-sample bias.

The sign indicates direction only after the group order and scale direction are known. An estimate of −0.40 could favor the intervention on a symptom scale where lower is better, but favor the comparator on a functioning scale where higher is better.

The denominator affects the meaning

A standardized effect is expressed in standard-deviation units. The same raw difference produces a smaller standardized effect in a more variable population and a larger effect in a less variable one.

Variability can reflect real population diversity, measurement reliability, eligibility restrictions, or study methods. Therefore, standardized effects are not independent of context even though they have no original measurement unit.

The Cochrane interpretation guidance recommends planning how an SMD will be made interpretable, such as re-expressing it on a familiar instrument or in relation to a meaningful-difference threshold.

Do not universalize small, medium, and large labels

Rules of thumb such as 0.2, 0.5, and 0.8 for standardized mean differences are conventions, not biological laws. An effect below a generic “small” threshold can matter for a common serious outcome, while a “large” surrogate change may not improve how patients feel or function.

Judge importance using:

  • the outcome’s role and measurement quality;
  • a prespecified minimal important difference when credible;
  • baseline severity and risk;
  • duration and durability;
  • adverse effects and treatment burden;
  • feasibility and cost;
  • the distribution of individual responses;
  • uncertainty and risk of bias.

If the meaningful threshold was defined after seeing results, it may have been chosen to favor the observed estimate.

Interpret correlations and regression coefficients carefully

A correlation describes the direction and strength of a linear association between two variables. It does not show that changing one will change the other. A value near zero can coexist with a strong nonlinear relationship, while a high correlation can reflect confounding, restricted sampling, or common measurement artifacts.

Regression coefficients depend on:

  • outcome and predictor units;
  • reference categories;
  • transformations and interactions;
  • adjustment variables;
  • model specification;
  • the population represented by the data.

An adjusted estimate and an unadjusted estimate answer different model-based questions. Do not choose the more favorable one without understanding the prespecified analysis and adjustment rationale.

Interpret hazard ratios as relative event rates over time

A hazard ratio compares instantaneous event rates under a time-to-event model. It is not a risk ratio, and it does not directly tell you how many participants experience the event by a particular time.

Ask whether:

  • the proportional-hazards assumption is reasonable;
  • survival curves cross;
  • censoring could be informative;
  • median survival or absolute survival probabilities are also shown;
  • follow-up is long enough for the decision;
  • competing events alter interpretation.

When hazards are not proportional, one summary hazard ratio may conceal effects that change over time. Differences in survival probability at relevant time points or restricted mean survival time may be more interpretable, depending on the question and analysis plan.

Pair every effect size with its confidence interval

The point estimate is one value supported by the sample and model. The confidence interval shows how imprecise that estimate remains.

Read the interval against three landmarks:

  1. the null value;
  2. the smallest important benefit;
  3. the smallest important harm.

An interval can exclude the null yet remain compatible only with trivial effects. Another can cross the null while including important benefit and harm. The confidence interval guide explains how to avoid turning those patterns into binary significance labels.

A worked interpretation across scales

Imagine a trial reports:

  • symptom mean difference: −3.2 points;
  • 95% confidence interval: −5.8 to −0.6;
  • prespecified meaningful difference: 5 points;
  • standardized mean difference: −0.28;
  • serious adverse-event risk: 3% versus 2%.

A defensible summary is:

The intervention improved the symptom score by an estimated 3.2 points. The interval excludes no mean difference but includes effects from 0.6 to 5.8 points, crossing the prespecified 5-point threshold. The standardized estimate is modest, and the adverse-event estimate is too sparse to establish safety. Practical importance therefore remains uncertain despite statistical evidence of a mean difference.

This preserves magnitude, uncertainty, threshold, and harm rather than announcing only that the result was significant.

Effect sizes in a meta-analysis

Before combining estimates, confirm that studies use compatible:

  • populations and interventions;
  • outcomes and time points;
  • group directions;
  • effect measures;
  • adjustment strategies;
  • study designs and estimands.

Standardization solves a unit problem, not clinical or methodological heterogeneity. Two symptom scales may target different constructs even if both produce an SMD.

The forest plot guide shows how individual effect sizes, weights, confidence intervals, heterogeneity, and a pooled estimate interact. A pooled average does not erase important study differences.

Separate magnitude from credibility

A large estimate is not automatically convincing. Ask whether bias could inflate, attenuate, or reverse it.

Check:

  • randomization or confounding control;
  • deviations from intended conditions;
  • missing outcome data;
  • outcome measurement;
  • selective analysis and reporting;
  • multiplicity across outcomes, time points, and subgroups;
  • model assumptions;
  • applicability to the decision population.

Use a design-specific risk-of-bias assessment rather than treating effect magnitude or sample size as a quality score.

Common effect-size mistakes

  • Calling a p-value an effect size.
  • Reporting “large” without naming the measure or outcome.
  • Treating an odds ratio as a risk ratio.
  • Presenting only relative change when baseline risk is available.
  • Using a generic standardized-effect threshold as a clinical cutoff.
  • Ignoring the sign convention for scales where lower is better.
  • Comparing standardized estimates that target different constructs.
  • Reading a hazard ratio as a difference in survival probability.
  • Treating a narrow interval as protection from bias.
  • Choosing the most favorable measure after seeing the results.
  • Interpreting an adjusted estimate without its adjustment set.
  • Ignoring harms because the primary benefit estimate is precise.

A practical effect-size checklist

For every central estimate, record:

  1. population, comparison, outcome, and time point;
  2. data type and effect measure;
  3. group ordering and direction of benefit;
  4. point estimate and confidence interval;
  5. null value;
  6. original units or baseline risk;
  7. practical threshold and who defined it;
  8. analysis population and adjustment set;
  9. model assumptions;
  10. bias, multiplicity, and missing-data concerns.

The CONSORT 2025 explanation emphasizes reporting the treatment-effect estimate, the chosen measure, and its precision for trial outcomes. Readers should demand the same clarity before converting an effect into a decision.

Effect size is the beginning of interpretation, not its conclusion. Preserve the scale, uncertainty, baseline, credibility, and consequence of being wrong.

Frequently asked questions

What is an effect size?

An effect size is a quantitative measure of the magnitude of a difference, association, or change. Examples include a mean difference, risk ratio, risk difference, odds ratio, hazard ratio, correlation, and standardized mean difference. Its meaning depends on the outcome, scale, comparison, and study design.

What is considered a small, medium, or large effect size?

There is no universal threshold that makes an effect small, medium, or large in every field. Generic conventions can help with orientation, but practical importance should be judged against the outcome scale, measurement reliability, baseline risk, harms, costs, and a context-specific meaningful threshold.

Is Cohen's d the same as an effect size?

Cohen's d is one standardized mean-difference effect size, but effect size is a broader category. Hedges' g applies a small-sample correction to a standardized mean difference, while binary and time-to-event outcomes use other measures such as risk ratios, risk differences, odds ratios, or hazard ratios.

Can an effect size be important without statistical significance?

Yes. A point estimate may be practically important while its confidence interval remains too wide to exclude the null. Conversely, a very small effect can be statistically precise in a large sample without being useful. Interpret magnitude and uncertainty together.

Continue exploring the methods and concepts used in this guide.

SinaPilot

Retrieve the effect before interpreting the claim

Use SinaPilot Paper Q&A to locate the outcome, effect measure, estimate, confidence interval, baseline risk, and analysis population, then verify them in the source table or figure.