10 min readStatistics

How to Read a Forest Plot Step by Step

Read a forest plot by checking the effect measure, null line, study estimates, confidence intervals, weights, heterogeneity, and pooled estimate.

Luminous study estimates and confidence intervals converging toward a pooled diamond around a reference line

To read a forest plot, first identify the outcome, comparison, effect measure, and null value. Then read each study’s point estimate and confidence interval, note which studies carry more weight, inspect the direction and spread of results, and only then interpret the pooled diamond and heterogeneity statistics.

The diamond is not the starting point. A pooled estimate can look precise while combining studies that answer different questions, use biased methods, or vary in clinically important ways.

The parts of a forest plot

Most intervention forest plots contain these elements:

ElementWhat it usually representsReader’s question
Study rowOne study, comparison, subgroup, or resultWhat exact result is shown?
Square or markerPoint estimate for that rowWhere is the estimated effect?
Marker sizeWeight assigned in the meta-analysisWhich rows influence the pooled result most?
Horizontal lineConfidence interval for the estimateWhich effects remain compatible with the data and model?
Vertical lineThe null or reference valueDoes the interval include no difference?
DiamondPooled estimate and its confidence intervalWhat summary did the chosen model calculate?
Direction labelsWhich side favors which groupIs the plot oriented as expected?
Heterogeneity statisticsBetween-study variation and inconsistency diagnosticsAre effects sufficiently coherent to summarize?

The Cochrane Handbook chapter on meta-analysis describes the conventional square, confidence-interval line, study weight, and summary diamond. Layout varies across software, so read the headings and footnotes rather than relying on shape alone.

Step 1: identify the outcome and comparison

A forest plot is meaningful only for a specified result. Record:

  • the population and included studies;
  • intervention or exposure and comparator;
  • outcome definition;
  • follow-up time;
  • analysis population;
  • effect measure;
  • subgroup or model, if any.

“Mortality” at 30 days and at five years are not interchangeable outcomes. Neither are symptom change and symptom response. A review may contain many forest plots because each outcome, time point, comparison, and subgroup answers a different question.

If the study descriptions are difficult to reconcile, use a structured research paper comparison matrix before interpreting the synthesis.

Step 2: find the effect measure and null value

The null depends on the scale:

  • for differences such as mean difference, standardized mean difference, or risk difference, the null is usually 0;
  • for ratios such as risk ratio, odds ratio, or hazard ratio, the null is usually 1.

Ratio measures are conventionally plotted on a logarithmic axis so reciprocal effects are spaced symmetrically. This means the visual distance from 0.5 to 1 matches the distance from 1 to 2, even though the numeric intervals are different on the printed scale.

Do not infer the measure from the axis alone. A risk ratio and odds ratio can occupy a similar visual position but answer different quantitative questions. The confidence interval guide explains null values, scales, and why an interval crossing the null does not prove equivalence.

Step 3: verify which side favors which group

Read the labels under the axis. Left does not universally mean benefit, and right does not universally mean harm. Direction depends on:

  • which group is in the numerator;
  • whether a higher outcome is desirable or undesirable;
  • whether the measure is a ratio or a difference;
  • how the software ordered the comparison.

For mortality, a risk ratio below 1 may favor the intervention if the numerator is intervention risk. For a beneficial recovery outcome, a risk ratio above 1 may favor the same intervention. Direction labels can also be misleading if the underlying outcome coding is unclear, so confirm the event definition.

Step 4: read each study estimate and interval

The marker gives the point estimate. Its horizontal line gives the confidence interval, usually 95%. Read both limits, not just whether the line touches the null.

Ask:

  1. What effect does the point estimate suggest?
  2. How precise is the estimate?
  3. Does the interval include the null?
  4. Does it include effects large enough to matter?
  5. Does it include important harm as well as benefit?

A small study with a wide interval may be inconclusive even if its point estimate looks dramatic. A large study can be precise but still biased. Statistical precision does not repair confounding, missing data, selective reporting, or poor outcome measurement.

Worked row example

Suppose one row reports a risk ratio of 0.82 with a 95% confidence interval from 0.64 to 1.05.

  • The point estimate suggests an 18% lower relative risk in the intervention group.
  • The interval includes 1, so the result does not exclude the null at the corresponding two-sided level.
  • The interval is also compatible with a 36% relative reduction and a 5% relative increase.
  • Whether that range is reassuring depends on baseline risk, outcome importance, bias, and the smallest effect that would change a decision.

The honest reading is not “no effect.” It is “the estimate is imprecise enough to include materially different possibilities.”

Step 5: understand study weights

In a conventional inverse-variance meta-analysis, more precise studies usually receive more weight. Their markers appear larger, and they influence the pooled estimate more strongly. Weight can depend on sample size, event counts, variance, and the synthesis model.

Do not interpret weight as study quality. A large biased study can receive substantial statistical weight. Risk of bias should affect interpretation and may inform sensitivity or subgroup analyses, but it is not automatically encoded in the square size.

Under a random-effects model, weights are often more balanced across studies than under a fixed-effect model because estimated between-study variance contributes to the weighting. The exact behavior depends on the estimator and data.

Step 6: interpret the pooled diamond

The diamond’s center represents the pooled point estimate; its horizontal tips represent the pooled confidence interval. Before using it, verify:

  • which rows contributed to the calculation;
  • whether the effect measures were made comparable;
  • whether the model was fixed effect or random effects;
  • whether clinical and methodological differences permit a meaningful summary;
  • whether influential studies have important bias concerns;
  • whether the confidence interval is precise enough for the decision.

If the diamond crosses the null, do not translate that into “the intervention does not work.” If it does not cross the null, do not translate it into “every patient benefits” or “every study agrees.” The pooled estimate describes the synthesis model, not a universal biological constant.

Fixed-effect and random-effects interpretations differ

A fixed-effect model estimates one common effect under its assumptions. A random-effects model estimates an average across a distribution of study effects. When genuine effects differ across settings, the average may not describe any specific setting well.

A prediction interval, when appropriately calculated, can help show the range in which an effect from a new similar setting might lie. It is not the same as the pooled confidence interval and can be much wider.

Step 7: assess heterogeneity beyond I-squared

Heterogeneity is variation in effects or study characteristics. Forest plots may report:

  • Chi-squared/Q test: tests whether observed variation exceeds sampling error under its assumptions, but often has low power with few studies and high power with many;
  • I²: estimates the proportion of observed variability attributed to heterogeneity rather than sampling error in the model;
  • Tau²: estimates between-study variance on the effect scale used by the model.

Do not use a universal I² threshold as an automatic pool-or-do-not-pool rule. Read the plot:

  • Do estimates point in the same direction?
  • Are differences small or decision-changing?
  • Do intervals overlap, and is that overlap informative?
  • Are populations, interventions, outcomes, follow-up, and designs comparable?
  • Can prespecified subgroups explain the pattern?
  • Would a single average hide clinically distinct effects?

The number of studies and their precision affect heterogeneity statistics. With only a few studies, uncertainty around I² and Tau² can be substantial even when software prints a single number.

Step 8: inspect subgroups and interaction tests

Subgroup panels may separate studies by dose, population, design, or risk of bias. A difference between “significant” and “not significant” subgroups does not itself show that subgroup effects differ.

Look for a formal interaction or test for subgroup differences, then assess:

  • whether the subgroup hypothesis was prespecified;
  • how many subgroup analyses were attempted;
  • whether the characteristic was measured at study or participant level;
  • whether within-subgroup evidence is credible;
  • whether the difference is large, consistent, and biologically or clinically plausible.

Subgroup analyses are especially vulnerable to the multiple comparisons problem. Treat unexpected splits as hypotheses unless the design and evidence support a stronger conclusion.

Step 9: connect the plot to risk of bias and missing evidence

A forest plot summarizes available estimates. It does not automatically reveal:

  • unpublished studies or outcomes;
  • outcome switching;
  • inaccessible data that could not be plotted;
  • duplicated participant cohorts;
  • errors in extraction;
  • biased measurements or analyses;
  • selective choice among eligible models.

Read the review’s methods, included-study table, exclusions, and risk-of-bias assessment. Compare the plotted outcome with the protocol and registry when these are available.

Funnel plots, when used appropriately, address a different question about small-study effects and possible reporting biases. They are not forest plots and do not prove publication bias or its absence.

Forest plots without a pooled estimate

A forest plot can display study estimates without a diamond. This is useful when studies can be expressed on a comparable scale but a pooled average would be misleading. The pattern can show direction, magnitude, precision, and outliers while preserving study identity.

The Cochrane guidance on synthesis without meta-analysis recommends structured displays rather than selective vote counting. Counting how many studies are “significant” discards effect magnitude and precision and gives small studies the same voice as informative ones.

Common forest plot interpretation mistakes

  • Reading the diamond before confirming the outcome and scale.
  • Assuming left always favors treatment.
  • Treating every square as equally important.
  • Treating larger squares as higher-quality studies.
  • Calling every interval that crosses the null “negative.”
  • Using I² as a standalone quality score or pooling rule.
  • Assuming overlapping study intervals prove homogeneity.
  • Claiming subgroup differences without an interaction test.
  • Ignoring studies excluded from the meta-analysis.
  • Assuming a precise pooled estimate overcomes risk of bias.

A 60-second forest plot checklist

Before accepting the headline conclusion, answer:

  1. What exact comparison, outcome, and time point is plotted?
  2. Which effect measure is used, and where is its null?
  3. Which side favors which group, and why?
  4. What do the largest studies estimate?
  5. Which intervals include meaningful benefit, no difference, or harm?
  6. What model produced the diamond?
  7. Are study differences small enough for the average to be useful?
  8. What do I² and Tau² add—and what uncertainty do they leave?
  9. Do subgroup claims use formal comparisons?
  10. Could bias, selective reporting, or missing estimates change the picture?

A forest plot is a compressed evidence table, not a verdict. Read it row by row, connect its statistical symbols to the research question, and carry design limitations into the final interpretation.

Frequently asked questions

What does the diamond mean on a forest plot?

The center of the diamond represents the pooled effect estimate, and its horizontal width usually represents the confidence interval for that pooled estimate. Interpret it only after confirming which studies, outcome, effect measure, and model produced it.

What does it mean when a confidence interval crosses the line of no effect?

For that estimate, the interval includes the null value at the stated confidence level. This does not prove no effect. The same interval may remain compatible with meaningful benefit, trivial difference, or harm, depending on its limits and the clinical threshold.

Is high I-squared always a reason to avoid meta-analysis?

No single I-squared cutoff determines whether pooling is defensible. Interpret inconsistency alongside the magnitude and direction of effects, confidence intervals, study characteristics, number of studies, tau-squared, model assumptions, and the planned synthesis question.

Can a forest plot be used without a meta-analysis?

Yes. A forest plot can display individual study estimates and confidence intervals without a summary diamond. This can reveal the range and pattern of results while avoiding an unjustified pooled estimate.

Continue exploring the methods and concepts used in this guide.

SinaPilot

Trace every plotted estimate back to its source

Use SinaPilot Paper Q&A to locate the outcome, effect measure, analysis population, and reported estimate in a readable paper, then verify each answer against the source table or figure.