03 · What You Need to Know
What Should You Actually Look for in a Systematic Review or Meta-Analysis?
First, Distinguish the Systematic Review From the Meta-Analysis
The terms are often used together, but they describe different things.
Systematic review
A structured evidence synthesis that uses explicit methods to identify, select, critically appraise, and synthesize research addressing a defined question.
Meta-analysis
A statistical synthesis that combines quantitative results from multiple studies when such combination is methodologically appropriate.
A systematic review does not have to contain a meta-analysis. Sometimes studies are too different in their populations, interventions or exposures, outcomes, designs, or effect measures for a pooled estimate to be meaningful. Cochrane guidance explicitly recognizes that not conducting a meta-analysis may be preferable when variation among study results makes an average effect misleading.
Conversely, calculating a pooled estimate from several studies does not by itself establish that the surrounding review process was systematic or rigorous.
Start With the Review Question and Eligibility Criteria
Before examining the forest plot, determine exactly what evidence the review intended to synthesize.
For intervention reviews, this is often expressed through elements such as population, intervention, comparator, and outcomes. Other types of systematic reviews may require different frameworks. Whatever structure is used, the eligibility criteria should follow logically from the review question.
Look for decisions about populations, settings, interventions or exposures, comparators, outcomes, study designs, publication status, language, and time period. Ask whether these criteria are sufficiently clear and whether important restrictions are justified.
Eligibility decisions determine what the review can eventually conclude. A beautifully executed meta-analysis of studies that poorly represent the question of interest will still answer the wrong or an unnecessarily narrow question.
Look for a Protocol or Prespecified Methods
Systematic reviews involve many methodological choices. Researchers decide what studies qualify, which outcomes matter, how results will be extracted, which analyses will be conducted, how subgroups will be defined, and how risk of bias will be handled.
Making important decisions after seeing the available evidence can create opportunities for selective analysis and interpretation. A prospectively specified protocol helps readers distinguish planned methods from decisions made after study results were known.
Look for protocol registration or publication and compare it with the completed review when feasible. Deviations are not automatically problematic. Research plans sometimes need to change. What matters is whether important changes are transparently reported and justified.
Ask Whether the Search Could Realistically Have Found the Relevant Evidence
A systematic review is only as comprehensive as the process used to identify eligible evidence. Missing relevant studies can change the apparent body of evidence, especially when availability is related to the direction or statistical significance of results.
Examine which bibliographic databases and other sources were searched, when the searches were conducted, and whether the search strategies appear appropriate to the topic. Depending on the question, additional sources might include trial registries, reference lists, citation searching, regulatory sources, conference material, preprints, or other grey literature.
Do not evaluate comprehensiveness merely by counting databases. Searching six poorly chosen sources is not necessarily better than searching three highly relevant ones with carefully developed strategies. The search needs to fit the research question and evidence base.
Watch Out
A systematic review cannot synthesize evidence it never finds. Studies or results may be easier to locate when their findings are statistically significant, favorable, rapidly published, or published in prominent journals. Missing evidence can therefore introduce systematic bias rather than merely reducing the number of studies available.
Study Selection Should Be Transparent and Reproducible
After searching, reviewers must decide which records meet the eligibility criteria. Look for a clear account of how records were screened and how disagreements or uncertainties were handled.
A PRISMA flow diagram can help you follow the movement of records through identification, screening, eligibility assessment, and final inclusion. PRISMA 2020 provides reporting guidance and flow-diagram templates intended to make these processes transparent.
But remember what PRISMA is: a reporting guideline. A completed PRISMA checklist or attractive flow diagram does not prove that the review methods were appropriate. Good reporting makes appraisal possible; it does not replace appraisal.
Read the Characteristics of the Included Studies Before the Pooled Estimate
One of the most useful habits when evaluating a meta-analysis is to resist jumping directly to the diamond at the bottom of the forest plot.
First ask what was actually combined. Are the studies investigating sufficiently similar populations, interventions or exposures, comparators, outcomes, time points, and designs? Are outcome definitions comparable? Are some studies randomized while others are observational? Do intervention doses, implementation conditions, or settings differ substantially?
Statistical software can calculate a pooled estimate whenever compatible numbers are supplied. The more difficult question is whether those numbers represent effects similar enough in meaning that averaging them produces an interpretable quantity.
Evaluate the Included Studies, Not Just the Review
A systematic review does not make methodological weaknesses in its included studies disappear. If most included studies have serious design problems, the synthesis may provide a more precise summary of biased evidence.
Look for a structured assessment of risk of bias using tools appropriate to the study designs and outcomes. For randomized trials, for example, Cochrane recommends its revised Risk of Bias tool, RoB 2, for assessing bias in trial results. Different designs require different approaches.
Risk-of-bias assessment should also affect interpretation. A review that identifies serious bias in the included evidence but then discusses the pooled result as if all studies were equally trustworthy has not completed the critical task.
When necessary, return to the principles involved in critically evaluating the individual research papers and consider whether their methodological problems merely qualify particular findings or fundamentally undermine them.
Do Not Assume More Studies Automatically Mean Better Evidence
A meta-analysis containing 50 studies may look more compelling than one containing five. The number alone tells you little about the credibility of the evidence.
Fifty small studies at high risk of bias do not automatically become high-quality evidence when pooled. Nor does an enormous combined participant count correct systematic problems shared across the studies.
The same logic applies to sample size more generally: large amounts of data do not automatically make research methodologically strong.
Understand What the Pooled Effect Represents
Most conventional meta-analysis methods calculate some form of weighted average of study effect estimates. Studies contribute different amounts of weight depending on the statistical method and information they provide.
Before interpreting the pooled estimate, identify the effect measure. It might be a risk ratio, odds ratio, hazard ratio, mean difference, standardized mean difference, correlation, prevalence estimate, or another statistic.
Then examine the confidence interval and the substantive magnitude of the effect. A statistically significant pooled result is not automatically important. Likewise, a non-significant pooled result does not necessarily demonstrate equivalence or absence of a meaningful effect.
The usual principles of recognizing overstated interpretations still apply to evidence synthesis.
Fixed-Effect and Random-Effects Models Answer Different Statistical Questions
You will often see a meta-analysis described as using a fixed-effect or random-effects model. These should not be treated as interchangeable buttons selected according to a single heterogeneity threshold.
A fixed-effect model assumes that the included study estimates address one common underlying effect, with observed differences attributed to sampling variation under the model. Random-effects approaches allow underlying effects to vary across studies according to a distribution and estimate an average effect across that distribution.
Random-effects meta-analysis does not make heterogeneity disappear. If effects differ substantially across settings or populations, an average may be statistically calculable yet have limited practical meaning unless the variation is understood and appropriately communicated.
Heterogeneity Is More Than an I-Squared Number
Heterogeneity refers to variation among study results. Some variation is expected, but substantial differences can affect what a pooled estimate means and how broadly it can be applied.
The I-squared statistic is commonly reported as an estimate of the proportion of observed variability in effect estimates that reflects between-study heterogeneity rather than sampling error under the model. It is useful, but it should not be interpreted through rigid thresholds alone.
Cochrane specifically cautions against using simple I-squared cutoffs to diagnose heterogeneity because interpretation depends on the magnitude and direction of effects, evidence for heterogeneity, and uncertainty in the statistic, particularly when few studies are available.
Statistical heterogeneity
Observed variation in effect estimates beyond what would be expected from sampling variation under the statistical model.
Clinical or methodological diversity
Differences among studies in populations, interventions, exposures, outcomes, settings, designs, procedures, or other characteristics that may help explain why effects differ.
When heterogeneity is important, ask what might explain it. Are effects different across populations? Intervention doses? Follow-up periods? Outcome definitions? Study quality? Settings?
If the effects point in substantially different directions, a single average deserves especially careful interpretation.
Prediction Intervals Can Be More Informative Than the Average Alone
In an appropriate random-effects meta-analysis, a confidence interval around the pooled effect describes uncertainty about the estimated mean effect. It does not directly describe the range within which effects in different settings might plausibly occur.
A prediction interval can help communicate the expected dispersion of underlying effects in a new study or setting under the model. This may substantially change interpretation when heterogeneity is considerable.
For example, the average effect might favor an intervention while a prediction interval includes effects ranging from meaningful benefit to little benefit or possible harm. In such a situation, “the meta-analysis shows the intervention works” would conceal consequential variation.
Subgroup Analyses Can Explain Heterogeneity, but They Can Also Generate Stories
When studies differ, review authors may investigate whether effects vary by population, intervention type, methodological characteristic, or another factor. Such analyses can be useful, but they are vulnerable to overinterpretation.
Subgroup analyses based on a small number of studies can be unstable. Multiple comparisons create opportunities for chance findings. Explanations developed only after inspecting the results are particularly vulnerable to data-driven interpretation.
Cochrane recommends caution even with prespecified investigations of heterogeneity and treats post hoc explorations primarily as hypothesis-generating.
When a review claims that an intervention “works only in younger participants” or “is effective only in high-income countries,” examine whether the evidence actually demonstrates a difference between subgroups rather than merely showing significance in one subgroup and non-significance in another.
Sensitivity Analyses Tell You Whether Reasonable Decisions Change the Answer
Systematic reviews contain judgment calls. Researchers may choose among statistical models, assumptions, eligibility decisions, definitions, methods for handling missing data, or approaches to studies at elevated risk of bias.
Sensitivity analyses examine whether the main finding remains similar when reasonable analytical decisions are changed.
If excluding studies at high risk of bias causes the pooled effect to disappear, that matters. If alternative defensible models produce substantially different estimates, that also matters. A conclusion that survives reasonable sensitivity analyses is generally more robust than one dependent on a narrow set of analytical choices.
Look for Evidence That Might Be Missing From the Synthesis
Publication bias is often discussed in meta-analysis, but the broader problem is missing evidence. Studies may remain unpublished, and particular outcomes or analyses within published studies may be selectively unavailable depending on their results.
Cochrane describes this more broadly as bias due to missing evidence. Results suggesting statistically significant or favorable effects may be more likely to become available rapidly, appear in prominent journals, or otherwise be easier to locate.
Funnel plots and statistical tests can sometimes help investigate small-study effects or asymmetry, but they are not definitive tests for publication bias. Their interpretation can be difficult when few studies are available or when asymmetry has causes other than selective publication.
Look at the review's broader strategy: Did the authors search trial registries or other relevant sources? Compare protocols or registrations with publications where appropriate? Consider whether missing results could materially change the synthesis?
Certainty of Evidence Is Not the Same as Statistical Significance
A pooled effect can be statistically significant while the overall certainty of evidence remains low.
The GRADE approach, widely used in systematic reviews of interventions, evaluates certainty for a body of evidence by considering domains including risk of bias, inconsistency, indirectness, imprecision, and publication bias. Certainty is categorized as high, moderate, low, or very low.
This distinction is crucial. The p-value associated with a pooled effect does not tell you whether the included studies were biased, whether their results are consistent, whether the evidence directly addresses the question, whether estimates are sufficiently precise, or whether relevant results may be missing.
If a review reports certainty assessments, examine the reasons for downgrading or upgrading rather than treating the final category as another unexplained score.
PRISMA Compliance Does Not Prove That the Review Is High Quality
PRISMA 2020 provides a 27-item reporting checklist, an expanded checklist, an abstract checklist, and flow-diagram templates. Its purpose is to improve transparent reporting of systematic reviews.
That transparency is extremely useful for appraisal. But PRISMA is not a methodological quality certificate.
A review could clearly report an inadequate search, inappropriate synthesis, or weak risk-of-bias assessment and still provide information corresponding to PRISMA items. Conversely, poor reporting can prevent you from determining whether methods were appropriate.
Use PRISMA to help determine whether the review tells you what was done. Then critically evaluate whether what was done was methodologically defensible.
The Review's Conclusion Should Reflect the Weakest Important Links in the Evidence Chain
At the end of the paper, compare the strength of the authors' language with the evidence you have just evaluated.
“The intervention reduces mortality” is a stronger claim than “the available evidence suggests that the intervention may reduce mortality.” The appropriate wording depends not merely on the pooled estimate but on risk of bias, precision, consistency, directness, missing evidence, and the broader certainty of the evidence.
Systematic reviews can look authoritative because they summarize many studies. That makes careful interpretation more important, not less.