Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Evaluate a Systematic Review or Meta-Analysis?

A systematic review can summarize an entire body of research, but its conclusions are only as credible as the methods used to find, select, evaluate, and synthesize the evidence. A meta-analysis adds statistical synthesis, not an automatic guarantee of stronger evidence.

159
How to Evaluate a Systematic Review or Meta-Analysis Guide 159 of 247
01 · The Question

Can You Trust a Paper More Because It Is a Systematic Review or Meta-Analysis?

You find a systematic review containing 35 studies and a meta-analysis involving thousands of participants. The forest plot shows a neat pooled estimate with a narrow confidence interval. Compared with a single study, this may look like much stronger evidence.

It can be. A well-conducted systematic review can identify and critically synthesize a body of research far more informatively than reading individual studies in isolation. When appropriate, meta-analysis can combine effect estimates to improve precision and help researchers investigate variation across studies.

But neither the label “systematic review” nor the presence of a meta-analysis establishes that the evidence is trustworthy. A review can miss relevant studies, use inappropriate eligibility criteria, combine studies that should not be pooled, inadequately address risk of bias, overlook missing results, or draw conclusions more certain than the underlying evidence permits.

The task is therefore not simply to inspect the pooled effect. You need to evaluate how the evidence entered the review, what happened to it during synthesis, and whether the final conclusion accurately represents what that body of evidence can support.

02 · The Short Answer

Evaluate the Evidence Pipeline, Not Just the Pooled Result

In Brief

To critically evaluate a systematic review or meta-analysis, examine whether the review asked a clear question, searched comprehensively, selected studies using defensible criteria, assessed risk of bias appropriately, synthesized sufficiently comparable evidence using suitable methods, investigated important heterogeneity and missing evidence, and drew conclusions proportionate to the certainty of the evidence.

A meta-analysis can increase precision by statistically combining compatible studies, but pooling does not erase weaknesses in the included research. A precise summary estimate can still be misleading when the evidence entering the analysis is biased, heterogeneous in consequential ways, incomplete, or inappropriate to combine.

03 · What You Need to Know

What Should You Actually Look for in a Systematic Review or Meta-Analysis?

First, Distinguish the Systematic Review From the Meta-Analysis

The terms are often used together, but they describe different things.

Systematic review A structured evidence synthesis that uses explicit methods to identify, select, critically appraise, and synthesize research addressing a defined question.
Meta-analysis A statistical synthesis that combines quantitative results from multiple studies when such combination is methodologically appropriate.

A systematic review does not have to contain a meta-analysis. Sometimes studies are too different in their populations, interventions or exposures, outcomes, designs, or effect measures for a pooled estimate to be meaningful. Cochrane guidance explicitly recognizes that not conducting a meta-analysis may be preferable when variation among study results makes an average effect misleading.

Conversely, calculating a pooled estimate from several studies does not by itself establish that the surrounding review process was systematic or rigorous.

Start With the Review Question and Eligibility Criteria

Before examining the forest plot, determine exactly what evidence the review intended to synthesize.

For intervention reviews, this is often expressed through elements such as population, intervention, comparator, and outcomes. Other types of systematic reviews may require different frameworks. Whatever structure is used, the eligibility criteria should follow logically from the review question.

Look for decisions about populations, settings, interventions or exposures, comparators, outcomes, study designs, publication status, language, and time period. Ask whether these criteria are sufficiently clear and whether important restrictions are justified.

Eligibility decisions determine what the review can eventually conclude. A beautifully executed meta-analysis of studies that poorly represent the question of interest will still answer the wrong or an unnecessarily narrow question.

Look for a Protocol or Prespecified Methods

Systematic reviews involve many methodological choices. Researchers decide what studies qualify, which outcomes matter, how results will be extracted, which analyses will be conducted, how subgroups will be defined, and how risk of bias will be handled.

Making important decisions after seeing the available evidence can create opportunities for selective analysis and interpretation. A prospectively specified protocol helps readers distinguish planned methods from decisions made after study results were known.

Look for protocol registration or publication and compare it with the completed review when feasible. Deviations are not automatically problematic. Research plans sometimes need to change. What matters is whether important changes are transparently reported and justified.

Ask Whether the Search Could Realistically Have Found the Relevant Evidence

A systematic review is only as comprehensive as the process used to identify eligible evidence. Missing relevant studies can change the apparent body of evidence, especially when availability is related to the direction or statistical significance of results.

Examine which bibliographic databases and other sources were searched, when the searches were conducted, and whether the search strategies appear appropriate to the topic. Depending on the question, additional sources might include trial registries, reference lists, citation searching, regulatory sources, conference material, preprints, or other grey literature.

Do not evaluate comprehensiveness merely by counting databases. Searching six poorly chosen sources is not necessarily better than searching three highly relevant ones with carefully developed strategies. The search needs to fit the research question and evidence base.

Watch Out

A systematic review cannot synthesize evidence it never finds. Studies or results may be easier to locate when their findings are statistically significant, favorable, rapidly published, or published in prominent journals. Missing evidence can therefore introduce systematic bias rather than merely reducing the number of studies available.

Study Selection Should Be Transparent and Reproducible

After searching, reviewers must decide which records meet the eligibility criteria. Look for a clear account of how records were screened and how disagreements or uncertainties were handled.

A PRISMA flow diagram can help you follow the movement of records through identification, screening, eligibility assessment, and final inclusion. PRISMA 2020 provides reporting guidance and flow-diagram templates intended to make these processes transparent.

But remember what PRISMA is: a reporting guideline. A completed PRISMA checklist or attractive flow diagram does not prove that the review methods were appropriate. Good reporting makes appraisal possible; it does not replace appraisal.

Read the Characteristics of the Included Studies Before the Pooled Estimate

One of the most useful habits when evaluating a meta-analysis is to resist jumping directly to the diamond at the bottom of the forest plot.

First ask what was actually combined. Are the studies investigating sufficiently similar populations, interventions or exposures, comparators, outcomes, time points, and designs? Are outcome definitions comparable? Are some studies randomized while others are observational? Do intervention doses, implementation conditions, or settings differ substantially?

Statistical software can calculate a pooled estimate whenever compatible numbers are supplied. The more difficult question is whether those numbers represent effects similar enough in meaning that averaging them produces an interpretable quantity.

Evaluate the Included Studies, Not Just the Review

A systematic review does not make methodological weaknesses in its included studies disappear. If most included studies have serious design problems, the synthesis may provide a more precise summary of biased evidence.

Look for a structured assessment of risk of bias using tools appropriate to the study designs and outcomes. For randomized trials, for example, Cochrane recommends its revised Risk of Bias tool, RoB 2, for assessing bias in trial results. Different designs require different approaches.

Risk-of-bias assessment should also affect interpretation. A review that identifies serious bias in the included evidence but then discusses the pooled result as if all studies were equally trustworthy has not completed the critical task.

When necessary, return to the principles involved in critically evaluating the individual research papers and consider whether their methodological problems merely qualify particular findings or fundamentally undermine them.

Do Not Assume More Studies Automatically Mean Better Evidence

A meta-analysis containing 50 studies may look more compelling than one containing five. The number alone tells you little about the credibility of the evidence.

Fifty small studies at high risk of bias do not automatically become high-quality evidence when pooled. Nor does an enormous combined participant count correct systematic problems shared across the studies.

The same logic applies to sample size more generally: large amounts of data do not automatically make research methodologically strong.

Understand What the Pooled Effect Represents

Most conventional meta-analysis methods calculate some form of weighted average of study effect estimates. Studies contribute different amounts of weight depending on the statistical method and information they provide.

Before interpreting the pooled estimate, identify the effect measure. It might be a risk ratio, odds ratio, hazard ratio, mean difference, standardized mean difference, correlation, prevalence estimate, or another statistic.

Then examine the confidence interval and the substantive magnitude of the effect. A statistically significant pooled result is not automatically important. Likewise, a non-significant pooled result does not necessarily demonstrate equivalence or absence of a meaningful effect.

The usual principles of recognizing overstated interpretations still apply to evidence synthesis.

Fixed-Effect and Random-Effects Models Answer Different Statistical Questions

You will often see a meta-analysis described as using a fixed-effect or random-effects model. These should not be treated as interchangeable buttons selected according to a single heterogeneity threshold.

A fixed-effect model assumes that the included study estimates address one common underlying effect, with observed differences attributed to sampling variation under the model. Random-effects approaches allow underlying effects to vary across studies according to a distribution and estimate an average effect across that distribution.

Random-effects meta-analysis does not make heterogeneity disappear. If effects differ substantially across settings or populations, an average may be statistically calculable yet have limited practical meaning unless the variation is understood and appropriately communicated.

Heterogeneity Is More Than an I-Squared Number

Heterogeneity refers to variation among study results. Some variation is expected, but substantial differences can affect what a pooled estimate means and how broadly it can be applied.

The I-squared statistic is commonly reported as an estimate of the proportion of observed variability in effect estimates that reflects between-study heterogeneity rather than sampling error under the model. It is useful, but it should not be interpreted through rigid thresholds alone.

Cochrane specifically cautions against using simple I-squared cutoffs to diagnose heterogeneity because interpretation depends on the magnitude and direction of effects, evidence for heterogeneity, and uncertainty in the statistic, particularly when few studies are available.

Statistical heterogeneity Observed variation in effect estimates beyond what would be expected from sampling variation under the statistical model.
Clinical or methodological diversity Differences among studies in populations, interventions, exposures, outcomes, settings, designs, procedures, or other characteristics that may help explain why effects differ.

When heterogeneity is important, ask what might explain it. Are effects different across populations? Intervention doses? Follow-up periods? Outcome definitions? Study quality? Settings?

If the effects point in substantially different directions, a single average deserves especially careful interpretation.

Prediction Intervals Can Be More Informative Than the Average Alone

In an appropriate random-effects meta-analysis, a confidence interval around the pooled effect describes uncertainty about the estimated mean effect. It does not directly describe the range within which effects in different settings might plausibly occur.

A prediction interval can help communicate the expected dispersion of underlying effects in a new study or setting under the model. This may substantially change interpretation when heterogeneity is considerable.

For example, the average effect might favor an intervention while a prediction interval includes effects ranging from meaningful benefit to little benefit or possible harm. In such a situation, “the meta-analysis shows the intervention works” would conceal consequential variation.

Subgroup Analyses Can Explain Heterogeneity, but They Can Also Generate Stories

When studies differ, review authors may investigate whether effects vary by population, intervention type, methodological characteristic, or another factor. Such analyses can be useful, but they are vulnerable to overinterpretation.

Subgroup analyses based on a small number of studies can be unstable. Multiple comparisons create opportunities for chance findings. Explanations developed only after inspecting the results are particularly vulnerable to data-driven interpretation.

Cochrane recommends caution even with prespecified investigations of heterogeneity and treats post hoc explorations primarily as hypothesis-generating.

When a review claims that an intervention “works only in younger participants” or “is effective only in high-income countries,” examine whether the evidence actually demonstrates a difference between subgroups rather than merely showing significance in one subgroup and non-significance in another.

Sensitivity Analyses Tell You Whether Reasonable Decisions Change the Answer

Systematic reviews contain judgment calls. Researchers may choose among statistical models, assumptions, eligibility decisions, definitions, methods for handling missing data, or approaches to studies at elevated risk of bias.

Sensitivity analyses examine whether the main finding remains similar when reasonable analytical decisions are changed.

If excluding studies at high risk of bias causes the pooled effect to disappear, that matters. If alternative defensible models produce substantially different estimates, that also matters. A conclusion that survives reasonable sensitivity analyses is generally more robust than one dependent on a narrow set of analytical choices.

Look for Evidence That Might Be Missing From the Synthesis

Publication bias is often discussed in meta-analysis, but the broader problem is missing evidence. Studies may remain unpublished, and particular outcomes or analyses within published studies may be selectively unavailable depending on their results.

Cochrane describes this more broadly as bias due to missing evidence. Results suggesting statistically significant or favorable effects may be more likely to become available rapidly, appear in prominent journals, or otherwise be easier to locate.

Funnel plots and statistical tests can sometimes help investigate small-study effects or asymmetry, but they are not definitive tests for publication bias. Their interpretation can be difficult when few studies are available or when asymmetry has causes other than selective publication.

Look at the review's broader strategy: Did the authors search trial registries or other relevant sources? Compare protocols or registrations with publications where appropriate? Consider whether missing results could materially change the synthesis?

Certainty of Evidence Is Not the Same as Statistical Significance

A pooled effect can be statistically significant while the overall certainty of evidence remains low.

The GRADE approach, widely used in systematic reviews of interventions, evaluates certainty for a body of evidence by considering domains including risk of bias, inconsistency, indirectness, imprecision, and publication bias. Certainty is categorized as high, moderate, low, or very low.

This distinction is crucial. The p-value associated with a pooled effect does not tell you whether the included studies were biased, whether their results are consistent, whether the evidence directly addresses the question, whether estimates are sufficiently precise, or whether relevant results may be missing.

If a review reports certainty assessments, examine the reasons for downgrading or upgrading rather than treating the final category as another unexplained score.

PRISMA Compliance Does Not Prove That the Review Is High Quality

PRISMA 2020 provides a 27-item reporting checklist, an expanded checklist, an abstract checklist, and flow-diagram templates. Its purpose is to improve transparent reporting of systematic reviews.

That transparency is extremely useful for appraisal. But PRISMA is not a methodological quality certificate.

A review could clearly report an inadequate search, inappropriate synthesis, or weak risk-of-bias assessment and still provide information corresponding to PRISMA items. Conversely, poor reporting can prevent you from determining whether methods were appropriate.

Use PRISMA to help determine whether the review tells you what was done. Then critically evaluate whether what was done was methodologically defensible.

The Review's Conclusion Should Reflect the Weakest Important Links in the Evidence Chain

At the end of the paper, compare the strength of the authors' language with the evidence you have just evaluated.

“The intervention reduces mortality” is a stronger claim than “the available evidence suggests that the intervention may reduce mortality.” The appropriate wording depends not merely on the pooled estimate but on risk of bias, precision, consistency, directness, missing evidence, and the broader certainty of the evidence.

Systematic reviews can look authoritative because they summarize many studies. That makes careful interpretation more important, not less.

04 · A Practical Example

A Significant Meta-Analysis Can Still Leave the Main Question Uncertain

Hypothetical Example

A Meta-Analysis of a New Educational Intervention

Suppose a systematic review identifies 12 studies examining whether a digital learning intervention improves academic performance. Ten studies are included in a meta-analysis.

Search and selection The review searches several major databases, reports explicit eligibility criteria, provides a reproducible search strategy, and documents study selection.
Pooled result The meta-analysis reports a statistically significant average effect favoring the intervention.
Look beneath the average The studies use different interventions, populations, follow-up periods, and outcome measures. Several have important risk-of-bias concerns, and the estimated effects vary substantially.
Sensitivity analysis When studies at high risk of bias are excluded, the estimated effect becomes smaller and substantially less precise.
Certainty The reviewers judge the overall evidence to have low certainty because of risk of bias, inconsistency, and imprecision.

The statistically significant primary meta-analysis is real, but “the intervention has been proven effective” would not accurately represent the evidence. A more defensible interpretation would acknowledge that the pooled findings suggest possible benefit while substantial uncertainty remains about the magnitude, consistency, and credibility of that benefit.

The critical appraisal changed because you examined the evidence underneath the pooled estimate rather than stopping at the diamond on the forest plot.

05 · What Researchers Often Get Wrong

Common Mistakes When Reading Systematic Reviews and Meta-Analyses

Misconception

A Meta-Analysis Is Automatically the Highest Level of Evidence

Evidence hierarchies can be useful shorthand, but the label does not determine credibility. A meta-analysis inherits the limitations of its evidence base and adds methodological decisions of its own. A poorly conducted synthesis of weak studies can provide less trustworthy evidence than a strong individual study addressing the question more directly.

Misconception

Combining Many Weak Studies Makes the Evidence Strong

Pooling can increase statistical precision, but it does not automatically eliminate systematic bias. If the included studies share serious methodological weaknesses, the pooled estimate may become more precise without becoming more credible.

Misconception

A Significant Pooled Effect Settles the Question

Statistical significance does not establish that the effect is important, unbiased, consistent, directly applicable, or supported by high-certainty evidence. Examine effect magnitude, uncertainty, risk of bias, heterogeneity, missing evidence, and certainty before deciding how strongly the result should influence your conclusion.

Misconception

A High I-Squared Value Automatically Invalidates the Meta-Analysis

Heterogeneity requires interpretation rather than a mechanical cutoff. Its importance depends on factors including the magnitude and direction of effects, uncertainty in the heterogeneity estimate, and the clinical or methodological differences among studies. In some cases pooling remains informative; in others an average may obscure differences that matter.

Misconception

A Random-Effects Model Solves Heterogeneity

A random-effects model allows for variation in underlying effects; it does not explain that variation or make it irrelevant. When effects differ meaningfully across studies, you still need to understand what the pooled average represents and whether the variation limits its applicability.

Misconception

A Funnel Plot Can Prove That There Is No Publication Bias

No. Funnel-plot asymmetry can have several causes, and an apparently symmetrical plot does not prove that all relevant evidence has been identified. These methods can contribute to an assessment of missing evidence but should not be interpreted as definitive tests.

Misconception

Following PRISMA Means the Review Is Methodologically Strong

PRISMA is primarily a reporting guideline. Good adherence makes important review methods and results more transparent, which helps you evaluate them. It does not certify that the search, study selection, risk-of-bias assessment, synthesis, or interpretation was methodologically appropriate.

06 · What This Means for You

Read the Review From the Evidence Up, Not From the Conclusion Down

A forest plot can compress years of research into one visually persuasive figure. Do not let that efficiency compress your appraisal as well.

Start with the question and eligibility criteria. Then examine how the evidence was found and selected, what kinds of studies entered the review, how trustworthy those studies were, whether synthesis was appropriate, how much the findings varied, whether important evidence may be missing, and how certain the resulting body of evidence actually is.

A simple decision framework

If the search appears incomplete
Treat the synthesis cautiously because the included studies may not adequately represent the available evidence.
If many included studies are at high risk of bias
Do not let the pooled sample size create false reassurance. Examine how risk of bias affects the synthesis and certainty of evidence.
If studies differ substantially
Ask whether pooling remains meaningful, investigate the direction and possible sources of heterogeneity, and examine prediction intervals where appropriate.
If the pooled effect is statistically significant
Examine its magnitude, precision, risk of bias, consistency, directness, possible missing evidence, and substantive importance before deciding how convincing it is.
If sensitivity analyses materially change the result
Recognize that the conclusion depends on methodological choices or particular studies and should be correspondingly less definitive.
If the review reports low or very low certainty
Interpret confident-sounding conclusions cautiously and examine which domains produced the uncertainty.

Do not reduce your appraisal to whether the review is “good” or “bad.” Different weaknesses affect different conclusions. A review may provide credible evidence for one outcome but uncertain evidence for another. It may establish an average effect reasonably well while providing little evidence about which populations benefit most.

When limitations are present, evaluate how much those limitations should change your confidence in the findings rather than treating every methodological imperfection as equally consequential.

07 · A Quick Checklist

What to Check Before Trusting a Systematic Review or Meta-Analysis

Before relying on the review's conclusions, check:
Is the review question clearly defined, and do the eligibility criteria logically match it?
Was a protocol or prespecified analysis plan available, and are important deviations explained?
Was the search sufficiently comprehensive for the topic, with reproducible strategies and appropriate sources?
Are study selection and reasons for exclusion reported transparently?
Was risk of bias assessed using methods appropriate to the included study designs and outcomes, and did those assessments affect interpretation?
If studies were meta-analyzed, are they sufficiently comparable for the pooled effect to have a meaningful interpretation?
Are heterogeneity and important differences among studies examined rather than reduced to a single I-squared threshold?
Were sensitivity analyses used where important methodological decisions could materially affect the result?
Did the reviewers consider bias from missing studies or results rather than assuming the identified literature is complete?
Are the final conclusions proportionate to the magnitude, precision, consistency, risk of bias, applicability, and overall certainty of the evidence?
08 · Frequently Asked Questions

Frequently Asked Questions About Systematic Reviews and Meta-Analyses

Is a systematic review the same as a meta-analysis?

No. A systematic review uses explicit methods to identify, select, appraise, and synthesize evidence addressing a defined question. Meta-analysis is a statistical method for quantitatively combining compatible study results. A systematic review may appropriately contain no meta-analysis.

Is a meta-analysis always stronger evidence than a single study?

No. A rigorous meta-analysis can provide a powerful synthesis of multiple studies, but its credibility depends on the quality, relevance, completeness, and comparability of the evidence and on the methods used to synthesize it. A poorly conducted meta-analysis of biased studies is not automatically superior to a strong individual study.

What does I-squared tell you in a meta-analysis?

I-squared helps describe heterogeneity among effect estimates, but it should not be interpreted through rigid thresholds alone. Its meaning depends on the magnitude and direction of effects, the number and precision of studies, and the clinical and methodological differences among them.

Does a random-effects meta-analysis solve heterogeneity?

No. A random-effects model incorporates an assumption that underlying effects vary across studies and estimates an average across that distribution. It does not explain why effects vary or guarantee that an average is clinically or scientifically meaningful.

What is publication bias in a meta-analysis?

Publication bias occurs when whether research becomes published is related to its results. The broader concern is bias due to missing evidence, which can also arise when particular outcomes or analyses are selectively unavailable. If missing evidence differs systematically from available evidence, the pooled estimate may be distorted.

What does a forest plot tell you?

A forest plot typically displays effect estimates and uncertainty intervals for individual studies and, when a meta-analysis is performed, a pooled estimate. It helps you inspect effect magnitude, precision, direction, study weights, and variation among results, but it does not by itself tell you whether the included studies are unbiased or appropriate to combine.

Does following PRISMA mean a systematic review is high quality?

No. PRISMA is designed to improve reporting of systematic reviews. Good reporting helps readers understand and evaluate what researchers did, but methodological quality still depends on decisions such as the search strategy, eligibility criteria, risk-of-bias assessment, synthesis methods, and interpretation.

What does certainty of evidence mean in a systematic review?

Certainty describes how confident reviewers can be in a body of evidence for a particular outcome. In the GRADE approach, considerations include risk of bias, inconsistency, indirectness, imprecision, and publication bias, leading to ratings of high, moderate, low, or very low certainty. Certainty therefore addresses more than whether a pooled result is statistically significant.

09 · The Bottom Line

A Pooled Estimate Is the End of a Methodological Process, Not the Beginning of Your Appraisal

The Bottom Line

A trustworthy systematic review or meta-analysis depends on how well researchers defined the question, found and selected the evidence, evaluated risk of bias, decided what could reasonably be synthesized, handled heterogeneity and missing evidence, and calibrated their conclusions to the certainty of the resulting evidence.

Do not equate more studies, more participants, a smaller p-value, or a narrower confidence interval with stronger evidence automatically. The pooled estimate is only as informative as the evidence and methodological decisions that produced it. Read the diamond, certainly, but read what went into the diamond first.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes