Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Handle Studies That Cannot Be Directly Compared?

Studies do not need to be directly comparable to contribute to the same literature review. The key is to identify what can legitimately be connected while preserving differences that make direct comparison inappropriate.

211
Synthesizing Noncomparable Studies Guide 211 of 247
01 · The Question

What If the Studies Are Related but Not Really Comparable?

You have several studies addressing the same broad research problem, but almost every important detail differs. One examines adolescents and another university students. One measures achievement and another engagement. Follow-up periods vary. Interventions that share the same label are implemented differently. Some studies are experimental, while others are observational.

You can see why all of them belong in the literature review, yet statements such as “the studies generally agree” begin to feel uncomfortable. Agree about what, exactly?

This is an important synthesis problem. Studies can be relevant to the same research question without being sufficiently comparable for every kind of cross-study conclusion. The solution is not necessarily to abandon synthesis. It is to narrow the level at which the evidence is connected.

02 · The Short Answer

Connect What Is Comparable and Preserve What Is Not

In Brief

When studies cannot be directly compared, do not force them into a common verdict. Identify the specific dimensions on which comparison is defensible, group studies accordingly, and synthesize complementary or distinct contributions separately where necessary.

Noncomparability does not make a study irrelevant. It changes the kinds of claims you can make across studies and may itself reveal important variation in how a phenomenon has been defined, investigated, implemented, or measured.

03 · What You Need to Know

Comparability Is Specific to the Claim You Want to Make

Studies are not simply comparable or incomparable

Comparability is rarely an all-or-nothing property. Two studies may be comparable for one purpose and inappropriate to combine for another.

Suppose two studies examine online peer feedback. One measures examination performance, while the other examines students' perceptions of feedback quality. They are not directly comparable for a claim about achievement. They may still contribute to a broader synthesis concerning how peer feedback functions in online learning.

This means the first question should not be, “Can these studies be compared?” It should be, “Can these studies be compared for the particular claim I am trying to make?”

Relevance The study contributes evidence related to the review question.
Comparability The study is sufficiently similar to other evidence on the dimensions necessary for a particular cross-study comparison or synthesis.

Identify exactly what differs

Calling studies “too different” is only the beginning of the analysis. Specify the source of the difference.

Studies may vary in population, context, intervention or exposure, comparator, outcome, measurement instrument, timing, research design, implementation, theoretical definition, analytical approach, or other characteristics. Each difference has different implications for synthesis.

Cochrane guidance similarly emphasizes defining comparisons for synthesis clearly enough that readers can understand which studies belong together and why. It also recommends specifying which outcomes and measures will be treated as belonging to the same outcome domain.

Conceptual similarity matters before statistical similarity

Two studies can report numbers that look easy to compare while measuring meaningfully different constructs. Conversely, two studies using different instruments may still operationalize a sufficiently similar construct for a particular synthesis.

Consider “engagement.” One study might define engagement as frequency of learning-management-system activity, another as behavioral participation during class, and another as a multidimensional construct involving behavioral, emotional, and cognitive engagement. Combining them merely because all three use the word engagement could create a misleading synthesis.

Before asking whether results point in the same direction, establish whether the underlying concepts are sufficiently aligned.

Population differences may change the meaning of a finding

A study of first-year university students and one involving experienced postgraduate students may investigate the same intervention but under meaningfully different conditions. Age, prior knowledge, institutional setting, socioeconomic context, discipline, or other characteristics may affect both implementation and outcomes.

This does not necessarily prevent any comparison. It means that a claim extending across populations should acknowledge whether the evidence supports such generalization.

Interventions with the same name may not be the same intervention

Broad labels can hide substantial variation. “Blended learning,” “peer mentoring,” “flipped classroom,” or “AI-assisted feedback” may refer to markedly different practices across studies.

If one intervention consists of weekly individualized mentoring for an entire semester and another involves two optional group sessions, describing their results as evidence for or against “mentoring” without examining implementation differences may be too coarse.

In such cases, intervention variation is not merely methodological clutter. It may help explain why outcomes differ.

Different outcomes should not be converted into one generic idea of success

A common synthesis error is to treat achievement, satisfaction, engagement, retention, confidence, and perceived usefulness as interchangeable indicators that an intervention “works.”

They are different outcomes. An intervention might improve satisfaction without changing achievement, or increase short-term engagement without affecting retention. A synthesis should preserve those distinctions.

Watch Out

Do not rescue a difficult synthesis by replacing several distinct outcomes with an undefined category such as “positive impact.” That can create apparent consistency by removing the very distinctions that make the evidence difficult to compare.

Create narrower synthesis groups when appropriate

When the full set of studies is too heterogeneous for one comparison, smaller groups may be more defensible.

You might group studies by population, intervention type, outcome domain, research design, context, follow-up period, or another characteristic that matters to the review question. The grouping criterion should have an analytical rationale rather than being chosen simply because it produces tidy categories.

Current Cochrane reporting guidance asks review authors to specify each comparison in enough detail that decisions about which studies belong in each synthesis could be replicated, along with the rationale for those comparisons.

Do not use “narrative synthesis” as permission to ignore comparability

If studies cannot be meta-analyzed, researchers sometimes assume they can simply describe all results together in prose. That does not solve the underlying problem.

Cochrane explicitly cautions that when meta-analysis is not possible, other synthesis methods should still be described in detail. The broad label “narrative synthesis” is insufficient because it does not specify how studies were grouped, compared, or interpreted. SWiM provides reporting guidance specifically for systematic reviews using synthesis methods other than meta-analysis.

Sometimes the studies are complementary rather than comparable

Not every useful relationship among studies is a comparison. One study may estimate whether an intervention changes an outcome. Another may examine how participants experience the intervention. A third may investigate implementation barriers.

Asking which study “found the stronger result” would make little sense. Yet together they can contribute to a broader understanding of effectiveness, experience, and implementation.

The synthesis then becomes integrative rather than directly comparative: what does each body of evidence contribute, and how do those contributions change the interpretation of the overall problem?

Noncomparability can itself be a finding about the literature

If researchers repeatedly study a nominally similar phenomenon using incompatible definitions, populations, measures, and outcomes, the inability to compare findings may reveal something consequential about the field.

Perhaps the construct lacks a stable operational definition. Perhaps researchers have prioritized different outcomes. Perhaps the intervention has no consistent implementation model. Perhaps evidence is fragmented across disciplinary traditions.

A literature review can legitimately identify this fragmentation as part of what is currently known about the evidence base.

Some comparisons should simply not be made

There is no requirement that every included study contribute to one overarching comparative statement. Cochrane guidance notes that studies whose results cannot be included in a synthesis should still be reported rather than disappearing from the review.

Sometimes the most defensible approach is to discuss a study separately, place it in another synthesis group, or state that differences prevent meaningful comparison on the outcome of interest.

04 · A Practical Example

When Five Studies Share a Topic but Not a Common Comparison

Hypothetical Example

Research on generative AI feedback in higher education

Imagine five hypothetical studies. Study A experimentally compares AI-generated and instructor feedback on essay scores. Study B surveys students about satisfaction with AI feedback. Study C interviews postgraduate students about trust in automated feedback. Study D examines revision behavior after students receive AI feedback in a programming course. Study E compares course completion before and after an AI feedback system is introduced.

The forced synthesis

“Most studies show that AI feedback has positive effects on students.”

The statement sounds synthetic but conceals several problems. The studies examine different outcomes, designs, disciplines, and forms of evidence. “Positive effects” collapses achievement, satisfaction, trust, revision behavior, and completion into one undefined outcome.

A more defensible synthesis

Separate the outcomes Study A contributes evidence about essay performance, Study E about completion, Study B about satisfaction, and Study D about revision behavior.
Preserve the qualitative contribution Study C provides evidence about how postgraduate students experience and evaluate the trustworthiness of automated feedback.
Identify legitimate connections Studies A and D may jointly inform how feedback relates to academic work, while Studies B and C illuminate different aspects of students' responses to the technology.
State the broader conclusion cautiously The hypothetical literature suggests that AI-generated feedback may influence several dimensions of students' learning and experience, but heterogeneity in outcomes and designs prevents those findings from being interpreted as one common effect.

The synthesis remains useful precisely because it does not pretend that all five studies answer the same question.

05 · What Researchers Often Get Wrong

Common Mistakes With Studies That Cannot Be Directly Compared

Misconception

“If the studies cannot be compared directly, they cannot be synthesized”

Direct comparison is only one form of synthesis. Studies may provide complementary evidence, contribute to separate subquestions, or reveal variation that becomes part of the larger interpretation.

Misconception

“If all studies concern the same topic, they belong in one synthesis”

Shared subject matter does not guarantee sufficient conceptual or methodological comparability. Determine whether the studies address sufficiently similar populations, constructs, interventions, outcomes, or questions for the particular claim being made.

Misconception

“Different positive outcomes can be summarized as an overall positive effect”

Improved satisfaction, achievement, retention, and confidence are not interchangeable outcomes. Combining them into a generic judgment of effectiveness can exaggerate consistency.

Misconception

“If meta-analysis is impossible, I should just count positive and negative studies”

Simple vote counting can be misleading because it ignores effect magnitude, precision, study size, and other important features. Cochrane specifically warns against inadvertent vote counting through statements such as “the majority of studies found a positive effect” unless an appropriate vote-counting method has actually been used.

Misconception

“Every included study has to fit one of my main comparison groups”

A relevant study may provide evidence that cannot be incorporated into a particular synthesis. It should still be reported transparently rather than forced into an inappropriate comparison or silently omitted.

06 · What This Means for You

Decide What Can Be Connected Before Deciding What the Literature Says

Before comparing results, create a structured overview of the studies. Record the dimensions that matter for your question: population, context, intervention or exposure, comparator, outcome, construct definition, measurement, timing, and study design where relevant.

Then examine comparability separately for each claim you intend to make. This prevents broad thematic similarity from being mistaken for evidential equivalence.

A simple decision framework

If studies address sufficiently similar questions and outcomes
Compare them using a synthesis method appropriate to the evidence.
If studies differ in one important characteristic that may influence findings
Consider defensible subgroups and examine whether the difference helps explain variation.
If studies address different outcomes
Synthesize the outcomes separately rather than converting them into a generic measure of success.
If studies answer different but related questions
Treat their contributions as complementary and integrate them at the broader conceptual level.
If differences undermine the particular comparison you want to make
Do not make that comparison. State the limitation and identify what narrower conclusions remain possible.

The important discipline is to define the synthesis before interpreting its result. In formal systematic reviews, transparent reporting of which studies belong to each synthesis and why helps readers judge whether the resulting conclusions are defensible.

07 · A Quick Checklist

Before Comparing Studies, Check Whether the Comparison Is Defensible

Before combining or comparing findings, check:
Are the studies addressing sufficiently similar research questions for the claim I want to make?
Are the underlying constructs defined in sufficiently comparable ways?
Do important population or contextual differences affect interpretation?
Are interventions or exposures that share a label actually similar enough to compare?
Am I keeping distinct outcomes separate rather than calling all favorable results “positive effects”?
Would narrower synthesis groups represent the evidence more accurately?
Could apparently noncomparable studies contribute complementary rather than directly comparative evidence?
Have I reported relevant studies that cannot legitimately enter a particular synthesis?
08 · Frequently Asked Questions

Frequently Asked Questions About Noncomparable Studies

Can studies be included in the same literature review if they cannot be directly compared?

Yes. Inclusion depends on relevance to the review's scope or question, whereas direct comparability depends on the particular synthesis or claim. Relevant studies can contribute separately or complementarily even when direct comparison is inappropriate.

Can I compare studies that measure different outcomes?

You may discuss how the different outcomes contribute to a broader research problem, but avoid treating them as equivalent. Achievement, engagement, satisfaction, retention, and other outcomes should remain distinguishable unless there is a defensible methodological basis for combining them.

Can I compare studies from different populations?

Sometimes. Ask whether population differences are likely to affect the phenomenon or intervention being studied and whether your intended conclusion extends beyond what the evidence supports. Population variation may also become something to examine rather than something to average away.

What if two studies use the same construct name but different measures?

Examine what each instrument actually measures. Shared terminology does not guarantee conceptual equivalence. If the measures capture meaningfully different dimensions, separate synthesis or explicit qualification may be necessary.

Does inability to perform a meta-analysis mean the studies cannot be synthesized?

No. Other synthesis methods may be appropriate. However, Cochrane recommends specifying the actual method used rather than relying on a vague label such as “narrative synthesis,” and SWiM provides reporting guidance for systematic reviews using synthesis methods other than meta-analysis.

Should I exclude studies that do not fit my synthesis groups?

Not merely because they are inconvenient to synthesize. Eligibility and synthesis are separate decisions. If a study satisfies the review's inclusion criteria but cannot enter a particular synthesis, report it appropriately and explain the limitation where relevant.

Can noncomparability itself be a finding?

Yes. Persistent variation in definitions, interventions, outcomes, populations, or methods may reveal fragmentation or lack of standardization in a field. That can be an important conclusion if it is demonstrated systematically rather than inferred from a few inconvenient studies.

09 · The Bottom Line

Do Not Demand One Comparison From Evidence That Cannot Support It

The Bottom Line

When studies cannot be directly compared, identify the dimensions on which comparison remains defensible, create narrower synthesis groups where appropriate, and treat genuinely different evidence as complementary or separate rather than forcing it into one conclusion.

The inability to make one broad comparison does not mean the literature has nothing to say. Often, the differences themselves reveal where evidence is fragmented, where conclusions apply only under particular conditions, and where future research needs greater conceptual or methodological consistency.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes