01 · The Question
What If the Studies Are Related but Not Really Comparable?
You have several studies addressing the same broad research problem, but almost every important detail differs. One examines adolescents and another university students. One measures achievement and another engagement. Follow-up periods vary. Interventions that share the same label are implemented differently. Some studies are experimental, while others are observational.
You can see why all of them belong in the literature review, yet statements such as “the studies generally agree” begin to feel uncomfortable. Agree about what, exactly?
This is an important synthesis problem. Studies can be relevant to the same research question without being sufficiently comparable for every kind of cross-study conclusion. The solution is not necessarily to abandon synthesis. It is to narrow the level at which the evidence is connected.
03 · What You Need to Know
Comparability Is Specific to the Claim You Want to Make
Studies are not simply comparable or incomparable
Comparability is rarely an all-or-nothing property. Two studies may be comparable for one purpose and inappropriate to combine for another.
Suppose two studies examine online peer feedback. One measures examination performance, while the other examines students' perceptions of feedback quality. They are not directly comparable for a claim about achievement. They may still contribute to a broader synthesis concerning how peer feedback functions in online learning.
This means the first question should not be, “Can these studies be compared?” It should be, “Can these studies be compared for the particular claim I am trying to make?”
Relevance
The study contributes evidence related to the review question.
Comparability
The study is sufficiently similar to other evidence on the dimensions necessary for a particular cross-study comparison or synthesis.
Identify exactly what differs
Calling studies “too different” is only the beginning of the analysis. Specify the source of the difference.
Studies may vary in population, context, intervention or exposure, comparator, outcome, measurement instrument, timing, research design, implementation, theoretical definition, analytical approach, or other characteristics. Each difference has different implications for synthesis.
Cochrane guidance similarly emphasizes defining comparisons for synthesis clearly enough that readers can understand which studies belong together and why. It also recommends specifying which outcomes and measures will be treated as belonging to the same outcome domain.
Conceptual similarity matters before statistical similarity
Two studies can report numbers that look easy to compare while measuring meaningfully different constructs. Conversely, two studies using different instruments may still operationalize a sufficiently similar construct for a particular synthesis.
Consider “engagement.” One study might define engagement as frequency of learning-management-system activity, another as behavioral participation during class, and another as a multidimensional construct involving behavioral, emotional, and cognitive engagement. Combining them merely because all three use the word engagement could create a misleading synthesis.
Before asking whether results point in the same direction, establish whether the underlying concepts are sufficiently aligned.
Population differences may change the meaning of a finding
A study of first-year university students and one involving experienced postgraduate students may investigate the same intervention but under meaningfully different conditions. Age, prior knowledge, institutional setting, socioeconomic context, discipline, or other characteristics may affect both implementation and outcomes.
This does not necessarily prevent any comparison. It means that a claim extending across populations should acknowledge whether the evidence supports such generalization.
Interventions with the same name may not be the same intervention
Broad labels can hide substantial variation. “Blended learning,” “peer mentoring,” “flipped classroom,” or “AI-assisted feedback” may refer to markedly different practices across studies.
If one intervention consists of weekly individualized mentoring for an entire semester and another involves two optional group sessions, describing their results as evidence for or against “mentoring” without examining implementation differences may be too coarse.
In such cases, intervention variation is not merely methodological clutter. It may help explain why outcomes differ.
Different outcomes should not be converted into one generic idea of success
A common synthesis error is to treat achievement, satisfaction, engagement, retention, confidence, and perceived usefulness as interchangeable indicators that an intervention “works.”
They are different outcomes. An intervention might improve satisfaction without changing achievement, or increase short-term engagement without affecting retention. A synthesis should preserve those distinctions.
Watch Out
Do not rescue a difficult synthesis by replacing several distinct outcomes with an undefined category such as “positive impact.” That can create apparent consistency by removing the very distinctions that make the evidence difficult to compare.
Create narrower synthesis groups when appropriate
When the full set of studies is too heterogeneous for one comparison, smaller groups may be more defensible.
You might group studies by population, intervention type, outcome domain, research design, context, follow-up period, or another characteristic that matters to the review question. The grouping criterion should have an analytical rationale rather than being chosen simply because it produces tidy categories.
Current Cochrane reporting guidance asks review authors to specify each comparison in enough detail that decisions about which studies belong in each synthesis could be replicated, along with the rationale for those comparisons.
Do not use “narrative synthesis” as permission to ignore comparability
If studies cannot be meta-analyzed, researchers sometimes assume they can simply describe all results together in prose. That does not solve the underlying problem.
Cochrane explicitly cautions that when meta-analysis is not possible, other synthesis methods should still be described in detail. The broad label “narrative synthesis” is insufficient because it does not specify how studies were grouped, compared, or interpreted. SWiM provides reporting guidance specifically for systematic reviews using synthesis methods other than meta-analysis.
Sometimes the studies are complementary rather than comparable
Not every useful relationship among studies is a comparison. One study may estimate whether an intervention changes an outcome. Another may examine how participants experience the intervention. A third may investigate implementation barriers.
Asking which study “found the stronger result” would make little sense. Yet together they can contribute to a broader understanding of effectiveness, experience, and implementation.
The synthesis then becomes integrative rather than directly comparative: what does each body of evidence contribute, and how do those contributions change the interpretation of the overall problem?
Noncomparability can itself be a finding about the literature
If researchers repeatedly study a nominally similar phenomenon using incompatible definitions, populations, measures, and outcomes, the inability to compare findings may reveal something consequential about the field.
Perhaps the construct lacks a stable operational definition. Perhaps researchers have prioritized different outcomes. Perhaps the intervention has no consistent implementation model. Perhaps evidence is fragmented across disciplinary traditions.
A literature review can legitimately identify this fragmentation as part of what is currently known about the evidence base.
Some comparisons should simply not be made
There is no requirement that every included study contribute to one overarching comparative statement. Cochrane guidance notes that studies whose results cannot be included in a synthesis should still be reported rather than disappearing from the review.
Sometimes the most defensible approach is to discuss a study separately, place it in another synthesis group, or state that differences prevent meaningful comparison on the outcome of interest.
04 · A Practical Example
When Five Studies Share a Topic but Not a Common Comparison
Hypothetical Example
Research on generative AI feedback in higher education
Imagine five hypothetical studies. Study A experimentally compares AI-generated and instructor feedback on essay scores. Study B surveys students about satisfaction with AI feedback. Study C interviews postgraduate students about trust in automated feedback. Study D examines revision behavior after students receive AI feedback in a programming course. Study E compares course completion before and after an AI feedback system is introduced.
The forced synthesis
“Most studies show that AI feedback has positive effects on students.”
The statement sounds synthetic but conceals several problems. The studies examine different outcomes, designs, disciplines, and forms of evidence. “Positive effects” collapses achievement, satisfaction, trust, revision behavior, and completion into one undefined outcome.
A more defensible synthesis
Separate the outcomes Study A contributes evidence about essay performance, Study E about completion, Study B about satisfaction, and Study D about revision behavior.
Preserve the qualitative contribution Study C provides evidence about how postgraduate students experience and evaluate the trustworthiness of automated feedback.
Identify legitimate connections Studies A and D may jointly inform how feedback relates to academic work, while Studies B and C illuminate different aspects of students' responses to the technology.
State the broader conclusion cautiously The hypothetical literature suggests that AI-generated feedback may influence several dimensions of students' learning and experience, but heterogeneity in outcomes and designs prevents those findings from being interpreted as one common effect.
The synthesis remains useful precisely because it does not pretend that all five studies answer the same question.
06 · What This Means for You
Decide What Can Be Connected Before Deciding What the Literature Says
Before comparing results, create a structured overview of the studies. Record the dimensions that matter for your question: population, context, intervention or exposure, comparator, outcome, construct definition, measurement, timing, and study design where relevant.
Then examine comparability separately for each claim you intend to make. This prevents broad thematic similarity from being mistaken for evidential equivalence.
A simple decision framework
If studies address sufficiently similar questions and outcomes
Compare them using a synthesis method appropriate to the evidence.
If studies differ in one important characteristic that may influence findings
Consider defensible subgroups and examine whether the difference helps explain variation.
If studies address different outcomes
Synthesize the outcomes separately rather than converting them into a generic measure of success.
If studies answer different but related questions
Treat their contributions as complementary and integrate them at the broader conceptual level.
If differences undermine the particular comparison you want to make
Do not make that comparison. State the limitation and identify what narrower conclusions remain possible.
The important discipline is to define the synthesis before interpreting its result. In formal systematic reviews, transparent reporting of which studies belong to each synthesis and why helps readers judge whether the resulting conclusions are defensible.