01 · The Question
What If There Seems to Be No Common Study to Compare?
You begin reviewing a research area and discover that almost every study is different. Populations vary. Researchers define the central construct differently. Interventions have the same name but different components. Outcomes range from achievement to satisfaction. Some studies last two weeks and others two years. Designs include experiments, surveys, interviews, case studies, and longitudinal observations.
At that point, synthesis can feel almost impossible. Any broad statement seems to ignore important differences, yet describing every study separately produces little more than an annotated bibliography.
The solution is not necessarily to eliminate the heterogeneity. In many literatures, variation is part of what needs to be understood. Your task is to determine which differences prevent comparison, which permit narrower comparisons, and which may actually help explain the pattern of findings.
03 · What You Need to Know
Heterogeneity Is Not One Kind of Difference
First identify what is actually heterogeneous
“The studies are very different” is too broad to guide a synthesis. You need to identify the dimensions of variation.
Studies may differ in participants, contexts, interventions, exposures, comparators, constructs, outcomes, measurement instruments, research designs, analytical procedures, implementation, follow-up periods, or risk of bias. These differences do not all have the same implications.
Cochrane distinguishes clinical diversity, such as variation in participants, interventions, and outcomes, from methodological diversity involving study design, outcome measurement, or risk of bias. Statistical heterogeneity refers more specifically to variation in estimated intervention effects beyond what would be expected from sampling error alone.
Diverse literature
The studies differ in characteristics such as populations, methods, contexts, constructs, or outcomes.
Uninterpretable literature
The available evidence is so disconnected, sparse, or incompatible that particular cross-study conclusions cannot be defended.
Heterogeneity does not automatically prevent synthesis
Variation can make some forms of synthesis inappropriate without making every form of synthesis impossible.
Suppose studies of peer mentoring differ in program duration, mentor training, student population, outcome measures, and research design. It may be inappropriate to speak of one uniform “effect of peer mentoring.” You may still be able to examine whether more intensive programs produce different patterns, whether findings vary across populations, which outcomes have been studied most consistently, or how participants experience different implementations.
The synthesis question changes from “What is the single result?” to “What pattern of results and variation does this evidence support?”
Find the highest defensible level of connection
When studies resist direct comparison, move upward conceptually until you reach a level at which meaningful relationships remain.
Two studies may not be comparable on effect magnitude because they measure different outcomes. They might still contribute to a broader question about whether an intervention influences academic functioning. An experiment and an interview study may not produce comparable results, but together they may inform effectiveness and implementation.
Be careful not to move so far upward that the synthesis becomes empty. “All studies concern education” is technically a commonality but analytically useless.
Create synthesis groups based on meaningful differences
Highly heterogeneous evidence often becomes more interpretable when divided into defensible groups. You might organize studies by intervention type, outcome domain, population, context, methodological design, construct definition, or another characteristic relevant to the review question.
The rationale matters. A subgroup should represent a distinction that could plausibly affect the finding or its interpretation, not merely a convenient way of distributing studies across headings.
For formal systematic reviews, Cochrane recommends considering whether planned comparisons need modification when clinical or methodological diversity makes the original grouping inappropriate. The rationale for such changes should be reported.
Do not confuse a broader synthesis method with permission to combine anything
If meta-analysis is inappropriate because the studies are too diverse, switching to prose does not make the diversity disappear. The conceptual problem remains.
Cochrane explicitly notes that concerns about diversity in populations, interventions, outcomes, or study designs apply to synthesis methods other than meta-analysis as well. In other words, an inappropriate comparison does not become appropriate because it is expressed narratively rather than statistically.
Watch Out
“The studies were heterogeneous, so a narrative synthesis was conducted” is not an analytical method by itself. You still need to explain which studies were considered together, on what basis, how their findings were compared, and how heterogeneity affected the conclusions.
Look for structured variation rather than perfect consistency
A heterogeneous literature may contain conditional patterns. Perhaps positive outcomes appear mainly in younger populations. Perhaps effects differ according to intervention intensity. Perhaps qualitative studies consistently identify an implementation problem that helps explain variation in quantitative findings.
These relationships can be more informative than an overall average because they begin to identify when, where, or under what conditions a phenomenon differs.
However, explanations for heterogeneity require caution. Cochrane warns that post-hoc subgroup analyses can generate hypotheses but may produce unreliable conclusions, particularly when there are few studies or many possible study characteristics that could explain the variation.
Do not discover an explanation merely because two patterns line up
Imagine that studies conducted in secondary schools report larger benefits than studies in universities. It is tempting to conclude that age explains the difference.
But perhaps the secondary-school studies also used more intensive interventions, shorter follow-up periods, or different outcome measures. Study-level characteristics are often correlated, making causal explanations for cross-study variation difficult to establish.
A responsible synthesis distinguishes an observed pattern from an explanation for that pattern.
Some heterogeneity reflects differences in terminology rather than phenomena
Researchers may use different labels for similar constructs or the same label for substantially different constructs. Before deciding that studies are fundamentally different, examine how key concepts are defined and operationalized.
Conceptual mapping can sometimes reveal that apparently fragmented studies are investigating related phenomena under different disciplinary vocabularies. The reverse is equally important: shared terminology should not be mistaken for conceptual equivalence.
Outcome heterogeneity requires particular care
A literature may appear consistently positive because every study reports some favorable result, even though each measures something different. One finds greater satisfaction, another improved achievement, another increased engagement, and another lower dropout.
Those outcomes should not be compressed into the conclusion that an intervention has a “positive impact.” Each outcome domain may require its own synthesis.
Statistical heterogeneity is not solved by an average
When quantitative studies estimate substantially different effects, an average can sometimes be useful, but it does not make the variation disappear. Cochrane notes that random-effects meta-analysis assumes studies estimate different but related effects and summarizes their average. It is not a substitute for investigating heterogeneity.
Where effects vary considerably, especially when their directions differ, reporting only the average may obscure important information. Cochrane advises that considerable variation can make an average effect misleading and identifies prediction intervals as one way of representing the spread of underlying effects when appropriate.
Heterogeneity can become a substantive conclusion
Sometimes the strongest finding of a review is that the field lacks sufficient conceptual or methodological consistency to support a broad conclusion.
If intervention definitions vary radically, outcomes are rarely repeated, populations are fragmented, and methods address different questions, that pattern says something important about the maturity of the evidence base.
The conclusion should be specific. “More research is needed” tells the reader little. “The literature cannot yet establish whether the intervention improves achievement because implementations and outcome measures are insufficiently comparable across studies” identifies what the problem actually is.
04 · A Practical Example
Finding Structure in a Literature Where Everything Seems Different
Hypothetical Example
Research on AI-supported learning in higher education
Imagine twelve hypothetical studies. Some examine generative AI tutors, others automated feedback or writing assistants. Participants come from medicine, programming, language learning, and general education. Outcomes include examination scores, writing quality, engagement, confidence, satisfaction, and self-regulation. Designs range from randomized experiments to interviews and observational studies.
The overly broad synthesis
“Overall, AI improves student learning, although findings vary across studies.”
The statement creates apparent simplicity by collapsing different technologies, outcomes, designs, and educational contexts into one undefined claim about “learning.”
A heterogeneity-sensitive synthesis
Map the variation Separate differences in AI function, learning outcome, discipline, study design, and implementation.
Find narrower comparisons Studies examining automated formative feedback may form one meaningful group, while conversational tutoring systems form another.
Separate outcome domains Achievement findings should not be merged conceptually with satisfaction, engagement, or perceived usefulness.
Integrate complementary evidence Qualitative studies may illuminate how students use these systems and identify implementation conditions relevant to interpreting quantitative outcomes.
Characterize the evidence base The hypothetical literature may support conclusions about particular AI functions and outcomes while remaining too heterogeneous to support a single conclusion about whether “AI improves learning.”
The heterogeneity has not been removed. It has been made analytically useful.