Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Could Different Populations Explain Why Two Studies Disagree?

Two well-conducted studies can reach different conclusions because they studied different people. Learn when population differences provide a plausible explanation, when they do not, and how to investigate the possibility without inventing subgroup effects.

166
Can Population Differences Explain Conflicting Results? Guide 166 of 247
01 · The Question

Could both studies be right for the populations they actually studied?

You find two studies investigating what appears to be the same intervention or relationship. One reports a meaningful effect. The other finds little or none. Their methods seem reasonably credible, their outcomes appear comparable, and yet their conclusions differ.

Before deciding that one study must be wrong, look closely at who participated. Perhaps one sample consisted mainly of younger adults while the other involved older adults. One population may have had substantially higher baseline risk, greater disease severity, more prior experience, different socioeconomic circumstances, or exposure to a different institutional environment.

Those differences do not automatically explain conflicting findings. But sometimes the effect itself varies across populations. The important task is to determine whether population differences provide a plausible and evidence-supported explanation rather than using them as a convenient story after the results are known.

02 · The Short Answer

Yes, an effect can genuinely differ across populations

In Brief

Different populations can explain why studies reach different conclusions when characteristics of those populations modify the effect being studied, alter baseline risk, change exposure to the intervention, or otherwise affect how the outcome arises. The relevant question is not simply whether the samples differ, but whether those differences could plausibly produce the observed difference in effects.

Compare population characteristics systematically, examine the magnitude and uncertainty of the study estimates, look for credible evidence of effect modification, and distinguish population-related explanations from other possibilities such as measurement, design, analysis, bias, or chance.

03 · What You Need to Know

How population differences can change what a study finds

Research findings are always findings about someone

A result such as "the intervention improved outcomes" sounds general, but the underlying evidence came from a particular sample drawn under particular eligibility and recruitment conditions.

Who participated matters because populations can differ in characteristics related to the outcome or to the intervention's effectiveness. Depending on the research question, potentially relevant characteristics might include age, sex, baseline risk, disease severity, prior knowledge, socioeconomic circumstances, comorbidities, previous treatment, language, culture, occupation, educational level, or other factors.

Not every difference matters. The analytical task is to identify characteristics that could reasonably affect the relationship being investigated.

Different baseline risks can produce different absolute effects

One particularly important population difference is baseline risk: the probability of an outcome without the intervention or exposure of interest.

Suppose a treatment reduces relative risk by approximately the same proportion in two populations, but one population begins with a much higher risk of the adverse outcome. The absolute reduction can then be substantially larger in the higher-risk population.

Hypothetical Example

The same relative effect can produce different absolute benefits

Imagine that an intervention reduces the risk of an undesirable outcome by 20% relative to the comparison condition.

Lower-risk population If baseline risk is 10%, a 20% relative reduction corresponds to a risk of 8%, an absolute reduction of 2 percentage points.
Higher-risk population If baseline risk is 40%, the same 20% relative reduction corresponds to a risk of 32%, an absolute reduction of 8 percentage points.

The relative effect is identical in this simplified example, but the practical benefit is not. Depending on which effect measure researchers emphasize, the populations may therefore appear to tell somewhat different stories.

AHRQ's guidance on applicability specifically recommends considering whether differences in population characteristics affect baseline risks for benefits or harms and whether those differences alter the applicability of evidence to a target population.

Sometimes the relative effect itself changes across populations

Population differences can do more than alter baseline risk. A characteristic may change the magnitude, or occasionally the direction, of the effect itself. This is generally described as effect modification or effect heterogeneity.

Suppose an educational intervention requires substantial self-regulation. Students with considerable prior experience using independent learning strategies might benefit more than students encountering those demands for the first time. Alternatively, an intervention designed to remedy a particular deficit might produce larger gains among participants who have more room to improve.

The underlying principle is broader than either example: an effect estimated in one population should not automatically be assumed to have exactly the same magnitude in every other population.

Different baseline risk Populations differ in how frequently the outcome would occur without the intervention or exposure. This can change absolute effects even if a relative effect remains similar.
Effect modification A population characteristic changes the effect itself, so the intervention or exposure has a different relative or absolute impact across levels of that characteristic.

Age can matter, but "different ages" is not an explanation by itself

Researchers often notice demographic differences first. One study involved younger participants and another older participants, so age becomes an obvious candidate explanation.

That observation is only the beginning. You still need a plausible reason why age would modify the particular effect. Age might correspond to biological changes, accumulated experience, different baseline risks, comorbidities, technological familiarity, or exposure histories. For another research question, the same age difference may be largely irrelevant.

The same caution applies to sex, ethnicity, education, income, geographic region, and other characteristics. A demographic difference should not be promoted into an effect modifier merely because it is visible in the study's descriptive table.

Eligibility criteria can create populations unlike those encountered elsewhere

Study populations are partly constructed by inclusion and exclusion criteria. Researchers may restrict eligibility to create a more homogeneous sample, protect participants, improve measurement, reduce confounding, or answer a deliberately narrow question.

The resulting sample may differ considerably from the population in which the findings are later interpreted or applied.

AHRQ's framework for assessing applicability highlights restrictive eligibility criteria, exclusion of people with comorbidities, demographic differences, disease severity, and other population characteristics as factors that may limit how directly evidence applies outside the original study.

This distinction is closely related to external validity and generalizability. A study can be well conducted for the population it enrolled while providing limited evidence about substantially different populations.

Selection into the study can matter even when eligibility criteria look broad

Formal eligibility criteria are not the only determinants of who ends up in a study. Recruitment methods, willingness to participate, access to the study site, consent procedures, incentives, attrition, and other selection processes can shape the analyzed sample.

For example, a study evaluating a demanding digital intervention may disproportionately attract participants already comfortable with technology. A voluntary workplace wellness study may enroll employees who are more health-conscious than those who decline participation.

These differences can matter when the characteristics associated with participation are also related to the intervention, exposure, or outcome.

Population and setting can be difficult to separate

Suppose one study was conducted in highly resourced universities and another in institutions with limited technological infrastructure. Are different results caused by the students or by the institutions?

Possibly both.

Populations are embedded in settings. Geographic location, health systems, educational institutions, workplaces, communities, and historical periods can influence who participates and the conditions under which an intervention operates.

AHRQ therefore uses a PICOS framework when considering applicability: Population, Intervention, Comparator, Outcome, and Setting. Separating population from setting analytically can help prevent every contextual difference from being attributed to participant characteristics.

Do not confuse population differences with different research questions

At some point, populations may differ enough that the studies are better understood as answering distinct questions rather than producing competing estimates of one universal effect.

A study of an intervention among children and another among adults can still be meaningfully compared when age-related variation is itself of interest. But if the mechanisms, implementation, outcomes, or substantive purpose of the intervention differ greatly between those populations, calling the findings contradictory may obscure more than it reveals.

Before searching for effect modification, establish whether the studies represent a genuine contradiction rather than different research questions.

Subgroup findings require particular caution

If populations might respond differently, it is tempting to divide participants into subgroups until an explanation appears. That strategy is statistically dangerous.

Cochrane's guidance on heterogeneity warns that subgroup analyses can be misleading. When many characteristics are examined, some apparent subgroup differences can emerge by chance. Comparisons based on relatively few studies are also imprecise, and relationships observed across studies can be confounded by other study characteristics.

A convincing effect-modification claim should therefore have more support than "the effect was significant in subgroup A but not in subgroup B."

The relevant question is whether the effects differ between the groups. That generally requires an appropriate test of interaction or another direct analysis of effect modification, together with consideration of the estimate and its uncertainty.

Watch Out

Do not infer a subgroup difference merely because one subgroup has a statistically significant result and another does not. Evidence that effects differ between groups requires a direct comparison of those effects, and even a statistically detectable interaction still needs substantive interpretation.

Within-study evidence is often more informative than comparing separate studies

Suppose Study A includes mostly younger participants and reports a benefit, while Study B includes mostly older participants and reports little effect. It is tempting to conclude that age explains the difference.

But age is not the only way those studies differ. They may also use different settings, measures, intervention procedures, follow-up periods, or analytical methods.

If a sufficiently large individual study includes both younger and older participants and directly examines whether the effect differs by age, that within-study comparison can provide more direct evidence about effect modification because many other study-level differences are held constant.

Cross-study patterns can still be useful, especially when many studies are available, but they require cautious interpretation.

Population differences can explain heterogeneity without making either study wrong

In systematic reviews, variation in effects across studies is commonly called heterogeneity. Population characteristics are one possible source of that variation.

If credible evidence shows that an intervention produces a larger effect among participants with one characteristic than another, different study estimates may reflect real variation rather than methodological failure.

The synthesis then becomes more informative. Instead of asking whether the intervention works universally, you may be able to ask for whom it appears more or less effective.

That does not mean every heterogeneous literature can be neatly explained. Sometimes population characteristics explain only part of the variation, and sometimes no convincing explanation emerges.

Consistency across diverse populations can strengthen generalizability

Population diversity does not matter only when findings conflict. It can also strengthen interpretation when findings remain reasonably consistent.

AHRQ notes that a collection of studies representing different populations and settings can provide broader evidence of applicability than any single study. If comparable effects repeatedly appear across populations differing in relevant characteristics, confidence that the finding extends across those conditions may increase.

This is an important counterpoint to the search for subgroup differences. Diversity is not merely a source of troublesome heterogeneity. It can also test the boundaries of a finding.

04 · A Practical Example

How population differences can turn an apparent contradiction into a testable explanation

Hypothetical Example

A learning intervention works in one university population but not another

Imagine two hypothetical randomized studies evaluating the same digital study-planning intervention using the same achievement measure after one semester.

Study A The participants are predominantly first-year students entering university. The intervention produces a moderate improvement in achievement compared with the control condition.
Study B The participants are predominantly final-year students. The estimated effect is close to zero.
Initial interpretation The studies appear contradictory because the intervention, comparator, outcome, and study design are broadly similar.
Population hypothesis First-year students may have less established study-planning routines and therefore more opportunity to benefit from structured planning support. Final-year students may already use effective strategies, leaving less room for the intervention to add value.

That explanation is plausible, but plausibility is not evidence. You would next examine whether the studies actually differ in prior study skills, baseline performance, or related characteristics; whether either study reports relevant subgroup analyses; whether other studies show a similar pattern; and whether methodological differences provide alternative explanations.

If several credible studies subsequently show larger effects among students with weaker baseline planning skills, the population explanation becomes more persuasive. If the pattern cannot be reproduced, the original difference may instead reflect chance or some other study characteristic.

The correct conclusion at the beginning is therefore not "the intervention works only for first-year students." It is more cautious: differences in participant characteristics provide one plausible explanation for the conflicting estimates and warrant direct investigation.

05 · What Researchers Often Get Wrong

Common mistakes when explaining disagreement through population differences

Misconception

If the samples are different, the population difference explains the results

Studies almost always differ in some participant characteristics. An explanation requires more than identifying a difference: the characteristic should plausibly influence the effect, and the observed evidence should support that interpretation over competing explanations.

Misconception

A study in one population tells you nothing about another population

Evidence can often generalize beyond the exact participants studied. The relevant issue is whether characteristics that differ between the study population and target population are likely to alter the effect or its practical consequences. Generalizability is a reasoned judgment, not a requirement that every future population perfectly resemble the original sample.

Misconception

If a treatment works in one subgroup but not another, the subgroups respond differently

Separate significance results do not establish a subgroup difference. One subgroup may simply provide a more precise estimate than another. The effects themselves need to be compared directly, preferably using an appropriate interaction analysis.

Misconception

A demographic difference must be the explanation

Age, sex, ethnicity, socioeconomic position, and other demographic characteristics can be important, but they should not become automatic explanations for every conflicting result. Researchers should articulate the mechanism or substantive reason the characteristic could modify the effect and consider alternative explanations.

Misconception

A more representative sample is always the better study

Representativeness concerns applicability, whereas internal validity concerns whether the study supports a credible inference under the conditions examined. A narrowly selected study can provide strong evidence for its target population, while a broadly representative study can still suffer from serious bias. The two issues should be assessed separately.

Misconception

Once you find a plausible population difference, the disagreement is resolved

A plausible explanation remains a hypothesis until the evidence supports it. Differences in measures, research designs, analyses, implementation, bias, and sampling variation may provide alternative explanations. Do not stop investigating simply because one explanation sounds reasonable.

06 · What This Means for You

How to investigate whether populations explain conflicting findings

When two studies disagree, compare their participants deliberately rather than scanning demographic tables for any available difference. Begin with characteristics that could reasonably influence the phenomenon under investigation.

A simple decision framework

If the populations are broadly comparable on characteristics likely to affect the outcome
Population differences become a less compelling explanation; examine measurement, design, analysis, bias, and sampling variability.
If the populations differ on a characteristic with a plausible effect-modifying mechanism
Treat population heterogeneity as a candidate explanation and look for direct evidence that the effect actually varies across that characteristic.
If only one study reports a post hoc subgroup effect
Interpret it cautiously, particularly if many subgroups were examined or the subgroup sample sizes were small.
If an interaction is observed within appropriately designed studies and recurs across independent evidence
Confidence increases that population characteristics genuinely modify the effect.
If populations differ so substantially that the underlying research questions are no longer equivalent
Describe the findings as applying to different populations rather than forcing them into a single contradictory conclusion.
If several plausible explanations remain
Preserve the uncertainty. State that population differences may contribute without claiming they have been established as the cause.

Compare characteristics that matter for the research question

A useful population comparison might include eligibility criteria, recruitment methods, baseline outcome levels, demographic composition, prior exposure or treatment, disease or condition severity, comorbidities, socioeconomic characteristics, and other theoretically relevant variables.

Which characteristics belong on that list depends on the topic. A variable that matters greatly in one field may have little relevance in another.

AHRQ's applicability guidance explicitly rejects the idea that one universal rating scale can determine whether evidence applies across populations. Applicability depends on the question and on which differences could meaningfully alter benefits, harms, or other outcomes.

Look for convergence rather than one convenient subgroup result

A population explanation becomes more persuasive when several pieces of evidence point in the same direction: the proposed effect modifier has a credible mechanism, the interaction is observed directly, the pattern recurs in independent studies, and alternative methodological explanations are less convincing.

None of those features alone guarantees that the explanation is correct. Together, however, they can make the interpretation more defensible than a post hoc comparison based on one pair of studies.

Keep population differences separate from other sources of disagreement

Studies with different populations frequently differ in other ways too. They may use different instruments, interventions, follow-up periods, designs, or analytical strategies.

If the measurement approach also differs, investigate whether different measures could explain the conflicting conclusions. If designs differ, consider whether research design changes the inference being made.

The aim is not to find one explanation at all costs. It is to determine which explanation, or combination of explanations, the evidence can actually support.

07 · A Quick Checklist

Before blaming conflicting findings on different populations, check these points

When comparing study populations, check:
Do the studies use materially different eligibility or exclusion criteria?
Do participants differ in baseline risk, severity, prior exposure, experience, or other characteristics that could plausibly alter the effect?
Could recruitment or participation processes have selected systematically different kinds of participants?
Is there a substantive reason the population characteristic should modify the effect rather than merely correlate with the study result?
Is the proposed effect modification supported by a direct comparison or interaction analysis rather than separate significance tests?
Was the subgroup hypothesis prespecified, or did it emerge only after many possible subgroups were examined?
Do independent studies show a similar population-related pattern?
Could differences in measures, settings, designs, analyses, or risk of bias explain the disagreement instead?
Are the populations still sufficiently comparable that describing the findings as contradictory is meaningful?
08 · Frequently Asked Questions

Questions about population differences and conflicting research

Can two good studies disagree simply because they studied different populations?

Yes. Effects can vary across populations because of differences in baseline risk or genuine effect modification. But population differences are only one possible explanation, so researchers should examine whether the relevant characteristics plausibly and demonstrably influence the effect.

What population characteristics should I compare?

Compare characteristics relevant to the specific research question. These might include age, baseline risk, condition severity, prior experience or treatment, comorbidities, socioeconomic characteristics, eligibility criteria, recruitment processes, and other theoretically relevant factors. There is no universal list that matters equally for every topic.

What is effect modification?

Effect modification occurs when the magnitude or direction of an effect differs according to another characteristic. For example, an intervention might have a larger effect among participants with high baseline risk than among those with low baseline risk. Demonstrating such variation requires appropriate comparison of effects, not merely separate significance tests within subgroups.

Is effect modification the same as confounding?

No. Confounding is a source of distortion in estimating an association or effect because another factor is related to the exposure and outcome. Effect modification describes genuine variation in the effect across levels of another characteristic. The distinction is important because confounding is something researchers generally try to control, whereas genuine effect modification is a substantive finding to understand and report.

If one subgroup has a significant effect and another does not, is that evidence of effect modification?

Not by itself. The difference between a significant result and a nonsignificant result is not necessarily statistically significant. Researchers should compare the subgroup effects directly, commonly through an interaction test or an appropriate model estimating effect modification.

Does a narrow study population make a study low quality?

Not necessarily. Narrow eligibility may limit applicability to other populations without undermining the credibility of the estimate for the population actually studied. Internal validity and applicability are related but distinct considerations.

Can consistent findings across very different populations strengthen the evidence?

Potentially, yes. If credible studies repeatedly produce compatible findings across populations that differ in relevant characteristics and settings, this can increase confidence that the finding applies across a broader range of conditions. The studies still need to be sufficiently comparable for that interpretation to be meaningful.

What if population differences explain only some of the conflicting results?

That is entirely possible. Research heterogeneity can have multiple sources. Population characteristics may explain part of the variation while differences in measurement, design, analysis, implementation, bias, or sampling variability explain the rest. You do not need to force the literature into a single explanation.

09 · The Bottom Line

Different people can produce different effects, but that explanation must be demonstrated

The Bottom Line

Different populations can explain why studies disagree when participant characteristics alter baseline risk, modify the effect being studied, or otherwise change how the intervention, exposure, or outcome operates. Simply observing that two samples differ, however, is not enough to establish the explanation.

Compare relevant population characteristics, examine effect estimates and uncertainty, seek direct and preferably replicated evidence of effect modification, and consider competing methodological explanations. Sometimes the disagreement reveals that an effect genuinely varies across people or contexts. When the evidence does not establish why the results differ, the appropriate conclusion is uncertainty rather than an invented subgroup story.

10 · Sources and Further Reading

Sources and further reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes