03 · What You Need to Know
How population differences can change what a study finds
Research findings are always findings about someone
A result such as "the intervention improved outcomes" sounds general, but the underlying evidence came from a particular sample drawn under particular eligibility and recruitment conditions.
Who participated matters because populations can differ in characteristics related to the outcome or to the intervention's effectiveness. Depending on the research question, potentially relevant characteristics might include age, sex, baseline risk, disease severity, prior knowledge, socioeconomic circumstances, comorbidities, previous treatment, language, culture, occupation, educational level, or other factors.
Not every difference matters. The analytical task is to identify characteristics that could reasonably affect the relationship being investigated.
Different baseline risks can produce different absolute effects
One particularly important population difference is baseline risk: the probability of an outcome without the intervention or exposure of interest.
Suppose a treatment reduces relative risk by approximately the same proportion in two populations, but one population begins with a much higher risk of the adverse outcome. The absolute reduction can then be substantially larger in the higher-risk population.
Hypothetical Example
The same relative effect can produce different absolute benefits
Imagine that an intervention reduces the risk of an undesirable outcome by 20% relative to the comparison condition.
Lower-risk population If baseline risk is 10%, a 20% relative reduction corresponds to a risk of 8%, an absolute reduction of 2 percentage points.
Higher-risk population If baseline risk is 40%, the same 20% relative reduction corresponds to a risk of 32%, an absolute reduction of 8 percentage points.
The relative effect is identical in this simplified example, but the practical benefit is not. Depending on which effect measure researchers emphasize, the populations may therefore appear to tell somewhat different stories.
AHRQ's guidance on applicability specifically recommends considering whether differences in population characteristics affect baseline risks for benefits or harms and whether those differences alter the applicability of evidence to a target population.
Sometimes the relative effect itself changes across populations
Population differences can do more than alter baseline risk. A characteristic may change the magnitude, or occasionally the direction, of the effect itself. This is generally described as effect modification or effect heterogeneity.
Suppose an educational intervention requires substantial self-regulation. Students with considerable prior experience using independent learning strategies might benefit more than students encountering those demands for the first time. Alternatively, an intervention designed to remedy a particular deficit might produce larger gains among participants who have more room to improve.
The underlying principle is broader than either example: an effect estimated in one population should not automatically be assumed to have exactly the same magnitude in every other population.
Different baseline risk
Populations differ in how frequently the outcome would occur without the intervention or exposure. This can change absolute effects even if a relative effect remains similar.
Effect modification
A population characteristic changes the effect itself, so the intervention or exposure has a different relative or absolute impact across levels of that characteristic.
Age can matter, but "different ages" is not an explanation by itself
Researchers often notice demographic differences first. One study involved younger participants and another older participants, so age becomes an obvious candidate explanation.
That observation is only the beginning. You still need a plausible reason why age would modify the particular effect. Age might correspond to biological changes, accumulated experience, different baseline risks, comorbidities, technological familiarity, or exposure histories. For another research question, the same age difference may be largely irrelevant.
The same caution applies to sex, ethnicity, education, income, geographic region, and other characteristics. A demographic difference should not be promoted into an effect modifier merely because it is visible in the study's descriptive table.
Eligibility criteria can create populations unlike those encountered elsewhere
Study populations are partly constructed by inclusion and exclusion criteria. Researchers may restrict eligibility to create a more homogeneous sample, protect participants, improve measurement, reduce confounding, or answer a deliberately narrow question.
The resulting sample may differ considerably from the population in which the findings are later interpreted or applied.
AHRQ's framework for assessing applicability highlights restrictive eligibility criteria, exclusion of people with comorbidities, demographic differences, disease severity, and other population characteristics as factors that may limit how directly evidence applies outside the original study.
This distinction is closely related to external validity and generalizability. A study can be well conducted for the population it enrolled while providing limited evidence about substantially different populations.
Selection into the study can matter even when eligibility criteria look broad
Formal eligibility criteria are not the only determinants of who ends up in a study. Recruitment methods, willingness to participate, access to the study site, consent procedures, incentives, attrition, and other selection processes can shape the analyzed sample.
For example, a study evaluating a demanding digital intervention may disproportionately attract participants already comfortable with technology. A voluntary workplace wellness study may enroll employees who are more health-conscious than those who decline participation.
These differences can matter when the characteristics associated with participation are also related to the intervention, exposure, or outcome.
Population and setting can be difficult to separate
Suppose one study was conducted in highly resourced universities and another in institutions with limited technological infrastructure. Are different results caused by the students or by the institutions?
Possibly both.
Populations are embedded in settings. Geographic location, health systems, educational institutions, workplaces, communities, and historical periods can influence who participates and the conditions under which an intervention operates.
AHRQ therefore uses a PICOS framework when considering applicability: Population, Intervention, Comparator, Outcome, and Setting. Separating population from setting analytically can help prevent every contextual difference from being attributed to participant characteristics.
Do not confuse population differences with different research questions
At some point, populations may differ enough that the studies are better understood as answering distinct questions rather than producing competing estimates of one universal effect.
A study of an intervention among children and another among adults can still be meaningfully compared when age-related variation is itself of interest. But if the mechanisms, implementation, outcomes, or substantive purpose of the intervention differ greatly between those populations, calling the findings contradictory may obscure more than it reveals.
Before searching for effect modification, establish whether the studies represent a genuine contradiction rather than different research questions.
Subgroup findings require particular caution
If populations might respond differently, it is tempting to divide participants into subgroups until an explanation appears. That strategy is statistically dangerous.
Cochrane's guidance on heterogeneity warns that subgroup analyses can be misleading. When many characteristics are examined, some apparent subgroup differences can emerge by chance. Comparisons based on relatively few studies are also imprecise, and relationships observed across studies can be confounded by other study characteristics.
A convincing effect-modification claim should therefore have more support than "the effect was significant in subgroup A but not in subgroup B."
The relevant question is whether the effects differ between the groups. That generally requires an appropriate test of interaction or another direct analysis of effect modification, together with consideration of the estimate and its uncertainty.
Watch Out
Do not infer a subgroup difference merely because one subgroup has a statistically significant result and another does not. Evidence that effects differ between groups requires a direct comparison of those effects, and even a statistically detectable interaction still needs substantive interpretation.
Within-study evidence is often more informative than comparing separate studies
Suppose Study A includes mostly younger participants and reports a benefit, while Study B includes mostly older participants and reports little effect. It is tempting to conclude that age explains the difference.
But age is not the only way those studies differ. They may also use different settings, measures, intervention procedures, follow-up periods, or analytical methods.
If a sufficiently large individual study includes both younger and older participants and directly examines whether the effect differs by age, that within-study comparison can provide more direct evidence about effect modification because many other study-level differences are held constant.
Cross-study patterns can still be useful, especially when many studies are available, but they require cautious interpretation.
Population differences can explain heterogeneity without making either study wrong
In systematic reviews, variation in effects across studies is commonly called heterogeneity. Population characteristics are one possible source of that variation.
If credible evidence shows that an intervention produces a larger effect among participants with one characteristic than another, different study estimates may reflect real variation rather than methodological failure.
The synthesis then becomes more informative. Instead of asking whether the intervention works universally, you may be able to ask for whom it appears more or less effective.
That does not mean every heterogeneous literature can be neatly explained. Sometimes population characteristics explain only part of the variation, and sometimes no convincing explanation emerges.
Consistency across diverse populations can strengthen generalizability
Population diversity does not matter only when findings conflict. It can also strengthen interpretation when findings remain reasonably consistent.
AHRQ notes that a collection of studies representing different populations and settings can provide broader evidence of applicability than any single study. If comparable effects repeatedly appear across populations differing in relevant characteristics, confidence that the finding extends across those conditions may increase.
This is an important counterpoint to the search for subgroup differences. Diversity is not merely a source of troublesome heterogeneity. It can also test the boundaries of a finding.
06 · What This Means for You
How to investigate whether populations explain conflicting findings
When two studies disagree, compare their participants deliberately rather than scanning demographic tables for any available difference. Begin with characteristics that could reasonably influence the phenomenon under investigation.
A simple decision framework
If the populations are broadly comparable on characteristics likely to affect the outcome
Population differences become a less compelling explanation; examine measurement, design, analysis, bias, and sampling variability.
If the populations differ on a characteristic with a plausible effect-modifying mechanism
Treat population heterogeneity as a candidate explanation and look for direct evidence that the effect actually varies across that characteristic.
If only one study reports a post hoc subgroup effect
Interpret it cautiously, particularly if many subgroups were examined or the subgroup sample sizes were small.
If an interaction is observed within appropriately designed studies and recurs across independent evidence
Confidence increases that population characteristics genuinely modify the effect.
If populations differ so substantially that the underlying research questions are no longer equivalent
Describe the findings as applying to different populations rather than forcing them into a single contradictory conclusion.
If several plausible explanations remain
Preserve the uncertainty. State that population differences may contribute without claiming they have been established as the cause.
Compare characteristics that matter for the research question
A useful population comparison might include eligibility criteria, recruitment methods, baseline outcome levels, demographic composition, prior exposure or treatment, disease or condition severity, comorbidities, socioeconomic characteristics, and other theoretically relevant variables.
Which characteristics belong on that list depends on the topic. A variable that matters greatly in one field may have little relevance in another.
AHRQ's applicability guidance explicitly rejects the idea that one universal rating scale can determine whether evidence applies across populations. Applicability depends on the question and on which differences could meaningfully alter benefits, harms, or other outcomes.
Look for convergence rather than one convenient subgroup result
A population explanation becomes more persuasive when several pieces of evidence point in the same direction: the proposed effect modifier has a credible mechanism, the interaction is observed directly, the pattern recurs in independent studies, and alternative methodological explanations are less convincing.
None of those features alone guarantees that the explanation is correct. Together, however, they can make the interpretation more defensible than a post hoc comparison based on one pair of studies.
Keep population differences separate from other sources of disagreement
Studies with different populations frequently differ in other ways too. They may use different instruments, interventions, follow-up periods, designs, or analytical strategies.
If the measurement approach also differs, investigate whether different measures could explain the conflicting conclusions. If designs differ, consider whether research design changes the inference being made.
The aim is not to find one explanation at all costs. It is to determine which explanation, or combination of explanations, the evidence can actually support.