03 · What You Need to Know
The Population Matters Because It Can Change What the Study Tests
Replication Already Requires New Data
A common misunderstanding is that replication requires studying the same participants again. It does not.
The National Academies of Sciences, Engineering, and Medicine defines replicability as obtaining consistent results across studies aimed at answering the same scientific question, with each study obtaining its own data. Under that framework, new data are fundamental to replication.
This immediately separates two ideas that are sometimes confused:
New sample
A different set of participants from the population relevant to the research question.
New population
A group that differs from the original target population in some potentially meaningful characteristic, context, or eligibility definition.
Suppose an original study surveys 400 undergraduate students from a university, and you recruit another 400 undergraduate students meeting comparable criteria. You necessarily have a new sample, but you may still be sampling from a broadly similar population.
Now suppose you recruit secondary school students, working adults, older adults, nurses, teachers, or students in a substantially different educational system. The population itself has changed in a potentially meaningful way.
That does not automatically prevent replication. It does, however, raise a further question: what is the population change intended to test?
First Ask Whether the Scientific Question Is Still the Same
The National Academies' definition provides a useful starting point: replication concerns studies aimed at answering the same scientific question with new data. Generalizability, by contrast, concerns the extent to which study results apply in other contexts or populations that differ from the original one.
These concepts can overlap within one project.
Suppose an original study asks whether academic self-efficacy is associated with persistence among university students. You investigate the same relationship among another group of university students.
Your principal question remains:
Does the relationship between academic self-efficacy and persistence receive support in new data?
That is straightforwardly replication-oriented.
Now suppose you deliberately recruit working adult learners because you want to know whether the relationship also appears in a population whose educational circumstances differ from those of conventional university students. You still test the original relationship, but the population variation now also provides evidence relevant to its generalizability.
The distinction is not that replication uses the same people while generalizability research uses different people. Both use new data. The distinction concerns what inference the population difference is intended to support.
Changing the Population Can Test Generalizability
A result obtained in one population does not establish that the same result will necessarily apply to all other populations.
This is where population variation becomes scientifically useful.
The National Academies notes that selective variation in experimental conditions can be intentional. When results remain consistent across studies using somewhat different methods or conditions, confidence in the validity and possible generality of the finding can increase. Systematically varying important parameters can help researchers identify the limits of an effect rather than discovering only what works under one narrow set of circumstances.
Population changes can serve exactly this purpose.
For example, imagine that an intervention improves mathematics performance among first-year university students. Repeating the study among senior university students could examine whether the result extends across stages of undergraduate education. Testing it among secondary school students introduces a larger developmental and institutional difference. Testing it among adult continuing-education learners changes the context further still.
Each population change potentially asks something about the scope of the original claim.
But this only works when the difference is meaningful. "Nobody has tested this at our university" does not by itself establish an important generalizability question.
A New Location Is Not Necessarily a New Population
Geography creates one of the most common sources of confusion.
Suppose a study conducted in Country A reports a relationship between students' perceived instructor support and academic engagement. You conduct a similar study in Country B.
Is that automatically a new population in a scientifically meaningful sense?
Not necessarily.
Country labels can correspond to differences in educational systems, language, culture, socioeconomic conditions, institutional structures, access to technology, curriculum, or other factors relevant to the phenomenon. But researchers should not assume that every geographical boundary creates a theoretically important difference.
The stronger rationale identifies what about the new population could plausibly matter to the claim.
| Weak Rationale |
Stronger Rationale |
| The previous study was conducted in another country. |
The new educational context differs in a way that theory or prior evidence suggests could affect the relationship being tested. |
| No study has examined students at our university. |
The institution serves a population with a characteristic relevant to the proposed boundary or generalizability of the finding. |
| This topic has not been studied in our city. |
The local context exposes participants to conditions that plausibly alter the phenomenon under investigation. |
Local replication can be worthwhile, especially when local decisions depend on evidence obtained elsewhere. But local relevance and scientific novelty are different arguments. State which one you are making.
Ask Whether the Population Is Part of the Original Claim
Some claims are deliberately broad. Others are population-specific.
Suppose an original paper makes the broad claim that a particular learning strategy improves retention. If the evidence comes only from undergraduate students, testing the strategy among another relevant population can probe the breadth of that claim.
But suppose the original claim specifically concerns first-year university students undergoing the transition into higher education. Repeating the study among experienced postgraduate students may not test the same claim because the population characteristic is integral to the phenomenon being investigated.
This gives you a useful diagnostic question:
If I replace the original population with mine, does the original scientific claim still make sense without substantial rewriting?
If yes, the new population may still provide replication or generalizability evidence about that claim.
If no, you may be asking a different question.
Theoretical Relevance Matters More Than Demographic Difference Alone
Researchers often describe populations using characteristics such as age, sex, occupation, educational level, nationality, ethnicity, socioeconomic status, clinical status, or institutional type. A difference on one of these dimensions does not automatically create a theoretically meaningful population contrast.
The important question is whether the characteristic could plausibly affect the phenomenon being studied.
Imagine an original study of interface design and task completion among undergraduate students. You repeat it among university faculty. The occupational and age distributions will probably differ, but those differences matter scientifically only if they are relevant to how participants interact with the interface or to the claim being tested.
Conversely, changing from novice to expert users may be highly consequential if expertise is expected to alter the cognitive processes underlying task performance.
Population differences should therefore be justified through theory, prior evidence, practical relevance, or a defensible question about the scope of the original finding. Demographics should not be treated as explanatory merely because they are easy to put in Table 1.
Population Changes Can Reveal Boundary Conditions
A finding does not have to apply universally to be scientifically useful.
Suppose a relationship appears among adolescents but not older adults. That discrepancy could indicate that the phenomenon depends on developmental stage. An intervention that succeeds among novice learners but not experts could suggest that prior knowledge modifies its effectiveness.
These differences are often described as boundary conditions: conditions under which a claim does or does not appear to hold.
A deliberately chosen new population can therefore contribute more than a generic claim that "the study was replicated elsewhere." It can help specify the scope of the original finding.
This is one reason replication in a new population can remain useful even when earlier replications have already produced consistent results. If those studies repeatedly sampled very similar populations, an important uncertainty about generalizability may remain. Whether another study is worthwhile therefore depends on what uncertainty remains after previous successful replications.
Changing the Population Can Also Create an Extension
Now consider a more ambitious design.
An original study finds that academic self-efficacy predicts persistence among conventional university students. You recruit working adult learners and propose that employment demands weaken this relationship.
Your project now contains at least two questions:
Replication-related question Does the self-efficacy-persistence relationship appear in the new population?
Extension question Do employment demands alter the strength of that relationship?
The first provides evidence about the original relationship in a new population. The second introduces a claim the original study did not test.
Your project can therefore be both replication and extension. The categories are not mutually exclusive. The broader distinction between replication and extension depends on which claims the new study actually evaluates.
Adding Population Comparisons Can Change the Question Substantially
Suppose instead of merely repeating the original study in Population B, you recruit both Population A and Population B and explicitly hypothesize that the effect differs between them.
You are no longer asking only whether the original finding recurs.
You are asking whether population membership modifies the effect.
For example:
Original question: Is academic self-efficacy associated with persistence among university students?
Replication question: Does that association also appear among working adult learners?
Extension question: Is the association weaker among working adult learners than among conventional university students?
The comparison introduces an additional inferential target. Evidence that an association is statistically detectable in one group but not another is not, by itself, evidence that the groups differ. A group difference should be tested directly using an appropriate comparative or interaction analysis.
This distinction prevents a common interpretive error: "significant here, not significant there" does not automatically mean "significantly different."
A Population Change Can Become a Genuinely New Study
At some point, the population change may alter the scientific problem so substantially that describing the project primarily as a replication becomes misleading.
Imagine that previous research examined whether parental involvement predicts academic engagement among primary school children. You propose studying workplace supervisor support and employee engagement among nurses.
There may be conceptual parallels, but this is not simply the same claim transported to another population. The actors, constructs, context, theoretical mechanisms, and substantive question have changed.
A more subtle example occurs when the defining characteristic of the original population is central to the phenomenon. A study of transition anxiety among first-year university students cannot necessarily be converted into a replication among graduating students simply by retaining the same questionnaire. The phenomenon itself is tied to a different transition.
Watch Out
Do not classify a study as replication merely because you reuse the original questionnaire, statistical model, or variable names. Methodological resemblance is not enough if the population change alters the constructs, theoretical mechanisms, or scientific question being investigated.
The Same Population Change Can Support Different Research Designs
Consider a study originally conducted among undergraduate students in one educational system. You now have access to secondary school students in another system.
Several research designs are possible:
| Purpose of the Population Change |
Likely Characterization |
Main Question |
| Obtain another independent test in a broadly comparable population |
Replication |
Does the finding recur? |
| Deliberately test the same claim in a meaningfully different population |
Replication with a generalizability focus, depending on disciplinary terminology |
Does the finding hold in this different population? |
| Test why the effect differs across populations |
Replication-extension or extension |
What explains variation in the effect? |
| Compare effects directly between populations |
Extension or comparative study with a replication component |
Does population membership modify the relationship? |
| Use the new population to investigate substantially different constructs or mechanisms |
New study informed by previous research |
What happens in this different phenomenon or theoretical context? |
These labels can vary among fields. The important distinction is the research logic underneath them.
Do Not Use a New Population Merely to Manufacture Novelty
A particularly common research rationale takes this form:
"Previous studies have investigated X among students in other countries, but no study has investigated X among students at University Y."
That establishes an absence in the literature. It does not yet establish why filling that absence matters.
Perhaps University Y genuinely represents an informative context. Perhaps local decision-makers need local evidence. Perhaps its students differ in a theoretically relevant way. If so, explain that.
But if there is no reason to expect the finding to behave differently and no consequential local decision depends on the result, changing the institution may add relatively little scientific information.
The same principle applies more broadly when deciding whether replication or a genuinely new research question would address the more important uncertainty.
Replication Results Across Populations Should Not Be Reduced to Same or Different
Suppose the original study reports a positive effect, and your new-population study estimates an effect in the same direction but with a wider confidence interval. Did it replicate?
There is no universal binary rule.
The National Academies emphasizes that replicability should not necessarily be treated as a simple pass-or-fail judgment. Results should be interpreted with their uncertainty, and different disciplines use different standards for assessing consistency.
This is especially important across populations. Effect sizes may genuinely vary. A finding can be broadly consistent while differing in magnitude, or the evidence can remain too imprecise to determine whether the populations meaningfully differ.
Rather than asking only whether both studies crossed a statistical significance threshold, compare effect estimates, uncertainty, study design, measurement, and the substantive meaning of any differences. Where enough studies exist, cumulative evidence may be more informative than treating each replication as an isolated verdict.