03 · What You Need to Know
A Credible Study Result Still Has Boundaries
Internal validity and generalizability ask different questions about evidence.
Internal validity concerns whether the design, conduct, measurement, and analysis support the inference being made for the study as conducted. In causal research, this commonly involves asking whether the estimated effect is distorted by confounding, selection processes, measurement problems, missing data, or other sources of bias.
Generalizability asks whether that internally credible inference can be extended from the study sample to a specified target population from which the sample is conceptually drawn. Closely related terminology distinguishes transportability when researchers want to extend the inference to a different external population.
Internal validity
Is the inference credible for the study participants and conditions actually investigated?
Generalizability
Can that inference reasonably be extended from the study sample to a specified target population?
A study can therefore perform very well on the first question without resolving the second.
Generalizable to Whom?
One of the most important questions in external validity is surprisingly easy to omit: what is the target population?
Saying that a study is “generalizable” without identifying where the inference is supposed to go is incomplete. Methodological work on target validity emphasizes that generalizability is a relationship between a study sample and a target population for a particular question, rather than an inherent characteristic that a study simply possesses.
A study might generalize reasonably well to students at similar urban universities but poorly to secondary-school students. It might apply to institutions with comparable technological infrastructure while providing little evidence for institutions with limited connectivity. The same study can therefore have different degrees of external validity for different targets.
This is why the broader distinction between internal and external validity should always be connected to the inference the researcher wants to make.
Why Can an Internally Valid Effect Fail to Generalize?
The most important reason is that effects may vary across people or circumstances.
Suppose an intervention works better among participants with substantial prior knowledge than among beginners. If the study sample contains mostly experienced participants but the target population contains mostly beginners, the average effect observed in the study may not equal the average effect in the target population.
Variables associated with such differences are often described as effect modifiers. Differences in the distribution of relevant effect modifiers between the study sample and target population can therefore create an external-validity problem even when the original study estimate is internally valid.
This connects generalizability directly to effect modification and heterogeneous effects. The issue is not merely whether study participants look demographically different from everyone else. The consequential question is whether those differences change the effect or relationship being generalized.
An Unrepresentative Sample Does Not Automatically Destroy Internal Validity
Researchers sometimes assume that a nonrepresentative sample makes a study invalid. That conclusion mixes two different inferential problems.
Suppose a randomized experiment recruits volunteers from one university. The volunteers may differ substantially from university students nationally. That can restrict population-level generalization.
Yet successful random assignment within the volunteer sample may still support a credible comparison between the intervention and control conditions for those participants, provided other important sources of bias are addressed.
Randomization and representative sampling serve different purposes. Random assignment concerns comparability of treatment groups and can strengthen causal inference within the study. Probability sampling concerns how observations represent a target population. One does not automatically provide the other.
A Representative Sample Does Not Automatically Create Internal Validity Either
The reverse mistake is also possible.
A study could obtain an excellent probability sample of a national population and still produce a biased causal estimate if the exposure groups differ because of confounding, the outcome is measured systematically differently, or other internal-validity threats remain.
Representativeness cannot rescue a distorted effect estimate.
Watch Out
Sampling and treatment assignment solve different methodological problems. A representative sample does not automatically establish a causal effect, and successful random assignment does not automatically establish that the effect applies to every population.
Eligibility Criteria Can Narrow the Population Represented
Studies often use eligibility criteria for good reasons. Researchers may exclude participants when an intervention could be unsafe, when a condition would make the outcome difficult to interpret, or when the study addresses a deliberately narrow population.
The resulting evidence may be internally credible but apply directly to a more restricted group than readers initially assume.
Consider a trial of a learning intervention that excludes students with limited proficiency in the language of instruction, students who have previously failed the course, and students who cannot attend every intervention session. Those restrictions may simplify implementation and interpretation, but the resulting sample no longer represents all students who might encounter the intervention in ordinary practice.
The appropriate response is not to declare the study invalid. It is to define the population to which the evidence applies and avoid extending the findings automatically to excluded groups.
Setting Can Matter as Much as Participant Characteristics
External validity is not only about who participated.
The intervention may depend on instructor expertise, institutional resources, implementation fidelity, class size, infrastructure, incentives, organizational culture, or other contextual conditions.
An intervention evaluated with intensive researcher supervision may perform differently when adopted routinely. A technology requiring high-speed connectivity may produce different results where access is unreliable. A teaching strategy implemented by specially trained instructors may not yield the same average effect when instructors receive minimal preparation.
Generalization therefore involves considering relevant variation in settings and implementation conditions as well as participant composition.
The Outcome and Intervention Can Also Define the Boundary
Researchers sometimes generalize beyond the intervention or outcome that was actually studied.
Suppose an eight-week intervention improves performance on a specific standardized assessment. That does not automatically establish that the intervention improves long-term retention, graduation, workplace performance, or every other plausible educational outcome.
Likewise, evidence for one implementation of an intervention does not necessarily transfer unchanged after substantial changes to intensity, duration, personnel, or delivery mode.
External-validity claims should therefore specify not only the target population but also what treatment, outcome, setting, and time frame are being generalized.
Generalizability and Transportability Are Related but Can Be Distinguished
Modern causal-inference literature often uses a more precise distinction between generalizability and transportability.
Generalizability typically refers to extending an inference from a study sample to a target population of which that sample is a subset. Transportability refers to extending an inference to an external target population from which the study sample was not drawn.
Suppose researchers recruit a subset of students from a national university system and want to estimate the effect for all eligible students in that system. That can be framed as a generalizability problem.
If they instead want to apply the result to students in another country with a different higher-education system, the problem is closer to transportability.
The terminology is not used identically in every field, so researchers should define the target population and inferential objective rather than relying on the label alone.
Generalizability Is Not Determined by Sample Size Alone
A sample of 20,000 participants can still provide weak evidence about a target population if important effect modifiers are distributed very differently between the sample and target.
Conversely, a smaller study may provide useful evidence for a narrowly defined target population when participants and relevant conditions correspond closely to that target.
Large samples primarily improve statistical precision under appropriate assumptions. They do not automatically solve selection into the study or differences between the study and target populations.
Researchers Can Sometimes Generalize Analytically
External validity is not always limited to a qualitative judgment that “more research is needed.”
When researchers have suitable information about the study and target populations, and relevant assumptions are defensible, methods such as standardization, weighting, and related generalizability or transportability approaches can estimate effects for a specified target population.
These methods require identifying and measuring variables related to differences in effects and study participation, along with assumptions concerning selection, consistency, positivity, and other aspects of the causal model.
They therefore do not magically convert any narrow study into universal evidence. Like statistical adjustment for confounding, external-validity adjustment is only as credible as the information and assumptions on which it depends.
Limited Generalizability Is Not Automatically a Design Flaw
A study may intentionally address a narrow scientific question.
An early explanatory experiment might ask whether an intervention can produce an effect under carefully specified conditions. A later pragmatic study might ask how well it works across routine settings. Both can be valuable because they answer different questions.
Problems arise when researchers make claims that extend substantially beyond what their design supports.
A narrowly focused study is not inherently defective because its direct external validity is limited. Whether that limitation represents a serious problem depends on the purpose of the study and the claims being made. This is part of the broader distinction between a design flaw and a defensible methodological trade-off.