03 · What You Need to Know
Internal Validity and Generalizability Are Different Goals, Not Opposite Ends of One Scale
Internal validity concerns whether the study supports a credible inference for the participants and conditions investigated. Generalizability concerns whether that inference extends to a specified target population.
Because the concepts address different inferential problems, they are not mathematically opposite quantities. Increasing one does not automatically decrease the other.
Still, some design choices affect both. A decision made to simplify causal interpretation may also change who is studied or how closely the study resembles routine conditions. That is where a genuine tension can arise.
Where Does the Trade-Off Idea Come From?
The traditional argument is straightforward.
Researchers seeking strong internal validity may create tightly controlled conditions. They might recruit a homogeneous sample, exclude participants with complicating conditions, standardize intervention delivery, use specially trained personnel, monitor adherence closely, or remove contextual variation.
These choices can reduce ambiguity about what happened inside the study. But if the intervention will eventually be used among heterogeneous populations under variable routine conditions, the experimental environment may represent only a narrow portion of that target.
This has led to the common contrast between explanatory studies, which tend to ask whether an intervention can work under selected or idealized conditions, and pragmatic studies, which tend to ask how an intervention performs under conditions closer to routine practice.
Restrictive Eligibility Criteria Can Narrow the Target Represented
Suppose researchers evaluate a digital learning intervention but enroll only students with high digital literacy, reliable personal computers, high-speed internet access, no previous course failures, and perfect availability for all intervention sessions.
Those restrictions may reduce variation and simplify implementation.
But if the intervention is intended for all students, the trial directly represents only a selected subgroup. Students who encounter the greatest implementation barriers have effectively disappeared from the evidence.
This is one reason a study can be internally valid without being broadly generalizable.
Not Every Restriction Improves Internal Validity
This point is easy to miss.
A narrow sample does not automatically make a study more internally valid. Restricting participants is useful only when the restriction addresses a genuine methodological, ethical, or scientific concern.
Excluding participants merely to create a more homogeneous sample may reduce heterogeneity without addressing bias. It can also reduce information about effect variation and narrow external validity unnecessarily.
Researchers should therefore be able to explain what each major restriction accomplishes. “To improve internal validity” is not sufficient if the connection between the restriction and a plausible threat to inference is unclear.
Standardization Can Strengthen Interpretation Without Necessarily Reducing Generalizability
Some forms of standardization are essential for fair comparisons and need not create unrealistic conditions.
Using the same outcome definition across groups, training assessors, applying consistent eligibility rules, or ensuring that measurements are taken at comparable times can reduce bias without making the intervention itself artificial.
Similarly, randomization can strengthen internal validity without requiring participants to be homogeneous or the setting to be laboratory-like.
The question is therefore not whether the study contains “control.” It is what is being controlled and whether that control changes conditions relevant to the effect in the target population.
Pragmatic Research Does Not Mean Methodologically Loose Research
A persistent misconception is that real-world research gains generalizability by accepting weaker internal validity.
That is not what pragmatic design requires.
Pragmatic randomized trials can preserve random allocation and rigorous outcome analysis while recruiting broader populations, using ordinary practitioners, allowing routine implementation flexibility, and selecting outcomes relevant to real-world decisions.
Empirical work associated with the PRECIS-2 framework has questioned the assumption that more pragmatic randomized trials necessarily have greater risk of bias than explanatory trials. A study can be pragmatic in its design choices without abandoning safeguards against systematic error.
Watch Out
Do not confuse realism with methodological weakness. Allowing routine variation in implementation may answer a different and highly relevant question about effectiveness. Internal validity concerns whether that question is answered without important bias, not whether every aspect of the study was held constant.
Efficacy and Effectiveness Ask Different Questions
The distinction between explanatory and pragmatic research is closely related to the distinction between efficacy and effectiveness.
An efficacy-oriented study asks whether an intervention can produce the intended effect under conditions designed to support its successful delivery. An effectiveness-oriented study asks what effect occurs under conditions closer to routine use.
These are not merely stronger and weaker versions of the same question.
An intervention could have substantial efficacy when delivered by highly trained specialists with excellent adherence yet produce a smaller average effect when implemented across ordinary settings. That difference may reflect real variation in implementation rather than bias.
For decision-makers, the routine-practice effect may be the quantity that matters most.
Variation Can Be Scientifically Informative
Researchers sometimes treat participant or setting heterogeneity as methodological noise that should be eliminated whenever possible.
But variation can reveal who benefits, where an intervention works, and under what conditions an effect changes.
Suppose a learning intervention works well for students with limited prior knowledge but provides little additional benefit for highly prepared students. A homogeneous sample could conceal that pattern.
Including meaningful variation and examining effect modification across relevant groups can therefore strengthen understanding of external validity rather than merely making the data messier.
Multisite Research Can Strengthen the External-Validity Argument
One way to investigate whether results persist across contexts is to conduct research across multiple settings.
A multisite study can include institutions differing in participant characteristics, staffing, resources, implementation practices, or geographic context while preserving a strong comparison within sites.
That design does not guarantee universal generalizability. Sites may still be selected nonrepresentatively, and effects may vary in unmeasured ways. But deliberately incorporating contextual variation can provide evidence that a result is not peculiar to one highly controlled environment.
Replication Can Be More Informative Than Trying to Make One Study Universal
Researchers sometimes expect a single study to achieve near-perfect causal control while representing every population and implementation condition of interest.
That may be neither realistic nor necessary.
A stronger evidence base can emerge from complementary studies: an explanatory experiment establishing that an effect can occur, pragmatic research evaluating routine effectiveness, observational studies examining rare or long-term outcomes, and replications across populations and settings.
External validity is often strengthened by accumulating evidence about where findings do and do not replicate rather than by stretching one study's conclusion far beyond its direct evidence.
Improving Generalizability Can Sometimes Threaten Internal Validity
The trade-off can also run in the other direction if attempts to create realism remove protections needed for credible inference.
For example, researchers might abandon an appropriate comparison group because it feels artificial, permit outcome assessment procedures to differ systematically across settings, or allow treatment assignment to follow routine preferences without adequately addressing the resulting confounding.
Those choices may make a study resemble practice more closely while introducing bias.
The solution is not to retreat automatically to the most controlled possible design. It is to preserve the methodological safeguards needed for the question while allowing realistic variation where that variation is part of what the research intends to study.
Generalizability Requires a Target, Not Just “Real-World” Conditions
A study conducted in routine practice is not automatically generalizable to every routine setting.
A pragmatic trial in urban hospitals may still generalize poorly to rural clinics. A classroom study using ordinary teachers may still represent only well-resourced schools. A national online survey may still exclude people without reliable internet access.
Generalizability remains a relationship between the study and a specified target population. Calling a design “real world” does not eliminate the need to define that target.
Sometimes the Apparent Trade-Off Is Really a Change in the Research Question
Consider an intervention that works when delivered exactly according to protocol but is difficult to implement consistently in ordinary practice.
A tightly controlled trial may estimate the effect under high adherence. A pragmatic trial may estimate the effect of offering or implementing the intervention under usual conditions.
If the estimates differ, it does not follow that one study has better internal validity and the other has better external validity. They may be estimating different effects under different intervention strategies.
This distinction is critical. Changing implementation conditions can change the treatment being evaluated and therefore change the estimand itself.
Design Choices Should Follow the Intended Inference
The best balance depends on what researchers and decision-makers need to know.
If the question is whether a biological or behavioral mechanism can produce an effect, substantial control may be appropriate. If the question is whether an intervention should be adopted across ordinary institutions, variation in participants, implementers, and settings may be essential rather than inconvenient.
A valid and defensible research design therefore does not maximize control for its own sake. It uses enough control to protect the intended inference while representing the conditions necessary to answer the actual research question.