03 · What You Need to Know
Why Can External Validation Remain Missing After a Successful Development Study?
Development Performance and External Performance Answer Different Questions
When researchers develop a statistical or machine-learning model, they use data to select predictors, estimate parameters, tune procedures, or otherwise construct the model. Performance evaluated using those same development data can be optimistic because the model has, in some sense, been tailored to those observations.
Internal validation methods such as bootstrapping or cross-validation can help estimate and adjust for optimism within the development process. They are important, but they do not answer the same question as external validation.
PROBAST guidance distinguishes prediction-model development with internal validation from external validation using data from different participants. External data may come from a later time period, another hospital or country, another relevant setting, or otherwise distinct participants.
Internal validation
Evaluates model performance while accounting for overfitting or optimism using the development data, commonly through resampling approaches.
External validation
Evaluates the developed model using data from participants who were not part of model development.
A Train-Test Split Is Not Necessarily the External Validation You Think It Is
Researchers sometimes randomly divide one dataset into a training set and a test set, develop the model in the first portion, and call evaluation in the second portion "external validation."
That terminology is problematic. PROBAST guidance treats random splitting of a single dataset as a form of internal validation and notes that it can be inefficient because it reduces the data available both for development and evaluation.
A held-out test set can still be useful, particularly in machine-learning workflows, but you should describe precisely where the test observations came from and what independence they provide. The word external should not conceal the structure of the data.
External Validation Tests More Than Whether the Model Still Produces Predictions
A model can generate predictions for a new dataset and still perform poorly. External validation therefore evaluates predictive performance, not simply technical portability.
For many prediction models, two central aspects are discrimination and calibration.
| Performance aspect |
Basic question |
| Discrimination |
Can the model distinguish participants with different outcomes or risks? |
| Calibration |
Do predicted probabilities correspond adequately to observed outcomes? |
| Overall predictive performance |
How accurately does the model perform when prediction error is considered more broadly? |
| Clinical or practical usefulness |
Would using the model improve decisions relative to relevant alternatives? |
The exact metrics depend on the type of model and outcome. A good external-validation study should therefore use performance measures appropriate to the model rather than reporting one convenient statistic and declaring the model "validated."
External Validation Is Not a Permanent Stamp of Approval
The word validated can create a misleading impression that a model has passed a one-time test and is now valid everywhere.
External performance is conditional on the populations, settings, periods, measurements, and implementation conditions represented in the validation data. A model that performs well in one external hospital, university, region, or period has gained important supporting evidence, but that result does not establish universal performance.
Validation is better understood as accumulating evidence about how a model behaves under relevant conditions.
The Validation Dataset Should Be Relevant to the Intended Use
Independence alone is not sufficient. Imagine a model intended for first-year university students that is externally tested only among postgraduate students. The validation sample is technically different, but it may not answer the most consequential question about the model's intended application.
The target population should guide validation priorities. Ask where the model is expected to be used, who will receive its predictions, how predictors are measured there, and whether important contextual differences could alter performance.
If the central concern is performance in a setting that has not been represented in existing evidence, external validation and contextual-gap arguments may overlap.
Temporal Validation Can Reveal Performance Drift
A model may be tested on participants observed later than those used for development. This temporal validation can be important when populations, practices, technologies, policies, prevalence, or relationships among predictors and outcomes change over time.
For example, a student-risk model developed before a major change in course-delivery practices may not retain the same calibration afterward. Even if predictor definitions remain unchanged, the relationships encoded in the original model can shift.
This means an externally validated model may still require later evaluation. Evidence about model performance has a temporal context.
Geographic Validation Is Not Simply About Crossing a Border
External validation may involve another institution, region, or country. The value comes from testing the model under relevant differences in case mix, measurement, practice, or context, not from geography itself.
A model developed at one university and validated at another nearby institution may face meaningful differences. Conversely, two institutions in different countries could share many characteristics relevant to the model.
The same principle applies to using a different country as a research gap: identify the substantive difference rather than relying on the location label.
Poor External Performance Does Not Necessarily Mean the Model Must Be Abandoned
External validation can reveal that a model discriminates reasonably well but is poorly calibrated, or that performance declines in particular populations. Depending on the model and intended use, updating or recalibration may improve performance.
PROBAST guidance recognizes that external validation may be followed by model adjustment, updating, or extension. Validation therefore serves not only as a pass-or-fail exercise but as evidence for deciding whether a model can be used as developed, requires modification, or should not be used in the target context.
Watch Out
Do not quietly modify a model using the external validation data and then report the modified model's performance on those same data as though it were an untouched external test. Once data contribute to model updating, they are no longer providing the same independent evaluation of the updated model.
External Validation Is Closely Related to Replication, but the Questions Differ
Both involve independent evidence. A replication generally asks whether a finding can be obtained again using new data. External validation asks whether a developed model or other research product performs adequately when applied to data independent of its development.
Suppose a study reports that study habits predict academic performance. Repeating the analysis in another sample might be a replication. If researchers use the original coefficients to create an individual prediction model and then evaluate that unchanged model in another sample, the study is performing external validation.
If the concern is that an empirical finding has received too little independent testing, the stronger framing may be lack of replication. If the concern is whether a developed model performs adequately beyond its development data, external validation is more precise.
External Validation Can Matter Even When Internal Validation Looks Excellent
Bootstrapping, cross-validation, and other internal procedures can provide valuable estimates of model performance and optimism. They should not be dismissed merely because external data are unavailable.
But a model can still encounter differences in predictor distributions, outcome prevalence, measurement procedures, case mix, or other conditions when moved to another population. External validation exposes the model to variation not necessarily represented by internal resampling.
This is why model-development research can be methodologically sophisticated yet still leave an external-validation gap.
External Validation and External Validity Are Related but Not Identical
The terms can be confusing. External validity is a broad concept concerning the applicability or generalizability of study findings beyond the conditions directly studied. External validation usually refers more specifically to testing the performance of a model, algorithm, instrument, or similar product using independent data.
A study can therefore raise questions about external validity without involving a prediction model at all. Likewise, an external-validation study may focus on highly specific performance characteristics of a model in a defined target population.
How Do You Establish That External Validation Is Actually Missing?
Start with the original development paper and trace subsequent studies citing or evaluating the model. Search the model's name, predictors, target outcome, validation terminology, and relevant populations or settings.
Systematic reviews of prediction models can be especially useful because they often distinguish development studies from validation studies. Determine whether purported validations use genuinely independent participants, whether the original model was applied without redevelopment, and whether appropriate performance measures were reported.
Then describe the gap precisely. "The model has been internally validated but has not been independently evaluated in the population for which it is now being considered" is more informative than saying simply that few validation studies exist.