03 · What You Need to Know
Instrument Validity and Study Validity Operate at Different Levels
The confusion begins with the phrase “validated instrument.” It sounds as though an instrument has passed a permanent quality inspection and can subsequently be inserted into any research project without further methodological concern.
Contemporary measurement theory is more cautious.
The Standards for Educational and Psychological Testing defines validity in relation to evidence and theory supporting interpretations of scores for proposed uses. In this framework, validity is not simply a permanent characteristic residing inside a questionnaire, test, rubric, or scale.
Educational measurement literature consequently cautions that describing a survey as “previously validated” can be misleading. Validity evidence is gathered for particular scores, interpretations, populations, contexts, and purposes. Validation is better understood as an accumulating body of evidence than as a one-time certification.
What Does a Validated Instrument Actually Give You?
A well-developed existing instrument can provide an important methodological advantage. Previous research may already offer evidence about its content, dimensional structure, reliability, relationships with other variables, scoring procedures, responsiveness, or other relevant measurement properties.
That evidence can make an existing instrument considerably more defensible than creating a new set of questions without a clear construct definition or measurement-development process.
But the precise advantage depends on what has actually been studied.
An instrument described as “validated” may have extensive evidence from multiple populations and settings. Another may have undergone only expert review and internal-consistency analysis in one small sample. The label alone does not tell you how strong or relevant the evidence is.
Watch Out
Do not treat the word “validated” in a previous article as sufficient evidence. Examine what measurement properties were actually evaluated, with which population, in what language and setting, using what methods, and for what intended interpretation or purpose.
Validity Evidence Belongs to an Interpretation and Use
Imagine a questionnaire developed to measure academic self-efficacy among undergraduate nursing students. Several studies provide evidence supporting the interpretation of its scores for that population.
You now use the questionnaire among experienced university professors and call the resulting score “research self-efficacy.”
The fact that the original instrument was well studied does not justify the new interpretation. The population has changed, and more importantly, the construct claim has changed.
This illustrates a central principle of contemporary validity theory: evidence supporting one score interpretation does not automatically support another.
Before adopting an instrument, researchers should therefore understand the validity evidence underlying the intended interpretation rather than simply searching the original article for the word “validated.”
Validated for Whom?
Population differences matter because people may understand and respond to the same item differently.
Age, educational background, profession, language, culture, clinical status, socioeconomic context, and familiarity with the subject matter can all affect how an instrument functions. COSMIN accordingly recommends evaluating measurement instruments in relation to the specific construct and population in which they are intended to be used.
Consider an item asking respondents whether they can “navigate a learning management system independently.” Among university students who use such systems daily, the wording may be immediately meaningful. Among another population with limited exposure to formal digital-learning platforms, the same item may invoke a different frame of reference.
The question is not whether every change of population automatically invalidates an instrument. It is whether the existing evidence remains sufficiently relevant to support the interpretation you intend to make. That question deserves particular attention when validity evidence is being transferred across populations.
Translation Is Not Just Replacing Words
A translated questionnaire may preserve literal wording while changing meaning.
Expressions, examples, response categories, social norms, and assumptions embedded in an item may function differently across linguistic or cultural settings. Research on cross-cultural adaptation repeatedly emphasizes that translation alone may not establish conceptual and measurement equivalence.
For example, an everyday activity used as an indicator in one country may be uncommon in another. Respondents could understand every translated word correctly while the item no longer represents the same experience or level of the construct.
When an instrument is translated or culturally adapted, researchers may therefore need evidence concerning comprehensibility, content relevance, structural properties, measurement invariance, reliability, or other properties appropriate to the intended use. The exact evaluation should depend on the nature of the adaptation and the claims being made.
Even Small Modifications Can Affect Existing Evidence
Researchers often modify published instruments for practical reasons. They shorten a questionnaire, remove “irrelevant” items, change a five-point response scale to seven points, replace examples, alter instructions, combine items from several scales, or change the mode of administration.
Some modifications may have little practical effect. Others can change the construct representation, score distribution, dimensional structure, reliability, or relationship with external variables.
The crucial point is that evidence collected for the original version does not automatically describe the modified version.
If modifications are necessary, document them transparently and consider which measurement properties need to be reevaluated. The greater the change to content, scoring, administration, language, or construct meaning, the weaker the assumption that previous evidence transfers unchanged.
Reliability in a Previous Study Does Not Guarantee Reliability in Yours
Reliability coefficients are not permanent specifications printed on an instrument like the capacity of a laboratory flask.
They are estimates obtained from particular data under particular measurement conditions. Score variability, sample characteristics, administration procedures, raters, timing, and measurement structure can affect reliability estimates.
You should therefore not simply copy a Cronbach's alpha, intraclass correlation coefficient, or other reliability estimate from the original validation paper and describe your own measurements as reliable.
The appropriate reliability evidence depends on how the instrument is used. The broader relationship between reliability and validity also means that even excellent reliability evidence cannot by itself establish that your intended interpretation is valid.
Strong Measurement Cannot Repair Weak Sampling
Suppose you use an exceptionally well-supported measure of student engagement but recruit participants exclusively through a voluntary online invitation that disproportionately attracts highly engaged students.
The instrument may measure engagement well. The sample may nevertheless provide a distorted picture of engagement in the population you intend to describe.
This is a selection problem, not an instrument-validation problem.
Similarly, a validated measure cannot correct a sampling frame that excludes important groups, substantial nonresponse that differs systematically across participants, or eligibility criteria inconsistent with the target population.
Strong Measurement Cannot Eliminate Confounding
Imagine an observational study examining whether frequent generative-AI use improves academic performance. Both AI use and academic performance are measured with strong instruments.
Students who use AI frequently may nevertheless differ from other students in prior achievement, digital competence, motivation, socioeconomic resources, course characteristics, or other factors associated with academic performance.
Accurately measuring the exposure and outcome does not automatically establish that the observed relationship is causal. The study must still address potential confounding that could distort the relationship.
Strong Measurement Cannot Fix an Inappropriate Research Design
An instrument can measure an outcome beautifully while the design remains incapable of answering the research question.
If a researcher wants to determine whether an intervention caused improvement but collects only a single post-intervention measurement from participants who received the intervention, even an excellent outcome instrument cannot create the missing comparison or establish what would have happened without the intervention.
This is why study validity depends on the logic of the entire research design. Measurement quality is necessary for many inferences, but it cannot substitute for appropriate design.
Strong Measurement Cannot Fix Inappropriate Analysis
The same principle applies after data collection.
A validated instrument does not protect against incorrect statistical models, failure to account for clustering or repeated observations, inappropriate handling of missing data, selective reporting, unjustified subgroup analyses, or conclusions based solely on statistical significance.
If an instrument contains several subscales, for example, collapsing them into a single total score may be inappropriate unless the scoring model and evidence support that interpretation. Researchers need to follow defensible scoring procedures and align the analysis with the measurement structure.
Strong Measurement Cannot Justify Overstated Conclusions
A final source of confusion occurs when researchers move from “we measured this construct well” to “therefore our substantive conclusion is valid.”
Measurement validity addresses what scores can reasonably be interpreted to represent. It does not determine whether an observed association is causal, whether results generalize to other populations, whether an intervention is practically important, or whether an observed difference arose without relevant bias.
A study may measure every variable extremely well and still support only a limited conclusion.
Keeping these levels separate is essential when distinguishing measurement validity from internal and external validity of the study's broader inferences.