03 · What You Need to Know
Measure Appropriateness Is About Fit Between Evidence and Use
“Validated” Is Not a Permanent Property of an Instrument
Researchers often describe an instrument as “validated” as though validation were a one-time certification. Contemporary measurement standards take a more qualified view.
The Standards for Educational and Psychological Testing frames validity in terms of evidence and theory supporting interpretations of scores for proposed uses. This means the relevant question is not simply, “Is this scale valid?” but rather, “What interpretation will I make from these scores, for what purpose and population, and what evidence supports that interpretation?”
An instrument can have substantial supporting evidence for one use without automatically having sufficient evidence for another.
Start With Construct Fit
Before examining reliability coefficients or factor models, determine whether the instrument actually represents the construct as you define it.
Two measures can share a title such as “digital competence,” “engagement,” or “well-being” while operationalizing the concept differently. One may emphasize particular dimensions that are central to your study, while another may omit them.
Compare the instrument's conceptual basis and content with your own conceptual and operational definitions. If the constructs do not align, population similarity cannot rescue the mismatch.
Ask Whether the Content Is Relevant to Your Population
An item can represent the intended construct in one population while being irrelevant, unfamiliar, inaccessible, or interpreted differently in another.
Suppose a technology-use measure contains items about workplace systems, organizational supervisors, and corporate technical support. Using it with secondary-school students would require more than changing “employee” to “student.” Some of the situations themselves may have no equivalent meaning.
Content validity therefore needs attention to the intended population. COSMIN's work on patient-reported outcome measures emphasizes relevance, comprehensiveness, and comprehensibility as central aspects of content validity. Although developed in a health measurement context, those questions are broadly instructive: Are the items relevant? Is important content represented? Can the intended respondents understand the items as required?
Age and Developmental Level Can Affect Measurement
Children, adolescents, university students, working adults, and older adults may differ in vocabulary, cognitive demands, life experiences, and familiarity with the situations described in an instrument.
An item requiring complex retrospective judgment may function differently for younger respondents. A construct may also manifest differently across developmental stages.
Researchers should therefore examine whether the measure has been used and evaluated with a population sufficiently similar to their own rather than assuming that a measure designed for “people” is developmentally neutral.
Language Adaptation Is More Than Translation
If the instrument was developed in another language, a literal translation may preserve words without preserving the construct meaning or response process.
Idioms may not transfer. A term may have several plausible translations. Response categories can carry different connotations. Examples familiar in one cultural setting may be confusing elsewhere.
The International Test Commission provides guidelines for translating and adapting tests that emphasize linguistic, cultural, and measurement considerations. Depending on the measure and intended use, adaptation may involve expert review, multiple translators, target-population input, pretesting, and empirical evaluation.
Watch Out
Do not describe a translated instrument as equivalent to the original merely because the wording was translated accurately. Linguistic accuracy is necessary, but measurement equivalence is an empirical and conceptual question.
Culture Can Affect More Than Language
Even when participants understand the same language, cultural context can affect how constructs, situations, and response options are interpreted.
Concepts such as autonomy, family obligation, authority, individual achievement, well-being, trust, or social support may not have identical meanings or manifestations across populations. Response tendencies can also vary.
This does not mean cross-cultural measurement is impossible. It means researchers should not assume that identical wording guarantees identical measurement.
The Setting Can Change What an Item Means
Context includes more than country or culture. An instrument developed for hospitals may function differently in community settings. A scale designed for face-to-face teaching may contain assumptions that do not apply to fully online education. A workplace measure developed for permanent employees may not fit temporary workers or volunteers.
The technology, institution, policy environment, and historical period can matter too. Measures containing references to particular technologies or practices may become less appropriate as those practices change.
This is especially relevant when choosing whether to use an existing measure, adapt one, or develop something new.
Administration Mode Can Affect the Response Process
An instrument may have been evaluated as a paper questionnaire administered under supervision but used in your study as an unsupervised mobile survey. A test may have been designed for individual administration but delivered to groups. An interviewer-administered measure may be converted to self-administration.
Such changes can affect item presentation, privacy, accessibility, missing responses, opportunities for clarification, timing, and other features of the response process.
Not every mode change will materially alter measurement, but neither should equivalence be assumed without considering the instrument and evidence.
Scoring and Interpretation Must Also Transfer
Suppose an instrument uses a published cutoff score to classify respondents. Even if the questionnaire items function adequately in your population, the cutoff may have been established for a different purpose or group.
Similarly, norms developed in one population cannot automatically be interpreted as norms for another. A score can be calculated correctly while its comparison or classification is inappropriate.
Check not only whether you can administer the instrument, but whether the scoring rules and interpretations you intend to use are supported.
Measurement Invariance Matters When You Want to Compare Groups
If your study compares groups, an additional question arises: does the measure function sufficiently similarly across those groups for the intended comparison?
Measurement invariance concerns whether measurement relationships are comparable across groups, occasions, or other conditions under a specified model. In latent-variable frameworks, researchers may examine increasingly restrictive forms of invariance depending on the comparisons they intend to make.
The precise procedures depend on the measurement model and research question. The important principle is that observed score differences do not automatically establish construct differences if the measure itself functions differently across groups.
Same questionnaire
Every group receives the same items and response options.
Comparable measurement
The measure functions sufficiently similarly for the particular cross-group interpretation or comparison being made.
These are not synonymous. Identical administration is useful, but it does not by itself establish measurement equivalence.
Reliability Should Be Examined in Your Application
A reliability estimate reported in an earlier study is evidence about scores obtained under those conditions. It is not an immutable characteristic carried by the questionnaire into every subsequent sample.
Your study may differ in population heterogeneity, administration, language, item interpretation, or other factors that affect score consistency. Appropriate reliability evidence should therefore be considered for the scores obtained in your application.
At the same time, a high reliability estimate in your sample does not establish that the instrument is appropriate. A measure can be consistent while representing the wrong construct or functioning systematically differently across groups.
How Much New Evidence Do You Need?
There is no universal requirement to revalidate every established instrument from the beginning whenever it is used in a new study.
The amount of additional evidence should be proportionate to the differences between the established use and your intended use, the stakes of the interpretation, the extent of any modification, and the measurement questions relevant to the study.
If the population, language, context, administration, and interpretation closely resemble well-supported previous applications, existing evidence may carry considerable weight. If you translate the instrument, substantially alter items, use it with a very different population, change its intended purpose, or make consequential group comparisons, more evidence may be needed.