03 · What You Need to Know
Measurement Quality Determines What Your Data Actually Represent
Start With the Construct, Not the Instrument
Before deciding that previous research used inadequate measurement, clarify what should have been measured.
A construct is the concept the researcher intends to investigate. Depending on the field, that might be anxiety, academic engagement, digital literacy, pain, physical activity, socioeconomic status, medication adherence, or countless other phenomena that cannot necessarily be observed directly.
The instrument or measurement procedure is how information about that construct is obtained.
COSMIN guidance emphasizes defining the outcome or construct of interest before selecting a measurement instrument, because researchers need that definition to judge whether an instrument's content is relevant and sufficiently comprehensive.
What you want to measure
The construct, outcome, exposure, behavior, or other quantity required by the research question.
How you measure it
The instrument, test, questionnaire, observation, device, record, coding procedure, or other method used to generate the data.
A new instrument is useful only if it provides better evidence about the thing that actually matters.
Validity Is About the Interpretation of Measurements
A measure can generate precise numbers without measuring the intended construct adequately.
In educational and psychological testing, the Standards for Educational and Psychological Testing provide a major framework for evaluating tests and their uses. In health measurement, COSMIN similarly treats validity as a central domain of measurement quality and defines construct validity in terms of whether scores are consistent with hypotheses based on the construct the instrument purports to measure.
This distinction matters when evaluating an existing literature. If researchers repeatedly make claims about a construct that their measurement approach represents poorly, another study using stronger measurement may revisit the same substantive question while producing meaningfully different evidence.
Reliability and Validity Are Related but Not Interchangeable
Researchers sometimes describe an instrument as “valid and reliable” as though the two properties were a single quality. They are not.
Reliability concerns consistency and the degree to which measurements are free from measurement error under specified conditions. Validity concerns whether the resulting scores support the intended interpretation of the construct being measured. COSMIN distinguishes reliability, validity, and responsiveness as separate domains of measurement properties.
A measure can be highly consistent while systematically representing the wrong thing. A bathroom scale that is always several kilograms off illustrates the intuition rather well, although actual research measurement is usually less cooperative than bathroom scales.
Consequently, a new study should not claim superior measurement merely because an instrument has a high reliability coefficient. The relevant measurement properties depend on what the study needs the scores to do.
Measurement Error Can Change the Estimated Relationship
Measurement error is not always harmless statistical noise. It can affect estimated associations, reduce precision, distort classification, and complicate causal interpretation.
Importantly, the common assumption that random measurement error simply weakens associations toward zero is not universally correct. The direction and magnitude of distortion can depend on what is measured, the error structure, the model, and other features of the analysis.
This creates a legitimate reason for another study when the existing evidence relies heavily on measurements known or plausibly expected to contain consequential error and the proposed study can reduce or characterize that error more adequately.
Better Measurement Can Change an Apparently Established Effect
Imagine that numerous studies report a relationship between two constructs, but both constructs are measured using self-report questionnaires administered at the same time. Part of the observed relationship could reflect shared measurement processes rather than the substantive relationship researchers intend to estimate.
A new study that obtains one construct from behavioral records, repeated observations, validated performance measures, or another independently justified source may test whether the relationship persists when the measurement process changes.
The point is not that objective measures are always superior to self-report. Self-report is appropriate when the phenomenon itself concerns perceptions, experiences, intentions, beliefs, or information that participants are uniquely positioned to provide. Measurement quality is question-dependent.
Better measurement means better alignment between the research question, construct, measurement procedure, and intended interpretation.
Content Validity Can Be More Important Than Adding More Items
A longer instrument does not automatically measure a construct better.
COSMIN defines content validity in terms of whether the content of a measurement instrument adequately reflects the construct it is intended to measure. Its guidance highlights relevance, comprehensiveness, and comprehensibility when considering content validity.
An instrument containing 50 poorly targeted items may therefore be less useful than a shorter measure whose content aligns closely with the intended construct and population.
If previous research systematically omitted an important dimension of a construct, another study may be justified when improved measurement captures that missing dimension and the omission matters to the substantive conclusion.
Responsiveness Matters When the Question Is About Change
Some instruments may distinguish people reasonably well at one point in time yet perform poorly when researchers need to detect meaningful change.
COSMIN treats responsiveness as the ability of an instrument to detect change over time in the construct being measured.
This distinction becomes important in intervention and longitudinal research. If a study asks whether an intervention changes a construct, the measurement approach needs to be appropriate for detecting that change. A measure that is poorly suited to longitudinal change can obscure an intervention effect even if it performs adequately for another purpose.
Measurement Properties Are Contextual
An instrument should not be treated as permanently “validated” for every possible use.
The adequacy of a measurement approach depends on the construct, population, language, context, intended interpretation, and purpose of measurement. COSMIN, for example, explicitly includes cross-cultural validity among relevant measurement properties and recommends selecting instruments based on evidence about their quality and intended use.
An instrument that performs well among adults in one linguistic context may require additional evidence before researchers assume equivalent performance among adolescents using a translated or culturally adapted version.
This does not mean every new population automatically requires a new validation study. It means that the evidence supporting the intended measurement interpretation should match the context in which researchers intend to use it.
A New Instrument Is Not Automatically a Better Instrument
Novelty is particularly seductive in measurement. Researchers may prefer a recently developed scale because it sounds contemporary, includes fashionable terminology, or was designed specifically for the emerging topic they are studying.
Yet an established instrument with substantial evidence supporting its measurement properties may be preferable to a new instrument with little evidence beyond the development paper.
COSMIN recommends considering evidence on reliability, validity, responsiveness, and feasibility when selecting an outcome measurement instrument, rather than choosing on novelty alone.
Watch Out
Do not justify another study merely by saying that you will use a “more comprehensive,” “updated,” or “validated” instrument. Specify which measurement property matters, what evidence supports the proposed measure for the intended use, and how the change addresses a limitation in the existing literature.
Better Technology Is Not Automatically Better Measurement
Digital sensors, learning analytics, administrative records, eye tracking, wearable devices, automated text analysis, and other technologies can provide forms of measurement unavailable to traditional instruments.
They can also introduce their own errors.
A digital trace may record clicks with extraordinary precision while remaining a poor measure of cognitive engagement. An automated classification system may produce reproducible labels while embedding systematic classification errors. Administrative records may avoid recall bias but contain missing or operationally defined variables that do not correspond neatly to the research construct.
Precision of data capture should therefore not be confused with validity of measurement.
Better Measurement Can Be More Valuable Than a Larger Sample
Suppose previous studies use a noisy or systematically problematic measure. Repeating the same procedure with 10,000 participants may provide highly stable estimates based on the same measurement limitation.
In that situation, a larger sample alone may not justify repeating the study. A smaller study with substantially better measurement might provide more useful evidence if measurement quality is what currently limits inference.
The reverse can also be true. If the existing measure is already well supported but estimates remain imprecise because studies are small, better measurement may not be the priority.
The methodological improvement should match the evidential problem.
Measurement Improvement Should Change the Inference, Not Merely the Methods Section
The strongest justification connects measurement directly to the conclusion.
Instead of writing, “Previous studies used Scale A, whereas this study uses the newer Scale B,” explain what Scale A could not adequately establish, what evidence supports the relevant properties of Scale B, and why that difference matters to the research question.
This is the same underlying test used to decide whether another study actually adds information. Methodological improvement is scientifically consequential when it changes what researchers can know, estimate, compare, or interpret.