03 · What You Need to Know
Why One Construct Can Produce Several Defensible Measurements
A Construct Is Not Identical to Any One Instrument
A construct is a theoretical concept. An instrument or measurement procedure is one way of obtaining empirical evidence about it.
This distinction is easy to lose when a particular scale becomes so widely used that researchers begin speaking as though the instrument and construct were synonymous. Yet replacing, revising, or supplementing an instrument does not necessarily create a new construct. Conversely, using the same instrument does not guarantee that researchers in different populations or contexts are making exactly the same measurement interpretation.
Keeping the construct separate from its measures, instruments, and indicators makes alternative operationalizations much easier to understand.
Operationalization Creates the Bridge From Construct to Evidence
Once you define a construct, you need to determine what observable evidence can represent it. That process can yield more than one reasonable answer.
Suppose the construct is physical activity. Researchers might obtain evidence through accelerometers, activity diaries, questionnaires, direct observation, or other methods. Each measurement approach can capture aspects of activity differently.
Likewise, academic achievement might be represented by performance on a standardized assessment, course examinations, grades, or other outcomes, depending on what “academic achievement” means in the particular study.
The possibility of several measures follows from the fact that turning a construct into a measurable variable requires theoretical and methodological choices rather than discovering one inevitable numerical form hidden inside the construct.
Different Measures May Capture Different Manifestations of the Same Construct
Many constructs can manifest through several kinds of evidence. Anxiety, for example, may involve self-reported experiences, observable behavior, physiological responses, or performance under particular conditions. Whether these are all indicators of one construct, dimensions of a broader construct, or distinct but related phenomena depends on the theoretical framework.
Researchers should therefore avoid assuming that two measures are interchangeable simply because both are described using the same construct label.
A self-report scale and a behavioral task may correlate imperfectly because one or both contain measurement error. But imperfect correspondence can also occur because they emphasize different aspects of the phenomenon.
Different Operational Definitions Can Answer Slightly Different Questions
Consider social media use. One study asks participants how many hours they believe they spend on social media. Another obtains device-recorded screen time. Both may reasonably be described as measures related to social media use, but they are not identical observations.
The first represents reported use. The second represents activity captured according to the device's recording rules. Differences between them may reflect recall error, system limitations, different definitions of use, or other factors.
The measurement choice therefore shapes the interpretation. Researchers should be alert to situations in which the way something is measured changes the research question actually being answered.
Several Measures Can Be Defensible Without Producing the Same Number
Validity should not be understood as requiring every defensible measure of a construct to yield numerically identical results. Measures may use different units, scales, response formats, indicators, observation periods, or methods.
Even when two instruments produce scores on superficially similar ranges, equal numerical scores need not have equivalent meanings.
What matters is whether evidence supports the interpretation and use of each measurement. The Standards for Educational and Psychological Testing emphasizes that validity concerns evidence and theory supporting interpretations of scores for proposed uses. This perspective makes it possible for different measurement approaches to be defensible while still requiring separate evidence for their intended interpretations.
“Valid in Different Ways” Does Not Mean “Anything Goes”
The existence of alternative operationalizations should not be mistaken for methodological relativism. Researchers cannot simply choose any convenient variable and declare it another valid way of measuring the construct.
A defensible measure needs a reasoned connection to the construct. Depending on the measurement context, relevant evidence may concern content representation, response processes, internal structure, relationships with other variables, consequences of testing or use, reliability, measurement error, or other measurement properties.
Watch Out
Two variables do not become valid measures of the same construct merely because researchers give them the same label. “Engagement,” “well-being,” “achievement,” and similar broad terms can conceal substantially different operationalizations. Examine what was actually measured before comparing findings.
Different Methods Have Different Error Structures
Alternative measures can differ not only in construct coverage but also in the errors they introduce.
Self-report can be influenced by recall, interpretation, or response tendencies. Observation can be affected by sampling of behavior and observer judgments. Administrative records can contain coding or coverage problems. Devices can have calibration, detection, and algorithmic limitations.
Consequently, using different measures can produce different estimates even when each provides relevant evidence about the same construct.
Understanding systematic and random measurement error helps explain why disagreement among measures does not automatically establish that one is valid and the others are not.
Multiple Methods Can Sometimes Strengthen the Measurement Argument
Researchers may deliberately obtain evidence through more than one measurement method. If theoretically related measures using different methods show expected patterns of association, that can contribute evidence about the interpretation of the construct.
The classic multitrait-multimethod framework developed by Campbell and Fiske was designed to examine convergent and discriminant validity by considering multiple traits measured through multiple methods. The underlying insight remains influential: apparent agreement may reflect the construct, but it may also reflect the method, while disagreement can reveal either construct differences or method effects.
Using multiple methods is therefore not simply a matter of “double-checking” one measure against another. The design needs a theoretical reason for what agreement and disagreement would mean.
Convergent Evidence Does Not Require Perfect Correlation
If two measures are intended to represent the same or closely related construct, researchers may expect them to relate. A near-perfect correlation, however, is not necessarily the criterion for validity.
Different methods contain different sources of error and may emphasize different manifestations of the construct. If two supposedly different methods correlate almost perfectly, researchers might even ask whether they are truly providing independent information.
The expected degree of convergence should therefore follow from theory, measurement design, and prior evidence rather than an arbitrary universal correlation threshold.
Population and Context Can Make One Measure More Appropriate Than Another
A measure that works well in one population may not be equally interpretable in another. Language, culture, developmental level, accessibility, technology access, institutional practices, and other contextual factors can affect how a measurement functions.
This means that two measures may both have strong evidence in general, yet one may be much more appropriate for a particular study. Researchers need to ask whether the measure fits the population and context in which it will actually be used.
Practical Constraints Can Legitimately Influence the Choice
The theoretically ideal measurement strategy may require expensive equipment, lengthy assessments, trained observers, proprietary instruments, repeated measurements, or access to records that researchers do not have.
Feasibility matters. A less resource-intensive measure can still be defensible if it adequately supports the intended interpretation. What researchers should avoid is allowing convenience to silently expand the meaning of the data.
If the feasible measure captures a narrower aspect of the construct, the defensible response may be to narrow the claim rather than pretend the measurement is more comprehensive than it is.