03 · What You Need to Know
Reliability and Validity Ask Different Questions About Measurement
Researchers often encounter reliability and validity together because both concern the quality of measurement. Yet treating them as interchangeable obscures an important distinction.
Reliability is primarily concerned with consistency and measurement error. Validity is concerned with the justification for interpreting measurements in a particular way and using them for a particular purpose.
| Question |
Reliability |
Validity |
| Central concern |
Are measurements sufficiently consistent? |
Does the evidence support the intended interpretation or use? |
| Typical problem |
Scores vary because of measurement error or inconsistent administration, items, occasions, or raters. |
The measure may not adequately represent the intended construct or may support a different interpretation from the one being claimed. |
| Typical evidence |
Internal consistency, test-retest reliability, inter-rater reliability, intra-rater reliability, and estimates of measurement error, depending on the instrument and purpose. |
Evidence concerning content, internal structure, relationships with other variables, consequences or other relevant sources, depending on the measurement framework and intended use. |
| What good evidence allows you to say |
The measurement procedure produces sufficiently consistent scores under specified conditions. |
The intended interpretation of the scores is supported by relevant evidence. |
| What it does not establish by itself |
That the intended construct is actually being measured appropriately. |
That every use, population, setting, or interpretation of the measure is justified. |
What Does Reliability Actually Mean?
Reliability concerns the extent to which observed measurements are sufficiently consistent rather than being dominated by measurement error. The exact form of consistency that matters depends on how the measure is used.
If the same relatively stable construct is measured on two occasions, you may be interested in test-retest reliability. If different raters evaluate the same performance, inter-rater reliability may matter. If a multi-item scale is intended to measure a common construct, researchers may examine internal consistency, although internal consistency is not interchangeable with every other form of reliability.
The COSMIN framework, developed for evaluating health measurement instruments, makes this contextual character explicit. Its extended definition of reliability concerns whether scores for people who have not changed remain similar across relevant measurement conditions, such as different occasions, raters, or sets of items.
Reliability is therefore not simply “getting the same answer twice.” Researchers need to specify what source of variation matters and under what conditions consistency is expected.
Internal Consistency Is Only One Form of Reliability
Cronbach's alpha is so frequently reported that it can appear to be synonymous with reliability. It is not.
Internal consistency concerns relationships among items intended to contribute to a scale. It does not tell you whether scores remain stable over time, whether two observers agree, or whether a measure adequately represents the intended construct.
Its interpretation also depends on the structure of the scale. COSMIN guidance, for example, evaluates internal consistency in conjunction with evidence about unidimensionality rather than treating an alpha coefficient in isolation. A high coefficient does not by itself demonstrate that a collection of items forms the measurement structure a researcher assumes.
Watch Out
Do not write “the questionnaire was valid because Cronbach's alpha was high.” Cronbach's alpha provides information about a particular aspect of internal consistency under particular assumptions. It is not a general test of validity, nor is it a complete evaluation of reliability.
Validity Is Not Simply Whether an Instrument “Works”
Validity is sometimes described informally as whether an instrument measures what it is supposed to measure. That is a useful starting intuition, but contemporary measurement theory is more precise.
The Standards for Educational and Psychological Testing, developed by the American Educational Research Association, American Psychological Association, and National Council on Measurement in Education, frame validity around the evidence and theory supporting interpretations of test scores for proposed uses. The emphasis is therefore not merely on attaching the label “valid” to an instrument. It is on the interpretation researchers intend to make from its scores.
Suppose a questionnaire produces a numerical score called “digital learning readiness.” The important question is not simply whether the questionnaire has previously been declared valid. Researchers need to consider what the score represents, what evidence supports that interpretation, for whom the interpretation is appropriate, and what uses of the score are justified.
This is why the different forms and sources of validity evidence matter. Content-related evidence, relationships with other measures or criteria, and evidence concerning the internal structure of a measure address different aspects of the validity argument.
Reliability Is Necessary for Many Interpretations but Is Not Sufficient for Validity
Consider a bathroom scale that consistently reports your weight several kilograms too high. If repeated measurements under unchanged conditions are almost identical, the scale may demonstrate considerable consistency. Yet its readings are systematically inaccurate.
The analogy is imperfect for complex psychological, educational, and social constructs, where there may be no simple physical “true value,” but it illustrates the central logic: consistency alone does not establish correctness of interpretation.
The same problem occurs in research instruments. A questionnaire may consistently capture respondents' confidence with technology while the researcher interprets its score as actual technological competence. The responses could be highly consistent while the intended interpretation remains poorly supported.
This is the central reason a measure can be reliable without being valid for the intended interpretation.
Low Reliability Can Undermine Validity Evidence
The relationship also works in the other direction. If measurements fluctuate substantially because of irrelevant sources of error, it becomes difficult to support stable interpretations from those scores.
Imagine a performance assessment in which a student's score depends heavily on which assessor happens to mark the work. If the construct being interpreted is the student's performance rather than the assessor's idiosyncratic judgment, substantial rater-related inconsistency weakens confidence in the score.
This does not mean there is a universal reliability coefficient that automatically determines whether a measure is valid. The adequacy of reliability depends on the measurement purpose, consequences of decisions, population, design, and form of reliability being evaluated. Numerical thresholds can be useful within specific methodological frameworks, but they should not replace substantive judgment about the measurement process.
Reliability Depends on the Source of Measurement Error You Care About
Different reliability studies are designed to investigate different sources of variation. The appropriate approach therefore depends on what could make the scores inconsistent.
| Form |
Question It Addresses |
Example |
| Test-retest reliability |
Are scores sufficiently consistent across occasions when the construct is expected to remain stable? |
Administering the same anxiety measure twice over an appropriate interval to participants whose anxiety is expected to be stable. |
| Inter-rater reliability |
Are scores sufficiently consistent across different raters? |
Two trained observers independently scoring the same classroom behavior. |
| Intra-rater reliability |
Does the same rater provide sufficiently consistent ratings across repeated assessments? |
An assessor rescoring the same performances under suitable conditions. |
| Internal consistency |
Are items intended to form a scale sufficiently interrelated, given the assumed measurement structure? |
Examining the items within a unidimensional attitude scale. |
The design of the reliability study matters as much as the statistic. For test-retest assessment, for instance, the interval should be appropriate: long enough to reduce simple recall effects where relevant, yet not so long that genuine change in the construct becomes a major source of score differences. COSMIN standards similarly emphasize appropriate intervals, comparable measurement conditions, and stability of the measured construct when evaluating test-retest reliability.
Validity Evidence Must Match the Interpretation You Intend to Make
Researchers sometimes ask, “What validity test should I run?” as though validity were established by a single statistical procedure.
Usually, the better question is: “What evidence would make my interpretation of these scores credible?”
If you are developing a questionnaire to assess students' experiences of feedback, you might need evidence that the items adequately represent the relevant aspects of that construct. You may need evidence about the scale's internal structure. You might examine theoretically expected relationships with other variables. Depending on the purpose, comparison with an appropriate criterion could also be relevant.
Which evidence matters depends on the construct and proposed use. COSMIN, for example, distinguishes content validity, structural validity, hypotheses testing for construct validity, cross-cultural validity or measurement invariance, and criterion validity within its framework for measurement properties. These are not interchangeable boxes that every instrument automatically needs to tick in exactly the same way.
A Published “Validated Instrument” Still Requires Judgment
Finding a published scale with evidence of reliability and validity is often preferable to inventing a new measure without justification. It does not, however, remove the researcher's responsibility to evaluate fit.
Evidence developed with one population, language, cultural context, administration mode, or purpose may not automatically justify every new use. Even seemingly modest adaptations, such as translating items, changing response options, removing items, or shifting from one population to another, may affect how scores should be interpreted.
This is why researchers should distinguish evidence that an instrument has performed well somewhere from evidence that its use is appropriate here. The question becomes especially important when considering whether validity evidence transfers across populations and contexts.
Measurement Validity Is Not the Same as Study Validity
An excellent measurement instrument does not make an entire research design valid. A study could use a well-supported measure yet suffer from selection bias, confounding, inappropriate timing, weak comparison conditions, substantial attrition, or an analysis that does not correspond to the research question.
Conversely, a carefully designed experiment can still produce questionable conclusions if its key outcome is measured poorly.
Measurement quality therefore forms one part of a larger chain of inference. A valid and defensible research design requires alignment among the research question, design, sampling, measurement, analysis, and conclusions rather than excellence in only one component.
06 · What This Means for You
Evaluate Reliability and Validity Separately, Then Consider Them Together
When choosing, adapting, developing, or evaluating a research measure, avoid beginning with the question, “Is this instrument reliable and valid?” That phrasing encourages a yes-or-no answer to what is usually a more contextual evaluation.
Instead, identify what you intend to measure, what interpretation you intend to make from the scores, and what sources of error or alternative interpretations could threaten that claim.
A simple decision framework
If you need to know whether scores are consistent across repeated occasions
Examine an appropriate form of test-retest reliability under conditions in which the construct should remain sufficiently stable.
If measurements depend on judgments by different raters
Evaluate inter-rater reliability or agreement using methods appropriate to the type of data and measurement design.
If multiple items are intended to represent a common underlying construct
Examine the measurement structure and appropriate internal-consistency evidence rather than interpreting an alpha coefficient in isolation.
If you want to claim that scores represent a particular construct
Identify the validity evidence needed to support that specific interpretation rather than relying on reliability coefficients.
If you are adopting an instrument validated in previous research
Check whether its evidence is relevant to your population, language, context, administration, purpose, and intended interpretation.
In a manuscript or thesis, report the evidence that is relevant to your actual measurement problem. A collection of coefficients is not a substitute for a measurement argument. Readers should be able to understand what each piece of evidence tells them and why it matters for the way you use the resulting scores.