03 · What You Need to Know
Reliability Cannot Tell You What a Measure Actually Represents
The possibility of being reliable but not valid becomes easier to understand once reliability and validity are treated as different questions rather than two interchangeable indicators of measurement quality.
Reliability concerns the extent to which measurement is free from relevant measurement error. Depending on the measurement procedure, researchers may examine consistency across occasions, raters, or items. COSMIN, for example, distinguishes internal consistency, test-retest reliability, inter-rater reliability, and intra-rater reliability as addressing different measurement conditions.
Validity asks a different question. The Standards for Educational and Psychological Testing frames validity in terms of the evidence and theory supporting interpretations of test scores for their proposed uses. The focus is therefore not simply whether an instrument produces stable numbers, but whether those numbers can reasonably be interpreted in the way the researcher proposes.
Reliable measurement
The measurement demonstrates sufficient consistency with respect to the sources of measurement error relevant to its use.
Valid interpretation
Available evidence supports interpreting and using the resulting scores in the way the researcher intends.
These properties are related, but one does not substitute for the other. The broader distinction between reliability and validity explains why reporting a reliability coefficient cannot answer the validity question.
Consistency Does Not Establish Correctness
The familiar analogy is a measurement device that produces the same incorrect result repeatedly. Imagine a scale that consistently reports every person's weight several kilograms too high. Its readings could be highly repeatable while remaining systematically inaccurate.
For psychological, educational, and social constructs, the problem is often subtler because there may be no directly observable “true value” against which the score can simply be checked.
Suppose researchers develop a questionnaire intended to measure students' digital competence. Most items ask whether students feel confident using computers, whether they enjoy experimenting with software, and whether they consider themselves technologically capable.
The items might be highly interrelated. Scores might also remain stable across repeated administrations. Yet if the intended construct includes demonstrated ability to evaluate information, solve technical problems, manage digital security, and use technology effectively, the questionnaire may primarily capture self-perceived technological confidence rather than competence.
The reliability evidence has not become false. It simply answers a different question.
High Internal Consistency Does Not Prove Validity
This issue appears frequently when researchers report Cronbach's alpha.
A high alpha may indicate strong interrelatedness among items under relevant assumptions. It does not establish that the items represent the intended construct. Ten questions asking nearly the same thing can be highly internally consistent while providing very narrow coverage of a multidimensional concept.
Imagine a scale intended to measure “research competence” that contains ten variations of the statement “I am confident conducting statistical analysis.” The items could produce an impressive internal-consistency coefficient. Yet research competence might also encompass problem formulation, literature synthesis, research design, data collection, qualitative reasoning, ethics, interpretation, and scholarly communication.
Consistency among narrowly focused items cannot compensate for inadequate representation of the intended construct.
Watch Out
Statements such as “Cronbach's alpha was.91; therefore, the questionnaire was valid” confuse reliability evidence with validity evidence. An internal-consistency coefficient does not establish content coverage, dimensional structure, expected relationships with other variables, or the appropriateness of the intended score interpretation.
Systematic Error Can Be Highly Consistent
Reliability is often associated with random measurement error. Validity can also be undermined by systematic influences that consistently push measurement away from the intended construct.
Consider a reading-comprehension assessment containing unnecessarily complex technical vocabulary. If the intended construct is general reading comprehension, performance may partly reflect prior knowledge of that technical vocabulary. The assessment could generate reproducible scores while systematically incorporating an unintended ability into the measurement.
Likewise, a questionnaire intended to measure teaching effectiveness may consistently capture how much students like their instructors. If likeability influences responses strongly, stable scores do not establish that the instrument supports the broader interpretation “teaching effectiveness.”
A consistent source of construct-irrelevant influence does not disappear merely because it is consistent.
A Measure Can Reliably Measure the Wrong Construct
This is perhaps the clearest way to understand the issue.
An instrument may be an excellent measure of construct A while being interpreted as construct B.
A self-report questionnaire may reliably assess perceived competence but be used as though it measured demonstrated competence. A knowledge test may consistently assess factual recall but be described as a measure of critical thinking. A student-satisfaction scale may produce stable scores but be interpreted as evidence of learning effectiveness.
In each case, the reliability evidence can be genuine. The problem lies in the inference attached to the score.
This is why the different sources and traditional types of validity evidence matter. Researchers need evidence concerning the interpretation they intend to make, not merely evidence that the resulting numbers are reproducible.
Reliability Is Still Important for Validity
Saying that reliability does not guarantee validity does not make reliability optional.
If measurements fluctuate substantially because of irrelevant error, it becomes difficult to make stable interpretations from them. Imagine two trained raters evaluating the same student performance but producing radically different scores. If the intended interpretation concerns student performance rather than rater preference, that inconsistency creates an obvious problem.
The relationship is therefore asymmetric in an important sense: good reliability evidence does not establish validity, but substantial unreliability can weaken the interpretations that validity evidence is meant to support.
How much reliability is sufficient depends on the measurement purpose, the source of error being investigated, the population, and the consequences attached to the scores. There is no single coefficient that converts a measure from unreliable to reliable for every use.
The Intended Use Determines What Validity Evidence You Need
A measure may support one interpretation but not another.
Suppose a short mathematics test was designed to identify which topics students should review before an examination. Evidence supporting that low-stakes formative use does not automatically justify using the same scores to make high-stakes admissions decisions.
The instrument has not physically changed. The claim attached to its scores has.
Modern measurement theory therefore emphasizes interpretations and uses rather than treating validity as a permanent property embedded inside an instrument. This is also why describing an instrument simply as “valid and reliable” can conceal important methodological details.
Reliability and Validity Evidence Are Population- and Context-Sensitive
Measurement evidence obtained in one population or setting may not behave identically elsewhere.
An instrument may produce consistent scores among adults but behave differently among adolescents. Items that work coherently in one language may not retain the same meaning after translation. A scale developed for one professional group may omit dimensions that matter in another.
COSMIN explicitly recommends examining measurement instruments for the specific construct and population of interest, and contemporary survey-methods literature similarly cautions that validity evidence is sensitive to target population, local context, and intended score use.
Accordingly, researchers should not assume that a published reliability coefficient or previous validity study settles the measurement question for every subsequent application. Whether validity evidence transfers across populations needs separate consideration.
Neither Reliability nor Validity of a Measure Validates the Whole Study
Measurement is only one component of research design.
A study could use a measure with excellent reliability and strong validity evidence but still suffer from severe selection bias, uncontrolled confounding, differential attrition, inappropriate analysis, or conclusions that exceed what the design supports.
Conversely, an otherwise strong research design can be compromised if its central variables are measured poorly.
A defensible research design therefore requires coherence across the entire chain from research question to sampling, measurement, procedures, analysis, and interpretation.
06 · What This Means for You
Do Not Stop Once Your Reliability Coefficient Looks Good
If your instrument produces encouraging reliability evidence, that answers an important measurement question. It does not finish the measurement argument.
Your next task is to ask what the scores actually represent and what evidence supports that interpretation. The answer should come from the construct definition, instrument content, measurement structure, relationships with other variables, response processes, relevant criteria, and other evidence appropriate to the intended use.
A simple decision framework
If your measure has strong internal consistency
Ask whether the items adequately represent the intended construct and whether the assumed dimensional structure is supported.
If test-retest scores are highly consistent
Treat this as evidence about temporal consistency under the studied conditions, not proof that the intended construct is being measured correctly.
If raters show strong agreement
Conclude that rater-related consistency is strong under those conditions, then separately evaluate whether the rating criteria represent the intended construct.
If reliability is strong but validity evidence contradicts your intended interpretation
Reconsider the interpretation, construct definition, instrument, or measurement model rather than allowing reliability to override the contradictory evidence.
Reliability should therefore be treated as evidence with a specific meaning. Report what kind of consistency was evaluated, why that form matters for your study, and what the result does and does not allow you to conclude.