Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can a Measure Be Reliable but Not Valid?

A measure can produce highly consistent results and still fail to support the interpretation a researcher wants to make from them. Reliability is important, but consistency alone cannot establish validity.

130
Reliable but Not Valid Guide 130 of 217
01 · The Question

How Can a Measure Be Consistent and Still Be Wrong?

A questionnaire produces almost identical scores when administered twice. A set of items has excellent internal consistency. Two observers independently rate the same performances and agree almost perfectly.

Those results sound reassuring. But do they establish that the measure is actually capturing what you intend to measure?

No. A measurement procedure can be impressively consistent while supporting the wrong interpretation. This is one of the most important distinctions in measurement: reliability can strengthen confidence in the consistency of scores, but it cannot tell you by itself whether those scores represent the intended construct.

02 · The Short Answer

Yes, Reliability Can Exist Without Validity

In Brief

Yes. A measure can be reliable but not valid for the interpretation you intend to make because consistency does not establish that the measure captures the intended construct.

Reliability concerns measurement consistency and error under specified conditions. Validity concerns whether evidence and theory support a particular interpretation and use of the resulting scores. A measure can therefore produce consistently similar results while systematically measuring something different from what the researcher claims.

03 · What You Need to Know

Reliability Cannot Tell You What a Measure Actually Represents

The possibility of being reliable but not valid becomes easier to understand once reliability and validity are treated as different questions rather than two interchangeable indicators of measurement quality.

Reliability concerns the extent to which measurement is free from relevant measurement error. Depending on the measurement procedure, researchers may examine consistency across occasions, raters, or items. COSMIN, for example, distinguishes internal consistency, test-retest reliability, inter-rater reliability, and intra-rater reliability as addressing different measurement conditions.

Validity asks a different question. The Standards for Educational and Psychological Testing frames validity in terms of the evidence and theory supporting interpretations of test scores for their proposed uses. The focus is therefore not simply whether an instrument produces stable numbers, but whether those numbers can reasonably be interpreted in the way the researcher proposes.

Reliable measurement The measurement demonstrates sufficient consistency with respect to the sources of measurement error relevant to its use.
Valid interpretation Available evidence supports interpreting and using the resulting scores in the way the researcher intends.

These properties are related, but one does not substitute for the other. The broader distinction between reliability and validity explains why reporting a reliability coefficient cannot answer the validity question.

Consistency Does Not Establish Correctness

The familiar analogy is a measurement device that produces the same incorrect result repeatedly. Imagine a scale that consistently reports every person's weight several kilograms too high. Its readings could be highly repeatable while remaining systematically inaccurate.

For psychological, educational, and social constructs, the problem is often subtler because there may be no directly observable “true value” against which the score can simply be checked.

Suppose researchers develop a questionnaire intended to measure students' digital competence. Most items ask whether students feel confident using computers, whether they enjoy experimenting with software, and whether they consider themselves technologically capable.

The items might be highly interrelated. Scores might also remain stable across repeated administrations. Yet if the intended construct includes demonstrated ability to evaluate information, solve technical problems, manage digital security, and use technology effectively, the questionnaire may primarily capture self-perceived technological confidence rather than competence.

The reliability evidence has not become false. It simply answers a different question.

High Internal Consistency Does Not Prove Validity

This issue appears frequently when researchers report Cronbach's alpha.

A high alpha may indicate strong interrelatedness among items under relevant assumptions. It does not establish that the items represent the intended construct. Ten questions asking nearly the same thing can be highly internally consistent while providing very narrow coverage of a multidimensional concept.

Imagine a scale intended to measure “research competence” that contains ten variations of the statement “I am confident conducting statistical analysis.” The items could produce an impressive internal-consistency coefficient. Yet research competence might also encompass problem formulation, literature synthesis, research design, data collection, qualitative reasoning, ethics, interpretation, and scholarly communication.

Consistency among narrowly focused items cannot compensate for inadequate representation of the intended construct.

Watch Out

Statements such as “Cronbach's alpha was.91; therefore, the questionnaire was valid” confuse reliability evidence with validity evidence. An internal-consistency coefficient does not establish content coverage, dimensional structure, expected relationships with other variables, or the appropriateness of the intended score interpretation.

Systematic Error Can Be Highly Consistent

Reliability is often associated with random measurement error. Validity can also be undermined by systematic influences that consistently push measurement away from the intended construct.

Consider a reading-comprehension assessment containing unnecessarily complex technical vocabulary. If the intended construct is general reading comprehension, performance may partly reflect prior knowledge of that technical vocabulary. The assessment could generate reproducible scores while systematically incorporating an unintended ability into the measurement.

Likewise, a questionnaire intended to measure teaching effectiveness may consistently capture how much students like their instructors. If likeability influences responses strongly, stable scores do not establish that the instrument supports the broader interpretation “teaching effectiveness.”

A consistent source of construct-irrelevant influence does not disappear merely because it is consistent.

A Measure Can Reliably Measure the Wrong Construct

This is perhaps the clearest way to understand the issue.

An instrument may be an excellent measure of construct A while being interpreted as construct B.

A self-report questionnaire may reliably assess perceived competence but be used as though it measured demonstrated competence. A knowledge test may consistently assess factual recall but be described as a measure of critical thinking. A student-satisfaction scale may produce stable scores but be interpreted as evidence of learning effectiveness.

In each case, the reliability evidence can be genuine. The problem lies in the inference attached to the score.

This is why the different sources and traditional types of validity evidence matter. Researchers need evidence concerning the interpretation they intend to make, not merely evidence that the resulting numbers are reproducible.

Reliability Is Still Important for Validity

Saying that reliability does not guarantee validity does not make reliability optional.

If measurements fluctuate substantially because of irrelevant error, it becomes difficult to make stable interpretations from them. Imagine two trained raters evaluating the same student performance but producing radically different scores. If the intended interpretation concerns student performance rather than rater preference, that inconsistency creates an obvious problem.

The relationship is therefore asymmetric in an important sense: good reliability evidence does not establish validity, but substantial unreliability can weaken the interpretations that validity evidence is meant to support.

How much reliability is sufficient depends on the measurement purpose, the source of error being investigated, the population, and the consequences attached to the scores. There is no single coefficient that converts a measure from unreliable to reliable for every use.

The Intended Use Determines What Validity Evidence You Need

A measure may support one interpretation but not another.

Suppose a short mathematics test was designed to identify which topics students should review before an examination. Evidence supporting that low-stakes formative use does not automatically justify using the same scores to make high-stakes admissions decisions.

The instrument has not physically changed. The claim attached to its scores has.

Modern measurement theory therefore emphasizes interpretations and uses rather than treating validity as a permanent property embedded inside an instrument. This is also why describing an instrument simply as “valid and reliable” can conceal important methodological details.

Reliability and Validity Evidence Are Population- and Context-Sensitive

Measurement evidence obtained in one population or setting may not behave identically elsewhere.

An instrument may produce consistent scores among adults but behave differently among adolescents. Items that work coherently in one language may not retain the same meaning after translation. A scale developed for one professional group may omit dimensions that matter in another.

COSMIN explicitly recommends examining measurement instruments for the specific construct and population of interest, and contemporary survey-methods literature similarly cautions that validity evidence is sensitive to target population, local context, and intended score use.

Accordingly, researchers should not assume that a published reliability coefficient or previous validity study settles the measurement question for every subsequent application. Whether validity evidence transfers across populations needs separate consideration.

Neither Reliability nor Validity of a Measure Validates the Whole Study

Measurement is only one component of research design.

A study could use a measure with excellent reliability and strong validity evidence but still suffer from severe selection bias, uncontrolled confounding, differential attrition, inappropriate analysis, or conclusions that exceed what the design supports.

Conversely, an otherwise strong research design can be compromised if its central variables are measured poorly.

A defensible research design therefore requires coherence across the entire chain from research question to sampling, measurement, procedures, analysis, and interpretation.

04 · A Practical Example

A Highly Reliable Questionnaire That Measures the Wrong Thing

Hypothetical Example

Measuring actual research competence with self-confidence items

A researcher wants to measure undergraduate students' research competence. The questionnaire asks students to rate statements such as “I feel confident conducting research,” “I believe I am good at research,” and “I am comfortable completing research activities.”

Reliability result The items are highly interrelated, and repeated measurement among students whose self-perceptions are expected to remain stable produces similar scores.
Tempting conclusion The researcher concludes that the questionnaire is a reliable measure of actual research competence.
Validity problem The items directly assess perceived confidence. Students can be confident without possessing strong methodological knowledge or practical research skills, while competent students may underestimate their ability.
More defensible interpretation The scores may provide a reliable measure of perceived research confidence if other evidence supports that interpretation. They cannot be treated automatically as demonstrated research competence.
What the researcher should do Either revise the construct claim to match what the questionnaire measures or gather evidence using measures that more directly represent the intended components of research competence.

Nothing in this example requires the reliability estimate to be poor. The instrument can consistently measure perceived confidence. The validity problem emerges only when researchers interpret that score as something broader.

05 · What Researchers Often Get Wrong

Common Mistakes When Interpreting Reliable Measurements

Misconception

Does High Reliability Mean High Validity?

No. Strong reliability evidence indicates consistency with respect to particular sources of measurement error. It does not establish that the intended construct is represented adequately or that the proposed interpretation of the scores is justified.

Misconception

If the Same Result Appears Repeatedly, Must It Be Accurate?

No. Systematic error can produce highly repeatable results. Repetition addresses consistency, not necessarily correctness of the interpretation attached to the measurement.

Misconception

Does a Cronbach's Alpha Above a Threshold Prove the Scale Works?

No universal alpha threshold establishes that a scale “works.” Alpha addresses internal consistency under particular assumptions and should be interpreted alongside the scale's dimensional structure, purpose, population, and other relevant measurement evidence.

Misconception

If a Measure Is Not Valid for One Purpose, Is It Useless?

Not necessarily. An instrument may fail to support one interpretation while supporting another. A questionnaire intended to measure competence might turn out to provide useful evidence about self-confidence. The appropriate response may be to change the interpretation rather than discard the instrument automatically.

Misconception

Does Reliability Have to Be Established Before Validity Can Even Be Studied?

Reliability and validity evidence are conceptually related, but validation is not necessarily a rigid two-stage sequence in which researchers first “finish” reliability and then begin validity. Instrument development and evaluation typically accumulate several forms of evidence iteratively. Relevant measurement error should nevertheless be understood because substantial inconsistency can constrain the interpretations that scores can support.

06 · What This Means for You

Do Not Stop Once Your Reliability Coefficient Looks Good

If your instrument produces encouraging reliability evidence, that answers an important measurement question. It does not finish the measurement argument.

Your next task is to ask what the scores actually represent and what evidence supports that interpretation. The answer should come from the construct definition, instrument content, measurement structure, relationships with other variables, response processes, relevant criteria, and other evidence appropriate to the intended use.

A simple decision framework

If your measure has strong internal consistency
Ask whether the items adequately represent the intended construct and whether the assumed dimensional structure is supported.
If test-retest scores are highly consistent
Treat this as evidence about temporal consistency under the studied conditions, not proof that the intended construct is being measured correctly.
If raters show strong agreement
Conclude that rater-related consistency is strong under those conditions, then separately evaluate whether the rating criteria represent the intended construct.
If reliability is strong but validity evidence contradicts your intended interpretation
Reconsider the interpretation, construct definition, instrument, or measurement model rather than allowing reliability to override the contradictory evidence.

Reliability should therefore be treated as evidence with a specific meaning. Report what kind of consistency was evaluated, why that form matters for your study, and what the result does and does not allow you to conclude.

07 · A Quick Checklist

When a Measure Looks Reliable, Check What That Really Means

Before interpreting strong reliability evidence, check:
Identify which form of reliability you actually evaluated: internal consistency, test-retest, inter-rater, intra-rater, or another relevant form.
Explain which source of measurement error that evidence addresses.
Define the construct independently of the questions or indicators used to measure it.
Check whether the content of the measure adequately represents that construct rather than merely producing similar responses.
Look for systematic influences that could consistently make scores represent something other than the intended construct.
Gather validity evidence appropriate to the interpretation and use you intend to make from the scores.
Avoid describing a reliability coefficient as proof that an instrument is valid.
Verify that existing reliability and validity evidence is relevant to your population, setting, language, and intended use.
08 · Frequently Asked Questions

Frequently Asked Questions About Reliable but Invalid Measures

Can something be 100% reliable but not valid?

Conceptually, perfect consistency would still not prove that the intended interpretation is correct. A measurement procedure could reproduce the same systematically distorted result. In real research, however, claims of perfect reliability should themselves be examined carefully because reliability estimates depend on the measurement design, population, and statistical model.

Why does reliability not guarantee validity?

Because reliability concerns consistency or measurement error, while validity concerns the interpretation and use of scores. Consistently measuring the wrong construct does not make that interpretation valid.

Can a high Cronbach's alpha coexist with poor validity?

Yes. Items can be highly interrelated while representing only a narrow part of the intended construct, reflecting a different construct, or containing systematic sources of bias. Alpha alone therefore cannot establish validity.

Can systematic error reduce validity without reducing reliability?

Yes. A systematic influence can shift measurements consistently in a way that undermines the intended interpretation while leaving repeatability relatively high. This is one reason consistency and validity need separate evaluation.

Is reliability necessary for validity?

Sufficient reliability is generally important because substantial measurement error limits the interpretations that scores can support. The amount and type of reliability required depend on the measurement procedure and intended use, so the relationship should not be reduced to a universal numerical threshold.

What should I do if my instrument is reliable but validity evidence is weak?

Determine which interpretation is poorly supported and why. You may need to revise items, redefine the construct, collect additional validity evidence, use another instrument, or narrow the claims you make from the scores. Strong reliability should not be used to dismiss contradictory validity evidence.

09 · The Bottom Line

A Consistent Measure Can Consistently Measure the Wrong Thing

The Bottom Line

A measure can be reliable but not valid because producing consistent scores does not establish that those scores represent the construct or support the interpretation the researcher intends.

Treat reliability as necessary evidence about specific sources of measurement error, not as a shortcut to validity. Once consistency is evaluated, you still need evidence showing what the scores mean, for whom that interpretation is appropriate, and whether it is justified for the way you intend to use them.

10 · Sources and Further Reading

Authoritative Resources on Reliability and Validity

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes