Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Know Whether the Measures Were Good Enough?

A familiar scale or statistically reliable instrument is not automatically a good measure for every study. Ask whether the measure actually captures the construct required by the research question, works adequately in the population and context, and supports the interpretation the authors make from its scores.

150
Were the Research Measures Good Enough? Guide 150 of 247
01 · The Question

How Do You Know Whether a Study Actually Measured What It Claims to Have Measured?

A paper says that an intervention improved learning.

What did the researchers measure?

Perhaps students completed an achievement test. Perhaps instructors rated their performance. Perhaps the researchers analyzed course grades. Or perhaps students simply indicated how much they believed they had learned.

Those measures may all provide useful information. They do not provide the same information.

Measurement sits between an abstract research idea and the evidence used to support a conclusion. If that connection is weak, a beautifully designed study can produce a very precise answer about something other than what you thought it was studying.

The central question is therefore not merely whether the authors used a questionnaire, test, sensor, rubric, database variable, or established scale. It is whether the evidence produced by that measure supports the interpretation the researchers place on it.

02 · The Short Answer

Judge the Interpretation of the Measurement, Not Just the Instrument

In Brief

A research measure is good enough when there is sufficient evidence that its scores or observations can be interpreted appropriately for the construct, population, context, and purpose relevant to the study, with measurement error and important limitations understood well enough for the conclusion being drawn.

Check what was actually measured, how it was operationalized, whether reliability and validity evidence is appropriate to this use, whether measurement was comparable across groups or time points, and whether the authors' conclusion is broader than the measure can support.

03 · What You Need to Know

A Measure Is Only as Useful as the Interpretation It Supports

Start with the construct, then look at the measure

Researchers often investigate concepts that cannot be observed directly.

Motivation, engagement, anxiety, learning, digital literacy, trust, socioeconomic status, quality of life, academic achievement, cognitive load, and many other constructs must be represented through observable information.

That representation is the operationalization.

Construct The theoretical or substantive concept the researchers want to understand.
Measure The observations, responses, scores, indicators, ratings, records, or procedures used to represent that concept empirically.

Critical evaluation begins by asking whether the second provides defensible evidence about the first.

For example, if the construct is academic achievement, a standardized achievement test, course grade, instructor rating, and self-reported academic performance may each capture different aspects of the broader idea. None should automatically be treated as interchangeable.

Translate the authors’ claim back into what was actually observed

When a paper says that "engagement increased," do not stop at the construct label.

Find out what changed in the data.

Was engagement measured through a validated self-report scale? Attendance? Number of clicks in a learning management system? Time on task? Classroom observations? Completion rates? A combination of indicators?

Then rewrite the finding in measurement-level language:

Broad claim: Students became more engaged.

Observed evidence: Students in the intervention condition scored an average of five points higher on a self-report behavioral-engagement scale.

The second statement is less elegant, but it tells you what the evidence actually consists of.

Your task is then to decide how confidently that observation can support the broader interpretation.

A measure can be reliable without being valid for your purpose

Reliability and validity are related but different ideas.

Reliability broadly concerns the consistency or precision of measurement under specified conditions. Validity concerns the degree to which evidence and theory support the interpretation and use of scores for an intended purpose.

The Standards for Educational and Psychological Testing, developed jointly by the American Educational Research Association, American Psychological Association, and National Council on Measurement in Education, emphasizes validity as concerning the interpretations of scores for proposed uses rather than treating validity as a permanent property attached to a test.

Reliability question Are the resulting scores sufficiently consistent or precise for this purpose?
Validity question What evidence supports interpreting these scores as representing the construct and use claimed in this study?

A bathroom scale that consistently reports everyone's weight ten kilograms too high is consistent but inaccurate. Research measurement is usually more complicated than that example, but the principle remains: consistency alone does not establish that you measured the intended thing appropriately.

Do not ask whether an instrument “is valid” as though validity never changes

Researchers often write that they used "a validated questionnaire."

That is useful information, but it is not the end of evaluation.

Validity evidence accumulates for particular interpretations and uses of scores in particular contexts. An instrument developed for one population, language, culture, age group, discipline, or purpose may require additional evidence when used somewhere substantially different.

For example, a scale developed to measure digital competence among practicing teachers may not automatically function identically among first-year university students. A clinical screening instrument may not be suitable as a diagnostic instrument merely because its psychometric properties were strong for screening.

Ask:

  • What interpretation was the instrument originally designed to support?
  • In what population was the evidence established?
  • How similar is that population to the current sample?
  • Was the instrument translated, shortened, modified, or adapted?
  • Does the current study provide additional evidence supporting its use here?

The phrase "previously validated" should open these questions, not close them.

Validity is supported by multiple forms of evidence

Modern measurement standards generally do not treat validity as a collection of completely separate permanent types attached to an instrument. Instead, different sources of evidence contribute to the argument supporting a proposed interpretation of scores.

The Standards for Educational and Psychological Testing discusses evidence based on areas such as test content, response processes, internal structure, relationships with other variables, and consequences of testing.

Source of evidence Practical question
Content Do the items, tasks, or observations adequately represent the domain the researchers claim to measure?
Response processes Are participants interpreting and responding to the measure in ways consistent with the intended construct?
Internal structure Does the relationship among items or components correspond to the structure implied by the construct and score interpretation?
Relations to other variables Do scores relate to other measures or outcomes in theoretically expected ways?
Consequences What intended or unintended consequences follow from the interpretation and use of scores, where relevant?

You will not need to investigate every source of validity evidence for every paper. Focus on evidence necessary for the way the measure contributes to the study's central claim.

Content matters when the construct has several dimensions

Suppose researchers want to measure "AI literacy."

The construct might include conceptual understanding, practical skills, critical evaluation, ethical awareness, and other dimensions depending on the theoretical framework.

If the instrument contains only questions about whether participants recognize AI terminology, it may capture knowledge of terminology without adequately representing the broader construct.

Similarly, a three-item questionnaire about enjoyment may not justify a conclusion about multidimensional student engagement if engagement is theoretically defined to include behavioral, emotional, and cognitive dimensions.

Ask whether the content of the measure covers enough of the construct to support the label used in the conclusion.

How participants interpret the questions matters

A questionnaire item does not necessarily mean to participants what researchers intend it to mean.

Consider:

"I frequently use AI in my academic work."

One student may interpret "AI" as ChatGPT. Another may include grammar checkers, recommendation systems, translation tools, and automated transcription. "Frequently" might mean several times per day to one respondent and several times per semester to another.

If those ambiguities matter to the research question, the resulting score may be difficult to interpret.

Response-process evidence can involve cognitive interviewing, pilot testing, observation, think-aloud procedures, or other methods for investigating how respondents understand and respond to items.

Even when a paper does not report such evidence, you can inspect whether central questions are sufficiently clear for the intended interpretation.

Internal consistency is not the same thing as overall reliability

Many papers report Cronbach's alpha and then declare the measure "reliable."

Alpha can provide information about the internal consistency of item responses under particular assumptions. It does not answer every reliability question, and a high alpha does not establish validity.

Other reliability considerations may include test-retest consistency, agreement among raters, inter-rater reliability, intra-rater reliability, or measurement precision across the score range, depending on the instrument and purpose.

For example:

Measurement situation Relevant reliability concern
Multi-item questionnaire Consistency and dimensional structure of item responses
Two independent raters scoring essays Agreement or consistency between raters
Stable construct measured twice Test-retest reliability or stability
Observer coding classroom behavior Consistency of coding across observers or occasions
Diagnostic classification Consistency and classification performance appropriate to the intended use

Ask what form of consistency the study actually requires.

A very high Cronbach’s alpha is not automatically better

Researchers sometimes treat higher internal-consistency coefficients as unconditionally desirable.

That can be misleading.

A very high alpha can sometimes reflect substantial item redundancy, particularly when several questions are near-duplicates. Alpha is also affected by the number of items and assumptions about the scale.

More importantly, alpha cannot tell you whether the instrument measures the intended construct. Ten highly similar items about satisfaction may be internally consistent while remaining a poor measure of learning.

Watch Out

Do not use a reliability coefficient as a shortcut for measurement quality. A measure can produce highly consistent scores for a construct that is poorly defined, incompletely represented, or inappropriate to the study's conclusion.

Reliability should be considered in the current sample when relevant

Reliability is not simply a permanent number inherited from the original instrument-development paper.

Score reliability can vary across populations, contexts, score distributions, administration conditions, raters, and other features.

If authors report that an instrument had alpha =.91 in an earlier study, that tells you something about that earlier application. It does not guarantee the same measurement properties in the current sample.

Where appropriate, look for reliability evidence from the present study, especially when the instrument has been translated, modified, shortened, or applied to a substantially different population.

Modification can change the instrument

Researchers frequently adapt established instruments.

They may remove items to shorten a survey, change wording to fit a new context, alter response options, translate items, combine subscales, or select only certain items from a longer instrument.

These changes may be sensible. They can also change what the resulting score represents.

For example, an established 24-item scale may have evidence supporting four related dimensions. If researchers select six convenient items and calculate one total score, the validity evidence for the original instrument does not automatically transfer to the new six-item composite.

Whenever you see phrases such as "adapted from," "modified version," or "selected items from," inspect what changed.

Translation is a measurement issue, not merely a language issue

Translating a questionnaire involves more than replacing words in one language with words in another.

Concepts may not have exact equivalents. Cultural conventions can change how questions are understood. Response categories may function differently. Examples or situations familiar in one context may be unusual in another.

Good adaptation may involve forward translation, back translation, expert review, cognitive testing, pilot work, and psychometric evaluation, depending on the instrument and intended use.

The precise procedure required varies, but the critical question remains: what evidence supports interpreting the translated scores in the intended way?

Measurement equivalence matters when comparing groups

Suppose a study reports that one demographic group has higher "academic motivation" scores than another.

Before interpreting the difference substantively, consider whether the instrument functions comparably across the groups.

If particular items are understood differently, have different relevance, or relate differently to the underlying construct across groups, observed score differences may partly reflect measurement rather than true differences in the construct.

Measurement invariance analyses are one approach used with latent-variable measures to investigate whether measurement structure is sufficiently comparable across groups or time points for particular comparisons.

You do not need to demand formal invariance testing in every paper. The need depends on the measure, design, claims, and consequences. But when group comparisons are central, ask whether there is reason to believe the measurement scale has the same meaning across those groups.

The same issue applies when measuring change over time

A study may report that motivation increased from the beginning to the end of a semester.

That interpretation assumes the measure remains sufficiently comparable across time.

But participants may learn how to interpret items differently, response standards may change, or the construct itself may be understood differently after an intervention.

If the central claim concerns change, consider whether the measurement procedure supports comparing scores across the relevant time points.

Self-report is not inherently weak

Self-report measures are often criticized too broadly.

If the research question concerns attitudes, perceptions, beliefs, intentions, subjective experiences, symptoms, or self-assessed states, participants' reports may be the most direct evidence available.

The problem arises when self-report is treated as interchangeable with another construct.

Appropriate interpretation Students reported greater confidence in their research skills.
Potential overinterpretation Students became objectively more competent researchers, based only on self-reported confidence.

Self-efficacy and demonstrated competence may be related. They are not the same outcome.

Objective-looking measures can also be problematic

The label "objective" can create unwarranted confidence.

Administrative records may contain coding errors. Sensor data depend on algorithms and thresholds. Learning-management-system logs record platform activity, not necessarily cognitive engagement. Course grades can combine exams, attendance, participation, assignments, and instructor judgment. Automated scoring systems can introduce model-specific measurement error.

No measure becomes automatically valid simply because a human respondent did not complete a questionnaire.

Ask what process generated the data and what the resulting variable actually represents.

Proxy measures need justification

Researchers sometimes cannot measure the target construct directly and use a proxy.

For example:

  • login frequency as a proxy for engagement;
  • publication count as a proxy for research productivity;
  • income as a proxy for socioeconomic status;
  • course grades as a proxy for learning;
  • time spent on a page as a proxy for attention.

Proxies can be useful, sometimes necessarily so. But the relationship between proxy and construct needs justification.

The question is not whether the proxy is correlated with the construct in some general sense. It is whether it represents enough of what matters for the particular inference being made.

Single-item measures are not automatically invalid

Multi-item scales can capture complex constructs more comprehensively and permit evaluation of internal structure and consistency. That does not mean every single-item measure is unacceptable.

Some constructs are sufficiently concrete that one well-defined question may be adequate for a particular purpose. Other constructs are multidimensional or difficult to represent with one response.

Ask whether the construct can reasonably be represented by the item and whether evidence supports that use.

For example, age may require one item. A complex construct such as research self-efficacy probably requires considerably more thought.

Researcher-developed instruments deserve additional attention

A paper may state:

"A researcher-made questionnaire was used."

That is not automatically a methodological problem. New research questions sometimes require new measures.

But a newly developed instrument has not inherited an established body of evidence supporting its interpretation.

Look for information about:

  • how the construct was defined;
  • how items were generated;
  • whether experts reviewed the content;
  • whether intended respondents were involved in testing comprehension;
  • whether the instrument was piloted;
  • how scoring was determined;
  • what reliability evidence was obtained;
  • what validity evidence supports the intended interpretation.

A questionnaire does not become validated because three experts signed a content-validation form and Cronbach's alpha exceeded.70. Those procedures can contribute evidence, but measurement quality is broader than either one.

Expert review is useful but cannot establish validity by itself

Subject-matter experts can judge whether items appear relevant, clear, representative, and aligned with a conceptual domain. That can provide valuable content-related evidence.

Experts cannot determine by judgment alone how respondents will actually interpret items, whether scores have the intended internal structure, whether the measure relates appropriately to external variables, or whether it performs adequately in the target population.

Content review should therefore be treated as one source of evidence rather than a complete validation process.

Factor analysis does not automatically prove that a scale is valid

Exploratory and confirmatory factor analyses can provide evidence about the internal structure of multi-item measures.

They are valuable when the proposed interpretation assumes that items reflect one or more latent dimensions.

But a good-fitting factor model does not establish every aspect of validity. Items can form a coherent factor while inadequately representing the intended construct. Model fit can also depend on analytical choices, sample characteristics, and assumptions.

Internal structure is part of the validity argument, not the entire argument.

Criterion comparisons need an appropriate criterion

Some measures are evaluated by comparing them with another measure or reference standard.

This can be useful when the external criterion is meaningful and sufficiently trustworthy.

But describing another measure as a "gold standard" does not make it perfect. If the reference itself contains substantial error or measures a somewhat different construct, the comparison becomes harder to interpret.

Ask what the criterion represents and why it is appropriate for validating the new measure.

For diagnostic or classification measures, accuracy has several dimensions

A measure intended to classify people or cases may require evaluation of sensitivity, specificity, predictive values, likelihood ratios, discrimination, calibration, or other properties depending on the task.

No single metric tells the whole story.

A test can have high sensitivity and relatively low specificity, making it useful for one screening purpose and poor for another decision. Predictive values also depend on prevalence in the population where the test is used.

When classification is central, evaluate the measure according to its intended decision context rather than asking only whether its overall accuracy sounds high.

Measurement timing matters

Even an excellent instrument can answer the wrong temporal question if administered at the wrong time.

An intervention study may measure learning immediately after instruction and conclude that the intervention improves long-term learning. The instrument itself may be perfectly sound. The timing does not support the long-term claim.

Similarly, measuring stress only after an intervention provides no direct evidence about whether stress changed unless an appropriate comparison or prior measurement allows that change to be inferred.

Ask whether measurement occurred at the points required by the study design and research question.

Measurement procedures should be comparable across groups

Suppose an intervention group completes an outcome assessment online at home while the control group completes it under supervised classroom conditions.

Any observed difference may reflect both the intervention and the measurement context.

Likewise, if assessors know which participants received the intervention, subjective ratings can potentially be influenced by that knowledge.

Blinding of outcome assessment is therefore important in some designs, particularly when outcomes involve judgment.

Ask whether groups were measured using sufficiently similar procedures and whether assessors could have been influenced by knowledge of exposure or condition.

Inter-rater reliability matters when humans make judgments

Many outcomes require human judgment: essay quality, classroom behavior, interview coding, clinical diagnosis, image interpretation, performance assessment, or qualitative categorization.

If multiple raters are involved, examine how they were trained, what scoring criteria they used, whether they rated independently, and whether agreement or consistency was evaluated where appropriate.

A detailed rubric does not guarantee that two raters will apply it identically.

If one rater scored everything, ask what procedures support the consistency and credibility of those judgments and whether blinding was relevant.

Floor and ceiling effects can hide meaningful differences

A measure can be poorly matched to the range of the population.

If nearly everyone scores close to the maximum before an intervention, there may be little room to detect improvement. This is a ceiling effect.

If nearly everyone scores near the minimum, a floor effect can similarly reduce sensitivity to differences at the lower end.

A study may then conclude that groups do not differ when the measure simply cannot discriminate effectively in the relevant range.

Inspect score distributions when available and ask whether the instrument is appropriately targeted to the sample.

Measurement error can weaken observed relationships

Measurement is rarely perfectly precise.

Random measurement error can reduce the precision of estimates and, in many common settings, attenuate observed associations. Systematic measurement error can bias results in more complicated directions.

The consequences depend on the design and how measurement error relates to groups, exposures, outcomes, and other variables.

For example, if participants in one study condition systematically exaggerate improvement because they know they received a novel intervention, the error is not merely random noise. It can favor a particular conclusion.

This is why measurement quality connects directly to risk of bias.

Ask whether measurement error differs across groups

Differential measurement can be especially problematic.

Suppose teachers know which students participated in a new educational intervention and subsequently rate those students' engagement. Even without deliberate bias, expectations may influence ratings.

Or suppose cases recall past exposure more carefully than controls because they are searching for explanations for an outcome.

When measurement quality differs according to study group or outcome status, observed differences may partly reflect the measurement process itself.

Composite scores need conceptual justification

Researchers sometimes combine several variables into one index.

This can simplify analysis and represent multidimensional constructs. But the combination should make substantive and measurement sense.

Ask:

  • Why were these components combined?
  • Are they supposed to represent one construct?
  • Were they weighted equally or differently?
  • Does a high score have a coherent interpretation?
  • Could different combinations of components produce the same total score while meaning very different things?

A mathematically convenient total is not necessarily a meaningful construct.

Changing a continuous measure into categories can discard information

Researchers sometimes transform continuous scores into categories such as "low," "moderate," and "high."

That may be useful when thresholds have established substantive meaning. Arbitrary cutoffs can also discard information and create artificial distinctions between participants whose scores are nearly identical.

For example, students scoring 69 and 70 may be classified into different categories even though their observed scores differ by only one point.

When categorical conclusions matter, ask where the thresholds came from and whether they have a defensible interpretation.

Do not confuse statistical differences with measurement differences

If groups have different average scores, you still need to know what those scores represent.

A statistically significant difference on a questionable measure remains a difference on that questionable measure.

Increasing the sample size can make the estimate more precise. Sophisticated analysis can model the scores more elegantly. Neither repairs a fundamental construct mismatch.

This is why measurement can become a limitation or a fatal flaw depending on whether the measure remains capable of supporting the central inference.

The strongest measurement question is often “What can I call this result?”

Suppose a study uses a self-report scale of perceived learning and finds a substantial intervention effect.

You could write:

The intervention improved perceived learning.

That may be defensible if the scale adequately measures perceived learning.

You might not be able to write:

The intervention improved learning.

unless additional evidence justifies treating perceived learning as evidence of actual learning.

Critical measurement appraisal often leads not to rejecting the study but to choosing a more accurate noun.

04 · A Practical Example

A Reliable Questionnaire Can Still Measure the Wrong Outcome

Hypothetical Example

Did an AI tutor improve student learning?

Imagine a randomized study comparing an AI tutoring system with conventional online learning. The authors conclude that the AI tutor significantly improves student learning. The study is otherwise carefully conducted and includes 600 students.

Research claim The AI tutoring system improves learning.
Outcome measure Learning is measured using an eight-item questionnaire asking students how much they believe they learned and how helpful they found the system.
Reliability evidence The authors report Cronbach's alpha =.92. The items therefore show high internal consistency in this sample under the assumptions relevant to that coefficient.
Measurement problem Internal consistency does not establish that the questionnaire measures demonstrated learning. Several items appear to concern perceived helpfulness and satisfaction.
What the study may support Students assigned to the AI tutor reported greater perceived learning or perceived helpfulness, depending on the exact content and structure of the instrument.
What remains unsupported Without an appropriate performance measure or other evidence, the study cannot establish from this questionnaire alone that students actually learned more.
Why sample size does not rescue it Six hundred participants can estimate the difference in questionnaire scores precisely. It cannot change what those questionnaire scores represent.

The study's problem is not that the questionnaire is necessarily bad. It may be an excellent measure of perceived learning. The problem appears when the construct in the conclusion becomes broader than the construct represented by the measurement.

05 · What Researchers Often Get Wrong

Common Mistakes When Evaluating Research Measures

Misconception

If Cronbach’s alpha is above.70, the instrument is valid

No. Alpha provides information about internal consistency under particular assumptions. It does not establish content coverage, construct interpretation, criterion relationships, measurement equivalence, or other validity evidence. Reliability is necessary for many interpretations but is not equivalent to validity.

Misconception

A previously validated instrument is valid everywhere

No. Existing validity evidence is useful, but the defensibility of score interpretation depends on population, context, language, administration, purpose, and any modifications made to the instrument. Evidence from one use does not automatically establish every later use.

Misconception

Self-report measures are weak

Not inherently. Self-report can be highly appropriate for perceptions, beliefs, experiences, attitudes, intentions, and other subjective constructs. The problem is interpreting self-report as direct evidence of a different construct, such as observed behavior or demonstrated performance.

Misconception

An objective measure is automatically better

No. Administrative data, behavioral logs, sensors, automated scores, and other apparently objective measures still depend on definitions, algorithms, recording systems, thresholds, and data quality. Evaluate what the variable represents rather than relying on the label "objective."

Misconception

Expert validation proves that a new questionnaire is valid

Expert review can provide useful evidence about content and clarity, but it cannot by itself establish how respondents interpret items, whether scores have the expected structure, whether they relate appropriately to other variables, or whether they support the intended use.

Misconception

Factor analysis proves that the instrument measures the construct

Factor analysis can contribute evidence about internal structure. A coherent factor structure does not establish every aspect of score interpretation or demonstrate that the items adequately represent the theoretical construct.

Misconception

A statistically significant result proves the measure worked

No. Statistical significance concerns an analysis of the scores produced. It does not establish that those scores represent the construct the authors claim to have measured. A precise difference on the wrong outcome remains a precise difference on the wrong outcome.

06 · What This Means for You

Trace the Claim All the Way Back to the Observation

When measurement matters to a paper's central conclusion, reconstruct the chain from concept to data.

A simple decision framework

First: What construct does the paper claim to study?
Define the substantive concept behind the outcome, predictor, exposure, or other variable that matters to the conclusion.
Next: What did the researchers actually observe?
Identify the questionnaire responses, test scores, ratings, records, behavioral traces, sensor outputs, classifications, or other data used to represent the construct.
Next: Why should those observations represent the construct?
Look for theoretical justification and relevant validity evidence rather than relying only on the instrument's name or previous use.
Next: Are the scores sufficiently reliable or precise?
Evaluate the type of reliability relevant to the measurement procedure, such as internal consistency, stability, or rater agreement.
Next: Does the measure work for this population and context?
Consider adaptations, translation, cultural context, age, administration mode, group comparability, and other differences from settings where prior evidence was established.
Next: Was measurement conducted comparably?
Check timing, administration conditions, rater knowledge, group differences, and other factors that could introduce differential measurement error.
Finally: What is the strongest accurate name for the result?
Describe the finding using the construct the measure genuinely supports rather than the broader construct the authors may prefer.

A useful final test is: If I replaced the construct label in the paper with a literal description of what was measured, would the conclusion still mean the same thing?

If the answer is no, investigate the measurement more closely.

07 · A Quick Checklist

Were the Measures Good Enough?

For every measure central to the conclusion, check:
What construct, outcome, exposure, predictor, or phenomenon is this measure supposed to represent?
What was actually observed, asked, scored, recorded, classified, or calculated?
Does the content of the measure adequately represent the construct required by the research question?
What validity evidence supports the interpretation of these scores for this purpose?
Is the relevant form of reliability or measurement precision adequate for the study's use?
Was the measure previously developed for a population, language, context, and purpose sufficiently similar to this study?
If the measure was translated, shortened, adapted, or modified, is there evidence supporting the altered version?
If groups or time points are compared, is the measurement sufficiently comparable across them?
Could rater knowledge, self-report bias, recording procedures, missingness, or other measurement processes systematically affect the result?
Does the conclusion describe what was actually measured rather than silently substituting a broader construct?
08 · Frequently Asked Questions

Questions About Evaluating Research Measures

How can I tell whether a research instrument is valid?

Look for evidence supporting the interpretation and use of the resulting scores for the construct, population, context, and purpose relevant to the study. This may include evidence concerning content, response processes, internal structure, relationships with other variables, and other considerations appropriate to the measurement purpose. Avoid treating validity as a permanent yes-or-no label attached to an instrument.

What is the difference between reliability and validity?

Reliability broadly concerns consistency or precision of measurement, while validity concerns whether evidence and theory support the interpretation and use made from the scores. A measure can produce consistent scores without adequately measuring the intended construct, so reliability alone cannot establish validity.

Is Cronbach’s alpha enough to show that a questionnaire is reliable?

Not for every measurement purpose. Alpha concerns internal consistency under particular assumptions. Depending on the measure, you may also need information about dimensionality, test-retest stability, inter-rater agreement, measurement error, or other forms of reliability. Alpha also provides no direct proof of validity.

What is a good Cronbach’s alpha?

There is no universal threshold that establishes measurement quality. Acceptability depends on the purpose, construct, number and nature of items, consequences of measurement error, and assumptions of the coefficient. Mechanical rules such as "above.70 is good" should not replace examination of the scale's structure and intended use.

Can I trust a questionnaire because it was validated in a previous study?

Previous evidence increases the information available for evaluating the measure, but it does not guarantee appropriateness in every new application. Check whether the current population, language, context, administration, scoring, and intended interpretation resemble those for which the earlier evidence was established.

Are researcher-made questionnaires acceptable?

Yes, when a research question requires a new measure and the development process is sufficiently justified. Examine how the construct was defined, how items were developed and reviewed, how respondents understood them, how scoring was determined, and what reliability and validity evidence supports the resulting interpretation.

Are self-report measures less reliable than objective measures?

Not as a universal rule. The appropriate measurement method depends on the construct. Self-report may be necessary for subjective experiences and beliefs, while behavioral or administrative measures may be preferable for other questions. Both can contain error and bias. Evaluate whether the method matches what researchers claim to measure.

What should I do if the study uses a questionable measure?

Determine how central the measure is to the paper's conclusion and what narrower interpretation remains defensible. A weak secondary measure may have little effect on the main finding, while a primary outcome that does not adequately represent the central construct can seriously undermine the study. Adjust the evidential weight and wording of any claim you take from the paper accordingly.

09 · The Bottom Line

Ask What Was Measured Before Deciding What Was Found

The Bottom Line

A research measure is good enough when the observations or scores it produces can be interpreted defensibly as evidence about the construct required by the research question, with adequate reliability, relevant validity evidence, and measurement procedures appropriate to the population, context, and intended use.

Do not let familiar instrument names, high Cronbach's alpha values, previous validation, large samples, or statistical significance substitute for that judgment. Trace the conclusion back to what researchers actually observed. Sometimes the study's evidence is sound and only the label needs narrowing; sometimes the measurement problem reaches the heart of the central claim.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes