Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Could Different Measures Explain Why Studies Reach Different Conclusions?

Studies can investigate what appears to be the same outcome yet measure it in substantially different ways. Understanding what was measured, how it was measured, and when it was measured can reveal whether conflicting conclusions reflect a genuine difference or a measurement problem.

167
Can Different Measures Explain Conflicting Results? Guide 167 of 247
01 · The Question

Are the studies disagreeing, or are they measuring different things?

Two studies evaluate the same intervention in similar populations. One reports that it improves "engagement." Another finds little effect on engagement. The conclusions appear contradictory until you inspect the outcome measures.

The first study defines engagement as participation recorded in a digital platform. The second uses a self-report questionnaire about students' emotional and cognitive involvement. Both measures may be defensible, but they are not necessarily capturing the same aspect of engagement.

Measurement differences can create apparent disagreement in many fields. Researchers may use different instruments, operational definitions, thresholds, data sources, scoring procedures, or measurement times for what appears to be the same outcome. Before treating the findings as contradictory, you need to determine whether the studies actually measured sufficiently comparable phenomena.

02 · The Short Answer

Yes, measurement choices can materially change study findings

In Brief

Different measures can help explain why studies reach different conclusions when they operationalize the outcome differently, capture different dimensions of a construct, vary in validity or reliability, use different thresholds or scales, or assess outcomes at different times. The same outcome label does not guarantee that two studies measured the same thing in the same way.

Compare the underlying construct, operational definition, measurement instrument, scoring and thresholds, measurement properties, data source, and timing before interpreting the findings as genuinely contradictory. Some different measures can legitimately be synthesized; others represent meaningfully different outcomes that should remain separate.

03 · What You Need to Know

How measurement choices can change the apparent pattern of evidence

Start by separating what was measured from how it was measured

A useful distinction is between an outcome or construct and an outcome measurement instrument. COSMIN defines the outcome as what is being measured and the measurement instrument as how that outcome is measured.

For example, depression might be the construct of interest, while a particular questionnaire is the instrument used to measure it. Academic achievement might be the outcome, while a standardized examination, researcher-developed test, or course grade provides the measurement.

Outcome or construct The phenomenon the researcher intends to measure, such as anxiety, achievement, physical functioning, engagement, or quality of life.
Measurement instrument The method or tool used to obtain information about that outcome, such as a questionnaire, test, rating scale, observation protocol, laboratory measurement, administrative record, or device.

This distinction matters because two studies can share an outcome label while operationalizing the construct differently. Conversely, two different instruments may be legitimate ways of measuring sufficiently similar versions of the same construct.

The same word can conceal different constructs

Broad constructs are particularly vulnerable to measurement ambiguity. Terms such as engagement, well-being, achievement, quality of life, digital literacy, socioeconomic status, research impact, or job performance can have multiple defensible definitions.

Consider student engagement. One study might measure behavioral engagement through attendance and platform activity. Another might measure emotional engagement through students' feelings of belonging and interest. A third might assess cognitive engagement through self-regulated learning strategies.

All three studies can reasonably discuss "engagement," but a finding about one dimension does not automatically establish the same effect on the others.

If studies reach different conclusions, therefore, inspect the construct definitions before comparing the numbers. Sometimes what looks like disagreement is really a conceptual mismatch hiding behind shared terminology.

Operational definitions determine what becomes observable

Researchers cannot usually measure an abstract construct directly. They operationalize it by specifying observable indicators or procedures.

"Academic success," for instance, might be operationalized as examination performance, grade point average, course completion, retention, or graduation. Those outcomes are related, but they are not interchangeable.

An intervention could improve examination performance without affecting course completion. Another could reduce dropout without changing average test scores among students who remain enrolled. If the studies use different operational definitions of success, their apparently conflicting conclusions may both accurately describe the outcomes they measured.

Different instruments can measure the same construct with different quality

Even when two instruments are intended to measure the same construct, their measurement properties may differ.

COSMIN emphasizes several properties when evaluating measurement instruments, including validity, reliability, measurement error, and responsiveness. Content validity is particularly fundamental because researchers need confidence that an instrument adequately reflects the construct it is supposed to measure.

An instrument with weak validity may capture something other than the intended construct. An unreliable instrument can introduce substantial measurement error. An instrument with limited responsiveness may fail to detect meaningful change over time.

Consequently, two studies can investigate the same underlying phenomenon yet differ partly because one measure detects the relevant effect more accurately or sensitively than another.

Reliability affects how much noise surrounds an estimate

Measurement error can obscure real relationships. If an instrument produces highly variable measurements that are not attributable to genuine changes in the construct, estimates based on it may be less precise or attenuated.

Suppose two studies investigate whether an intervention improves a particular skill. One uses a well-validated multi-item assessment with good reliability. Another relies on a single noisy indicator. The second study might find a weaker or less precise association even if the underlying effect is similar.

This does not mean that every null finding can be blamed on poor reliability. The quality of the instrument needs evidence, not speculation. But when measures differ, their measurement properties belong in the comparison.

Responsiveness matters when the question involves change

An instrument can be useful for distinguishing among people at one point in time yet perform poorly at detecting change. COSMIN refers to the ability of an instrument to detect change over time in the construct being measured as responsiveness.

This becomes particularly important in intervention studies. If the intervention produces a modest but meaningful change and one measure is insensitive to that change, a study using it may report little apparent effect. Another instrument designed to capture changes in the relevant range may detect one.

The conclusion should not automatically be that the intervention works according to one study and fails according to another. The possibility that the measures differ in responsiveness needs examination.

Objective and self-reported measures may answer related but distinct questions

Researchers sometimes treat "objective" and self-reported measures as though one simply replaces the other. In practice, they can capture different aspects of a phenomenon.

For example, self-reported physical activity reflects participants' perceptions and recall, while wearable devices capture particular forms of recorded movement according to device and algorithm specifications. Self-reported learning may reflect perceived understanding, while a performance assessment tests demonstrated knowledge or skill.

Neither category is universally superior. The appropriate measure depends on the construct and research question. What matters when studies disagree is whether the measures are sufficiently comparable for their results to be interpreted as competing evidence.

Different thresholds can turn similar measurements into different categories

Some outcomes are measured continuously and then converted into categories. A scale might classify participants as having or not having a condition according to a cutoff score. Researchers can also define improvement as exceeding a particular threshold.

Different thresholds can change the proportion of participants classified as cases, responders, high performers, or otherwise meeting the outcome definition.

Imagine two studies using the same underlying test but defining "successful improvement" differently. One requires a five-point increase; another requires ten points. The studies could report different success rates even if the distributions of score changes were very similar.

Always inspect the outcome definition behind categorical labels.

The scale can affect how results are expressed

Two studies may measure the same continuous construct using different scales. One depression scale might range from 0 to 27 while another ranges from 0 to 60. An educational test might report raw scores while another uses standardized scores.

Raw mean differences on such scales cannot simply be compared as though the units were identical. When different instruments genuinely measure the same underlying construct, systematic reviews may use standardized effect measures to place results on a common metric.

Cochrane notes that the standardized mean difference is commonly used when studies assess the same continuous outcome using different measurement scales. This statistical transformation can facilitate synthesis, but it does not prove that the instruments are conceptually interchangeable.

Watch Out

Standardizing numerical results cannot repair a conceptual mismatch. Before combining studies that use different instruments, establish that the measures represent sufficiently similar outcomes. A common statistical scale does not turn different constructs into the same construct.

When an outcome is measured can matter as much as how

Suppose an intervention improves symptoms immediately after treatment, but the difference diminishes after six months. One study measuring the immediate outcome could conclude that the intervention is beneficial, while another measuring only long-term outcomes could find little difference.

These findings do not necessarily conflict. They may reveal the time course of the effect.

The same issue occurs in education, psychology, public health, organizational research, and other fields. Immediate learning, delayed retention, behavior change, and long-term outcomes are not interchangeable merely because they all follow the same intervention.

Record the measurement time whenever you compare study findings.

Who provides the measurement can change what is observed

Outcomes may be reported by participants, parents, teachers, clinicians, supervisors, independent assessors, administrative systems, or automated devices. Different informants have access to different information and may apply different judgments.

For example, a child may report improvements in emotional well-being that are not detected by a parent questionnaire. A teacher's assessment of classroom engagement may not correspond exactly with platform activity logs.

Disagreement across informants can sometimes be substantive rather than merely measurement error. Each source may be observing a different manifestation of the phenomenon.

Measurement procedures can introduce systematic differences

Even use of the same nominal instrument does not guarantee identical measurement. Translation, administration mode, instructions, scoring, assessor training, cultural adaptation, device calibration, and data-processing procedures can differ.

A questionnaire administered privately online may elicit different responses from the same questionnaire administered in a face-to-face interview. A performance assessment scored by trained blinded assessors may behave differently from one scored by individuals who know participants' treatment assignments.

When conflicting findings persist despite use of apparently identical measures, examine how the measurement was actually implemented.

Measurement differences are a recognized source of methodological heterogeneity

Cochrane distinguishes clinical diversity from methodological diversity when considering heterogeneity across studies. Differences in outcome measurement tools and in how outcomes are defined or measured can contribute to methodological diversity and may lead to differences in observed intervention effects.

This is important because heterogeneity associated with outcome assessment does not necessarily mean that the underlying true effect genuinely differs between populations. The observed difference may arise because studies measured the phenomenon differently.

That possibility should be investigated alongside population differences that might genuinely modify the effect.

Different measures do not automatically prevent synthesis

If studies use different instruments to assess the same sufficiently similar outcome domain, it may still be reasonable to synthesize them.

Cochrane describes several possible approaches. Researchers may combine different measures of the same outcome, perhaps examining measurement method through subgroup or sensitivity analyses, or synthesize individual outcome measures separately when that is more appropriate.

The choice depends on whether the measures are conceptually compatible and whether combining them produces a meaningful answer.

The important distinction is between different ways of measuring the same thing and measures that only appear to represent the same thing because researchers use the same broad label.

04 · A Practical Example

How one intervention can improve one measure but not another

Hypothetical Example

Does a digital learning intervention improve student engagement?

Imagine three hypothetical studies evaluating the same digital learning intervention among broadly similar undergraduate populations.

Study A Engagement is measured using platform logs. Students receiving the intervention complete more activities and access the learning platform more frequently than students in the comparison group.
Study B Engagement is measured using a self-report scale focused on interest, emotional involvement, and perceived connection to learning. The groups differ very little.
Study C Engagement is measured through independent classroom observations of participation. A small improvement is observed, but the estimate is imprecise.

A simplistic review might classify Study A as positive, Study B as null, and Study C as inconclusive, then declare the evidence inconsistent.

A more useful interpretation begins by examining the measures. Platform activity captures observable digital behavior. The questionnaire captures students' subjective experience. Classroom observation captures another form of behavioral engagement in a different environment.

The intervention could plausibly increase interaction with the platform without meaningfully changing students' emotional engagement. The findings therefore need not be contradictory.

The synthesis might instead conclude: the intervention appears to increase some forms of behavioral participation, while current evidence does not establish a comparable improvement in students' subjective engagement.

That conclusion is more specific than saying simply that the studies disagree. Measurement differences have helped identify what the intervention may actually be changing.

05 · What Researchers Often Get Wrong

Common mistakes when comparing studies that use different measures

Misconception

If studies use the same outcome label, they measured the same thing

Broad labels can conceal important conceptual differences. Always inspect the construct definition, operationalization, instrument, scoring, and measurement time rather than relying on terminology in titles, abstracts, or tables.

Misconception

If studies use different instruments, their results cannot be compared

Different instruments can sometimes measure sufficiently similar versions of the same construct. Appropriate statistical methods may permit synthesis, particularly for continuous outcomes measured on different scales. Conceptual comparability must be established before statistical comparability becomes useful.

Misconception

An objective measure is always more valid than a self-report measure

Validity depends on the construct and purpose of measurement. A device may be preferable for measuring some behaviors but incapable of measuring a person's subjective experience. Self-report can be indispensable when the construct itself concerns perceptions, symptoms, attitudes, or experiences.

Misconception

A validated instrument is valid for every population and purpose

Evidence supporting a measurement instrument is contextual. Measurement properties may need consideration for the population, language, setting, and intended use. The fact that an instrument has previously been called "validated" does not eliminate the need to examine whether it is suitable for the current application.

Misconception

Standardizing scores makes different outcomes equivalent

A standardized effect size can place results measured on different scales onto a common statistical metric. It does not establish that the underlying instruments measure the same construct. Conceptually different outcomes should not be combined merely because their numbers can be transformed.

Misconception

If the measures differ, measurement must explain the conflicting findings

Measurement is only one candidate explanation. Studies may also differ in populations, research designs, implementation, analyses, risk of bias, and sampling variability. Measurement differences should be evaluated rather than assumed to be causal.

06 · What This Means for You

How to determine whether measurement explains the disagreement

When studies report different findings, reconstruct the measurement process before comparing their conclusions. Begin with the construct and work outward to the instrument, scoring, timing, and interpretation.

A simple decision framework

If the studies use different labels for clearly comparable measures of the same construct
The findings may still be meaningfully compared or synthesized using an appropriate effect measure.
If the studies use the same label but operationalize substantially different constructs
Do not treat the findings as direct contradictions. Describe what each measure actually captures.
If instruments differ in validity, reliability, or responsiveness
Consider whether measurement quality could contribute to differences in observed effect magnitude or precision.
If studies use different thresholds or categorical definitions
Inspect the underlying continuous data or definitions where available before comparing rates or labels.
If outcomes are measured at different time points
Consider whether the findings describe different stages of the effect rather than contradictory evidence.
If measurement appears comparable but results still differ materially
Investigate other explanations, including population, research design, analysis, bias, and sampling variation.

Build a measurement comparison table

For each study, record the construct, operational definition, instrument, data source, respondent or assessor, scale, scoring procedure, cutoff where relevant, measurement time, and available evidence about measurement properties.

This simple exercise can expose differences that disappear in narrative summaries. Two papers may both report "anxiety," for example, while one measures current symptoms and another measures a broader trait. Two educational studies may both report "performance" while one assesses immediate recall and the other measures transfer to a novel task.

Ask whether the measure was capable of detecting the effect of interest

A measure should be appropriate not merely because it is familiar or widely used. Consider whether it captures the relevant construct, has adequate measurement properties for the intended population and purpose, and can detect the type of difference or change the study is designed to investigate.

COSMIN's guidance is useful when measurement instruments themselves require evaluation because it emphasizes systematic consideration of their quality rather than treating instrument selection as a procedural afterthought.

Keep measurement explanations proportional to the evidence

If the studies use clearly different constructs, measurement provides a strong reason not to call the findings directly contradictory. If they use two established instruments intended to measure the same construct, however, simply observing that the instrument names differ is much weaker evidence that measurement caused the discrepancy.

In the latter situation, examine whether findings differ systematically by measurement method across multiple studies or whether sensitivity analyses change the synthesis. Cochrane explicitly identifies such analyses as possible ways to investigate whether results are modified by the type of measurement method or tool.

Return to the larger pattern of evidence

Measurement rarely exists in isolation. If studies also differ in participants, design, and analysis, several explanations may operate simultaneously.

If design differences remain important, examine whether different research designs could explain the conflicting findings. When several factors remain plausible, avoid pretending that measurement alone has resolved the disagreement.

The broader goal is to determine whether the literature is truly inconsistent or simply reflects a more complex pattern of evidence.

07 · A Quick Checklist

Before attributing conflicting findings to measurement, check these points

When comparing outcome measures, check:
Do the studies define the underlying outcome or construct in the same way?
Do the instruments capture the same dimension of that construct?
Are the operational definitions, scoring procedures, scales, and thresholds comparable?
Who or what provided the measurement: participant, observer, clinician, administrative record, test, or device?
Is there appropriate evidence concerning validity, reliability, measurement error, and responsiveness for the instrument and intended use?
Were outcomes measured at comparable time points?
Could administration, translation, assessor training, scoring, calibration, or data-processing procedures differ across studies?
If different scales measure the same construct, has an appropriate effect measure been used to compare or synthesize them?
Could population, design, analysis, bias, or sampling variation provide a stronger explanation for the disagreement?
08 · Frequently Asked Questions

Questions about different measures and conflicting research findings

Can two studies measure the same outcome differently?

Yes. The same construct can often be measured using different instruments, scales, data sources, assessors, or procedures. Whether the resulting findings are comparable depends on whether those approaches genuinely capture sufficiently similar versions of the underlying outcome.

What is the difference between an outcome and an outcome measurement instrument?

The outcome or construct is what you want to measure, such as depression, achievement, or physical functioning. The outcome measurement instrument is how you measure it, such as a questionnaire, examination, observation protocol, laboratory test, or device.

Can different measurement scales be combined in a meta-analysis?

Sometimes. When studies measure the same continuous outcome using different scales, standardized effect measures such as the standardized mean difference may permit synthesis. This requires the underlying outcomes to be sufficiently comparable; statistical standardization cannot compensate for conceptually different constructs.

Does using a validated instrument guarantee accurate measurement?

No. Measurement properties should be considered in relation to the construct, population, setting, language, and intended purpose. Evidence that an instrument performed well in one context does not automatically establish that it is optimal in every other context.

Can measurement error cause a study to find no effect?

Measurement error can reduce precision and, depending on the measurement process and study design, can distort estimated associations or effects. However, a null finding should not automatically be attributed to measurement error. Evidence about the instrument and measurement process is needed.

Can one study find an immediate effect while another finds no long-term effect without contradicting it?

Yes. Effects can change over time. Studies measuring different follow-up periods may be describing different stages of the same phenomenon rather than providing incompatible conclusions.

Are self-reported outcomes less trustworthy than objective measures?

Not categorically. The appropriate measurement approach depends on the construct. Some phenomena, such as subjective symptoms, perceptions, or attitudes, inherently require participant report. Other outcomes may be better captured through direct observation, records, tests, or devices. Each approach has its own measurement considerations.

How do I know whether measurement really explains conflicting findings?

Look for more than the existence of different instruments. Determine whether the measures capture different constructs or dimensions, whether their measurement properties differ meaningfully, and whether results systematically vary according to measurement approach. Then compare this explanation with alternatives such as population, design, analysis, and bias.

09 · The Bottom Line

Before comparing results, make sure the studies measured comparable outcomes

The Bottom Line

Different measures can explain why studies reach different conclusions when they capture different constructs or dimensions, differ in measurement quality or responsiveness, use different thresholds or data sources, or assess outcomes at different times. The same outcome label does not guarantee equivalent measurement.

Start with what each study intended to measure, then examine how that construct was operationalized and measured. Different instruments can sometimes provide comparable evidence and be synthesized appropriately; in other cases, the measurement difference reveals that the studies were never estimating quite the same outcome. The distinction should be established before calling the evidence contradictory.

10 · Sources and Further Reading

Sources and further reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes