01 · The Question
What If Your Instrument Is Not Measuring as Well as You Expected?
You selected or developed a measure, included it in the protocol, and started collecting data. Then problems appear. Participants consistently misunderstand an item. A scale has unexpectedly poor reliability. A device produces implausible readings. Interview questions fail to elicit information needed for the research question. A measure that performed adequately elsewhere seems poorly suited to your population.
The temptation is to fix the instrument immediately. Rewrite confusing questions. Remove troublesome items. substitute another scale. Change the administration procedure. These responses may improve measurement going forward, but they can also create a second version of the measurement process midway through the study.
The central question is therefore not simply whether the instrument should be changed. You first need to determine what is failing, whether that failure threatens the interpretation of the measurements, and what a correction would mean for observations already collected.
03 · What You Need to Know
First Determine What “Not Working” Actually Means
An Instrument Can Fail in Several Different Ways
A measure can produce numbers or responses perfectly well and still fail scientifically. The important question is whether those observations support the interpretation you need to make.
SPIRIT 2025 emphasizes that the quality of outcome data depends on properties such as reliability, validity, and responsiveness of measurement instruments. Low reliability can reduce statistical power, while poor validity means that an instrument may not accurately measure the intended outcome.
| Observed problem |
Possible issue |
What to investigate |
| Participants misunderstand questions |
Comprehension or population fit |
Wording, language, literacy demands, cultural interpretation, administration |
| Items do not behave consistently |
Reliability or scale structure |
Whether the expected measurement model is appropriate for this population and purpose |
| Scores do not appear to reflect the intended construct |
Validity of score interpretation |
Whether the instrument provides appropriate evidence for the intended use |
| Most participants receive very high or very low scores |
Range, targeting, ceiling, or floor problem |
Whether the measure can distinguish participants at the relevant levels |
| A device gives unstable or implausible values |
Technical, calibration, procedural, or environmental problem |
Equipment function, calibration, operator procedure, data transfer |
| Interview questions produce little relevant information |
Question design or methodological fit |
Whether prompts elicit evidence capable of addressing the research question |
| Different assessors produce inconsistent observations |
Inter-rater reliability or implementation |
Training, coding rules, administration consistency, construct ambiguity |
These diagnoses lead to different solutions. Retraining assessors may address inconsistent administration. It will not repair an instrument that measures the wrong construct.
A Published or Previously Validated Instrument Is Not Valid for Every Use
Researchers sometimes speak of a questionnaire as though validity were a permanent property attached to its name. The situation is more nuanced. Evidence supporting measurement in one population, language, setting, administration mode, or purpose does not automatically establish that every new interpretation and use is equally well supported.
A measure developed with adults may behave differently with adolescents. Translation can affect item meaning. Moving from interviewer administration to self-administration can alter responses. A scale designed for clinical screening may not necessarily be appropriate for measuring subtle change in a nonclinical population.
This does not mean that every instrument must be redeveloped from scratch for every study. It means that researchers should examine whether the existing evidence is relevant to the population and use in question rather than relying only on the label “validated.”
Do Not Confuse Reliability With Validity
Reliability
Concerns the consistency or reproducibility of measurements under relevant conditions.
Validity
Concerns whether evidence supports the interpretation and use of the measurements for the intended purpose.
A measure can produce highly consistent scores while systematically measuring something other than the construct the researcher intends. Conversely, an apparently low reliability estimate does not tell you, by itself, exactly why measurement is failing.
A single coefficient should therefore not become an automatic trigger for deleting items or replacing an instrument. Examine the measurement model, instrument design, population, administration, scoring, and substantive meaning of the items before deciding what the statistic implies.
Check Whether the Instrument Failed or Its Administration Failed
Before changing the measure itself, investigate whether the problem arose in implementation.
Were instructions followed consistently? Were assessors trained appropriately? Was the device calibrated? Were questionnaires displayed correctly? Did a software implementation reverse or omit response options? Were translations administered as intended? Did the setting influence responses? Were scoring rules implemented correctly?
Correcting an implementation error may preserve the intended instrument. Rewriting the instrument when administration is the actual problem may create unnecessary methodological disruption.
Changing Items Mid-Study Can Create Two Different Measures
Suppose 150 participants answer one questionnaire and you then rewrite three items before another 150 participants complete it. It may be tempting to combine all 300 responses because the questionnaire still has the same title.
But the later participants did not receive exactly the same measurement procedure. A change in wording, response options, scoring, administration mode, item order, prompts, or timing can alter how responses are generated.
The critical question becomes whether scores from the two versions are sufficiently comparable for the intended analysis. That cannot be assumed merely because the revised instrument seems better.
Watch Out
Do not overwrite the original instrument or recode earlier observations to make them look as though they were collected using the revised version. Preserve version history and distinguish what each participant actually received.
Replacing an Instrument Can Change the Outcome Itself
Two instruments can carry similar labels while operationalizing a construct differently. They may use different items, scoring systems, time frames, response scales, thresholds, or conceptual definitions.
In research where an instrument defines a prespecified outcome, replacing that instrument can therefore be more consequential than changing the data-collection tool. CONSORT and SPIRIT guidance specifies that outcomes should be defined in sufficient detail to identify the measurement variable, analysis metric, aggregation method, and time point. Changes to prespecified outcomes should be disclosed and justified.
If the instrument is central to the primary outcome, replacing it may affect sample-size assumptions, analysis, interpretation, and the comparability of observations collected before and after the change.
Do Not Drop Items Merely Because They Hurt a Reliability Statistic
Removing items after examining the study data can improve a statistic such as internal consistency while simultaneously narrowing or altering the construct represented by the scale.
An item may appear statistically inconvenient precisely because it captures an aspect of the construct that differs from the others. Conversely, a problematic item may genuinely reflect poor wording or inappropriate content. The decision requires substantive and measurement reasoning, not simply optimization of a coefficient.
If items are removed or scoring is modified after examining the data, preserve the original specification and disclose the change. Analyses based on revised scoring may need to be distinguished from those that were prespecified.
Consider Whether Existing Data Can Still Be Used
A measurement problem does not automatically require discarding everything collected before it was discovered. The appropriate treatment depends on the nature and severity of the problem.
If a device malfunctioned only during a known period, affected observations may be identifiable. If one questionnaire item was displayed incorrectly, other items may remain usable. If the entire measure lacks validity for the intended construct, however, technically complete responses may still fail to answer the research question.
The important distinction is between data existence and data fitness. Having a value in every row of a dataset does not guarantee that those values represent what the study needs to know.
A Measurement Problem May Reveal a Larger Methodological Mismatch
Sometimes the instrument is not the real problem. The study may be attempting to answer a question that the chosen form of evidence cannot adequately address.
For example, repeatedly modifying a short self-report questionnaire may not solve a research question that fundamentally requires behavioral observation, longitudinal measurement, qualitative inquiry, or another source of evidence. If repairing the instrument still leaves a gap between the evidence and the question, consider whether the planned method no longer fits the research question.
Any Mid-Study Modification Should Remain Visible
Once data collection has started, a consequential instrument modification becomes part of the methodological history of the study. Record what malfunctioned, how the problem was identified, which participants were affected, what changed, when the revised version began, and what consequences were considered.
If the modification changes approved participant-facing procedures, study outcomes, or other protocol elements, determine whether ethics approval, the protocol, or preregistration needs to be revisited before implementing it.
The final report should likewise distinguish planned and subsequent decisions when the difference matters. APA's Journal Article Reporting Standards emphasize transparent reporting of methods and distinguishing primary, secondary, and exploratory analyses, while CONSORT requires important post-commencement protocol changes and their reasons to be reported.
06 · What This Means for You
Repair the Measurement Process Without Erasing Its History
When an instrument begins behaving unexpectedly, resist the urge to optimize it immediately against the accumulating dataset. Diagnose the problem first, preferably using evidence that distinguishes technical failure, administration failure, population mismatch, and a genuine problem with the measure itself.
A simple decision framework
If the problem is an administration or implementation error
Correct the procedure prospectively where appropriate and identify which earlier observations may have been affected.
If one component of the instrument appears problematic
Evaluate its substantive and measurement role before removing, rewriting, or rescoring it.
If the instrument appears unsuitable for the study population
Assess whether another measure is defensible and whether switching instruments would compromise comparability with data already collected.
If the instrument defines a primary or important outcome
Treat replacement or substantial modification as a potentially consequential methodological change requiring careful documentation and any applicable approvals.
If no available measurement approach can adequately address the original question
Reconsider the method or study design rather than continuing to collect measurements that cannot support the intended inference.
If the solution requires modifying the instrument after participants have already been assessed, treat it as a potential change to the study after data collection has started, not merely as an editorial correction to the questionnaire.
And if you do make the change, keep enough documentation for another researcher to reconstruct exactly which instrument version generated which observations. Methodological archaeology is much less entertaining when you are excavating your own dataset six months later.
07 · A Quick Checklist
Before Changing a Research Measure, Check These Points
When an instrument is not working as expected, check:
Define the measurement problem precisely rather than relying on a general impression that the instrument is performing poorly.
Check administration, training, scoring, software, calibration, data entry, and other implementation issues before changing the instrument itself.
Review whether existing evidence for the instrument is relevant to your population, setting, language, administration mode, and intended use.
Assess what any modification would change about the construct, outcome, score, or interpretation.
Determine whether measurements collected before and after a change can legitimately be compared or combined.
Preserve the original instrument, scoring rules, data, and version history.
Verify whether modifying or replacing the measure requires ethics, protocol, registry, preregistration, sponsor, or other approval.
Document and report consequential measurement changes and their implications for analysis and interpretation.