03 · What You Need to Know
Research Is Designed to Manage the Possibility of Being Wrong
Research Usually Observes a Sample, Not the Entire Phenomenon
Many studies cannot observe every person, event, location, or possible instance relevant to a research question. Researchers therefore examine samples and use them to learn about a broader population or process.
Samples vary.
Imagine that an intervention has a modest beneficial effect on average. Two carefully selected random samples may nevertheless produce somewhat different estimates simply because different individuals happened to be included. In one sample, the estimated effect may be larger than the underlying average. In another, it may be smaller. Occasionally, the observed result may even point in the wrong direction.
This variability does not necessarily indicate that anyone made a mistake. It is one reason researchers quantify sampling uncertainty where appropriate rather than treating every sample estimate as the exact population value.
A Result Can Be Unusual Even When the Study Is Conducted Correctly
Random variation occasionally produces observations that are not representative of the broader pattern.
Suppose two treatments genuinely have similar average effects. A properly conducted comparison can still produce a difference between the groups in a particular sample. Statistical methods help researchers evaluate uncertainty around such differences, but they do not make misleading samples impossible.
This is one reason research produces evidence rather than absolute proof. An empirical result changes what conclusions are reasonable without making error logically impossible.
Measurements Can Be Imperfect Without Being Obviously Bad
Researchers rarely observe every concept of interest directly. They use instruments, assessments, survey items, behavioral indicators, laboratory procedures, coding schemes, sensors, administrative records, or other operational measures.
Even good measurements have limitations.
A test may capture some dimensions of learning better than others. A questionnaire may imperfectly represent motivation. A sensor may have measurement error. Participants may not recall past behavior accurately. Researchers coding complex qualitative material may encounter ambiguous cases.
These limitations can affect the resulting evidence even when researchers followed accepted procedures carefully.
A Valid Measure Can Still Have Limited Scope
A related problem occurs when researchers measure something accurately but interpret it too broadly.
For example, an intervention might genuinely improve performance on a particular assessment without improving every dimension of the broader construct called “learning.” A study may accurately document short-term behavioral change without establishing that the change persists for years.
The finding itself may be correct while the broader conclusion is not.
This distinction illustrates why a finding, evidence, and a conclusion are not the same thing. Error can enter when researchers move from what was observed to what they infer that observation means.
Research Designs Depend on Assumptions
Every research design relies on assumptions, although their form differs across methodologies.
A statistical model may assume a particular relationship among variables. A causal analysis may depend on assumptions about confounding or treatment assignment. A survey may assume that respondents interpret questions as intended. A qualitative interpretation may depend on the adequacy of sampling, data generation, analytical procedures, and contextual understanding.
Some assumptions can be examined directly. Others are only partially testable from the available data.
A study may therefore be conducted competently under reasonable assumptions and still reach an incorrect conclusion if an important assumption turns out not to hold.
Unknown Factors Can Matter
Researchers can address only the alternative explanations and relevant variables they can identify and investigate.
A previously unknown biological mechanism, social process, environmental factor, behavioral response, technological change, or interaction among variables may influence the phenomenon in ways that the original study could not anticipate.
Later research sometimes changes a conclusion precisely because it identifies something earlier researchers had no adequate reason or ability to measure.
Good methodology can protect against many known threats. It cannot control a variable that nobody yet realizes is consequential.
Bias and Random Error Are Different Problems
It is useful to distinguish systematic error from random variation.
Random error
Variation that causes observations or estimates to fluctuate unpredictably around an underlying value or pattern.
Systematic error or bias
A process that tends to shift observations, estimates, or conclusions in a particular direction because of features of the design, measurement, sampling, analysis, or other procedures.
Increasing the amount of appropriate data can often improve precision and reduce some forms of random error. It does not automatically eliminate systematic bias. A very large biased sample can produce a highly precise estimate of the wrong quantity.
Conversely, a well-designed study that minimizes important biases can still produce an unusual estimate because random variation remains.
Statistical Decisions Carry Error Possibilities
In statistical inference, researchers often make decisions under uncertainty. Depending on the analytical framework, this can include the possibility of concluding that an effect or relationship exists when it does not, or failing to identify one that does exist.
These possibilities are not evidence that statistical reasoning is useless. They reflect the unavoidable difficulty of inferring an underlying process from limited observations.
Problems arise when statistical procedures are treated as mechanisms for producing certainty. A threshold such as statistical significance cannot guarantee that a substantive conclusion is correct. The National Academies has specifically cautioned against inappropriate reliance on statistical significance in evaluating replicability and scientific findings.
Researchers May Draw a Conclusion That Is Too Broad
Sometimes the problem lies less in the observed result than in the scope assigned to it.
A study conducted among university students in one country may accurately describe those participants. The conclusion becomes vulnerable if it is generalized without justification to all students, all educational systems, or entirely different age groups.
Similarly, evidence of an association does not automatically justify a causal conclusion, and evidence of a short-term effect does not automatically establish long-term effectiveness.
The strength of a conclusion therefore depends partly on whether the research claim matches the strength and scope of the evidence.
Rigor Reduces Avoidable Sources of Error
Scientific rigor matters precisely because research can go wrong. The U.S. National Institutes of Health defines scientific rigor as the strict application of the scientific method to support robust and unbiased design, methodology, analysis, interpretation, and reporting.
Depending on the study, rigorous practices may include appropriate controls, randomization, blinding, careful measurement, adequate sample-size justification, transparent analytical procedures, consideration of relevant variables, sensitivity analyses, or other safeguards.
These practices can substantially improve the reliability of evidence. They should not be interpreted as a certification that a result must be correct.
Watch Out
“The conclusion was later shown to be incorrect” and “the original researchers conducted poor research” are different claims. Sometimes both are true. Sometimes only the first is.
Later Research Provides Information the Original Study Did Not Have
A subsequent investigation may use a larger sample, improved measurement, another population, a stronger design, additional controls, longer follow-up, or a method capable of detecting something the original study could not.
Researchers can then evaluate the earlier conclusion with more information than was available when it was first proposed.
This is one reason scientific conclusions can change. The question is not whether researchers should somehow have known all future evidence in advance. It is whether the original conclusion was reasonable given the evidence available and whether it is revised appropriately when better evidence emerges.
A Failed Replication Does Not Automatically Prove the Original Study Wrong
The reverse caution is equally important. If a later study obtains a different result, the original conclusion does not instantly become false.
The National Academies notes that even rigorously conducted and correctly analyzed research may fail to replicate. Non-replication can arise from previously unrecognized variability, differences in methods or conditions, limitations of measurement, problems in one or both studies, or other causes.
Understanding why well-conducted studies can reach different conclusions therefore requires examining the studies rather than deciding that disagreement itself identifies which result is wrong.