Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Why Can Good Research Still Produce an Incorrect Conclusion?

A well-designed and carefully conducted study can still reach a conclusion that later evidence does not support. Research reduces the risk of error, but empirical evidence rarely eliminates uncertainty completely.

35
Why Good Research Can Be Wrong Guide 35 of 533
01 · The Question

If the Research Was Done Properly, How Can the Conclusion Still Be Wrong?

It is easy to assume that an incorrect research conclusion must have resulted from bad research. Perhaps the researchers used an inappropriate method, analyzed the data incorrectly, ignored contradictory evidence, or allowed bias to influence the study.

Those problems certainly can produce unreliable conclusions. But they are not the only possibilities.

Even a carefully designed, competently executed, transparently reported study can produce a result that later turns out not to represent the underlying phenomenon accurately. Samples vary. Measurements are imperfect. Statistical estimates contain uncertainty. Study conditions capture only part of a complex world. Some influential variables may remain unknown.

Good research manages these sources of uncertainty. It does not make researchers infallible.

02 · The Short Answer

Good Methods Reduce Error; They Cannot Guarantee a Correct Conclusion

In Brief

Good research can produce an incorrect conclusion because empirical studies operate with incomplete observations, variable samples, imperfect measurements, methodological assumptions, and uncertainty that rigorous methods can reduce but not always eliminate.

An incorrect conclusion therefore does not automatically imply misconduct or incompetence. The important distinction is between conclusions that turn out to be wrong despite reasonable research practices and conclusions made unreliable by avoidable flaws in design, conduct, analysis, interpretation, or reporting.

03 · What You Need to Know

Research Is Designed to Manage the Possibility of Being Wrong

Research Usually Observes a Sample, Not the Entire Phenomenon

Many studies cannot observe every person, event, location, or possible instance relevant to a research question. Researchers therefore examine samples and use them to learn about a broader population or process.

Samples vary.

Imagine that an intervention has a modest beneficial effect on average. Two carefully selected random samples may nevertheless produce somewhat different estimates simply because different individuals happened to be included. In one sample, the estimated effect may be larger than the underlying average. In another, it may be smaller. Occasionally, the observed result may even point in the wrong direction.

This variability does not necessarily indicate that anyone made a mistake. It is one reason researchers quantify sampling uncertainty where appropriate rather than treating every sample estimate as the exact population value.

A Result Can Be Unusual Even When the Study Is Conducted Correctly

Random variation occasionally produces observations that are not representative of the broader pattern.

Suppose two treatments genuinely have similar average effects. A properly conducted comparison can still produce a difference between the groups in a particular sample. Statistical methods help researchers evaluate uncertainty around such differences, but they do not make misleading samples impossible.

This is one reason research produces evidence rather than absolute proof. An empirical result changes what conclusions are reasonable without making error logically impossible.

Measurements Can Be Imperfect Without Being Obviously Bad

Researchers rarely observe every concept of interest directly. They use instruments, assessments, survey items, behavioral indicators, laboratory procedures, coding schemes, sensors, administrative records, or other operational measures.

Even good measurements have limitations.

A test may capture some dimensions of learning better than others. A questionnaire may imperfectly represent motivation. A sensor may have measurement error. Participants may not recall past behavior accurately. Researchers coding complex qualitative material may encounter ambiguous cases.

These limitations can affect the resulting evidence even when researchers followed accepted procedures carefully.

A Valid Measure Can Still Have Limited Scope

A related problem occurs when researchers measure something accurately but interpret it too broadly.

For example, an intervention might genuinely improve performance on a particular assessment without improving every dimension of the broader construct called “learning.” A study may accurately document short-term behavioral change without establishing that the change persists for years.

The finding itself may be correct while the broader conclusion is not.

This distinction illustrates why a finding, evidence, and a conclusion are not the same thing. Error can enter when researchers move from what was observed to what they infer that observation means.

Research Designs Depend on Assumptions

Every research design relies on assumptions, although their form differs across methodologies.

A statistical model may assume a particular relationship among variables. A causal analysis may depend on assumptions about confounding or treatment assignment. A survey may assume that respondents interpret questions as intended. A qualitative interpretation may depend on the adequacy of sampling, data generation, analytical procedures, and contextual understanding.

Some assumptions can be examined directly. Others are only partially testable from the available data.

A study may therefore be conducted competently under reasonable assumptions and still reach an incorrect conclusion if an important assumption turns out not to hold.

Unknown Factors Can Matter

Researchers can address only the alternative explanations and relevant variables they can identify and investigate.

A previously unknown biological mechanism, social process, environmental factor, behavioral response, technological change, or interaction among variables may influence the phenomenon in ways that the original study could not anticipate.

Later research sometimes changes a conclusion precisely because it identifies something earlier researchers had no adequate reason or ability to measure.

Good methodology can protect against many known threats. It cannot control a variable that nobody yet realizes is consequential.

Bias and Random Error Are Different Problems

It is useful to distinguish systematic error from random variation.

Random error Variation that causes observations or estimates to fluctuate unpredictably around an underlying value or pattern.
Systematic error or bias A process that tends to shift observations, estimates, or conclusions in a particular direction because of features of the design, measurement, sampling, analysis, or other procedures.

Increasing the amount of appropriate data can often improve precision and reduce some forms of random error. It does not automatically eliminate systematic bias. A very large biased sample can produce a highly precise estimate of the wrong quantity.

Conversely, a well-designed study that minimizes important biases can still produce an unusual estimate because random variation remains.

Statistical Decisions Carry Error Possibilities

In statistical inference, researchers often make decisions under uncertainty. Depending on the analytical framework, this can include the possibility of concluding that an effect or relationship exists when it does not, or failing to identify one that does exist.

These possibilities are not evidence that statistical reasoning is useless. They reflect the unavoidable difficulty of inferring an underlying process from limited observations.

Problems arise when statistical procedures are treated as mechanisms for producing certainty. A threshold such as statistical significance cannot guarantee that a substantive conclusion is correct. The National Academies has specifically cautioned against inappropriate reliance on statistical significance in evaluating replicability and scientific findings.

Researchers May Draw a Conclusion That Is Too Broad

Sometimes the problem lies less in the observed result than in the scope assigned to it.

A study conducted among university students in one country may accurately describe those participants. The conclusion becomes vulnerable if it is generalized without justification to all students, all educational systems, or entirely different age groups.

Similarly, evidence of an association does not automatically justify a causal conclusion, and evidence of a short-term effect does not automatically establish long-term effectiveness.

The strength of a conclusion therefore depends partly on whether the research claim matches the strength and scope of the evidence.

Rigor Reduces Avoidable Sources of Error

Scientific rigor matters precisely because research can go wrong. The U.S. National Institutes of Health defines scientific rigor as the strict application of the scientific method to support robust and unbiased design, methodology, analysis, interpretation, and reporting.

Depending on the study, rigorous practices may include appropriate controls, randomization, blinding, careful measurement, adequate sample-size justification, transparent analytical procedures, consideration of relevant variables, sensitivity analyses, or other safeguards.

These practices can substantially improve the reliability of evidence. They should not be interpreted as a certification that a result must be correct.

Watch Out

“The conclusion was later shown to be incorrect” and “the original researchers conducted poor research” are different claims. Sometimes both are true. Sometimes only the first is.

Later Research Provides Information the Original Study Did Not Have

A subsequent investigation may use a larger sample, improved measurement, another population, a stronger design, additional controls, longer follow-up, or a method capable of detecting something the original study could not.

Researchers can then evaluate the earlier conclusion with more information than was available when it was first proposed.

This is one reason scientific conclusions can change. The question is not whether researchers should somehow have known all future evidence in advance. It is whether the original conclusion was reasonable given the evidence available and whether it is revised appropriately when better evidence emerges.

A Failed Replication Does Not Automatically Prove the Original Study Wrong

The reverse caution is equally important. If a later study obtains a different result, the original conclusion does not instantly become false.

The National Academies notes that even rigorously conducted and correctly analyzed research may fail to replicate. Non-replication can arise from previously unrecognized variability, differences in methods or conditions, limitations of measurement, problems in one or both studies, or other causes.

Understanding why well-conducted studies can reach different conclusions therefore requires examining the studies rather than deciding that disagreement itself identifies which result is wrong.

04 · A Practical Example

A Carefully Conducted Study Can Still Misestimate an Effect

Hypothetical Example

A New Teaching Strategy Appears Highly Effective

Suppose a carefully designed randomized study tests a new teaching strategy among 120 students. The intervention group performs substantially better than the comparison group.

Design Participants are appropriately randomized, the outcome is measured consistently, the planned analysis is followed, and the study is transparently reported.
Finding The observed difference is considerably larger than researchers expected.
Initial conclusion The evidence reasonably suggests that the strategy improves performance under the studied conditions.
Later evidence Several larger studies also find a benefit, but their estimates consistently indicate a much smaller average effect.
Revised understanding The original study may have correctly detected a beneficial effect while substantially overestimating its typical magnitude because its particular sample produced an unusually large result.

Calling the first study “bad research” would miss the point. Its methods may have been sound. What changed was the amount of evidence available for estimating the underlying effect.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Incorrect Research Conclusions

Misconception

If the Study Was Rigorous, Its Conclusion Must Be Correct

Rigor reduces avoidable bias and error but cannot eliminate every source of uncertainty. Random variation, imperfect measurement, unrecognized variables, limited scope, and assumptions can still affect conclusions.

Misconception

If a Conclusion Is Wrong, Someone Must Have Made a Mistake

Some incorrect conclusions do result from mistakes, questionable practices, or misconduct. Others arise despite reasonable methods because empirical inference is made from incomplete and variable observations. Those situations should not be conflated.

Misconception

A Larger Sample Guarantees the Correct Answer

Larger appropriate samples can improve precision and reduce some forms of sampling uncertainty, but they do not automatically repair biased sampling, invalid measurement, confounding, inappropriate analysis, or other systematic problems.

Misconception

A Failed Replication Shows That the Original Researchers Were Wrong

A different result warrants investigation rather than an automatic verdict. The National Academies emphasizes that a single failed replication does not conclusively refute an original claim. Differences can arise for scientifically informative as well as problematic reasons.

Misconception

If Science Can Be Wrong, Research Cannot Be Trusted

The possibility of error is one reason research relies on transparency, criticism, replication, synthesis, and continuing investigation. Scientific reliability does not depend on every individual study being correct; it depends partly on having mechanisms through which claims can be tested and revised.

06 · What This Means for You

Design Against Error, Then Remain Proportionate About What Your Study Establishes

You cannot design a study that guarantees the truth of its conclusion. You can design one that makes important errors less likely, makes remaining uncertainty visible, and allows other researchers to evaluate how the conclusion was reached.

That changes the goal from “make sure I cannot possibly be wrong” to “use the strongest defensible procedures for this question and make claims proportionate to the resulting evidence.”

A simple decision framework

If an important source of bias can be anticipated
Address it through appropriate design, measurement, analysis, or transparency rather than assuming it will be negligible.
If your estimate is uncertain
Represent that uncertainty rather than treating the observed value as exact.
If your conclusion depends on important assumptions
State and examine those assumptions where possible.
If your evidence applies to a limited population or context
Keep the conclusion within those boundaries unless broader generalization is independently justified.
If later evidence conflicts with your finding
Examine the disagreement as new evidence rather than defending the original result merely because it was published first.
07 · A Quick Checklist

Before Assuming Your Conclusion Must Be Correct, Check:

Before making the final claim, check:
Could sampling variation materially affect the observed result?
Are the measurements sufficiently valid and reliable for the interpretation I am making?
Have important known sources of bias been addressed?
What assumptions does the analysis or interpretation depend on?
Could plausible alternative explanations account for the finding?
Does my conclusion extend beyond the population, setting, measurement, or period actually studied?
Have I distinguished an observed finding from the broader interpretation attached to it?
Does the wording of the conclusion reflect the uncertainty that remains?
How does the result compare with the wider body of relevant evidence?
08 · Frequently Asked Questions

Frequently Asked Questions About Research Being Wrong

Can a perfectly designed study still be wrong?

No empirical study is literally perfect, but even a study conducted according to strong methodological standards can obtain a result that later evidence does not support. Random variation, limited observations, imperfect measurement, and unrecognized features of the phenomenon can remain.

Does an incorrect conclusion mean the data were wrong?

No. The recorded data may be accurate while the inference drawn from them is incorrect or too broad. Researchers must distinguish the observations themselves from the conclusions they believe those observations support.

Can random chance make a study appear to find something that is not there?

Yes. Random variation can occasionally produce patterns that do not adequately represent the underlying phenomenon. Statistical methods can help characterize this uncertainty, but no procedure makes misleading samples impossible.

Does a statistically significant result mean the conclusion is probably correct?

Not by itself. Statistical significance addresses a limited statistical question under specified assumptions. The credibility of the substantive conclusion also depends on design, measurement, bias, analytical choices, effect estimates, uncertainty, and prior or subsequent evidence.

Does failure to replicate mean the first study was wrong?

Not necessarily. A single unsuccessful replication does not conclusively refute an earlier result. Researchers need to examine differences in samples, methods, measurements, conditions, uncertainty, and possible errors in both studies before explaining the discrepancy.

How can researchers reduce the chance of reaching the wrong conclusion?

Appropriate strategies depend on the methodology but can include stronger study designs, adequate samples, appropriate controls, careful measurement, transparent analysis, attention to bias, robustness checks, replication, and interpretation alongside the broader evidence.

Why should I trust research if studies can be wrong?

Scientific confidence should not rest solely on the assumption that every individual study is correct. It develops through critical scrutiny, methodological rigor, independent investigation, replication where appropriate, research synthesis, and the accumulation of multiple lines of evidence.

09 · The Bottom Line

Good Research Makes Error Less Likely, Not Impossible

The Bottom Line

A carefully designed and competently conducted study can still reach an incorrect conclusion because empirical research operates under uncertainty, and rigorous methods reduce rather than eliminate every possible source of error.

The appropriate response is neither to expect infallibility nor to treat all research as unreliable. Evaluate how well a study managed avoidable sources of error, how much uncertainty remains, and how its findings compare with evidence accumulated through further investigation.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes