01 · The Question
What Result Would Make You Say, “The Study Worked”?
Imagine that you are studying whether AI-generated feedback improves students' academic writing. Before collecting data, you already expect the AI group to perform better. The hypothesis predicts it, the theoretical rationale supports it, and perhaps the practical appeal of the study depends partly on demonstrating that benefit.
Now suppose the difference is negligible. Or the comparison group performs better. Would you regard those findings as legitimate answers, or would the study suddenly feel unsuccessful?
That reaction is worth examining before the study begins. Researchers can have well-founded expectations, and confirmatory research necessarily tests predictions. The problem arises when one preferred outcome becomes the only result treated as scientifically successful, while credible alternatives are implicitly regarded as failures to be explained away, ignored, or replaced with more favorable analyses.
03 · What You Need to Know
Separate a Successful Study From a Successful Prediction
Researchers do not approach every study without expectations. Theory, previous evidence, pilot work, and substantive reasoning often support directional hypotheses. Confirmatory research can be especially valuable when a prediction is specified clearly enough to face a genuine empirical test.
Yet two kinds of success should be distinguished. A hypothesis succeeds as a prediction when the evidence supports what was predicted. A study succeeds scientifically when its design and evidence allow the research question to be addressed credibly, whether or not the prediction survives.
A Supported Hypothesis Is One Possible Outcome, Not the Definition of Research Success
Suppose a researcher predicts that students receiving AI-generated formative feedback will outperform students receiving conventional instructor feedback.
If the prediction is supported, that matters. If the evidence instead suggests little meaningful difference, that also matters if the study was capable of evaluating such a difference. If instructor feedback performs better, the theory or assumptions motivating the original prediction may require reconsideration.
All three possibilities can contribute knowledge. What changes is the answer, not whether the study deserves to exist.
Watch for Questions That Contain the Preferred Result
Some wording makes one outcome appear normal before evidence has been collected:
- How does generative AI improve students' academic writing?
- Why does gamification increase student motivation?
- How does flexible work enhance employee productivity?
- Why does social media reduce adolescents' attention?
These questions do more than express an expectation. They treat improvement, increase, enhancement, or reduction as the starting premise.
If those propositions have not already been established appropriately, the question may contain an assumption that should itself remain open to investigation.
Ask What Would Count as a Legitimate Answer Before Seeing the Data
Before data collection, write down several possible findings and the conclusions each would permit. For a comparative study, these might include a meaningful advantage for the focal condition, a meaningful disadvantage, evidence that the difference is practically negligible, heterogeneous effects, or evidence too imprecise to distinguish among important possibilities.
The last possibility is different from the others. An inconclusive study does not become informative merely because every outcome must be called a success. If measurement is inadequate or estimates are too uncertain, the study may genuinely fail to resolve the question.
The point is narrower: the direction of a credible finding should not determine whether the finding is accepted as an answer.
Statistical Significance Can Quietly Become the Definition of Success
In quantitative research, the problem often appears as a binary expectation: p <.05 means success; p >.05 means failure.
That framing is too crude. Statistical significance does not establish theoretical importance, practical importance, causality, measurement quality, or replication. Conversely, conventional nonsignificance does not establish that an effect is absent.
The estimated magnitude, uncertainty, design, assumptions, and substantive threshold for what matters should influence interpretation. If evidence of a negligible effect would itself be important, the study should be designed to evaluate that proposition rather than simply hoping not to reject a null hypothesis.
One-Sided Expectations Require Particular Care
A strong theoretical prediction may concern one direction. That does not make directional hypotheses inappropriate. It does mean researchers should decide in advance how evidence in the opposite direction would be handled.
An unexpected direction may indicate measurement problems, implementation failure, model misspecification, an incorrect theory, an overlooked mechanism, or a genuine phenomenon. Those possibilities need investigation. The opposite result should not automatically be discarded simply because it was not the one predicted.
Outcome-Dependent Success Can Encourage Selective Analysis
If only one result feels publishable or worthwhile, researchers may face incentives to search for analyses that produce it. Choices involving exclusions, transformations, covariates, outcomes, subgroups, statistical models, and stopping rules can sometimes alter results.
This does not mean every analytical change is improper. Data problems and unexpected methodological issues genuinely arise. The concern is allowing the desired conclusion to determine which defensible-looking analysis becomes the one reported as though it were inevitable.
Prespecification and preregistration can help distinguish analyses planned before outcomes were known from later exploratory decisions. Preregistration does not eliminate researcher judgment, but it can make the chronology of those decisions more transparent.
Unexpected Findings Are Not Automatically Confirmatory Findings
Researchers often discover interesting patterns that were not predicted. Exploration is a legitimate and productive part of science. The problem arises when an unexpected finding is retrospectively presented as though it had been predicted from the beginning.
Kerr termed this practice HARKing, or “Hypothesizing After the Results are Known.” The central issue is not generating hypotheses from results. Researchers do that all the time. The problem is obscuring that chronology by presenting a post hoc hypothesis as an a priori prediction.
Unexpected results can therefore generate valuable hypotheses for subsequent research without being rewritten into a story in which the researcher somehow knew the answer all along.
A Study Can Fail Methodologically Even When It Gets the Expected Result
Obtaining the preferred result does not rescue a weak study.
A statistically significant difference measured with an inappropriate instrument, produced by a severely biased comparison, or interpreted causally from evidence incapable of supporting causation remains problematic. Likewise, an expected result produced only after extensive undisclosed analytical searching should not be treated as stronger simply because it matches the theory.
This is why the appropriate criterion for success is evidentiary: did the study generate credible evidence relevant to the research question?
A Study Can Succeed Scientifically While Refuting Its Hypothesis
Suppose the theory clearly predicts that an intervention will increase an outcome. A rigorous study produces precise evidence inconsistent with that predicted increase.
The hypothesis was not supported. Yet the study may have done exactly what a good empirical test should do: expose a prediction to evidence capable of contradicting it.
This distinction becomes especially important when evaluating whether the question remains useful when the expected relationship is absent.
Define Success Before the Results Are Known
A useful exercise is to write two definitions before data collection:
Prediction success
The observed evidence supports the prespecified hypothesis or expected direction under the planned inferential criteria.
Study success
The study produces evidence of sufficient quality and informativeness to address the research question, including when that evidence contradicts the prediction.
Keeping those definitions separate makes it harder for a disappointing hypothesis test to be mistaken for a failed research project.
07 · A Quick Checklist
Check Whether You Have Defined Only One Result as Success
Before collecting or examining outcome data, check:
Write down the result your hypothesis predicts without treating that prediction as established fact.
Describe what a credible result in the opposite direction would contribute to the research question.
Determine whether evidence of a negligible or absent relationship would matter and what evidence would justify that conclusion.
Distinguish an unsupported hypothesis from an inconclusive study.
Avoid defining statistical significance alone as the criterion for a successful study.
Prespecify important hypotheses and analytical decisions when appropriate to the research design.
Label analyses or hypotheses developed after seeing the results as exploratory or post hoc rather than presenting them as originally predicted.
Define study success in terms of credible evidence about the question rather than confirmation of the preferred answer.