03 · What You Need to Know
Studying Similarity Requires Different Thinking From Failing to Find a Difference
First, distinguish your expectation from the evidence you want to obtain
You may expect two groups to perform similarly for several reasons. Previous studies may repeatedly show small differences. A theory may predict that a particular factor should not matter under the conditions you are studying. Preliminary data may suggest little separation between groups. Or you may simply suspect that a widely assumed difference has been exaggerated.
Those expectations can motivate a study. They do not determine its conclusion.
The research still needs to produce evidence capable of changing your confidence in the claim. If a meaningful difference appears, your interpretation should be capable of changing accordingly.
Expecting no meaningful difference
A prediction made before examining the study's results, based on theory, prior evidence, reasoning, or other relevant information.
Evidence of no meaningful difference
Data sufficiently informative to rule out differences large enough to matter according to a defensible criterion.
The first is a hypothesis or expectation. The second is an inferential conclusion. Research is needed precisely because the two are not interchangeable.
“No statistically significant difference” does not mean “the groups are the same”
This is the most important statistical distinction in this guide.
In conventional null-hypothesis significance testing, researchers often begin with a null hypothesis specifying no difference and ask whether the data provide sufficient evidence against it. If the resulting test is not statistically significant, the researcher has failed to reject that null hypothesis.
That does not establish that the null hypothesis is true.
A non-significant result can occur because the true difference is small. It can also occur because the estimate is imprecise, the sample is too small, measurements are noisy, or the study otherwise lacks the ability to distinguish plausible effects from zero.
Methodological work on null results repeatedly emphasizes this distinction: absence of evidence for a difference is not automatically evidence that an important difference is absent.
Watch Out
Do not write “there was no difference between the groups” merely because a conventional test produced p >.05. Inspect the estimated difference and its uncertainty, and use an inferential approach appropriate to the question you actually want to answer.
A study can be inconclusive in both directions
Suppose two instructional approaches produce nearly identical sample means and the comparison is not statistically significant. It is tempting to conclude that they perform equally well.
Now suppose the confidence interval around the estimated difference is very wide. The data may be compatible with one approach being meaningfully better, the other approach being meaningfully better, or the two being practically similar.
The study has not established equivalence. It has produced an imprecise answer.
This is why recent work examining replication of null findings distinguishes between studies that provide evidence for a negligible effect and studies that are simply inconclusive. A sufficiently underpowered study can easily produce a non-significant result even when a meaningful effect exists.
The distinction matters because a weak study should not be rewarded for failing to detect anything.
Ask what size of difference would actually matter
If your substantive question is whether two conditions are meaningfully similar, you need to think beyond whether the true difference is exactly zero.
Exact equality is rarely the practical question.
Suppose two teaching approaches differ by 0.1 percentage points in average examination performance. Technically, that is a difference. But if differences smaller than two percentage points would have no meaningful educational consequence, the more useful question may be whether the true difference is sufficiently small to fall within that range.
This introduces the idea of a smallest effect size of interest or another substantively justified threshold defining what would count as a meaningful difference.
The threshold should come from the research context rather than being chosen after seeing the results simply because it produces a convenient conclusion.
Equivalence testing can address whether meaningful differences can be ruled out
When the research question concerns whether an effect is small enough to be considered practically negligible, equivalence testing is one available frequentist approach.
The researcher specifies an equivalence range before interpreting the data. Effects within that range are regarded as too small to be substantively meaningful for the purpose of the study. The analysis then asks whether effects at least as large as those boundaries can be rejected.
One common implementation is the two one-sided tests procedure, often abbreviated TOST. At a 5% significance level, equivalence can also be assessed by determining whether the corresponding 90% confidence interval lies entirely inside the prespecified equivalence bounds.
The interpretation is not that the true effect is proven to be exactly zero. Rather, the evidence supports ruling out effects at least as large as the threshold you defined as meaningful.
The equivalence margin is a substantive decision, not statistical decoration
The hardest part of equivalence testing is often not calculating the test. It is deciding what difference is small enough not to matter.
That judgment might be informed by previous research, theory, clinical or educational relevance, costs, measurement properties, policy thresholds, stakeholder judgments, or other domain-specific considerations.
For example, a one-point difference on a 100-point examination might be negligible for one research question but consequential if a one-point threshold determines whether students qualify for something important.
There is no universal “small effect” that can be declared irrelevant independently of context.
This is one reason equivalence margins should ideally be justified before the results are known. Otherwise, researchers could widen the acceptable range until the data appear equivalent, which would reverse the logic of the test.
Bayesian approaches can address a related but different question
Equivalence testing is not the only way to examine evidence concerning small or absent effects.
Bayesian methods can compare how well observed data are predicted under competing hypotheses or estimate the distribution of plausible effect sizes given a model and prior assumptions. Bayes factors, for example, can be used to quantify relative evidence for a null hypothesis versus a specified alternative.
These methods answer somewhat different inferential questions and depend on choices such as the prior distribution under the alternative hypothesis. They should not be treated as interchangeable buttons for making a null result sound stronger.
The broader principle is more important than allegiance to a particular framework: if absence or practical equivalence is the claim you care about, choose an analysis capable of providing evidence about that claim.
Finding no meaningful difference can change a practical decision
Consider two interventions that appear to produce comparable outcomes, but one is substantially cheaper, safer, faster, easier to administer, or less burdensome.
If decision-makers need to know whether the simpler alternative sacrifices an important amount of effectiveness, establishing that differences are sufficiently small can have considerable practical value.
This logic underlies equivalence and non-inferiority research in clinical trials, where investigators may wish to determine whether a new treatment is sufficiently similar to, or not unacceptably worse than, an established treatment according to a prespecified margin.
The same general reasoning can arise in education, organizational research, technology, and other fields, although the appropriate design and inferential standards depend on the discipline and question.
A study does not need to identify a winner to be useful.
No meaningful difference can challenge an assumption
Sometimes the important contribution is not choosing between two alternatives but questioning a distinction researchers or practitioners have assumed to be consequential.
Imagine that a theory predicts substantially different responses between two groups, yet a sufficiently precise study rules out differences large enough to support that prediction.
That result may weaken the theory, suggest that the proposed mechanism is less important than assumed, or redirect attention toward other explanations.
Similarly, institutions may invest resources in customizing programs for two populations because they assume those populations respond differently. Evidence that any difference is too small to matter for the relevant outcome could prompt reconsideration of that assumption.
The absence of an important difference can therefore itself be a theoretical or practical contribution.
No difference can prevent unnecessary complexity
Research often asks whether increasingly elaborate approaches outperform simpler ones.
Suppose a complex analytical model, instructional intervention, diagnostic procedure, or technological system requires substantially more resources than a simpler alternative. If the added complexity produces no meaningful improvement, that knowledge can matter.
The contribution is not “nothing happened.” It is evidence relevant to whether the extra complexity is justified.
This connects to the broader question of whether research can be valuable without solving a problem. Sometimes the useful outcome of research is ruling out a presumed advantage rather than creating a new solution.
Expected similarity can provide a demanding test of theory
Not every theory predicts differences everywhere. A theory may explicitly predict that a variable matters under one condition but becomes irrelevant under another.
Testing the condition under which an effect should disappear can therefore provide information about the theory's boundary conditions.
For example, suppose a theory predicts that prior experience should affect performance when users receive minimal guidance but should cease to matter when comprehensive guidance is provided. Demonstrating that the experience gap becomes negligibly small under comprehensive guidance could be theoretically informative.
The expected absence of a difference is part of the prediction rather than a disappointing failure to obtain an effect.
Expecting no difference is different from hoping for no difference
Researchers should also examine why they expect similarity.
If a new intervention is cheaper, easier, or personally favored, you may want it to perform just as well as the alternative. That preference can influence how results are interpreted.
The remedy is not to avoid the question. It is to make the criterion for success explicit before examining the results and design the study so that the preferred conclusion is genuinely capable of being challenged.
A wide, underpowered study that predictably produces p >.05 is not persuasive evidence of similarity.
Research designed to demonstrate equivalence should make equivalence difficult enough to establish that obtaining it is informative.
Sometimes another study of “no difference” adds almost nothing
Expected similarity does not automatically make a question worthwhile.
Suppose multiple rigorous studies have already shown that any difference between two approaches is smaller than a well-justified threshold across the populations and conditions that matter. Conducting another nearly identical comparison may have little marginal value.
As with questions whose expected result seems obvious, you should ask what uncertainty remains.
Perhaps evidence is lacking for an important population. Perhaps previous estimates are imprecise. Perhaps the equivalence margin used previously was inappropriate for your decision context. Perhaps new technology has changed the conditions sufficiently to reopen the question.
If none of these applies, the expectation of no difference may simply reflect that the question has already been answered adequately.
Null results should not disappear simply because they are less exciting
Research literatures can become distorted when studies with statistically significant or favorable results are more likely to be disseminated than studies producing null or negative findings.
A null finding that was generated by a rigorous and informative design can contribute to cumulative knowledge. It can constrain meta-analytic estimates, challenge theoretical expectations, prevent unnecessary duplication, or show that an intervention did not produce the anticipated benefit under the studied conditions.
This does not mean every non-significant analysis deserves publication. A result that is too imprecise to distinguish meaningful effects from negligible ones may simply be inconclusive.
The value lies in the information the study provides, not in whether the p-value falls on one side or the other of an arbitrary threshold.
“No difference” and “nothing happens” are related but not identical questions
A comparison between two groups or conditions asks whether they differ on an outcome. Another kind of research question asks whether an event, exposure, intervention, or change produces any meaningful response at all.
The statistical and conceptual issues overlap, but the scientific questions can differ.
For example, “Do two feedback methods produce meaningfully different learning outcomes?” is comparative. “Does introducing automated feedback meaningfully change students' revision behavior?” asks whether a change occurs relative to an appropriate reference or expectation.
The next guide examines more directly whether a question remains worth studying when the answer might be “nothing happens”.
The question is worthwhile when establishing similarity changes what we know or do
Before committing to a study in which you expect no difference, imagine two possible conclusions.
First, suppose the data are sufficiently precise to rule out differences large enough to matter. What changes? Does a theory become more constrained? Can a simpler alternative be considered? Does an assumed distinction become less defensible? Can resources be redirected? Does a future research program avoid pursuing an effect unlikely to be consequential?
Second, suppose a meaningful difference appears despite your expectation. Would that also matter?
If both possibilities provide useful information, you probably have a genuine empirical question.
If neither outcome changes much, expecting no difference may not be the main problem. The underlying question may simply not be important enough.