03 · What You Need to Know
How analytical decisions can change the story told by the same evidence
Data do not analyze themselves
Once researchers have collected data, numerous decisions may remain. Which participants belong in the analysis? How should variables be coded? Which covariates should be included? What statistical model is appropriate? How should missing observations be handled? Should a continuous variable remain continuous or be categorized? Which time point should represent the outcome?
Some decisions are largely determined by the research design and question. Others involve legitimate methodological judgment. Still others may be inappropriate.
The resulting estimate is therefore not produced by the dataset alone. It arises from the combination of data, research question, design, assumptions, and analytical choices.
Different models can estimate different relationships
A statistical model formalizes assumptions about how variables relate to one another. Changing those assumptions can change the resulting estimate.
For example, one analysis might model a continuous outcome using linear regression, while another transforms the outcome because its distribution or substantive interpretation warrants doing so. Researchers might model time as linear in one analysis but allow a nonlinear pattern in another. A longitudinal dataset might be analyzed using methods that make different assumptions about repeated observations and correlation within participants.
Different models are not automatically competing attempts to estimate exactly the same quantity. Before comparing their results, determine what each model estimates and whether its assumptions are appropriate for the question and data.
Covariate adjustment can change an estimate substantially
In observational research, researchers frequently adjust for other variables to address confounding or improve estimation. The choice of variables matters.
Suppose an unadjusted analysis finds that participation in an optional academic support program is strongly associated with higher achievement. After adjusting for prior academic performance and other relevant pre-existing characteristics, the association becomes considerably smaller.
The two estimates need not represent contradictory evidence. The unadjusted estimate describes the observed difference between participants and nonparticipants. The adjusted estimate attempts to estimate a relationship conditional on the included covariates, under the assumptions of the model.
Unadjusted estimate
Describes the observed relationship or difference without conditioning on the covariates subsequently included in an adjusted model.
Adjusted estimate
Estimates a relationship after conditioning on specified variables, with its interpretation depending on why those variables were selected and whether the analytical assumptions are appropriate.
Adjustment should not be judged simply by counting covariates. Including inappropriate variables can also introduce bias or change the estimand. Covariate selection should be guided by the research question, design, substantive knowledge, and an appropriate causal or statistical rationale.
More adjustment is not automatically better adjustment
A common assumption is that a model becomes more credible whenever researchers add another control variable. That is not generally true.
Adjusting for a genuine pre-exposure confounder may reduce bias. Adjusting for a variable that lies on the causal pathway between an exposure and outcome can instead remove part of the effect the researcher intends to estimate. Conditioning on certain variables can also introduce bias through other causal structures.
Consequently, two studies can reach different conclusions because they condition on different sets of variables, but deciding which analysis is more appropriate requires understanding why those variables were included.
The question is not "Which study controlled for more things?" It is "Which adjustment strategy corresponds to the effect being estimated and the assumptions needed to identify it?"
Participant exclusions can change the analyzed population
Analytical samples are often smaller than the original study samples. Researchers may exclude participants because of missing information, protocol deviations, ineligibility discovered after enrollment, implausible measurements, inadequate exposure to an intervention, or other reasons.
Some exclusions are justified and prespecified. Others can create selection problems, particularly when inclusion in the final analysis is related to intervention, exposure, prognosis, or outcome.
Two studies with apparently similar recruited populations can therefore analyze meaningfully different groups.
When results conflict, compare both the original sample and the final analytical sample. Ask who disappeared between recruitment and analysis and why.
How missing data are handled can affect the result
Missing data are not merely an inconvenience that software needs to tidy up. The implications depend on why observations are missing and how the analysis handles that missingness.
A complete-case analysis includes only observations with the information required for the analysis. Other approaches may use multiple imputation, likelihood-based methods, weighting, or other techniques under particular assumptions.
No method can recover information without assumptions. If participants with missing outcomes differ systematically from those observed, the way missingness is handled may affect both the estimate and its uncertainty.
Cochrane treats missing outcome data as a distinct potential source of bias in randomized trials and emphasizes that the consequences depend on the amount and causes of missingness as well as its relationship to the true outcome.
Changing the outcome definition can change the apparent conclusion
Analytical decisions sometimes transform a measured outcome before analysis.
Suppose researchers collect a continuous symptom score. One study analyzes the mean difference between groups. Another converts the scale into a binary outcome using a threshold for "response." A third examines change from baseline.
These analyses can answer related but non-identical questions. A modest shift in the full score distribution may produce a noticeable difference in the percentage crossing a particular threshold, or the reverse.
If studies use substantially different measures in the first place, the problem extends beyond analysis. In that situation, examine whether different measurement choices explain the conflicting conclusions.
Categorizing continuous variables can discard information
Researchers sometimes convert continuous variables such as age, income, test scores, biomarker values, or exposure duration into categories. This may simplify presentation or correspond to meaningful clinical or policy thresholds, but arbitrary categorization can discard information and make results depend on chosen cut points.
Two analyses using different thresholds can therefore classify the same observations differently and potentially reach different conclusions.
This does not mean continuous variables should never be categorized. It means that cut points require substantive justification, particularly when conclusions change depending on where the boundary is placed.
Different effect measures can make the same evidence look different
The same underlying data can sometimes be summarized using relative risks, odds ratios, risk differences, hazard ratios, mean differences, standardized mean differences, or other effect measures.
These quantities are not interchangeable. They answer different statistical questions and can create different impressions of practical magnitude.
For example, a relative reduction can appear substantial while the corresponding absolute reduction is small when baseline risk is low. Neither measure is necessarily wrong. Each expresses a different aspect of the effect.
When studies appear to disagree, first determine whether the reported effect measures can be meaningfully compared.
Statistical significance can change even when the substantive estimate barely changes
One analysis produces a p-value of 0.04. A slightly different but defensible model produces 0.06. It would be misleading to describe the first as demonstrating an effect and the second as demonstrating no effect.
The estimates may be nearly identical. The difference may simply reflect small changes in standard errors, model assumptions, sample composition, or degrees of freedom.
The American Statistical Association has emphasized that scientific conclusions should not be based only on whether a p-value crosses a particular threshold and that statistical significance does not measure the size or importance of an effect.
Watch Out
An analysis changing from "statistically significant" to "not statistically significant" does not necessarily mean the substantive finding reversed. Inspect the effect estimate, its uncertainty, and how much those quantities changed before calling the results contradictory.
Researchers often have more than one reasonable analytical path
Many datasets permit multiple plausible analyses. Researchers may reasonably disagree about model form, covariate selection, missing-data assumptions, transformations, inclusion criteria, or operational definitions.
Steegen and colleagues use the term multiverse analysis for an approach that examines how conclusions vary across alternative datasets created by different reasonable data-processing choices. Simonsohn, Simmons, and Nelson similarly proposed specification curve analysis for examining results across a set of reasonable statistical specifications.
The broader insight is important even if you never conduct either formal procedure: a single reported specification may not reveal how dependent a finding is on analytical choices.
Robustness asks whether the conclusion survives reasonable alternatives
A result is analytically more reassuring when its substantive interpretation remains reasonably stable across plausible analytical choices.
That does not require every model to produce the same coefficient or p-value. Different specifications can legitimately estimate somewhat different quantities. The important question is whether reasonable choices repeatedly support a broadly similar substantive conclusion or whether the result changes dramatically depending on how the analysis is configured.
Sensitivity analyses are one way to examine this issue. Depending on the study, researchers might vary assumptions about missing data, model specifications, participant exclusions, outcome definitions, or other analytical choices.
A sensitivity analysis is useful only when the alternative assumptions or analyses are themselves meaningful. Running dozens of arbitrary models does not automatically make a finding robust.
Analytical flexibility can become a problem when results guide the choices
Having multiple defensible analytical options is not itself misconduct. The problem becomes more serious when researchers repeatedly try alternatives and selectively report the analysis that produces the most favorable or statistically significant result.
Simmons, Nelson, and Simonsohn demonstrated how undisclosed flexibility in data collection and analysis can substantially increase the probability of obtaining statistically significant findings even when the underlying hypothesis is false.
This is one reason transparency about analytical decisions matters. Readers need to know which analyses were planned, which were exploratory, which alternatives were considered, and whether conclusions depend heavily on one particular specification.
Preregistration can help distinguish planned from data-responsive analyses
Preregistration records aspects of the research plan before outcomes are known or analyzed, depending on the form and timing of registration. It can help readers distinguish confirmatory analyses specified in advance from exploratory analyses developed after examining the data.
Preregistration does not make an analysis correct. Researchers can preregister inappropriate methods, encounter unforeseen data problems, or have legitimate reasons to deviate from a plan.
The value lies partly in transparency. When deviations occur, they can be reported and justified rather than disappearing into the final analysis as though they had always been intended.
Exploratory analysis is not inferior, but it answers a different evidential question
Researchers often discover interesting patterns that were not anticipated. Exploratory analysis is essential for developing hypotheses and understanding data.
The problem arises when exploratory findings are presented with the same evidential status as hypotheses and analyses specified independently of the observed results.
If one study reports a prespecified primary analysis and another emphasizes a subgroup discovered after extensive exploration, the conclusions should not automatically receive equal interpretive weight simply because both appear in peer-reviewed papers.
Different software is rarely the important explanation by itself
Researchers may analyze data using R, Stata, SPSS, SAS, Python, or other software. Properly implemented equivalent methods should generally not produce substantively different answers merely because the software name changed.
Differences can arise from default settings, estimation algorithms, treatment of missing observations, reference categories, numerical procedures, or implementation choices. But the relevant issue is the statistical specification and its implementation, not the brand of software.
"They used different software" is therefore usually an incomplete explanation. Ask what the software was instructed to do differently.
Analytical differences can coexist with design differences
Analysis cannot be interpreted independently of study design. Randomized trials, observational studies, clustered designs, longitudinal data, case-control studies, and other designs create different analytical requirements.
An adjustment strategy appropriate for one design may be unnecessary or inappropriate for another. An intention-to-treat analysis answers a different question from some per-protocol analyses. Observational causal estimates require assumptions that randomized comparisons may address differently.
If studies use fundamentally different designs, first examine whether research design itself could explain the conflicting findings.
Analytical disagreement can sometimes be scientifically informative
If an effect appears only under one narrow specification and disappears across many other reasonable analyses, that fragility is information. It tells you that the conclusion depends strongly on analytical assumptions.
Conversely, if different reasonable methods repeatedly produce estimates of similar magnitude and direction, confidence in the substantive pattern may increase even if individual p-values vary.
The goal is therefore not to find the one analysis that makes disagreement disappear. It is to understand how much of the conclusion is supported by the data and how much depends on particular analytical choices.
06 · What This Means for You
How to determine whether analysis explains conflicting findings
When studies appear similar but their results differ, compare the analytical pipeline rather than stopping at the methods section's statistical test name.
A simple decision framework
If the studies use different covariate sets
Determine why each variable was included and whether the models estimate the same substantive effect.
If the analytical samples differ because of exclusions or missing data
Compare who was included, who was excluded, why observations were missing, and whether those decisions could bias or alter the estimate.
If different statistical models produce different estimates
Examine their assumptions, estimands, fit to the research design, and substantive justification rather than choosing the model with the preferred result.
If the main difference is significant versus nonsignificant
Compare effect estimates and uncertainty before deciding that the conclusions actually conflict.
If reasonable alternative analyses repeatedly support a similar conclusion
The finding has stronger evidence of analytical robustness, although other sources of bias may still remain.
If the conclusion changes substantially across reasonable specifications
Report the finding as analytically sensitive and avoid presenting one specification as uniquely definitive without a strong justification.
Reconstruct the analytical pipeline
When enough information is available, record the primary outcome definition, analytical sample, exclusion criteria, covariates, missing-data method, statistical model, effect measure, transformations, interaction terms, clustering or repeated-measures approach, and sensitivity analyses for each study.
This often reveals differences that disappear in broad descriptions such as "multiple regression was used."
Two papers can both report regression analyses while conditioning on different variables, analyzing different subsets of participants, and estimating different quantities.
Look for prespecified and alternative analyses
Check protocols, registrations, statistical analysis plans, supplementary materials, and methods sections where available. Determine whether the reported analysis was specified before researchers examined the relevant results or whether it emerged during exploration.
An exploratory analysis can still be valuable. The distinction matters because a result selected after examining many alternatives warrants different evidential interpretation from a prespecified test that was not chosen because of its outcome.
Use robustness as evidence, not as a ritual
Ask whether the conclusion survives alternative analyses that are genuinely plausible for the research question. Useful sensitivity analyses target consequential assumptions or decisions rather than generating arbitrary model variations for decoration.
If conclusions remain similar across defensible specifications, that stability is informative. If they change substantially, the analytical sensitivity should become part of the reported conclusion.
Do not let analysis distract from deeper differences
Analytical choices are only one source of disagreement. If studies use different populations, instruments, interventions, or research designs, a statistical explanation may address only part of the discrepancy.
The broader task remains to understand why the studies reached different conclusions and determine which differences actually matter.
Give greater weight to findings whose analytical support you can inspect
Transparent reporting makes it easier to assess whether conclusions depend on particular choices. When authors clearly report exclusions, covariates, missing-data procedures, model specifications, deviations from plans, and sensitivity analyses, readers can evaluate the analytical path rather than accepting the final coefficient on trust.
Analytical transparency is not proof of correctness. But when deciding which evidence deserves greater weight, the ability to scrutinize how an estimate was produced is a meaningful consideration.