03 · What You Need to Know
How to Choose Among Several Defensible Analyses
First Check Whether the Analyses Really Answer the Same Question
Two methods can use the same dataset and variables without estimating the same thing.
Suppose a study measures an outcome before and after an intervention. Comparing post-intervention scores while adjusting for baseline, comparing raw change scores, and modeling repeated measurements may all appear to address “whether the groups changed differently.” Depending on the model and design, however, they can differ in what they estimate, what assumptions they make, and how efficiently they use the available information.
Before treating methods as alternatives, write down the quantity each one estimates. If the quantities differ, you may not actually be choosing among interchangeable analyses. You may be choosing among different versions of the research question.
This is why the first check remains whether the analysis matches the research question and design.
Define the Primary Analytical Target Before Comparing Methods
Ask what result would constitute the most direct answer to the research question. Is it a mean difference, ratio, probability, change, association, interaction, predicted outcome, or another quantity?
Once that target is explicit, some candidate analyses may become less attractive because they answer a neighboring question rather than the intended one.
In clinical trials, ICH E9(R1) formalizes this distinction by separating the estimand, which describes the treatment effect of interest, from the estimator used to estimate it. The framework is specific to clinical trials, but the reasoning travels well: define what you want to estimate before deciding among methods for estimating it.
Compare the Assumptions Each Analysis Requires
Two analyses may address essentially the same target while relying on different assumptions.
Those assumptions may concern distributions, functional form, missing data, independence, variance structure, censoring, confounding, measurement, or other features of the analytical model. The appropriate comparison is not simply which method has fewer assumptions. Every method makes assumptions, sometimes implicitly.
Ask instead which assumptions are most defensible for the design and scientific context, which assumptions matter most to the conclusion, and whether those assumptions can be examined through diagnostics or sensitivity analyses.
Consider the Data Structure the Design Created
An analysis that ignores how the observations were generated may be less appropriate even if it is easier to run.
Repeated measurements from the same participant are related. Students within classrooms may be correlated. Matched observations differ from independent samples. Complex survey designs may involve weights, strata, or clusters.
Candidate methods should therefore be compared according to how well they represent the actual structure of the data. This is not a matter of choosing the most sophisticated model. It is a matter of avoiding a simpler model that obtains its simplicity by pretending important design features do not exist.
Consider Precision and Statistical Efficiency, but Not in Isolation
When two methods target the same quantity under defensible assumptions, one may use the available information more efficiently and therefore provide more precise estimates.
Efficiency is a legitimate consideration, particularly during study planning. It should not, however, be confused with choosing whichever analysis happens to produce the narrowest confidence interval or smallest p-value in the observed dataset.
The methodological question is whether the procedure is expected to use the design and available information efficiently under reasonable assumptions. The result itself should not become the selection criterion.
Prefer Interpretability When Other Considerations Are Comparable
An analytically sophisticated method is not automatically preferable if its output is difficult to connect to the research question.
When two methods are similarly defensible, consider which produces an estimate that researchers, practitioners, or decision-makers can understand. An interpretable effect estimate with an appropriate measure of uncertainty may be more useful than a complicated parameter whose substantive meaning is obscure.
This does not justify using an inappropriate simple method. It means interpretability is one legitimate criterion among several when methodological fit is otherwise comparable.
Designate a Primary Analysis When the Study Requires One
When several analyses are possible, a confirmatory study often benefits from identifying one as primary before the relevant results are known.
The primary analysis is the analysis intended to carry the main inferential burden. Other analyses can then have clearly defined roles. They might examine robustness, investigate secondary outcomes, explore heterogeneity, or answer additional questions.
This hierarchy reduces the temptation to run several plausible analyses and retrospectively promote whichever one produces the most compelling result.
The primary analysis should be specified as part of the broader analysis plan developed before data collection when the study's purpose permits that level of advance specification.
Use Sensitivity Analysis to Ask Whether the Conclusion Depends on Important Assumptions
A sensitivity analysis is not merely “another way to analyze the data.” Its purpose is to examine how robust the primary conclusion is to meaningful changes in assumptions, methods, or analytical decisions.
CONSORT 2025 describes sensitivity analyses as additional analyses used to examine robustness when assumptions about the data, methods, or models differ from those used in the primary analysis. A useful sensitivity analysis should therefore have a reason for existing.
One principled approach asks whether the proposed sensitivity analysis addresses the same question as the primary analysis, whether it could plausibly produce a different result, and whether such a difference would create genuine uncertainty about which conclusion to believe.
Primary analysis
The prespecified analysis intended to provide the principal answer to the research question.
Sensitivity analysis
An alternative analysis used to examine whether the main conclusion depends materially on important assumptions or analytical choices.
Do Not Call Every Alternative a Sensitivity Analysis
Running several models with slightly different variable combinations does not automatically make them sensitivity analyses.
If an alternative model answers a different question, it may be a secondary analysis. If it investigates an unplanned relationship suggested by the observed data, it may be exploratory. If it examines whether the same substantive conclusion survives a defensible change in an important assumption, sensitivity analysis may be the appropriate description.
Clear labels help readers understand why multiple analyses were conducted and what differences among them mean.
What If Different Reasonable Analyses Produce Different Conclusions?
This is where multiple analyses become scientifically informative.
Suppose one defensible analysis suggests an important association while another, based on a similarly plausible assumption, produces a much smaller and highly uncertain estimate. The correct response is not automatically to decide which analysis is “right” and hide the other.
Investigate why the results differ. Is the difference driven by missing-data assumptions? Treatment of an influential observation? Functional form? Adjustment choices? Outcome definition? Different analytical populations?
If the substantive conclusion depends heavily on a reasonable analytical choice, that dependence is itself part of the result. The evidence may simply be less robust than one analysis alone would suggest.
Watch Out
If several defensible analyses are tried after the results are visible, selecting only the analysis that gives the strongest support for the preferred conclusion conceals analytical uncertainty. A favorable result is not a methodological criterion.
Multiple Analyses Can Create Multiplicity Concerns
When researchers repeatedly test outcomes, contrasts, subgroups, models, or hypotheses, conventional significance testing can generate opportunities for chance findings. The implications depend on why the analyses are being conducted and how the resulting claims are interpreted.
NIH methodological guidance notes that studies with multiple treatment arms, outcomes, or interim analyses may need to address multiple comparisons and recommends detailing the adjustment procedure where relevant.
There is no single multiplicity rule appropriate to every study. Confirmatory analyses, exploratory analyses, sensitivity analyses, and descriptive estimates serve different purposes. What matters is that researchers do not treat a large collection of analyses as though each were the sole planned test.
Sometimes the Best Choice Is to Seek Methodological Input
When candidate analyses differ in their estimands, assumptions, treatment of clustering or missingness, or interpretation, the choice can require expertise beyond an introductory statistical decision tree.
Consultation is particularly useful before the results are examined. A statistician or methodologist involved during study design can help identify the primary analytical target and compare methods without being influenced by which one ultimately produces the preferred result.