01 · The Question
How Can You Choose a Statistical Test Before You Have Seen the Data?
Researchers are often told to decide their statistical analysis before collecting data. Then they encounter another familiar piece of advice: examine the data, check assumptions, and choose an appropriate statistical procedure based on what you find.
At first, those recommendations can appear contradictory. How can you choose between methods that depend partly on data characteristics such as distribution, variance, missingness, or model fit when those data do not yet exist?
The apparent contradiction comes from treating statistical planning as an all-or-nothing decision. You can usually determine the primary analytical strategy, and often the primary statistical method, before collection while explicitly planning how you will respond to data characteristics that cannot reasonably be known in advance.
03 · What You Need to Know
What You Can Decide About Statistical Testing Before the Data Exist
Start With the Question, Design, and Variables
A statistical test should not be chosen by looking at a list of procedures and selecting the one you recognize. The appropriate options are constrained by the research question, study design, analytical target, variables, and structure of the observations.
Are you describing one population or comparing groups? Are observations independent, paired, repeated, or clustered? What is the outcome and how is it measured? Is the question about a difference, association, prediction, change, interaction, or another quantity?
These characteristics exist before the final dataset does. They often narrow the reasonable analytical approaches considerably. Guidance on statistical planning similarly emphasizes the research question, study design, and measurement characteristics when identifying suitable statistical methods.
This is why the study design should not be built around a preferred statistical test. The question and design lead; the analysis follows from them.
You Usually Know More Before Data Collection Than You Think
Consider a planned randomized study with two conditions and an outcome measured before and after an intervention. Before collecting data, you already know that observations from the same participant are related. You know which condition represents the intervention, what the primary outcome is intended to be, when measurements occur, and what comparison the research question requires.
You may not yet know the exact distribution of the observed outcome or the extent of missing data. But those unknowns do not prevent you from defining the main analytical target and selecting an appropriate primary strategy.
The CDC's Field Epidemiology Manual explicitly recommends deciding what data will be collected and how they will be analyzed before the data-collection instrument is designed. Its analysis-plan guidance connects research questions and hypotheses to variables, data handling, measures of association, significance tests, confidence intervals, and other analytical decisions.
Choosing an Analytical Strategy Is Broader Than Naming a Test
“We will use a t-test” may sound specific, but it leaves many important questions unanswered. Which outcome will be tested? Which observations will be included? Are the observations independent or paired? What comparison does the test represent? What happens if important assumptions are unreasonable? What estimate and measure of uncertainty will be reported?
A useful plan therefore moves from the scientific question to the analytical target and then to the statistical method.
Analytical strategy
The broader plan connecting the question, design, variables, analytical target, data structure, assumptions, and interpretation.
Statistical test
A particular inferential procedure used to evaluate a statistical hypothesis or quantity within that strategy.
For many studies, it is more important to have the first clearly specified than to attach a familiar test name prematurely.
Confirmatory Research Usually Calls for Greater Advance Specification
The stronger the confirmatory purpose of the analysis, the stronger the case for specifying consequential analytical choices before the relevant results are known.
In clinical trials, for example, ICH E9 states that the principal features of the statistical analysis should be described in the statistical section of the protocol and that the principal statistical analysis should be specified before breaking the blind. These formal requirements belong to a particular regulatory setting, so they should not be applied mechanically to every research project.
The underlying reasoning is more general. If researchers can choose among many plausible analyses after learning which one produces the most favorable result, the observed result can influence the analytical decision that supposedly evaluates it.
Advance specification reduces that particular source of analytical flexibility. It does not guarantee that the chosen method is correct.
Assumptions Do Not Necessarily Prevent Advance Planning
A common objection is that you cannot select a test until you have checked its assumptions. Some assumptions do concern observed data characteristics, but others follow largely from the design or analytical model.
Whether observations are paired because the same participants are measured twice is known from the design. Whether students are clustered within classrooms is known from the sampling or intervention structure. Whether an outcome is binary or continuous is usually known from its measurement plan.
Other issues may genuinely require observed data. The appropriate response is not necessarily to leave the analysis completely unspecified. You can often identify the primary method and define what would constitute a consequential problem and what alternative strategy would then be considered.
Watch Out
Avoid creating a mechanical chain in which one preliminary assumption test automatically determines the main analysis without considering the robustness of the planned method, the study design, sample size, purpose of the analysis, and consequences of the departure. Assumption assessment requires statistical judgment, not merely another p-value.
Do Not Reduce the Choice to “Parametric or Nonparametric”
Introductory decision trees often suggest choosing a parametric procedure when assumptions are met and a nonparametric alternative when they are not. Such heuristics can be useful educational starting points, but real analytical decisions may be more nuanced.
Different procedures test or estimate different quantities. Some commonly used parametric methods are reasonably robust to particular departures under some conditions. Transformations, generalized models, robust methods, permutation approaches, or other techniques may sometimes be more appropriate than automatically replacing one familiar test with its presumed nonparametric counterpart.
The question should remain: which method appropriately answers the research question for this design and these data?
Plan the Primary Analysis and Reasonable Contingencies
Suppose your planned primary analysis relies on a model whose assumptions could be seriously compromised by particular data characteristics. Instead of writing “the statistical test will be chosen after examining the data,” the analysis plan can specify the intended primary method and identify how consequential departures will be evaluated.
Where appropriate, it can also identify sensitivity analyses or alternative approaches. The exact level of detail depends on the research context.
This preserves legitimate flexibility in the analysis plan without giving the researcher unrestricted freedom to search for whichever analysis produces the preferred conclusion.
The Choice of Test Can Affect Sample-Size Planning
Another reason not to postpone all statistical decisions is sample size. Sample-size calculations commonly depend on the primary outcome, target effect or precision, design, allocation, variability, significance criterion, statistical model, and other assumptions.
If you do not know what the primary analysis is trying to estimate, it may be difficult to justify the sample size needed to estimate it with adequate precision or power.
This does not mean choosing a test simply because its sample-size formula is convenient. It means the analysis, question, and design need to be aligned while the study can still be modified.
What If More Than One Statistical Method Is Defensible?
Many research questions do not have one uniquely correct statistical procedure. Several models may be defensible while differing in assumptions, efficiency, interpretability, robustness, or the exact quantity they estimate.
When that happens, compare the alternatives before seeing the results. Ask which best represents the research question and design, which assumptions are most defensible, and which output is easiest to interpret substantively.
If genuine uncertainty remains, the analysis plan can specify a primary approach and, when useful, one or more sensitivity analyses. The broader issue of choosing among multiple analyses that could answer the same question deserves more care than simply trying all of them and retaining the most favorable result.
Changing the Planned Test Is Not Automatically Wrong
Sometimes the planned analysis turns out to be inappropriate. An unexpected data problem may appear. A model may fail to converge. A measurement may behave differently from what was reasonably anticipated. A mistake in the original analytical plan may be discovered.
In those circumstances, methodological validity matters more than ceremonial adherence to a bad plan. Use a more appropriate method when necessary.
The important practice is transparency. Preserve the original plan, record what changed and when, explain why the change was justified, and distinguish the revised analysis from what was specified before the relevant results were examined.