Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Should You Choose the Statistical Test Before Collecting the Data?

You should usually determine your main analytical approach before collecting data, and often the primary statistical method as well. But good planning does not require pretending that every analytical decision can be made before you see the data.

153
Choose the Statistical Test Before Data Collection? Guide 153 of 217
01 · The Question

How Can You Choose a Statistical Test Before You Have Seen the Data?

Researchers are often told to decide their statistical analysis before collecting data. Then they encounter another familiar piece of advice: examine the data, check assumptions, and choose an appropriate statistical procedure based on what you find.

At first, those recommendations can appear contradictory. How can you choose between methods that depend partly on data characteristics such as distribution, variance, missingness, or model fit when those data do not yet exist?

The apparent contradiction comes from treating statistical planning as an all-or-nothing decision. You can usually determine the primary analytical strategy, and often the primary statistical method, before collection while explicitly planning how you will respond to data characteristics that cannot reasonably be known in advance.

02 · The Short Answer

Choose the Main Analytical Approach in Advance, With Justified Contingencies

In Brief

Yes, you should usually identify the statistical approach for your main research questions before collecting data, and in many confirmatory studies the primary statistical model or test should be specified in advance.

You do not need to predict every characteristic of data that have not yet been observed. When a choice legitimately depends on features such as serious assumption violations or unforeseen data problems, plan the decision rule or reasonable alternatives in advance where possible and document any later changes.

03 · What You Need to Know

What You Can Decide About Statistical Testing Before the Data Exist

Start With the Question, Design, and Variables

A statistical test should not be chosen by looking at a list of procedures and selecting the one you recognize. The appropriate options are constrained by the research question, study design, analytical target, variables, and structure of the observations.

Are you describing one population or comparing groups? Are observations independent, paired, repeated, or clustered? What is the outcome and how is it measured? Is the question about a difference, association, prediction, change, interaction, or another quantity?

These characteristics exist before the final dataset does. They often narrow the reasonable analytical approaches considerably. Guidance on statistical planning similarly emphasizes the research question, study design, and measurement characteristics when identifying suitable statistical methods.

This is why the study design should not be built around a preferred statistical test. The question and design lead; the analysis follows from them.

You Usually Know More Before Data Collection Than You Think

Consider a planned randomized study with two conditions and an outcome measured before and after an intervention. Before collecting data, you already know that observations from the same participant are related. You know which condition represents the intervention, what the primary outcome is intended to be, when measurements occur, and what comparison the research question requires.

You may not yet know the exact distribution of the observed outcome or the extent of missing data. But those unknowns do not prevent you from defining the main analytical target and selecting an appropriate primary strategy.

The CDC's Field Epidemiology Manual explicitly recommends deciding what data will be collected and how they will be analyzed before the data-collection instrument is designed. Its analysis-plan guidance connects research questions and hypotheses to variables, data handling, measures of association, significance tests, confidence intervals, and other analytical decisions.

Choosing an Analytical Strategy Is Broader Than Naming a Test

“We will use a t-test” may sound specific, but it leaves many important questions unanswered. Which outcome will be tested? Which observations will be included? Are the observations independent or paired? What comparison does the test represent? What happens if important assumptions are unreasonable? What estimate and measure of uncertainty will be reported?

A useful plan therefore moves from the scientific question to the analytical target and then to the statistical method.

Analytical strategy The broader plan connecting the question, design, variables, analytical target, data structure, assumptions, and interpretation.
Statistical test A particular inferential procedure used to evaluate a statistical hypothesis or quantity within that strategy.

For many studies, it is more important to have the first clearly specified than to attach a familiar test name prematurely.

Confirmatory Research Usually Calls for Greater Advance Specification

The stronger the confirmatory purpose of the analysis, the stronger the case for specifying consequential analytical choices before the relevant results are known.

In clinical trials, for example, ICH E9 states that the principal features of the statistical analysis should be described in the statistical section of the protocol and that the principal statistical analysis should be specified before breaking the blind. These formal requirements belong to a particular regulatory setting, so they should not be applied mechanically to every research project.

The underlying reasoning is more general. If researchers can choose among many plausible analyses after learning which one produces the most favorable result, the observed result can influence the analytical decision that supposedly evaluates it.

Advance specification reduces that particular source of analytical flexibility. It does not guarantee that the chosen method is correct.

Assumptions Do Not Necessarily Prevent Advance Planning

A common objection is that you cannot select a test until you have checked its assumptions. Some assumptions do concern observed data characteristics, but others follow largely from the design or analytical model.

Whether observations are paired because the same participants are measured twice is known from the design. Whether students are clustered within classrooms is known from the sampling or intervention structure. Whether an outcome is binary or continuous is usually known from its measurement plan.

Other issues may genuinely require observed data. The appropriate response is not necessarily to leave the analysis completely unspecified. You can often identify the primary method and define what would constitute a consequential problem and what alternative strategy would then be considered.

Watch Out

Avoid creating a mechanical chain in which one preliminary assumption test automatically determines the main analysis without considering the robustness of the planned method, the study design, sample size, purpose of the analysis, and consequences of the departure. Assumption assessment requires statistical judgment, not merely another p-value.

Do Not Reduce the Choice to “Parametric or Nonparametric”

Introductory decision trees often suggest choosing a parametric procedure when assumptions are met and a nonparametric alternative when they are not. Such heuristics can be useful educational starting points, but real analytical decisions may be more nuanced.

Different procedures test or estimate different quantities. Some commonly used parametric methods are reasonably robust to particular departures under some conditions. Transformations, generalized models, robust methods, permutation approaches, or other techniques may sometimes be more appropriate than automatically replacing one familiar test with its presumed nonparametric counterpart.

The question should remain: which method appropriately answers the research question for this design and these data?

Plan the Primary Analysis and Reasonable Contingencies

Suppose your planned primary analysis relies on a model whose assumptions could be seriously compromised by particular data characteristics. Instead of writing “the statistical test will be chosen after examining the data,” the analysis plan can specify the intended primary method and identify how consequential departures will be evaluated.

Where appropriate, it can also identify sensitivity analyses or alternative approaches. The exact level of detail depends on the research context.

This preserves legitimate flexibility in the analysis plan without giving the researcher unrestricted freedom to search for whichever analysis produces the preferred conclusion.

The Choice of Test Can Affect Sample-Size Planning

Another reason not to postpone all statistical decisions is sample size. Sample-size calculations commonly depend on the primary outcome, target effect or precision, design, allocation, variability, significance criterion, statistical model, and other assumptions.

If you do not know what the primary analysis is trying to estimate, it may be difficult to justify the sample size needed to estimate it with adequate precision or power.

This does not mean choosing a test simply because its sample-size formula is convenient. It means the analysis, question, and design need to be aligned while the study can still be modified.

What If More Than One Statistical Method Is Defensible?

Many research questions do not have one uniquely correct statistical procedure. Several models may be defensible while differing in assumptions, efficiency, interpretability, robustness, or the exact quantity they estimate.

When that happens, compare the alternatives before seeing the results. Ask which best represents the research question and design, which assumptions are most defensible, and which output is easiest to interpret substantively.

If genuine uncertainty remains, the analysis plan can specify a primary approach and, when useful, one or more sensitivity analyses. The broader issue of choosing among multiple analyses that could answer the same question deserves more care than simply trying all of them and retaining the most favorable result.

Changing the Planned Test Is Not Automatically Wrong

Sometimes the planned analysis turns out to be inappropriate. An unexpected data problem may appear. A model may fail to converge. A measurement may behave differently from what was reasonably anticipated. A mistake in the original analytical plan may be discovered.

In those circumstances, methodological validity matters more than ceremonial adherence to a bad plan. Use a more appropriate method when necessary.

The important practice is transparency. Preserve the original plan, record what changed and when, explain why the change was justified, and distinguish the revised analysis from what was specified before the relevant results were examined.

04 · A Practical Example

Planning the Test Without Pretending You Know the Future Data

Hypothetical Example

Comparing an Intervention With a Control Condition

A researcher plans a randomized study comparing an educational intervention with a control condition. A continuous primary outcome will be measured at baseline and after the intervention.

Before collection The researcher defines the primary outcome, intervention contrast, baseline measurement, analysis population, and a primary model appropriate to the randomized design.
Sample-size planning The proposed primary analysis and assumptions about variability inform the sample-size calculation rather than leaving sample size disconnected from the main inferential objective.
Contingency planning The analysis plan identifies important assumptions and describes how serious departures or unexpected missingness will be investigated, including any prespecified sensitivity analyses that are warranted.
After collection The researcher examines data quality and relevant model diagnostics without treating every minor imperfection as an automatic reason to search through alternative tests.
If the plan must change A justified alternative is used and reported as a deviation from the planned analysis, with the reason documented.

The researcher did not need to know the exact future distribution of every variable. What mattered was specifying the analytical logic before the results could influence it and planning how genuine uncertainty would be handled.

05 · What Researchers Often Get Wrong

Common Mistakes When Choosing Statistical Tests in Advance

Misconception

“I Cannot Choose Anything Until I Test the Data for Normality”

You can usually determine much of the analytical strategy from the research question, design, variable roles, measurement, and dependency structure. Distributional characteristics may influence particular decisions, but they do not require leaving the entire analysis unspecified until the dataset arrives.

Misconception

“If I Specify a Test in Advance, I Must Use It No Matter What”

No. If the planned procedure becomes demonstrably inappropriate, methodological validity should take priority. Change the analysis when justified, document the reason, and report the deviation transparently.

Misconception

“Choosing the Test in Advance Guarantees an Unbiased Analysis”

Advance specification can constrain some forms of data-dependent decision-making, but a prespecified method can still be poorly chosen, incorrectly implemented, or interpreted beyond what the design supports. Pre-specification complements methodological quality; it does not replace it.

Misconception

“The Statistical Test Can Be Chosen From the Variable Types Alone”

Measurement type matters, but so do the research question, study design, analytical target, number and relationship of groups, repeated or clustered observations, assumptions, and intended interpretation. Knowing that an outcome is continuous does not tell you enough to choose the analysis.

Misconception

“I Can Try Several Tests and Report the One That Works”

Trying alternative analyses can be useful for sensitivity or exploration, but selecting the reported analysis according to which produces the most desirable result creates a different evidential situation. If several methods are genuinely plausible, define their roles and interpretation rather than silently choosing among them after seeing the results.

06 · What This Means for You

Decide What Can Be Decided Before the Results Are Known

The practical goal is not to freeze every statistical choice. It is to prevent decisions that could reasonably be made from the design and research question from being postponed until the results themselves can influence those decisions.

A simple decision framework

If the research question, outcome, design, and data structure clearly identify an appropriate primary method
Specify that method before data collection and use it in sample-size and study planning where relevant.
If a choice genuinely depends on an unknown data characteristic
Define the relevant contingency or decision principle in advance where feasible rather than leaving the choice unrestricted.
If several methods are defensible
Choose a justified primary approach and consider prespecified sensitivity analyses when they would meaningfully test robustness.
If the planned method becomes inappropriate after collection
Use a defensible alternative and transparently document the deviation and its rationale.
If you cannot determine a defensible analysis from the proposed question and design
Revisit the design or seek statistical or methodological input before collecting the data.

Complex designs, unfamiliar models, clustered or longitudinal observations, multiple primary outcomes, and specialized sample-size calculations are good reasons to consider when a statistician or methodologist should become involved. Statistical consultation is considerably more useful when the study can still be changed.

07 · A Quick Checklist

Before You Commit to a Statistical Test

Before data collection begins, check:
The statistical approach follows from the research question rather than familiarity with a particular test.
The primary outcome, comparison, association, effect, or other analytical target is defined.
The method reflects whether observations are independent, paired, repeated, clustered, or otherwise dependent.
The measurement and analytical roles of the important variables are clear.
Important assumptions of the proposed method have been considered rather than deferred entirely until analysis.
Reasonable contingencies or sensitivity analyses are specified where important choices genuinely depend on unknown data characteristics.
The planned primary analysis is compatible with the sample-size justification where relevant.
There is a plan for documenting justified deviations rather than quietly replacing the planned method after results are known.
08 · Frequently Asked Questions

Frequently Asked Questions About Choosing Statistical Tests Before Data Collection

How can I choose a statistical test if I do not know whether my data will be normally distributed?

Normality is only one consideration in statistical selection, and its relevance depends on the particular method and what aspect of the model is assumed to follow a distribution. You can usually specify the primary analytical strategy from the question and design, then plan how consequential assumption violations will be assessed and handled.

Should I run a normality test and let its p-value decide which statistical test to use?

Not as a universal rule. Assumption assessment should consider the statistical method, sample size, graphical and diagnostic information, robustness of the procedure, and consequences of departures. Automatically switching analyses according to one preliminary significance test can oversimplify the problem.

Do exploratory studies need to choose every statistical test in advance?

No. Exploratory research can legitimately preserve greater analytical flexibility. Researchers should still define the study's analytical purpose and distinguish exploratory findings from analyses that were specified as confirmatory before the relevant results were examined.

Can I change a statistical test after data collection?

Yes, when there is a defensible methodological reason. Preserve the planned analysis, explain why it became inappropriate, identify the alternative used, and report the change transparently rather than implying that the replacement was always planned.

Is choosing a statistical test the same as writing an analysis plan?

No. A statistical test is only one component. An analysis plan should address the broader path from research questions and data to analysis and interpretation, including variables, analytical populations, data handling, assumptions, missingness, and other consequential decisions where relevant.

What if I do not know enough statistics to choose the test beforehand?

Clarify the question, design, variables, outcome, and data structure first, then seek appropriate statistical support. Avoid collecting data on the assumption that someone can always identify a suitable test afterward. Some analytical problems are actually design problems that become much harder to repair once collection is complete.

09 · The Bottom Line

Choose What You Can in Advance Without Pretending the Data Are Already Known

The Bottom Line

You should usually identify your primary statistical approach before collecting data, and in confirmatory research you will often want to specify the principal statistical method in considerable detail before the relevant results are known.

Advance planning does not require blind loyalty to a method that later proves inappropriate. Plan the main analysis, anticipate important contingencies where possible, and document justified changes. That gives you both methodological discipline and enough flexibility to deal responsibly with real data.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes