03 · What You Need to Know
The Decisions a Useful Analysis Plan Should Make
Start With the Research Questions, Not the Statistical Software
The first task is to identify what the analysis must answer. An analysis plan should therefore begin with the research questions, objectives, or hypotheses rather than with a menu of statistical procedures.
The U.S. Centers for Disease Control and Prevention describes an analysis plan as a guide from raw data to the final report and notes that research questions or hypotheses ordinarily lead to the variables and methods that need to be analyzed. Similarly, the National Center for Advancing Translational Sciences describes a data analysis plan as a roadmap for organizing and analyzing data and presenting the results.
For each substantive question, you should be able to trace a path from the question to the evidence needed to answer it. If that path cannot be articulated before collection, the problem may lie in the question, measurement, design, or proposed analysis rather than in the wording of the plan.
This is why it helps to establish a clear connection between each research question and its planned analysis rather than writing one generic analysis section for the entire study.
Define the Outcome, Exposure, Predictor, Grouping, and Other Variable Roles
A variable name by itself tells you surprisingly little. The plan should establish what role each important variable plays in the analysis.
Depending on the design, a variable may function as an outcome, predictor, exposure, treatment or condition indicator, covariate, confounder, moderator, mediator, stratification variable, clustering variable, or another analytically relevant quantity. Those roles should follow from the research question and design rather than being assigned opportunistically after examining associations in the data.
You should also define how important variables will be represented. Will an outcome be analyzed as its final value, change from baseline, a proportion, a category, a count, or time to an event? If a continuous measurement will be categorized, what determines the categories? If a composite score will be created, how will its components be combined?
These decisions matter because the roles assigned to variables shape what analyses are appropriate and what conclusions those analyses can support.
Specify the Unit of Analysis
The unit of analysis is the entity about which the analysis produces information. It might be an individual student, patient, classroom, school, organization, household, document, interview, observation period, or another unit.
This can differ from the unit at which data are collected or an intervention is assigned. For example, researchers may collect responses from individual students while assigning an intervention to entire classrooms. Treating every student as statistically independent would then ignore the clustering created by the design.
The analysis plan should therefore identify the relevant unit or units and any dependency structure that the analysis must accommodate, including repeated measurements, nested observations, paired data, or clusters.
Define the Analysis Population or Inclusion Rules
Who or what contributes to each analysis? The answer is not always simply “everyone in the dataset.”
A plan may need to define eligibility for a particular analysis, how repeated observations contribute, which observations occur within the relevant time window, and whether any exclusions are analytically justified. In randomized trials, formal analysis populations may also be specified according to the study's estimand and protocol.
These rules are particularly important when excluding observations could materially change the result. Decisions should not quietly become more permissive or restrictive after researchers learn which observations favor a particular conclusion.
Specify the Main Comparison, Association, or Quantity You Want to Estimate
Before choosing a technique, state what you actually want the analysis to produce.
Are you estimating the difference in mean outcomes between groups? A change over time? An association between two variables? The probability of an event? A relative risk? The effect of an intervention under a particular set of conditions? A predictive quantity?
This distinction is easy to overlook because two analyses may use the same variables while answering different questions. In clinical trials, ICH E9(R1) makes this issue particularly explicit through the concept of an estimand: a precise description of the treatment effect corresponding to the clinical question. Although that regulatory framework is specific to clinical trials, its broader lesson is useful elsewhere. Researchers should be clear about the quantity or comparison their analysis is intended to estimate.
Choose the Analytical Method Because It Fits the Question and Design
Once the target of the analysis is clear, specify the method that will be used to estimate, test, describe, classify, or interpret it.
For a quantitative study, this may involve descriptive statistics, regression models, group comparisons, longitudinal models, survival analysis, multilevel models, or other techniques. The choice depends on the research question, design, measurement level, data structure, intended inference, and relevant assumptions.
This is more defensible than beginning with a preferred statistical test and designing the study around it. The specific statistical test may sometimes be selected in advance, but selecting a test is only one part of planning the analysis.
State Which Analyses Are Primary and Which Are Secondary or Exploratory
If a study has several outcomes, hypotheses, comparisons, time points, subgroups, or models, the plan should distinguish the analyses that directly address its central claims from analyses serving secondary or exploratory purposes.
This is especially important when many analytical choices are available. Without priorities established in advance, researchers may be tempted to emphasize whichever result appears most favorable after analysis.
Not every study requires the terminology of “primary” and “secondary” analysis. In studies where that distinction is meaningful, however, it can clarify which result carries the main evidential burden and which findings should be interpreted more cautiously.
Plan How Data Will Be Prepared Before Analysis
Raw data rarely move directly from collection into a final model. The plan should identify consequential preparation steps when they can reasonably be anticipated.
These may include coding categorical responses, calculating scale or composite scores, transforming variables, identifying impossible values, reconciling duplicate observations, deriving new variables, or applying predefined rules to measurements.
Routine data cleaning need not become an encyclopedic list of every possible typo. The important decisions are those that could change which observations enter an analysis or what values they contribute.
Decide How Missing Data Will Be Addressed
Missing data should not be treated as a problem to think about only after opening the dataset. The consequences depend on why data are missing, how much is missing, where the missingness occurs, and what assumptions the analytical approach requires.
An advance plan might describe procedures intended to minimize missingness, how the extent and patterns of missing data will be summarized, the primary analytical approach when values are missing, and any sensitivity analyses needed to examine how conclusions depend on assumptions about the missingness.
CONSORT 2025, for randomized trials, explicitly asks authors to report how missing data were handled, while its explanatory guidance emphasizes the assumptions and methods underlying those decisions. The precise requirements vary across research designs, but the planning principle is broader: do not allow the treatment of missing observations to become an undocumented choice made after seeing how alternatives affect the result.
Identify Assumptions and What You Will Do if They Are Not Reasonable
Analytical procedures depend on assumptions. The analysis plan should identify assumptions that are consequential enough to influence whether the proposed method remains appropriate.
It should also distinguish assumption checking from automatic decision rules. A single diagnostic result does not necessarily dictate that one method must be abandoned for another. The relevant question is whether departures are substantial enough to undermine the intended inference and, if so, what alternative or sensitivity analysis is defensible.
Where reasonable, specify those contingencies before the results are known. This preserves flexibility without turning the analysis into improvisation.
Plan Adjustments, Subgroups, and Multiple Analyses Deliberately
If the primary analysis will adjust for covariates, identify them and explain why they belong in the model. If subgroup analyses are substantively important, specify which subgroups matter and what question each comparison addresses. If multiple outcomes or comparisons create a multiplicity issue, determine how that will affect testing or interpretation.
These choices can become highly result-sensitive when made after the data are examined. A variable should not become a covariate merely because adjustment improves a result, nor should a subgroup become theoretically important only after producing a striking difference.
Decide What Will Be Reported, Not Merely What Will Be Tested
An analysis plan should anticipate the form in which results will be communicated. For quantitative analyses, this often means identifying the relevant effect estimates, measures of uncertainty such as confidence intervals, and descriptive information needed to interpret the sample and outcome.
Planning table shells or anticipated figures can be surprisingly useful. They force you to ask what numbers must eventually occupy each cell and whether your proposed data collection can actually generate them. The CDC's Field Epidemiology Manual specifically identifies table shells as a practical way to plan an analysis.
Document Which Decisions Can Still Change
An analysis plan is not useful because it predicts every problem. It is useful because it makes the analytical reasoning visible before the results can influence that reasoning.
Some decisions may legitimately remain conditional. The plan can specify what is fixed, what depends on observable data characteristics, what conditions would justify an alternative approach, and how deviations will be documented.
This makes the distinction between planned adaptation and retrospective optimization much clearer. It also creates a natural record when the analysis plan needs legitimate flexibility.