Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Should an Analysis Plan Decide Before Data Collection Begins?

A useful analysis plan does more than name a statistical test. Before data collection, it should connect each research question to the data, comparisons, analytical methods, assumptions, and decisions needed to answer it.

147
What an Analysis Plan Should Decide Guide 147 of 217
01 · The Question

What Exactly Should You Decide Before the Data Exist?

Knowing that you should plan your analysis before collecting data raises a more difficult question: how much should you actually decide?

Writing “the data will be analyzed using appropriate statistical methods” is not much of a plan. Neither is listing software and several familiar tests without explaining what each one will answer. At the other extreme, trying to anticipate every possible feature of a dataset that does not yet exist can create unnecessary rigidity.

A useful analysis plan sits between those extremes. It should specify the decisions that determine how your research questions will be translated into evidence, while identifying reasonable contingencies for analytical decisions that may depend on what happens during the study.

02 · The Short Answer

An Analysis Plan Should Map Questions to Evidence

In Brief

Before data collection, your analysis plan should specify what questions or hypotheses will be analyzed, which data and variables will answer them, who or what will be included in each analysis, how the data will be prepared and analyzed, how important complications will be handled, and how the resulting estimates or findings will be interpreted.

The required level of detail depends on the study. A confirmatory quantitative study may require substantial pre-specification, whereas exploratory or iterative designs can legitimately preserve more flexibility. The plan should nevertheless be detailed enough to expose mismatches between the question, design, data collection, and intended analysis before those mismatches become irreversible.

03 · What You Need to Know

The Decisions a Useful Analysis Plan Should Make

Start With the Research Questions, Not the Statistical Software

The first task is to identify what the analysis must answer. An analysis plan should therefore begin with the research questions, objectives, or hypotheses rather than with a menu of statistical procedures.

The U.S. Centers for Disease Control and Prevention describes an analysis plan as a guide from raw data to the final report and notes that research questions or hypotheses ordinarily lead to the variables and methods that need to be analyzed. Similarly, the National Center for Advancing Translational Sciences describes a data analysis plan as a roadmap for organizing and analyzing data and presenting the results.

For each substantive question, you should be able to trace a path from the question to the evidence needed to answer it. If that path cannot be articulated before collection, the problem may lie in the question, measurement, design, or proposed analysis rather than in the wording of the plan.

This is why it helps to establish a clear connection between each research question and its planned analysis rather than writing one generic analysis section for the entire study.

Define the Outcome, Exposure, Predictor, Grouping, and Other Variable Roles

A variable name by itself tells you surprisingly little. The plan should establish what role each important variable plays in the analysis.

Depending on the design, a variable may function as an outcome, predictor, exposure, treatment or condition indicator, covariate, confounder, moderator, mediator, stratification variable, clustering variable, or another analytically relevant quantity. Those roles should follow from the research question and design rather than being assigned opportunistically after examining associations in the data.

You should also define how important variables will be represented. Will an outcome be analyzed as its final value, change from baseline, a proportion, a category, a count, or time to an event? If a continuous measurement will be categorized, what determines the categories? If a composite score will be created, how will its components be combined?

These decisions matter because the roles assigned to variables shape what analyses are appropriate and what conclusions those analyses can support.

Specify the Unit of Analysis

The unit of analysis is the entity about which the analysis produces information. It might be an individual student, patient, classroom, school, organization, household, document, interview, observation period, or another unit.

This can differ from the unit at which data are collected or an intervention is assigned. For example, researchers may collect responses from individual students while assigning an intervention to entire classrooms. Treating every student as statistically independent would then ignore the clustering created by the design.

The analysis plan should therefore identify the relevant unit or units and any dependency structure that the analysis must accommodate, including repeated measurements, nested observations, paired data, or clusters.

Define the Analysis Population or Inclusion Rules

Who or what contributes to each analysis? The answer is not always simply “everyone in the dataset.”

A plan may need to define eligibility for a particular analysis, how repeated observations contribute, which observations occur within the relevant time window, and whether any exclusions are analytically justified. In randomized trials, formal analysis populations may also be specified according to the study's estimand and protocol.

These rules are particularly important when excluding observations could materially change the result. Decisions should not quietly become more permissive or restrictive after researchers learn which observations favor a particular conclusion.

Specify the Main Comparison, Association, or Quantity You Want to Estimate

Before choosing a technique, state what you actually want the analysis to produce.

Are you estimating the difference in mean outcomes between groups? A change over time? An association between two variables? The probability of an event? A relative risk? The effect of an intervention under a particular set of conditions? A predictive quantity?

This distinction is easy to overlook because two analyses may use the same variables while answering different questions. In clinical trials, ICH E9(R1) makes this issue particularly explicit through the concept of an estimand: a precise description of the treatment effect corresponding to the clinical question. Although that regulatory framework is specific to clinical trials, its broader lesson is useful elsewhere. Researchers should be clear about the quantity or comparison their analysis is intended to estimate.

Choose the Analytical Method Because It Fits the Question and Design

Once the target of the analysis is clear, specify the method that will be used to estimate, test, describe, classify, or interpret it.

For a quantitative study, this may involve descriptive statistics, regression models, group comparisons, longitudinal models, survival analysis, multilevel models, or other techniques. The choice depends on the research question, design, measurement level, data structure, intended inference, and relevant assumptions.

This is more defensible than beginning with a preferred statistical test and designing the study around it. The specific statistical test may sometimes be selected in advance, but selecting a test is only one part of planning the analysis.

State Which Analyses Are Primary and Which Are Secondary or Exploratory

If a study has several outcomes, hypotheses, comparisons, time points, subgroups, or models, the plan should distinguish the analyses that directly address its central claims from analyses serving secondary or exploratory purposes.

This is especially important when many analytical choices are available. Without priorities established in advance, researchers may be tempted to emphasize whichever result appears most favorable after analysis.

Not every study requires the terminology of “primary” and “secondary” analysis. In studies where that distinction is meaningful, however, it can clarify which result carries the main evidential burden and which findings should be interpreted more cautiously.

Plan How Data Will Be Prepared Before Analysis

Raw data rarely move directly from collection into a final model. The plan should identify consequential preparation steps when they can reasonably be anticipated.

These may include coding categorical responses, calculating scale or composite scores, transforming variables, identifying impossible values, reconciling duplicate observations, deriving new variables, or applying predefined rules to measurements.

Routine data cleaning need not become an encyclopedic list of every possible typo. The important decisions are those that could change which observations enter an analysis or what values they contribute.

Decide How Missing Data Will Be Addressed

Missing data should not be treated as a problem to think about only after opening the dataset. The consequences depend on why data are missing, how much is missing, where the missingness occurs, and what assumptions the analytical approach requires.

An advance plan might describe procedures intended to minimize missingness, how the extent and patterns of missing data will be summarized, the primary analytical approach when values are missing, and any sensitivity analyses needed to examine how conclusions depend on assumptions about the missingness.

CONSORT 2025, for randomized trials, explicitly asks authors to report how missing data were handled, while its explanatory guidance emphasizes the assumptions and methods underlying those decisions. The precise requirements vary across research designs, but the planning principle is broader: do not allow the treatment of missing observations to become an undocumented choice made after seeing how alternatives affect the result.

Identify Assumptions and What You Will Do if They Are Not Reasonable

Analytical procedures depend on assumptions. The analysis plan should identify assumptions that are consequential enough to influence whether the proposed method remains appropriate.

It should also distinguish assumption checking from automatic decision rules. A single diagnostic result does not necessarily dictate that one method must be abandoned for another. The relevant question is whether departures are substantial enough to undermine the intended inference and, if so, what alternative or sensitivity analysis is defensible.

Where reasonable, specify those contingencies before the results are known. This preserves flexibility without turning the analysis into improvisation.

Plan Adjustments, Subgroups, and Multiple Analyses Deliberately

If the primary analysis will adjust for covariates, identify them and explain why they belong in the model. If subgroup analyses are substantively important, specify which subgroups matter and what question each comparison addresses. If multiple outcomes or comparisons create a multiplicity issue, determine how that will affect testing or interpretation.

These choices can become highly result-sensitive when made after the data are examined. A variable should not become a covariate merely because adjustment improves a result, nor should a subgroup become theoretically important only after producing a striking difference.

Decide What Will Be Reported, Not Merely What Will Be Tested

An analysis plan should anticipate the form in which results will be communicated. For quantitative analyses, this often means identifying the relevant effect estimates, measures of uncertainty such as confidence intervals, and descriptive information needed to interpret the sample and outcome.

Planning table shells or anticipated figures can be surprisingly useful. They force you to ask what numbers must eventually occupy each cell and whether your proposed data collection can actually generate them. The CDC's Field Epidemiology Manual specifically identifies table shells as a practical way to plan an analysis.

Document Which Decisions Can Still Change

An analysis plan is not useful because it predicts every problem. It is useful because it makes the analytical reasoning visible before the results can influence that reasoning.

Some decisions may legitimately remain conditional. The plan can specify what is fixed, what depends on observable data characteristics, what conditions would justify an alternative approach, and how deviations will be documented.

This makes the distinction between planned adaptation and retrospective optimization much clearer. It also creates a natural record when the analysis plan needs legitimate flexibility.

04 · A Practical Example

Turning a Research Question Into an Analysis Plan

Hypothetical Example

Does a New Teaching Approach Improve Student Performance?

Suppose a researcher will compare two instructional approaches across several classes and measure student achievement before and after the intervention. The research question asks whether students receiving the new approach improve more than students receiving the comparison approach.

Question Is improvement in achievement different between the instructional conditions?
Outcome The plan defines how achievement is measured, identifies the relevant post-intervention time point, and specifies how the baseline measure will enter the analysis.
Data structure Students are nested within classes, so the plan recognizes that observations from students in the same class may not be independent.
Primary analysis The researcher identifies an analytical approach appropriate to the group comparison, baseline information, and clustered structure rather than treating all student observations as unrelated.
Missing data The plan states how missing outcome observations will be described and how the primary analysis and any necessary sensitivity analyses will address them.
Reporting The researcher plans to report the estimated difference between conditions with an appropriate measure of uncertainty, rather than relying only on whether a significance threshold is crossed.

Notice what the plan accomplishes before a single outcome is collected. It reveals what must be measured, how the structure of the design affects the analysis, and what the final result is intended to mean. The specific model still has to be technically justified, but the analytical logic is no longer waiting to be invented after the results appear.

05 · What Researchers Often Get Wrong

What an Analysis Plan Should Not Become

Misconception

“I Listed the Statistical Tests, So I Have an Analysis Plan”

A list of tests does not establish which question each procedure answers, which variables enter it, what comparison is being made, which observations are included, or how complications will be handled. Start with the analytical question and work toward the method.

Misconception

“I Can Decide the Covariates After Seeing Which Ones Matter”

Some model modifications may be defensible after data inspection, but choosing adjustment variables solely according to which specification produces a preferred result creates substantial analytical flexibility. When covariate adjustment is important to the main analysis, its rationale should generally come from the design, substantive knowledge, or a defensible analytical strategy rather than the desirability of the observed result.

Misconception

“Missing Data Are a Data-Cleaning Problem”

Missingness can affect the validity and interpretation of an analysis. How missing observations are handled may depend on assumptions about why they are missing, not merely on software settings. Important missing-data decisions belong in analytical planning.

Misconception

“More Detail Always Makes an Analysis Plan Better”

Detail is useful when it constrains consequential analytical choices or makes the analysis reproducible. Specifying arbitrary rules for every imaginable event can instead make the plan brittle or obscure the decisions that actually matter. The appropriate level of specification depends on the study's purpose and methodology.

Misconception

“Anything Not Specified in Advance Is Invalid”

Unplanned analyses are not inherently invalid. They may be necessary responses to unforeseen problems or valuable exploratory investigations. The key is to distinguish them from pre-specified analyses and explain why important deviations occurred.

06 · What This Means for You

Build the Plan by Working Backward From the Answer You Need

A practical way to write an analysis plan is to imagine the results section before the dataset exists. What conclusion would answer the research question? What estimate, comparison, pattern, or interpretation would support that conclusion? What analysis produces that evidence? What variables and observations does that analysis require?

A simple decision framework

If a decision could materially change the answer to your primary research question
Try to specify it before examining the relevant results.
If a decision depends on a data characteristic that cannot reasonably be known beforehand
Specify the principle or contingency that will guide the decision rather than inventing an arbitrary fixed choice.
If several analyses could defensibly answer the same question
Identify the primary approach and explain why it best matches the question, design, and intended inference.
If you cannot specify an analysis without changing the proposed data collection
Treat that discovery as useful design feedback and revise the study before collection begins.

The aim is not to produce paperwork for its own sake. A good plan should make the study easier to reason about. If writing it reveals that you do not know which variable answers a question, which observations belong in the comparison, or what the result will actually estimate, you have found a design problem at precisely the right time.

07 · A Quick Checklist

What to Check in Your Analysis Plan Before Data Collection

Before finalizing the analysis plan, check:
Each central research question or hypothesis is connected to an identifiable analysis.
The outcome, predictor, exposure, grouping, adjustment, and other important variable roles are defined where relevant.
The unit of analysis and important dependencies such as clustering or repeated measurements are recognized.
The main comparison, association, effect, or other quantity of interest is clear.
The primary analytical method is justified by the question and design rather than familiarity alone.
Consequential data preparation, exclusion, and missing-data decisions are specified or governed by clear principles.
Important assumptions and reasonable contingencies have been considered.
Primary, secondary, sensitivity, subgroup, and exploratory analyses are distinguished where those categories are relevant.
The planned results include the estimates and uncertainty needed to answer the question, not only significance tests.
Any decisions intentionally left flexible can be explained and documented later.
08 · Frequently Asked Questions

Frequently Asked Questions About Analysis Plans

Does an analysis plan have to name the exact statistical test?

Not always. The level of specificity should match the study. For a confirmatory quantitative study, the primary model or procedure may need to be specified in considerable detail. Earlier-stage or exploratory work may appropriately specify the analytical strategy while leaving some choices conditional on data characteristics.

Should the analysis plan include descriptive statistics?

Usually, when they are relevant to understanding the sample, variables, data quality, or results. The plan can specify which characteristics and outcomes will be summarized and how, but descriptive analysis should not be confused with the analysis required to answer an inferential research question.

Should I specify software in the analysis plan?

You can, particularly when software or package versions affect reproducibility, but naming software does not substitute for describing the analytical method. “Data will be analyzed in R” or “using SPSS” explains the computational environment, not the analytical reasoning.

What if more than one method could answer my question?

That is common. Compare the alternatives according to the question, design, data structure, assumptions, estimand or target quantity, and interpretability. When several approaches remain defensible, the choice among multiple possible analyses should be justified rather than made according to which result looks best.

Is a data analysis plan the same as a statistical analysis plan?

Not necessarily. “Data analysis plan” is a broader term used across many forms of research. A statistical analysis plan commonly refers to a more detailed specification of statistical analyses, particularly in clinical and other quantitative research contexts. Terminology and formal requirements vary by discipline, institution, funder, and regulatory setting.

Can I change the plan after data collection starts?

Sometimes. A legitimate methodological or operational reason may require amendment. Preserve the original plan, document what changed and when, explain why, and distinguish analyses specified before relevant results were examined from analyses developed afterward.

09 · The Bottom Line

A Good Analysis Plan Makes the Path From Question to Answer Explicit

The Bottom Line

Before data collection, an analysis plan should specify enough of the path from research question to data to analysis to interpretation that you can determine whether the study will actually produce the evidence needed to answer its questions.

That usually means defining the relevant variables and observations, analytical target, main methods, important data-handling decisions, assumptions, and contingencies. The objective is not to eliminate every future judgment. It is to make consequential analytical choices deliberately, while you can still improve the study rather than merely work around its limitations.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes