03 · What You Need to Know
How to Test Whether the Question, Design, and Analysis Fit Together
First Identify What Kind of Answer the Research Question Requires
Begin with the verb hidden inside the question. Are you trying to describe something, compare groups, estimate change, examine an association, predict an outcome, evaluate an intervention, explain a process, or interpret experiences and meanings?
These are not interchangeable analytical goals. “What proportion of students use generative AI?” asks for a description. “Is AI use associated with academic performance?” asks about a relationship. “Does allowing AI improve academic performance?” asks a causal question that places much stronger demands on the design and analysis.
The variables may look similar across all three questions. The evidence required is not.
Research question
Defines what you want to know.
Study design
Determines what evidence the study can generate and which interpretations that evidence can support.
Analysis
Uses the evidence generated by the design to estimate, compare, describe, predict, or interpret the quantity or phenomenon relevant to the question.
Ask What Quantity, Comparison, or Pattern Would Actually Answer the Question
Once you know the type of question, define the analytical target more precisely. If you want to compare two groups, what exactly is being compared? Their final scores? Their changes from baseline? Their probability of reaching an outcome? Their trajectories across several time points?
This step prevents a subtle but common mismatch: using a plausible analysis that answers a neighboring question rather than the one that was asked.
In clinical trials, ICH E9(R1) formalizes this idea through the estimand framework. It distinguishes the clinical question and target of estimation from the estimator, which is the analytical method used to estimate that target. The guidance emphasizes alignment among objectives, design, conduct, analysis, and interpretation. Although the framework has a specific regulatory context, the underlying reasoning is widely useful: decide what you are trying to estimate before deciding how to estimate it.
Check Whether the Design Can Produce the Required Evidence
Analysis cannot create design features that never existed.
If your question asks about change, the design needs information capable of representing change. If it asks about differences between conditions, the study needs a defensible comparison. If it asks about an intervention's effect, the design must address alternative explanations strongly enough to support the intended causal interpretation. If observations are clustered, longitudinal, matched, or otherwise dependent, that structure must be reflected in the analysis.
This is also why a statistical test should not be allowed to dictate the study design. A familiar technique cannot determine what scientific question the design is capable of answering.
Watch Out
A more sophisticated statistical model does not automatically strengthen a weak design. Statistical adjustment may address particular measured differences under particular assumptions, but it does not retroactively randomize participants, create missing comparison groups, recover unmeasured variables, or establish temporal order that was never observed.
Check Whether the Variables Represent the Concepts in the Question
Even when the design is appropriate, the analysis can be misaligned if the variables do not operationalize the concepts the question refers to.
Suppose a question asks whether “academic success” differs between groups, but the study measures only one quiz score. The statistical comparison may be perfectly valid for that quiz score. The broader claim about academic success may nevertheless exceed what was measured.
Alignment therefore requires both statistical and conceptual correspondence. Ask whether each important construct in the research question has an appropriate observable representation and whether the proposed analysis uses that representation in the role the question implies.
The roles of variables in the analysis should reflect the conceptual and design logic of the study rather than labels assigned after the dataset is assembled.
Check the Unit of Analysis Against the Unit of the Question and Design
Ask who or what the conclusion is about. Is it students, classrooms, schools, hospitals, countries, repeated observations, documents, or something else?
Then ask how those units entered the study and how they relate to one another. Students within the same classroom may share an instructor and learning environment. Measurements from the same participant over time are related. Patients treated at the same hospital may share institutional conditions.
An analysis that treats dependent observations as independent can misrepresent uncertainty and, depending on the context, alter the substantive interpretation. The appropriate method should therefore reflect clustering, matching, repeated measurement, nesting, or other dependencies created by the design.
Make Sure the Analysis Respects How the Data Were Obtained
Sampling and assignment matter. A probability sample, convenience sample, randomized experiment, observational cohort, case-control study, cross-sectional survey, and purposive qualitative sample generate different forms of evidence.
The same numerical procedure does not acquire the same interpretation in every design. For example, an association estimated from a cross-sectional observational dataset is not transformed into a causal effect simply because the analysis uses regression and adjusts for several variables.
Interpretation should therefore inherit the limitations of the design. Statistical analysis can refine the evidence the design provides; it cannot grant the study evidential properties it does not possess.
Check Whether the Method Matches the Measurement and Data Structure
Analytical methods make assumptions about the form and structure of data. Relevant considerations may include whether an outcome is continuous, binary, ordinal, nominal, a count, or time-to-event; whether measurements are independent or repeated; whether observations are censored; and whether data are nested or clustered.
Measurement level alone does not mechanically dictate a single test. Several methods may be defensible for the same broad type of variable depending on the question and design. The point is to rule out methods whose mathematical and inferential requirements conflict with the data you intend to collect.
This is one reason that simply choosing a statistical test before collecting data is not enough. The test has to sit within a coherent analytical strategy.
Check Whether the Analysis Answers Every Part of the Question, but Nothing More
Research questions sometimes contain more analytical content than their authors realize.
“Is there a difference?” asks something different from “How large is the difference?” “Which variables predict the outcome?” differs from “Which variables cause the outcome?” “Does the relationship differ by gender?” introduces an effect-modification question rather than merely asking whether gender is associated with the outcome.
Break complex questions into their analytical components. Then check whether the planned analysis produces evidence for each component. Conversely, do not allow the interpretation to expand beyond the question and evidence simply because the software produced additional coefficients.
Match Each Hypothesis to the Quantity Being Tested
When a study contains formal hypotheses, identify exactly what statistical quantity corresponds to each hypothesis. A hypothesis about a difference between conditions should map to the comparison that represents that difference. A hypothesis about an interaction requires an analysis capable of evaluating that interaction rather than separate significance tests within each subgroup.
It should therefore be possible to trace each substantive hypothesis to a corresponding analysis without inventing the connection after the results are known.
Do Not Confuse Statistical Significance With Answering the Research Question
A small p-value does not prove that the analysis matches the question. It tells you something about the observed data under a specified statistical model and null hypothesis. It does not determine whether the outcome was measured appropriately, whether the design supports a causal claim, whether the effect is practically important, or whether the sample represents the population about which you want to speak.
Similarly, failure to cross a significance threshold does not necessarily mean “there is no effect” or “there is no relationship.” The estimate, its uncertainty, the design, sample size, measurement quality, and substantive context all matter.
Plan what estimate and uncertainty would answer the question, not merely which test will generate a p-value.
Perform the Conclusion Test Before Collecting Data
One of the simplest alignment checks is to write the kind of conclusion you expect the study could legitimately support, leaving the result itself blank.
For example: “Students assigned to the intervention had, on average, ___ points higher post-intervention scores than students assigned to the comparison condition, after accounting for ___ as specified.” Then ask whether the design and analysis genuinely support every phrase in that sentence.
If you find yourself writing “caused” when there was no design basis for causal inference, “improved” when there was no baseline or longitudinal information, or “students generally” when the sample came from one highly selected setting, the mismatch becomes visible before any statistical output can disguise it.