03 · What You Need to Know
Analysis Should Follow the Design, Not Rescue It
Start with the research question, not the name of the statistical test
A statistical method is appropriate only relative to the analytical task it is being asked to perform.
Before deciding whether a t-test, regression, ANOVA, mixed model, survival analysis, structural equation model, thematic analysis, or another method is appropriate, ask:
What exactly are the researchers trying to estimate, compare, predict, classify, explain, or interpret?
For example:
| Research task |
Analytical question |
| Compare two groups |
How different are the groups on the relevant outcome? |
| Estimate change |
How does the outcome change across relevant time points? |
| Examine association |
How are two or more variables related? |
| Predict an outcome |
How accurately can specified information predict new or future observations? |
| Estimate an intervention effect |
What contrast represents the outcome under intervention versus the relevant alternative? |
| Analyze clustered observations |
How should dependence among observations within groups be handled? |
| Interpret qualitative material |
How were data transformed into categories, themes, patterns, explanations, or interpretations? |
The technical procedure comes after the analytical question.
The analysis cannot change what kind of study was conducted
This is one of the most important principles in critical appraisal.
A sophisticated statistical model cannot retroactively randomize participants, create a missing control group, measure a variable that was never collected, make a convenience sample representative, or transform a cross-sectional dataset into longitudinal evidence.
Analysis can sometimes address particular biases or improve estimation under explicit assumptions. It cannot manufacture design features that do not exist.
Design
Determines how observations, comparisons, exposures, interventions, measurements, and time points enter the study.
Analysis
Uses the resulting data to estimate, compare, model, classify, summarize, or interpret quantities relevant to the research question.
The second should respect the first.
This is why analysis cannot be evaluated independently of whether the study design itself is capable of answering the research question.
Identify the unit of analysis
Ask what counts as one observation in the analysis.
Is it one student? One classroom? One school? One patient? One hospital? One country? One document? One measurement occasion? One eye from each patient? Multiple observations from the same person?
This matters because the unit analyzed should correspond appropriately to the way the data were generated.
Suppose an intervention is assigned to entire classrooms, but the analysis treats 500 students as though all 500 were independently randomized.
The apparent sample size may then exaggerate the amount of independent information in the study because students within the same classroom share the same intervention assignment and may resemble one another for other reasons.
That is not a cosmetic statistical issue. It changes the uncertainty attached to the intervention comparison.
Independence is an assumption you should not grant automatically
Many common statistical procedures assume that observations are independent or otherwise require dependence to be modeled appropriately.
Research data frequently violate simple independence:
- students are nested within classrooms;
- patients are nested within hospitals;
- employees are nested within organizations;
- repeated measurements belong to the same participant;
- siblings belong to the same family;
- multiple observations may come from the same geographical area;
- two eyes, limbs, lesions, or other units may belong to the same person.
When observations share a context or individual, their values may be more similar than observations selected independently.
Appropriate analyses may use multilevel models, mixed-effects models, generalized estimating equations, cluster-robust standard errors, design-based survey methods, paired procedures, or other approaches depending on the design and inferential goal.
You do not need to decide which specific method is best before noticing the underlying problem: are observations related in a way the analysis needs to recognize?
Repeated measures should usually not be analyzed as unrelated observations
Suppose researchers measure the same students before and after an intervention.
Each student's post-test score is related to that student's pretest score. Treating the pretest and post-test samples as though they came from completely different people discards that pairing and misrepresents the data structure.
Depending on the question and design, appropriate approaches might include paired comparisons, repeated-measures models, analysis of covariance, mixed models, or other longitudinal methods.
The precise choice requires methodological judgment. The critical-reading principle is simpler:
If the same units were measured repeatedly, does the analysis acknowledge that the observations belong together?
The outcome type constrains which analyses make sense
Different outcomes contain different kinds of information.
A continuous test score, binary disease outcome, ordered rating, count of events, time-to-event outcome, and nominal category should not automatically be analyzed as though they were the same kind of variable.
| Outcome |
Example |
Analytical feature that matters |
| Continuous |
Writing score from 0 to 100 |
Distribution, scale, variance, model form, and other assumptions |
| Binary |
Passed or failed |
Probability or odds of the outcome and appropriate binary-outcome modeling |
| Count |
Number of hospital admissions |
Non-negative integer structure, dispersion, exposure time, and count-model assumptions |
| Ordinal |
Five-level severity rating |
Ordering without automatically assuming equal distances between categories |
| Time-to-event |
Time until relapse |
Censoring and timing of events |
| Nominal category |
Choice among several unrelated options |
Multiple categories without natural ordering |
An analysis can sometimes accommodate these outcomes in several defensible ways. What matters is whether the chosen approach preserves the important structure of the data and supports the interpretation made from it.
Ask whether the analysis matches how groups were created
Group comparisons mean different things depending on how people entered the groups.
If participants were randomized, the analysis should preserve the logic of random assignment. If groups arose naturally, such as users versus non-users of a technology, the analysis must contend with the possibility that the groups differ for reasons other than the exposure of interest.
Suppose students choose whether to use an optional AI tutoring tool. Users subsequently achieve higher grades.
A regression model adjusting for prior achievement, age, and study time may reduce some observed differences between groups. But the analysis does not automatically eliminate unmeasured differences such as motivation or instructor encouragement.
The model can adjust for measured covariates. It cannot guarantee that observational groups have become equivalent to randomized groups.
Adjustment is not automatically better than no adjustment
Regression models often include covariates, and readers may assume that a more heavily adjusted model must provide a more accurate answer.
That is not necessarily true.
Variables should be included for defensible reasons related to the research question, causal structure, precision, design, or analysis plan. Adjusting for inappropriate variables can introduce bias rather than remove it.
For causal questions, variables affected by the exposure, colliders, or variables selected solely because they are statistically significant can create interpretive problems.
The exact causal logic can become technically demanding, but the practical reading question is:
Why were these variables adjusted for, and does that choice correspond to the causal or analytical question?
A table containing fifteen covariates is not evidence that confounding has been solved.
Baseline adjustment should correspond to the design
In intervention studies with baseline and follow-up measurements, authors may compare post-intervention scores, change scores, or use models adjusting follow-up outcomes for baseline values.
These approaches are not interchangeable in every setting.
Baseline adjustment can improve precision and address chance baseline differences under appropriate conditions. In randomized trials, analysis of covariance is commonly used for continuous outcomes when baseline values are available.
The relevant question is whether the chosen approach corresponds to the study design and estimand rather than whether one method is universally superior.
If groups differ substantially at baseline, simply comparing post-test means without considering the design and baseline information may be difficult to interpret.
Ask whether the analysis follows the original assignment in randomized trials
Randomization creates the basis for causal comparison by assigning participants to conditions.
Problems can arise if the analysis later reorganizes participants according to what they actually did rather than how they were assigned.
For example, suppose participants randomized to an exercise intervention are analyzed only if they attended at least 80% of sessions. Those highly adherent participants may differ systematically from participants who did not adhere.
Intention-to-treat analysis generally preserves participants in the groups to which they were randomized and is often important for estimating the effect of assignment to an intervention under trial conditions. Other estimands and analyses can also be legitimate, but they answer different questions and require corresponding assumptions.
Do not treat "per-protocol" and "intention-to-treat" as merely two technical labels. Ask what population and treatment effect each analysis is attempting to estimate.
Missing data are part of the analysis, not an inconvenience outside it
Suppose 20% of participants lack follow-up measurements.
What did the researchers do?
They may analyze only complete cases, use multiple imputation, model outcomes under assumptions about missingness, perform sensitivity analyses, or use another strategy.
No method can recover missing information without assumptions.
The seriousness depends partly on why data are missing and whether missingness is related to variables or outcomes relevant to the analysis.
A complete-case analysis can be reasonable under some conditions and biased under others. Multiple imputation can be useful when its model and assumptions are appropriate, but it does not magically recreate unknowable data.
Ask:
- How much data are missing?
- Which variables and groups are affected?
- Why might the data be missing?
- What assumptions does the analytical strategy make?
- Would plausible alternative assumptions materially change the conclusion?
Dropping participants can change the question the study answers
Researchers may exclude participants with incomplete data, protocol deviations, outlying observations, low adherence, or other characteristics.
Some exclusions are necessary.
But the analytical sample can gradually become different from the sample defined by the original design.
If an intervention is difficult to tolerate and researchers exclude everyone who failed to complete it, the remaining sample may disproportionately contain people for whom the intervention was easiest or most successful.
When exclusions are substantial, examine who ultimately contributed to the analysis and how that differs from the sample originally recruited or randomized.
Multiple outcomes create multiple opportunities for interesting results
A study may measure several outcomes, analyze multiple subscales, examine several time points, fit different models, and test numerous subgroups.
That flexibility can increase the chance of obtaining apparently noteworthy findings even when no strong underlying effect exists.
Appropriate approaches depend on the study and inferential framework. Researchers may prespecify primary outcomes, adjust for multiple comparisons, distinguish confirmatory from exploratory analyses, or interpret secondary findings cautiously.
The key question is not simply whether a multiple-testing correction was used.
Ask:
How many analytical opportunities existed, which results were planned as primary, and is the prominence of the reported finding proportionate to its role in the original analysis?
Do not judge multiple testing from the paper’s main table alone
A paper may prominently display only a few statistically significant results while many other analyses appear in supplementary material or are mentioned briefly.
Try to understand the full analytical landscape.
If 30 outcomes were tested and two produced p <.05, those two findings require different interpretation from a study with one prespecified primary outcome that produced the same p-value.
The number itself has not changed. The analytical context has.
Subgroup analysis requires testing the subgroup difference
Suppose an intervention is statistically significant among women but not among men.
Can the authors conclude that the intervention works differently by sex?
Not merely from those two significance labels.
The relevant question is whether the estimated intervention effects differ between the groups. That usually requires an interaction test or another direct comparison of the subgroup effects appropriate to the model.
Incorrect shortcut
Significant in Group A + not significant in Group B = groups differ.
Relevant question
Is the estimated effect in Group A statistically and substantively different from the estimated effect in Group B?
This is a classic example of why reading statistical results requires understanding the comparison actually being tested.
Post hoc subgroup findings deserve more caution
Subgroup analyses can reveal meaningful heterogeneity. They can also generate unstable findings, especially when many subgroups are examined after researchers have seen the data.
Ask whether the subgroup was prespecified, theoretically motivated, adequately represented, and directly tested.
An unexpected subgroup finding can be scientifically useful as a hypothesis for future research without deserving the same confidence as a prespecified primary analysis.
Dichotomizing continuous variables can change the analysis unnecessarily
Researchers sometimes transform continuous variables into categories.
Age becomes "young" versus "old." Test scores become "high" versus "low." Engagement becomes "engaged" versus "not engaged."
If a clinically or substantively meaningful threshold exists, categorization may be appropriate. Arbitrary cutoffs can discard information, reduce statistical efficiency, and create artificial distinctions between observations close to the threshold.
If a continuous variable was categorized, ask why and whether the cutoff was defined before examining the data.
Model assumptions matter, but you do not need to worship assumption checklists
Statistical models depend on assumptions. These can concern distributional form, functional relationships, variance, independence, censoring, missingness, measurement, model specification, or other features.
The consequences of assumption violations vary.
Some methods are reasonably robust to modest deviations. Other violations can seriously distort estimates or uncertainty.
Do not simply look for a sentence saying "all assumptions were met." Ask which assumptions are consequential to the estimate you care about and whether the authors provide evidence or sensitivity analyses relevant to them.
Methodological appraisal is not improved by converting statistical assumptions into another ritual checklist.
Linear models assume more than “the variables are numbers”
A linear regression coefficient has a particular interpretation under the model specification.
If the relationship between predictor and outcome is strongly nonlinear but the model assumes a simple linear relationship, the coefficient may provide a poor summary.
Researchers may examine residuals, transformations, nonlinear terms, splines, interactions, or alternative models when appropriate.
You do not need to inspect every diagnostic plot in every paper. But if a conclusion depends on a strong linear relationship, ask whether the assumed functional form is plausible.
Time-to-event data require methods that handle censoring
Some participants may not experience the event of interest before a study ends.
For example, researchers studying time until relapse cannot simply assign everyone who has not relapsed a fixed outcome equivalent to "no relapse." Their true event time may occur after follow-up ends.
Survival-analysis methods are designed to handle this censoring under relevant assumptions.
If the research question concerns when an event occurs, ask whether the analysis uses the timing information and handles incomplete follow-up appropriately.
Prediction models should be evaluated as prediction models
A paper may build a model to predict academic failure, disease, employee turnover, or another outcome.
If prediction is the goal, model performance on the same data used to build the model can be overly optimistic.
Appropriate evaluation may involve internal validation through resampling or cross-validation and, ideally, external validation in different data when the intended use requires generalization.
Relevant performance dimensions can include discrimination, calibration, prediction error, and decision usefulness depending on the task.
A model with several statistically significant predictors is not necessarily a good predictive model.
Explanatory models and predictive models should not be judged identically
An explanatory model may focus on estimating a relationship and testing theoretical expectations. A predictive model may prioritize accurate predictions for new observations.
The variables, evaluation criteria, validation procedures, and interpretation can differ.
A variable can be a useful predictor without being causal. Conversely, a theoretically important causal variable may contribute little incremental predictive accuracy.
Ask what the model was built to do before judging whether its analysis succeeded.
Complex models need enough information to estimate what they contain
Adding predictors, interactions, random effects, latent variables, nonlinear terms, or other parameters increases the amount of information the model must estimate.
A model can become too complex for the available data, leading to unstable estimates or overfitting.
This is especially important when outcomes are rare or when subgroup analyses involve only a small fraction of the total sample.
Do not let a large overall N distract you from asking how much information actually contributes to the parameters or events central to the model.
Model fit does not prove that the model is true
Statistical models may report fit indices, information criteria, likelihood statistics, pseudo-R
2 values, or other indicators.
Good fit means something specific within the modeling framework. It does not establish that the model is the only explanation, that causal directions are correct, that measurement is valid, or that predictions will generalize.
This is particularly important in structural equation modeling, where a model may fit the observed covariance structure reasonably well while alternative models could also fit.
Model fit contributes evidence. It is not a truth detector.
Qualitative analysis also needs to match the design
The analysis-design relationship is not exclusively statistical.
Qualitative studies may use thematic analysis, grounded theory, phenomenological analysis, discourse analysis, qualitative content analysis, narrative analysis, framework analysis, or other approaches.
These are not interchangeable labels for "finding themes."
Ask whether the analytical approach fits the research question, methodological orientation, type of data, and claims the researchers want to make.
If the study claims to use grounded theory, for example, examine whether the sampling and analytical procedures correspond meaningfully to that methodology rather than merely using open coding followed by the grounded-theory label.
Likewise, phenomenological inquiry should be evaluated according to what kind of experiential understanding the chosen approach seeks to produce.
This is why qualitative analysis should be evaluated using standards appropriate to qualitative methodology rather than by asking where the p-values went.
Coding frequency is not automatically qualitative importance
Qualitative researchers may report how often themes or codes occur. Frequency can sometimes provide useful descriptive context.
But an idea mentioned by fewer participants can still be analytically important, especially when the research aims to understand variation, mechanisms, exceptions, meanings, or theoretically significant experiences.
Do not assume that qualitative analysis is appropriate only when themes are ranked by how many participants mentioned them.
The analytical logic should follow the qualitative question and methodology.
Mixed-methods analysis requires integration, not merely two parallel analyses
A mixed-methods study may contain a competent statistical analysis and a competent qualitative analysis yet still provide little mixed-methods insight if the two components never meaningfully interact.
Ask what integration accomplishes.
Do qualitative findings explain an unexpected quantitative pattern? Does one component help develop measures or hypotheses for the other? Are results compared, merged, connected, or interpreted together?
If the central claim depends on integration, evaluate whether the mixed-methods analysis produces insight beyond simply placing two studies in the same paper.
Systematic reviews have their own analysis-design relationship
In a systematic review or meta-analysis, the "data" are results from other studies.
The analysis should therefore correspond to the review question, eligibility criteria, study designs, outcomes, effect measures, heterogeneity, and risk of bias among included studies.
Pooling studies statistically is not automatically appropriate simply because several papers report numerical results.
Studies may differ in populations, interventions, measures, comparators, follow-up periods, designs, and methodological quality.
A meta-analysis can calculate a precise pooled estimate from studies that should not meaningfully have been combined.
When reviewing synthesized evidence, the relevant question becomes whether the review's analytical synthesis is justified by the studies it combines.
Prespecified analysis deserves different confidence from data-driven discovery
Exploratory analysis is scientifically valuable. The problem is presenting exploratory findings as though they were confirmatory tests specified independently of the observed data.
Preregistration, protocols, and statistical analysis plans can help readers distinguish analyses planned before outcomes were examined from those developed afterward.
If the paper reports a surprising result, ask whether that analysis was prespecified.
A post hoc finding can still be interesting and real. It generally deserves more cautious interpretation and independent confirmation.
Sensitivity analyses help reveal whether conclusions depend on one analytical choice
Many analytical decisions have reasonable alternatives.
Researchers may vary assumptions about missing data, inclusion criteria, model specification, variable definitions, outliers, adjustment sets, or other choices to see whether the central result changes materially.
If the conclusion survives several defensible alternatives, confidence may increase.
If a result exists only under one narrow analytical specification and disappears under equally plausible alternatives, the finding is more fragile.
Sensitivity analysis does not prove that the preferred model is correct. It helps reveal how dependent the conclusion is on analytical choices.
The analysis should answer the question, not merely produce significance
A common analytical failure is choosing or interpreting methods around whether they generate p <.05.
Suppose a study asks whether two interventions produce clinically meaningful differences. The relevant analysis should estimate the difference and its uncertainty in relation to what counts as meaningful.
A statistically significant difference of negligible magnitude may not answer the substantive question affirmatively.
Likewise, a non-significant result with a wide confidence interval may not establish equivalence.
The analysis succeeds when it produces information capable of answering the research question, not when it produces a desirable significance label.
Ask whether the analysis supports the language in the conclusion
Analysis and interpretation are connected.
A regression coefficient showing association should not automatically become "X causes Y." A model developed in one dataset should not automatically become "the model accurately predicts future cases" without appropriate validation. A statistically significant interaction should not automatically become a practically important subgroup difference.
Once you understand what the analysis actually estimates, compare that quantity with the wording in the discussion.
This helps you recognize when authors are asking the analysis to support more than it actually establishes.