03 · What You Need to Know
Baseline Differences Mean Different Things in Different Designs
Baseline refers to the relevant point before the intervention, treatment, or exposure being evaluated. Researchers often record characteristics at this point so they can describe the groups and understand their starting conditions.
These characteristics might include age, prior achievement, disease severity, previous experience, socioeconomic characteristics, an earlier measure of the study outcome, or other variables relevant to the research question.
The important issue is not whether every number is identical across groups. It is why differences exist and whether they threaten the intended interpretation.
Why pre-existing differences can distort a comparison
Suppose students receiving a new mathematics intervention score an average of 85 on the final assessment, while students in the comparison group score 78.
A seven-point difference appears to favor the intervention.
But suppose the intervention group averaged 80 before the program and the comparison group averaged 70. The intervention group was already performing better.
The final difference cannot simply be treated as seven points created by the intervention. The groups began from different positions, and that starting difference must be incorporated into the interpretation and, where appropriate, the analysis.
The problem becomes even more consequential when baseline characteristics influence both which condition participants enter and the outcome. This can create confounding in a nonrandomized comparison.
Baseline differences are especially consequential in nonrandomized studies
When participants are not randomly assigned, the mechanism that placed them into groups can create systematic differences.
Students may choose whether to join an optional tutoring program. Physicians may determine which patients receive a treatment. Schools may decide whether to adopt an educational intervention. Institutions with greater resources may be more likely to implement a new technology.
In each case, intervention status may be connected to characteristics that also affect the outcome.
This is why establishing baseline comparability receives considerable attention in quasi-experimental research. The What Works Clearinghouse, for example, evaluates baseline equivalence between intervention and comparison groups for particular nonrandomized education studies.
If you are still selecting the comparator, these issues are part of determining whether a comparison group is appropriate in the first place.
Randomized groups can also differ at baseline
A common misunderstanding is that successful random assignment should produce groups with identical baseline characteristics.
It does not.
Random assignment uses chance to allocate participants or other units to study conditions. It prevents systematic assignment based on baseline characteristics when properly implemented, but any particular random allocation can still produce numerical differences between groups.
Imagine repeatedly flipping a fair coin 20 times. You would not expect exactly 10 heads every time. The same logic applies to random allocation. Particularly in smaller samples, chance may produce noticeable imbalances in some characteristics.
The crucial distinction is that, in a properly randomized study, such baseline differences arise through chance rather than a systematic process that assigned certain kinds of participants to one condition.
Randomized and nonrandomized baseline differences should not be interpreted in the same way
| Issue |
Properly randomized study |
Nonrandomized study |
| Why might groups differ at baseline? |
Chance imbalance can occur after random allocation. |
Differences may reflect selection, treatment decisions, eligibility, self-selection, institutional processes, or other systematic mechanisms. |
| What created the groups? |
A random allocation mechanism |
A nonrandom mechanism |
| What does baseline similarity establish? |
It describes the realized groups but is not what makes the allocation randomized. |
It can provide evidence about measured comparability but does not establish equivalence on unmeasured characteristics. |
| What is the main concern? |
Chance imbalance, precision, prespecified adjustment, attrition, or problems with randomization |
Confounding and selection processes that may explain outcome differences |
This distinction is fundamental. The credibility of randomization comes from the assignment procedure, not from obtaining visually identical baseline tables afterward.
A nonsignificant baseline test does not prove that groups are equivalent
Researchers sometimes compare every baseline variable using hypothesis tests and conclude that groups are “equivalent” because none of the p-values falls below.05.
That reasoning is problematic.
A nonsignificant test means that the chosen statistical test did not reject its null hypothesis at the selected threshold. It does not establish that the groups are substantively identical or sufficiently comparable for every purpose.
With small samples, substantial differences may fail to reach statistical significance because estimates are imprecise. With very large samples, relatively small differences can become statistically significant.
For randomized trials, significance testing of baseline differences is particularly difficult to justify because randomization itself tells us that any baseline differences were generated by chance if the allocation was properly implemented. Baseline characteristics are still worth reporting, but the question is not whether randomization “worked” because all p-values exceed.05.
Watch Out
Do not use “p >.05 at baseline” as a universal certificate that two groups are equivalent. Examine the size, direction, and substantive importance of relevant differences, and interpret them in light of how the groups were created.
Not every baseline characteristic matters equally
A difference becomes particularly relevant when the characteristic is strongly related to the outcome, treatment assignment, or both.
Suppose two groups differ slightly in average height in a study of academic-writing performance. Unless height is relevant to the study's causal structure or outcome, that imbalance may have little substantive consequence.
A difference in prior writing proficiency is another matter. Because previous proficiency is likely to predict later proficiency, failing to account appropriately for a substantial baseline difference could distort the interpretation of post-intervention scores.
Researchers should therefore identify important baseline variables using substantive knowledge and the study's causal logic rather than treating every available characteristic as equally consequential.
The baseline value of the outcome can be especially informative
When the outcome can be measured before an intervention, its baseline value often provides important information about participants' starting position.
In an educational intervention measuring mathematics achievement, prior mathematics achievement may strongly predict later performance. In a health intervention targeting blood pressure, baseline blood pressure may be directly relevant to subsequent measurements.
This is one reason a baseline measurement can be valuable even in a randomized study. It can improve description and, when incorporated appropriately into a prespecified analysis, may improve precision.
Adjusting for baseline differences can help, but it is not a universal repair
Statistical methods can account for measured baseline variables in various ways. Depending on the design and estimand, researchers may use regression adjustment, analysis of covariance, stratification, weighting, matching, or other approaches.
These methods do not all solve the same problem, and their validity depends on assumptions.
In nonrandomized studies, adjustment can address measured confounding when the relevant variables have been identified and adequately measured and modeled. It cannot guarantee that important unmeasured confounding has disappeared.
If intervention and comparison groups have almost no overlap on important characteristics, statistical modeling may also require extrapolation into regions where little or no comparable information exists.
Design problems are often easier to prevent than to repair.
Attrition can create a new comparability problem after baseline
Even groups created through random assignment can become harder to compare if participants leave the study differently across conditions.
Suppose students with low achievement disproportionately withdraw from one condition. The participants remaining at the end may no longer reflect the original randomized comparison in a straightforward way.
This is why evidence standards such as those of the What Works Clearinghouse consider attrition when evaluating randomized studies. Baseline equivalence may become relevant when attrition is high or there are concerns about the integrity of randomization.
Baseline differences do not automatically tell you whether an outcome difference is causal
The presence or absence of obvious baseline differences is only one part of causal reasoning.
Groups that appear similar on measured characteristics can still differ on unmeasured factors. Groups that differ on some measured characteristics may still support useful inference if the design and analysis appropriately address those differences under defensible assumptions.
The deeper question remains whether the study can distinguish the exposure or intervention from plausible competing explanations. This is why researchers should avoid treating a baseline table as a substitute for understanding when a design supports association rather than causation.
04 · A Practical Example
When the Final Scores Hide Where the Groups Started
Hypothetical Example
Evaluating an academic-writing program
A university offers an optional eight-week academic-writing program. Eighty students participate, while another 80 eligible students do not. At the end of the semester, both groups complete the same writing assessment.
Post-intervention result Program participants average 84 points. Nonparticipants average 78. Looking only at the final scores suggests a six-point advantage for participants.
Baseline information Before the program, participants averaged 79 and nonparticipants averaged 72. Participants therefore entered the study with substantially higher writing performance.
The interpretation problem Because students chose whether to participate, baseline proficiency may be part of a broader selection process involving motivation, academic preparation, help-seeking, schedules, or other characteristics. The six-point final difference cannot simply be attributed to the program.
What the researchers should do They should analyze the data using a method appropriate to the design and research question, account for relevant measured baseline variables where justified, and acknowledge that residual confounding may remain. The strongest remedy would have been to address group formation during study design rather than relying solely on post hoc statistical correction.
Notice something else: subtracting 72 from 78 and 79 from 84 gives changes of six and five points respectively. That comparison is more informative than looking only at final scores, but it still does not automatically identify a one-point causal effect. Because participation was self-selected, other differences between the groups may influence their trajectories.
Baseline data reveal the problem. They do not, by themselves, solve it.