03 · What You Need to Know
Comparability Is About the Counterfactual You Need, Not Superficial Resemblance
When researchers search for a comparison group, the natural instinct is often to look for people who resemble the intervention group.
Similarity matters, but that intuition is incomplete.
The first question is not “Who looks like my participants?” It is “What alternative condition does my research question require?”
Start by defining what the comparison is supposed to represent
Suppose you are evaluating a new academic-support program.
If your question is whether the program is better than receiving no additional support, you need a comparator representing no additional support. If the question is whether the program performs better than the university's existing support service, the relevant comparator is usual practice. If you want to compare two competing programs, the alternative program itself becomes the comparator.
These are not interchangeable research questions.
Comparator-selection guidance developed by an NIH expert panel emphasizes that comparator choice should follow from the research question because different comparator conditions support different interpretations.
This means an appropriate group cannot be identified independently of the claim you want to make.
The groups should be comparable on factors that could explain the outcome
Imagine that students voluntarily enroll in an optional academic-support program. Those who enroll may differ from those who do not in motivation, prior achievement, available time, academic difficulty, help-seeking behavior, or other characteristics.
If program participants later achieve higher grades, the outcome difference could reflect the program, those pre-existing characteristics, or both.
This is the central problem of confounding and selection in nonrandomized comparisons.
What matters is not whether the groups have identical characteristics overall. It is whether differences relevant to the outcome and treatment-selection process undermine the contrast you intend to interpret.
Baseline equivalence can provide useful evidence, but it is not a guarantee
In nonrandomized education evaluations, baseline equivalence receives particular attention. The What Works Clearinghouse assesses whether intervention and comparison groups are similar at baseline on key observable characteristics expected to influence outcomes.
That makes sense: if one group begins with substantially higher achievement, an outcome difference at the end of the study cannot simply be interpreted as emerging during the intervention period.
But baseline equivalence has limits.
Two groups can appear similar on measured variables while differing on variables that were not measured. Conversely, a baseline difference does not necessarily make a study unusable if the design and analysis can appropriately address it under defensible assumptions.
Baseline assessment should therefore be treated as evidence about comparability, not a certificate that the groups are exchangeable.
Random assignment changes how the comparison group is constructed
In a properly implemented randomized experiment, participants or other units are assigned to conditions by chance. Random assignment makes treatment assignment independent of baseline characteristics in expectation, providing a strong basis for comparing outcomes across conditions.
This does not guarantee that every measured characteristic will be numerically identical in a particular sample. Chance imbalances can occur.
The important point is that randomization creates the groups through a process designed to prevent systematic allocation based on participant characteristics. This is fundamentally different from selecting a nonrandomized comparison group because its observed characteristics happen to look similar.
Understanding what random assignment actually accomplishes therefore helps explain why randomized comparisons receive special methodological status.
In nonrandomized research, how people entered the groups matters enormously
Consider three potential comparison groups for students who voluntarily joined a tutoring program:
| Potential comparison |
Why it might seem reasonable |
Potential problem |
| Students who chose not to participate |
They were present at the same institution during the same period |
The decision to participate may reflect motivation, need, schedules, prior achievement, or other outcome-related factors. |
| Students from another university |
A similar program is not available there |
Institutional context, admissions, curriculum, resources, assessment, and student composition may differ. |
| Students from the previous academic year |
They experienced the same institution before the program existed |
Policies, instructors, cohorts, assessments, external events, and secular trends may have changed over time. |
None of these groups is automatically invalid. Each simply introduces different assumptions and possible alternative explanations.
An appropriate comparator is therefore one for which those assumptions are plausible enough for the intended inference and can be investigated or addressed as far as the design permits.
The comparison groups should experience comparable measurement
Even well-matched populations can produce misleading contrasts if outcomes are measured differently.
Suppose intervention students take a standardized assessment administered under supervised conditions, while comparison students' outcomes are extracted from course grades assigned by different instructors. The groups may resemble each other demographically, but the outcome does not have the same operational meaning.
Ideally, the outcome definition, measurement instrument, measurement schedule, follow-up period, and data-collection procedures should be sufficiently consistent to make the contrast interpretable.
Time matters when choosing a comparator
Historical comparison groups are attractive because the data may already exist. They can also be problematic.
Students enrolled three years earlier may have experienced different curricula, technologies, instructors, institutional policies, assessments, economic conditions, or educational disruptions. A difference between historical and current cohorts could therefore reflect calendar time rather than the intervention.
A concurrent comparator often avoids some of these temporal differences because both groups are observed over the same period, although concurrent observation does not solve every source of confounding.
Setting matters too
External comparison groups can introduce contextual differences.
A school implementing a digital-learning intervention might be compared with another school that appears similar in enrollment size and student demographics. Yet the schools may differ in teacher experience, leadership, internet infrastructure, curriculum implementation, local resources, or prior technology use.
Context is not background noise when it influences the outcome.
Matching can improve comparability on measured variables, but it does not create randomization
Researchers sometimes use matching or propensity-score methods to construct a comparison group that resembles the intervention group on observed covariates.
These approaches can be useful when their assumptions and implementation are appropriate. They can improve balance on variables included in the matching procedure.
They cannot generally guarantee balance on important characteristics that were never measured or adequately modeled.
Watch Out
A matched comparison group should not be described as equivalent to a randomized control group merely because measured baseline characteristics are balanced. Matching addresses comparability with respect to observed information used in the procedure; unmeasured confounding may remain.
Statistical adjustment cannot rescue every poor comparison
Regression adjustment, weighting, matching, stratification, and related methods can address particular forms of measured confounding under assumptions. They do not make every available comparison scientifically defensible.
If groups have little overlap in relevant characteristics, if important confounders were not measured, or if the comparator represents the wrong alternative altogether, increasingly elaborate analysis may not solve the underlying design problem.
This is why comparator selection should happen while designing the study, not after data collection when researchers discover that the available groups are difficult to compare.
Appropriateness depends on the claim
A group can be suitable for one research question and unsuitable for another.
Suppose students who voluntarily use an online tutoring system are compared with non-users. That comparison may adequately answer a descriptive question about whether outcomes differ between users and non-users.
It is much more demanding to interpret the same difference as the causal effect of tutoring.
This follows from the distinction between association and causation. The stronger the intended claim, the more consequential the assumptions behind the comparison become.
04 · A Practical Example
Choosing Among Three Plausible Comparison Groups
Hypothetical Example
Evaluating a first-year retention program
A university introduces an intensive mentoring program for first-year students identified as being at elevated risk of dropping out. The researchers want to estimate whether the program improves retention.
Option 1: All students who did not receive mentoring This group is easy to obtain, but many students were never eligible for the program because they had a much lower predicted dropout risk. Their expected retention would therefore differ even without mentoring.
Option 2: Eligible students who declined mentoring These students satisfy the same eligibility criteria, which improves comparability in one respect. However, accepting or declining mentoring may itself reflect motivation, schedules, perceived need, employment commitments, or other factors related to retention.
Option 3: Eligible students assigned to conditions by an appropriate random procedure If random assignment is feasible and ethically acceptable, this creates a stronger basis for estimating the program's causal effect because assignment is not systematically determined by the students' baseline characteristics.
The design decision The researchers should choose the comparator according to the causal question, feasibility, ethics, and assumptions they can defend. If randomization is impossible, the nonrandomized comparison requires explicit attention to selection, baseline differences, overlap, measurement, and residual confounding.
Notice that Option 2 might initially look like the “most similar” observational group. It is still not automatically exchangeable with participants because the decision to participate may contain exactly the information that matters for the outcome.
This is why selecting a comparator requires reasoning about how the data were generated, not merely comparing a demographic table after recruitment.