Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Makes a Comparison Group Appropriate?

An appropriate comparison group is not simply a group that looks similar to the intervention group. It must represent the alternative required by the research question and support a credible interpretation of the resulting contrast.

33
What Makes a Comparison Group Appropriate? Guide 33 of 217
01 · The Question

What Makes One Group a Credible Comparison for Another?

Suppose you want to evaluate a new teaching program. You have 200 students who received it and need another group for comparison.

Would students from another university work? Students from last year's cohort? Students who declined the program? Students from the same university who were not eligible? What about simply finding a group with approximately the same average age and proportion of women and men?

All of these groups could technically be compared with the intervention group. That does not make all of them appropriate comparators.

A comparison becomes scientifically useful when it represents the alternative relevant to your research question and allows differences in outcomes to be interpreted with as few plausible competing explanations as the design can reasonably achieve.

02 · The Short Answer

An Appropriate Comparison Group Must Represent the Right Alternative

In Brief

An appropriate comparison group represents the alternative condition required by the research question and is sufficiently comparable with the focal group that the intended contrast can be interpreted credibly.

Similarity on a few demographic characteristics is not enough. Researchers should consider how groups were formed, whether they differed before the exposure or intervention, whether outcomes were measured consistently, whether they were observed during comparable periods and settings, and which sources of confounding or selection remain.

03 · What You Need to Know

Comparability Is About the Counterfactual You Need, Not Superficial Resemblance

When researchers search for a comparison group, the natural instinct is often to look for people who resemble the intervention group.

Similarity matters, but that intuition is incomplete.

The first question is not “Who looks like my participants?” It is “What alternative condition does my research question require?”

Start by defining what the comparison is supposed to represent

Suppose you are evaluating a new academic-support program.

If your question is whether the program is better than receiving no additional support, you need a comparator representing no additional support. If the question is whether the program performs better than the university's existing support service, the relevant comparator is usual practice. If you want to compare two competing programs, the alternative program itself becomes the comparator.

These are not interchangeable research questions.

Comparator-selection guidance developed by an NIH expert panel emphasizes that comparator choice should follow from the research question because different comparator conditions support different interpretations.

This means an appropriate group cannot be identified independently of the claim you want to make.

The groups should be comparable on factors that could explain the outcome

Imagine that students voluntarily enroll in an optional academic-support program. Those who enroll may differ from those who do not in motivation, prior achievement, available time, academic difficulty, help-seeking behavior, or other characteristics.

If program participants later achieve higher grades, the outcome difference could reflect the program, those pre-existing characteristics, or both.

This is the central problem of confounding and selection in nonrandomized comparisons.

What matters is not whether the groups have identical characteristics overall. It is whether differences relevant to the outcome and treatment-selection process undermine the contrast you intend to interpret.

Baseline equivalence can provide useful evidence, but it is not a guarantee

In nonrandomized education evaluations, baseline equivalence receives particular attention. The What Works Clearinghouse assesses whether intervention and comparison groups are similar at baseline on key observable characteristics expected to influence outcomes.

That makes sense: if one group begins with substantially higher achievement, an outcome difference at the end of the study cannot simply be interpreted as emerging during the intervention period.

But baseline equivalence has limits.

Two groups can appear similar on measured variables while differing on variables that were not measured. Conversely, a baseline difference does not necessarily make a study unusable if the design and analysis can appropriately address it under defensible assumptions.

Baseline assessment should therefore be treated as evidence about comparability, not a certificate that the groups are exchangeable.

Random assignment changes how the comparison group is constructed

In a properly implemented randomized experiment, participants or other units are assigned to conditions by chance. Random assignment makes treatment assignment independent of baseline characteristics in expectation, providing a strong basis for comparing outcomes across conditions.

This does not guarantee that every measured characteristic will be numerically identical in a particular sample. Chance imbalances can occur.

The important point is that randomization creates the groups through a process designed to prevent systematic allocation based on participant characteristics. This is fundamentally different from selecting a nonrandomized comparison group because its observed characteristics happen to look similar.

Understanding what random assignment actually accomplishes therefore helps explain why randomized comparisons receive special methodological status.

In nonrandomized research, how people entered the groups matters enormously

Consider three potential comparison groups for students who voluntarily joined a tutoring program:

Potential comparison Why it might seem reasonable Potential problem
Students who chose not to participate They were present at the same institution during the same period The decision to participate may reflect motivation, need, schedules, prior achievement, or other outcome-related factors.
Students from another university A similar program is not available there Institutional context, admissions, curriculum, resources, assessment, and student composition may differ.
Students from the previous academic year They experienced the same institution before the program existed Policies, instructors, cohorts, assessments, external events, and secular trends may have changed over time.

None of these groups is automatically invalid. Each simply introduces different assumptions and possible alternative explanations.

An appropriate comparator is therefore one for which those assumptions are plausible enough for the intended inference and can be investigated or addressed as far as the design permits.

The comparison groups should experience comparable measurement

Even well-matched populations can produce misleading contrasts if outcomes are measured differently.

Suppose intervention students take a standardized assessment administered under supervised conditions, while comparison students' outcomes are extracted from course grades assigned by different instructors. The groups may resemble each other demographically, but the outcome does not have the same operational meaning.

Ideally, the outcome definition, measurement instrument, measurement schedule, follow-up period, and data-collection procedures should be sufficiently consistent to make the contrast interpretable.

Time matters when choosing a comparator

Historical comparison groups are attractive because the data may already exist. They can also be problematic.

Students enrolled three years earlier may have experienced different curricula, technologies, instructors, institutional policies, assessments, economic conditions, or educational disruptions. A difference between historical and current cohorts could therefore reflect calendar time rather than the intervention.

A concurrent comparator often avoids some of these temporal differences because both groups are observed over the same period, although concurrent observation does not solve every source of confounding.

Setting matters too

External comparison groups can introduce contextual differences.

A school implementing a digital-learning intervention might be compared with another school that appears similar in enrollment size and student demographics. Yet the schools may differ in teacher experience, leadership, internet infrastructure, curriculum implementation, local resources, or prior technology use.

Context is not background noise when it influences the outcome.

Matching can improve comparability on measured variables, but it does not create randomization

Researchers sometimes use matching or propensity-score methods to construct a comparison group that resembles the intervention group on observed covariates.

These approaches can be useful when their assumptions and implementation are appropriate. They can improve balance on variables included in the matching procedure.

They cannot generally guarantee balance on important characteristics that were never measured or adequately modeled.

Watch Out

A matched comparison group should not be described as equivalent to a randomized control group merely because measured baseline characteristics are balanced. Matching addresses comparability with respect to observed information used in the procedure; unmeasured confounding may remain.

Statistical adjustment cannot rescue every poor comparison

Regression adjustment, weighting, matching, stratification, and related methods can address particular forms of measured confounding under assumptions. They do not make every available comparison scientifically defensible.

If groups have little overlap in relevant characteristics, if important confounders were not measured, or if the comparator represents the wrong alternative altogether, increasingly elaborate analysis may not solve the underlying design problem.

This is why comparator selection should happen while designing the study, not after data collection when researchers discover that the available groups are difficult to compare.

Appropriateness depends on the claim

A group can be suitable for one research question and unsuitable for another.

Suppose students who voluntarily use an online tutoring system are compared with non-users. That comparison may adequately answer a descriptive question about whether outcomes differ between users and non-users.

It is much more demanding to interpret the same difference as the causal effect of tutoring.

This follows from the distinction between association and causation. The stronger the intended claim, the more consequential the assumptions behind the comparison become.

04 · A Practical Example

Choosing Among Three Plausible Comparison Groups

Hypothetical Example

Evaluating a first-year retention program

A university introduces an intensive mentoring program for first-year students identified as being at elevated risk of dropping out. The researchers want to estimate whether the program improves retention.

Option 1: All students who did not receive mentoring This group is easy to obtain, but many students were never eligible for the program because they had a much lower predicted dropout risk. Their expected retention would therefore differ even without mentoring.
Option 2: Eligible students who declined mentoring These students satisfy the same eligibility criteria, which improves comparability in one respect. However, accepting or declining mentoring may itself reflect motivation, schedules, perceived need, employment commitments, or other factors related to retention.
Option 3: Eligible students assigned to conditions by an appropriate random procedure If random assignment is feasible and ethically acceptable, this creates a stronger basis for estimating the program's causal effect because assignment is not systematically determined by the students' baseline characteristics.
The design decision The researchers should choose the comparator according to the causal question, feasibility, ethics, and assumptions they can defend. If randomization is impossible, the nonrandomized comparison requires explicit attention to selection, baseline differences, overlap, measurement, and residual confounding.

Notice that Option 2 might initially look like the “most similar” observational group. It is still not automatically exchangeable with participants because the decision to participate may contain exactly the information that matters for the outcome.

This is why selecting a comparator requires reasoning about how the data were generated, not merely comparing a demographic table after recruitment.

05 · What Researchers Often Get Wrong

Common Mistakes When Choosing Comparison Groups

Misconception

The Comparison Group Just Needs to Look Demographically Similar

Demographic similarity can be relevant, but it is not sufficient. Researchers need to consider characteristics related to treatment selection and the outcome, including prior values of the outcome when appropriate, as well as contextual and temporal differences.

Misconception

No Significant Baseline Differences Means the Groups Are Equivalent

Failure to detect a statistically significant baseline difference does not demonstrate equivalence. Significance tests depend partly on sample size and precision, and they do not address unmeasured characteristics. Comparability should be evaluated substantively and according to design-specific standards rather than inferred from a collection of nonsignificant p-values.

Misconception

Matching Makes a Comparison Group the Same as a Randomized Control Group

No. Matching can improve balance on measured characteristics used in the procedure. Random assignment operates through the assignment mechanism itself and balances baseline characteristics in expectation, including unmeasured ones. The inferential assumptions are therefore different.

Misconception

A Historical Group Is Appropriate Because It Came From the Same Institution

Institutional location does not eliminate changes over time. Cohort composition, policies, instructors, assessment, technologies, external events, and other conditions may differ between historical and current groups.

Misconception

Regression Can Correct Any Difference Between the Groups

Statistical adjustment depends on measured variables, sufficient overlap, appropriate models, and other assumptions. It cannot automatically recover a credible causal comparison when the groups are fundamentally incomparable or when important confounders are unavailable.

06 · What This Means for You

Choose the Comparator by Working Backward From the Claim

Before searching your dataset for a convenient second group, define the alternative condition your research question requires.

Then ask what would make observed outcomes under that alternative informative about what would have happened to the focal group under the same condition.

A simple decision framework

If the research question compares an intervention with usual practice
Select a comparator that genuinely represents usual practice in the relevant population and setting.
If groups cannot be randomized
Investigate how participants entered each group and identify characteristics that could influence both group membership and the outcome.
If the proposed groups differ substantially before the intervention
Determine whether those baseline differences threaten the intended comparison and whether the design and analysis can address them credibly.
If you plan to use historical or external controls
Examine differences in calendar time, setting, eligibility, measurement, follow-up, data quality, and contextual conditions.
If matching or adjustment is planned
Specify which variables are needed and why, assess overlap and balance appropriately, and remain explicit about the possibility of unmeasured confounding.

An appropriate comparison group is therefore not something you discover by sorting a spreadsheet until two columns look similar. It emerges from the research question, the data-generating process, and the assumptions required for the intended inference.

07 · A Quick Checklist

Before Accepting a Group as an Appropriate Comparator

Before using a comparison group, check:
Does the group represent the alternative condition required by your research question?
Were intervention and comparison participants drawn from populations that are sufficiently relevant to the intended contrast?
Do you understand how participants entered each group and what selection processes may have operated?
Have you examined important baseline characteristics, including prior outcome measures when appropriate?
Were outcomes defined and measured comparably across groups?
Were groups observed over comparable periods and under sufficiently comparable contextual conditions?
If matching or statistical adjustment is used, have you identified its assumptions and what sources of bias it cannot address?
Is there sufficient overlap between groups to support the comparisons you intend to make?
Would your interpretation remain defensible if unmeasured differences between nonrandomized groups exist?
08 · Frequently Asked Questions

Frequently Asked Questions About Choosing Comparison Groups

Do comparison groups need to have the same number of participants?

No. Equal group sizes are not required for groups to be scientifically comparable. Sample allocation affects precision and statistical efficiency, but similarity in sample size does not establish similarity in baseline characteristics or remove confounding.

Should the comparison group have the same demographics as the intervention group?

Relevant demographic characteristics may matter, but demographic similarity alone is insufficient. Researchers should focus on variables and selection processes that could affect the outcome or the probability of entering the intervention condition.

Can I use students from another university as my comparison group?

Possibly, but institutional differences can complicate the comparison. Examine differences in student populations, eligibility, curriculum, resources, assessment, teaching, policies, timing, and other contextual factors relevant to the outcome before deciding whether the external group is suitable.

Can last year's participants be my comparison group?

Historical comparison groups can sometimes be informative, but they introduce the possibility that changes over time rather than the intervention explain outcome differences. Consider changes in participants, policies, measurement, practice, external events, and other temporal factors.

Does matching eliminate selection bias?

Not necessarily. Matching can improve comparability on measured variables included in the procedure. Bias may remain because of unmeasured characteristics, measurement error, model specification, inadequate overlap, or other aspects of the selection process.

How do I know whether groups are equivalent at baseline?

Use criteria appropriate to your study design and field rather than relying only on whether baseline p-values exceed a significance threshold. Examine substantively important baseline measures, the magnitude of differences, balance, overlap, and the design process that generated the groups. Some evidence frameworks, such as the What Works Clearinghouse in education, specify formal baseline-equivalence standards for particular designs.

Can statistical adjustment fix baseline differences?

Adjustment can address certain measured differences under appropriate assumptions. It cannot guarantee that groups are comparable on unmeasured factors, and severe lack of overlap or an inappropriate comparator may be a design problem that modeling alone cannot adequately repair.

09 · The Bottom Line

The Best Comparison Group Is the One That Makes the Intended Contrast Credible

The Bottom Line

An appropriate comparison group represents the alternative condition required by your research question and provides a sufficiently credible basis for interpreting differences between groups.

Do not judge comparability from labels, equal sample sizes, or a few nonsignificant baseline tests. Examine how the groups were formed, what they experienced, when and where they were observed, how outcomes were measured, and which plausible alternative explanations remain.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes