Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Does the Research Question Assume a Comparison That Is Not Actually Meaningful?

Two groups can be easy to compare statistically while making little sense as a scientific comparison. A meaningful research question requires a comparator that helps answer the substantive question rather than merely providing a second group.

335
Is Your Research Comparison Meaningful? Guide 335 of 533
01 · The Question

Just Because Two Groups Can Be Compared, Should They Be?

Researchers frequently build questions around contrasts: users versus non-users, online versus face-to-face students, public versus private institutions, before versus after an intervention, younger versus older participants, high versus low performers.

Statistically, comparing two groups may be straightforward. Scientifically, the comparison can be much harder to justify.

The crucial question is not simply whether the groups differ. It is what their difference would mean. If the comparison groups differ in several important ways, represent fundamentally different populations, experience different contexts, or fail to represent the alternative condition implied by the research question, even an accurately estimated difference may provide a weak answer.

02 · The Short Answer

A Useful Comparator Must Help Answer the Actual Question

In Brief

A comparison in a research question is meaningful when the groups, conditions, exposures, interventions, or periods being contrasted correspond to the substantive difference the researcher wants to understand and are sufficiently interpretable for that contrast to support the intended conclusion.

Meaningful does not mean that the groups must be identical. Some differences are the very reason for the comparison. The problem arises when additional differences, ambiguous reference conditions, selection processes, or incompatible contexts make it unclear what the observed contrast actually represents.

03 · What You Need to Know

What Makes One Comparison More Informative Than Another?

Comparator selection is part of the substantive design of a study, not merely an analytical convenience. In comparative-effectiveness research, methodological guidance explicitly links the choice of comparator to the research question and emphasizes that comparison groups should represent meaningful alternatives. The same reasoning is useful well beyond health research.

Before asking whether Group A differs from Group B, ask why Group B is the appropriate reference for understanding Group A.

Begin With the Contrast You Actually Care About

Suppose a researcher asks: “Do students who use generative AI perform better than students who do not?”

What is the substantive contrast? Any AI use versus absolutely no use? Regular use versus no use? AI-supported study versus conventional study strategies? Permitted AI assistance versus an otherwise equivalent task completed without AI?

Those comparisons answer different questions. “Non-users” may seem like the obvious comparator, but it may be a heterogeneous group containing students who avoid AI for very different reasons and use many alternative forms of assistance.

A meaningful comparison begins by defining what alternative condition would make the observed difference informative.

The Comparator Defines the Meaning of the Effect or Difference

Consider a study of a new teaching approach. Comparing it with no instruction asks a different question from comparing it with standard instruction. Comparing it with another active teaching approach asks yet another question.

The intervention itself has not changed, but the comparison has. Consequently, the meaning of “better,” “effective,” or “different” changes as well.

This is particularly important for causal questions. A causal effect is inherently comparative: it concerns outcomes under one condition relative to outcomes under an alternative condition. A poorly specified comparator therefore means a poorly specified causal effect.

Existing Groups May Differ for Reasons Central to the Outcome

Imagine comparing students who voluntarily attend optional tutoring with students who do not. If tutoring attendees obtain higher grades, the difference might reflect tutoring, but it might also reflect motivation, prior difficulties, available time, help-seeking behavior, instructor recommendations, or other factors associated with both attendance and performance.

The comparison is not meaningless simply because confounding exists. Observational research routinely addresses imperfect comparability. The issue is whether the comparison, together with the design and available evidence, can support the intended interpretation.

If the question asks what tutoring causes, this becomes part of the broader need to identify a causal commitment hidden in the research question.

“Users Versus Non-Users” Is Often Less Simple Than It Looks

Technology studies frequently compare users with non-users. Yet adoption is rarely random. People who choose a technology may differ from non-users in digital competence, resources, motivation, prior performance, attitudes, task requirements, institutional support, or opportunity to use the technology.

The meaning of “use” can also vary. A student who asks an AI system to clarify one concept and another who delegates entire assignments to it may both be classified as users.

Before accepting a user/non-user contrast, define both conditions and ask whether the resulting comparison corresponds to the phenomenon you actually want to investigate.

A Before-and-After Comparison Changes More Than Time

Researchers sometimes compare outcomes before and after a policy, technology, curriculum, or institutional intervention and interpret the difference as its impact.

The difficulty is that time itself carries other changes. Participants may mature, external events may occur, institutional practices may shift, measurement may change, and broader trends may affect the outcome. A pre-post contrast can be useful, but “before” does not automatically provide the counterfactual for what would have happened “after” without the intervention.

If the research question is causal, the study needs a defensible strategy for distinguishing the intervention-related change from other temporal changes.

Historical Comparators May Come From a Different World

Comparing a contemporary group with participants or records from several years earlier can sometimes be informative. It can also introduce differences in technology, policy, curriculum, assessment, demographics, economic conditions, professional practice, or data quality.

The question is whether those historical differences interfere with the interpretation of the contrast. A convenient archive is not automatically a suitable comparison condition.

Groups Defined by an Outcome Can Create Circular Comparisons

Suppose researchers divide students into “successful” and “unsuccessful” groups based on grades and then ask whether successful students have higher grades. The answer is built into the grouping rule.

Less obvious versions occur when group definitions contain a variable closely tied to the outcome being compared. Researchers should inspect how categories were created and ensure that the comparison does not merely rediscover the criterion used to define them.

Arbitrary Categories Can Destroy Useful Information

Researchers sometimes convert continuous variables into “high” and “low” groups using a median split or another convenient threshold, then formulate the research question around those categories.

This may be justified when the threshold has substantive meaning, but arbitrary categorization can discard information, reduce statistical power, and create groups that do not correspond to meaningful distinctions in the phenomenon.

Before asking whether “high AI users” differ from “low AI users,” ask what makes those categories substantively different and whether treating use as a continuous or more nuanced exposure would better represent the research problem.

Comparison Requires Comparable Measurement

Even well-chosen groups may be difficult to compare if the outcome does not have sufficiently comparable meaning across them.

For example, comparing course grades across two institutions can be problematic if grading standards, assessments, curricula, or grade distributions differ substantially. The numerical values may share the same scale while representing different academic performances.

This links comparator selection to measurement. Researchers should check whether the variables can be measured in a way that supports the intended comparison.

Different Populations May Answer a Different Question

Suppose one group consists of first-year engineering students from a highly selective university and another consists of graduating business students from an open-admission institution. Comparing their AI use may yield a numerical difference, but attributing that difference to discipline would be difficult because discipline is entangled with year level, institution, admissions context, and potentially many other characteristics.

A comparison does not become scientifically meaningful merely because both groups can be placed in the same table.

Meaningful Does Not Mean Perfectly Matched

Researchers sometimes conclude that groups must be identical on every characteristic except the focal variable. In many real studies, particularly observational research, that is neither possible nor conceptually necessary.

The relevant standard is whether the comparison supports the intended inference under defensible design and analytical assumptions. Some baseline differences can be measured and addressed. Others may be so fundamental that the comparison ceases to illuminate the question.

Statistically possible comparison Two groups have data on the same nominal outcome and can be entered into an analysis.
Scientifically meaningful comparison The contrast between those groups represents an interpretable difference relevant to the research question, with other consequential differences appropriately handled or acknowledged.
04 · A Practical Example

Why “AI Users Versus Non-Users” May Be the Wrong Comparison

Hypothetical Example

Comparing academic performance by generative AI use

A researcher asks: “Does generative AI use improve academic performance among undergraduate students?” The plan is to compare the final grades of students who report using generative AI with those who report no use.

Clarify the intended contrast The researcher wants to know whether using generative AI for coursework changes academic performance compared with completing comparable academic work without generative AI assistance.
Inspect the observed groups Students chose whether to use AI. Users and non-users may differ in prior achievement, digital competence, course demands, attitudes toward AI, access, motivation, or study strategies.
Inspect exposure heterogeneity The “AI user” group combines students who use AI for brainstorming, explanations, proofreading, coding, answer generation, and other activities with potentially different educational consequences.
Inspect the outcome Final grades may come from different courses with different assessment and grading practices, weakening the meaning of a direct comparison.
Redesign the contrast Depending on the scientific objective, the researcher may need a more specific exposure, a more comparable reference condition, a common outcome measure, and a design capable of addressing important pre-existing differences.

The original comparison could still support a descriptive question about observed differences between self-selected users and non-users. What it cannot do automatically is answer the stronger causal question implied by “improve.”

05 · What Researchers Often Get Wrong

Common Mistakes When Choosing What to Compare

Misconception

Every Study Needs a Control Group

No. Whether a comparison group is necessary depends on the research question and design. Descriptive, exploratory, qualitative, methodological, and other studies may address important questions without a conventional control group.

Misconception

The Most Convenient Available Group Is a Suitable Comparator

Convenience does not establish conceptual relevance. The comparator should represent an alternative condition that makes the observed contrast informative for the research question.

Misconception

If Two Groups Differ Significantly, the Comparison Was Meaningful

Statistical significance concerns the observed statistical contrast under a model. It does not establish that the groups were substantively comparable or that their difference answers the intended scientific question.

Misconception

Matching Two Groups Makes Them Equivalent

Matching can improve comparability on measured characteristics used in the matching procedure. It does not guarantee equivalence on unmeasured characteristics or automatically justify every causal interpretation.

Misconception

Before-and-After Data Automatically Reveal an Intervention's Effect

A temporal change may reflect the intervention, but other events, maturation, trends, measurement changes, or contextual shifts can also contribute. The appropriate interpretation depends on the design and assumptions.

Misconception

More Comparison Groups Always Produce a Stronger Study

Additional groups are useful only when each addresses a meaningful contrast. Adding weakly justified comparisons can increase complexity, multiplicity, and interpretive ambiguity without improving the answer.

06 · What This Means for You

Choose the Comparator by Working Backward From the Difference You Want to Interpret

Write down the contrast implied by your question in plain language: “I want to know what differs when ___ is present rather than ___.” Then examine whether the proposed groups genuinely represent those two conditions.

Next, list other ways the groups differ that could affect the outcome or alter the meaning of the comparison. The purpose is not to demand impossible perfection. It is to determine whether the comparison can answer the question you intend to ask.

A simple decision framework

If the comparator represents the substantive alternative in your question
Evaluate whether the design can handle other consequential differences between the groups.
If the comparator was chosen mainly because data are available
Ask what question that comparison actually answers before retaining it.
If the groups differ systematically on factors closely related to the outcome
Strengthen the design, measure and address relevant differences, reconsider the comparator, or narrow the inference.
If the groups use different measures or the outcome has different meanings across them
Establish measurement comparability before interpreting the numerical contrast.
If no available comparator represents the contrast your question requires
Change the design or revise the question rather than forcing an uninformative comparison.
Watch Out

Do not let statistical software decide the comparison for you. A difference between two available groups can be calculated almost instantly. Determining what that difference means is the research-design problem.

07 · A Quick Checklist

Check Whether Your Proposed Comparison Makes Scientific Sense

Before building a question around two or more groups, check:
State the exact substantive contrast you want the comparison to represent.
Explain why the proposed comparator is the relevant alternative rather than merely an available group.
Define exposure, intervention, group, or condition membership clearly enough that the contrast is interpretable.
Identify pre-existing differences between groups that could affect the outcome or interpretation.
Check whether both groups' outcomes are measured in sufficiently comparable ways.
Avoid categories created from arbitrary thresholds unless those thresholds have a defensible substantive purpose.
For before-and-after comparisons, identify other changes over time that could explain the observed difference.
For causal questions, determine whether the comparator represents the alternative condition needed for the causal contrast.
Revise the comparison or question if the resulting difference would be difficult to interpret even if estimated perfectly.
08 · Frequently Asked Questions

Questions About Comparison Groups in Research

What makes a good comparison group?

A useful comparison group represents a substantively relevant alternative to the focal condition and permits an interpretable contrast for the research question. For causal questions, researchers must also consider whether differences between groups threaten identification of the causal effect.

Do comparison groups need to be identical at baseline?

No. Exact equality is neither expected nor always necessary. What matters is whether differences that threaten the intended inference are adequately addressed through design, measurement, analysis, and appropriately qualified interpretation.

Is a control group the same as a comparison group?

“Comparison group” is the broader term for a group against which another group or condition is contrasted. “Control group” is often used for a specific reference condition in experimental or quasi-experimental research, although terminology varies across disciplines and designs.

Can I compare public and private university students?

Yes, if that comparison addresses a meaningful research question. However, public and private university students may differ in characteristics beyond institution type. Your design and interpretation should account for the factors relevant to the particular conclusion you intend to draw.

Can I compare two groups if their sample sizes are different?

Unequal sample sizes do not automatically make a comparison invalid. Their consequences depend on the analytical method, variability, design, precision, and other assumptions. Conceptual comparability should be evaluated separately from numerical balance.

Is comparing users and non-users enough to estimate a technology's effect?

Not automatically. Users and non-users may self-select into those conditions and differ in characteristics related to the outcome. Estimating a causal effect requires a design and assumptions capable of addressing those differences and defining the relevant alternative condition.

Can I compare the same group before and after an intervention?

Yes, and such designs can answer useful questions. If the goal is causal inference, however, the researcher must consider whether other changes over time could account for the observed difference rather than attributing every change to the intervention.

09 · The Bottom Line

The Right Comparator Is the One That Makes the Difference Interpretable

The Bottom Line

A meaningful research comparison is not simply a contrast between two available groups; it is a contrast in which the comparator represents a relevant alternative and the resulting difference can reasonably inform the research question.

Choose the comparison by asking what difference you actually want to interpret. Then examine how the groups were formed, what else differs between them, whether the outcome has comparable meaning, and whether the design supports the strength of conclusion you intend to draw.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes