Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Between-Subjects vs. Within-Subjects Designs: Which Comparison Makes More Sense?

Between-subjects designs compare different participants across conditions, while within-subjects designs compare conditions within the same participants. The better choice depends on what is being studied and whether experiencing one condition can influence another.

37
Between-Subjects vs. Within-Subjects Designs Guide 37 of 217
01 · The Question

Should Different People Experience Each Condition, or Should the Same People Experience Both?

Suppose you want to compare two interfaces for an educational application. You could assign one group of students to interface A and another group to interface B.

Or you could ask every student to use both interfaces and compare each person's performance under the two conditions.

Both designs create a comparison. They do not create the same comparison.

The first is a between-subjects design. The second is a within-subjects design. Choosing between them affects what variation enters the comparison, how many observations each participant contributes, which biases become plausible, and how the data should be analyzed.

02 · The Short Answer

The Better Design Depends on Whether Participants Can Meaningfully Serve as Their Own Comparison

In Brief

A between-subjects design compares outcomes from different participants assigned or belonging to different conditions, whereas a within-subjects design compares outcomes from the same participants across two or more conditions or occasions.

Within-subjects designs can reduce variation attributable to stable differences between people and may improve statistical efficiency, but they can be vulnerable to order, learning, fatigue, period, and carryover effects. Between-subjects designs avoid many cross-condition carryover problems but require comparisons across different people.

03 · What You Need to Know

The Key Difference Is Where the Comparison Occurs

Researchers sometimes choose between these designs by asking which one requires fewer participants. That can matter, but it should not be the first question.

First ask whether the same participant can meaningfully experience all conditions without one condition changing the response to another.

Between-subjects designs compare different participants

In a between-subjects design, each participant typically contributes data under one study condition relevant to the comparison.

For example, 100 students might be assigned to one of two feedback systems:

  • 50 students use feedback system A;
  • 50 different students use feedback system B.

The outcome comparison occurs between the two groups.

In a randomized experiment, participants can be randomly assigned to those conditions. In observational research, naturally occurring groups can also be compared, although the causal implications are different because the groups were not created through random assignment.

Within-subjects designs compare participants with themselves

In a within-subjects design, the same participant contributes outcome information under multiple conditions or occasions.

Suppose every student completes one task using interface A and another using interface B. The researcher can compare each student's performance under A with that same student's performance under B.

This pairing can be valuable because stable participant characteristics are held constant within that comparison. A student's general ability, educational background, or other relatively stable characteristics do not differ between that student's own measurements in the way they can differ across two separate groups.

Crossover trials are a formal example. Cochrane describes a simple two-period crossover design in which participants receive sequences such as A then B or B then A, allowing each participant to act as their own control.

Between-subjects comparison Condition A is evaluated using one set of participants and condition B using another set, so the treatment contrast includes between-person variation.
Within-subjects comparison The same participants contribute measurements under multiple conditions or occasions, allowing comparisons within individuals.

Within-subjects designs can reduce between-person variability

People differ. Sometimes considerably.

If two different groups are compared, natural variation in ability, prior experience, motivation, physiology, preferences, or other characteristics contributes to variability in the outcome.

Within-subjects designs can remove stable between-person differences from the treatment contrast because each participant is compared with themselves. Cochrane notes that this can reduce between-participant variation and often allows crossover designs to achieve a given precision with fewer participants than an analogous parallel-group design.

That efficiency can be attractive, especially when participants are difficult to recruit.

It is not free.

The first condition can affect what happens in the second

This is the central difficulty of many within-subjects designs.

Suppose participants learn how to solve a task while using interface A. When they later use interface B, they may perform better simply because they have already practiced the task.

Or the first intervention may produce an effect that persists into the second period.

In crossover terminology, a carryover effect occurs when the effect of an intervention from one period persists into a subsequent period and interferes with the response to the later intervention.

If the effect is durable or irreversible, a crossover design may make little sense. You cannot meaningfully return a participant to their original state after an intervention that permanently changes the outcome.

Order effects are broader than carryover

Participants' responses may change simply because of the order in which conditions occur.

They may become more practiced, tired, bored, familiar with the equipment, anxious, or strategic as the study progresses. These changes can influence later measurements even when the substantive intervention itself has no lingering effect.

Researchers can sometimes address order through counterbalancing or randomizing condition sequence. In a simple two-condition design, for example, some participants might receive A then B while others receive B then A.

Counterbalancing does not make every order problem disappear. It distributes order systematically across conditions so that treatment effects can be separated more effectively from order when the design and analysis support that separation.

Period effects can occur when the background changes over time

Cochrane distinguishes carryover from period effects.

A period effect occurs when responses differ systematically between study periods for reasons other than the treatment being compared. Participants may change over time, the environment may change, or external conditions may differ between periods.

Imagine evaluating two learning applications during an academic semester. Students might naturally become more knowledgeable as the semester progresses. If every student uses application A first and B later, application B is confounded with being observed later in the course.

Randomizing or balancing sequence can help separate treatment from period in suitable designs, while the analysis must reflect the repeated structure.

A washout period can help only when the effect can actually wash out

Crossover studies sometimes place time between conditions so that the effects of the first intervention can dissipate before the next condition is evaluated.

This is commonly called a washout period.

The concept is sensible only when the carryover process is sufficiently understood and reversible. Waiting a week cannot erase a skill that participants learned permanently. Nor can a washout restore a participant to a pre-intervention state after an irreversible outcome.

The existence of a proposed washout period therefore does not automatically make a crossover design appropriate.

Some interventions simply cannot be studied meaningfully within subjects

Suppose you want to compare two semester-long teaching curricula. Once students have completed curriculum A, their knowledge has changed. Asking them to complete curriculum B afterward does not recreate the state they would have been in had they received B first.

A between-subjects design may therefore be much more defensible.

Similarly, interventions involving surgery, vaccination, permanent training, substantial learning, or other durable changes may be unsuitable for crossover designs because the first condition cannot be removed.

Cochrane specifically cautions against crossover designs when interventions can lead to permanent or long-term modification or when the condition changes substantially over time.

Some comparisons are naturally well suited to within-subject designs

Other research questions fit within-subject comparisons very naturally.

A researcher might compare how quickly the same participant completes short standardized tasks under two interface layouts, provided exposure to one does not meaningfully alter performance under the other. Researchers might compare responses to transient stimuli, alternative display formats, or short-lived conditions when carryover can reasonably be controlled.

The crucial issue is whether each participant can provide valid outcome information under each condition.

Pretest-posttest is within-subject, but it is not necessarily a crossover design

The terminology can become confusing here.

If the same participants are measured before and after an intervention, the observations are repeated within participants. The design therefore contains a within-subject temporal comparison.

But participants have not necessarily crossed between two alternative interventions.

A crossover design usually refers to participants receiving multiple treatment conditions in different periods or sequences. A simple pretest-posttest study measures change over time under an intervention but does not automatically provide the same causal comparison as a randomized crossover trial.

This distinction matters because a baseline measurement is not itself another treatment condition.

The statistical analysis must respect repeated observations

Measurements from the same participant are generally correlated. Treating them as if they came from unrelated individuals discards the pairing built into the design and can produce incorrect uncertainty estimates.

For simple situations, paired analyses may be appropriate. More complex repeated-measures, longitudinal, or crossover designs may require models that account explicitly for within-participant correlation, treatment sequence, period, clustering, missing data, or other design features.

The correct analysis depends on the research question, outcome, number of conditions and measurements, randomization structure, and assumptions.

Consideration Between-subjects design Within-subjects design
Who experiences each condition? Different participants The same participants experience multiple conditions or occasions
Stable between-person differences Contribute to group variation Can be controlled through within-person comparison
Carryover between conditions Usually not a cross-condition problem for an individual Can be a major threat
Order and practice effects Usually less central across separate participants Often require explicit design attention
Participant burden Each participant may experience fewer conditions Each participant may complete more conditions or sessions
Analysis Must reflect independent or clustered group structure as appropriate Must account for repeated or paired observations

The choice is not simply about which design is “stronger”

Neither design dominates universally.

A within-subject design can be exceptionally efficient for a reversible, short-lived intervention with stable participants and manageable order effects. The same design can be indefensible for an intervention whose effects persist.

A between-subjects design avoids exposing the same participant to competing conditions but may require more participants or greater precision to separate the treatment effect from natural between-person variation.

The better design is the one whose comparison matches the phenomenon being studied.

04 · A Practical Example

Choosing a Design for Two Educational Interfaces

Hypothetical Example

Comparing two interfaces for locating scholarly sources

A researcher develops two interfaces for a literature-search system and wants to determine which allows postgraduate students to complete standardized search tasks more efficiently.

Between-subjects option Sixty students are randomly assigned to interface A and 60 different students to interface B. Each student completes the same type of search task using only the assigned interface. Performance is compared across groups.
Within-subjects option The same 60 students use both interfaces on different but comparable task sets. Half use A first and B second, while the other half use B first and A second. Each student's performance under A can be compared with that student's performance under B.
Why within-subjects might help Search skill varies substantially among students. Having each student use both interfaces reduces the influence of stable between-student differences on the interface comparison.
Why within-subjects might hurt Completing the first set of searches may teach participants strategies they use in the second condition. Fatigue or increasing familiarity with the study procedure may also affect later performance. The researcher must decide whether sequence balancing and suitable tasks can manage these effects sufficiently.

Now change the intervention. Suppose the researcher wants to compare two semester-long information-literacy curricula.

A crossover becomes far less attractive. Completing the first curriculum is intended to change students' knowledge permanently. That learning would carry directly into the second period.

The same general research topic therefore supports different designs depending on what the intervention actually does.

05 · What Researchers Often Get Wrong

Common Mistakes When Choosing Between- and Within-Subjects Designs

Misconception

Within-Subjects Designs Are Always Better Because Each Person Is Their Own Control

Within-person comparison can reduce between-person variability, but the advantage can be outweighed by carryover, learning, fatigue, period effects, or an intervention that permanently changes participants. Suitability depends on the phenomenon and timing.

Misconception

Between-Subjects Designs Are Automatically Weaker

No. Separate groups are often necessary when interventions have lasting effects, conditions cannot be repeated, or exposure to one condition would compromise another. Random assignment can provide a strong causal basis for between-group comparison.

Misconception

Counterbalancing Eliminates All Order Effects

Counterbalancing can distribute order across conditions and help separate treatment from sequence or period effects. It cannot make an irreversible intervention reversible or guarantee that complex carryover processes disappear.

Misconception

A Washout Period Always Solves Carryover

No. Washout is useful only when enough time can reasonably remove the relevant residual effect before the next outcome measurement. Durable learning, permanent treatment effects, and irreversible outcomes cannot simply be washed out by waiting.

Misconception

Pretest-Posttest and Crossover Mean the Same Thing

No. Both involve repeated observations, but a crossover design exposes participants to multiple intervention conditions across periods. A simple pretest-posttest study can measure the same participants twice while exposing them to only one intervention.

Misconception

You Can Analyze Repeated Measurements as If They Came From Different People

Measurements from the same participant are correlated. Analyses should account for the paired or repeated structure. Ignoring that dependence can waste information and produce inappropriate estimates of uncertainty.

06 · What This Means for You

Ask Whether the First Condition Changes the Participant for the Second

If you are deciding between these designs, begin with the intervention rather than the sample-size calculation.

Can participants experience both conditions while still providing a meaningful comparison? If the answer is no, the apparent efficiency of a within-subject design is beside the point.

A simple decision framework

If the intervention has a durable or irreversible effect
Prefer a between-subjects comparison unless a specialized design provides a defensible alternative.
If the condition is temporary and participants can return sufficiently to an appropriate state before another condition
A within-subject or crossover design may be efficient if order, period, and carryover effects can be managed.
If stable differences between participants create substantial outcome variability
A within-subject design may gain precision by comparing each participant with themselves, provided its other assumptions are plausible.
If performing one condition teaches participants how to perform better in another
Consider whether counterbalancing and alternative tasks can adequately address learning; otherwise a between-subjects design may provide the cleaner comparison.
If each participant would face excessive burden from completing every condition
Consider a between-subjects or other design that reduces repeated exposure while preserving the comparison required by the research question.
Watch Out

Do not choose a within-subject design merely because it appears to require fewer participants. If the first condition changes how participants respond to the second, the resulting comparison may be more precise statistically while answering the scientific question less credibly.

Whichever design you choose, first ensure that the research question actually requires the comparison you are building. Design efficiency matters only after the intended contrast is clear.

07 · A Quick Checklist

Before Choosing Between- or Within-Subjects Comparison

Before finalizing the design, check:
Can the same participant meaningfully experience every condition required by the research question?
Could the effect of one condition persist into a later condition?
Could practice, learning, fatigue, boredom, or familiarity influence later measurements?
Could outcomes change systematically over study periods for reasons unrelated to the interventions?
If sequence matters, have you planned an appropriate randomization or counterbalancing strategy?
If a washout period is proposed, is there a defensible reason to believe the relevant effects will actually dissipate?
Have you considered participant burden, study duration, attrition, and the feasibility of completing multiple conditions?
Does the planned analysis account for paired or repeated observations when the same participants contribute multiple measurements?
Does the chosen design produce the comparison your substantive research question actually requires?
08 · Frequently Asked Questions

Frequently Asked Questions About Between- and Within-Subjects Designs

Is repeated measures the same as within-subjects?

The terms overlap substantially because repeated-measures designs collect multiple observations from the same units. However, repeated measurements can refer broadly to outcomes collected over multiple times, while “within-subjects” often emphasizes comparisons of conditions or occasions within the same participant. Describe the actual measurement and treatment structure rather than relying on the label alone.

Is a crossover design a within-subjects design?

Yes. In a crossover study, participants receive multiple intervention conditions in different periods or sequences, creating within-participant treatment comparisons. The design requires particular attention to carryover, period, sequence, and repeated-measurement issues.

Is pretest-posttest a within-subjects design?

When the same participants provide both pretest and posttest measurements, the temporal comparison is within subjects. It is not necessarily a crossover design because participants may receive only one intervention rather than multiple alternative treatments.

Do within-subjects designs always need fewer participants?

They can be more statistically efficient when within-participant measurements are strongly correlated and the design is appropriate, but required sample size depends on the expected effect, variability, within-participant correlation, attrition, analysis, number of conditions, and other design features. A formal sample-size calculation should reflect the actual design.

What is a carryover effect?

Carryover occurs when the effect of an earlier intervention persists into a later study period and influences the response observed under the subsequent condition. It is a central concern in crossover designs.

What is a period effect?

A period effect is a systematic difference in outcomes between study periods that is not caused by the intervention being compared. It may arise because participants, environments, background practices, or other conditions change over time.

Can counterbalancing solve learning effects?

It can help prevent treatment condition from being perfectly confounded with order by varying the sequence across participants. Whether that is sufficient depends on the nature and persistence of learning. If exposure fundamentally changes participants, a within-subject comparison may remain inappropriate.

Which design is better for educational interventions?

Neither is universally better. Short-lived conditions such as interface presentations or transient task manipulations may suit within-subject comparisons. Educational interventions intended to produce durable learning may be better suited to between-subjects designs because learning from the first condition cannot simply be removed before the second.

09 · The Bottom Line

Choose the Design That Preserves a Meaningful Comparison

The Bottom Line

Use a between-subjects design when different participants should provide the alternative conditions; consider a within-subjects design when the same participants can validly experience multiple conditions and serve as their own comparison.

Within-subject designs can reduce between-person variability, but their efficiency is useful only when carryover, learning, fatigue, order, and period effects do not undermine the contrast. Start with what the intervention does to participants, then choose the design.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes