03 · What You Need to Know
The Key Difference Is Where the Comparison Occurs
Researchers sometimes choose between these designs by asking which one requires fewer participants. That can matter, but it should not be the first question.
First ask whether the same participant can meaningfully experience all conditions without one condition changing the response to another.
Between-subjects designs compare different participants
In a between-subjects design, each participant typically contributes data under one study condition relevant to the comparison.
For example, 100 students might be assigned to one of two feedback systems:
- 50 students use feedback system A;
- 50 different students use feedback system B.
The outcome comparison occurs between the two groups.
In a randomized experiment, participants can be randomly assigned to those conditions. In observational research, naturally occurring groups can also be compared, although the causal implications are different because the groups were not created through random assignment.
Within-subjects designs compare participants with themselves
In a within-subjects design, the same participant contributes outcome information under multiple conditions or occasions.
Suppose every student completes one task using interface A and another using interface B. The researcher can compare each student's performance under A with that same student's performance under B.
This pairing can be valuable because stable participant characteristics are held constant within that comparison. A student's general ability, educational background, or other relatively stable characteristics do not differ between that student's own measurements in the way they can differ across two separate groups.
Crossover trials are a formal example. Cochrane describes a simple two-period crossover design in which participants receive sequences such as A then B or B then A, allowing each participant to act as their own control.
Between-subjects comparison
Condition A is evaluated using one set of participants and condition B using another set, so the treatment contrast includes between-person variation.
Within-subjects comparison
The same participants contribute measurements under multiple conditions or occasions, allowing comparisons within individuals.
Within-subjects designs can reduce between-person variability
People differ. Sometimes considerably.
If two different groups are compared, natural variation in ability, prior experience, motivation, physiology, preferences, or other characteristics contributes to variability in the outcome.
Within-subjects designs can remove stable between-person differences from the treatment contrast because each participant is compared with themselves. Cochrane notes that this can reduce between-participant variation and often allows crossover designs to achieve a given precision with fewer participants than an analogous parallel-group design.
That efficiency can be attractive, especially when participants are difficult to recruit.
It is not free.
The first condition can affect what happens in the second
This is the central difficulty of many within-subjects designs.
Suppose participants learn how to solve a task while using interface A. When they later use interface B, they may perform better simply because they have already practiced the task.
Or the first intervention may produce an effect that persists into the second period.
In crossover terminology, a carryover effect occurs when the effect of an intervention from one period persists into a subsequent period and interferes with the response to the later intervention.
If the effect is durable or irreversible, a crossover design may make little sense. You cannot meaningfully return a participant to their original state after an intervention that permanently changes the outcome.
Order effects are broader than carryover
Participants' responses may change simply because of the order in which conditions occur.
They may become more practiced, tired, bored, familiar with the equipment, anxious, or strategic as the study progresses. These changes can influence later measurements even when the substantive intervention itself has no lingering effect.
Researchers can sometimes address order through counterbalancing or randomizing condition sequence. In a simple two-condition design, for example, some participants might receive A then B while others receive B then A.
Counterbalancing does not make every order problem disappear. It distributes order systematically across conditions so that treatment effects can be separated more effectively from order when the design and analysis support that separation.
Period effects can occur when the background changes over time
Cochrane distinguishes carryover from period effects.
A period effect occurs when responses differ systematically between study periods for reasons other than the treatment being compared. Participants may change over time, the environment may change, or external conditions may differ between periods.
Imagine evaluating two learning applications during an academic semester. Students might naturally become more knowledgeable as the semester progresses. If every student uses application A first and B later, application B is confounded with being observed later in the course.
Randomizing or balancing sequence can help separate treatment from period in suitable designs, while the analysis must reflect the repeated structure.
A washout period can help only when the effect can actually wash out
Crossover studies sometimes place time between conditions so that the effects of the first intervention can dissipate before the next condition is evaluated.
This is commonly called a washout period.
The concept is sensible only when the carryover process is sufficiently understood and reversible. Waiting a week cannot erase a skill that participants learned permanently. Nor can a washout restore a participant to a pre-intervention state after an irreversible outcome.
The existence of a proposed washout period therefore does not automatically make a crossover design appropriate.
Some interventions simply cannot be studied meaningfully within subjects
Suppose you want to compare two semester-long teaching curricula. Once students have completed curriculum A, their knowledge has changed. Asking them to complete curriculum B afterward does not recreate the state they would have been in had they received B first.
A between-subjects design may therefore be much more defensible.
Similarly, interventions involving surgery, vaccination, permanent training, substantial learning, or other durable changes may be unsuitable for crossover designs because the first condition cannot be removed.
Cochrane specifically cautions against crossover designs when interventions can lead to permanent or long-term modification or when the condition changes substantially over time.
Some comparisons are naturally well suited to within-subject designs
Other research questions fit within-subject comparisons very naturally.
A researcher might compare how quickly the same participant completes short standardized tasks under two interface layouts, provided exposure to one does not meaningfully alter performance under the other. Researchers might compare responses to transient stimuli, alternative display formats, or short-lived conditions when carryover can reasonably be controlled.
The crucial issue is whether each participant can provide valid outcome information under each condition.
Pretest-posttest is within-subject, but it is not necessarily a crossover design
The terminology can become confusing here.
If the same participants are measured before and after an intervention, the observations are repeated within participants. The design therefore contains a within-subject temporal comparison.
But participants have not necessarily crossed between two alternative interventions.
A crossover design usually refers to participants receiving multiple treatment conditions in different periods or sequences. A simple pretest-posttest study measures change over time under an intervention but does not automatically provide the same causal comparison as a randomized crossover trial.
This distinction matters because a baseline measurement is not itself another treatment condition.
The statistical analysis must respect repeated observations
Measurements from the same participant are generally correlated. Treating them as if they came from unrelated individuals discards the pairing built into the design and can produce incorrect uncertainty estimates.
For simple situations, paired analyses may be appropriate. More complex repeated-measures, longitudinal, or crossover designs may require models that account explicitly for within-participant correlation, treatment sequence, period, clustering, missing data, or other design features.
The correct analysis depends on the research question, outcome, number of conditions and measurements, randomization structure, and assumptions.
| Consideration |
Between-subjects design |
Within-subjects design |
| Who experiences each condition? |
Different participants |
The same participants experience multiple conditions or occasions |
| Stable between-person differences |
Contribute to group variation |
Can be controlled through within-person comparison |
| Carryover between conditions |
Usually not a cross-condition problem for an individual |
Can be a major threat |
| Order and practice effects |
Usually less central across separate participants |
Often require explicit design attention |
| Participant burden |
Each participant may experience fewer conditions |
Each participant may complete more conditions or sessions |
| Analysis |
Must reflect independent or clustered group structure as appropriate |
Must account for repeated or paired observations |
The choice is not simply about which design is “stronger”
Neither design dominates universally.
A within-subject design can be exceptionally efficient for a reversible, short-lived intervention with stable participants and manageable order effects. The same design can be indefensible for an intervention whose effects persist.
A between-subjects design avoids exposing the same participant to competing conditions but may require more participants or greater precision to separate the treatment effect from natural between-person variation.
The better design is the one whose comparison matches the phenomenon being studied.