03 · What You Need to Know
Random Assignment Changes How Groups Are Created
The National Institutes of Health describes randomization as assigning treatments to participants by chance rather than by choice. The National Library of Medicine similarly defines random allocation as a chance-based process for allocating experimental subjects among treatment or control groups.
That chance mechanism is the defining feature.
Without random assignment, group membership may tell you something about the participants
Imagine evaluating an optional tutoring program. Students decide for themselves whether to participate.
Those who join may be more motivated, more concerned about their performance, more available after class, more comfortable seeking help, or academically different from those who decline. If the two groups later have different examination scores, the tutoring program is not the only plausible explanation.
Group membership itself may reflect characteristics related to the outcome.
Random assignment interrupts that selection process. Rather than allowing participant characteristics, preferences, instructor judgments, or other systematic factors to determine condition membership, the allocation mechanism uses chance.
Random assignment makes baseline characteristics balanced in expectation, not identical in every study
One of the most persistent misconceptions about random assignment is that successful randomization must produce two groups with exactly the same characteristics.
It does not.
Under proper randomization, treatment assignment is independent of baseline characteristics according to the randomization mechanism. Across repeated random allocations, this produces balance in expectation between conditions. In any particular study, however, chance can produce numerical imbalances.
With 40 participants, for example, one randomized condition might happen to contain more participants with high prior achievement. That does not automatically mean the randomization failed.
This is why baseline differences between randomized groups should be interpreted differently from systematic baseline differences created through self-selection or investigator assignment.
Balance in expectation
The randomization mechanism does not systematically direct particular baseline characteristics into particular treatment conditions.
Exact balance in the realized sample
The observed groups happen to have identical or nearly identical distributions of particular characteristics. Random assignment does not guarantee this.
The major benefit is protection against confounding by baseline characteristics
Suppose motivation affects examination performance. In a self-selected intervention, highly motivated students might disproportionately choose the new teaching program, creating confounding.
With valid random assignment, motivation does not determine which intervention participants are assigned to, even if motivation was never measured.
This point is especially important. Unlike statistical adjustment in an observational study, randomization does not require researchers to identify and measure every baseline characteristic that might affect the outcome in order for the assignment mechanism to operate independently of those characteristics.
That is one reason randomized experiments provide such a strong design for estimating causal effects of assigned interventions.
Random assignment helps construct the comparison needed for causal inference
Causal questions ask what would happen under one condition compared with an alternative. For an individual participant, however, we generally cannot observe the outcome under two mutually exclusive conditions at the same time.
Randomization helps address this counterfactual problem by creating groups whose treatment assignments arise through chance. Outcomes under the alternative assigned conditions can then be compared to estimate treatment effects under the design and analysis.
This does not mean every difference observed after randomization must be caused by treatment. Sampling variability remains, which is why statistical uncertainty is quantified. The causal strength comes from the design of the assignment mechanism, not from the disappearance of chance.
Random assignment is not random sampling
Random assignment answers:
Which condition will participating units receive?
Random sampling answers:
Which units from a population will enter the sample?
A study can recruit volunteers from one university and then randomly assign them to two teaching methods. The treatment comparison can still be randomized even though the participants were not randomly sampled from all university students.
Conversely, researchers could draw a probability sample from a population and merely observe participants' naturally occurring exposures. That would not create randomized exposure groups.
Keeping random assignment separate from random sampling prevents causal inference and population generalization from being conflated.
Random assignment does not make your sample representative
Suppose 120 volunteers from one university are randomly assigned to an intervention and comparison condition.
The randomization governs allocation among those 120 participants. It says nothing by itself about whether they resemble students at other universities, older learners, working students, or even nonparticipating students at the same institution.
Questions about who participated and to whom findings apply concern sampling, recruitment, eligibility, setting, external validity, and the target population.
Random assignment cannot retroactively change any of those features.
Random assignment does not prevent every kind of bias
Randomization primarily protects the treatment-assignment process. Many things can still go wrong afterward.
| Problem |
Does random assignment automatically prevent it? |
Why? |
| Systematic baseline treatment assignment |
Yes, when randomization is properly implemented |
Assignment is determined by the specified chance mechanism rather than baseline participant characteristics or investigator preference. |
| Chance baseline imbalance |
No |
Random allocation can still produce numerical differences in a particular sample. |
| Attrition after assignment |
No |
Participants may leave conditions at different rates or for different reasons. |
| Nonadherence |
No |
Participants may not receive or follow the intervention as assigned. |
| Contamination |
No |
Participants may be exposed to elements of another condition. |
| Biased outcome measurement |
No |
Knowledge of assignment or different measurement procedures may influence assessment. |
| Selective reporting |
No |
Researchers can still selectively report outcomes or analyses. |
| Population representativeness |
No |
Randomization allocates conditions; it does not select the study sample from the population. |
A randomized design is therefore not synonymous with a bias-free study.
Randomization must be genuinely random
Alternating participants between groups may look balanced. Assigning participants according to birth date, record number, day of attendance, or another predictable rule may also appear impartial.
These procedures are not equivalent to genuine random assignment merely because the researcher does not personally choose each condition.
A proper random allocation sequence uses an appropriate chance mechanism. Depending on the design, this might involve computer-generated random numbers or another defensible randomization procedure.
The methods section should describe the mechanism rather than simply state that participants “were randomized.”
Generating a random sequence and concealing it are different problems
Even a genuinely random allocation sequence can be undermined if upcoming assignments are known before participants are enrolled or allocated.
Suppose a recruiter knows that the next assignment is the intervention condition. If that knowledge influences whether a particular participant is enrolled at that moment, the resulting groups may no longer reflect the intended randomization process.
Allocation concealment addresses this problem by preventing those involved in enrollment or assignment from knowing upcoming allocations before the participant is irreversibly entered into the trial.
Random sequence generation and allocation concealment therefore protect related but distinct parts of the allocation process.
Random assignment does not eliminate the need for blinding
After allocation, participants, intervention providers, researchers, or outcome assessors may know which condition was assigned.
That knowledge can sometimes influence behavior, co-interventions, adherence, reporting, or measurement.
Randomization does not prevent these post-assignment mechanisms. When knowledge of condition could introduce consequential bias and masking is feasible, blinding may address a different methodological problem.
Random assignment can occur at levels other than the individual
Sometimes individuals cannot or should not be assigned independently.
An educational intervention might be delivered to entire classrooms. A public-health intervention might operate at the community level. Hospitals or clinics might be randomized as clusters.
Random assignment can therefore occur at the individual or cluster level. NIH definitions explicitly recognize prospective assignment of research participants individually or in clusters.
The analysis must respect the unit and structure of randomization. Randomizing 20 classrooms does not create the same statistical structure as independently randomizing every student within those classrooms.
Restricted randomization can be used when particular forms of balance matter
Researchers do not always rely on unrestricted allocation.
Methods such as block randomization can help maintain allocation numbers across conditions, while stratified randomization can help achieve balance on prespecified important characteristics. Cluster trials may use other appropriate randomization strategies.
These procedures are still random assignment when treatment allocation contains the required chance mechanism. They simply constrain the randomization according to a prespecified design.
The choice should be made before outcomes are observed and reported transparently.