03 · What You Need to Know
Baseline Is a Reference Point, Not a Particular Statistical Test
The National Institute on Aging describes baseline as an initial measurement made at an early point in a study, before participants begin receiving the intervention under investigation, at which measurable values such as assessments and laboratory tests may be recorded.
In practice, baseline information can include much more than one pretest score. What counts as relevant baseline information depends on the study.
Baseline characteristics describe who participants were when the study began
Baseline characteristics can include demographic, clinical, educational, behavioral, or other pre-intervention information.
In an educational study, researchers might record students' age, program, year level, prior academic achievement, previous experience with the technology being studied, or other characteristics relevant to the research question.
In a randomized trial, CONSORT recommends reporting important baseline demographic and clinical characteristics for each group. This allows readers to understand the participants who actually entered the trial and to inspect the realized characteristics of the randomized groups.
These variables do not all have to be outcomes. They describe relevant features of participants at the starting point.
A baseline outcome is the outcome measured before the intervention or follow-up
Suppose your primary outcome is academic-writing proficiency measured at the end of a training program.
If you administer the same or an appropriately comparable writing assessment before the program begins, you have a baseline measure of that outcome.
This can answer an important question that a post-intervention score alone cannot:
Where did participants start?
CONSORT specifically notes that baseline information can be particularly valuable for outcomes that can also be measured at the beginning of a trial.
Baseline and pretest often overlap, but they are not perfect synonyms
In many educational and behavioral studies, a pretest is a baseline measurement because it assesses the outcome before the intervention.
But baseline is broader.
A study can collect baseline age, prior achievement, socioeconomic information, previous technology use, or other characteristics that are not “pretests” in the conventional sense.
Baseline
The broader starting reference point and the relevant characteristics or measurements recorded at that point.
Pretest
A measurement administered before an intervention, often corresponding directly to an outcome that will be assessed again later.
A baseline is not a control group
This distinction causes considerable confusion.
Suppose one group completes an achievement test, receives an intervention, and then completes another achievement test. The first measurement is the baseline. It tells you where that group started.
It does not tell you what would have happened over the same period if those participants had experienced another condition.
A control or comparison group addresses that different problem by providing observations under an alternative condition.
| Feature |
Baseline measurement |
Control or comparison group |
| Main purpose |
Establishes a starting reference |
Provides an alternative condition for comparison |
| When observed |
Typically before intervention or follow-up |
Usually observed concurrently across the relevant study period |
| Can involve the same participants? |
Yes |
A separate group is common, although other comparative designs are possible |
| Shows whether participants changed? |
Can support assessment of change when followed by later measurement |
Can show how outcomes differ under alternative conditions |
| Shows what would have happened without the intervention? |
Not by itself |
Can contribute to estimating that alternative when the design supports the comparison |
This is why having a baseline does not answer the separate question of whether you need a control group.
A baseline lets you distinguish starting level from later outcome
Imagine two groups both score 80 at the end of a program.
Without baseline information, they appear identical on the final outcome.
Now suppose one group began at 60 and the other at 78. The same final score tells a very different story about their trajectories.
Baseline information can therefore reveal patterns that endpoint measurements alone conceal.
That does not mean researchers should automatically analyze simple change scores. The appropriate analysis depends on the design, outcome, estimand, measurement properties, and statistical assumptions. Baseline-adjusted analyses are often preferable in randomized trials when appropriately prespecified.
Baseline information is useful even when participants are randomized
If random assignment is intended to create comparable groups, why collect baseline information at all?
Because randomization and baseline measurement do different jobs.
Random assignment determines treatment allocation through a chance mechanism. Baseline data describe the participants and can provide prognostic information, support planned covariate adjustment, and help readers understand the realized groups.
Proper randomization does not guarantee that every baseline characteristic will be numerically identical. CONSORT explicitly notes that chance differences can occur even when random assignment has been correctly implemented.
Baseline significance tests are generally not how you determine whether randomization worked
A familiar table in research papers lists baseline characteristics followed by p-values comparing randomized groups. Researchers may then declare randomization successful because every p-value exceeds.05.
CONSORT advises against significance testing of baseline differences in randomized trials. If randomization was properly implemented, observed baseline differences arise by chance. Testing whether they could have arisen by chance therefore adds little and can mislead.
Instead, researchers should consider the magnitude of relevant imbalances and the prognostic importance of the variables.
This connects to the broader issue of interpreting baseline differences between study groups. The assignment mechanism matters more than a collection of baseline p-values.
Baseline measurements can improve statistical precision
When a baseline measure strongly predicts the later outcome, incorporating it appropriately into the analysis can improve precision.
For example, if prior mathematics achievement is strongly related to post-intervention mathematics achievement, an analysis that appropriately adjusts for baseline achievement may estimate the treatment contrast more precisely than an analysis using only post-intervention scores.
The analytical strategy should be selected because it matches the design and estimand, preferably before outcome data are examined, rather than chosen retrospectively according to which analysis produces the most favorable result.
Baseline data are particularly important in many nonrandomized comparisons
When groups are not randomly assigned, baseline information can help researchers examine whether intervention and comparison groups differed before treatment.
Suppose students who voluntarily participate in tutoring already have lower achievement than nonparticipants. Without baseline information, a final difference could easily be misinterpreted.
Baseline measurement can reveal such measured differences and support appropriate analytical adjustment where the necessary assumptions are plausible.
It cannot establish that nonrandomized groups are equivalent on unmeasured characteristics. A detailed baseline table is not a substitute for randomization.
Not every study needs a baseline outcome measurement
Consider a cross-sectional survey asking what proportion of university faculty currently use generative AI for lesson planning.
There may be no meaningful “before” measurement. The research question concerns current prevalence.
Similarly, a qualitative study exploring how researchers experience journal peer review does not automatically require a baseline interview conducted before they ever encountered peer review. That would be a different study.
Baseline data are most useful when a starting condition is relevant to the inference you intend to make.
Sometimes a baseline would be useful but cannot be measured
Researchers do not always have the luxury of measuring participants before an exposure occurs.
An observational study may begin after exposure has already happened. Researchers may use existing records or earlier measurements if valid data are available. In other cases, a genuine pre-exposure baseline simply does not exist.
That limitation should be acknowledged rather than disguising a later measurement as “baseline.”
The timing of baseline matters
A baseline should correspond to a scientifically meaningful starting point.
If an intervention begins on Monday but the supposed baseline outcome is measured two weeks later, participants have already been exposed to the intervention. Calling that measurement baseline does not make it pre-intervention.
Similarly, if groups are assessed at substantially different times relative to intervention initiation, the measurements may not represent comparable starting conditions.
Researchers should therefore define when baseline occurs and ensure that the measurement precedes the change whose effect or trajectory they want to study.