03 · What You Need to Know
Start With the Mechanism, Not Just the Bias Label
Bias is systematic error that can distort the result or inference a study is intended to produce. It is therefore different from ordinary sampling variability. Increasing sample size may improve precision, but it does not necessarily eliminate systematic distortion.
Classic epidemiologic classifications often distinguish selection bias, information or observation bias, and confounding. CDC materials, for example, describe selection bias in terms of how study subjects are identified or included and information bias as systematic differences in how exposure or outcome information is obtained across study groups.
Contemporary risk-of-bias frameworks classify problems somewhat differently depending on the study design and causal question. Cochrane's ROBINS-I framework separates domains such as confounding, participant selection, intervention classification, missing data, outcome measurement, and selective reporting. The categories are useful precisely because different mechanisms can threaten the same final estimate in different ways.
| Problem |
Where the Distortion Arises |
Typical Question |
| Selection bias |
Who enters, remains in, or is included in the relevant comparison or analysis |
Did the selection process create systematic differences related to the exposure, intervention, outcome, or prognosis? |
| Information bias |
How information about participants or variables is obtained, recalled, recorded, or classified |
Was information collected or classified systematically differently or inaccurately? |
| Measurement bias |
How a particular construct, exposure, outcome, or covariate is assessed |
Did the measurement process systematically distort the value being observed? |
What Is Selection Bias?
Selection bias occurs when the process determining inclusion in the study, comparison, follow-up, or analysis creates systematic distortion in the relationship being estimated.
The crucial point is that selection bias is not simply “the sample is not representative.” A nonrepresentative sample may primarily restrict generalizability without necessarily biasing a particular association or causal effect estimate. Selection bias becomes an internal-validity problem when the selection mechanism distorts the comparison relevant to the research question.
Suppose researchers compare academic performance between students who voluntarily participate in an intensive study-skills program and students who do not. If students who volunteer are systematically more motivated, the groups may differ in ways relevant to academic performance before the program begins.
Whether this particular problem is best characterized as selection, confounding, or both depends on the design and causal structure. That qualification matters. Bias terminology is not a substitute for explaining the mechanism.
Selection Bias Can Occur After Recruitment
Selection does not stop when participants enroll.
Suppose a longitudinal study successfully recruits comparable groups, but participants experiencing poor outcomes are substantially more likely to leave one group during follow-up. An analysis restricted to those with complete outcome data may no longer preserve the original comparison.
Cochrane's risk-of-bias framework explicitly treats missing outcome data as a potential source of bias when missingness depends on factors related to the outcome or intervention. In other settings, researchers may describe related mechanisms using terms such as attrition bias or loss-to-follow-up bias.
The methodological question remains the same: has the process determining which observations remain available for analysis changed the estimate?
Volunteer Samples Are Not Automatically Biased
Researchers sometimes label every convenience or volunteer sample as selection bias. That is too broad.
A volunteer sample may differ substantially from a target population, creating legitimate concerns about generalizability. Yet whether volunteering biases a particular association depends on how participation relates to the variables involved in that association.
For example, an online survey of university students recruited through social media may overrepresent highly active social-media users. That could severely distort an estimate of social-media usage prevalence. Whether it also biases the relationship between two other variables requires additional reasoning about the selection process.
This distinction prevents internal validity and external validity from being collapsed into the same question.
What Is Information Bias?
Information bias arises when information about exposures, outcomes, participant characteristics, or other variables is systematically inaccurate or obtained differently across relevant study groups.
The umbrella can include several familiar problems. Participants may remember previous exposures differently. Interviewers may question one group more intensively. Medical records may classify conditions inconsistently. Researchers may use different data sources for different groups.
CDC epidemiologic materials describe information or observation bias as systematic differences in how exposure or outcome data are obtained from study groups and include examples such as recall bias, interviewer bias, and misclassification.
The common feature is not merely that measurement is imperfect. The information process systematically distorts the comparison or estimate.
What Is Recall Bias?
Recall bias occurs when the accuracy or completeness of remembered information differs systematically between groups in a way relevant to the research question.
Consider a case-control study asking participants about exposures that occurred several years earlier. People who have developed the outcome may search their memories more intensively for possible causes than people who have not.
If that process leads cases to report past exposure differently from controls even when their actual exposure histories were comparable, the exposure-outcome association may become distorted.
Recall bias is therefore more specific than ordinary forgetting. Everyone forgetting equally does not automatically produce the same bias mechanism.
What Is Interviewer or Observer Bias?
Researchers themselves can influence the information collected.
An interviewer who knows which participants experienced an outcome might probe them more extensively about possible exposures. An observer who knows which treatment a participant received may rate an ambiguous outcome more favorably. A researcher expecting a particular result might unconsciously interpret borderline observations differently.
Standardized procedures, assessor training, objective measurement where appropriate, and blinding to information that could influence assessment can reduce some of these risks.
Blinding is not always possible or even relevant, however. The question is whether knowledge available to the person collecting or judging the information could systematically affect the measurement.
What Is Measurement Bias?
Measurement bias refers to systematic distortion introduced by the way a variable is measured. Depending on the field, it may be discussed as a form of information bias, measurement error, detection bias, observer bias, or outcome-assessment bias.
Suppose a study compares two instructional approaches but one group's examinations are scored by instructors who know which teaching method students received while the other group's assessments are scored independently. If that knowledge systematically influences scoring, outcome measurement may favor one group.
Measurement bias can also arise from instruments themselves. A device may be systematically miscalibrated. A questionnaire may systematically undercapture a behavior because respondents interpret its wording differently. An administrative database may classify an outcome using rules that change during the study.
This is why using a previously validated instrument does not automatically eliminate measurement bias. Administration, scoring, context, and implementation still matter.
Measurement Error Can Be Differential or Non-Differential
An important distinction concerns whether measurement error differs according to other variables involved in the study.
Differential misclassification occurs when classification error depends on another relevant variable, such as outcome status affecting how exposure is reported or intervention status influencing how an outcome is assessed.
Non-differential misclassification generally refers to classification error that does not differ according to the comparison variable in the specified way.
A familiar shortcut says that non-differential misclassification always biases results toward the null. That is not a universal rule. The direction of bias depends on the structure of the variables, the form of misclassification, and the effect measure involved. Researchers should therefore avoid predicting the direction of bias mechanically.
Watch Out
Do not assume that every measurement error merely makes it harder to find an effect. Depending on how measurement errors arise, estimates can be attenuated, exaggerated, or distorted in less predictable ways.
Misclassification Can Affect Exposures, Outcomes, and Covariates
Researchers often think about misclassification only in relation to outcomes, but any important variable can be classified incorrectly.
An exposure may be categorized incorrectly. A participant may be placed in the wrong intervention group. A confounder may be measured so crudely that adjustment is incomplete. Even variables used to determine eligibility can be misclassified.
Cochrane's ROBINS-I framework, for example, separately evaluates bias in classification of interventions and bias in measurement of outcomes because those errors occur at different points in the causal process and may affect estimates differently.
The Same Problem Can Receive Different Labels Across Frameworks
Bias terminology is not perfectly standardized across disciplines.
A problem described as information bias in epidemiology may appear as outcome-measurement bias in a trial risk-of-bias tool. Differential loss to follow-up may be described as attrition bias, missing-data bias, or a selection process depending on the framework and inferential question.
This does not mean the concepts are arbitrary. It means the mechanism should be described rather than relying on a label alone.
When writing a thesis or article, specify what happened: who was systematically excluded, what was measured incorrectly, which group was affected, and how the process could distort the estimate. That explanation is more informative than simply adding “selection bias may be present” to a limitations section.
Bias Prevention Usually Begins Before Statistical Analysis
Many selection and information problems are easier to prevent than to repair after data collection.
Researchers can define eligibility procedures prospectively, recruit comparison groups using compatible procedures, standardize measurement, use appropriate instruments, train assessors, blind outcome assessment where feasible, record reasons for nonparticipation or attrition, and collect information needed to evaluate missingness.
Some biases can be addressed partly through analysis, but statistical adjustment cannot recreate information that was never collected accurately. This is part of the broader reason threats to valid research should be anticipated during study design.