03 · What You Need to Know
Four Probability Sampling Methods and the Problems They Solve
What Makes These Methods Probability Sampling?
A probability sampling design uses a random mechanism that gives population or frame units known, nonzero probabilities of selection under the design. The U.S. Census Bureau includes random, systematic, and stratified sampling among probabilistic methods and emphasizes that the probabilities of selection and other design information must be retained for estimation and variance calculation.
This matters because “random” should describe an actual probability mechanism, not an informal attempt to recruit a varied group.
If you post a survey link and accept whoever volunteers, the sample does not become random merely because you did not personally choose the respondents. Likewise, selecting people who happen to be available in several departments is not stratified probability sampling simply because several departments appear in the dataset.
Before choosing among specific designs, make sure the broader distinction between probability and non-probability sampling is clear.
What Is Simple Random Sampling?
In simple random sampling, a specified number of units is selected from a population or sampling frame using a random mechanism so that the design gives each eligible unit an equal probability of selection and each possible sample of the specified size the appropriate probability under the design.
Suppose a university has a complete frame of 10,000 eligible students and you need a simple random sample of 500. Each student can be assigned a unique identifier, and a properly implemented random procedure can select 500 identifiers from the frame.
The conceptual appeal is obvious: the design is comparatively easy to understand, and analysis is often more straightforward than with complex samples.
The practical difficulty is that simple random sampling may require a sufficiently complete list of individual units. It can also be inefficient operationally when the population is geographically dispersed or when reliable estimates are required for relatively small subgroups.
What Is Stratified Sampling?
In stratified sampling, the population or frame is divided into non-overlapping groups called strata, and probability samples are selected within those strata.
The strata are formed before sample selection using information already available on the frame. A university study might stratify students by college, degree level, campus, or another characteristic relevant to the design.
Why do this? One reason is to ensure that important subgroups receive planned sample allocations rather than leaving their realized sample counts entirely to chance. Stratification can also improve statistical efficiency when units within strata are relatively similar with respect to variables related to key estimates.
Sample allocation across strata does not have to be proportional to population size. Researchers may deliberately oversample a small subgroup to obtain enough observations for reliable subgroup analysis. If selection probabilities differ across strata, appropriate weights and design-aware analysis may be needed for population estimates.
Watch Out
Oversampling a subgroup does not mean pretending that the subgroup is equally common in the population. The sampling design and analysis need to preserve the actual selection probabilities when producing population estimates.
What Is Cluster Sampling?
Sometimes selecting individuals directly is expensive or operationally difficult because the population is naturally organized into groups. Those groups can sometimes be used as clusters.
Schools, classrooms, villages, hospitals, households, geographic areas, and workplaces are common examples of naturally occurring clusters, depending on the study.
In a cluster design, clusters are selected through a probability mechanism. Researchers may then study all eligible units within selected clusters or select additional probability samples within them. The latter produces a multistage design.
For example, instead of drawing individual students from every school in a province, researchers might first select schools and then sample students within selected schools.
The major attraction is operational efficiency. Data collection can be concentrated in fewer locations, and a complete list of every individual in the entire target population may not be required at the first stage.
There is a statistical trade-off. People within the same cluster often resemble one another. Students in the same school, for example, may share institutional conditions and demographic characteristics. Such within-cluster similarity can reduce the amount of independent information obtained from a given number of observations compared with a simple random sample, which is one reason cluster designs often require design-effect considerations in sample-size planning and variance estimation.
What Is Systematic Sampling?
Systematic sampling selects units at regular intervals from an ordered sampling frame after a probability-based starting point is established.
Suppose a frame contains 10,000 units and the design calls for approximately 500 selections. The sampling interval would be around 20. After choosing an appropriate random start, the procedure selects units according to that interval.
Systematic selection can be easier to implement than independently generating a random number for every selected unit. The Census Bureau includes systematic sampling among methods used in statistically sound sample designs.
The ordering of the frame deserves attention. If the list contains a periodic pattern that aligns unfavorably with the sampling interval, systematic selection can produce undesirable results. Conversely, deliberate sorting before systematic selection can sometimes provide useful implicit stratification.
How Do the Four Methods Compare?
| Method |
Core idea |
Particularly useful when |
Important consideration |
| Simple random sampling |
Select individual units directly through a simple random mechanism |
A suitable individual-level frame exists and the population is operationally manageable |
May provide inadequate realized numbers for small subgroups and can be costly for dispersed populations |
| Stratified sampling |
Divide the frame into strata and sample within each |
Important subgroups require planned representation or stratification can improve precision |
Allocation and unequal selection probabilities must be reflected appropriately in estimation |
| Cluster sampling |
Select naturally occurring groups, sometimes followed by sampling within groups |
Individual units are geographically or organizationally dispersed and clustering reduces fieldwork costs |
Within-cluster similarity can reduce statistical efficiency |
| Systematic sampling |
Select units at regular intervals after an appropriate random start |
An ordered frame is available and an efficient selection procedure is useful |
Frame ordering and periodic patterns need consideration |
Stratified Sampling and Cluster Sampling Are Almost Opposite Ideas
These two methods are frequently confused because both divide a population into groups.
With stratified sampling, researchers intentionally sample from the strata. Ideally, stratification creates groups within which units are relatively similar on useful design variables while ensuring that different strata are represented according to the allocation plan.
With cluster sampling, researchers select clusters themselves, and only selected clusters may contribute observations. Operationally, the design often benefits from having units geographically or organizationally concentrated, although similarity within clusters can reduce precision.
Stratification
Divide the population into groups and deliberately sample within the groups according to the design.
Clustering
Use groups themselves as sampling units at one stage, then observe or sample units within the selected groups.
Can You Combine Probability Sampling Methods?
Yes. Real survey designs often combine them.
A national education study might stratify geographic areas, select schools as clusters within strata, and then systematically select students within sampled schools. Such a design is both stratified and multistage, with clustering introduced through school selection.
The Census Bureau explicitly treats stratification, clustering, systematic selection, oversampling, probabilities of selection, and multistage sampling as design elements that can be combined to meet statistical and operational requirements.
This is why reducing an entire methodology to “random sampling was used” is inadequate. Readers need enough information to understand what was randomized, at which stage, and with what probabilities.
Your Sampling Frame May Determine Which Designs Are Feasible
A method that looks attractive theoretically may be impossible with the information available.
Simple random sampling of individuals generally requires an individual-level frame. Stratified sampling additionally requires reliable information for assigning frame units to strata. Cluster sampling may be useful when a complete individual-level frame is unavailable nationally but lists of schools, villages, or other clusters exist.
Before selecting a design, examine whether your sampling frame adequately corresponds to the population and contains the variables needed to implement the proposed design.
The Analysis Must Remember How the Sample Was Selected
Sampling design does not end when data collection begins.
Stratification, clustering, unequal probabilities of selection, and multistage selection can affect weights, variance estimates, standard errors, and confidence intervals. An analysis that treats a complex sample as though it were a simple random sample may calculate uncertainty incorrectly.
The CDC Field Epidemiology Manual specifically advises consulting a survey-sampling expert for probability procedures beyond simple random sampling. For complex surveys, that is often prudent. The clever sampling design should not disappear mysteriously when the dataset reaches the statistics software.