03 · What You Need to Know
Size the Pilot for the Question the Pilot Must Answer
A Pilot Has Its Own Research Objectives
A common mistake is to think of the pilot sample as a miniature version of the definitive sample-size calculation. If the main study requires 500 participants, perhaps the pilot should use 50 because 10% sounds reasonable.
That logic does not explain what 50 participants allow researchers to learn.
A pilot should have explicit preparatory objectives. These might concern recruitment, retention, intervention delivery, assessment completion, participant burden, data quality, staff procedures, or the operation of a complete study workflow. The sample should provide enough observations to investigate those objectives.
Before choosing a number, therefore, return to what the pilot is actually supposed to test.
There Is No Universal Minimum Sample Size
Methodological literature does not support one sample size that makes every pilot adequate. Different objectives require different amounts and types of information.
A pilot designed mainly to identify obvious procedural problems may require a different justification from one intended to estimate a recruitment proportion or the variability of an outcome. A pilot involving repeated observations may obtain substantial process information from each participant, while another design may depend on observing events that occur infrequently.
The research question for the pilot comes first. The number follows.
Precision Matters When You Are Estimating a Rate or Proportion
Suppose you want to estimate the proportion of enrolled participants who complete follow-up. A pilot estimate of 80% sounds informative, but its usefulness depends heavily on how many participants produced that estimate.
If four of five participants complete follow-up, the observed completion proportion is 80%. If 80 of 100 complete follow-up, the observed proportion is also 80%. Those identical percentages carry very different levels of statistical uncertainty.
For feasibility parameters such as recruitment, retention, adherence, completion, or missingness, sample-size planning can therefore consider the precision required around the estimate. Confidence intervals can make that uncertainty explicit.
Think in Terms of Opportunities to Observe the Process
Not every pilot objective is about estimating a proportion. Sometimes you need enough opportunities for a process to fail.
If the pilot tests a complex participant visit, each completed visit gives researchers an opportunity to observe misunderstandings, equipment problems, scheduling difficulties, protocol deviations, or data-recording errors. If only two visits are conducted, many plausible problems may never occur.
The relevant unit may not always be participants. It could be recruitment weeks, intervention sessions, study sites, interviewers, clusters, data transfers, laboratory runs, or repeated assessments.
Ask how many times the uncertain process must occur before you can reasonably learn how it behaves.
Complexity and Heterogeneity Can Require More Testing
A procedure may work with the first few participants and fail when circumstances change.
If the main study includes participants with substantially different characteristics, several sites, multiple intervention providers, different technologies, or varied settings, a very small homogeneous pilot may not expose the relevant implementation challenges.
This does not mean the pilot must statistically represent every subgroup in the definitive population. It means the preliminary design should include enough variation to investigate the uncertainties that matter.
For example, if researchers specifically need to know whether procedures work at both urban and rural sites, testing ten participants at one urban site cannot answer the site-related feasibility question regardless of how carefully those ten participants are studied.
Duration Can Matter as Much as Participant Number
A pilot can have many participants and still be inadequate if it is too short.
Suppose a main study requires twelve months of follow-up. A pilot involving 100 participants but observing them for only two weeks cannot establish whether twelve-month retention procedures are workable. Similarly, a recruitment pilot conducted during an unusually favorable two-week period may provide weak evidence about recruitment across an academic year or seasonal clinical population.
Sample size should therefore be considered alongside duration, number of repeated observations, sites, and exposure to the processes being tested.
Recruitment Feasibility Requires Enough Time and Opportunity
If recruitment is an objective, the pilot needs to observe the recruitment pathway under conditions informative about the main study.
A tiny recruitment target can be misleading. A research team may recruit the first ten highly accessible volunteers quickly and conclude that enrolling 500 participants will be straightforward. Later recruitment may slow as the easiest participants are exhausted.
Assessing whether recruitment can realistically support the main study may require considering recruitment rate, eligibility, consent, available population, duration, sites, and possible changes over time rather than only the pilot's final sample size.
A Pilot Can Be Too Small to Reveal Data Problems
Some data problems appear only after enough observations accumulate. Rare response categories may not occur. Certain combinations of variables may never be seen. Missingness patterns may remain invisible. Data linkage may appear straightforward because only a handful of records are involved.
If data collection and management are important pilot objectives, researchers should ensure that preliminary testing generates enough realistic data to exercise the relevant workflow.
Simulated data can supplement participant data when testing databases, coding, scoring, validation, or analysis scripts, but simulation cannot reproduce every real-world pattern of participant behavior or missingness.
Do Not Size a Pilot to Test Effectiveness Unless Effectiveness Is Truly Its Design Question
A feasibility-focused pilot is generally not intended to provide a definitive test of whether an intervention works. Attempting to power it for conventional hypothesis testing can transform the study into something substantially larger and conceptually different from the preliminary research originally proposed.
Conversely, choosing a small sample because “it is only a pilot” and then interpreting p-values from that sample as evidence of effectiveness is equally problematic.
The sample-size rationale and the conclusions should align. If the study is sized around feasibility, interpret it as feasibility research and respect the limits on substantive conclusions from pilot data.
Published Rules of Thumb Need Context
Researchers will encounter various recommendations for pilot sample sizes in methodological literature. Some concern estimating a standard deviation for planning a future trial. Others concern detecting common procedural problems, assessing questionnaire performance, or obtaining a desired level of precision for a feasibility parameter.
These recommendations answer different methodological questions. A number derived for one purpose should not be detached from that purpose and presented as a universal minimum.
If you use a published recommendation, explain why the underlying rationale applies to your pilot objective.
A Small Pilot Can Still Be Useful
Being small does not automatically make a pilot poor. A small pilot may be highly informative when the objective is focused and the problems being investigated are readily observable.
For example, a small number of carefully observed participant sessions may reveal that instructions are consistently misunderstood or that a device cannot function in the intended environment. Once a decisive problem is identified, recruiting many additional participants simply to confirm that the same broken procedure remains broken may add little.
Usefulness depends on information gained relative to the decision, not on reaching an impressive-looking sample.
Watch Out
Do not justify a pilot sample solely by saying it represents a particular percentage of the main-study sample. A percentage describes the relationship between two numbers; it does not explain why the pilot contains enough participants, observations, sites, or time to answer its feasibility questions.