01 · The Question
How Far Can You Take the Findings From a Pilot Study?
A pilot study produces data. Once those data exist, it is tempting to analyze them in much the same way as data from the future main study. Perhaps the intervention group performed better than the comparison group. Perhaps a correlation was statistically significant. Perhaps no adverse events occurred. Perhaps recruitment went smoothly.
Can you conclude that the intervention works, the relationship is real, the procedure is safe, or the main study will succeed?
Usually, those conclusions go beyond what the pilot was designed to establish. A pilot is conducted principally to reduce uncertainty about whether and how a future study can be carried out. Its sample size, objectives, and analytical strategy should reflect that purpose. The presence of substantive outcome data does not automatically make the pilot capable of answering the substantive research question.
03 · What You Need to Know
The Pilot’s Objectives Define the Boundaries of Its Conclusions
A Pilot Is Primarily About the Future Study
The defining logic of a pilot is preparatory. Researchers conduct the future study, or relevant parts of it, on a smaller scale to investigate uncertainties before proceeding to the definitive research.
Those uncertainties might concern whether participants can be recruited, whether randomization works, whether an intervention can be delivered consistently, whether participants complete the required assessments, whether follow-up procedures retain participants, or whether the resulting data can be managed as planned.
This is why the objectives should be established before the pilot begins. The analyses should then address those objectives. If the pilot was designed to determine whether recruitment and follow-up are workable, discovering a difference in an outcome between two small groups does not transform effectiveness into the study's primary question.
Before interpreting results, return to what the pilot was actually designed to test.
Do Not Conclude That the Intervention Is Effective
One of the most consequential errors is treating a favorable outcome in a small pilot as evidence that an intervention works.
A feasibility-focused pilot is generally not designed or powered to provide a definitive test of efficacy or effectiveness. Small samples produce substantial statistical uncertainty. An apparently large difference may be exaggerated by sampling variation, while a small or absent difference may simply reflect imprecision.
Even when the pilot uses randomization and measures the same primary outcome intended for the definitive trial, its purpose may still be feasibility. The fact that a treatment comparison can be calculated does not establish that the study was designed to interpret that comparison definitively.
Pilot question
Can the procedures required for the future study operate sufficiently well, and what should be changed before proceeding?
Definitive study question
What does adequately designed evidence show about the substantive research question, such as the efficacy or effectiveness of an intervention?
Do Not Conclude That an Intervention Is Ineffective Because the Pilot Is Not Statistically Significant
The opposite mistake is equally problematic. Researchers may perform a hypothesis test in a small pilot, obtain a non-significant result, and decide that a larger study is not worth conducting.
A non-significant result does not establish the absence of a meaningful effect. In a small pilot, the study may simply have insufficient precision to distinguish among a wide range of plausible effects.
This can create a peculiar situation: the pilot successfully demonstrates that recruitment, intervention delivery, measurement, and follow-up are workable, yet the researchers abandon the definitive study because an inferential analysis that the pilot was never designed to support happens to produce a large p-value.
The decision to progress should instead be based on the feasibility objectives and any prespecified progression process, together with other relevant scientific evidence.
Do Not Treat Statistical Significance as Confirmation of the Main Hypothesis
A statistically significant pilot result is not automatically more trustworthy. With small samples, estimates can be unstable, and a statistically significant finding may correspond to an unusually large observed effect.
There is also a broader design problem. If the pilot was sized around feasibility rather than the substantive hypothesis, conventional hypothesis testing answers a question the study was not designed to answer reliably.
Researchers should therefore resist both directions of overinterpretation: “the pilot was significant, so the intervention works” and “the pilot was not significant, so there is no effect.” Neither follows simply from a feasibility-focused pilot.
Do Not Assume the Observed Effect Size Is a Reliable Planning Estimate
Researchers sometimes use the observed treatment effect from a small pilot to determine the sample size of the definitive study. This can be risky because effect estimates from small samples are often imprecise.
If the observed effect happens to be unusually large, the resulting main-study sample size may be too small. If it happens to be unusually small, the calculation may suggest a much larger study than necessary or discourage further research.
Sample-size planning for the definitive study should therefore use the strongest relevant evidence available and a scientifically meaningful target effect, rather than assuming that the pilot's point estimate represents the true effect.
Pilot data may sometimes inform other quantities needed for planning, such as aspects of variability or event rates, but those estimates also carry uncertainty and should be handled accordingly.
Do Not Conclude That No Observed Harm Means the Procedure Is Safe
Suppose a small pilot involving 20 participants produces no serious adverse events. That may be reassuring within the limited experience obtained, but it does not establish that a rare adverse event cannot occur.
Small studies have little opportunity to observe uncommon events. Safety conclusions therefore require particular caution and should reflect the intervention, existing evidence, sample size, duration of exposure, and nature of the risks being considered.
A pilot can identify obvious or frequent safety and tolerability problems. Absence of such observations should not be converted into a broad declaration of safety that the study cannot support.
Do Not Assume Feasibility in the Pilot Guarantees Feasibility at Full Scale
A pilot may show that the procedures work with a small number of participants, a highly engaged research team, one site, or a short follow-up period. Scaling changes the conditions.
Recruitment may become more difficult once the most accessible participants have enrolled. Additional sites may implement procedures differently. Staff training may become harder to standardize. Data-management systems may behave differently when the volume increases. Problems that are rare may never appear during the pilot.
Feasibility evidence is therefore conditional on what was actually tested. A successful pilot reduces uncertainty about the future study; it does not certify that the definitive study will encounter no operational problems.
Do Not Generalize Beyond the Population and Conditions Tested Without Justification
If a pilot was conducted at one institution, with a particular participant group, using a particular delivery system, its feasibility findings arise from those conditions.
Recruitment success among university students at one campus does not automatically establish recruitment feasibility across universities with different populations or administrative arrangements. Acceptability among participants who volunteered for a pilot may not represent all eligible participants in the main study.
The appropriate conclusion should preserve these boundaries. State what worked, for whom, where, and under what conditions, then explain how confidently that evidence can inform the planned main study.
Do Not Confuse “We Completed the Pilot” With “Everything Was Feasible”
Completing data collection is not itself proof of feasibility. Researchers may finish a pilot despite recruitment taking much longer than planned, extensive staff intervention being required, substantial missing data, repeated protocol deviations, or poor participant retention.
Feasibility should be judged against the objectives and evidence, not against the mere fact that the research team managed to reach the end.
Watch Out
A pilot can be operationally completed while simultaneously providing strong evidence that the proposed main study should not proceed unchanged. Treating completion itself as success can obscure precisely the problems the pilot was intended to reveal.
A Pilot Can Support Strong Conclusions About the Questions It Was Designed to Answer
Caution does not mean pilot studies tell researchers nothing. The issue is alignment between question and conclusion.
If the pilot was explicitly designed to investigate whether the recruitment pathway works, its recruitment evidence can inform that question. If it examined whether assessments can be completed, those findings can inform measurement procedures. If it tested whether intervention providers can follow the protocol, implementation findings may identify necessary changes.
The correct response is not to weaken every conclusion with vague language. It is to make the conclusion proportional to the design and the uncertainty actually investigated.