Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Shouldn’t You Conclude From a Pilot Study?

A pilot study can provide valuable evidence about whether and how a future main study can be conducted, but it usually cannot support the substantive conclusions expected from that definitive study. Small samples and feasibility-focused objectives place important limits on what pilot results mean.

166
What Not to Conclude From a Pilot Study Guide 166 of 217
01 · The Question

How Far Can You Take the Findings From a Pilot Study?

A pilot study produces data. Once those data exist, it is tempting to analyze them in much the same way as data from the future main study. Perhaps the intervention group performed better than the comparison group. Perhaps a correlation was statistically significant. Perhaps no adverse events occurred. Perhaps recruitment went smoothly.

Can you conclude that the intervention works, the relationship is real, the procedure is safe, or the main study will succeed?

Usually, those conclusions go beyond what the pilot was designed to establish. A pilot is conducted principally to reduce uncertainty about whether and how a future study can be carried out. Its sample size, objectives, and analytical strategy should reflect that purpose. The presence of substantive outcome data does not automatically make the pilot capable of answering the substantive research question.

02 · The Short Answer

Do Not Treat a Pilot as a Small Definitive Study

In Brief

You generally should not conclude from a feasibility-focused pilot study that an intervention is effective or ineffective, that a hypothesis has been confirmed or rejected, that the observed effect size is a reliable estimate of the true effect, or that successful piloting guarantees the main study will succeed.

Interpret the findings against the objectives the pilot was actually designed to address. A pilot can provide valuable evidence about recruitment, retention, procedures, implementation, measurement, data collection, and other feasibility questions without being capable of supporting definitive substantive conclusions.

03 · What You Need to Know

The Pilot’s Objectives Define the Boundaries of Its Conclusions

A Pilot Is Primarily About the Future Study

The defining logic of a pilot is preparatory. Researchers conduct the future study, or relevant parts of it, on a smaller scale to investigate uncertainties before proceeding to the definitive research.

Those uncertainties might concern whether participants can be recruited, whether randomization works, whether an intervention can be delivered consistently, whether participants complete the required assessments, whether follow-up procedures retain participants, or whether the resulting data can be managed as planned.

This is why the objectives should be established before the pilot begins. The analyses should then address those objectives. If the pilot was designed to determine whether recruitment and follow-up are workable, discovering a difference in an outcome between two small groups does not transform effectiveness into the study's primary question.

Before interpreting results, return to what the pilot was actually designed to test.

Do Not Conclude That the Intervention Is Effective

One of the most consequential errors is treating a favorable outcome in a small pilot as evidence that an intervention works.

A feasibility-focused pilot is generally not designed or powered to provide a definitive test of efficacy or effectiveness. Small samples produce substantial statistical uncertainty. An apparently large difference may be exaggerated by sampling variation, while a small or absent difference may simply reflect imprecision.

Even when the pilot uses randomization and measures the same primary outcome intended for the definitive trial, its purpose may still be feasibility. The fact that a treatment comparison can be calculated does not establish that the study was designed to interpret that comparison definitively.

Pilot question Can the procedures required for the future study operate sufficiently well, and what should be changed before proceeding?
Definitive study question What does adequately designed evidence show about the substantive research question, such as the efficacy or effectiveness of an intervention?

Do Not Conclude That an Intervention Is Ineffective Because the Pilot Is Not Statistically Significant

The opposite mistake is equally problematic. Researchers may perform a hypothesis test in a small pilot, obtain a non-significant result, and decide that a larger study is not worth conducting.

A non-significant result does not establish the absence of a meaningful effect. In a small pilot, the study may simply have insufficient precision to distinguish among a wide range of plausible effects.

This can create a peculiar situation: the pilot successfully demonstrates that recruitment, intervention delivery, measurement, and follow-up are workable, yet the researchers abandon the definitive study because an inferential analysis that the pilot was never designed to support happens to produce a large p-value.

The decision to progress should instead be based on the feasibility objectives and any prespecified progression process, together with other relevant scientific evidence.

Do Not Treat Statistical Significance as Confirmation of the Main Hypothesis

A statistically significant pilot result is not automatically more trustworthy. With small samples, estimates can be unstable, and a statistically significant finding may correspond to an unusually large observed effect.

There is also a broader design problem. If the pilot was sized around feasibility rather than the substantive hypothesis, conventional hypothesis testing answers a question the study was not designed to answer reliably.

Researchers should therefore resist both directions of overinterpretation: “the pilot was significant, so the intervention works” and “the pilot was not significant, so there is no effect.” Neither follows simply from a feasibility-focused pilot.

Do Not Assume the Observed Effect Size Is a Reliable Planning Estimate

Researchers sometimes use the observed treatment effect from a small pilot to determine the sample size of the definitive study. This can be risky because effect estimates from small samples are often imprecise.

If the observed effect happens to be unusually large, the resulting main-study sample size may be too small. If it happens to be unusually small, the calculation may suggest a much larger study than necessary or discourage further research.

Sample-size planning for the definitive study should therefore use the strongest relevant evidence available and a scientifically meaningful target effect, rather than assuming that the pilot's point estimate represents the true effect.

Pilot data may sometimes inform other quantities needed for planning, such as aspects of variability or event rates, but those estimates also carry uncertainty and should be handled accordingly.

Do Not Conclude That No Observed Harm Means the Procedure Is Safe

Suppose a small pilot involving 20 participants produces no serious adverse events. That may be reassuring within the limited experience obtained, but it does not establish that a rare adverse event cannot occur.

Small studies have little opportunity to observe uncommon events. Safety conclusions therefore require particular caution and should reflect the intervention, existing evidence, sample size, duration of exposure, and nature of the risks being considered.

A pilot can identify obvious or frequent safety and tolerability problems. Absence of such observations should not be converted into a broad declaration of safety that the study cannot support.

Do Not Assume Feasibility in the Pilot Guarantees Feasibility at Full Scale

A pilot may show that the procedures work with a small number of participants, a highly engaged research team, one site, or a short follow-up period. Scaling changes the conditions.

Recruitment may become more difficult once the most accessible participants have enrolled. Additional sites may implement procedures differently. Staff training may become harder to standardize. Data-management systems may behave differently when the volume increases. Problems that are rare may never appear during the pilot.

Feasibility evidence is therefore conditional on what was actually tested. A successful pilot reduces uncertainty about the future study; it does not certify that the definitive study will encounter no operational problems.

Do Not Generalize Beyond the Population and Conditions Tested Without Justification

If a pilot was conducted at one institution, with a particular participant group, using a particular delivery system, its feasibility findings arise from those conditions.

Recruitment success among university students at one campus does not automatically establish recruitment feasibility across universities with different populations or administrative arrangements. Acceptability among participants who volunteered for a pilot may not represent all eligible participants in the main study.

The appropriate conclusion should preserve these boundaries. State what worked, for whom, where, and under what conditions, then explain how confidently that evidence can inform the planned main study.

Do Not Confuse “We Completed the Pilot” With “Everything Was Feasible”

Completing data collection is not itself proof of feasibility. Researchers may finish a pilot despite recruitment taking much longer than planned, extensive staff intervention being required, substantial missing data, repeated protocol deviations, or poor participant retention.

Feasibility should be judged against the objectives and evidence, not against the mere fact that the research team managed to reach the end.

Watch Out

A pilot can be operationally completed while simultaneously providing strong evidence that the proposed main study should not proceed unchanged. Treating completion itself as success can obscure precisely the problems the pilot was intended to reveal.

A Pilot Can Support Strong Conclusions About the Questions It Was Designed to Answer

Caution does not mean pilot studies tell researchers nothing. The issue is alignment between question and conclusion.

If the pilot was explicitly designed to investigate whether the recruitment pathway works, its recruitment evidence can inform that question. If it examined whether assessments can be completed, those findings can inform measurement procedures. If it tested whether intervention providers can follow the protocol, implementation findings may identify necessary changes.

The correct response is not to weaken every conclusion with vague language. It is to make the conclusion proportional to the design and the uncertainty actually investigated.

04 · A Practical Example

The Same Pilot Results Can Support Some Conclusions but Not Others

Hypothetical Example

A Small Pilot of a Digital Learning Intervention

A research team pilots a digital learning intervention with 30 students before a substantially larger randomized study. The pilot was designed to test recruitment, intervention delivery, completion of assessments, and follow-up.

Finding Recruitment reaches the pilot target within the planned period, most participants complete the intervention, and follow-up data are obtained from 27 of the 30 participants.
Reasonable conclusion The results provide preliminary evidence that the tested recruitment, intervention, and follow-up procedures can operate under the pilot conditions, while identifying any modifications required before scaling.
Additional observation Students receiving the intervention have a higher average outcome score than students in the comparison condition.
Unreasonable conclusion The intervention is effective and will improve learning in the main study.
Why The pilot was designed around feasibility objectives rather than a sufficiently precise test of effectiveness. The observed difference may be reported cautiously where appropriate, but it does not carry the evidential weight of the future definitive analysis.

The pilot has produced useful evidence. The methodological mistake would be asking its data to answer a question that its design was not built to answer.

05 · What Researchers Often Get Wrong

Common Ways Pilot Findings Become Overinterpreted

Misconception

The Intervention Group Improved, So the Intervention Works

Improvement within a group does not by itself establish that the intervention caused the improvement, and a small feasibility-focused pilot is generally not designed to provide definitive evidence of efficacy or effectiveness. Interpret substantive outcomes according to the actual design and objectives.

Misconception

The Result Was Statistically Significant, So the Pilot Confirmed the Hypothesis

Statistical significance does not repair a design that was not intended to test the substantive hypothesis definitively. Small-sample estimates can also be unstable. A surprising p-value should not redefine the purpose of the study after the data have been examined.

Misconception

The Result Was Not Significant, So a Main Study Is Unnecessary

A feasibility-focused pilot may have inadequate precision for detecting the effect the main study is intended to investigate. Lack of statistical significance therefore should not ordinarily be used as evidence that the substantive question has been answered negatively.

Misconception

No Participants Experienced the Problem, So It Will Not Occur

A small sample can easily miss uncommon events, rare procedural failures, or problems that emerge only at scale. Absence of an observation in a pilot establishes only that it was not observed under the conditions tested.

Misconception

Everything Worked in the Pilot, So the Main Study Can Proceed Unchanged

Progression should be judged against the pilot objectives, evidence, and relevant progression criteria. Scaling, longer follow-up, additional sites, and larger research teams can introduce problems that a small pilot could not reveal.

Misconception

The Pilot Effect Size Is the Effect Size We Should Expect in the Main Study

Small-sample effect estimates can be highly imprecise. Treating the pilot point estimate as though it were the true effect can distort planning for the definitive study. Sample-size decisions should incorporate stronger external evidence and scientifically meaningful effects where available.

06 · What This Means for You

Match Every Conclusion to a Pilot Objective

When interpreting a pilot, write down each conclusion you want to make and ask what part of the study was designed to support it. This simple check can expose overreach quickly.

A simple decision framework

If the conclusion concerns recruitment, retention, acceptability, implementation, or another prespecified feasibility objective
Interpret the relevant pilot evidence with appropriate attention to precision and the conditions under which it was collected.
If the conclusion claims that an intervention is effective or ineffective
Ask whether the pilot was genuinely designed and adequately powered to answer that substantive question rather than assuming that outcome data make the claim valid.
If the conclusion depends mainly on whether p < 0.05
Return to the pilot objectives and determine whether hypothesis testing was actually appropriate for the question being investigated.
If the conclusion treats the observed effect size as a precise planning value
Consider the uncertainty of the estimate and use stronger external evidence or a scientifically meaningful target effect where appropriate.
If the conclusion claims the definitive study is guaranteed to be feasible
Restrict the interpretation to the procedures, population, setting, duration, and scale actually tested.

If the findings reveal serious procedural problems, do not reinterpret them away simply because you want the main study to proceed. A pilot may legitimately show that the original design needs substantial revision before the main study.

Likewise, if the pilot supports progression, explain why in terms of its feasibility evidence rather than presenting a favorable substantive outcome as permission to continue.

07 · A Quick Checklist

Before Interpreting Pilot Study Results, Check

Before drawing conclusions from a pilot, check:
What were the prespecified objectives of the pilot?
Was the sample size justified for those pilot objectives rather than for a definitive test of the main hypothesis?
Are you interpreting recruitment, retention, implementation, measurement, or other feasibility outcomes according to the conditions actually tested?
Are you avoiding claims of efficacy or effectiveness that the pilot was not designed to establish?
Are you avoiding both positive and negative conclusions based merely on statistical significance?
Have you considered the uncertainty around small-sample effect estimates before using them for planning?
Are safety or rare-event claims appropriately limited by the pilot's sample size and duration?
Does the decision about progression follow the pilot's feasibility evidence rather than an incidental substantive result?
08 · Frequently Asked Questions

Frequently Asked Questions About Pilot Study Conclusions

Can a pilot study prove that an intervention works?

A feasibility-focused pilot generally should not be interpreted as definitive evidence of efficacy or effectiveness. Its primary purpose is to resolve uncertainty about the future study. A study specifically designed and adequately powered for a substantive effectiveness question is methodologically different, even if its sample happens to be relatively small.

Should I report p-values from a pilot study?

That depends on the objective and analysis, but conventional hypothesis testing of effectiveness is generally discouraged when the pilot was designed around feasibility. Analyses should primarily address the pilot objectives, with estimates and their uncertainty often being more informative than treating statistical significance as the decision criterion.

Can I use the pilot effect size to calculate the main-study sample size?

Using a small pilot's observed treatment effect as the expected true effect can be unreliable because the estimate may be highly imprecise. Where possible, planning should draw on stronger external evidence and a scientifically meaningful target effect. Pilot information may still inform other design quantities, but its uncertainty should be acknowledged.

Can a pilot tell me that recruitment is feasible?

Yes, if recruitment was a genuine pilot objective and the design tested it under sufficiently relevant conditions. Interpret the result in relation to recruitment rate, eligibility, time, setting, available sites, and the scale required for the future study rather than simply noting that the pilot reached its target.

Can I conclude that an intervention is safe if nobody was harmed in the pilot?

Not broadly. The absence of observed harm provides information about what occurred in the tested sample, but a small pilot may have little ability to detect uncommon adverse events. Safety conclusions should therefore remain proportionate to the sample, exposure, follow-up, intervention, and wider evidence.

What if my pilot unexpectedly finds a statistically significant effect?

Report and interpret the finding according to the study's prespecified objectives and analytical plan rather than allowing the unexpected result to redefine the pilot retrospectively. A significant result does not convert a feasibility-focused pilot into a definitive effectiveness study.

What should the conclusion section of a pilot study emphasize?

It should emphasize what the pilot revealed about its stated feasibility objectives, the uncertainties that remain, modifications required for the future study, and the resulting progression decision. Substantive outcome findings should not overshadow the questions the pilot was actually designed to answer.

09 · The Bottom Line

Do Not Ask Pilot Data to Answer the Definitive Study’s Question

The Bottom Line

Do not treat a feasibility-focused pilot as evidence that an intervention works or does not work, that a hypothesis has been confirmed or rejected, or that a small observed effect estimate represents what the definitive study will find.

Use the pilot to answer the questions it was designed for: whether relevant study processes can work, what needs modification, and whether the evidence supports progression. A pilot is valuable precisely because it can improve the future study without pretending to replace it.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes