01 · The Question
What Are You Actually Supposed to Test in a Pilot Study?
Researchers sometimes describe a pilot as “a small version of the main study.” That description is useful up to a point, but it can create the wrong objective. If you simply run the intended study with fewer participants and analyze the same outcomes, what exactly have you learned about whether the larger study will work?
A useful pilot is more deliberate. It identifies areas of uncertainty about the future study and tests the procedures or processes needed to resolve them. Recruitment may need testing. So might consent, randomization, intervention delivery, participant adherence, measurement, follow-up, data management, or the planned analytical workflow.
The pilot does not need to test everything. It should test what you genuinely need to know before committing to the main study.
03 · What You Need to Know
Choose Pilot Objectives From the Weakest Assumptions in Your Design
Start With Areas of Uncertainty, Not a Standard Pilot Checklist
The most defensible pilot objectives arise from uncertainties in the future study. Ask what must happen successfully for the main study to work and which of those conditions have not yet been demonstrated adequately.
If recruitment has already been established in comparable studies using the same population and procedures, it may not deserve most of the pilot's attention. If your study introduces a new recruitment channel, however, recruitment could become one of the central pilot objectives.
The same principle applies to intervention delivery, assessments, follow-up, data handling, and other procedures. Pilot only what needs piloting.
If the uncertainty can be investigated without implementing the future study procedures, broader feasibility work may answer the question without requiring a full pilot.
Can Participants Be Recruited Through the Planned Pathway?
Recruitment is often tested because the main study's sample-size plan implicitly assumes that enough eligible participants can be identified and enrolled within a particular period.
A pilot can examine the complete recruitment pathway: how potential participants are identified, how many satisfy eligibility criteria, how many can be contacted, how many consent, how long recruitment takes, and where potential participants are lost.
This provides more information than reporting “we recruited 30 participants.” Thirty participants in two weeks and 30 participants in twelve months imply very different prospects for a future study.
If recruitment represents a major threat to progression, the pilot should produce evidence that can inform whether recruitment at the required scale is realistic.
Do Consent and Enrollment Procedures Work as Intended?
The journey between identifying an eligible participant and beginning study procedures contains several steps that can create difficulties. Participant information may be too long or unclear. Consent processes may be cumbersome. Eligibility assessment may take too much staff time. Randomization or allocation systems may not operate smoothly.
A pilot allows these procedures to be rehearsed under conditions closer to the main study. Researchers can identify delays, ambiguities, technical failures, or procedural deviations while changes remain relatively manageable.
Where participant understanding is uncertain, researchers may need to test whether participants interpret study instructions and procedures as intended, rather than assuming successful completion demonstrates genuine understanding.
Can an Intervention Be Delivered With Adequate Fidelity?
For intervention research, it may be necessary to determine whether the intervention can be delivered as specified in the protocol. Relevant considerations can include training, timing, dose, attendance, adherence, consistency across providers or sites, equipment, and deviations from the intended procedure.
Implementation problems can otherwise become difficult to distinguish from intervention effects in the main study. If participants do not receive the intervention as intended, an eventual null finding may be difficult to interpret.
The pilot can also reveal whether the intervention is excessively burdensome or incompatible with the setting. These are feasibility questions about implementation, not preliminary tests of whether the intervention ultimately works.
Can Participants Complete the Assessments?
A measure can have strong evidence of validity and reliability while still being impractical within a particular study workflow. A pilot can examine how long assessments take, whether participants skip items, whether instructions are understood, whether repeated assessments cause fatigue, and whether the timing of measurement is workable.
This is particularly important when several individually reasonable instruments are combined. A 10-minute questionnaire may be easy to complete. Six such instruments plus an interview and a performance task may create a very different participant experience.
The relevant pilot question is therefore not simply “Does the instrument work?” but “Can this measurement strategy operate as planned within this study?”
Will Participants Remain in the Study?
Longitudinal and repeated-measures studies depend on retention. A pilot can test follow-up procedures, reminder systems, scheduling, participant burden, and adherence to repeated study activities.
Researchers should pay attention not only to how many participants are lost, but when and possibly why they disengage. A pattern of missed assessments after a particularly burdensome procedure may suggest a different solution from loss caused by ineffective contact information.
However, short pilots may provide limited information about long-term retention. If the future study follows participants for a year but the pilot observes them for only a few weeks, the pilot cannot establish one-year retention simply by extrapolation.
Can Researchers Deliver the Protocol Consistently?
Participants are not the only source of uncertainty. Research staff must also implement the design reliably.
A pilot may reveal that eligibility decisions are inconsistent, interviewers interpret instructions differently, intervention providers require additional training, laboratory procedures vary among sites, or case-report forms encourage inconsistent recording.
Testing staff-facing procedures can therefore identify where the protocol needs clarification, training, standardization, or monitoring before scaling up.
Does the Data-Collection and Management Workflow Actually Work?
Data quality problems are often discovered too late because researchers focus on the participant-facing procedure but not on what happens to the data afterward.
A pilot can test whether identifiers link correctly across systems, coding conventions are workable, electronic forms enforce appropriate ranges, timestamps are recorded correctly, files transfer as expected, missing values are distinguishable from true zeroes, and datasets can be assembled into the intended analytical structure.
Researchers should examine the entire pathway from data generation to an analysis-ready dataset. Testing whether data collection is practical under realistic conditions can reveal problems that are invisible in a protocol diagram.
Can the Planned Analysis Be Implemented With the Data Structure?
A pilot may also be useful for testing the mechanics of the planned analytical workflow. This can include checking variable coding, data transformations, scoring procedures, software scripts, linkage, model implementation, and whether the collected data have the expected structure.
This is different from using a small pilot to determine whether the main hypothesis is supported. The purpose is to test the analytical process and identify practical problems before the definitive dataset arrives.
Where this is relevant, it is worth separating testing whether the planned analysis works technically from using pilot data to make substantive statistical conclusions.
Test the Whole Workflow When Interactions Between Procedures Matter
Individual components can work perfectly in isolation and still fail when combined.
A consent process may take 20 minutes, an assessment 45 minutes, and an intervention session 60 minutes. Each may be acceptable separately. Requiring participants to complete all three consecutively may not be.
A pilot is particularly useful for detecting these interactions because it can reproduce the sequence intended for the main study. This is one reason researchers may test several connected parts of the study before full implementation.
Decide in Advance What the Pilot Results Will Mean
Testing a procedure is useful only if the result can inform a decision. For important objectives, researchers may establish progression criteria before collecting pilot data.
These might concern recruitment, retention, intervention fidelity, assessment completion, or other feasibility outcomes. Depending on the results, the decision might be to proceed, proceed with modifications, undertake further preliminary work, or reconsider the proposed main study.
Watch Out
Do not make formal hypothesis testing of effectiveness the primary purpose of a pilot designed to assess feasibility. Pilot samples are usually not selected to provide adequate power for definitive effectiveness testing, and apparently promising or disappointing results can be highly unstable.