03 · What You Need to Know
What Should a Research Question Survive Before You Design the Study?
Established frameworks already provide useful criteria for evaluating research questions. The widely used FINER framework, for example, asks whether a question is feasible, interesting, novel, ethical, and relevant. Other question-formulation frameworks help researchers specify important elements such as populations, exposures or interventions, comparisons, and outcomes. These frameworks are valuable, but a stress test serves a slightly different purpose: it treats the question as something to challenge rather than merely something to describe.
A useful stress test therefore asks not just, “Is this a good question?” but, “What could prevent this question from producing a convincing answer?” That shift matters because apparently minor wording choices can carry substantial methodological commitments.
1. Can Different Readers Tell What the Question Is Actually Asking?
Start with interpretation. Give the question to someone who understands the field but has not participated in developing the study. Ask that person to explain, in ordinary language, what would have to be investigated.
If two informed readers identify different populations, variables, relationships, outcomes, or units of analysis, the wording may be carrying more ambiguity than you realized. A question does not have to specify every methodological detail, but its central meaning should not depend on the reader guessing what you intended.
This becomes particularly important when terms such as “effectiveness,” “engagement,” “success,” “quality,” “impact,” or “use” appear without sufficient conceptual boundaries. Before proceeding, ask whether different researchers could reasonably interpret the question differently.
2. What Has to Be True for the Question to Make Sense?
Next, identify the premises hidden inside the wording. Some questions quietly assume that a phenomenon exists, that two groups are comparable, that a variable varies sufficiently to study, or that a proposed relationship is plausible.
For example, consider the question: “Why does the use of generative AI reduce students' critical-thinking ability?” The wording does not merely ask about a relationship. It presupposes that generative AI use reduces critical thinking and moves immediately to explaining why.
A stronger stress test would ask what evidence establishes that reduction in the first place. If it has not been established in the relevant context, the researcher may need to reformulate the question so that the presumed relationship becomes something to investigate rather than something embedded as fact.
This is why it is useful to inspect whether the question assumes something that has not yet been established.
3. Is the Question Quietly Making a Causal Claim?
Words such as “effect,” “impact,” “influence,” “leads to,” “results in,” and “because of” can imply causation. That implication matters because causal questions generally require stronger design and inferential conditions than descriptive or associational questions.
A researcher may intend only to examine whether two variables are associated while phrasing the question as though one produces changes in the other. If the eventual design cannot support that inference, the problem begins before data collection. Check explicitly for a hidden causal assumption in the research question rather than waiting for it to surface during interpretation.
4. Can the Important Concepts Actually Be Observed or Measured?
A concept can be theoretically interesting without being straightforward to investigate empirically. Ask what observable evidence would represent every important construct in the question.
If the question refers to “learning,” for example, what would count as evidence of learning in this particular study? Examination performance? Conceptual understanding? Skill demonstration? Retention? Transfer? Self-reported learning? These are not interchangeable.
The stress test is not simply whether an instrument exists. It is whether the intended construct can be represented in a way that is sufficiently valid for the claim you hope to make. If you cannot explain how the central concepts could become defensible observations, the question may be asking more than the study can establish. This warrants a separate check of whether the variables can actually be measured.
5. Can You Identify the Population the Question Refers To?
Questions sometimes refer to apparently recognizable populations such as “online learners,” “AI users,” “working students,” “high-performing researchers,” or “at-risk students.” Operationally identifying membership in those populations may be much harder.
Ask who qualifies, who does not, where eligible participants can be found, and whether the inclusion criteria would identify the population consistently. Population definition also affects the scope of the conclusions that can reasonably be drawn from the eventual study.
6. Is Every Comparison Defensible?
Comparative questions deserve particular scrutiny. Two groups can be statistically compared without the comparison necessarily answering a meaningful scientific question.
Suppose you want to compare students who voluntarily use an educational technology with students who do not. Those groups may already differ in motivation, digital competence, prior achievement, access to technology, or other characteristics relevant to the outcome. The mere availability of two groups does not make them equivalent counterfactuals or even necessarily informative comparison groups.
Ask what the comparison is intended to reveal and what alternative explanations would remain. If the comparison itself is conceptually weak, more sophisticated analysis may not rescue the underlying question.
7. Is There Really One Question Here?
Long research questions often conceal several investigations. A question may ask about prevalence, causes, experiences, group differences, and consequences in a single sentence. Each component may require different evidence and possibly a different design.
Try removing each clause. If the remaining parts still constitute independent research questions, you may be dealing with multiple questions bundled together. In that case, determine whether they should become a primary question with carefully justified subquestions or separate investigations. A dedicated check can help determine whether the wording combines several different questions into one.
8. What Evidence Would Convince You That the Question Has Been Answered?
This is one of the strongest tests you can apply before designing a study. Imagine that data collection is complete. What evidence would allow you to write a defensible answer to the research question?
Do not answer with a method such as “a survey,” “interviews,” or “regression analysis.” Those are ways of producing or analyzing evidence. Instead, describe what the evidence itself would need to establish.
If you cannot specify what evidence would count as an answer, you are not yet ready to decide confidently how that evidence should be generated.
9. Could the Study Realistically Produce That Evidence?
Once the required evidence is clear, test whether it is obtainable. Feasibility is a recognized criterion for evaluating research questions and can involve participant availability, expertise, resources, time, funding, equipment, institutional support, data access, and manageable scope.
There is also a deeper form of feasibility. A question may require information that no realistic version of your proposed study could produce. A cross-sectional survey, for example, can provide useful evidence about variables measured at one period, but it cannot automatically establish temporal ordering or eliminate plausible alternative explanations required for a strong causal conclusion.
If the evidence demanded by the wording exceeds the evidence the project could plausibly generate, either the design or the question must change.
10. Would Every Plausible Result Still Produce an Informative Answer?
Imagine several possible outcomes before collecting data. The expected relationship appears. It is absent. The relationship runs in the opposite direction. Results differ across subgroups. Estimates are too imprecise to support a confident conclusion.
Does the question remain scientifically meaningful across these possibilities?
A robust research question should not depend on obtaining the result the researcher hopes to see. If only one direction feels like a successful outcome, the wording may be functioning more like a prediction than an open empirical question. Stress-testing therefore includes asking whether the question can produce a meaningful answer regardless of the direction of the result.
11. Is the Question Asking Evidence to Do More Than Evidence Can Do?
Some questions move from empirical investigation directly to prescription: “What strategy should universities adopt?” or “Which policy should the government implement?” Evidence can inform such decisions, but recommendations usually depend on values, priorities, costs, feasibility, stakeholder preferences, and acceptable trade-offs in addition to empirical findings.
If the real objective is to generate evidence that informs a decision, frame the empirical question around the outcomes, experiences, relationships, mechanisms, or trade-offs that can actually be investigated. Then make the recommendation at the appropriate stage of interpretation rather than building it into the question itself.