03 · What You Need to Know
First Diagnose Why the Question Cannot Be Answered Directly
“Impossible to answer” can describe several very different problems. The solution depends on which problem you actually have.
A question may be answerable in principle but infeasible for your particular study. The necessary population may be inaccessible. The outcome may take decades to emerge. The required data may be proprietary. The ideal intervention may be unethical. A rare event may require a sample far larger than you can recruit.
In other cases, the problem is conceptual. The question may ask about something that cannot be observed or measured in the form in which it has been stated.
Before rewriting anything, identify the obstacle.
| Why direct investigation is difficult |
Example |
Possible response |
| Time |
Long-term consequences cannot be observed within the project period |
Study intermediate outcomes, use longitudinal archival data, or narrow the time horizon |
| Access |
Required participants, institutions, records, or platforms cannot be accessed |
Use another defensible population, data source, setting, or indirect indicator |
| Ethics |
The exposure or condition cannot ethically be assigned |
Use naturally occurring variation, observational evidence, natural experiments, or another ethical design |
| Measurement |
The construct cannot be observed directly |
Develop or identify defensible indicators, measures, or multiple sources of evidence |
| Resources |
The ideal study exceeds available funding, expertise, equipment, or sample size |
Reduce scope, collaborate, use existing data, or investigate a more feasible component |
| Counterfactual impossibility |
You cannot observe the same unit simultaneously under two alternative conditions |
Use a design and inferential framework capable of estimating the relevant causal contrast |
These problems are not interchangeable. A measurement problem is not solved simply by recruiting more participants, and an ethical problem is not solved by obtaining more funding. Methodological obstacles deserve methodological diagnoses.
Distinguish the big question from the empirical question
Researchers often have a larger question motivating their work than any single study can answer.
For example:
Big question: “Will widespread generative AI use weaken university students' capacity to think and write independently over the long term?”
A single short-term study is unlikely to settle that question. It involves long-term development, changing technologies, different patterns of use, institutional responses, and difficult questions about what “independent thinking” means.
But the larger question can generate empirically tractable questions:
“Does access to AI-generated feedback during practice affect students' subsequent performance on an independently completed writing task?”
“How do students describe changes in their writing practices after sustained use of generative AI?”
“Is frequency of generative AI use associated with performance on independently completed writing tasks over one academic year?”
Each addresses a particular piece of the larger problem. None should be presented as the definitive answer to whether AI will weaken independent thinking over decades.
Motivating question
The larger scientific, theoretical, practical, or societal problem that makes the research important.
Empirical research question
The specific question the present study has evidence and methods capable of answering.
Keeping these levels separate can preserve ambition without requiring one study to make claims it cannot support.
Ask which part of the question is actually inaccessible
A large question may contain several components, only one of which creates the impossibility.
Suppose you ask:
“How does generative AI affect students' independent writing ability throughout their university education and after graduation?”
Perhaps the problem is not measuring writing ability. The problem is the long follow-up period.
You could investigate writing performance during one academic year without claiming to have answered the post-graduation question.
This is different from simply making the question smaller for convenience. You are isolating the component for which credible evidence is available.
Use an intermediate outcome when the final outcome cannot yet be observed
Some outcomes take too long to occur for a particular study.
Suppose an educational intervention is intended ultimately to improve university completion. A short-term study cannot observe graduation among first-year students. It might instead examine theoretically relevant intermediate outcomes such as persistence into the following year, credit accumulation, or another proximal outcome.
This can be useful, but an intermediate outcome is not automatically equivalent to the final outcome.
Improving an intermediate measure does not guarantee improvement in the ultimate outcome unless the relationship between them is sufficiently established.
Watch Out
Do not quietly replace the outcome you care about with an easier short-term measure and then discuss the findings as though the original outcome had been observed. State clearly that the study examines an intermediate or proximal outcome and justify why it is informative.
A proxy can make an unobservable concept researchable, but only if the proxy is defensible
Many research concepts are not directly observable.
Motivation, socioeconomic status, cognitive load, academic engagement, trust, well-being, institutional culture, and countless other constructs are investigated through indicators rather than direct observation of some perfectly tangible entity.
Using a proxy is therefore not inherently a methodological compromise. Research routinely depends on operationalization.
The problem arises when the proxy is treated as interchangeable with the underlying construct without adequate justification.
Suppose you want to investigate “student learning” but use final course grade as the only indicator. Grades may contain information about learning, but they can also reflect attendance, assignment completion, grading policies, prior achievement, participation, and other influences.
Your study needs a defensible argument connecting the observed indicator to the concept in the research question.
As discussed in the previous guide on whether a research question can actually be answered with data, convenient variables should not silently redefine the constructs researchers intended to study.
Multiple indicators may be better than one weak proxy
If no single measure captures the concept adequately, several sources of evidence may provide a more defensible representation.
For example, a study of academic engagement might draw on self-report, learning-management-system activity, attendance, classroom observation, or other evidence depending on how engagement is conceptualized.
Using several indicators does not automatically solve validity problems. The indicators may represent different dimensions rather than interchangeable measurements of one thing.
The value comes from explicitly theorizing what each indicator contributes and how the evidence relates to the construct.
Study observable implications of an unobservable explanation
Sometimes you cannot observe the process you ultimately care about directly, but a theory implies patterns that should be observable if the explanation is plausible.
Suppose you hypothesize that students rely on generative AI because uncertainty about assessment expectations creates decision pressure. “Decision pressure” may not be directly observable as an object.
You could investigate observable implications: students' accounts of uncertainty, variation in AI use across assignments with different guidance, changes after clearer policies are introduced, or patterns in help-seeking behavior.
No single observation necessarily proves the theoretical explanation. But systematically examining predicted implications can provide evidence for or against parts of it.
This is a common logic across empirical research: theories concern entities and processes that may not always be directly visible, while evidence concerns observations that should differ depending on which explanations are plausible.
Use a related population when the ideal population is inaccessible, but narrow the claim
Suppose your real interest is students who permanently left university after experiencing severe academic difficulty. You cannot identify or recruit them reliably.
You may be able to study currently enrolled students who considered withdrawing, students who temporarily stopped out and returned, or institutional staff who work with students at risk of leaving.
Those populations can provide relevant evidence. They do not become substitutes for the inaccessible population merely because recruitment is easier.
A study of students who considered leaving can answer questions about those students' experiences. It cannot automatically tell you whether people who actually left had the same experiences.
The appropriate response is to make the population difference explicit and explain what part of the larger problem the accessible population can illuminate.
Use retrospective evidence when prospective observation is impossible, but understand the tradeoff
If you cannot follow a process forward in time, you may sometimes reconstruct it retrospectively.
For example, instead of observing the development of institutional AI policy over three years, you might interview participants involved in the process and analyze drafts, meeting minutes, email records, policy documents, and other archival materials.
This can produce a rich reconstruction.
But retrospective evidence has limitations. Memories may be incomplete or reconstructed. Documents may preserve only parts of the process. Participants may interpret earlier events through the lens of what eventually happened.
These limitations do not make retrospective research invalid. They define what kind of reconstruction the evidence can support.
Use existing longitudinal data when your own study cannot wait
Your project may last two years while the outcome of interest takes ten years to develop. Existing longitudinal datasets, registries, cohort studies, administrative records, or archived data may allow you to investigate longer-term questions without waiting a decade.
Secondary data can therefore transform an otherwise infeasible time horizon.
However, you inherit the design decisions of the original data collection. The required constructs may not have been measured, the population may differ from your intended target, and variables may have been defined for another purpose.
Existing longitudinal data solve the problem of waiting. They do not automatically solve measurement, sampling, or causal-identification problems.
Natural variation can sometimes replace an unethical intervention
Some causal questions concern exposures that researchers cannot ethically assign.
You cannot randomly expose children to harmful pollutants to determine their developmental effects. You cannot deliberately deprive students of necessary educational support merely to create a control condition.
Researchers may instead examine naturally occurring variation, policy changes, eligibility thresholds, natural experiments, longitudinal cohorts, or other observational situations that provide informative contrasts.
Modern causal-inference methods explicitly address how observational data can sometimes be used to estimate causal effects under specified assumptions. The absence of randomization does not automatically eliminate the causal question, although it generally makes identification more dependent on assumptions about confounding, selection, measurement, and the data-generating process.
This is why a causal-effect question does not necessarily require an experiment, even though an ordinary observed association should not be relabeled as an effect.
Natural experiments can be valuable when circumstances create the comparison
Sometimes policies, administrative rules, geographic boundaries, timing, lotteries, thresholds, or other external circumstances generate variation that approximates aspects of an experiment.
For example, a university might introduce a new AI policy in some faculties before others for administrative reasons unrelated to student outcomes. Depending on the details, that staggered implementation could provide opportunities for stronger causal analysis than a simple comparison between students who voluntarily follow different practices.
Calling something a “natural experiment” does not make it one, of course. The credibility of the design depends on why the exposure differs and whether the assumptions required for the intended comparison are plausible.
Nature, like reviewers, rarely provides perfect experiments on demand.
Ask participants about perceived causes when objective causal identification is impossible, but label them as perceptions
Suppose you cannot establish what objectively caused students to leave a degree program, but you can interview former students about what they believe contributed to their decisions.
A legitimate question might be:
“How do former students explain their decisions to leave the program?”
The study can identify perceived reasons, experiences, and decision processes.
It should not automatically conclude:
“These factors caused university withdrawal.”
This distinction follows from the earlier discussion of asking “why” when a study cannot establish causation. Participants' explanations are valuable evidence about how they understand their decisions, but that evidentiary claim differs from estimating causal effects across a population.
Study mechanisms when the ultimate outcome is too distant
Suppose the ultimate question is whether a teaching practice improves long-term educational attainment. You cannot observe attainment within the present project.
You might investigate a theoretically specified mechanism expected to connect the intervention to later outcomes.
For example, does the practice improve retrieval, self-regulation, persistence, or another intermediate process?
This can strengthen understanding of how an intervention might work, but mechanism evidence should not be treated as proof of the eventual distal outcome. A mechanism can operate without producing a meaningful final effect, and an intervention can produce an effect through pathways different from those hypothesized.
Study necessary conditions without claiming they are sufficient
Some large questions can be decomposed by asking what must be true for the larger outcome to occur.
Suppose you ask whether an institutional AI policy can improve responsible student use. Before evaluating long-term behavioral effects, you might investigate whether students know the policy exists, understand its provisions, perceive it as applicable, and encounter consistent implementation.
If students have never heard of the policy, some proposed mechanisms of influence become difficult to sustain.
But awareness alone does not establish that the policy changes behavior.
Evidence about necessary or enabling conditions can therefore narrow plausible explanations without answering the entire outcome question.
Use simulation or modeling when direct observation is unavailable, but distinguish model results from observed reality
Some questions concern systems, scenarios, or future conditions that cannot yet be observed directly. Researchers may use statistical models, simulations, agent-based models, mathematical models, or scenario analyses to investigate what follows under specified assumptions.
For example, a university might model how different student-retention interventions could affect enrollment under alternative assumptions about uptake and effectiveness.
Such models can be extremely useful for exploring implications and decision scenarios.
The conclusions remain conditional on the model structure, parameter estimates, and assumptions. A simulated outcome is not an observed future outcome.
Good modeling makes that conditionality visible rather than allowing a plausible-looking graph to acquire supernatural predictive powers.
Use analogue settings cautiously
Sometimes the exact phenomenon is new, but a related phenomenon has already been studied.
Generative AI may create novel educational practices, yet earlier research on automated writing evaluation, intelligent tutoring systems, calculator adoption, internet search, or other technologies may provide useful conceptual analogues.
An analogue can help generate hypotheses, identify mechanisms, or suggest measurements.
It cannot establish that the new phenomenon behaves identically. The more important the differences between contexts, technologies, populations, or mechanisms, the weaker the analogy becomes.
Evidence synthesis may answer a question your individual study cannot
Sometimes no single study can provide enough evidence, but a body of research can.
A systematic review, meta-analysis, qualitative evidence synthesis, or other form of evidence synthesis may be appropriate when the question concerns what existing studies collectively show.
This is especially useful when conducting another primary study would add little or when the relevant evidence is dispersed across populations, settings, and methods.
Evidence synthesis has its own research questions, eligibility criteria, search strategies, appraisal procedures, and inferential limitations. It should not be treated as a shortcut for a question that has not been formulated clearly.
Sometimes the correct study is a feasibility or pilot study
You may know what definitive study would answer the question but not whether that study can actually be conducted.
For example, you may want to evaluate a complex intervention in a large randomized trial but remain uncertain whether participants can be recruited, whether institutions will implement the intervention, whether the outcome can be measured reliably, or whether adherence will be adequate.
A feasibility or pilot study can investigate those uncertainties.
The research question changes from:
“Does the intervention improve the outcome?”
to questions such as:
“Can eligible participants be recruited and retained at the required rate?”
“Can the intervention be delivered with acceptable fidelity?”
“Can the proposed outcome be collected sufficiently completely?”
The pilot should not be interpreted as a small, underpowered definitive effectiveness study. Its purpose is to determine whether and how the definitive question can later be answered.
Sometimes methodological research is the necessary first step
Your substantive question may be impossible to answer because no adequate measurement instrument, coding scheme, data linkage procedure, or analytical method exists.
In that situation, the first study may need to solve the methodological problem.
For example, before estimating how often students delegate substantive authorship to generative AI, researchers may need a defensible way to distinguish editing assistance, idea generation, rewriting, and text generation.
Developing and evaluating that measurement approach is not a detour from the substantive research agenda. It may be what makes the later question researchable.
Do not confuse an indirect question with the original question
This is the central discipline required by indirect research.
Suppose the original question is:
“Does generative AI use reduce students' long-term independent writing ability?”
Your study examines:
“Is frequent generative AI use associated with performance on one independently completed writing task at the end of a semester?”
The second study is relevant to the first question. It does not answer it completely.
The appropriate conclusion might be:
“These findings provide evidence about the relationship between AI-use frequency and short-term independent writing performance.”
It should not become:
“Generative AI harms students' long-term writing development.”
The inferential distance between the evidence and the larger question should remain visible.
Think in terms of an evidence chain
When direct evidence is impossible, map how the evidence you can obtain connects to the question you ultimately care about.
| Level |
Example |
What it contributes |
| Large motivating question |
Does long-term generative AI use weaken independent writing development? |
Defines the broader scientific problem |
| Mechanism |
Does AI use reduce opportunities for independent drafting and revision? |
Examines a plausible pathway |
| Intermediate outcome |
Do students who receive different forms of AI support differ in subsequent unaided writing performance? |
Provides evidence about a nearer outcome |
| Observable behavior |
How frequently do students delegate drafting or revision tasks to AI? |
Documents behavior relevant to the proposed mechanism |
| Participant explanation |
How do students describe changes in their own writing practices after using AI? |
Provides evidence about perceived processes and experiences |
No single row necessarily settles the first question. Together, however, a program of research can progressively strengthen or weaken particular explanations.
Some questions really should be left unanswered for now
Not every interesting question can be rescued by a clever proxy.
If the necessary evidence does not exist, cannot ethically be generated, cannot be approximated credibly, and no defensible indirect question would illuminate the problem, the appropriate conclusion may simply be that the question cannot currently be answered.
That is methodologically preferable to producing an answer from evidence that has only a superficial relationship with the question.
Research has genuine epistemic limits. Acknowledging them is part of rigor, not a failure of imagination.
An unanswered big question can still guide a research program
Large questions are often answered cumulatively.
One study describes a phenomenon. Another develops a measure. A third investigates an association. A fourth examines mechanisms. A fifth exploits a policy change for stronger causal evidence. A later synthesis integrates findings across settings.
No individual study needs to carry the entire burden.
This perspective can be especially useful for students who feel that narrowing a question somehow betrays the importance of the original problem. It does not. The narrower empirical question can be one defensible contribution to a larger research agenda.