01 · The Question
Should the Question Follow the Problem or the Data You Already Have?
Researchers are often caught between two sensible principles.
The first says that research should begin with an important unanswered question. Decide what you need to know, then determine what evidence could answer it.
The second is more pragmatic: your study can only make claims supported by the data you can actually obtain. If the ideal variables, participants, measurements, or records are unavailable, asking a question that requires them will not make those limitations disappear.
The tension becomes especially obvious in secondary-data research. You may want to understand why students leave university, for example, but the institutional dataset contains only enrollment records, grades, demographics, and learning-management-system activity. Should you ask the important question you cannot fully answer, or replace it with whatever the database makes convenient?
Neither extreme is satisfactory.
03 · What You Need to Know
A Good Study Aligns the Question You Care About With the Evidence You Can Obtain
In an idealized research sequence, a researcher identifies an important problem, formulates a question, and then selects a design and data capable of answering it. In practice, research is often more iterative. Existing datasets, participant access, available instruments, ethical constraints, funding, and time may influence what can realistically be investigated.
This does not make the research illegitimate. Secondary analysis of existing data, for example, may proceed from a prior research question, from exploration of available variables, or through an iterative combination of the two. The critical issue is whether the eventual question and the available evidence genuinely correspond.
Established criteria for evaluating research questions explicitly include feasibility. FINER asks whether a question is Feasible, Interesting, Novel, Ethical, and Relevant. Data availability is therefore a legitimate consideration, but it is only one consideration. A question can fit your dataset beautifully while failing the equally important tests of relevance or scientific value.
Separate the Question You Care About From the Question Your Data Can Answer
Suppose you want to know:
“Why do students drop out of university?”
Your dataset contains enrollment status, grades, age, program, scholarship status, and learning-management-system activity. It contains no information about students' motivations, family responsibilities, employment, health, experiences with instructors, financial pressures beyond scholarship status, or reasons for leaving.
The dataset may support questions about which recorded characteristics are associated with subsequent withdrawal. It cannot, by itself, provide a complete answer to why students leave.
Question you ultimately care about
Why do students leave university?
Question these data may support
Which measured characteristics in the available institutional records are associated with subsequent withdrawal?
The second question can contribute to the first without being equivalent to it. Keeping that distinction explicit prevents a common inferential error: presenting the answer to an available-data question as though it resolved the larger phenomenon.
Available Data Can Legitimately Refine a Research Question
Researchers should not treat every constraint as methodological defeat.
If an important question cannot be investigated exactly as originally imagined, you may narrow the population, revise the outcome, modify the comparison, change the timeframe, use a defensible proxy, or focus on one component of the larger question. Secondary-data researchers routinely evaluate whether existing datasets contain variables suitable for answering proposed questions and refine those questions when necessary.
This is part of determining whether a question is actually answerable with the evidence available.
The important qualification is that the revised question must remain scientifically worthwhile. Feasibility should refine the inquiry, not reduce it to whatever happens to be easiest to calculate.
Do Not Rename a Proxy as the Phenomenon You Really Wanted
Data constraints often encourage researchers to use proxy measures. This can be defensible when the relationship between the proxy and the intended construct is theoretically and empirically justified.
Problems arise when the distinction disappears.
If a dataset contains login counts but not direct measures of learning, “learning-management-system activity” should not quietly become “learning.” If records contain attendance but not engagement, attendance should not automatically be described as engagement. If a survey measures self-reported intention to use a technology, the result should not be reported as actual adoption.
Your question should reflect what the evidence measures, not what you wish it measured.
Watch Out
Changing the label does not change the evidence. If your dataset measures a proxy, frame the research question and subsequent claims around what that proxy can defensibly represent, and acknowledge important limitations in interpretation.
Data-Driven Questions Are Not Automatically Bad Questions
Beginning with a dataset is sometimes portrayed as inferior to beginning with a hypothesis or research problem. That judgment is too broad.
Existing data can reveal valuable opportunities for secondary analysis. Researchers may identify patterns requiring explanation, evaluate questions that would otherwise be expensive or unethical to investigate prospectively, or reuse data to address new questions efficiently.
What matters is how the question is developed and evaluated.
If you notice that a dataset contains information about student withdrawal and financial aid, asking whether financial-aid status is associated with withdrawal may be reasonable if the relationship has a defensible rationale and the data are appropriate. It becomes much weaker if you mechanically test every available variable against withdrawal until something produces an interesting result and then write the study as though that hypothesis had been specified from the beginning.
Exploration and Confirmation Should Not Be Confused
Available data may generate hypotheses. That is useful. But questions discovered through exploration of the same dataset should be represented honestly as exploratory rather than retroactively presented as entirely prespecified.
This matters because repeated searching across variables, subgroups, outcomes, and models can produce apparently interesting patterns by chance. Statistical methods and reporting practices need to reflect the exploratory or confirmatory character of the analysis.
The broader lesson for question formulation is simple: data may inspire a question, but the provenance of that question matters when interpreting the resulting evidence.
Sometimes the Right Decision Is to Change the Data, Not the Question
If the question is important and your current dataset cannot answer it, you do not always have to surrender the question.
You might collect additional variables, recruit a different population, add interviews, obtain another dataset, link multiple data sources, extend the observation period, use a different instrument, or redesign the study entirely.
This is especially important when narrowing the question to fit the existing data would remove the very phenomenon that made the research worthwhile.
When the desired question cannot be answered directly, the relevant decision is whether to change the evidence strategy or reformulate the question into something that can be answered defensibly.
Sometimes the Right Decision Is to Change the Question
The opposite is also true.
A researcher may want to estimate the causal effect of an educational intervention but possess only cross-sectional observational data. Collecting experimental or sufficiently informative longitudinal evidence may be impossible within the project.
In that situation, retaining an “effect” question does not make the evidence causal. The researcher may need to ask about an association instead, assuming the resulting question remains worthwhile.
This is one reason to be cautious when asking about an effect without a design capable of supporting the intended causal inference.
The Available-Data Question Still Has to Pass the “So What?” Test
Suppose you narrow an ambitious question until your dataset can answer it perfectly. You are not finished.
Ask what the narrower answer contributes.
If the revised question is technically feasible but no longer addresses a meaningful uncertainty, you may have solved the methodological problem by creating an intellectual one. This is precisely how a study can end up with an answerable question that is nevertheless the wrong question to ask.