03 · What You Need to Know
Good Research Can Begin With Data, but It Should Not End With Availability
The familiar ideal of research design begins with a problem, develops a research question, and then identifies the evidence required to answer it. This sequence is valuable because it makes methodological decisions responsive to the question rather than allowing the available data to dictate what researchers claim to be interested in.
Real research is sometimes less linear. Researchers encounter public datasets, institutional records, archived interviews, longitudinal cohorts, administrative databases, or unused variables that reveal questions they had not previously considered. Secondary data analysis explicitly works with data that already exist, and methodological literature recognizes both question-driven and data-driven routes to developing secondary analyses.
The important distinction is between allowing data to inspire a worthwhile question and allowing data availability to become the entire justification for that question.
Question-Driven and Data-Driven Research Are Not Simple Opposites
In a question-driven approach, researchers identify a question and then seek data capable of answering it. In a data-driven approach, researchers inspect the information available in a dataset and identify potentially worthwhile questions that the data could address.
In practice, the process can be iterative. A researcher may begin with a broad area of interest, discover a high-quality dataset, examine its measures, refine the question, and then evaluate whether the resulting analysis can make a meaningful contribution.
Data-informed question development
Available evidence helps refine or inspire a scientifically meaningful question whose rationale can be defended independently of the dataset's convenience.
Data dredging
Large numbers of relationships are searched primarily in the hope of finding an interesting or statistically significant result, without adequate scientific rationale or transparent reporting.
The boundary is not always perfectly sharp. Exploratory analysis can be valuable precisely because researchers do not yet know which patterns deserve explanation. The problem arises when exploration is disguised as confirmation or when a statistically interesting pattern is treated as sufficient scientific justification after the fact.
Ask Whether the Question Matters Without the Convenience
A simple thought experiment can expose weak additions. Imagine that the variable you want to analyze had not already been collected. Would you consider the question important enough to design a study around it or make a serious effort to obtain the necessary evidence?
If the answer is clearly yes, data availability may simply make a worthwhile project more feasible.
If the answer is no and the main argument is, "We already have the variable," the scientific rationale may need more work.
This test should not be interpreted too rigidly. Some worthwhile secondary questions would never justify expensive new data collection because existing data already provide an efficient way to answer them. The point is to separate scientific relevance from mere analytical convenience.
Available Data May Not Be Appropriate Data
The presence of a variable in a dataset can create false confidence. A column with a promising label does not necessarily measure the construct your research question requires.
Methodological guidance on secondary analysis emphasizes examining the original study design, sampling, measures, data quality, missingness, statistical power, and other features before deciding whether existing data can answer a new question.
Suppose a dataset contains one item asking students how confident they feel using digital tools. You are interested in digital competence. It would be convenient to treat that item as a measure of competence, but confidence and competence are not necessarily equivalent constructs. The research question should not be rewritten around a weak proxy simply because the proxy is available.
| Check |
Promising situation |
Warning sign |
| Scientific rationale |
The question addresses a meaningful problem or gap. |
The main justification is that the variables are available. |
| Construct fit |
The existing measures appropriately operationalize the constructs. |
Available variables are weak proxies being stretched to fit the question. |
| Population fit |
The sampled population is appropriate for the inference. |
The question requires a population the dataset does not adequately represent. |
| Temporal fit |
Variables were measured at appropriate times for the question. |
The question implies temporal ordering the data cannot establish. |
| Analytical adequacy |
The sample and data structure can support the proposed analysis. |
The question requires underpowered subgroups or unavailable covariates. |
| Transparency |
The analysis is described according to when and how it was developed. |
A post hoc question is presented as though it had been prespecified. |
Availability Does Not Establish Causal Evidence
Existing datasets frequently contain variables that are statistically associated. That does not mean their relationships can answer causal questions.
For example, a cross-sectional student survey might contain measures of generative AI use and academic performance. Finding an association does not by itself establish whether AI use affected performance, whether academic performance influenced AI use, whether both are related to other factors, or whether the observed association has another explanation.
Researchers should therefore formulate questions that match what the design can support. If the available data permit an associational analysis, do not convert convenience into causal language that the study design cannot justify.
The Question May Not Belong in the Current Study
Suppose the additional question is genuinely worthwhile and the data can answer it. There is still another issue: does it belong in the study you are currently conducting?
A large survey may collect measures relevant to several independent problems. Adding every answerable question to one manuscript or protocol can make the study conceptually incoherent. A scientifically valuable question may be better treated as a separate study.
This is particularly important when the new question requires a different literature, conceptual framework, outcome, analytical strategy, or interpretation. Data collection can be shared even when research projects should remain separate.
The Same Dataset Can Legitimately Support Another Project
Moving a question out of the current study does not necessarily mean collecting new data. Existing datasets can support multiple distinct investigations when each has a substantive purpose and the data are appropriate for the question.
That distinction is especially useful with large surveys, cohorts, registries, administrative databases, and other rich sources. Several questions may share the same dataset while representing different research projects.
Researchers should still consider relevant consent, ethics, governance, licensing, privacy, and reporting requirements for secondary use.
Do Not Confuse Exploration With Confirmation
One of the most consequential issues arises when researchers inspect data, discover an interesting association, develop a hypothesis explaining it, and then report that hypothesis as though it preceded the analysis.
This practice is commonly called HARKing, or hypothesizing after the results are known. The problem is not that researchers learned something from their data. Science should generate new ideas. The problem is concealing the sequence of events and thereby making an exploratory result appear to provide a stronger test of a prior prediction than it actually did.
Watch Out
An unexpected association can be worth reporting and may lead to an important research question. Describe it according to how it was discovered. Do not retrospectively present a data-inspired hypothesis as though the study had been designed to test it from the beginning.
Preregistration Can Clarify What Was Planned
Where appropriate to the methodology, preregistration can help distinguish decisions made before analysis from those made after researchers had access to the data or results. A preregistration can specify research questions, hypotheses, operationalizations, exclusion criteria, and planned analyses before those analyses are conducted.
Preregistration does not make exploratory research illegitimate, nor does the absence of preregistration make research invalid. Its value here is transparency. Readers can more readily distinguish confirmatory tests from analyses that emerged during exploration.
Preexisting datasets require additional care because researchers may already know something about the data. Preregistration frameworks for secondary analysis therefore recommend disclosing prior access, previous analyses, and relevant knowledge of the dataset rather than pretending the data are entirely unseen.
Exploratory Questions Can Be Planned Before Data Collection
Researchers sometimes know in advance that certain relationships are worth exploring but do not have sufficiently developed theory or evidence for strong confirmatory hypotheses. Those questions can still be specified prospectively as exploratory.
Doing so can improve measurement planning without falsely converting exploration into confirmation. Whether you should add exploratory questions before data collection depends on their value, feasibility, and fit with the study.
More Available Variables Mean More Analytical Choices
Rich datasets create analytical flexibility. Researchers may choose among many outcomes, predictors, covariates, subgroups, exclusions, transformations, interaction terms, and models. If enough combinations are tried, some apparently interesting patterns may emerge by chance.
This does not make large datasets undesirable. It means researchers should be especially disciplined about distinguishing prespecified analyses from exploration, accounting for multiplicity where required by the inferential framework, and avoiding selective reporting of only the most favorable findings.
The appropriate response to abundant data is better analytical reasoning, not an obligation to analyze everything.