03 · What You Need to Know
Let the dataset reveal possibilities without letting it dictate the science
Secondary data analysis is a legitimate research strategy
Researchers frequently analyze data originally collected for another study, administrative purpose, surveillance system, repository, registry, or large research program. Contemporary research infrastructures can make extensive data available without requiring each investigator to recruit a new sample.
For example, the U.S. National Institutes of Health's All of Us Research Program provides eligible researchers with access to data that include electronic health records, surveys, physical measurements, wearable-device data, and, at a controlled access level, genomic information. The program illustrates how a shared dataset can support many research questions beyond a single original study.
Likewise, scholarly infrastructures can themselves become data sources. Crossref exposes deposited scholarly metadata through a public REST API, while DataCite provides public retrieval of DOI metadata. Such resources can enable bibliometric, metadata-quality, scholarly-communication, and related studies without researchers assembling every record manually.
The legitimacy of the research does not depend on whether you personally collected the observations. It depends on whether the data are appropriate for the question and whether your analysis and interpretation respect how those data were produced.
Start by asking what is newly possible
A dataset becomes interesting when it changes the set of questions you can realistically investigate.
| What is distinctive about the dataset? |
What research opportunity might it create? |
| Large sample |
Estimate uncommon outcomes or subgroup patterns with greater precision, when the sample supports such inference |
| Longitudinal observations |
Study change, trajectories, sequencing, or temporal relationships |
| Multiple linked data sources |
Examine relationships that cannot be observed within one source alone |
| Previously underrepresented population |
Investigate questions for populations poorly represented in earlier evidence |
| Fine-grained behavioral records |
Study patterns that retrospective self-report may not capture well |
| Long historical coverage |
Examine trends, transitions, or responses to events across time |
| Newly available variables |
Test relationships or explanations that earlier datasets could not operationalize |
This reasoning is stronger than “the dataset contains many variables.” The opportunity lies in what those features allow you to learn.
Data availability can come before the final research question
It is sometimes taught that researchers should formulate the question first and only then consider data. That sequence is often sensible for primary research, but existing-data research can be more iterative.
You may discover a dataset, inspect its documentation, identify what is measurable, review the relevant literature, and gradually formulate a question that is both important and answerable with those data. There is nothing inherently illegitimate about this sequence.
The risk appears when the process becomes an unrestricted search for statistically interesting patterns followed by a retrospective story about why the discovered relationship was supposedly the question all along.
Watch Out
Exploring a dataset to generate hypotheses is legitimate. Presenting a relationship discovered through extensive exploration as though it were a prespecified confirmatory test is not. Keep exploratory discovery and confirmatory testing distinguishable in your reasoning and reporting.
Understand why and how the data were collected
Existing data were produced for a purpose, and that purpose shapes what they contain. Administrative records are designed primarily to support administration. Electronic health records support care. Learning-management-system logs record particular interactions with a platform. Bibliographic metadata depend on what participating organizations deposit and how metadata fields are populated.
These sources can be extremely useful, but researchers should not assume that a field in a database perfectly represents the theoretical construct they wish to study.
Before analyzing an unfamiliar dataset, investigate:
- the original purpose of data collection;
- the target and observed populations;
- sampling, recruitment, inclusion, and exclusion processes;
- how each important variable was generated or measured;
- changes in measurement or data systems over time;
- missingness and data-quality procedures;
- linkage procedures when multiple sources are combined;
- known limitations and access restrictions.
A codebook is not optional bedtime reading in secondary-data research. It is part of understanding what your evidence actually means.
A variable name is not a construct
Suppose a dataset contains a variable called “engagement.” Before using it as a measure of student engagement, determine how it was constructed. It might represent login frequency, assignment completion, a survey scale, attendance, or a composite score produced by an undocumented algorithm.
Each operationalization supports different interpretations.
The fact that a dataset contains a conveniently named variable does not establish construct validity. If the available measure only partially represents your intended concept, you may need to narrow the claim, choose another variable, combine appropriate indicators, or conclude that the dataset cannot answer that particular question.
Large datasets do not automatically eliminate bias
A dataset can contain hundreds of thousands of observations and still represent its target population poorly. Increasing sample size reduces some forms of sampling uncertainty, but it does not automatically repair systematic selection, measurement error, missingness, confounding, or other sources of bias.
Ask who enters the dataset and who does not. Participation may depend on healthcare access, institutional enrollment, platform use, consent, geography, technology access, administrative procedures, or other mechanisms related to the phenomenon being studied.
The NIH All of Us program, for example, explicitly describes its dataset as drawing on a diverse participant cohort and uses tiered access and privacy safeguards. Those features are important to understand when determining what research is feasible and how findings should be interpreted.
A new dataset does not automatically create causal evidence
Observational data can support important descriptive, associational, predictive, and, under suitable designs and assumptions, causal analyses. But a dataset does not become causal merely because it is large, longitudinal, detailed, or computationally impressive.
If your question asks whether one exposure causes an outcome, you need to consider confounding, selection, temporal ordering, measurement, missing data, and the assumptions required by your analytical strategy.
Sometimes the best research question supported by a dataset is descriptive rather than causal. That is not a methodological consolation prize. Accurate description can be an important scientific contribution when the phenomenon is poorly characterized.
New data can make subgroup analysis possible, but caution remains necessary
Large datasets can provide enough observations to examine variation across groups that smaller studies could not estimate reliably. This may create valuable questions about heterogeneity or equity.
However, repeatedly dividing data into subgroups until an interesting difference appears increases the opportunity for chance findings. Some categories may also contain too few observations for stable estimates, even in an otherwise large dataset.
Subgroup questions should therefore be driven by substantive reasoning where possible, and uncertainty should remain visible in the interpretation.
Access does not mean unrestricted use
Datasets may be publicly downloadable, available only through controlled environments, subject to data-use agreements, or restricted according to sensitivity. Access conditions can determine whether a proposed study is feasible.
Researchers should examine requirements involving ethics review, institutional agreements, researcher registration, privacy, disclosure control, publication review, security, permitted analyses, and data retention before designing a project around a resource.
For example, access to the All of Us Researcher Workbench requires institutional and researcher conditions, and its registered and controlled tiers provide different levels of data. A theoretically perfect question is not yet a feasible project if you cannot lawfully or practically obtain the necessary variables.
The dataset may create a methodological question rather than a substantive one
Sometimes the most interesting feature of new data is not a previously unstudied relationship but a new way of observing something. Researchers might compare a digital trace with a conventional self-report measure, evaluate missing-data patterns, investigate metadata completeness, or test whether a new data source adequately represents an established construct.
If the central opportunity concerns how something can now be measured, the research problem may be closer to one created by a new measurement tool than by the substantive contents of the dataset itself.
A dataset can also expose questions you did not know to ask
Exploratory visualization and descriptive analysis can reveal unexpected patterns. Perhaps a trend reverses after a particular period, a subgroup behaves differently, or a distribution looks nothing like the literature led you to expect.
That discovery can legitimately generate a new question. The reasoning then resembles an unexpected observation becoming a research idea: verify the pattern, consider artifacts and alternative explanations, examine prior evidence, and determine whether independent or confirmatory analysis is warranted.
Data can therefore participate in question generation without being allowed to manufacture certainty after the fact.
04 · A Practical Example
From “we have years of student data” to a defensible research question
Hypothetical Example
A university opens a longitudinal learning dataset
A university makes an approved de-identified dataset available to researchers containing several years of enrollment information, course outcomes, selected student characteristics, and learning-management-system activity. A researcher initially thinks, “There must be dozens of papers in this dataset.”
Inspect the data-generating process The researcher learns which students are represented, when the learning platform was introduced, how activity logs are generated, which courses use the platform consistently, and where substantial missingness occurs.
Identify a distinctive opportunity The longitudinal structure allows patterns of course participation to be observed across multiple semesters rather than at one point in time.
Review the literature The researcher examines prior evidence on academic persistence, online activity, course engagement, and related constructs rather than assuming that the available variables define the theoretical framework.
Reject an overclaim Login frequency alone is not treated as a comprehensive measure of student engagement, and observational associations are not automatically interpreted as causal effects.
Refine the question The researcher asks whether changes in patterns of course-platform activity across successive semesters precede changes in course completion among students with different prior academic trajectories.
Define the contribution The dataset matters because repeated observations make temporal patterns visible that a single-semester survey could not examine in the same way.
The research opportunity came from a distinctive property of the dataset. The question still required literature, conceptual reasoning, and methodological restraint.
07 · A Quick Checklist
Before building a study around an existing dataset
Before committing to the analysis, check:
Identify what is distinctive about the dataset and what substantive question that feature makes possible.
Read the codebook, technical documentation, data dictionary, and relevant methodological documentation before finalizing the question.
Determine why the data were originally collected and how that purpose shaped the available variables.
Examine who is represented, who is excluded or missing, and what selection processes generated the observed sample.
Verify how your central variables were defined, measured, coded, and, where relevant, validated.
Assess missingness, changes in data collection, linkage quality, and other relevant data-quality issues.
Match causal, descriptive, predictive, or associational claims to what the dataset and design can actually support.
Distinguish exploratory analyses from confirmatory tests and document important analytical decisions.
Verify ethics, access, privacy, security, licensing, attribution, and data-use requirements before beginning the project.