03 · What You Need to Know
How to Check Whether Existing Data Can Answer Your Question
Start with the question, not the dataset
There are legitimate research projects in which researchers explore an existing dataset and identify questions that its variables can address. There are also question-driven projects in which the research question comes first and researchers search for suitable data. In practice, the process can become iterative: the question guides the search for data, and the realities of available data lead to refinement of the question.
Whichever route you take, avoid allowing the mere availability of a large dataset to substitute for a meaningful research question. Existing-data research still requires alignment among the question, population, measures, design, and analysis.
A useful first step is to translate the question into concrete requirements. Suppose you want to ask whether sustained use of a learning-management system during the first semester predicts second-year student retention. That question immediately implies several data needs: student-level records, a defensible indicator of system use, dates or another measure of timing, first-semester enrollment, subsequent enrollment status, and a way to link those observations for the same students over time.
Once the question has been decomposed in this way, you can investigate whether those elements really exist.
A database existing is not the same as the required data existing
Researchers sometimes reason from the existence of a system to the existence of a variable.
"The university has an LMS, so it must have detailed historical activity logs."
"The hospital has electronic records, so it must have consistent information on this clinical characteristic."
"The government collects employment statistics, so it must have data for this occupation in my municipality."
None of those conclusions necessarily follows. Operational systems are usually created to serve administrative, clinical, educational, financial, regulatory, or other purposes. What they record, how long they retain it, and how consistently they define it may differ considerably from what your research question requires.
A data source exists
A database, registry, survey, archive, administrative system, repository, or collection contains information relevant to the general topic.
The required research data exist
The source contains the particular variables, observations, population, timeframe, level of detail, and relationships among data elements needed to address your question.
Write a data requirements list before searching
Do not begin by asking vaguely whether a provider has "data about student performance" or "health data." Specify what your proposed analysis would require.
Depending on the question, your requirements might include an exposure or predictor, outcome, relevant covariates, participant characteristics, dates, geographical identifiers, repeated observations, identifiers for linkage, institutional characteristics, survey weights, or other design variables.
For each requirement, ask what would count as an acceptable measure. If your question concerns academic achievement, for example, do you need course grades, cumulative grade-point average, standardized test scores, pass or fail status, or another measure? Those variables are not interchangeable simply because all relate broadly to achievement.
This exercise prevents a common problem: finding a dataset that contains variables with promising names and only later discovering that their operational definitions do not match the constructs in the research question.
Read the documentation before assuming a variable is usable
A dataset should be evaluated through its documentation, not merely through its filename or variable list.
Useful documentation may include a data dictionary, codebook, questionnaire, interview schedule, technical report, user guide, sampling documentation, metadata, processing notes, or information about known errors and changes across releases. The UK Data Service specifically advises prospective users to consult study documentation before downloading data so they can assess whether the information collected, population, timing, location, and processing are suitable for their research.
Methodological guidance on secondary analysis makes the same point. Researchers need to understand the population studied, sampling strategy, data-collection period, assessment instruments, response levels, quality-control procedures, and other features of the dataset rather than treating the file as self-explanatory.
A variable called engagement, for example, tells you almost nothing until you know how engagement was defined, measured, coded, and collected.
Check the exact variables, not approximate substitutes
Existing datasets rarely contain every measure a researcher would ideally choose. Sometimes a closely related variable can reasonably operationalize the construct of interest. Sometimes it cannot.
Suppose your question concerns students' critical-thinking ability, but the dataset contains only final course grades. Grades may be important educational outcomes, but they are not automatically valid measures of critical thinking. Renaming the variable in your manuscript does not change what was measured.
Similarly, a dataset containing whether a patient ever received a diagnosis may not support a question about disease severity. Annual household income categories may not support an analysis requiring precise income values. A single survey question about technology use may not support a construct that originally required a multidimensional validated scale.
If the available measure differs from the construct in your question, decide explicitly whether the substitute is conceptually and methodologically defensible. If not, change the question or find another source rather than quietly changing what the variable means.
Check who is represented in the data
The right variables are not enough if they were collected from the wrong population.
A national dataset might exclude particular regions, age groups, institutions, occupations, or people outside formal systems. Administrative records may cover only individuals who used a service. A survey may include university students but not distinguish the subgroup required by your question. A clinical database may contain patients treated at participating facilities rather than all people with the condition.
Study documentation should therefore tell you who was included, how cases entered the dataset, and what population the data were intended to represent.
This distinction matters because a dataset can contain thousands or millions of observations while still containing too few cases from the subgroup you need.
Check whether there are enough relevant cases
Large datasets can create a false sense of abundance.
Imagine a dataset with 100,000 respondents. Your question concerns participants who satisfy four characteristics simultaneously. After restricting the data to the correct age, occupation, geographical area, and exposure, only 83 cases remain. The original dataset is still large; your analytical sample is not.
Where possible, examine frequency tables, documentation, online data explorers, published reports, or preliminary counts to estimate how many observations meet the complete criteria for your intended analysis.
This is particularly important when the study involves rare outcomes, small subgroups, interactions, longitudinal follow-up, or analyses requiring complete information across several variables.
Check the timeframe carefully
Data may exist for the right variables and population but not for the period required by your question.
A variable might have been introduced in 2022 even though the database dates back to 2010. A survey item may have changed wording between waves. An administrative system may retain detailed logs for only a limited period. An institution may have migrated to a new information system, leaving earlier records in a different format.
Longitudinal questions require particular care. The fact that a dataset contains observations from several years does not necessarily mean the same individuals can be followed across those years.
Ask when each required variable was collected, whether definitions changed, whether the relevant cases can be followed over time, and whether the period actually corresponds to the phenomenon named in the research question.
Check the unit and level of detail
Another common mismatch occurs when the data are stored at a different level from the intended analysis.
You may need student-level information but find only school-level averages. You may need monthly observations but receive annual totals. You may need exact ages but the public dataset reports broad age categories. You may need municipality-level geography but the released data identify only regions.
Aggregation is not necessarily a defect. It may be necessary for confidentiality, operational reporting, or the original purpose of the data. It does mean that some questions cannot be answered from that version of the dataset.
Be particularly careful about moving between levels of analysis. Group-level data cannot simply be treated as though they describe individual-level relationships.
Check whether the data can be linked
Some questions require information from more than one file, system, wave, or source.
Suppose one database contains students' learning-platform activity while another contains their academic outcomes. Both required data elements exist. Your study still fails if there is no valid way to connect the same student's records across the two sources.
Linkage may depend on unique identifiers, probabilistic matching, dates, institutional codes, or other linking information. Privacy-protective public-use datasets may intentionally remove identifiers needed for linkage. Data-use rules may also prohibit linking even when it is technically possible.
Before assuming that two useful datasets can become one useful analytical dataset, verify whether linkage is technically possible, methodologically defensible, and permitted.
Check whether the public version differs from the restricted version
Finding documentation showing that a variable was collected does not necessarily mean that variable appears in the dataset you can obtain.
Data providers may create different releases. A public-use file may suppress detailed geography, exact dates, rare categories, identifiers, sensitive measures, or other information to reduce disclosure risk. More detailed data may exist only through a restricted-access process or secure research environment.
This distinction is easy to miss when researchers read a questionnaire showing that a measure was collected and assume it will therefore be present in the downloadable file.
Verify the contents of the specific release you would actually be permitted to use.
Check missingness, completeness, and consistency
A variable can technically exist while being practically unusable.
Perhaps it is missing for most participants. Perhaps one site did not collect it. Maybe the measure became mandatory only halfway through the study period. Administrative data may contain values entered inconsistently because the field was not important for operational purposes.
Do not stop at "Is there a column with this variable name?" Ask how complete the observations are and whether missingness is concentrated in particular years, sites, groups, or circumstances.
The implications depend on the pattern and purpose of the analysis. A variable with 15% missing data is not automatically unusable, just as one with 2% missing data is not automatically harmless. What matters is why information is missing, where the missingness occurs, and how it affects the intended analysis.
Understand how the data were originally produced
Existing data were collected through a particular process for a particular purpose. That history matters.
Survey data depend on questionnaire design, sampling, response, weighting, and field procedures. Administrative records reflect operational processes and definitions. Clinical records reflect healthcare encounters and documentation practices. Platform logs reflect what the software records rather than every behavior a researcher might conceptually care about.
Researchers using existing data should therefore understand how observations entered the dataset, what the original measures meant, and what quality-control or processing procedures were applied.
Good documentation provides the context necessary to interpret data correctly. The UK Data Service emphasizes that documentation should allow a new user to understand how the research was conducted and what the data mean. Without that context, a dataset can be technically available but scientifically difficult to use.
Do not confuse data existence with data access
This distinction is fundamental.
Existence
The required information has actually been collected or generated and is retained in a usable form.
Access
You are permitted and practically able to obtain or use that information for the proposed research.
A hospital may possess exactly the records you need while refusing external research access. A government agency may hold detailed microdata but release only aggregated statistics publicly. A commercial platform may retain extensive behavioral data without providing researchers a mechanism to obtain them.
Confirming that the required data exist is only one part of the feasibility assessment. If the information is available but another person or organization controls whether you can use it, you then need to determine what to do when data access depends on someone else's permission.
Watch Out
Do not write a proposal around a dataset based only on its title, webpage description, or the assumption that an organization "must collect" the information. Inspect the documentation for the specific data release and verify the variables, population, timeframe, level of detail, and other features your analysis actually requires.
Sometimes the data tell you that the question needs to change
Existing-data research often involves iteration between the ideal question and the information actually available. Methodological guidance on secondary analysis recognizes that researchers may need to modify a research question or analytical plan when suitable datasets do not contain all of the required variables.
That does not mean allowing the dataset to dictate an arbitrary question simply because certain columns happen to be available. The revised question still needs intellectual and methodological justification.
Suppose you intended to study whether daily social-media use predicts academic performance, but the dataset measures only whether participants use social media at all. You have at least three choices: find a better dataset, reformulate the question around the measure that actually exists if that question remains worthwhile, or collect new data.
The wrong choice is to keep the original question and behave as though the weaker measure answers it.