Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Know Whether the Data You Need Actually Exist?

Knowing that an organization collects data is not enough. Learn how to verify whether the specific variables, cases, time periods, detail, and documentation your research question requires actually exist.

444
Do the Data You Need Actually Exist? Guide 444 of 533
01 · The Question

You Know Data Are Collected, but Are the Data You Need Actually There?

Your research question appears perfectly suited to existing data. The university has student records. The hospital uses an electronic health record. A government agency maintains a national database. A large survey has been conducted for years. The organization surely tracks the information you need.

Then you inspect the actual dataset.

The variable you assumed would be there was never collected. The outcome exists only for recent years. A measure is available, but only as a broad category rather than the detail your analysis requires. Two variables are stored in separate systems without a reliable way to link them. The population you care about cannot be identified. Or the public-use version omits precisely the information your question depends on.

This is a common feasibility problem in research based on existing records or datasets. Before designing a study around data that supposedly exist, you need to determine whether the specific evidence required by your research question exists in a usable form.

02 · The Short Answer

Verify the Variables, Cases, Timeframe, and Detail Before You Commit

In Brief

To determine whether the data you need actually exist, translate your research question into specific data requirements and verify those requirements against the dataset's documentation, including its variables, measures, population, sampling or coverage, timeframe, unit of analysis, level of detail, missingness, and relevant data-quality information.

Do not rely on the dataset's title, a general description, or the fact that an organization routinely collects data. A dataset can be highly valuable and still be incapable of answering your particular question.

03 · What You Need to Know

How to Check Whether Existing Data Can Answer Your Question

Start with the question, not the dataset

There are legitimate research projects in which researchers explore an existing dataset and identify questions that its variables can address. There are also question-driven projects in which the research question comes first and researchers search for suitable data. In practice, the process can become iterative: the question guides the search for data, and the realities of available data lead to refinement of the question.

Whichever route you take, avoid allowing the mere availability of a large dataset to substitute for a meaningful research question. Existing-data research still requires alignment among the question, population, measures, design, and analysis.

A useful first step is to translate the question into concrete requirements. Suppose you want to ask whether sustained use of a learning-management system during the first semester predicts second-year student retention. That question immediately implies several data needs: student-level records, a defensible indicator of system use, dates or another measure of timing, first-semester enrollment, subsequent enrollment status, and a way to link those observations for the same students over time.

Once the question has been decomposed in this way, you can investigate whether those elements really exist.

A database existing is not the same as the required data existing

Researchers sometimes reason from the existence of a system to the existence of a variable.

"The university has an LMS, so it must have detailed historical activity logs."

"The hospital has electronic records, so it must have consistent information on this clinical characteristic."

"The government collects employment statistics, so it must have data for this occupation in my municipality."

None of those conclusions necessarily follows. Operational systems are usually created to serve administrative, clinical, educational, financial, regulatory, or other purposes. What they record, how long they retain it, and how consistently they define it may differ considerably from what your research question requires.

A data source exists A database, registry, survey, archive, administrative system, repository, or collection contains information relevant to the general topic.
The required research data exist The source contains the particular variables, observations, population, timeframe, level of detail, and relationships among data elements needed to address your question.

Write a data requirements list before searching

Do not begin by asking vaguely whether a provider has "data about student performance" or "health data." Specify what your proposed analysis would require.

Depending on the question, your requirements might include an exposure or predictor, outcome, relevant covariates, participant characteristics, dates, geographical identifiers, repeated observations, identifiers for linkage, institutional characteristics, survey weights, or other design variables.

For each requirement, ask what would count as an acceptable measure. If your question concerns academic achievement, for example, do you need course grades, cumulative grade-point average, standardized test scores, pass or fail status, or another measure? Those variables are not interchangeable simply because all relate broadly to achievement.

This exercise prevents a common problem: finding a dataset that contains variables with promising names and only later discovering that their operational definitions do not match the constructs in the research question.

Read the documentation before assuming a variable is usable

A dataset should be evaluated through its documentation, not merely through its filename or variable list.

Useful documentation may include a data dictionary, codebook, questionnaire, interview schedule, technical report, user guide, sampling documentation, metadata, processing notes, or information about known errors and changes across releases. The UK Data Service specifically advises prospective users to consult study documentation before downloading data so they can assess whether the information collected, population, timing, location, and processing are suitable for their research.

Methodological guidance on secondary analysis makes the same point. Researchers need to understand the population studied, sampling strategy, data-collection period, assessment instruments, response levels, quality-control procedures, and other features of the dataset rather than treating the file as self-explanatory.

A variable called engagement, for example, tells you almost nothing until you know how engagement was defined, measured, coded, and collected.

Check the exact variables, not approximate substitutes

Existing datasets rarely contain every measure a researcher would ideally choose. Sometimes a closely related variable can reasonably operationalize the construct of interest. Sometimes it cannot.

Suppose your question concerns students' critical-thinking ability, but the dataset contains only final course grades. Grades may be important educational outcomes, but they are not automatically valid measures of critical thinking. Renaming the variable in your manuscript does not change what was measured.

Similarly, a dataset containing whether a patient ever received a diagnosis may not support a question about disease severity. Annual household income categories may not support an analysis requiring precise income values. A single survey question about technology use may not support a construct that originally required a multidimensional validated scale.

If the available measure differs from the construct in your question, decide explicitly whether the substitute is conceptually and methodologically defensible. If not, change the question or find another source rather than quietly changing what the variable means.

Check who is represented in the data

The right variables are not enough if they were collected from the wrong population.

A national dataset might exclude particular regions, age groups, institutions, occupations, or people outside formal systems. Administrative records may cover only individuals who used a service. A survey may include university students but not distinguish the subgroup required by your question. A clinical database may contain patients treated at participating facilities rather than all people with the condition.

Study documentation should therefore tell you who was included, how cases entered the dataset, and what population the data were intended to represent.

This distinction matters because a dataset can contain thousands or millions of observations while still containing too few cases from the subgroup you need.

Check whether there are enough relevant cases

Large datasets can create a false sense of abundance.

Imagine a dataset with 100,000 respondents. Your question concerns participants who satisfy four characteristics simultaneously. After restricting the data to the correct age, occupation, geographical area, and exposure, only 83 cases remain. The original dataset is still large; your analytical sample is not.

Where possible, examine frequency tables, documentation, online data explorers, published reports, or preliminary counts to estimate how many observations meet the complete criteria for your intended analysis.

This is particularly important when the study involves rare outcomes, small subgroups, interactions, longitudinal follow-up, or analyses requiring complete information across several variables.

Check the timeframe carefully

Data may exist for the right variables and population but not for the period required by your question.

A variable might have been introduced in 2022 even though the database dates back to 2010. A survey item may have changed wording between waves. An administrative system may retain detailed logs for only a limited period. An institution may have migrated to a new information system, leaving earlier records in a different format.

Longitudinal questions require particular care. The fact that a dataset contains observations from several years does not necessarily mean the same individuals can be followed across those years.

Ask when each required variable was collected, whether definitions changed, whether the relevant cases can be followed over time, and whether the period actually corresponds to the phenomenon named in the research question.

Check the unit and level of detail

Another common mismatch occurs when the data are stored at a different level from the intended analysis.

You may need student-level information but find only school-level averages. You may need monthly observations but receive annual totals. You may need exact ages but the public dataset reports broad age categories. You may need municipality-level geography but the released data identify only regions.

Aggregation is not necessarily a defect. It may be necessary for confidentiality, operational reporting, or the original purpose of the data. It does mean that some questions cannot be answered from that version of the dataset.

Be particularly careful about moving between levels of analysis. Group-level data cannot simply be treated as though they describe individual-level relationships.

Check whether the data can be linked

Some questions require information from more than one file, system, wave, or source.

Suppose one database contains students' learning-platform activity while another contains their academic outcomes. Both required data elements exist. Your study still fails if there is no valid way to connect the same student's records across the two sources.

Linkage may depend on unique identifiers, probabilistic matching, dates, institutional codes, or other linking information. Privacy-protective public-use datasets may intentionally remove identifiers needed for linkage. Data-use rules may also prohibit linking even when it is technically possible.

Before assuming that two useful datasets can become one useful analytical dataset, verify whether linkage is technically possible, methodologically defensible, and permitted.

Check whether the public version differs from the restricted version

Finding documentation showing that a variable was collected does not necessarily mean that variable appears in the dataset you can obtain.

Data providers may create different releases. A public-use file may suppress detailed geography, exact dates, rare categories, identifiers, sensitive measures, or other information to reduce disclosure risk. More detailed data may exist only through a restricted-access process or secure research environment.

This distinction is easy to miss when researchers read a questionnaire showing that a measure was collected and assume it will therefore be present in the downloadable file.

Verify the contents of the specific release you would actually be permitted to use.

Check missingness, completeness, and consistency

A variable can technically exist while being practically unusable.

Perhaps it is missing for most participants. Perhaps one site did not collect it. Maybe the measure became mandatory only halfway through the study period. Administrative data may contain values entered inconsistently because the field was not important for operational purposes.

Do not stop at "Is there a column with this variable name?" Ask how complete the observations are and whether missingness is concentrated in particular years, sites, groups, or circumstances.

The implications depend on the pattern and purpose of the analysis. A variable with 15% missing data is not automatically unusable, just as one with 2% missing data is not automatically harmless. What matters is why information is missing, where the missingness occurs, and how it affects the intended analysis.

Understand how the data were originally produced

Existing data were collected through a particular process for a particular purpose. That history matters.

Survey data depend on questionnaire design, sampling, response, weighting, and field procedures. Administrative records reflect operational processes and definitions. Clinical records reflect healthcare encounters and documentation practices. Platform logs reflect what the software records rather than every behavior a researcher might conceptually care about.

Researchers using existing data should therefore understand how observations entered the dataset, what the original measures meant, and what quality-control or processing procedures were applied.

Good documentation provides the context necessary to interpret data correctly. The UK Data Service emphasizes that documentation should allow a new user to understand how the research was conducted and what the data mean. Without that context, a dataset can be technically available but scientifically difficult to use.

Do not confuse data existence with data access

This distinction is fundamental.

Existence The required information has actually been collected or generated and is retained in a usable form.
Access You are permitted and practically able to obtain or use that information for the proposed research.

A hospital may possess exactly the records you need while refusing external research access. A government agency may hold detailed microdata but release only aggregated statistics publicly. A commercial platform may retain extensive behavioral data without providing researchers a mechanism to obtain them.

Confirming that the required data exist is only one part of the feasibility assessment. If the information is available but another person or organization controls whether you can use it, you then need to determine what to do when data access depends on someone else's permission.

Watch Out

Do not write a proposal around a dataset based only on its title, webpage description, or the assumption that an organization "must collect" the information. Inspect the documentation for the specific data release and verify the variables, population, timeframe, level of detail, and other features your analysis actually requires.

Sometimes the data tell you that the question needs to change

Existing-data research often involves iteration between the ideal question and the information actually available. Methodological guidance on secondary analysis recognizes that researchers may need to modify a research question or analytical plan when suitable datasets do not contain all of the required variables.

That does not mean allowing the dataset to dictate an arbitrary question simply because certain columns happen to be available. The revised question still needs intellectual and methodological justification.

Suppose you intended to study whether daily social-media use predicts academic performance, but the dataset measures only whether participants use social media at all. You have at least three choices: find a better dataset, reformulate the question around the measure that actually exists if that question remains worthwhile, or collect new data.

The wrong choice is to keep the original question and behave as though the weaker measure answers it.

04 · A Practical Example

When the Dataset Exists but the Study Still Does Not

Hypothetical Example

Using university records to study student engagement and retention

A researcher proposes the question: "Does first-year engagement with the university's learning-management system predict whether students remain enrolled in their third year?"

The university has used the learning-management system for many years and maintains enrollment records. At first glance, the study appears ideal for existing-data analysis.

Translate the question into data requirements The researcher needs student-level LMS activity during the first year, an identifier that can connect those records to enrollment data, third-year enrollment status, the relevant cohort years, and suitable covariates for the planned analysis.
Inspect the documentation The researcher learns that detailed LMS activity logs are retained for only two years, although the platform itself has been used much longer.
Check linkage Recent LMS records can be linked to student enrollment records through an institutional identifier under appropriate procedures.
Check the outcome period Students in the cohorts with detailed first-year LMS logs have not yet reached the third year, so the proposed outcome does not yet exist for them.
Reach the feasibility conclusion Both information systems exist, but the combination of exposure and outcome data required by the original longitudinal question does not yet exist for the same cohort.
Decide what to change The researcher can wait until the outcome becomes available, identify another historical source, choose an earlier outcome that is substantively defensible, or reformulate the research question.

The important discovery occurred before analysis and, ideally, before the proposal became fixed. Asking whether "LMS data and enrollment data exist" would have produced a misleading yes. Asking whether the specific longitudinal observations required by the question coexist for the same students produces the answer that matters.

05 · What Researchers Often Get Wrong

Common Mistakes When Checking Whether Data Exist

Misconception

If an organization uses a digital system, the historical data must be available

Systems may overwrite, aggregate, archive, or delete information according to operational and retention requirements. Features can also change over time. The current system may display information that was not recorded in the same way historically. Verify actual retention and coverage rather than inferring them from the age of the system.

Misconception

If the dataset contains a variable with the right name, I have the measure I need

Variable labels can conceal important differences in definition, measurement, timing, coding, and source. Read the codebook, questionnaire, technical documentation, and other relevant materials before deciding that a variable represents your construct.

Misconception

A huge dataset will contain enough cases for almost any analysis

The relevant number is not the dataset's total size but the number of observations satisfying your analytical criteria. Rare populations, outcomes, combinations of characteristics, missing values, and longitudinal requirements can reduce a very large dataset to a small usable sample.

Misconception

If two required variables exist somewhere, I can combine them

Variables stored in separate datasets or systems may not share a usable identifier, may refer to different populations or periods, or may be prohibited from linkage. Verify the linkage mechanism before designing an analysis that depends on combining sources.

Misconception

If a variable was collected, it will be in the dataset I download

Public-use and restricted-access releases may differ. Sensitive variables, detailed geography, dates, identifiers, and rare categories may be suppressed or transformed. Check the documentation for the exact version of the data you expect to use.

Misconception

If the data exist, my study is feasible

Existence is only one requirement. You may still need permission, sufficient documentation, appropriate analytical expertise, adequate computing resources, and enough time to obtain, clean, understand, and analyze the data. Data existence should therefore be treated as one part of the broader feasibility assessment.

06 · What This Means for You

Conduct a Data Feasibility Check Before Building the Study Around the Dataset

Before committing to an existing-data study, create a simple map from the research question to the evidence required to answer it. For every important concept, identify the variable or data element that would represent it. Then verify the population, timeframe, unit, measurement, completeness, linkage, and documentation associated with those data.

The purpose is not to demand a perfect dataset. Perfect datasets are rare enough to deserve their own mythology. The purpose is to identify compromises before they quietly become methodological limitations you discover halfway through the analysis.

A simple decision framework

If all essential variables exist with appropriate definitions, population coverage, timeframe, and usable detail
The dataset may be suitable enough to evaluate further, including its quality, access conditions, analytical requirements, and other limitations.
If an ideal measure is absent but a defensible alternative exists
Determine whether the alternative genuinely represents the construct and whether the research question should be revised to match what is actually measured.
If essential variables exist but cannot be linked across the required files, systems, or waves
Do not design the analysis as though linkage were available. Investigate another source, another design, or a revised question.
If the necessary information was collected but is absent from the public-use release
Investigate whether an appropriate restricted-access version exists and whether you are eligible and realistically able to obtain it.
If a variable central to the question was never collected or does not exist for the required population or period
Find another dataset, collect new data, or reformulate the question rather than pretending an inadequate proxy answers the original question.

Once you have established that a promising dataset really does contain the necessary information, the next step is not immediately to analyze it. You still need to determine what should be checked before building the study around that existing dataset, including how it was produced, its quality, documentation, limitations, and suitability for the intended analysis.

07 · A Quick Checklist

Do the Data Required by Your Research Question Actually Exist?

Before designing the study around existing data, check:
Translate every essential part of the research question into a specific variable, measure, observation, or data element you would need.
Inspect the data dictionary, codebook, questionnaire, technical documentation, or equivalent source rather than relying on the dataset's title or description.
Verify how each important variable was defined, measured, coded, and collected.
Confirm that the dataset includes the population and enough relevant cases required by your intended analysis.
Check whether every required variable exists for the same relevant timeframe, cohort, wave, or observation period.
Verify that the unit of analysis and level of detail match the question, such as individual rather than aggregate data where individual-level analysis is required.
If multiple files or sources are required, confirm that they can be validly and permissibly linked.
Check whether important variables have substantial or systematic missingness, inconsistent collection, or changes in definition over time.
Verify the contents of the specific public, licensed, or restricted data release you could actually use.
Keep data existence separate from data access: after confirming the information exists, determine whether you can realistically obtain permission to use it.
08 · Frequently Asked Questions

Frequently Asked Questions About Research Data Availability

How can I find out what variables an existing dataset contains?

Look for the dataset's data dictionary, codebook, questionnaire, user guide, metadata, technical report, variable catalogue, or online data explorer. These resources can show not only variable names but also definitions, response categories, coding, collection periods, sampling information, and other details needed to judge suitability.

Is a data dictionary enough to decide whether a dataset is suitable?

Usually not by itself. A data dictionary can identify variables and coding, but you may also need documentation about sampling, data collection, instruments, weighting, processing, missing values, changes across waves, quality control, and known limitations. Suitability depends on how the data were produced as well as which columns exist.

What if the dataset has a similar variable but not exactly the measure I wanted?

Determine whether the available variable is a defensible operationalization of the construct in your research question. If it measures something meaningfully different, do not simply relabel it. Find another source or revise the question so that your claims correspond to what was actually measured.

What if the data exist only for some years?

Check whether the available period can still answer a worthwhile version of the research question. If the missing years are essential to the intended comparison, trend, exposure, or follow-up, the dataset may not support the original study. Also verify whether variable definitions and collection procedures changed across the years that are available.

Can I combine two datasets if each contains part of what I need?

Possibly, but only if there is a defensible linkage or integration strategy. Check whether the datasets refer to compatible populations and periods, whether appropriate identifiers or linking variables exist, how accurately records can be matched, and whether the applicable data-use conditions permit linkage.

What if the dataset is enormous but my subgroup is very small?

Evaluate the number of cases remaining after applying all eligibility criteria and analytical requirements. The total dataset size does not determine whether your subgroup analysis is viable. Rare subgroups, missing data, longitudinal requirements, and other restrictions can sharply reduce the usable analytical sample.

Does publicly available mean all of the collected data are public?

No. Providers may release public-use versions that suppress, aggregate, recode, or remove information that could create disclosure or other risks. More detailed information may exist in restricted files or secure environments. Always verify the contents of the particular release available to you.

What if the data I need do not exist?

You generally have three broad options: locate another suitable source, collect the required data yourself if feasible, or revise the research question so that it can be answered credibly with evidence that actually exists. What you should not do is preserve the original question while substituting variables that do not adequately represent what it asks.

09 · The Bottom Line

Do Not Ask Only Whether a Dataset Exists

The Bottom Line

The data your research requires exist only when the specific variables, measures, cases, population, timeframe, level of detail, and relationships among data elements needed to answer your question have actually been collected or generated and retained in a usable form.

Verify those requirements against documentation before committing to the study. A database can exist without containing the evidence your question requires, and a variable can exist without being measured in the way your question assumes. Discovering that mismatch early gives you time to find another source, collect new data, or revise the question before the dataset starts making methodological decisions on your behalf.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes