Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Should You Check Before Building a Study Around an Existing Dataset?

Finding a dataset with the variables you need is only the beginning. Learn how to evaluate its design, population, measures, quality, missingness, documentation, access conditions, and analytical fit before building your study around it.

445
What to Check Before Using an Existing Dataset Guide 445 of 533
01 · The Question

You Found a Promising Dataset. Should You Build Your Study Around It?

An existing dataset can make an ambitious research project suddenly look feasible. The participants have already been recruited. The observations have already been collected. Perhaps the dataset contains thousands of cases, years of records, expensive measurements, or a population you could never realistically recruit yourself.

That is an opportunity, but it also changes the research problem.

You did not design how those data were collected. You may not have chosen the population, measures, timing, sampling strategy, instruments, coding rules, or procedures for handling missing information. The dataset carries those decisions into your study whether you would have made them yourself or not.

Before building a thesis, dissertation, article, or other project around an existing dataset, you therefore need to ask more than whether the required variables are present. You need to determine whether the dataset is methodologically suitable for the particular question you want to answer.

02 · The Short Answer

Evaluate the Dataset as Carefully as You Would Design New Data Collection

In Brief

Before building a study around an existing dataset, examine why and how the data were collected, who and what they represent, how the important variables were measured, whether the design fits your question, how complete and reliable the data appear to be, what documentation is available, whether the sample supports your analysis, and whether you can legally and practically use the data as intended.

Existing data can save substantial time and resources, but convenience does not establish suitability. The dataset should fit the research question closely enough that its original design and limitations do not undermine the conclusions you intend to draw.

03 · What You Need to Know

How to Evaluate an Existing Dataset Before Committing to It

First confirm that the data you need actually exist

Before evaluating a dataset in depth, establish that it contains the essential information required by your research question.

That means checking the actual variables and measures, relevant population, timeframe, unit of analysis, level of detail, and any linkage required across files or waves. A dataset can be excellent in every other respect and still be unsuitable because it does not contain the evidence your particular question requires.

If you have not completed that check, first determine whether the required research data actually exist. Once they do, the next question is whether those data are good enough and appropriately structured for the study you intend to conduct.

Understand why the data were originally collected

Existing data were created for a purpose, and that purpose influences what was recorded and how.

A national survey may have been designed to estimate population characteristics. Clinical records primarily support patient care. University records support academic administration. Learning-management-system logs are generated by platform activity. Government administrative databases may exist for regulatory or service-delivery purposes. Data from an earlier research project were collected to answer that project's questions, not necessarily yours.

This matters because data collected for another purpose may emphasize different constructs, populations, time intervals, or levels of precision from those your study requires.

Secondary-analysis guidance therefore recommends understanding the original study or data-generation process rather than treating an existing file as a neutral collection of variables. Cheng and Phillips, for example, emphasize examining the study population, sampling strategy, assessment methods, response levels, quality-control procedures, and other features of the original data before conducting a new analysis.

Reconstruct how the dataset came into existence

A useful evaluation begins with provenance: where did these observations come from?

For a research dataset, examine the original study design, sampling strategy, recruitment, instruments, data-collection procedures, study period, follow-up, and processing. For administrative or operational data, determine how records enter the system, who records them, why they are recorded, what changes occurred over time, and what quality controls apply.

Ask whether the analytical file is raw, cleaned, transformed, imputed, weighted, aggregated, de-identified, or otherwise processed. If variables were derived, find out how.

You should be able to trace important variables back to the process that produced them. If you cannot determine what an observation represents or how it was created, interpreting it confidently becomes difficult.

Check whether the population matches your question

A dataset may contain the right variables but represent the wrong people, institutions, events, locations, or period.

Suppose you want to investigate digital-learning behavior among university students generally, but the dataset contains only students enrolled in online degree programs. Or perhaps you want to study employees nationally but the administrative dataset contains only workers who applied for a particular government benefit.

The question is not simply whether your desired subgroup appears somewhere in the file. You need to understand how cases entered the dataset and what population the data can reasonably represent.

For survey data, this includes examining the target population, sampling frame, sampling design, response, and weights where applicable. For administrative records, it includes understanding the service, institution, or process that generated inclusion in the database.

Do not let a huge sample distract you from poor population fit

Large datasets are seductive. A file containing 500,000 observations can look inherently more impressive than a study of 500 carefully selected participants.

Sample size does not solve a mismatch between the dataset and the research question.

If the database systematically excludes people relevant to your target population, increasing the number of observations from the included population does not make those excluded people appear. Likewise, a massive convenience or administrative dataset does not become population-representative simply because it is massive.

Methodological guidance on secondary dataset analysis makes a similar point: a large sample is valuable only when the dataset fits a meaningful question, appropriate population, measures, and analytical strategy.

Examine how the important constructs were measured

The variables in an existing dataset were defined before your current research question existed. You therefore inherit their operationalizations.

Suppose your question concerns student engagement. The dataset contains number of logins. Does login frequency adequately represent the type of engagement your theoretical argument concerns? Perhaps. Perhaps not.

Or suppose you want to study psychological distress but have only one general well-being item. You cannot turn that item into a validated distress scale by giving the variable a more ambitious label.

For each central construct, examine the original instrument, question wording, response categories, scoring, timing, units, and evidence supporting the measure when relevant. If you use a proxy, explain why the proxy is reasonable and what it does not capture.

Smith and colleagues note that available measures in secondary datasets rarely correspond perfectly to every construct researchers might ideally want. The question is whether the available measurement is sufficiently appropriate for the specific inference you intend to make.

Check whether definitions changed over time

Long-running datasets deserve special scrutiny because apparent continuity can conceal measurement changes.

A survey question may be reworded. Diagnostic criteria can change. Administrative categories may be reclassified. A university may alter grading policies. A platform update can change what counts as an activity event. A variable may move from optional to mandatory entry.

If you intend to examine trends or combine multiple waves, determine whether important variables are comparable across the entire period. A column with the same name in several years does not guarantee that it represents exactly the same construct in each year.

When harmonization is required, document what was combined, recoded, or excluded and what comparability assumptions remain.

Inspect missing data before deciding the dataset is usable

Missingness should be investigated variable by variable and, when relevant, across subgroups, sites, waves, or periods.

A dataset may contain your outcome variable for 98% of observations but an essential predictor for only 60%. Or the overall missingness may look modest while nearly all missing observations come from one region or participant group.

Start descriptively. Determine how much information is missing, where it is missing, and whether important variables are missing together. Then consider plausible reasons for those patterns and what they mean for your analysis.

Do not assume that deleting incomplete cases is automatically appropriate. Neither should you assume that a sophisticated missing-data technique can repair every problem. The defensibility of an approach depends on the data-generating and missingness processes, analytical method, and assumptions involved.

Look for impossible, inconsistent, or suspicious values

Data quality assessment should go beyond missing cells.

Check ranges, categories, units, duplicates, dates, identifiers, internal consistency, and relationships among variables. An age of 240 years, a graduation date before enrollment, duplicate identifiers assigned to different people, or a score outside an instrument's possible range warrants investigation.

Not every surprising observation is an error. Unusual cases can be real and scientifically important. The purpose of checking is to distinguish plausible extremes from coding, entry, linkage, or processing problems rather than automatically deleting anything inconvenient.

When a provider supplies data-quality notes, cleaning documentation, validation reports, or known-issues files, read them before inventing your own explanation for anomalies.

Understand derived variables and recoding

Some of the most convenient variables in an existing dataset may not have been collected directly.

Variables such as socioeconomic status, disease classifications, composite scores, risk groups, geographic categories, engagement indices, or household measures may have been derived from several source fields. Their definitions may contain assumptions that matter for your analysis.

Inspect how important derived variables were constructed. If the documentation provides the underlying components, determine whether using the provider's derived variable or constructing your own version better matches the research question and permitted use of the data.

Do not assume that a professionally prepared derived variable is conceptually appropriate for every research purpose simply because it is convenient.

Check whether the design supports the relationship you want to study

A dataset may contain both your predictor and outcome without supporting the inference implied by your question.

For example, a cross-sectional survey may contain measures of generative-AI use and academic confidence collected at the same time. It can support analyses of association under appropriate assumptions. It does not by itself establish that AI use changed later confidence or that one variable caused the other.

Likewise, repeated cross-sectional surveys conducted every year are not the same as longitudinal data following the same individuals over time.

Variables coexist in the dataset The dataset contains measures of both concepts required for the proposed analysis.
The design supports the intended inference The timing, sampling, measurement, and structure of the data are appropriate for the relationship or change the research question asks you to estimate.

The second requirement is more demanding. Existing data do not become longitudinal, experimental, representative, or causal merely because the desired variables happen to appear together.

Understand the sampling design before analyzing survey data

Large surveys may use stratification, clustering, unequal selection probabilities, oversampling, multistage sampling, and survey weights. Those design features can affect estimation and uncertainty.

If a dataset provides weights, strata, clusters, replicate weights, or other design variables, determine why they exist and whether your planned analysis needs to account for them. Ignoring a complex sampling design can produce estimates or standard errors that do not correspond to the way the sample was actually obtained.

The relevant documentation should explain the sampling and weighting procedures. If you cannot understand how to incorporate them, that may be a methodological skill gap to address before proceeding rather than a reason to quietly analyze the file as though it came from a simple random sample.

Check whether you have enough usable observations for the planned analysis

The total number of rows in the dataset is rarely the number available to every analysis.

Your usable sample may shrink after restricting to the appropriate population, selecting the relevant years, requiring valid observations on essential variables, linking records, or applying design criteria.

More complex analyses can create additional requirements. A multilevel analysis needs sufficient information at relevant levels. Longitudinal models require repeated observations. Rare outcomes may leave few events despite a large overall sample.

Doolan and Froelicher specifically identify available statistical power and data quality as considerations when determining whether an existing dataset is adequate for a proposed secondary analysis.

Evaluate the analytical sample you will actually have, not the sample size printed on the dataset's homepage.

Determine whether documentation is sufficient to interpret the data

A dataset without adequate documentation can be difficult to use responsibly even when the file itself opens perfectly.

Look for a codebook or data dictionary, questionnaires or instruments, study-design documentation, sampling information, fieldwork procedures, variable descriptions, coding information, derivation rules, weighting guidance, missing-value codes, version history, and known limitations where applicable.

Good documentation allows another researcher to understand what the variables mean and how the observations were generated. Metadata and data dictionaries are particularly important when datasets are shared or combined because they preserve the context required for reuse.

If you cannot tell whether 9 means "missing," "not applicable," or an actual response category, you do not yet understand the dataset well enough to analyze that variable.

Check the dataset version and release history

Datasets can change after release. Providers may correct coding errors, update weights, add cases, revise derived variables, or publish new documentation.

Record the version, release date, DOI or persistent identifier when available, and the date you obtained the data. Review release notes or errata before beginning substantive analysis.

This is important for reproducibility. If another researcher later downloads a revised version, apparently identical code may produce different results because the underlying file has changed.

Find out what you are actually allowed to do with the data

A downloadable file is not automatically unrestricted research material.

Datasets may be governed by licenses, data-use agreements, repository terms, consent conditions, confidentiality protections, ethical requirements, or restrictions on linkage, redistribution, publication, commercial use, geographical detail, or attempts to identify participants.

Restricted datasets may require an application, institutional affiliation, ethics documentation, secure computing environment, training, or disclosure review before outputs can leave the environment.

Read the terms applicable to the specific dataset and release. If access requires another organization's authorization, the study also raises the separate issue of what to do when data access depends on someone else's permission.

Check whether access will arrive in time

Even a methodologically ideal dataset can be a poor choice for a time-limited project if obtaining it takes longer than the project allows.

Public-use data may be available immediately. Restricted data can involve applications, agreements, ethics review, institutional signatures, security assessments, fees, training, or waiting periods. Some secure environments also constrain which software can be used or how outputs are released.

Investigate these requirements before designing the entire project around the dataset. A dataset you are theoretically eligible to use but cannot obtain until after your submission deadline is not practically available for the current study.

Estimate the work required after you receive the dataset

Existing data are sometimes described as though analysis begins immediately after download. Often it does not.

You may need to understand dozens of files, merge waves, reshape records, recode variables, reproduce derived measures, apply weights, clean dates, resolve duplicates, harmonize classifications, construct analytical variables, document exclusions, and learn specialized software.

A very large or complex dataset may also require computing resources beyond an ordinary laptop.

Secondary analysis can save substantial data-collection time, but those savings should not be confused with zero preparation time. The data already exist; your analytical dataset may not.

Consider whether the dataset has already been heavily studied

Widely used datasets are attractive partly because their quality and documentation may be strong. They also attract many researchers.

Before committing, search the literature for studies using the same dataset, population, variables, and relationships. This can reveal established analytical conventions, known limitations, previous operationalizations, and whether your proposed question has already been answered.

The objective is not to avoid a dataset because others have used it. Reanalysis, replication, updated periods, new theoretical questions, subgroup analyses, and different methods can all be worthwhile. You should simply know what contribution remains to be made.

Do not let the dataset manufacture the research question

Existing-data research often involves iteration between the question and available information. Cheng and Phillips describe both question-driven and data-driven approaches, noting that researchers commonly move between them as they identify what available datasets can support.

That flexibility is useful. Unstructured searching for statistically interesting relationships is a different matter.

If you begin with thousands of variables and repeatedly test combinations until something becomes statistically significant, you increase the risk of generating findings that reflect analytical searching rather than a well-founded research question. Exploratory analysis can be legitimate, but it should be identified and interpreted as exploratory rather than retrospectively presented as though every hypothesis had been specified in advance.

Watch Out

Do not choose an existing dataset merely because it is large, free, familiar, or easy to download. The most convenient dataset may still have the wrong population, weak measures, unsuitable timing, problematic missingness, inadequate documentation, or a design that cannot support the inference your question requires.

04 · A Practical Example

A Dataset Has Every Variable You Want, but Is It the Right Dataset?

Hypothetical Example

Studying generative AI use and academic performance

A graduate student discovers a large institutional dataset containing information on students' use of generative-AI tools, self-reported study behavior, demographic characteristics, and semester grades. The file contains more than 20,000 records and is already cleaned.

The student proposes to investigate whether generative-AI use improves academic performance and initially concludes that the dataset is ideal.

Check how the data were generated The documentation shows that AI use came from an optional student survey administered near the end of the semester rather than from direct usage records.
Check who responded Only a subset of enrolled students completed the survey, and response differed across academic programs.
Check measurement The AI variable records frequency categories but does not distinguish what students used AI for, how extensively they used it, or whether use occurred before the academic outcomes were determined.
Check the design AI use and other student characteristics were observational rather than experimentally assigned. The dataset therefore does not, by itself, establish the causal effect implied by the word "improves."
Check the analytical sample After restricting the analysis to students with complete survey and grade information in the relevant programs, the usable sample remains substantial but is much smaller than the headline figure of 20,000.
Refine the question The researcher reformulates the project around the association between reported generative-AI use and academic performance within the population and period supported by the dataset, while addressing relevant confounding and selection limitations as far as the design permits.

The dataset did not become worse during this evaluation. The researcher simply learned what it could and could not support. That is precisely what should happen before the dataset becomes the foundation of the study.

05 · What Researchers Often Get Wrong

Common Mistakes When Choosing an Existing Dataset

Misconception

A well-known dataset must be suitable for my study

Reputation can indicate that a dataset is widely used or well documented, but suitability is question-specific. An excellent national survey may still lack the measure, subgroup, timing, geographical detail, or design required by your particular research question.

Misconception

A very large sample can compensate for weak measurement

More observations can improve precision under appropriate conditions, but they do not transform a poor measure into a valid one. If your variable does not adequately represent the construct in your question, increasing the sample size does not repair that conceptual mismatch.

Misconception

If the dataset is already cleaned, I do not need to inspect data quality

"Cleaned" usually means that particular processing decisions have already been applied. You still need to understand those decisions, examine values relevant to your analysis, check missingness and consistency, and determine whether the provider's cleaning rules are appropriate for your research purpose.

Misconception

If my predictor and outcome are both present, I can answer the question

Variable availability is necessary but not sufficient. Timing, measurement, sampling, confounding, clustering, repeated observations, selection, and the overall study design determine which relationships can be estimated and how they should be interpreted.

Misconception

Publicly available data have no access or ethics issues

Public availability does not mean unrestricted use in every sense. Data may still have licenses, citation requirements, restrictions on redistribution or linkage, prohibitions on re-identification, or institutional requirements for research use. Check the conditions attached to the specific release.

Misconception

Using existing data automatically makes a study faster and easier

Secondary analysis can avoid primary data collection and may save considerable time and cost. Complex datasets can nevertheless require substantial work to understand, clean, merge, harmonize, weight, document, and analyze. Feasibility depends on the dataset and your intended analysis, not merely on the fact that the observations already exist.

06 · What This Means for You

Audit the Dataset Before It Becomes the Foundation of Your Study

Treat a promising existing dataset as a candidate rather than a commitment. Before finalizing the study, conduct a structured audit of the data source, design, population, variables, documentation, quality, analytical sample, access conditions, and practical requirements.

Then ask the decisive question: If I had known all of these limitations before finding the dataset, would I still consider it appropriate evidence for this research question?

A simple decision framework

If the dataset closely matches the required population, measures, timeframe, design, and analytical needs
Proceed to develop the analysis plan while documenting the dataset's remaining limitations and applicable access conditions.
If the dataset has minor limitations that do not fundamentally change the question
Use it if the limitations can be handled transparently and the resulting evidence still addresses the intended question credibly.
If an imperfect measure can serve as a defensible proxy
Define exactly what the proxy measures, justify its use, and narrow the interpretation so that the claim does not exceed the measure.
If the population, measurement, timing, or design materially conflicts with the original question
Find another dataset or revise the research question rather than forcing the existing data to answer something they were not capable of answering.
If access is uncertain, highly restricted, or unlikely to arrive within your project period
Do not treat the dataset as guaranteed. Resolve the access pathway or maintain an alternative before making the entire project depend on it.
If you cannot adequately understand how important variables or observations were produced
Seek better documentation or another data source. A dataset you cannot interpret is a weak foundation for claims you will ultimately have to defend.

The aim is not to find a flawless dataset. Few datasets survive that standard. The aim is to find one whose imperfections you understand and whose strengths are sufficient for the question you are asking.

07 · A Quick Checklist

Before Building Your Study Around an Existing Dataset

Before committing to the dataset, check:
Confirm that the dataset actually contains the essential variables, population, timeframe, observations, and level of detail required by your research question.
Read the codebook, data dictionary, questionnaires, technical documentation, sampling information, processing notes, and known limitations that are available.
Understand why and how the original data were collected or generated and how cases entered the dataset.
Verify how every central construct was measured, coded, derived, and timed rather than relying only on variable names.
Inspect missingness, unusual values, consistency, duplicates, changes across periods, and other data-quality issues relevant to your analysis.
Determine whether the sampling and study design support the population and type of inference implied by your research question.
Calculate or estimate the usable analytical sample after applying your actual population, timeframe, variable, linkage, and completeness requirements.
Record the dataset version and check release notes, corrections, or errata before analysis.
Review licenses, data-use agreements, confidentiality requirements, ethics considerations, linkage restrictions, and other conditions governing the specific data release.
Estimate how much time, expertise, software, computing capacity, cleaning, merging, harmonization, and preparation the dataset will require before substantive analysis can begin.
08 · Frequently Asked Questions

Frequently Asked Questions About Using Existing Datasets

What is secondary analysis of existing data?

It generally refers to using data that have already been collected or generated to address a research question beyond the original analysis or purpose. Existing data may come from previous research studies, surveys, administrative systems, registries, electronic records, organizations, platforms, or other sources.

Should I choose the research question or the dataset first?

Either route can occur, and in practice the process is often iterative. A question-driven study searches for data capable of answering a prior question, while a data-driven approach may identify worthwhile questions after examining available information. What matters is that the final question is substantively meaningful and that the dataset can address it without post hoc analytical searching being misrepresented as prior hypothesis testing.

How do I know whether an existing dataset is good quality?

There is no single quality score suitable for every dataset. Examine how the data were generated, sampling and coverage, measurement procedures, missingness, consistency, validation or quality-control processes, documentation, processing history, and known limitations. Quality should ultimately be judged in relation to the particular research question and intended analysis.

Can I use a proxy when the dataset does not contain my ideal variable?

Sometimes. A proxy should have a defensible conceptual relationship to the construct and be appropriate for the particular question. Explain what it actually measures and constrain your interpretation accordingly. A convenient variable should not simply be renamed as the construct you originally wanted.

Can I use an existing cross-sectional dataset to study causation?

The presence of an exposure and outcome in cross-sectional observational data does not by itself establish causal direction or temporal ordering. Causal inference requires assumptions and designs appropriate to the causal question. If the available design cannot support the intended claim, revise the question or interpretation rather than allowing causal language to outrun the evidence.

Do I need to worry about missing data if the dataset has already been cleaned?

Yes. Cleaning does not imply complete data or make missingness irrelevant to every new analysis. Inspect missingness in the variables and population you actually plan to use, understand the provider's missing-value codes and processing decisions, and select an analytical approach appropriate to the missing-data problem.

Do publicly available datasets require ethics approval?

The answer depends on the data, jurisdiction, institution, identifiability of the information, terms of access, and proposed use. Public availability should not be treated as a universal exemption from ethics or institutional requirements. Check the dataset's conditions and your institution's applicable research-ethics procedures.

What if the dataset is ideal but I have not yet been granted access?

Treat access as a project dependency rather than an accomplished fact. Investigate eligibility, application requirements, approval timelines, costs, security requirements, and alternatives. If the entire project depends on that dataset, consider carefully whether you should base the research project on data you have not yet been guaranteed access to.

09 · The Bottom Line

The Dataset Must Fit the Question, Not Merely Contain Useful Variables

The Bottom Line

Before building a study around an existing dataset, make sure you understand how the data were produced, whom they represent, what the variables actually measure, how complete and reliable the relevant observations are, whether the design supports your intended analysis, and what restrictions govern their use.

Existing data can make otherwise expensive or time-consuming research possible, but they also make you inherit methodological decisions you did not choose. A strong secondary analysis is built not on finding a perfect dataset, but on understanding an imperfect one well enough to know exactly which questions it can credibly answer.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes