03 · What You Need to Know
How to Evaluate an Existing Dataset Before Committing to It
First confirm that the data you need actually exist
Before evaluating a dataset in depth, establish that it contains the essential information required by your research question.
That means checking the actual variables and measures, relevant population, timeframe, unit of analysis, level of detail, and any linkage required across files or waves. A dataset can be excellent in every other respect and still be unsuitable because it does not contain the evidence your particular question requires.
If you have not completed that check, first determine whether the required research data actually exist. Once they do, the next question is whether those data are good enough and appropriately structured for the study you intend to conduct.
Understand why the data were originally collected
Existing data were created for a purpose, and that purpose influences what was recorded and how.
A national survey may have been designed to estimate population characteristics. Clinical records primarily support patient care. University records support academic administration. Learning-management-system logs are generated by platform activity. Government administrative databases may exist for regulatory or service-delivery purposes. Data from an earlier research project were collected to answer that project's questions, not necessarily yours.
This matters because data collected for another purpose may emphasize different constructs, populations, time intervals, or levels of precision from those your study requires.
Secondary-analysis guidance therefore recommends understanding the original study or data-generation process rather than treating an existing file as a neutral collection of variables. Cheng and Phillips, for example, emphasize examining the study population, sampling strategy, assessment methods, response levels, quality-control procedures, and other features of the original data before conducting a new analysis.
Reconstruct how the dataset came into existence
A useful evaluation begins with provenance: where did these observations come from?
For a research dataset, examine the original study design, sampling strategy, recruitment, instruments, data-collection procedures, study period, follow-up, and processing. For administrative or operational data, determine how records enter the system, who records them, why they are recorded, what changes occurred over time, and what quality controls apply.
Ask whether the analytical file is raw, cleaned, transformed, imputed, weighted, aggregated, de-identified, or otherwise processed. If variables were derived, find out how.
You should be able to trace important variables back to the process that produced them. If you cannot determine what an observation represents or how it was created, interpreting it confidently becomes difficult.
Check whether the population matches your question
A dataset may contain the right variables but represent the wrong people, institutions, events, locations, or period.
Suppose you want to investigate digital-learning behavior among university students generally, but the dataset contains only students enrolled in online degree programs. Or perhaps you want to study employees nationally but the administrative dataset contains only workers who applied for a particular government benefit.
The question is not simply whether your desired subgroup appears somewhere in the file. You need to understand how cases entered the dataset and what population the data can reasonably represent.
For survey data, this includes examining the target population, sampling frame, sampling design, response, and weights where applicable. For administrative records, it includes understanding the service, institution, or process that generated inclusion in the database.
Do not let a huge sample distract you from poor population fit
Large datasets are seductive. A file containing 500,000 observations can look inherently more impressive than a study of 500 carefully selected participants.
Sample size does not solve a mismatch between the dataset and the research question.
If the database systematically excludes people relevant to your target population, increasing the number of observations from the included population does not make those excluded people appear. Likewise, a massive convenience or administrative dataset does not become population-representative simply because it is massive.
Methodological guidance on secondary dataset analysis makes a similar point: a large sample is valuable only when the dataset fits a meaningful question, appropriate population, measures, and analytical strategy.
Examine how the important constructs were measured
The variables in an existing dataset were defined before your current research question existed. You therefore inherit their operationalizations.
Suppose your question concerns student engagement. The dataset contains number of logins. Does login frequency adequately represent the type of engagement your theoretical argument concerns? Perhaps. Perhaps not.
Or suppose you want to study psychological distress but have only one general well-being item. You cannot turn that item into a validated distress scale by giving the variable a more ambitious label.
For each central construct, examine the original instrument, question wording, response categories, scoring, timing, units, and evidence supporting the measure when relevant. If you use a proxy, explain why the proxy is reasonable and what it does not capture.
Smith and colleagues note that available measures in secondary datasets rarely correspond perfectly to every construct researchers might ideally want. The question is whether the available measurement is sufficiently appropriate for the specific inference you intend to make.
Check whether definitions changed over time
Long-running datasets deserve special scrutiny because apparent continuity can conceal measurement changes.
A survey question may be reworded. Diagnostic criteria can change. Administrative categories may be reclassified. A university may alter grading policies. A platform update can change what counts as an activity event. A variable may move from optional to mandatory entry.
If you intend to examine trends or combine multiple waves, determine whether important variables are comparable across the entire period. A column with the same name in several years does not guarantee that it represents exactly the same construct in each year.
When harmonization is required, document what was combined, recoded, or excluded and what comparability assumptions remain.
Inspect missing data before deciding the dataset is usable
Missingness should be investigated variable by variable and, when relevant, across subgroups, sites, waves, or periods.
A dataset may contain your outcome variable for 98% of observations but an essential predictor for only 60%. Or the overall missingness may look modest while nearly all missing observations come from one region or participant group.
Start descriptively. Determine how much information is missing, where it is missing, and whether important variables are missing together. Then consider plausible reasons for those patterns and what they mean for your analysis.
Do not assume that deleting incomplete cases is automatically appropriate. Neither should you assume that a sophisticated missing-data technique can repair every problem. The defensibility of an approach depends on the data-generating and missingness processes, analytical method, and assumptions involved.
Look for impossible, inconsistent, or suspicious values
Data quality assessment should go beyond missing cells.
Check ranges, categories, units, duplicates, dates, identifiers, internal consistency, and relationships among variables. An age of 240 years, a graduation date before enrollment, duplicate identifiers assigned to different people, or a score outside an instrument's possible range warrants investigation.
Not every surprising observation is an error. Unusual cases can be real and scientifically important. The purpose of checking is to distinguish plausible extremes from coding, entry, linkage, or processing problems rather than automatically deleting anything inconvenient.
When a provider supplies data-quality notes, cleaning documentation, validation reports, or known-issues files, read them before inventing your own explanation for anomalies.
Understand derived variables and recoding
Some of the most convenient variables in an existing dataset may not have been collected directly.
Variables such as socioeconomic status, disease classifications, composite scores, risk groups, geographic categories, engagement indices, or household measures may have been derived from several source fields. Their definitions may contain assumptions that matter for your analysis.
Inspect how important derived variables were constructed. If the documentation provides the underlying components, determine whether using the provider's derived variable or constructing your own version better matches the research question and permitted use of the data.
Do not assume that a professionally prepared derived variable is conceptually appropriate for every research purpose simply because it is convenient.
Check whether the design supports the relationship you want to study
A dataset may contain both your predictor and outcome without supporting the inference implied by your question.
For example, a cross-sectional survey may contain measures of generative-AI use and academic confidence collected at the same time. It can support analyses of association under appropriate assumptions. It does not by itself establish that AI use changed later confidence or that one variable caused the other.
Likewise, repeated cross-sectional surveys conducted every year are not the same as longitudinal data following the same individuals over time.
Variables coexist in the dataset
The dataset contains measures of both concepts required for the proposed analysis.
The design supports the intended inference
The timing, sampling, measurement, and structure of the data are appropriate for the relationship or change the research question asks you to estimate.
The second requirement is more demanding. Existing data do not become longitudinal, experimental, representative, or causal merely because the desired variables happen to appear together.
Understand the sampling design before analyzing survey data
Large surveys may use stratification, clustering, unequal selection probabilities, oversampling, multistage sampling, and survey weights. Those design features can affect estimation and uncertainty.
If a dataset provides weights, strata, clusters, replicate weights, or other design variables, determine why they exist and whether your planned analysis needs to account for them. Ignoring a complex sampling design can produce estimates or standard errors that do not correspond to the way the sample was actually obtained.
The relevant documentation should explain the sampling and weighting procedures. If you cannot understand how to incorporate them, that may be a methodological skill gap to address before proceeding rather than a reason to quietly analyze the file as though it came from a simple random sample.
Check whether you have enough usable observations for the planned analysis
The total number of rows in the dataset is rarely the number available to every analysis.
Your usable sample may shrink after restricting to the appropriate population, selecting the relevant years, requiring valid observations on essential variables, linking records, or applying design criteria.
More complex analyses can create additional requirements. A multilevel analysis needs sufficient information at relevant levels. Longitudinal models require repeated observations. Rare outcomes may leave few events despite a large overall sample.
Doolan and Froelicher specifically identify available statistical power and data quality as considerations when determining whether an existing dataset is adequate for a proposed secondary analysis.
Evaluate the analytical sample you will actually have, not the sample size printed on the dataset's homepage.
Determine whether documentation is sufficient to interpret the data
A dataset without adequate documentation can be difficult to use responsibly even when the file itself opens perfectly.
Look for a codebook or data dictionary, questionnaires or instruments, study-design documentation, sampling information, fieldwork procedures, variable descriptions, coding information, derivation rules, weighting guidance, missing-value codes, version history, and known limitations where applicable.
Good documentation allows another researcher to understand what the variables mean and how the observations were generated. Metadata and data dictionaries are particularly important when datasets are shared or combined because they preserve the context required for reuse.
If you cannot tell whether 9 means "missing," "not applicable," or an actual response category, you do not yet understand the dataset well enough to analyze that variable.
Check the dataset version and release history
Datasets can change after release. Providers may correct coding errors, update weights, add cases, revise derived variables, or publish new documentation.
Record the version, release date, DOI or persistent identifier when available, and the date you obtained the data. Review release notes or errata before beginning substantive analysis.
This is important for reproducibility. If another researcher later downloads a revised version, apparently identical code may produce different results because the underlying file has changed.
Find out what you are actually allowed to do with the data
A downloadable file is not automatically unrestricted research material.
Datasets may be governed by licenses, data-use agreements, repository terms, consent conditions, confidentiality protections, ethical requirements, or restrictions on linkage, redistribution, publication, commercial use, geographical detail, or attempts to identify participants.
Restricted datasets may require an application, institutional affiliation, ethics documentation, secure computing environment, training, or disclosure review before outputs can leave the environment.
Read the terms applicable to the specific dataset and release. If access requires another organization's authorization, the study also raises the separate issue of what to do when data access depends on someone else's permission.
Check whether access will arrive in time
Even a methodologically ideal dataset can be a poor choice for a time-limited project if obtaining it takes longer than the project allows.
Public-use data may be available immediately. Restricted data can involve applications, agreements, ethics review, institutional signatures, security assessments, fees, training, or waiting periods. Some secure environments also constrain which software can be used or how outputs are released.
Investigate these requirements before designing the entire project around the dataset. A dataset you are theoretically eligible to use but cannot obtain until after your submission deadline is not practically available for the current study.
Estimate the work required after you receive the dataset
Existing data are sometimes described as though analysis begins immediately after download. Often it does not.
You may need to understand dozens of files, merge waves, reshape records, recode variables, reproduce derived measures, apply weights, clean dates, resolve duplicates, harmonize classifications, construct analytical variables, document exclusions, and learn specialized software.
A very large or complex dataset may also require computing resources beyond an ordinary laptop.
Secondary analysis can save substantial data-collection time, but those savings should not be confused with zero preparation time. The data already exist; your analytical dataset may not.
Consider whether the dataset has already been heavily studied
Widely used datasets are attractive partly because their quality and documentation may be strong. They also attract many researchers.
Before committing, search the literature for studies using the same dataset, population, variables, and relationships. This can reveal established analytical conventions, known limitations, previous operationalizations, and whether your proposed question has already been answered.
The objective is not to avoid a dataset because others have used it. Reanalysis, replication, updated periods, new theoretical questions, subgroup analyses, and different methods can all be worthwhile. You should simply know what contribution remains to be made.
Do not let the dataset manufacture the research question
Existing-data research often involves iteration between the question and available information. Cheng and Phillips describe both question-driven and data-driven approaches, noting that researchers commonly move between them as they identify what available datasets can support.
That flexibility is useful. Unstructured searching for statistically interesting relationships is a different matter.
If you begin with thousands of variables and repeatedly test combinations until something becomes statistically significant, you increase the risk of generating findings that reflect analytical searching rather than a well-founded research question. Exploratory analysis can be legitimate, but it should be identified and interpreted as exploratory rather than retrospectively presented as though every hypothesis had been specified in advance.
Watch Out
Do not choose an existing dataset merely because it is large, free, familiar, or easy to download. The most convenient dataset may still have the wrong population, weak measures, unsuitable timing, problematic missingness, inadequate documentation, or a design that cannot support the inference your question requires.