03 · What You Need to Know
Available Data Set Boundaries, but They Should Not Dictate the Research Uncritically
Start by specifying what evidence the original question requires
Before deciding that a dataset is "good enough," identify the minimum evidence necessary to answer the question.
What population needs to be represented? Which exposure, outcome, construct, or phenomenon must be measured? What covariates or contextual information are necessary for the intended analysis? What timeframe is required? Does the question depend on longitudinal observations, comparisons, particular subgroups, or sufficient numbers of cases?
Guidance on selecting data sources recommends specifying these minimum data requirements before committing to a dataset. If a critical element is absent, the appropriate response may be another data source, linkage with additional data, primary data collection, or revision of the research question.
This prevents a common mistake: becoming attached to an accessible dataset first and gradually redefining the research problem until almost any available variable appears adequate.
Inspect the dataset before deciding what it can answer
A dataset is more than a spreadsheet containing familiar column names.
Before relying on existing data, examine how the data were generated, which population was sampled, the sampling strategy, when collection occurred, what instruments or administrative processes produced the variables, response levels, quality-control procedures, coding rules, and the extent and pattern of missing information. Researchers conducting secondary analysis are advised to study the available documentation, including instruments, codebooks, manuals, and related methodological materials.
The aim is to understand what each variable actually represents rather than what its label appears to represent.
A field called "AI use," for example, might record whether someone has ever used an AI tool. That is not automatically an adequate measure of frequency, purpose, intensity, quality, or type of generative AI use. If your question concerns how frequently students use generative AI for academic writing, a yes-or-no item about any previous AI use may not support the construct you intended to investigate.
Identify exactly where the mismatch occurs
"The data are limited" is too vague to guide a methodological decision. Determine which dimension of the scope is affected.
| Data problem |
What it may affect |
Possible consequence |
|
Required population absent or poorly represented
|
Population scope |
The intended population claim may need narrowing |
|
Required variable was never collected
|
Substantive or analytical scope |
The question may be unanswerable with this dataset |
|
Available measure is only a weak proxy
|
Construct validity |
The question or interpretation may need reformulation |
|
Required years or waves are unavailable
|
Temporal scope |
Trend or longitudinal questions may no longer be supportable |
|
Relevant subgroup is too small
|
Comparison or subgroup scope |
Estimates may lack adequate precision or the comparison may need to be removed |
|
Important contextual information is absent
|
Interpretation |
Context-dependent conclusions may be difficult to support |
|
Substantial missing data affect key variables
|
Analysis and effective evidence |
Bias, precision, and the appropriate analytical population require assessment |
Different mismatches require different solutions. Missing an optional descriptive variable is not equivalent to missing the primary outcome. Having fewer years than expected is not equivalent to having no measure of the construct the study exists to investigate.
A missing essential variable may mean the dataset is unsuitable
Researchers sometimes assume that once a dataset has been obtained, the project must use it. That is not necessarily so.
If a question requires a variable that the dataset does not contain, and no defensible measure or linkage can supply it, the dataset may simply be incapable of answering that question. Guidance on observational data-source selection explicitly notes that when critical information is unavailable, primary data collection or linkage with other datasets may be needed.
Suppose you want to investigate whether generative AI use improves students' writing performance, but the dataset contains only students' attitudes toward AI and no measure of writing performance. Attitude toward AI is not a substitute for writing performance simply because both concern the same topic.
You have at least three defensible options: find data containing the required outcome, collect it yourself if feasible and appropriate, or ask a different question that the available measures can genuinely answer.
Do not rename a proxy until it becomes the construct you wanted
Secondary data often provide imperfect measures of the constructs researchers would ideally study. Proxy measures can sometimes be defensible, but their adequacy must be evaluated rather than assumed.
If a dataset records whether students logged into an online platform, for example, that variable may measure platform access or activity under a particular operational definition. It does not automatically measure engagement, learning, motivation, or meaningful participation.
A proxy is most defensible when there is a substantive and methodological rationale linking it to the construct, its limitations are understood, and the interpretation uses language appropriate to what was actually measured.
Watch Out
Do not solve a data mismatch linguistically. If the dataset measures platform logins, calling the variable "student engagement" in the manuscript does not make it a validated measure of engagement. Narrow or reformulate the construct when necessary and describe the operational measure accurately.
Population coverage can force the question to become narrower
Suppose you intended to study university students nationally but discover that the available dataset includes only students from public universities in selected regions.
You cannot recover the missing population merely through broader wording in the title or discussion.
The study may still be valuable. You could reformulate the question around the population actually represented, provided that this population remains substantively meaningful and the dataset's sampling design supports the intended inference.
The important step is to carry the revised boundary throughout the study. The research question, objectives, methods, interpretation, and conclusions should not continue to imply that private-university students, unrepresented regions, or other absent populations were studied.
This is a case in which population and place become explicit boundaries of the revised scope.
Missing years can change a longitudinal question into a different question
Time coverage is another common constraint in existing datasets.
Suppose your original plan is to examine a ten-year trend, but only four comparable annual waves are available. Simply conducting the analysis on those four years does not mean you have answered the original ten-year question.
Perhaps the four-year period remains theoretically or practically meaningful. If so, narrow the temporal scope and explain why that period can still address a worthwhile question.
If the missing years remove an essential pre-policy baseline, intervention period, or longitudinal sequence, the original analysis may no longer be possible. In that case, finding another data source may be preferable to forcing the available period into a design it cannot support.
Too few cases can remove a subgroup or comparison from the feasible scope
A dataset can be enormous overall and still contain too little evidence for the particular question you want to ask.
Suppose a national dataset contains 100,000 respondents, but only 60 belong to the specific subgroup central to your proposed analysis. The impressive total sample size does not solve the evidential problem within that subgroup.
Secondary-analysis guidance emphasizes evaluating whether the available dataset has adequate sample size, power, and data quality for the proposed question. Recent methodological discussions similarly caution that secondary analyses of small subgroups may be insufficiently powered even when the parent study was large.
Your options may include combining substantively defensible categories, selecting a broader population, treating the analysis as exploratory, obtaining another dataset, or abandoning the subgroup question. The appropriate response depends on the design and purpose.
Do not combine groups solely to increase numbers if doing so destroys a distinction that matters to the research question.
Missing data are not the same as an absent variable
It is useful to distinguish two problems.
An absent variable was not collected at all. The dataset cannot directly provide that information.
Missing values occur when the variable exists but observations are unavailable for some cases.
The methodological responses differ. Missing-data strategies may include complete-case analysis, imputation, weighting, model-based approaches, sensitivity analysis, or other techniques appropriate to the design and missingness mechanism. STROBE reporting guidance requires observational studies to explain how missing data were addressed.
But no missing-data technique can reconstruct an entire construct that was never measured without additional information and assumptions. Do not treat "not collected" as an ordinary missing-value problem.
Ask whether another data source can preserve the original question
Narrowing should not be automatic merely because the first dataset is inadequate.
Before changing the question, consider whether a better dataset exists. Could another source provide the missing population, variable, timeframe, or outcome? Could datasets be linked legitimately and methodologically? Would primary data collection be feasible? Could the question be addressed through a different research design?
Data-source guidance emphasizes matching data to the question and considering primary data collection when existing data do not contain the required information.
The decision is partly practical. Obtaining another dataset may involve cost, access restrictions, linkage difficulties, ethical requirements, or substantial delays. But the existence of those obstacles does not make an inadequate dataset adequate.
Sometimes revising the question is exactly the right methodological response
Secondary-data analysis reverses part of the usual research-design sequence because the evidence already exists. Researchers therefore often move iteratively between the research question and the characteristics of available datasets.
Methodological reviews acknowledge that when existing datasets do not contain all required variables, researchers may need to modify the research question or analytical plan based on the best available data.
That is not inherently poor practice.
The critical distinction is between adapting the question to what valid evidence can support and searching through available variables until an interesting association appears.
A revised question should still emerge from a meaningful research problem, existing literature, and defensible reasoning. High-value secondary-analysis guidance recommends starting with a substantive topic or question and remaining flexible to the strengths and limitations of potential datasets, while warning against unfocused data dredging.
Do not let the dataset become the research problem
Large datasets are seductive. Hundreds of variables invite hundreds of possible associations.
But "these two variables happen to be available" is not a sufficient research rationale.
After adapting the question, return to the literature. Ask whether the revised question addresses genuine uncertainty, whether the available measures provide a defensible operationalization, and whether either a positive or negative finding would contribute something worth knowing.
If the revised question exists only because two columns can be correlated, you may have solved the feasibility problem by creating a significance problem.
This connects directly to the risk that narrowing the study can eventually make the research question too trivial.
Do not hide the change when the scope was originally broader
If the study was already planned, registered, approved, or otherwise documented before the data problem became apparent, preserve the distinction between the original and revised scope.
For example, you might state that the original analysis intended to include a particular variable, but examination of the data documentation showed that the measure was unavailable or unsuitable, leading to a revised research question.
The exact reporting requirements depend on the methodology, institutional requirements, preregistration, protocol, and publication venue. The general principle is transparency: do not rewrite an unplanned data-driven restriction as though it had always been an elegant delimitation.
If the study has already commenced formally, follow the broader principles for documenting and evaluating changes to research scope after the project has begun.
Reassess the analysis after narrowing the scope
Changing the scope can alter the analytical problem.
If you remove a population, the sample composition changes. If you shorten the timeframe, the number of observations and temporal structure may change. If you remove an outcome, the objectives and statistical testing plan may change. If you substitute a different operational measure, the interpretation of coefficients or themes changes.
Do not revise the research question while leaving the original analysis plan untouched.
For observational studies, transparent reporting includes defining outcomes, exposures, predictors, potential confounders and effect modifiers, explaining study size, describing missing-data handling, and reporting subgroup and sensitivity analyses where applicable.
The revised analysis should correspond to the revised question.
Reassess whether the narrowed study is still worth doing
Data availability can make a study feasible while simultaneously making it less meaningful.
Suppose your original question examines whether generative AI use predicts writing quality. The dataset lacks writing-quality measures but includes whether students say they "like technology." You could construct a new study examining AI use and liking technology, but the mere availability of both variables does not establish that the new question matters.
Secondary-data guidance emphasizes that a good research question remains central even when a very large dataset is available. Large samples do not rescue unimportant questions.
After every major narrowing decision, ask:
If this is now the question, is the answer still worth obtaining?
Sometimes the correct decision is not to proceed
Researchers understandably dislike abandoning a study after investing time in finding data, obtaining access, cleaning files, or developing an analysis plan.
But sunk effort does not make a dataset capable of answering a question.
If the essential population is absent, the primary construct cannot be measured defensibly, the required comparison is impossible, data quality is inadequate, or the narrowed question no longer has meaningful value, stopping or changing datasets may be the most methodologically responsible choice.
A smaller answer is often better than an overstated answer. Sometimes, however, no answer from the available data is better than a confident answer to a question the data were never capable of addressing.