03 · What You Need to Know
How Data Availability Should Influence Research Topic Selection
Data Access Is a Real Feasibility Constraint
Research questions do not exist independently of the evidence required to answer them.
If your question concerns employee turnover but you cannot obtain employment records, recruit relevant participants, or generate another appropriate source of evidence, the proposed project may not be executable. If your study requires a rare clinical population but recruitment cannot realistically reach an adequate sample, the question may need to change.
The FINER framework makes this explicit. Feasibility includes practical resources such as time, expertise, funding, institutional support, and the availability of data or participants. A question can be intellectually well constructed while remaining impossible within a particular research environment.
Checking data access early is therefore not compromising research quality. It is part of designing research that can actually be completed.
But “I Have the Data” Is Not a Research Rationale
Imagine receiving a dataset containing age, occupation, commute length, household size, income, exercise frequency, sleep duration, job satisfaction, and dozens of other variables.
You could test hundreds of relationships.
That does not mean hundreds of worthwhile research questions are waiting inside the spreadsheet.
A research question should still arise from a meaningful problem, uncertainty, theoretical issue, evidence need, or other defensible research purpose. Guidance on research-question formulation evaluates questions not only for feasibility but also for interest, novelty, ethics, and relevance.
Data availability solves one part of the research problem: obtaining evidence. It does not establish why analyzing that evidence would matter.
Question-First and Data-First Research Are Both Possible
Research does not always develop in exactly the same direction.
In a question-first project, you identify an important research problem, formulate a question, and determine what evidence would be required to answer it. You then collect new data or find existing data that fit those requirements.
In a data-opportunity project, you encounter a valuable dataset, archive, cohort, administrative system, collection, or other source and ask what meaningful questions it can legitimately help answer.
Question-first
What evidence do I need to answer this worthwhile question?
Data-opportunity
What worthwhile questions can this evidence credibly help answer?
The requirement in both
The final question and evidence must fit each other.
The danger is not that data ever inspire research. Data can absolutely reveal patterns, possibilities, and questions. The danger is confusing analytical possibility with research importance.
Know What the Data Were Originally Created to Do
Secondary data were usually generated for some purpose, and that purpose affects what they contain.
Administrative records may have been designed to process transactions rather than measure research constructs. Customer databases may support operations. Government statistics may follow reporting definitions created for policy or administration. Platform data may reflect product architecture. Historical archives preserve what institutions or individuals chose to record.
Monash University's guidance on external data recommends understanding why a dataset was produced and for whom, while also clarifying the research question and unit of analysis before selecting a source.
Ask why each variable exists, how it was generated, whose information is absent, what definitions were used, and what changes in collection procedures may have occurred.
A Variable Name Is Not the Same as the Concept You Want to Study
Suppose your research question concerns socioeconomic disadvantage and your dataset contains “annual household income.” Is income an adequate measure of the concept your question requires?
Perhaps. Perhaps not.
Or imagine a dataset contains the number of times students logged into an online learning system. That variable may measure login frequency. It does not automatically measure “student engagement,” even if engagement is the concept you want to investigate.
This distinction is crucial in secondary-data research because you cannot redesign measurements after the fact. You inherit operational definitions, categories, missingness, collection procedures, and measurement limitations.
Do not rename an available variable to make it sound like the construct your question needs.
Check the Unit of Analysis Before You Fall in Love With the Dataset
What exactly does each row or observation represent?
A person? Household? School? Hospital admission? Transaction? Company? Country? Social-media post? Measurement occasion?
Monash specifically recommends clarifying the unit of analysis, location, and timeframe when determining what external data are needed to answer a research question.
A dataset can contain fascinating information while operating at the wrong level for your intended inference. Country-level statistics, for example, cannot automatically establish relationships that hold among individuals within those countries.
The available unit of analysis needs to match the question you are actually asking.
Check Whether the Population Matches the Population in Your Question
Data access often tempts researchers to redefine the target population around whoever happens to appear in the dataset.
Suppose you want to understand working adults generally but possess records from employees of one large technology company. That may support a worthwhile study of that workforce. It does not automatically support conclusions about all workers.
Research-question guidance emphasizes specifying the population to which the question is relevant, while feasibility includes whether appropriate participants or observations are available.
If the accessible population is narrower than the population you originally wanted to study, you have several choices: narrow the question, obtain additional evidence, redesign the project, or make appropriately limited claims.
Check Whether the Timeframe Can Answer the Question
A dataset may contain exactly the variables you want but cover the wrong period.
Cross-sectional data collected once cannot directly reveal within-person change over several years. Records beginning after a policy was introduced cannot establish the pre-policy baseline needed for some evaluations. A dataset ending in 2018 may be poorly suited to a question specifically concerning post-pandemic behavior.
Do not let the presence of variables distract you from the temporal structure of the question.
Ask when observations were collected, how often, whether the same units were followed over time, whether definitions changed, and whether the relevant event occurred inside the observation window.
Data Structure Limits the Claims You Can Make
Researchers can sometimes formulate causal-sounding questions around data that support only weaker inferences.
Suppose an observational dataset shows that employees who work remotely more often report greater job satisfaction. That association does not, by itself, establish that remote work caused the difference. Employees may self-select into remote arrangements, job types may differ, organizational policies may matter, and other factors may affect both variables.
The research question, design, and analysis need to match. Research-methods guidance emphasizes that the question should guide study design rather than allowing an available dataset to determine claims the design cannot support.
When the data cannot answer the causal question, change the claim or obtain evidence capable of addressing it.
Data Quality Matters as Much as Data Availability
Accessible data are not necessarily good data.
Before committing to a project, investigate missing values, coding practices, measurement error, duplicate records, changes in collection systems, coverage, representativeness, unusual exclusions, reliability, and documentation.
A dataset can be enormous and still be poorly suited to your question.
Monash's external-data guidance explicitly recommends evaluating data quality and credibility rather than treating availability as sufficient.
A smaller dataset collected specifically for the relevant construct may sometimes provide stronger evidence than millions of convenient observations measuring the wrong thing.
“Big Data” Does Not Automatically Mean a Strong Study
Large sample size can improve precision for some estimates, but it cannot repair every design problem.
With very large datasets, tiny associations can become statistically detectable even when they are substantively unimportant. Measurement bias does not disappear because the dataset contains millions of rows. Selection bias can remain. Confounding can remain. Missing variables can remain.
More data are useful when they are appropriate data.
When comparing potential topics, do not rank a project automatically above another because its dataset is larger. Ask what the data allow you to learn credibly.
Existing Data Can Make Otherwise Impossible Research Possible
The caution against data-led questions should not obscure the enormous advantages of secondary data.
Existing datasets can provide large samples, long time periods, rare outcomes, population coverage, historical information, or observations that would be prohibitively expensive for an individual researcher to collect.
Government data, repositories, linked datasets, licensed sources, and other researchers' archived data can therefore transform feasibility. Monash identifies all of these as potential external sources for secondary research.
For a student or early-career researcher with limited funding, a high-quality existing dataset can make a sophisticated question realistically answerable.
Data Access Can Be a Strategic Tie-Breaker
Suppose you are choosing between two questions that are both significant, interesting, ethically acceptable, and methodologically defensible.
Topic A requires uncertain access to an organization that has not yet approved your request. Topic B can be investigated using a well-documented dataset for which access is already secured.
Choosing Topic B can be entirely reasonable. Data availability is a legitimate component of feasibility, and feasibility increases the likelihood that a study can actually be completed.
This becomes especially useful when choosing between two otherwise strong research topics.
But “Accessible” Should Mean More Than “I Know Where the File Is”
Researchers often underestimate what data access entails.
A dataset may exist without being immediately usable. Access can depend on licenses, data-use agreements, ethics approval, organizational authorization, secure computing environments, fees, training, confidentiality restrictions, or approval from the data custodian.
Some datasets allow analysis but restrict publication of small cells or particular variables. Others may permit access only for a specified research purpose.
Before treating data as available, verify the actual conditions under which you can obtain, analyze, link, store, and report them.
Ethical and Legal Permission Are Part of Data Feasibility
The technical ability to obtain data does not automatically create ethical permission to use them.
Personal, sensitive, confidential, proprietary, or linked data may involve requirements concerning consent, privacy, governance, storage, access control, de-identification, data-use agreements, and institutional review.
The exact requirements vary by jurisdiction, institution, dataset, and research design. Check the applicable rules rather than assuming that existing data are exempt from research governance because somebody else originally collected them.
Feasibility frameworks explicitly include ethics alongside practical research considerations.
Do Not Build the Question After Looking for Significant Correlations
One particularly risky form of data-first research is searching through many variables, discovering an interesting association, and then presenting the resulting relationship as though it had been the original research question.
Exploratory analysis can be scientifically useful. The problem is disguising exploration as prespecified confirmatory testing.
If the dataset generates an unexpected hypothesis, say so where appropriate and test it independently when the research claim requires confirmation. Keep exploratory and confirmatory purposes conceptually distinct.
The issue is not that researchers are forbidden to learn from patterns in data. It is that the evidentiary strength of a result depends partly on how the question and analysis arose.
Available Data Can Reveal Better Questions Than the One You Started With
Question refinement does not need to flow in only one direction.
You may begin with a theoretical question and discover that the ideal data do not exist. While investigating alternatives, you find a dataset containing repeated observations that allow you to examine a related but potentially more informative question.
That can be excellent research development.
The important step is to return to the literature and ask whether the revised question is meaningful on its own merits. Do not treat the new question as justified merely because the data made it possible.
Sometimes You Should Change the Question to Fit the Evidence
Researchers occasionally treat changing a question because of data limitations as a methodological failure.
It is not necessarily one.
If your original question requires evidence you cannot obtain, a narrower or differently framed question may still make a useful contribution. Research-question guidance explicitly recognizes that feasibility depends on the resources and observations researchers can access.
The critical requirement is transparency: the revised question should still address something worth knowing, and your conclusions should match what the available evidence can establish.
Sometimes You Should Reject the Dataset Instead
Not every feasibility problem should be solved by changing the research question until it fits the data.
If the only way to use a dataset is to replace the concept you care about with a poor proxy, abandon the relevant population, ignore the necessary timeframe, or make a much weaker inference, the dataset may simply be unsuitable.
Researchers can become attached to data because obtaining them required effort. That sunk cost should not determine the research question.
Ask whether you would still consider the study worthwhile if somebody else handed you the same dataset today. If not, access may be driving the project more than the research problem.
Pilot Data Can Help You Decide Before Committing
Sometimes you cannot tell whether a project is feasible from documentation alone.
A pilot or feasibility assessment may reveal whether enough participants can be recruited, whether a measure works in the intended setting, whether records contain the necessary information, or whether missingness is manageable. Research guidance specifically recommends considering pilot or proof-of-concept work when feasibility is uncertain.
This can be particularly useful when data access is the main uncertainty separating a promising idea from an executable study.
Accessible Data Can Support a Smaller Study Within a Larger Research Agenda
You may have a large research question that the available data cannot answer completely.
Instead of forcing one dataset to do everything, ask whether it can answer one useful precursor question.
For example, existing administrative data might establish the scale and distribution of a phenomenon. A later qualitative study could investigate mechanisms. Another study might evaluate an intervention.
This is one way one research topic can become several linked studies, each using evidence appropriate to its particular question.
Think in Terms of Question–Data Fit
The most useful principle is not “question first” or “data first.” It is alignment.
| Question |
Data |
Decision |
| Important and answerable |
Appropriate and accessible |
Strong candidate for research |
| Important |
Inaccessible |
Find another evidence source, redesign, defer, or change the question |
| Important |
Accessible but poorly matched |
Do not force the dataset to answer the question |
| Weak or trivial |
Excellent and accessible |
Find a better question rather than analyzing data merely because they exist |
| Promising question discovered through data |
Appropriate |
Check the literature, clarify exploratory versus confirmatory aims, and develop the research rationale |
When question and data align, accessibility becomes an advantage. When they do not, something needs to change.