01 · The Question
Your Supervisor Has a Dataset Ready to Use. Should That Decide Your Topic?
Your supervisor offers you access to data from an earlier project, an ongoing survey, laboratory work, institutional records, interviews, or a larger research programme. The data already exist, so you may be able to avoid months of recruitment and collection. Perhaps the sample is larger than anything you could gather independently. The practical advantages are obvious.
But there is an equally important question: are you choosing a research problem that the data can answer, or are you searching for any question that will justify using convenient data?
Existing data can provide an excellent foundation for original research. They can also constrain what you are able to ask, measure, infer, and claim. The distinction matters before you commit your thesis, dissertation, or research project to a dataset simply because it is already available.
03 · What You Need to Know
What You Should Establish Before Building a Study Around Existing Data
Using existing data can be a legitimate research design
Research does not need to involve collecting a new dataset to be original. Secondary analysis involves analysing data that were collected previously, often to address a new research question. The UK Data Service describes secondary analysis as the reanalysis of existing qualitative or quantitative data, typically by a researcher seeking to address a new question.
Existing data can therefore support substantive new research. Depending on the dataset and discipline, researchers may investigate questions that were not central to the original study, apply different analytical approaches, examine subgroups, test alternative theoretical explanations, or combine existing evidence with other sources.
The relevant question is not “Did I personally collect these data?” It is “Can these data provide appropriate evidence for the question I am asking?”
Start with the research problem, then test the dataset against it
The availability of data can inspire questions, and there is nothing inherently wrong with data-driven topic development. Seeing a rich dataset may reveal patterns, populations, variables, or possibilities you had not previously considered.
Still, the resulting study needs a defensible research problem and research question. “My supervisor already has the data” establishes feasibility. It does not establish significance.
A useful discipline is to write down the proposed question before examining whether the dataset can answer it. Then map each part of the question to the actual variables, measures, observations, documents, interviews, specimens, time points, or other evidence available. If the match is poor, revise the question only when the revised version remains intellectually worthwhile.
Available data are not necessarily suitable data
A dataset can be large, clean, and immediately accessible yet still be inappropriate for your question. Variables may measure related concepts rather than the constructs you actually need. Important confounders may be absent. The population may not match the population about which you want to make claims. The timing of measurement may make temporal questions impossible to answer.
For qualitative data, the original interview questions may not have elicited sufficient depth on the issue you now want to investigate. For administrative records, variables may have been created for operational rather than research purposes. For historical data, definitions and measurement practices may have changed.
This is why choosing a topic partly around accessible data can be sensible while still requiring a careful assessment of what those data can and cannot support.
You need to understand how the data were produced
Before deciding on the topic, learn the provenance of the dataset. Who collected the data? For what original purpose? Which population was sampled? How were participants recruited or records selected? What instruments or protocols were used? When were the data collected? What cleaning, coding, transformation, exclusion, or imputation has already occurred?
Without that information, it may be difficult to evaluate bias, measurement quality, missingness, representativeness, or the defensibility of your planned analysis.
Your supervisor's familiarity with the dataset can be extremely valuable here. It can also create a temptation to treat undocumented knowledge as sufficient. Where possible, examine the original protocol, questionnaires or interview guides, codebook, data dictionary, consent materials, documentation of transformations, and relevant prior publications rather than relying solely on an informal explanation.
Permission to access data is not automatically permission to use them however you want
This is one of the most consequential distinctions. A supervisor possessing a dataset does not necessarily mean every member of the research group may use it for every new research question.
Restrictions can arise from informed consent, ethics or institutional review requirements, privacy obligations, data-use agreements, repository conditions, contractual arrangements, collaborative agreements, applicable law, or commitments made when the data were originally collected.
For research involving human data, the exact requirements depend on jurisdiction, institution, identifiability, the original consent, and the proposed secondary use. U.S. federal guidance, for example, distinguishes circumstances in which secondary research involving coded private information or biospecimens may or may not constitute human-subjects research. NIH guidance also emphasises that limitations on future data use can arise from informed consent and other conditions.
Watch Out
Do not assume that “the data are already collected” means “no ethics review is needed.” Secondary use can raise different regulatory and ethical questions from primary collection. Ask the appropriate ethics committee, institutional review board, data custodian, or equivalent institutional authority what applies to your proposed use rather than making the determination yourself.
The original consent can limit what secondary researchers may do
Where human-participant data are involved, the terms under which information was originally collected may affect future use. NIH's intramural guidance, for example, states that secondary research using existing specimens or data should be consistent with the terms of the original consent and gives examples of consent language that can restrict future topics or sharing.
These issues cannot be solved simply by changing the title of the project. If the original permission limits data use to a particular disease, population, research purpose, institution, or research team, a proposed secondary analysis outside those conditions may require further review or may not be permissible.
Requirements differ across jurisdictions and institutions, so treat this as an issue to verify, not as a universal rule that can be resolved from a generic checklist.
Convenient data can quietly reshape the question
Imagine that your real interest is in why students abandon online courses. Your supervisor has an existing dataset containing demographics, final grades, login counts, and completion status, but nothing about students' motivations, competing responsibilities, experiences, or reasons for leaving.
You could ask a narrower question about associations between recorded characteristics and course completion. What you cannot do is pretend that the dataset answers the original “why” question merely because it contains variables that are easy to analyse.
This is a common danger in existing-data research: the available variables begin to define the phenomenon. Methodological convenience can then masquerade as conceptual adequacy.
Your contribution needs to be distinguishable from your supervisor's project
If the dataset comes from your supervisor's ongoing programme, clarify what intellectual territory remains available for your work. Has your proposed question already been analysed? Is another student or collaborator working on it? Are manuscripts already in preparation? Which variables, subsamples, or analyses are reserved for other projects?
You should also discuss expectations regarding authorship, data access, analysis, publication, confidentiality, and what happens if you leave the research group. These arrangements may differ across disciplines and institutions, so the appropriate solution is explicit discussion rather than assumption.
A supervisor's expertise can strengthen a project considerably, but you should still consider how much that expertise should influence your topic without allowing the supervisor's existing programme to become your only reason for choosing it.
Existing data may make a more ambitious study feasible
The advantages can be substantial. A well-documented dataset may contain a larger sample, longer follow-up period, more expensive measurements, rarer populations, or more extensive observations than a student could realistically collect independently. Reusing data can also reduce unnecessary duplication of collection and allow researchers to extract additional scientific value from resources already created.
In these cases, data availability is not merely a shortcut. It may enable a better question than the researcher could otherwise answer.
The important qualification is that the research design must respect the properties and limitations of the existing evidence. You inherit the strengths of the dataset, but you inherit its weaknesses too.
07 · A Quick Checklist
Before Choosing a Topic Around Your Supervisor's Dataset, Check These
Before committing to the topic, check:
Can I state a meaningful research problem and question independently of the fact that these data are available?
Do the actual variables, measures, observations, or materials provide appropriate evidence for that question?
Have I examined how the data were collected, sampled, coded, cleaned, transformed, and documented?
Do I understand important missing data, measurement limitations, selection issues, and other weaknesses that I will inherit?
Have I verified that my proposed use is permitted under applicable consent terms, ethics requirements, data-use agreements, institutional policies, and other relevant restrictions?
Have I clarified whether the proposed question or analysis overlaps with work already completed or planned by my supervisor, other students, or collaborators?
Are expectations about data access, authorship, publication, confidentiality, and my own contribution sufficiently clear?
Would I still consider the resulting research question worthwhile if I had encountered the same dataset somewhere other than my supervisor's research group?