03 · What You Need to Know
Existing Data Are Not Automatically Available Research Data
Separate data existence from data accessibility
Researchers sometimes make an early feasibility assumption: if the information exists somewhere, it can probably be obtained.
That assumption can derail an otherwise promising project.
A dataset can exist yet remain unavailable because the organization holding it has no authority to disclose it for your proposed purpose, because participant consent does not cover the intended secondary use, because the information is subject to privacy or confidentiality protections, because a data-use agreement restricts further disclosure, or because the data holder simply does not permit external research access.
Data exist
The information has been collected or is held somewhere.
Data are accessible for your research
You have a legitimate and appropriately authorized route to obtain and use the information for the particular research purpose you propose.
The second does not follow from the first.
Technical access is not the same as legitimate access
You may be technically capable of viewing, copying, downloading, scraping, linking, or otherwise obtaining information without having an ethical or legal basis for using it in research.
This distinction is particularly important with digital data. Information may be visible through an online account, obtainable through an application programming interface, discoverable through a search engine, accessible to an employee as part of their job, or retrievable from an institutional system. None of those facts alone establishes that a researcher may repurpose the information for a study.
Similarly, having access to confidential records through your professional role does not necessarily authorize you to use those records for research. Clinical, administrative, educational, employment, and research access can be governed by different permissions.
Watch Out
Never treat "I can access the data" as equivalent to "I am authorized to use the data for this research." Verify the authority, permissions, ethical requirements, and conditions governing the proposed research use before obtaining or analyzing restricted information.
Different restrictions come from different authorities
There is no single universal rule governing research data access. The relevant requirements depend on what the data are, who holds them, how they were collected, where the research takes place, and what the researcher intends to do with them.
Possible constraints may arise from privacy and data-protection law, health-information regulations, research ethics requirements, consent agreements, contractual restrictions, professional duties, institutional policy, Indigenous or community governance arrangements, intellectual property, database terms, national-security controls, or sector-specific rules.
A data holder may also impose conditions that are more restrictive than what the law would theoretically permit.
For this reason, "Is it legal?" and "Will the data holder give me access?" are separate questions. So are "Is access legally permissible?" and "Is this research use ethically defensible?"
Ethics approval does not automatically give you a right to the data
Researchers sometimes assume that once a research ethics committee or institutional review board approves the protocol, organizations holding the necessary records must provide them.
That is generally not how access works.
Ethics review addresses the responsibilities and requirements within its authority. A hospital, government agency, company, school, repository, archive, platform, or other data controller or custodian may have separate legal, governance, contractual, or institutional obligations governing disclosure.
You may therefore need ethics approval and data-holder authorization, not one instead of the other. Depending on the jurisdiction and data, additional approvals or agreements may also be required.
Conversely, permission from the data holder does not necessarily eliminate applicable ethics-review or consent requirements.
Secondary research can be subject to different rules from primary data collection
Reusing information originally collected for another purpose can reduce participant burden and make otherwise difficult research possible. But "secondary data" does not mean "unrestricted data."
Under the U.S. Common Rule, for example, particular categories of secondary research using identifiable private information or identifiable biospecimens can qualify for exemption when specified conditions are met. These include certain uses of publicly available information, certain situations in which investigators record information so that identities cannot readily be ascertained and agree not to contact or re-identify participants, certain research involving identifiable health information regulated under HIPAA, and particular federally conducted or supported uses of government information. Other secondary research provisions involve broad consent and limited IRB review.
These are specific U.S. regulatory provisions, not general permission for researchers everywhere to reuse existing records. Whether a secondary study requires consent, qualifies for an exemption, requires review, or falls outside a particular human-participant regulation should be determined under the rules applicable to the project rather than by the researcher informally.
Coded data may create an access route without giving researchers the identities
Sometimes the research requires information about the same individuals across records but does not require researchers to know who those individuals are.
That distinction can make an otherwise difficult study possible.
For example, a trusted data holder could link records using identifiers and then provide researchers with coded information under conditions that prevent the researchers from obtaining the linkage key. OHRP guidance under the U.S. Common Rule recognizes that, under specified circumstances, secondary research involving coded private information or biospecimens may not involve human subjects as defined by that regulation when investigators cannot readily ascertain the individuals' identities.
This does not mean coding universally removes every privacy, legal, or ethics requirement. Other regulatory systems may define identifiability differently, and the data holder may impose additional restrictions.
The practical lesson is narrower: if your analysis requires linkage rather than identity, ask whether an authorized intermediary can provide the necessary linked information without giving the research team access to identities.
Controlled access can sometimes replace unrestricted access
Data access is not always binary. The alternatives are not necessarily "give the researcher the complete dataset" and "deny the research."
Repositories and organizations may use controlled-access arrangements in which researchers apply for permission, justify the proposed use, agree to conditions, access data in secure environments, receive only approved variables, or are prohibited from attempting re-identification or redistributing the data.
Such arrangements can preserve research utility while reducing privacy and confidentiality risks.
But controlled access is useful only if it provides the evidence your analysis actually needs. If key variables, linkage capabilities, geographic detail, time resolution, or other essential information remain unavailable, you need to assess what question the permitted dataset can genuinely answer.
De-identification can solve one problem and create another
Suppose a hospital is willing to provide data only after removing dates, detailed locations, rare diagnoses, and several demographic characteristics that could increase identification risk.
That may substantially improve privacy protection. It may also eliminate variables central to your study.
If your question concerns whether a particular exposure precedes an outcome, removing dates may prevent temporal analysis. If your question concerns geographic inequalities, broad geographic aggregation may remove the relevant variation. If your study concerns a rare subgroup, suppressing rare categories may make that subgroup impossible to identify analytically.
Privacy-preserving data are not automatically analytically equivalent to the original records.
If the protections necessary for privacy and confidentiality remove information essential to the question, the problem may need to be addressed at the level of research design or question formulation.
Consent for the original activity may not cover your new research
Existing research datasets and biospecimens may have been collected under consent terms specifying particular purposes, types of future use, data-sharing conditions, or other limitations.
Researchers should therefore examine what participants were actually told and what permissions apply rather than assuming that prior consent authorizes any scientifically related secondary project.
Under the revised U.S. Common Rule, broad consent is one permissible mechanism for certain storage, maintenance, and secondary research uses of identifiable private information or identifiable biospecimens. It is not the only possible regulatory pathway for secondary research, nor is it a universal requirement.
Other frameworks may use different concepts and requirements. The governing consent and data-use conditions should be verified for the specific source.
A waiver may exist in some systems, but inconvenience is not enough
Some ethical and regulatory frameworks allow consent or authorization requirements to be waived under specified conditions. Researchers should not interpret this as a general mechanism for accessing any dataset when obtaining consent would be difficult.
For research governed by the U.S. Common Rule, for example, waiver or alteration of informed consent is subject to specified criteria. These include no more than minimal risk, protection of participants' rights and welfare, impracticability of carrying out the research without the waiver or alteration, and, when identifiable private information or identifiable biospecimens are involved, impracticability of conducting the research without using them in an identifiable format.
Other regulatory regimes use different standards.
The important point is that "contacting everyone would take too long" is not itself a universal entitlement to a waiver. Researchers should determine whether a waiver mechanism actually applies and whether the competent reviewing authority finds that its requirements are satisfied.
Publicly available information still requires careful interpretation
Some regulatory frameworks distinguish publicly available information from private information. Under the U.S. Common Rule, for example, one secondary-research exemption applies when identifiable private information or identifiable biospecimens are publicly available.
But "publicly available" is not a magic phrase that resolves every research-ethics question in every jurisdiction.
Online environments are particularly complicated. Information may be technically visible while users perceive the context as socially private. Large-scale aggregation, linkage, quotation, or re-identification can create consequences very different from those associated with an individual post being casually viewed.
Researchers using internet data should therefore assess the applicable ethical, legal, platform, and institutional requirements for the proposed collection and use rather than relying solely on whether a web page can be opened without a password.
Data access agreements are substantive research constraints
A data-use or access agreement may specify what researchers may analyze, where data may be stored, who may access them, whether linkage is permitted, whether results require disclosure review, when files must be destroyed, and whether data may be shared with collaborators.
These conditions can affect the research design itself.
For example, if an agreement prohibits attempting to identify individuals, a planned analysis that effectively depends on identifying particular cases is incompatible with the access conditions. If the agreement permits use only for a specified research purpose, a later secondary question may require new authorization.
Researchers should read access conditions before finalizing the analysis plan rather than treating the agreement as administrative paperwork to sign after the methodology has been settled.
Do not substitute a convenient dataset without checking what question it answers
When the ideal dataset is inaccessible, researchers often search for whatever dataset is available. That can be sensible, but it creates a methodological risk.
Suppose you want to study whether disciplinary actions at universities disproportionately affect a particular student population. Individual administrative records would provide relevant evidence, but you cannot legitimately access them. You locate a public dataset containing university-level disciplinary counts and demographic composition.
The public dataset may support a useful institutional-level study. It does not necessarily answer whether individual students from a particular group face different probabilities of disciplinary action.
The replacement dataset has changed the unit of analysis.
Before adopting alternative data, ask what constructs it measures, which population it represents, what time period it covers, what selection processes generated it, what variables are missing, and what inference its level of analysis supports.
Primary data collection may be an alternative, but not always
If existing records are inaccessible, you might collect new data directly from participants.
That can sometimes work. A researcher denied access to institutional records might survey participants about their own experiences. A study unable to obtain administrative outcomes might collect self-reported outcomes prospectively.
But primary collection produces different evidence. Self-report may not substitute adequately for verified administrative records. Recruitment may introduce selection bias. Participants may not know or remember the information contained in official records. Sensitive questions may also create new participant risks.
Primary data collection therefore needs its own scientific and ethical evaluation. It should not be treated as a universal replacement for inaccessible records.
Indirect evidence can still be valuable
A research question does not always require one perfect dataset.
Researchers may sometimes combine evidence from several imperfect but legitimate sources. Surveys, interviews, aggregate statistics, public records, existing literature, natural experiments, archival materials, or other sources may illuminate different parts of the phenomenon.
This approach can be particularly useful when direct individual-level evidence is unavailable. Multiple sources may reveal whether different observations converge on a similar interpretation.
But triangulation does not manufacture information that none of the sources contains. Several indirect datasets do not automatically become equivalent to direct evidence. Researchers should remain explicit about which parts of the original question the available evidence can and cannot address.
Sometimes you should change the question rather than keep searching for the forbidden data
There is a point at which repeated attempts to obtain restricted data stop being productive.
If the research question requires a specific form of evidence, no legitimate access mechanism exists, alternatives cannot provide an adequate basis for the intended inference, and the access restriction is unlikely to change within the project timeframe, the question may simply be infeasible for the present study.
This does not necessarily mean the question lacks scientific value. It means you cannot currently answer it responsibly with the evidence available to you.
When the constraint fundamentally determines what can be known, changing the question may be more appropriate than repeatedly changing the method.
And in some cases, the question may remain worth asking even when no ethical study can answer it directly. Recognizing that boundary is a legitimate scholarly conclusion, not a failure of imagination.