Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can Privacy or Confidentiality Requirements Make a Research Question Infeasible?

Some research questions require information that cannot responsibly be collected, linked, accessed, or reported in the form researchers would prefer. Before abandoning the question, determine whether you can reduce identifiability, change the data source or level of analysis, or narrow the question without losing what you actually need to know.

467
Privacy, Confidentiality, and Research Feasibility Guide 467 of 533
01 · The Question

What If Answering Your Question Requires Information You Cannot Adequately Protect?

Many research questions become more informative when researchers can connect detailed information to individual people. You may want to link educational records with socioeconomic characteristics, follow patients across time, compare employee reports with workplace outcomes, map experiences geographically, or combine multiple datasets to understand how different factors interact.

The difficulty is that detailed data can also make people identifiable. Removing names may not solve the problem if combinations of age, occupation, location, diagnosis, dates, quotations, or other characteristics can still reveal who someone is.

At some point, privacy and confidentiality stop being matters to address after the research design has been chosen. They become constraints on what information can responsibly be obtained and, in some cases, on whether the original research question can be answered at all.

02 · The Short Answer

Yes, Privacy Requirements Can Limit What a Study Can Realistically Answer

In Brief

Yes. Privacy and confidentiality requirements can make a research question infeasible when answering it depends on identifiable or sensitive information that researchers cannot legitimately access, collect, link, protect, retain, or disclose with acceptable risk.

That does not mean every privacy problem requires abandoning the question. Data minimization, coding, de-identification, restricted access, aggregation, alternative data sources, or a different research design may reduce the problem. But if the protections necessary to safeguard participants also remove information essential to answering the question, the question itself may need to be narrowed or changed.

03 · What You Need to Know

Privacy and Confidentiality Affect What Evidence You Can Obtain

Privacy and confidentiality are related but different

Researchers sometimes use privacy and confidentiality as though they mean the same thing. They address different parts of the research relationship.

Privacy Concerns people's interests in controlling access to themselves, their activities, circumstances, spaces, and personal information.
Confidentiality Concerns how information entrusted to or obtained by researchers is protected from unauthorized access, use, or disclosure.

The distinction becomes clearer in practice. Secretly observing behavior in a setting where people reasonably expect not to be observed raises a privacy issue at the point of data collection. An interviewer who legitimately obtains sensitive information but later allows an unauthorized person to access the interview file has a confidentiality problem.

A study can involve both. Researchers therefore need to ask not only whether they can keep information secure after obtaining it, but whether they should obtain the information in the proposed manner in the first place.

Removing names does not necessarily make data non-identifiable

A dataset does not become anonymous simply because it lacks a column labelled "Name."

People may be identifiable through direct identifiers such as names, email addresses, identification numbers, photographs, or contact details. They may also be identifiable through combinations of indirect information. A precise date of birth, rare occupation, small workplace, unusual diagnosis, geographic location, or distinctive sequence of life events can sometimes narrow a record to one person.

This problem becomes more important as datasets grow richer and as information can potentially be combined with other available sources.

NIH's current Certificate of Confidentiality guidance, for example, defines identifiable, sensitive information broadly enough to include information for which there is at least a very small risk that a combination of the research information, a request for it, and other available data sources could be used to deduce an individual's identity.

That definition applies within the specific Certificate of Confidentiality framework, not as a universal definition for every privacy law or ethics system. It nevertheless illustrates why identifiability should be assessed in context rather than equated with the presence of obvious identifiers.

More detailed data are not automatically better data

Researchers often want maximum detail because it preserves analytical possibilities. Exact ages are preferable to age bands. Precise addresses enable spatial analysis. Detailed job titles permit occupational comparisons. Dates allow temporal linkage. Multiple demographic variables permit subgroup analysis.

Each additional detail can also increase identifiability.

This creates a methodological trade-off. Removing detail may reduce privacy risk but also reduce analytical precision. Retaining detail may improve the analysis while increasing the consequences of unauthorized disclosure.

A useful principle is data minimization: collect or retain only the personal information genuinely needed for the research purpose, subject to the requirements governing the study.

This does not mean mechanically deleting every variable that looks sensitive. If a variable is essential to the research question, removing it may make the study scientifically pointless. The researcher instead needs to determine whether the variable can be collected or used with appropriate protections or whether another design can answer the question.

Ask whether you need identity or merely linkage

Sometimes researchers do not actually need to know who participants are. They need only to connect information belonging to the same person across different records or time points.

That distinction creates useful design possibilities.

A study might use participant codes while storing the linkage key separately. A trusted data holder might perform linkage and provide researchers with a dataset in which identities cannot readily be ascertained by the research team. Researchers might receive only the variables necessary for the analysis rather than the complete source records.

Whether such arrangements remove a project from a particular legal or regulatory category depends on the applicable framework. Under current OHRP guidance for the U.S. Common Rule, for example, coded private information may in some circumstances not be considered individually identifiable to investigators when they cannot readily ascertain the identities associated with the coded information.

That is a specific regulatory interpretation. Researchers should not assume that coding automatically makes data anonymous, de-identified under another law, or exempt from all ethical obligations.

Privacy protections can change the level of analysis

Suppose you want to know whether individual employees who report workplace harassment subsequently experience poorer promotion outcomes. Answering that question requires linking an individual's report with an individual's employment trajectory.

If the organization will release only department-level counts because individual linkage would create unacceptable privacy or confidentiality risks, you may still be able to examine whether departments with more reported harassment have different aggregate promotion patterns.

But that is a different question.

Aggregate relationships do not automatically establish individual-level relationships. Protecting privacy by changing the level of analysis can therefore alter the unit of analysis, the statistical model, and the inference the study supports.

This is an important recurring principle: privacy-preserving redesign should be accompanied by inferential redesign when necessary.

Data access and confidentiality are not the same problem

You may be fully capable of protecting a dataset that an organization still has no obligation or authority to provide to you.

Hospitals, schools, government agencies, employers, platforms, archives, and other data holders operate under legal, ethical, contractual, professional, and institutional rules governing information access. A researcher's promise of secure storage does not override those rules.

Conversely, being technically able to obtain information does not necessarily mean you are ethically or legally entitled to use it for research.

If the central difficulty is that the required data cannot legally or ethically be accessed, the problem extends beyond confidentiality management. You may need another legitimate source of evidence or a different question.

Health information illustrates how privacy rules can shape research design

In the United States, the HIPAA Privacy Rule provides a useful example of how a specific legal framework can affect research using identifiable health information. It applies to covered entities and their business associates rather than to every researcher or every health-related dataset.

For research uses and disclosures of protected health information by covered entities, individual authorization is one route. The Privacy Rule also permits certain research uses or disclosures without individual authorization under specified conditions, including appropriately documented waivers or alterations approved by an IRB or Privacy Board.

A waiver is not simply granted because obtaining authorization would be inconvenient. HHS states that the relevant criteria include no more than minimal risk to privacy, an adequate plan to protect identifiers, an adequate plan to destroy identifiers at the earliest appropriate opportunity unless retention is justified or legally required, and adequate assurances concerning reuse and disclosure. The research must also be impracticable without the waiver or alteration and without access to and use of the protected health information.

These are U.S. HIPAA requirements for covered situations. Other countries and sectors operate under different privacy and data-protection regimes. The broader lesson is that researchers should identify the rules governing the actual data source rather than assume that general research ethics approval settles every data-access question.

Consent does not eliminate confidentiality obligations

A participant may knowingly agree to provide sensitive information, but that does not mean researchers can subsequently handle the information however they wish.

The consent process should accurately describe relevant data practices, and researchers remain responsible for the safeguards promised to participants and required by applicable policies, law, and ethical review.

This becomes particularly important when data may be retained for future research, shared with collaborators, deposited in repositories, transferred across institutions, or used for secondary analyses.

Researchers should avoid vague promises such as "your information will be completely confidential" when the actual study involves multiple authorized users, external services, data-sharing requirements, or legal limitations. Participants need an accurate account of the protections that actually exist.

Sometimes stronger confidentiality protection is legally available

In some jurisdictions and research systems, particular legal protections can restrict disclosure of identifiable research information.

In the United States, NIH Certificates of Confidentiality provide an important example. NIH states that Certificates protect identifiable, sensitive research information by prohibiting disclosure to people not connected with the research except in specified circumstances. Since 2017, qualifying NIH-funded research collecting or using identifiable, sensitive information is automatically deemed to have a Certificate.

These protections are significant, but they are not a universal confidentiality mechanism for all research everywhere. They also do not eliminate the need for sound data security, appropriate consent information, risk minimization, or compliance with other applicable requirements.

Researchers should determine whether any specific legal or institutional protections apply to their project rather than promising protections they do not actually possess.

Security is necessary, but security alone does not solve privacy

A perfectly encrypted dataset can still contain information that should never have been collected.

Technical safeguards such as encryption, access controls, secure transfer, authentication, audit mechanisms, and controlled storage can reduce the likelihood of unauthorized access. They do not answer whether the data collection itself is justified, whether every variable is necessary, whether consent is adequate, or whether the research team has legitimate access.

Privacy protection therefore begins before the first file is stored.

Question What information is actually needed to answer the research question?
Collection Can unnecessary identifiers and sensitive variables be avoided from the beginning?
Access Who genuinely needs access to identifiable or linkable information?
Protection What technical, organizational, procedural, and legal safeguards apply while the data exist?
Dissemination Could published tables, quotations, case descriptions, or shared datasets allow identities to be inferred?

Small samples can make confidentiality especially difficult

Confidentiality becomes challenging when the research population is small or participants have distinctive characteristics.

Imagine interviewing six department heads within one named institution. Calling them Participant A through Participant F may provide little meaningful protection if readers already know who holds each role.

The same problem can occur with rare conditions, specialized professions, small communities, senior organizational roles, unusual demographic combinations, or highly distinctive experiences.

This is particularly important when research involves small, vulnerable, or hard-to-reach populations. Removing names may protect against casual identification while doing little against identification by insiders who know the setting.

Qualitative quotations can function as identifiers

Rich qualitative evidence creates its own privacy tension. Detailed quotations give readers access to participants' voices and support transparent interpretation, but distinctive language or events can reveal identity.

A participant may describe an incident that only a handful of colleagues know about. A quotation may mention a rare role, exact location, distinctive project, or unusual personal history. Even after obvious identifiers are removed, someone familiar with the participant may recognize the account.

Researchers therefore need a publication-stage confidentiality strategy, not merely a storage-stage strategy.

Depending on the methodology, participant expectations, and approved protocol, researchers may consider removing unnecessary identifying details, masking particular contextual information, limiting quotations, or using other appropriate approaches. Any modification should preserve the substantive meaning of the data rather than silently manufacturing a cleaner narrative.

Data sharing can create a second privacy decision

Increasing expectations for research transparency and data sharing can be valuable for verification, reuse, and cumulative science. Human participant data, however, cannot always be made openly available without qualification.

NIH's current best practices for protecting participant privacy in data sharing emphasize understanding applicable laws, regulations, policies, informed consent, identifiability, and appropriate data-access controls. Depending on the data, controlled access may be more appropriate than unrestricted public release.

"Open data" should therefore not be interpreted as "upload every participant-level file to the internet." Responsible sharing depends on what participants agreed to, what the data contain, the likelihood of identification, applicable requirements, and the available access mechanisms.

Privacy can become a scientific limitation rather than a problem to engineer away

Suppose your question requires precise geographic trajectories linked to sensitive health outcomes for members of a very small population. Aggregating locations sufficiently to protect participants destroys the spatial resolution needed for the analysis. Removing the health variables eliminates the outcome. Breaking the linkage prevents the longitudinal analysis.

You have reached a genuine constraint.

The responsible response is not necessarily to weaken privacy protection until the original analysis becomes possible. Instead, consider whether another population, dataset, analytic level, design, or question can preserve the underlying research objective.

If every credible alternative changes what you are studying, then the ethical constraint may need to change the research question rather than merely the method.

Watch Out

Do not assume that information is ethically safe to use simply because it is technically obtainable, stored securely, stripped of names, or available somewhere online. Privacy, confidentiality, legal access, identifiability, participant expectations, and research ethics are related questions that may require separate assessment.

04 · A Practical Example

When Protecting Identity Changes the Question You Can Answer

Hypothetical Example

Linking reports of workplace harassment to career outcomes

A researcher wants to determine whether employees who formally report workplace harassment subsequently experience different promotion outcomes from comparable employees who do not report harassment.

Evidence needed The study requires individual-level linkage between harassment reports, employment characteristics, and subsequent promotion records.
Privacy problem The organization is small, harassment reports are highly sensitive, and relatively few employees have filed them. Combining department, role, report date, and promotion history could make participants readily identifiable.
First redesign The researcher removes names and employee numbers but discovers that the remaining variables still permit likely identification of several individuals.
Stronger protection considered The organization offers only aggregated department-level information and broader time periods.
Scientific consequence Those data can no longer establish whether individual employees who reported harassment experienced different promotion outcomes. They can support a different analysis concerning department-level patterns.
Decision The researcher must either establish a legitimate and adequately protected arrangement for the necessary individual-level linkage, find another suitable source of evidence, or revise the research question to match what the protected data can actually answer.

Simply calling the aggregated dataset a privacy-preserving version of the original data would conceal an important methodological change. The unit of analysis and inferential target have changed.

Privacy protection can therefore affect not just how data are stored, but what claim the research is capable of making.

05 · What Researchers Often Get Wrong

Common Mistakes About Privacy, Confidentiality, and Research Feasibility

Misconception

If You Remove Names, Are the Data Anonymous?

Not necessarily. Indirect identifiers and combinations of variables can sometimes reveal identity. The likelihood depends on the data, population, available external information, and who may attempt identification. Researchers should evaluate identifiability based on the actual dataset and context.

Misconception

Are Anonymous and Confidential Research the Same Thing?

No. In confidential research, researchers may know participants' identities or possess information that can be linked to them but restrict access and disclosure. Truly anonymous collection does not give the researcher the same ability to link responses to identifiable individuals. Researchers should not promise anonymity when they actually mean confidentiality.

Misconception

If Participants Consent, Can You Collect Any Personal Information You Need?

No. Consent does not by itself establish that every requested variable is necessary, that risks are adequately minimized, that the researcher has legal authority to access third-party records, or that the proposed data handling satisfies applicable requirements. Consent is important, but it is not a universal exemption from privacy obligations.

Misconception

If the Data Holder Gives Permission, Is the Privacy Problem Solved?

Not necessarily. Organizational permission, research ethics requirements, participant consent or authorization, data-protection law, confidentiality obligations, contractual restrictions, and other requirements may all be relevant. Which apply depends on the jurisdiction and data source.

Misconception

Does Encryption Make Sensitive Data Safe to Collect?

Encryption can reduce particular security risks, but it does not establish that collecting the information is necessary or legitimate. Strong cybersecurity cannot compensate for unjustified collection, inappropriate access, misleading consent, or an analytically unnecessary identifier.

Misconception

Can You Always Solve Privacy Problems by Aggregating the Data?

No. Aggregation can reduce identification risk, but it may also change the unit of analysis and remove information required to answer an individual-level question. When that happens, the research question and intended conclusions must change with the data.

06 · What This Means for You

Determine the Minimum Information Your Question Actually Requires

When privacy or confidentiality begins to threaten feasibility, return to the research question before adding more security technology.

Write down exactly which information must be available for each planned inference. Distinguish variables that are scientifically necessary from variables that are merely useful to have. Then identify which of those data need to remain identifiable, which require only linkage, and which could be aggregated, generalized, or omitted.

A simple decision framework

If direct identifiers are not analytically necessary
Avoid collecting them where appropriate or separate them from research data according to the approved data-management plan.
If identity is unnecessary but record linkage is necessary
Consider coding, trusted linkage, controlled access, or another arrangement that limits researchers' access to identity while preserving the required linkage.
If detailed variables create re-identification risk
Determine whether less granular information can support the analysis without changing the intended inference.
If aggregation changes the unit of analysis
Revise the question and conclusions rather than treating aggregate evidence as though it answers the original individual-level question.
If the current data source cannot legitimately provide what you need
Look for another lawful and ethically appropriate source rather than attempting to bypass the data holder's restrictions.
If no available protection preserves both participant confidentiality and the information essential to the analysis
Narrow, reformulate, or replace the research question.

These decisions are easier when privacy is considered before the research question and data plan become fixed. Once researchers have built an entire project around a dataset they cannot legitimately obtain, methodological flexibility becomes considerably more expensive.

If your question concerns particularly sensitive information, also examine whether the proposed data collection creates risks that cannot be adequately minimized or justified. Privacy is one dimension of participant protection, not the whole ethical assessment.

The practical objective is not maximum privacy regardless of scientific consequence, nor maximum data richness regardless of privacy. It is a defensible alignment between the information genuinely required, the protections participants deserve, the permissions and rules governing access, and the claims the resulting evidence can support.

07 · A Quick Checklist

Before Building a Study Around Sensitive or Identifiable Data

Before finalizing the data requirements, check:
Identify exactly which personal, sensitive, or potentially identifying variables are necessary to answer the research question.
Remove variables collected merely because they might be useful later when they are not justified by the research purpose.
Assess indirect identification from combinations of demographics, dates, locations, roles, diagnoses, quotations, or other contextual details.
Determine whether you need to know participant identity or merely need to link records belonging to the same individual.
Verify that you have legitimate authority to obtain and use the data rather than assuming that secure storage is sufficient.
Check which ethics, privacy, data-protection, contractual, institutional, and sector-specific requirements apply to the actual data source and jurisdiction.
Plan who will access identifiable or linkable information and whether each person's access is genuinely necessary.
Examine whether published quotations, tables, case descriptions, subgroup results, or shared datasets could reveal participants indirectly.
Determine whether privacy-preserving aggregation or de-identification would change the unit of analysis or inference your question requires.
If adequate privacy protection removes information essential to the analysis, reconsider the method, data source, or research question before collecting data.
08 · Frequently Asked Questions

Frequently Asked Questions About Privacy and Confidentiality in Research

What is the difference between privacy and confidentiality in research?

Privacy concerns access to people and information about them, while confidentiality concerns how information obtained by researchers is subsequently protected from unauthorized use or disclosure. A study may therefore create a privacy problem during data collection, a confidentiality problem during data handling, or both.

Does removing names make research data anonymous?

Not necessarily. Other variables or combinations of variables may permit identification. Whether data are considered identifiable, coded, de-identified, or anonymous also depends on the definition used by the applicable ethical, legal, institutional, or regulatory framework.

Can I use coded data instead of identifiable data?

Often, if the research requires linkage but researchers do not need direct access to identity. Coding can reduce some risks, but coded data are not automatically anonymous. Who holds the key, whether researchers can access it, and which regulatory definitions apply all matter.

Can I use medical records for research without asking every patient?

That depends on the applicable jurisdiction and regulatory framework. In the United States, for example, the HIPAA Privacy Rule permits certain research uses and disclosures of protected health information without individual authorization when specified requirements are met, including some appropriately approved waivers. Research ethics and Common Rule requirements may also apply separately. Do not assume that the same rules apply to every institution, dataset, or country.

Are public online data free to use for research?

Not automatically. Public accessibility does not settle every question concerning reasonable expectations of privacy, platform terms, identifiability, sensitivity, research ethics, intellectual property, or applicable law. The ethical and legal analysis depends on what information is being used, how it was obtained, the research context, and the governing requirements.

Can I publish quotations from confidential interviews?

Potentially, if doing so is consistent with the study's consent and confidentiality arrangements and applicable requirements. Researchers should assess whether the wording or surrounding details could identify the participant or another person, particularly in small populations or distinctive cases.

Does a Certificate of Confidentiality make research data completely secret?

No. In the United States, Certificates of Confidentiality provide important statutory protections against certain disclosures of identifiable, sensitive research information, but specified exceptions exist. They do not replace appropriate data security, consent information, ethics review, or other legal and institutional requirements.

What if I cannot answer the question without identifiable data?

Determine whether an appropriately protected and legitimately authorized arrangement can permit the necessary use. If not, consider another data source, design, level of analysis, or narrower question. If every ethical alternative removes information essential to the inference, the original question may not currently be answerable through an ethical study.

09 · The Bottom Line

Sometimes Protecting Participants Changes What You Can Know

The Bottom Line

Yes, privacy or confidentiality requirements can make a research question infeasible when the evidence needed to answer it depends on information that cannot be legitimately obtained, adequately protected, or used without unacceptable identification or disclosure risk.

Before abandoning the question, determine whether you truly need identity, whether less detailed or differently controlled data can preserve the necessary inference, and whether another legitimate data source or design is available. If adequate protection removes the very information the question requires, revise the question rather than weakening participant privacy merely to preserve the original analysis.

10 · Sources and Further Reading

Authoritative Sources on Research Privacy and Confidentiality

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes