03 · What You Need to Know
Privacy and Confidentiality Affect What Evidence You Can Obtain
Privacy and confidentiality are related but different
Researchers sometimes use privacy and confidentiality as though they mean the same thing. They address different parts of the research relationship.
Privacy
Concerns people's interests in controlling access to themselves, their activities, circumstances, spaces, and personal information.
Confidentiality
Concerns how information entrusted to or obtained by researchers is protected from unauthorized access, use, or disclosure.
The distinction becomes clearer in practice. Secretly observing behavior in a setting where people reasonably expect not to be observed raises a privacy issue at the point of data collection. An interviewer who legitimately obtains sensitive information but later allows an unauthorized person to access the interview file has a confidentiality problem.
A study can involve both. Researchers therefore need to ask not only whether they can keep information secure after obtaining it, but whether they should obtain the information in the proposed manner in the first place.
Removing names does not necessarily make data non-identifiable
A dataset does not become anonymous simply because it lacks a column labelled "Name."
People may be identifiable through direct identifiers such as names, email addresses, identification numbers, photographs, or contact details. They may also be identifiable through combinations of indirect information. A precise date of birth, rare occupation, small workplace, unusual diagnosis, geographic location, or distinctive sequence of life events can sometimes narrow a record to one person.
This problem becomes more important as datasets grow richer and as information can potentially be combined with other available sources.
NIH's current Certificate of Confidentiality guidance, for example, defines identifiable, sensitive information broadly enough to include information for which there is at least a very small risk that a combination of the research information, a request for it, and other available data sources could be used to deduce an individual's identity.
That definition applies within the specific Certificate of Confidentiality framework, not as a universal definition for every privacy law or ethics system. It nevertheless illustrates why identifiability should be assessed in context rather than equated with the presence of obvious identifiers.
More detailed data are not automatically better data
Researchers often want maximum detail because it preserves analytical possibilities. Exact ages are preferable to age bands. Precise addresses enable spatial analysis. Detailed job titles permit occupational comparisons. Dates allow temporal linkage. Multiple demographic variables permit subgroup analysis.
Each additional detail can also increase identifiability.
This creates a methodological trade-off. Removing detail may reduce privacy risk but also reduce analytical precision. Retaining detail may improve the analysis while increasing the consequences of unauthorized disclosure.
A useful principle is data minimization: collect or retain only the personal information genuinely needed for the research purpose, subject to the requirements governing the study.
This does not mean mechanically deleting every variable that looks sensitive. If a variable is essential to the research question, removing it may make the study scientifically pointless. The researcher instead needs to determine whether the variable can be collected or used with appropriate protections or whether another design can answer the question.
Ask whether you need identity or merely linkage
Sometimes researchers do not actually need to know who participants are. They need only to connect information belonging to the same person across different records or time points.
That distinction creates useful design possibilities.
A study might use participant codes while storing the linkage key separately. A trusted data holder might perform linkage and provide researchers with a dataset in which identities cannot readily be ascertained by the research team. Researchers might receive only the variables necessary for the analysis rather than the complete source records.
Whether such arrangements remove a project from a particular legal or regulatory category depends on the applicable framework. Under current OHRP guidance for the U.S. Common Rule, for example, coded private information may in some circumstances not be considered individually identifiable to investigators when they cannot readily ascertain the identities associated with the coded information.
That is a specific regulatory interpretation. Researchers should not assume that coding automatically makes data anonymous, de-identified under another law, or exempt from all ethical obligations.
Privacy protections can change the level of analysis
Suppose you want to know whether individual employees who report workplace harassment subsequently experience poorer promotion outcomes. Answering that question requires linking an individual's report with an individual's employment trajectory.
If the organization will release only department-level counts because individual linkage would create unacceptable privacy or confidentiality risks, you may still be able to examine whether departments with more reported harassment have different aggregate promotion patterns.
But that is a different question.
Aggregate relationships do not automatically establish individual-level relationships. Protecting privacy by changing the level of analysis can therefore alter the unit of analysis, the statistical model, and the inference the study supports.
This is an important recurring principle: privacy-preserving redesign should be accompanied by inferential redesign when necessary.
Data access and confidentiality are not the same problem
You may be fully capable of protecting a dataset that an organization still has no obligation or authority to provide to you.
Hospitals, schools, government agencies, employers, platforms, archives, and other data holders operate under legal, ethical, contractual, professional, and institutional rules governing information access. A researcher's promise of secure storage does not override those rules.
Conversely, being technically able to obtain information does not necessarily mean you are ethically or legally entitled to use it for research.
If the central difficulty is that the required data cannot legally or ethically be accessed, the problem extends beyond confidentiality management. You may need another legitimate source of evidence or a different question.
Health information illustrates how privacy rules can shape research design
In the United States, the HIPAA Privacy Rule provides a useful example of how a specific legal framework can affect research using identifiable health information. It applies to covered entities and their business associates rather than to every researcher or every health-related dataset.
For research uses and disclosures of protected health information by covered entities, individual authorization is one route. The Privacy Rule also permits certain research uses or disclosures without individual authorization under specified conditions, including appropriately documented waivers or alterations approved by an IRB or Privacy Board.
A waiver is not simply granted because obtaining authorization would be inconvenient. HHS states that the relevant criteria include no more than minimal risk to privacy, an adequate plan to protect identifiers, an adequate plan to destroy identifiers at the earliest appropriate opportunity unless retention is justified or legally required, and adequate assurances concerning reuse and disclosure. The research must also be impracticable without the waiver or alteration and without access to and use of the protected health information.
These are U.S. HIPAA requirements for covered situations. Other countries and sectors operate under different privacy and data-protection regimes. The broader lesson is that researchers should identify the rules governing the actual data source rather than assume that general research ethics approval settles every data-access question.
Consent does not eliminate confidentiality obligations
A participant may knowingly agree to provide sensitive information, but that does not mean researchers can subsequently handle the information however they wish.
The consent process should accurately describe relevant data practices, and researchers remain responsible for the safeguards promised to participants and required by applicable policies, law, and ethical review.
This becomes particularly important when data may be retained for future research, shared with collaborators, deposited in repositories, transferred across institutions, or used for secondary analyses.
Researchers should avoid vague promises such as "your information will be completely confidential" when the actual study involves multiple authorized users, external services, data-sharing requirements, or legal limitations. Participants need an accurate account of the protections that actually exist.
Sometimes stronger confidentiality protection is legally available
In some jurisdictions and research systems, particular legal protections can restrict disclosure of identifiable research information.
In the United States, NIH Certificates of Confidentiality provide an important example. NIH states that Certificates protect identifiable, sensitive research information by prohibiting disclosure to people not connected with the research except in specified circumstances. Since 2017, qualifying NIH-funded research collecting or using identifiable, sensitive information is automatically deemed to have a Certificate.
These protections are significant, but they are not a universal confidentiality mechanism for all research everywhere. They also do not eliminate the need for sound data security, appropriate consent information, risk minimization, or compliance with other applicable requirements.
Researchers should determine whether any specific legal or institutional protections apply to their project rather than promising protections they do not actually possess.
Security is necessary, but security alone does not solve privacy
A perfectly encrypted dataset can still contain information that should never have been collected.
Technical safeguards such as encryption, access controls, secure transfer, authentication, audit mechanisms, and controlled storage can reduce the likelihood of unauthorized access. They do not answer whether the data collection itself is justified, whether every variable is necessary, whether consent is adequate, or whether the research team has legitimate access.
Privacy protection therefore begins before the first file is stored.
Question What information is actually needed to answer the research question?
Collection Can unnecessary identifiers and sensitive variables be avoided from the beginning?
Access Who genuinely needs access to identifiable or linkable information?
Protection What technical, organizational, procedural, and legal safeguards apply while the data exist?
Dissemination Could published tables, quotations, case descriptions, or shared datasets allow identities to be inferred?
Small samples can make confidentiality especially difficult
Confidentiality becomes challenging when the research population is small or participants have distinctive characteristics.
Imagine interviewing six department heads within one named institution. Calling them Participant A through Participant F may provide little meaningful protection if readers already know who holds each role.
The same problem can occur with rare conditions, specialized professions, small communities, senior organizational roles, unusual demographic combinations, or highly distinctive experiences.
This is particularly important when research involves small, vulnerable, or hard-to-reach populations. Removing names may protect against casual identification while doing little against identification by insiders who know the setting.
Qualitative quotations can function as identifiers
Rich qualitative evidence creates its own privacy tension. Detailed quotations give readers access to participants' voices and support transparent interpretation, but distinctive language or events can reveal identity.
A participant may describe an incident that only a handful of colleagues know about. A quotation may mention a rare role, exact location, distinctive project, or unusual personal history. Even after obvious identifiers are removed, someone familiar with the participant may recognize the account.
Researchers therefore need a publication-stage confidentiality strategy, not merely a storage-stage strategy.
Depending on the methodology, participant expectations, and approved protocol, researchers may consider removing unnecessary identifying details, masking particular contextual information, limiting quotations, or using other appropriate approaches. Any modification should preserve the substantive meaning of the data rather than silently manufacturing a cleaner narrative.
Data sharing can create a second privacy decision
Increasing expectations for research transparency and data sharing can be valuable for verification, reuse, and cumulative science. Human participant data, however, cannot always be made openly available without qualification.
NIH's current best practices for protecting participant privacy in data sharing emphasize understanding applicable laws, regulations, policies, informed consent, identifiability, and appropriate data-access controls. Depending on the data, controlled access may be more appropriate than unrestricted public release.
"Open data" should therefore not be interpreted as "upload every participant-level file to the internet." Responsible sharing depends on what participants agreed to, what the data contain, the likelihood of identification, applicable requirements, and the available access mechanisms.
Privacy can become a scientific limitation rather than a problem to engineer away
Suppose your question requires precise geographic trajectories linked to sensitive health outcomes for members of a very small population. Aggregating locations sufficiently to protect participants destroys the spatial resolution needed for the analysis. Removing the health variables eliminates the outcome. Breaking the linkage prevents the longitudinal analysis.
You have reached a genuine constraint.
The responsible response is not necessarily to weaken privacy protection until the original analysis becomes possible. Instead, consider whether another population, dataset, analytic level, design, or question can preserve the underlying research objective.
If every credible alternative changes what you are studying, then the ethical constraint may need to change the research question rather than merely the method.
Watch Out
Do not assume that information is ethically safe to use simply because it is technically obtainable, stored securely, stripped of names, or available somewhere online. Privacy, confidentiality, legal access, identifiability, participant expectations, and research ethics are related questions that may require separate assessment.