03 · What You Need to Know
Disaggregating Data Is Useful Only When the Comparison Is Meaningful
There are at least two different reasons researchers may report results separately by sex or gender. One is descriptive: readers need to see how participants or outcomes are distributed across groups. The other is analytical: researchers want to know whether an association, intervention effect, risk, experience, or other result actually differs across groups.
Those purposes should not be confused. Showing group-specific estimates is not the same as demonstrating that the groups differ statistically or substantively.
Begin by Establishing Whether Sex or Gender Is Relevant
The strongest reason to conduct a subgroup analysis comes from the research question, theory, previous evidence, biological plausibility, social mechanism, clinical importance, or another substantive rationale established before the results are examined.
For example, sex-related biological processes may plausibly influence pharmacokinetics, disease processes, or responses to some interventions. Gender-related roles, exposures, expectations, access, or social experiences may plausibly influence other outcomes.
These are different hypotheses. Researchers should first understand whether sex, gender, or a more specific underlying construct is relevant before deciding how to analyze the data.
An Overall Average Can Conceal Meaningful Heterogeneity
Suppose an intervention produces a moderate average improvement across all participants. That overall estimate may be appropriate if the effect is reasonably similar across the population. But if the intervention has substantially different effects across relevant groups, the pooled estimate can conceal information important to science, practice, or policy.
Sex- or gender-specific estimates may therefore help researchers examine heterogeneity rather than assuming that the average describes everyone equally well.
This principle is particularly important in areas where evidence has historically been generated disproportionately from one group. The U.S. Food and Drug Administration, for example, identifies analysis of data by sex as important for assessing possible sex differences in response to medical treatments. Current FDA guidance for medical device clinical studies also addresses sex-specific enrollment, analysis, and reporting.
Including Multiple Groups Does Not Guarantee Enough Data to Compare Them
A study may include participants of different sexes or genders while having very little information for one or more groups. This creates a distinction between inclusion and analytical adequacy.
Suppose 480 participants belong to one group and 20 to another. Reporting descriptive information for both may be straightforward. Estimating a group-specific treatment effect or testing an interaction may be much less precise for the smaller group.
This is one reason diversity in a research sample should not be confused with having sufficient statistical information for every subgroup analysis.
If subgroup comparisons are important to the study's objectives, sample-size and recruitment planning should address them before data collection whenever feasible.
“Significant in One Group but Not the Other” Does Not Establish a Difference Between Groups
This is one of the most consequential mistakes in subgroup analysis.
Imagine that an intervention has a statistically significant association with the outcome among women but not among men. It may be tempting to conclude that the intervention works for women but not for men.
That conclusion does not follow merely from the two p-values.
The estimates could be very similar while one group has a wider confidence interval because its sample is smaller. To evaluate whether the association or effect differs between groups, researchers generally need an analysis that directly evaluates the between-group difference, such as an appropriate interaction or effect-modification analysis, depending on the design and model.
Watch Out
A statistically significant result in one subgroup and a non-significant result in another subgroup is not, by itself, evidence that the subgroup effects are statistically different. Compare the effects directly rather than comparing whether their individual p-values cross a significance threshold.
Interaction Analysis Asks a Different Question From Separate Group Analyses
Suppose researchers model an outcome as a function of treatment, sex, and a treatment-by-sex interaction. The interaction addresses whether the treatment association differs according to sex under the specified model.
That is different from asking whether the treatment estimate is statistically distinguishable from zero separately within each sex.
The exact analytical approach depends on the study design, outcome, model, estimand, and scale on which effect modification is scientifically meaningful. Researchers should therefore avoid treating “run the model twice” as a universal method for subgroup analysis.
Exploratory and Confirmatory Analyses Should Be Distinguished
A sex- or gender-based comparison specified in advance because it addresses a central hypothesis is different from searching through many possible subgroup comparisons after seeing the results.
Exploratory subgroup analysis can be valuable. It can reveal patterns worth investigating in future research. The difficulty arises when an unexpected exploratory finding is presented with the certainty of a prespecified confirmatory result.
As the number of subgroup comparisons increases, the possibility of chance findings also increases. Researchers should report whether analyses were prespecified or exploratory and interpret unexpected subgroup patterns accordingly.
Measurement Must Support the Analysis
Before stratifying by sex or gender, researchers need to know what the variable actually represents.
If an instrument asks “Gender: Male/Female,” it may be unclear whether the intended construct was sex, gender identity, an administrative classification, or something else. Secondary datasets can present the same problem.
Splitting the data into categories does not repair ambiguous measurement. Researchers should inspect how the variable was collected and use terminology consistent with that measurement.
Sex- or Gender-Specific Analysis Does Not Automatically Explain the Difference
Suppose a study finds that an intervention has different outcomes across groups. That finding may justify further investigation, but the grouping variable does not necessarily identify the mechanism.
A sex-associated difference could reflect biological mechanisms, social exposures, treatment patterns, environmental factors, or correlated characteristics. A gender-associated pattern could likewise reflect specific social mechanisms that were never directly measured.
Researchers should therefore separate three questions:
Are the group-specific estimates different? This is the descriptive or statistical comparison.
Is the difference sufficiently precise and meaningful? This requires attention to uncertainty, magnitude, study design, and context rather than p-values alone.
Why are they different? This is a mechanistic or causal question that requires evidence about the proposed explanation.
Some Funders, Regulators, and Journals Have Specific Expectations
Analysis decisions are not governed only by researcher preference. Requirements can arise from funders, regulators, study protocols, reporting guidelines, or journals.
The U.S. National Institutes of Health currently expects sex as a biological variable to be factored into the design, analysis, and reporting of applicable NIH-funded vertebrate animal and human studies. NIH emphasizes that considering sex begins during research design rather than being added only after data collection.
In the regulatory context, the FDA's 2025 final guidance for medical device clinical studies addresses evaluation and reporting of sex-specific data. The FDA also issued draft guidance in December 2025 concerning sex differences in the clinical evaluation of medical products more broadly. Because that latter document remains draft guidance, it should not be described as a final binding requirement.
Researchers publishing health research may also encounter the SAGER guidelines, which provide recommendations for reporting sex and gender throughout research reports.
Requirements differ across organizations and research contexts. Always verify the current policy that actually applies to the study.
07 · A Quick Checklist
Before Running a Sex- or Gender-Specific Analysis, Check the Design
Before disaggregating or comparing results, check:
Identify the scientific, clinical, social, policy, or reporting reason for examining sex or gender.
Verify what the sex or gender variable actually measured and how it was collected.
Distinguish descriptive reporting of group-specific results from a statistical test of whether groups differ.
Determine whether each group contains enough information for the intended analysis and precision.
Use an appropriate direct comparison or interaction analysis when the research question concerns whether an effect differs across groups.
Do not infer a subgroup difference merely because one p-value is significant and another is not.
Distinguish prespecified analyses from exploratory subgroup findings.
Consider multiplicity and the risk of chance findings when conducting numerous subgroup analyses.
Check the current requirements of applicable funders, regulators, protocols, reporting guidelines, and journals.
Interpret observed differences without claiming a biological or social mechanism that the study did not establish.