Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

When Should Research Analyze Results by Sex or Gender?

Analyzing results by sex or gender can reveal important differences that an overall average conceals, but not every dataset supports a meaningful subgroup comparison. The decision should follow from the research question, study design, sample size, and relevant scientific or reporting requirements.

100
Analyzing Results by Sex or Gender Guide 100 of 217
01 · The Question

Should You Split Your Results by Sex or Gender Just Because You Collected the Variable?

A researcher completes the primary analysis and then notices that the dataset also contains sex or gender. Should the results now be divided into groups? Should separate significance tests be run? If one group shows a statistically significant result and another does not, does that mean the effect differs between them?

Not necessarily.

Analyzing results by sex or gender can be scientifically important when these variables are relevant to the phenomenon, intervention, exposure, or outcome. But subgroup analysis also creates methodological problems when it is poorly planned, inadequately powered, or interpreted through separate significance tests rather than a direct assessment of whether groups differ.

02 · The Short Answer

Analyze by Sex or Gender When the Question and Design Give You a Reason to Do So

In Brief

Research should analyze results by sex or gender when there is a scientific, clinical, social, policy, or reporting reason to examine whether findings differ across those groups and the study contains sufficient information for a meaningful analysis.

Disaggregating results can reveal patterns hidden by an overall average, but small subgroups, multiple exploratory comparisons, ambiguous measurement, and incorrect interpretation of separate significance tests can produce misleading conclusions. The analysis should therefore be planned and justified whenever possible.

03 · What You Need to Know

Disaggregating Data Is Useful Only When the Comparison Is Meaningful

There are at least two different reasons researchers may report results separately by sex or gender. One is descriptive: readers need to see how participants or outcomes are distributed across groups. The other is analytical: researchers want to know whether an association, intervention effect, risk, experience, or other result actually differs across groups.

Those purposes should not be confused. Showing group-specific estimates is not the same as demonstrating that the groups differ statistically or substantively.

Begin by Establishing Whether Sex or Gender Is Relevant

The strongest reason to conduct a subgroup analysis comes from the research question, theory, previous evidence, biological plausibility, social mechanism, clinical importance, or another substantive rationale established before the results are examined.

For example, sex-related biological processes may plausibly influence pharmacokinetics, disease processes, or responses to some interventions. Gender-related roles, exposures, expectations, access, or social experiences may plausibly influence other outcomes.

These are different hypotheses. Researchers should first understand whether sex, gender, or a more specific underlying construct is relevant before deciding how to analyze the data.

An Overall Average Can Conceal Meaningful Heterogeneity

Suppose an intervention produces a moderate average improvement across all participants. That overall estimate may be appropriate if the effect is reasonably similar across the population. But if the intervention has substantially different effects across relevant groups, the pooled estimate can conceal information important to science, practice, or policy.

Sex- or gender-specific estimates may therefore help researchers examine heterogeneity rather than assuming that the average describes everyone equally well.

This principle is particularly important in areas where evidence has historically been generated disproportionately from one group. The U.S. Food and Drug Administration, for example, identifies analysis of data by sex as important for assessing possible sex differences in response to medical treatments. Current FDA guidance for medical device clinical studies also addresses sex-specific enrollment, analysis, and reporting.

Including Multiple Groups Does Not Guarantee Enough Data to Compare Them

A study may include participants of different sexes or genders while having very little information for one or more groups. This creates a distinction between inclusion and analytical adequacy.

Suppose 480 participants belong to one group and 20 to another. Reporting descriptive information for both may be straightforward. Estimating a group-specific treatment effect or testing an interaction may be much less precise for the smaller group.

This is one reason diversity in a research sample should not be confused with having sufficient statistical information for every subgroup analysis.

If subgroup comparisons are important to the study's objectives, sample-size and recruitment planning should address them before data collection whenever feasible.

“Significant in One Group but Not the Other” Does Not Establish a Difference Between Groups

This is one of the most consequential mistakes in subgroup analysis.

Imagine that an intervention has a statistically significant association with the outcome among women but not among men. It may be tempting to conclude that the intervention works for women but not for men.

That conclusion does not follow merely from the two p-values.

The estimates could be very similar while one group has a wider confidence interval because its sample is smaller. To evaluate whether the association or effect differs between groups, researchers generally need an analysis that directly evaluates the between-group difference, such as an appropriate interaction or effect-modification analysis, depending on the design and model.

Watch Out

A statistically significant result in one subgroup and a non-significant result in another subgroup is not, by itself, evidence that the subgroup effects are statistically different. Compare the effects directly rather than comparing whether their individual p-values cross a significance threshold.

Interaction Analysis Asks a Different Question From Separate Group Analyses

Suppose researchers model an outcome as a function of treatment, sex, and a treatment-by-sex interaction. The interaction addresses whether the treatment association differs according to sex under the specified model.

That is different from asking whether the treatment estimate is statistically distinguishable from zero separately within each sex.

The exact analytical approach depends on the study design, outcome, model, estimand, and scale on which effect modification is scientifically meaningful. Researchers should therefore avoid treating “run the model twice” as a universal method for subgroup analysis.

Exploratory and Confirmatory Analyses Should Be Distinguished

A sex- or gender-based comparison specified in advance because it addresses a central hypothesis is different from searching through many possible subgroup comparisons after seeing the results.

Exploratory subgroup analysis can be valuable. It can reveal patterns worth investigating in future research. The difficulty arises when an unexpected exploratory finding is presented with the certainty of a prespecified confirmatory result.

As the number of subgroup comparisons increases, the possibility of chance findings also increases. Researchers should report whether analyses were prespecified or exploratory and interpret unexpected subgroup patterns accordingly.

Measurement Must Support the Analysis

Before stratifying by sex or gender, researchers need to know what the variable actually represents.

If an instrument asks “Gender: Male/Female,” it may be unclear whether the intended construct was sex, gender identity, an administrative classification, or something else. Secondary datasets can present the same problem.

Splitting the data into categories does not repair ambiguous measurement. Researchers should inspect how the variable was collected and use terminology consistent with that measurement.

Sex- or Gender-Specific Analysis Does Not Automatically Explain the Difference

Suppose a study finds that an intervention has different outcomes across groups. That finding may justify further investigation, but the grouping variable does not necessarily identify the mechanism.

A sex-associated difference could reflect biological mechanisms, social exposures, treatment patterns, environmental factors, or correlated characteristics. A gender-associated pattern could likewise reflect specific social mechanisms that were never directly measured.

Researchers should therefore separate three questions:

Are the group-specific estimates different? This is the descriptive or statistical comparison.
Is the difference sufficiently precise and meaningful? This requires attention to uncertainty, magnitude, study design, and context rather than p-values alone.
Why are they different? This is a mechanistic or causal question that requires evidence about the proposed explanation.

Some Funders, Regulators, and Journals Have Specific Expectations

Analysis decisions are not governed only by researcher preference. Requirements can arise from funders, regulators, study protocols, reporting guidelines, or journals.

The U.S. National Institutes of Health currently expects sex as a biological variable to be factored into the design, analysis, and reporting of applicable NIH-funded vertebrate animal and human studies. NIH emphasizes that considering sex begins during research design rather than being added only after data collection.

In the regulatory context, the FDA's 2025 final guidance for medical device clinical studies addresses evaluation and reporting of sex-specific data. The FDA also issued draft guidance in December 2025 concerning sex differences in the clinical evaluation of medical products more broadly. Because that latter document remains draft guidance, it should not be described as a final binding requirement.

Researchers publishing health research may also encounter the SAGER guidelines, which provide recommendations for reporting sex and gender throughout research reports.

Requirements differ across organizations and research contexts. Always verify the current policy that actually applies to the study.

04 · A Practical Example

Why Two Separate P-Values Can Give the Wrong Impression

Hypothetical Example

An intervention appears significant in one group but not another

Researchers evaluate an intervention and report the outcome separately for two groups. In Group A, the estimated improvement is 5.0 points with a relatively narrow confidence interval and a p-value below the study's significance threshold. In Group B, the estimated improvement is 4.5 points, but the confidence interval is wider and the p-value is above the threshold because substantially fewer participants were recruited.

Tempting conclusion The intervention works in Group A but not in Group B.
What the estimates show The estimated improvements, 5.0 and 4.5 points, are actually quite similar.
Why the p-values differ Group B has less information and therefore greater uncertainty around its estimate.
Relevant comparison If the research question concerns whether the intervention effect differs between groups, the analysis should directly evaluate that difference using a method appropriate to the design and model.
Interpretation Different significance classifications do not establish different effects. The estimated difference between groups and its uncertainty are what matter for the subgroup question.

This is why subgroup analysis needs to be designed around the comparison researchers actually want to make. Splitting a dataset and inspecting which p-values become significant is not a substitute for testing the relevant hypothesis.

05 · What Researchers Often Get Wrong

Common Mistakes in Sex- and Gender-Based Analysis

Misconception

If You Collected Sex or Gender, You Must Always Test for Group Differences

Not every recorded demographic variable requires inferential subgroup analysis. The comparison should have a scientific, clinical, social, policy, or reporting rationale and should be supported by the study design and available data.

Misconception

Significant in One Group and Not Significant in Another Means the Groups Differ

This is incorrect. The difference between a significant and non-significant result is not itself necessarily statistically significant. Researchers need to evaluate the between-group difference directly.

Misconception

Equal Numbers in Every Group Are Always Required

No. Appropriate allocation depends on the research objectives, population, expected effects, sampling design, analytical plan, and precision requirements. Equal group sizes can sometimes improve efficiency for comparisons, but they are not a universal requirement.

Misconception

Any Detected Difference Can Be Explained by Sex or Gender

A group difference is not automatically a mechanistic explanation. Researchers should avoid attributing differences to biological sex or gender-related processes unless those explanations are supported by the study and relevant evidence.

Misconception

More Subgroup Analyses Always Produce a More Complete Study

Running many poorly motivated comparisons can increase false-positive findings and make interpretation unstable. A smaller number of well-justified analyses is often more informative than searching every demographic variable for statistically significant differences.

06 · What This Means for You

Plan the Comparison Before the Dataset Makes It Tempting

If sex or gender could plausibly modify an important association or intervention effect, consider that possibility while designing the study. Decide what variable needs to be measured, whether subgroup estimates or a direct comparison is required, and whether the proposed sample can support the analysis.

A simple decision framework

If previous evidence, theory, biology, or social mechanisms suggest an important difference
Prespecify the relevant sex- or gender-based analysis when feasible and design recruitment and sample size with that analysis in mind.
If a funder, regulator, protocol, or reporting standard requires particular analyses
Follow the applicable current requirements and document how they were addressed.
If one subgroup is too small for a reliable comparison
Report what the data can support, show uncertainty where appropriate, and avoid treating an imprecise estimate as evidence of no effect.
If an important subgroup is expected to be too small during study planning
Consider whether a recruitment strategy such as deliberate oversampling is appropriate for the intended analysis.
If a subgroup pattern appears unexpectedly after data collection
Treat it as exploratory when appropriate, report that status transparently, and avoid presenting it as though it were a prespecified confirmatory hypothesis.
If separate analyses produce different p-value classifications
Do not infer a subgroup difference from that alone; directly evaluate the contrast relevant to the research question.

There is also a design implication. If sex- or gender-specific inference is genuinely important, who is able to enter the study becomes part of the analytical problem. A sophisticated interaction model cannot recover information about a relevant population that was scarcely recruited in the first place.

07 · A Quick Checklist

Before Running a Sex- or Gender-Specific Analysis, Check the Design

Before disaggregating or comparing results, check:
Identify the scientific, clinical, social, policy, or reporting reason for examining sex or gender.
Verify what the sex or gender variable actually measured and how it was collected.
Distinguish descriptive reporting of group-specific results from a statistical test of whether groups differ.
Determine whether each group contains enough information for the intended analysis and precision.
Use an appropriate direct comparison or interaction analysis when the research question concerns whether an effect differs across groups.
Do not infer a subgroup difference merely because one p-value is significant and another is not.
Distinguish prespecified analyses from exploratory subgroup findings.
Consider multiplicity and the risk of chance findings when conducting numerous subgroup analyses.
Check the current requirements of applicable funders, regulators, protocols, reporting guidelines, and journals.
Interpret observed differences without claiming a biological or social mechanism that the study did not establish.
08 · Frequently Asked Questions

Questions About Analyzing Results by Sex or Gender

Should every study report results separately by sex?

Not every study requires the same analysis. The appropriate approach depends on the research question, study design, available data, and applicable requirements. In some biomedical research contexts, funders or regulators have explicit expectations concerning consideration or reporting of sex.

What if my sample is too small for sex-specific analysis?

Do not force a poorly supported comparison. Report relevant descriptive information and uncertainty where appropriate, acknowledge the limitation, and avoid interpreting an underpowered analysis as proof that no difference exists. If the comparison is central to the research question, future studies should plan recruitment and sample size accordingly.

If a result is significant for women but not men, can I say the effect exists only for women?

No, not from those two significance tests alone. The relevant question is whether the estimated effects differ between the groups. That requires a direct comparison appropriate to the study design and statistical model.

What is an interaction between treatment and sex?

In a suitable statistical model, a treatment-by-sex interaction evaluates whether the treatment association or effect differs according to sex on the scale represented by that model. The interpretation depends on the design, model, outcome, and effect measure.

Should sex- and gender-based analyses be planned before data collection?

When they address important hypotheses or study objectives, planning them in advance is preferable because recruitment, sample size, measurement, and analysis can then be designed appropriately. Post hoc analyses can still generate useful hypotheses but should be identified and interpreted as exploratory when appropriate.

Does NIH require analysis by sex?

NIH expects sex as a biological variable to be considered in the design, analysis, and reporting of applicable NIH-funded vertebrate animal and human studies. The exact expectations depend on the research and funding context, so investigators should consult the current NIH policy and application guidance directly.

Does a sex difference tell me why the groups differ?

No. Detecting a difference establishes a pattern under the study design; it does not by itself identify the biological, social, environmental, behavioral, or other mechanism responsible for that difference.

09 · The Bottom Line

Disaggregate When It Answers a Real Question, Not Merely Because You Can

The Bottom Line

Analyze results by sex or gender when doing so addresses a scientifically or practically meaningful question, meets an applicable requirement, and is supported by appropriate measurement, study design, and sufficient data.

When the goal is to determine whether results differ across groups, compare those results directly. Do not mistake separate significance tests for evidence of a subgroup difference, and do not interpret a detected difference as a biological or social mechanism unless the research actually supports that explanation.

10 · Sources and Further Reading

Authoritative Sources on Sex- and Gender-Specific Analysis

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes