03 · What You Need to Know
Translate the Statistics Back Into the Research Question
Start with the question before looking at the numbers
Statistics are easier to interpret when you know what problem they are being used to solve.
Before decoding a test statistic, return briefly to the research question and methods. What variables were measured? Which groups, conditions, or time points are being compared? Is the analysis examining a difference, association, prediction, change over time, or something else?
For example, "independent-samples t-test" is a statistical label. "Comparing the mean writing scores of students who received AI-assisted feedback with those who received conventional feedback" tells you what the analysis is doing in the study.
That translation should come first.
Statistical language
A regression coefficient was statistically significant.
Research language
Higher values of one variable were associated with higher or lower values of another, by an estimated amount, under the specified model.
If you cannot state in ordinary language what is being estimated, compared, or tested, interpreting the rest of the output becomes much harder.
Identify the descriptive statistics before the inferential statistics
Many results sections begin with descriptive statistics. These summarize what was observed in the sample before inferential procedures are used to make broader statistical judgments.
| Common notation |
Usually means |
What it helps you understand |
|
n
|
Number of observations or participants |
How much data contributed to a result |
|
M or mean |
Arithmetic average |
The average observed value |
| Median |
Middle value after ordering observations |
The center of a distribution without relying on the arithmetic mean |
|
SD
|
Standard deviation |
How dispersed individual observations are around the mean |
| % |
Percentage |
The proportion of observations with a particular characteristic or outcome |
These quantities may already tell you something important. If one group has a mean score of 84 and another has a mean of 83.8, you know the observed difference is small before seeing any hypothesis test.
Descriptive statistics also provide context for later inferential results. A standardized effect size of 0.50, for example, is easier to understand when you also know what the groups actually scored on the original measurement scale.
Find the effect or association the researchers estimated
Once you know what was compared, look for the estimate that describes what happened.
Depending on the analysis, this could be a mean difference, correlation coefficient, regression coefficient, risk difference, risk ratio, odds ratio, hazard ratio, standardized mean difference, or another measure.
Do not worry initially about memorizing every possible statistic. Ask two simpler questions:
- What quantity is being estimated?
- What would larger, smaller, positive, negative, or null values mean?
Suppose an intervention group scores an average of 4 points higher than a comparison group. The estimated difference is 4 points. That is the substantive quantity you want to understand. A subsequent statistical test provides additional information about that estimate; it does not replace it.
This emphasis on estimation is important because a p-value alone cannot tell you the magnitude of an effect. Statistical guidance has long recommended paying attention to estimated effects and their uncertainty rather than reducing interpretation to hypothesis testing alone.
Ask whether the effect is large enough to matter
Statistical significance and substantive importance are different questions.
With a sufficiently large sample, a very small difference may produce a small p-value. Conversely, a potentially meaningful effect estimated from a small or noisy sample may fail to cross a conventional significance threshold.
So after identifying the estimate, ask whether its magnitude matters in the context of the research.
A two-point increase could be trivial on one outcome and important on another. An odds ratio of 1.20 might have very different implications depending on the baseline risk and outcome. A standardized effect size may help comparison, but it does not determine practical importance automatically.
Watch Out
Do not translate "statistically significant" into "large," "important," "meaningful," or "useful." Statistical significance does not by itself establish the substantive importance of an effect.
Use confidence intervals to examine uncertainty
A point estimate gives you one estimated value. A confidence interval provides additional information about the precision or uncertainty surrounding that estimate under the statistical procedure used.
For example, suppose a study estimates that an intervention increases a test score by 5 points:
Mean difference = 5 points, 95% CI [4, 6]
Compare that with:
Mean difference = 5 points, 95% CI [-2, 12]
The point estimate is identical, but the second interval reflects much greater uncertainty. It includes values corresponding to a modest negative difference as well as a substantially larger positive difference.
Confidence intervals therefore help you ask questions that a point estimate alone cannot answer: How precise is this estimate? What effect sizes remain reasonably compatible with the data and statistical model? Does the interval include effects that would lead to materially different interpretations?
For many conventional comparisons, the confidence interval can also indicate whether the corresponding null value is excluded. For differences, that null value is commonly zero. For ratio measures such as risk ratios and odds ratios, it is commonly one.
Do not turn that observation into another mechanical significance test, however. The interval's value lies partly in showing the range and precision of plausible effect estimates, not merely whether it crosses a particular threshold.
Do not interpret a 95% confidence interval as a 95% probability statement about the parameter
A conventional frequentist confidence interval is often described casually as though there were a 95% probability that the true parameter lies inside the particular interval you see. That is not the standard frequentist interpretation.
The 95% refers to the long-run performance of the interval-producing procedure under its assumptions: across repeated samples generated under the same conditions, approximately 95% of intervals constructed by that procedure would contain the true parameter.
For practical reading, you do not need to rehearse that definition every time you encounter an interval. You should avoid converting the interval into a probability claim that the method itself does not provide.
Understand what a p-value does and does not tell you
A p-value is calculated under a specified statistical model that includes a null hypothesis. In broad terms, it describes how incompatible the observed data, or more extreme data according to the test statistic, are with that model.
It does not tell you the probability that the null hypothesis is true. A value of p =.03 does not mean there is a 3% probability that the null hypothesis is correct.
It also does not tell you how large or important the effect is. Two studies can produce similar p-values for effects of different magnitudes, and similar effect estimates can produce different p-values when their precision differs.
This is why reading only the p-value is inadequate.
| A p-value can contribute information about... |
A p-value alone does not tell you... |
| Compatibility of the observed data with a specified null model |
The probability that the hypothesis is true |
| Evidence assessed through a particular statistical test |
How large the effect is |
| Whether a conventional significance threshold is crossed |
Whether the result is practically important |
| A component of the reported statistical analysis |
Whether the study design, measurement, or analysis is otherwise valid |
Do not let p =.049 and p =.051 become different scientific worlds
The convention of using p <.05 as a threshold can encourage a sharp distinction between results just below and just above.05.
That distinction is often much stronger than the underlying evidence warrants. Results with p =.049 and p =.051 are statistically very similar, even though one may be labeled "significant" and the other "not significant."
Read the effect estimate, confidence interval, study design, sample size, and broader evidence rather than allowing a threshold to make the interpretation for you.
A non-significant result does not prove that there is no effect
This is one of the most consequential statistical reading errors.
Suppose a study reports p =.18. That result does not automatically demonstrate that two groups are equivalent or that an intervention has no effect. The study may simply provide insufficiently precise evidence to distinguish among several possibilities.
The confidence interval can be particularly informative here. A wide interval may contain both effects large enough to matter and the null value. In that situation, "no effect" would be much stronger than the data justify.
Classic statistical guidance summarizes the problem succinctly through the distinction between absence of evidence and evidence of absence. Demonstrating equivalence or sufficiently excluding an important effect generally requires methods and evidence appropriate to that question, not merely a conventional non-significant test.
Check the sample size, but do not use it as a quality score
Sample size affects statistical precision and often affects the ability of an analysis to detect or estimate effects. Larger samples can produce narrower confidence intervals when other conditions are comparable.
But "large sample" does not mean "good study."
A large dataset cannot automatically repair poor measurement, systematic selection problems, confounding, an inappropriate comparison, a badly specified model, or a design incapable of supporting the claimed inference.
Similarly, a small sample is not automatically worthless. It may produce imprecise estimates that should be interpreted cautiously, but the appropriate judgment depends on the design, research question, expected effect, variability, and analysis.
Sample size is part of the interpretation, not a substitute for methodological evaluation.
Read the test statistic only after you know what was tested
Results sections may report statistics such as t, F, χ², z, r, regression coefficients, or model-specific quantities. These values have technical meanings within their respective procedures.
You do not necessarily need to calculate them yourself to understand the paper. You do need to connect them to the research question.
For example:
t(118) = 2.47, p =.015
Knowing that this is a t-test result is useful, but your interpretation still depends on what was compared. Which groups? Which outcome? What were their means? What was the estimated difference? How uncertain was it?
The statistical symbol should never become detached from the substantive comparison it represents.
Regression results require you to identify the outcome, predictor, and adjustment
Regression tables often intimidate non-statisticians because they can contain many rows of coefficients, standard errors, confidence intervals, test statistics, and p-values.
Start by finding the outcome variable. What is the model trying to explain or predict?
Then locate the predictor you care about. What does its coefficient represent? In a simple linear regression, for example, a coefficient can represent the expected change in the outcome associated with a one-unit increase in a predictor. In multiple regression, that interpretation is conditional on the other variables included in the model.
Then ask what was adjusted for. An unadjusted association and an association estimated after including several covariates are not necessarily the same quantity.
For logistic, survival, multilevel, generalized, or other models, coefficients may be reported on scales that require more specific interpretation. If that scale is central to your use of the finding, this is a point where additional statistical explanation may be necessary.
Correlation is about association, not automatically causation
A correlation coefficient describes the direction and strength of an association under the particular correlation measure used. A positive coefficient indicates that higher values of one variable tend to occur with higher values of the other; a negative coefficient indicates an inverse relationship.
That does not establish that one variable causes the other.
The same caution applies more broadly to regression coefficients and other associations from observational data. Statistical adjustment can address specified variables under particular assumptions, but the appearance of an adjusted coefficient does not automatically transform an observational relationship into a causal effect.
When causal language matters, return to the study design and methods that generated the statistical result.
Look for effect sizes, but interpret them in context
Effect sizes quantify the magnitude of a difference or relationship. Depending on the study, the natural effect estimate may already be directly interpretable, such as a mean difference, risk difference, or risk ratio. Standardized measures such as Cohen's d can help express effects relative to variability and may facilitate certain comparisons.
Do not treat conventional labels such as "small," "medium," and "large" as universal laws. Whether an effect matters depends on the outcome, research context, measurement scale, costs, risks, baseline conditions, and substantive consequences.
A statistically small effect may matter when applied to a large population or consequential outcome. A numerically larger effect may matter little if it concerns an outcome with limited practical significance.
Be careful when a paper reports many statistical tests
A long results section may contain dozens or hundreds of comparisons. The more opportunities an analysis creates to find apparently interesting results, the more important it becomes to understand how those analyses were planned and interpreted.
Ask whether the reported analysis was central to the original research question, whether multiple outcomes or subgroup analyses were examined, whether relevant adjustments were made when appropriate, and whether the paper distinguishes prespecified analyses from exploratory ones.
This is one of the places where the results section cannot be interpreted independently of the methods.
Subgroup claims require direct evidence about differences between subgroups
A common mistake is to see that a result is statistically significant in one subgroup but not another and conclude that the effect differs between the groups.
That conclusion does not follow automatically.
Different p-values can arise because subgroup estimates have different precision, even when the estimated effects themselves are similar. The appropriate question is whether the effects differ from each other, which generally requires a direct comparison or interaction analysis rather than comparing two separate significance labels.
This is an excellent example of why statistical interpretation should focus on estimates and the comparison actually being tested rather than on isolated p-values.
Do not interpret statistics independently of the methods
Even perfectly calculated statistics cannot rescue a study whose design does not support the intended conclusion.
If a paper reports a strong association, you may need to know how the variables were measured. If an intervention produces a large estimated effect, you may need to know how participants entered the groups. If an adjusted regression coefficient is central to the conclusion, you may need to understand what variables were included and why.
Statistical results are downstream of research design, measurement, data collection, and analytical decisions.
If you reach a point where you cannot interpret a central statistic without understanding the analysis that generated it, work through that methodological uncertainty rather than treating the number as self-validating.