03 · What You Need to Know
The Discussion Interprets the Evidence; It Does Not Replace It
Start by separating results from interpretation
The distinction between the results and discussion is fundamental to critical reading. In a conventional empirical paper, the results section primarily reports what was observed or estimated, while the discussion explains what the authors think those findings mean, how they relate to previous research, and what implications may follow.
Guidance on reading scientific papers explicitly distinguishes these functions. Carey and colleagues characterize the discussion as the authors' perspective on the results, while StatPearls recommends independently interpreting and critically appraising the findings before reading the authors' discussion.
That distinction gives you a useful first question whenever you encounter a major statement in the discussion:
Is this statement a result, or is it an interpretation of a result?
Result
The intervention group had a higher mean score than the comparison group under the reported analysis.
Interpretation
The intervention improved learning because it increased students' engagement with the material.
The second statement contains considerably more than the first. "Improved" may imply a causal interpretation, "learning" must correspond to what was actually measured, and the proposed mechanism involving engagement may or may not have been tested.
Recognizing that movement from observation to explanation is the beginning of critical reading.
Try to interpret the important results before reading the discussion closely
One useful strategy is to examine the relevant results, figures, and tables and form a provisional interpretation before allowing the discussion to frame them for you.
You do not need to produce an elaborate independent theory. Ask simpler questions. What is the major pattern? How large is the effect or difference? How uncertain is it? Which findings appear robust, ambiguous, or unexpected? What does the design allow you to conclude?
Then read the discussion and compare.
Read the result What does the study actually report?
Form a provisional interpretation What seems justified from the evidence and design?
Read the discussion How do the authors explain the same result?
Compare Where does their interpretation match yours, add useful context, or go beyond what you think the evidence supports?
This is one reason there can be value in reading the results before the discussion when your goal is critical evaluation. It reduces the chance that the authors' narrative becomes your interpretation before you have inspected the underlying evidence.
Ask whether every important claim can be traced back to evidence
A strong discussion should remain anchored to the study's actual findings.
When you encounter an important interpretive statement, mentally trace it backward. Which result supports this? Which table, figure, qualitative theme, estimate, or analysis is relevant? Was the outcome actually measured? Was the proposed relationship tested?
If you cannot locate the evidential basis, several possibilities exist. The authors may be drawing on previous literature, proposing a hypothesis, offering a plausible explanation, or extending beyond what their own study directly demonstrates.
Those moves are not inherently inappropriate. Discussion sections are supposed to interpret. Problems arise when speculation, explanation, and empirical finding blur together so completely that readers cannot tell how strongly each statement is supported.
| Discussion statement |
Question to ask |
What to verify |
| "X improved Y" |
Does the design support a causal claim? |
Study design, comparison, temporality, confounding, and relevant result |
| "X is associated with Y" |
What was the magnitude and uncertainty? |
Effect estimate, confidence interval, model, and measurement |
| "The effect occurred because of Z" |
Was Z actually measured or tested as a mechanism? |
Relevant variables, analysis, and whether the explanation is empirical or speculative |
| "These findings apply to..." |
How far does the study population and context support that extension? |
Sample, setting, eligibility criteria, and contextual differences |
| "Our findings are consistent with previous research" |
Are the studies genuinely comparable? |
Differences in populations, designs, measures, outcomes, and analytical approaches |
Check whether the language is stronger than the study design
One of the most consequential interpretive problems occurs when the wording implies more than the design can establish.
Suppose an observational study finds that students who use a particular educational technology more frequently also report higher academic engagement. The discussion might reasonably describe an association. It would require stronger assumptions to conclude that using the technology caused greater engagement.
Why? Students who use the technology more frequently may differ in other ways. They might already be more motivated, have stronger digital skills, receive different instructional support, or differ on variables that were not measured adequately.
Statistical adjustment can address specified factors under particular assumptions, but it does not automatically transform observational evidence into experimental evidence.
Watch Out
Pay close attention when interpretive language shifts from "associated with," "related to," or "observed alongside" toward "caused," "led to," "resulted in," or "improved." Ask whether the study design and analysis genuinely support that stronger inference.
If you are uncertain about what the design permits, return to the methodological features that determine what kind of evidence was produced.
Check whether the interpretation matches the magnitude of the result
A statistically detectable finding can still be small. A discussion may accurately report that an effect exists while giving the reader an exaggerated impression of its substantive importance.
Return to the effect estimate.
How large was the difference, association, or change? What does that magnitude mean on the original scale? How uncertain is the estimate? Does the confidence interval include effects that would lead to substantially different practical interpretations?
Terms such as "substantial," "meaningful," "important," "strong," and "promising" deserve scrutiny when they are not accompanied by a clear substantive basis.
This is especially important when a discussion emphasizes statistical significance. A small p-value does not establish that an effect is large or practically important. If necessary, return to the effect estimate and uncertainty rather than relying on the significance label.
Look for findings that receive less attention than they deserve
The discussion is necessarily selective. Authors cannot interpret every number with equal emphasis. That selection can shape the story readers take away from the paper.
Suppose a study measures five outcomes. Two show statistically significant differences and three do not. If the discussion concentrates almost entirely on the two favorable outcomes, the overall pattern may sound more consistent than the results actually are.
Similarly, subgroup findings, secondary outcomes, sensitivity analyses, or contradictory observations may receive less narrative attention than the headline result.
Compare the discussion with the full set of results relevant to the research question. Ask not only whether the statements made are defensible, but whether important findings that complicate the interpretation have been adequately acknowledged.
Ask whether alternative explanations remain plausible
An observed pattern may have more than one explanation.
If students using a learning platform achieve higher grades, perhaps the platform helped. Perhaps more motivated students used it more frequently. Perhaps instructors who encouraged its use also provided other forms of support. Perhaps the measured association partly reflects prior achievement.
A good discussion may consider plausible alternatives, but you should not assume that every important one will be identified.
Critical reading therefore asks:
- What else could have produced this result?
- Does the study design rule out those alternatives, reduce their plausibility, or leave them largely unresolved?
- Did the authors measure or analyze variables relevant to the proposed explanation?
- Are they distinguishing a demonstrated mechanism from a plausible one?
Carey and colleagues explicitly recommend asking whether the data support the authors' interpretations and what alternative explanations exist. Critical appraisal similarly requires assessing whether the study design and methods support the conclusions being drawn.
Take the limitations section seriously, but do not assume it is complete
Authors often discuss limitations, and those disclosures are useful. They can reveal measurement problems, sampling constraints, potential biases, analytical uncertainty, implementation difficulties, or restrictions on generalizability.
But the authors' limitations section is not an exhaustive certificate of everything that could affect the study.
Carey and colleagues specifically note that discussions often describe some, but not necessarily all, strengths and limitations. Their recommendation is to consider whether you agree with the authors' self-assessment and whether you would add anything.
When you encounter a limitation, ask what it does to the conclusion. Does it merely add a minor caveat? Does it reduce precision? Does it limit generalizability? Could it offer an alternative explanation for the finding? Could it substantially weaken the inference?
A limitation that can change the meaning of the result should not disappear simply because it appears in a paragraph labeled "limitations."
Do not treat mentioning a limitation as solving it
There is an important difference between acknowledging a limitation and neutralizing its consequences.
Imagine that authors state that their convenience sample limits generalizability. That acknowledgment is appropriate. It does not suddenly make the sample representative.
Similarly, acknowledging that a study is cross-sectional does not permit causal conclusions that the design could not otherwise support. Noting that a measure is self-reported does not remove the measurement limitations associated with self-report.
Acknowledged limitation
The authors have correctly identified a potential weakness or boundary.
Resolved limitation
The study design, additional analysis, evidence, or another methodological feature has actually addressed the problem sufficiently for the relevant inference.
The first does not imply the second.
Check whether the discussion generalizes beyond the sample and setting
Researchers often want their findings to matter beyond the exact participants and conditions studied. That is reasonable. The question is how far the evidence can travel.
If a study involved students from one institution, one discipline, one country, or a highly selected sample, claims about "university students" generally may require caution. A laboratory result may not transfer directly to authentic practice. A short intervention may not establish long-term effects.
Generalizability is not determined by sample size alone. You need to consider who or what was studied, how they were selected, the context in which the research occurred, how the intervention or exposure was implemented, and how similar those conditions are to the population or setting to which the authors extend their claims.
Sometimes narrow evidence is entirely appropriate. The problem is not specificity; it is making the conclusion sound broader than the evidence warrants.
Check whether recommendations are one step beyond the evidence
Discussions often move from findings to recommendations: institutions should adopt a technology, clinicians should change practice, educators should implement an intervention, policymakers should revise a program.
Recommendations involve additional judgments beyond whether an effect was observed. Costs, feasibility, risks, alternatives, equity, implementation conditions, durability of effects, and the quality of the broader evidence may all matter.
A single positive study can contribute to a recommendation without being sufficient to establish it.
When authors move from "we observed X" to "therefore practitioners should do Y," ask what additional assumptions connect those statements.
Compare the interpretation with the wider literature, not just the citations chosen by the authors
The discussion commonly positions findings relative to previous studies. This helps readers understand whether the result confirms, extends, or challenges existing knowledge.
Remember that this account is necessarily selective. Authors choose which prior studies to discuss and how to characterize them.
If the paper is important to your work, follow consequential citations and examine the broader literature yourself. A claim that findings are "consistent with previous studies" may conceal important differences in population, measurement, design, or effect magnitude. A claim that the result is "novel" may depend on how the relevant literature was defined.
Reading cited sources is particularly valuable when a discussion uses previous literature to support an explanation that the current study did not directly test.
Consider conflicts of interest without using them as a shortcut to judgment
Funding sources and competing interests can be relevant to critical appraisal. Young and Solomon include potential conflicts of interest among the factors that readers should consider when assessing research.
A declared conflict does not prove that the findings or interpretation are wrong. Likewise, absence of a declared financial conflict does not guarantee freedom from every source of bias.
Use disclosure information as context. Then return to the evidence: Are the methods appropriate? Are all relevant outcomes reported? Does the interpretation remain proportionate? Are limitations acknowledged? Are alternative explanations considered?
The aim is evaluation, not guilt by association.
Your interpretation can differ from the authors' without making either side obviously wrong
Research findings often admit more than one reasonable interpretation, particularly when evidence is incomplete or several mechanisms could explain the same pattern.
You may judge an effect less practically important than the authors do. You may think a methodological limitation deserves more weight. You may see an alternative explanation as more plausible. Another knowledgeable reader may disagree with both of you.
That does not make interpretation arbitrary. Interpretations should still be constrained by the data, study design, relevant theory, and existing evidence.
The goal of critical reading is not to replace author authority with reader authority. It is to make the reasoning between evidence and conclusion visible enough to evaluate.