03 · What You Need to Know
Uncertainty Is Part of the Finding, Not a Writing Defect
Separate what was observed from what you infer
A study may observe an association between two variables. Writing that one variable causes the other adds an inference the design may not support.
A qualitative study may show that participants commonly describe a particular experience. Writing that the experience is universal extends beyond the evidence.
A literature review should therefore preserve the distinction between the result and the interpretation placed on that result.
What the evidence shows
The result, pattern, estimate, observation, or experience supported by the studies.
What you infer from it
The broader explanation, causal claim, generalization, prediction, or practical implication you believe the evidence supports.
Certainty is not the same as consistency
Ten studies reporting similar findings may appear convincing, but consistency is only one consideration. If all ten share serious methodological limitations, examine an indirect population, or produce very imprecise estimates, the body of evidence may remain uncertain.
Formal GRADE assessment makes this distinction explicit. For a body of evidence concerning an outcome, certainty judgments consider domains including risk of bias, inconsistency, indirectness, imprecision, and publication bias.
The broader lesson is useful even when you are not conducting a formal GRADE assessment: ask what could make the apparent conclusion less trustworthy.
Risk of bias should change the conclusion, not just the appraisal table
Researchers sometimes perform critical appraisal carefully and then write the synthesis as though the appraisal never happened.
If studies supporting a conclusion have substantial risk of bias, that limitation should affect the certainty of the final interpretation. Cochrane guidance emphasizes incorporating risk-of-bias judgments into interpretation and certainty assessment rather than allowing compromised evidence to contribute to an apparently precise conclusion without qualification.
Imprecision means you may know less than the point estimate suggests
A study reports an estimated effect, but that estimate is uncertain. Confidence intervals help represent the range of values compatible with the data under the statistical model.
A point estimate may look beneficial while the confidence interval remains compatible with little effect or a meaningfully different effect. Cochrane guidance emphasizes interpreting estimates together with their confidence intervals rather than relying on a binary significant-versus-non-significant distinction.
When uncertainty is substantial, your prose should preserve it.
“No statistically significant effect” does not mean “no effect”
This is one of the easiest ways to make evidence sound more certain than it is.
A statistically non-significant result may arise because the estimated effect is small, because the estimate is imprecise, or both. It does not automatically demonstrate that the true effect is zero.
Cochrane explicitly warns against mistaking lack of evidence of an effect for evidence of a lack of effect and recommends focusing interpretation on effect estimates and confidence intervals rather than significance labels.
Watch Out
“There was no effect” is a strong conclusion. If the evidence is simply too imprecise to determine whether an important effect exists, say that the effect remains uncertain instead.
Statistical significance does not eliminate uncertainty either
The reverse error is equally common. A small p-value does not tell you whether the effect is important, unbiased, generalizable, or estimated precisely enough for the decision at hand.
Large studies can detect very small differences that have limited practical importance. A statistically significant result can also come from evidence at substantial risk of bias.
The certainty of a conclusion therefore cannot be read directly from a p-value.
Indirect evidence should produce narrower claims
Suppose every study in your review involves undergraduate students at highly resourced universities, but your conclusion concerns university students generally. The evidence may be internally credible while remaining indirect for the broader population.
Cochrane describes indirectness in terms of how well the available evidence matches the population, intervention, comparator, and outcome question to which the evidence is being applied.
The solution is not necessarily to discard the studies. It is to limit the scope of the conclusion and make the uncertainty about transferability visible.
Unexplained inconsistency should reduce confidence
If credible and sufficiently comparable studies produce substantially different findings, the synthesis should not simply average away the disagreement and proceed with a confident generalization.
Sometimes variation can be explained by populations, interventions, contexts, or methodological differences. When important inconsistency remains unexplained, confidence in a broad conclusion should usually decrease.
Cochrane's methodological standards require statistical heterogeneity to be considered when interpreting results, particularly when effects vary in direction because such variation affects the extent to which conclusions can be generalized.
Missing evidence can create false certainty
The published literature is not necessarily a complete representation of all research conducted. Studies or outcomes with disappointing or null findings may be less likely to appear in the accessible evidence base, and selective reporting can distort what seems to be known.
This is why publication bias and other forms of missing-results bias matter to certainty assessment. A literature that appears remarkably consistent may deserve additional scrutiny if the available evidence could systematically exclude unfavorable results.
Certainty can differ by outcome
Do not assign one certainty label to an entire topic when the evidence differs across outcomes.
An intervention might have relatively credible evidence for improving short-term performance, uncertain evidence concerning long-term retention, and sparse evidence concerning harms. The sentence “the intervention is well established” would blur those distinctions.
Cochrane's GRADE-based approach assesses certainty for individual outcomes rather than treating an intervention or review as possessing one universal level of certainty.
Generalization adds another layer of uncertainty
A result can be credible for the participants studied without being equally applicable everywhere.
Differences in culture, institutions, resources, implementation, adherence, or baseline conditions may affect whether a finding transfers to another setting. Cochrane's guidance on interpretation emphasizes the distinction between evidence within the studied context and judgments about applicability elsewhere.
Causal language requires causal evidence
Words such as “causes,” “leads to,” “results in,” and “improves” imply relationships that may exceed what observational evidence can establish.
When the evidence is associational, verbs such as “is associated with,” “correlates with,” or “is linked to” may be more accurate. When causal evidence is credible, stronger language can be justified.
Hedging should therefore be methodological rather than habitual. The verb should reflect the inference.
Use calibrated language rather than vague caution
Not every sentence needs to sound uncertain. When evidence is strong, clear language is appropriate. When certainty is lower, the language should signal that limitation.
Cochrane provides examples of wording aligned with certainty assessments. Moderate-certainty evidence may be expressed with terms such as “likely” or “probably,” while low-certainty evidence may use “may.” For very low-certainty evidence, the conclusion should make the substantial uncertainty explicit.
| Evidence situation |
Potential wording |
Wording to reconsider |
| Strong and directly relevant evidence |
“The evidence indicates...” or a direct effect statement where justified |
Unnecessary layers of “possibly” and “might” |
| Moderate uncertainty |
“The intervention likely...” or “probably...” where appropriate |
“Proves” or “definitively demonstrates” |
| Substantial uncertainty |
“The evidence suggests...” or “may...” |
Unqualified causal or universal claims |
| Very uncertain evidence |
“The evidence is very uncertain about...” |
A precise conclusion unsupported by the evidence |
| Observational association |
“X was associated with Y” |
“X caused Y” without adequate causal support |
Do not confuse evidence with recommendations
Even reasonably certain evidence about an effect does not automatically determine what people should do. Decisions may depend on benefits, harms, costs, feasibility, acceptability, equity, and values or preferences.
Cochrane explicitly distinguishes review conclusions from recommendations, noting that recommendations generally require judgments and evidence beyond the effect estimates contained in a systematic review.
Moving from “the intervention probably improves X” to “therefore all institutions should adopt it” introduces another inferential step that needs its own justification.
04 · A Practical Example
How Certainty Can Disappear Between the Evidence and the Sentence
Hypothetical Example
Does generative AI feedback improve student writing?
Imagine six hypothetical studies. Four report better writing scores among students receiving AI-generated feedback. Two report little difference. Most studies are small, follow students for only one assignment, use different writing rubrics, and compare different AI systems. Only one uses random assignment, and none examines whether improvements persist after students stop using the tool.
The overstated version
“Research demonstrates that generative AI feedback improves students' writing ability.”
The sentence compresses several uncertain steps. It treats short-term scores as writing ability, generalizes across different systems and measures, implies a causal conclusion stronger than most designs support, and says nothing about the inconsistency or short follow-up.
A calibrated synthesis
State the observed pattern Several hypothetical studies report higher writing scores when students receive AI-generated feedback.
Identify limitations The evidence is based largely on small, short-duration studies using heterogeneous systems and outcome measures.
Limit the inference The literature provides some evidence of short-term improvement in assessed writing performance but offers less certainty about durable writing development.
Preserve the unresolved question Whether improvements persist without AI assistance remains unestablished in this hypothetical evidence base.
The calibrated version is not weaker because it contains uncertainty. It is stronger because each claim corresponds more closely to what the evidence can actually support.