Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Avoid Making the Literature Look More Certain Than It Really Is?

A literature review should distinguish what the evidence suggests from what it firmly establishes. Certainty depends on more than how many studies report similar findings.

216
Avoiding Overstated Certainty Guide 216 of 247
01 · The Question

How Certain Should Your Literature Review Sound?

Academic writing often rewards clarity. That can create an unfortunate temptation: remove the qualifications, choose the strongest verb, and make the literature sound as though it has reached a settled conclusion.

“Studies demonstrate that...” sounds cleaner than “available evidence suggests that...” Yet the stronger sentence may quietly discard uncertainty arising from study limitations, imprecise estimates, inconsistent findings, indirect evidence, or gaps in the literature.

The opposite problem is possible too. Filling every sentence with “may,” “might,” and “possibly” does not produce methodological sophistication. The goal is calibrated certainty: language that is as confident as the evidence permits and no more.

02 · The Short Answer

Match the Strength of the Claim to the Strength of the Evidence

In Brief

Avoid overstating the literature by distinguishing findings from interpretations, considering risk of bias, consistency, directness, precision, and possible missing evidence, and using language whose level of certainty matches what the body of evidence can reasonably support.

Do not infer certainty from publication count, peer review, statistical significance, or apparent consensus alone. In formal evidence synthesis, frameworks such as GRADE provide explicit methods for assessing certainty across outcomes.

03 · What You Need to Know

Uncertainty Is Part of the Finding, Not a Writing Defect

Separate what was observed from what you infer

A study may observe an association between two variables. Writing that one variable causes the other adds an inference the design may not support.

A qualitative study may show that participants commonly describe a particular experience. Writing that the experience is universal extends beyond the evidence.

A literature review should therefore preserve the distinction between the result and the interpretation placed on that result.

What the evidence shows The result, pattern, estimate, observation, or experience supported by the studies.
What you infer from it The broader explanation, causal claim, generalization, prediction, or practical implication you believe the evidence supports.

Certainty is not the same as consistency

Ten studies reporting similar findings may appear convincing, but consistency is only one consideration. If all ten share serious methodological limitations, examine an indirect population, or produce very imprecise estimates, the body of evidence may remain uncertain.

Formal GRADE assessment makes this distinction explicit. For a body of evidence concerning an outcome, certainty judgments consider domains including risk of bias, inconsistency, indirectness, imprecision, and publication bias.

The broader lesson is useful even when you are not conducting a formal GRADE assessment: ask what could make the apparent conclusion less trustworthy.

Risk of bias should change the conclusion, not just the appraisal table

Researchers sometimes perform critical appraisal carefully and then write the synthesis as though the appraisal never happened.

If studies supporting a conclusion have substantial risk of bias, that limitation should affect the certainty of the final interpretation. Cochrane guidance emphasizes incorporating risk-of-bias judgments into interpretation and certainty assessment rather than allowing compromised evidence to contribute to an apparently precise conclusion without qualification.

Imprecision means you may know less than the point estimate suggests

A study reports an estimated effect, but that estimate is uncertain. Confidence intervals help represent the range of values compatible with the data under the statistical model.

A point estimate may look beneficial while the confidence interval remains compatible with little effect or a meaningfully different effect. Cochrane guidance emphasizes interpreting estimates together with their confidence intervals rather than relying on a binary significant-versus-non-significant distinction.

When uncertainty is substantial, your prose should preserve it.

“No statistically significant effect” does not mean “no effect”

This is one of the easiest ways to make evidence sound more certain than it is.

A statistically non-significant result may arise because the estimated effect is small, because the estimate is imprecise, or both. It does not automatically demonstrate that the true effect is zero.

Cochrane explicitly warns against mistaking lack of evidence of an effect for evidence of a lack of effect and recommends focusing interpretation on effect estimates and confidence intervals rather than significance labels.

Watch Out

“There was no effect” is a strong conclusion. If the evidence is simply too imprecise to determine whether an important effect exists, say that the effect remains uncertain instead.

Statistical significance does not eliminate uncertainty either

The reverse error is equally common. A small p-value does not tell you whether the effect is important, unbiased, generalizable, or estimated precisely enough for the decision at hand.

Large studies can detect very small differences that have limited practical importance. A statistically significant result can also come from evidence at substantial risk of bias.

The certainty of a conclusion therefore cannot be read directly from a p-value.

Indirect evidence should produce narrower claims

Suppose every study in your review involves undergraduate students at highly resourced universities, but your conclusion concerns university students generally. The evidence may be internally credible while remaining indirect for the broader population.

Cochrane describes indirectness in terms of how well the available evidence matches the population, intervention, comparator, and outcome question to which the evidence is being applied.

The solution is not necessarily to discard the studies. It is to limit the scope of the conclusion and make the uncertainty about transferability visible.

Unexplained inconsistency should reduce confidence

If credible and sufficiently comparable studies produce substantially different findings, the synthesis should not simply average away the disagreement and proceed with a confident generalization.

Sometimes variation can be explained by populations, interventions, contexts, or methodological differences. When important inconsistency remains unexplained, confidence in a broad conclusion should usually decrease.

Cochrane's methodological standards require statistical heterogeneity to be considered when interpreting results, particularly when effects vary in direction because such variation affects the extent to which conclusions can be generalized.

Missing evidence can create false certainty

The published literature is not necessarily a complete representation of all research conducted. Studies or outcomes with disappointing or null findings may be less likely to appear in the accessible evidence base, and selective reporting can distort what seems to be known.

This is why publication bias and other forms of missing-results bias matter to certainty assessment. A literature that appears remarkably consistent may deserve additional scrutiny if the available evidence could systematically exclude unfavorable results.

Certainty can differ by outcome

Do not assign one certainty label to an entire topic when the evidence differs across outcomes.

An intervention might have relatively credible evidence for improving short-term performance, uncertain evidence concerning long-term retention, and sparse evidence concerning harms. The sentence “the intervention is well established” would blur those distinctions.

Cochrane's GRADE-based approach assesses certainty for individual outcomes rather than treating an intervention or review as possessing one universal level of certainty.

Generalization adds another layer of uncertainty

A result can be credible for the participants studied without being equally applicable everywhere.

Differences in culture, institutions, resources, implementation, adherence, or baseline conditions may affect whether a finding transfers to another setting. Cochrane's guidance on interpretation emphasizes the distinction between evidence within the studied context and judgments about applicability elsewhere.

Causal language requires causal evidence

Words such as “causes,” “leads to,” “results in,” and “improves” imply relationships that may exceed what observational evidence can establish.

When the evidence is associational, verbs such as “is associated with,” “correlates with,” or “is linked to” may be more accurate. When causal evidence is credible, stronger language can be justified.

Hedging should therefore be methodological rather than habitual. The verb should reflect the inference.

Use calibrated language rather than vague caution

Not every sentence needs to sound uncertain. When evidence is strong, clear language is appropriate. When certainty is lower, the language should signal that limitation.

Cochrane provides examples of wording aligned with certainty assessments. Moderate-certainty evidence may be expressed with terms such as “likely” or “probably,” while low-certainty evidence may use “may.” For very low-certainty evidence, the conclusion should make the substantial uncertainty explicit.

Evidence situation Potential wording Wording to reconsider
Strong and directly relevant evidence “The evidence indicates...” or a direct effect statement where justified Unnecessary layers of “possibly” and “might”
Moderate uncertainty “The intervention likely...” or “probably...” where appropriate “Proves” or “definitively demonstrates”
Substantial uncertainty “The evidence suggests...” or “may...” Unqualified causal or universal claims
Very uncertain evidence “The evidence is very uncertain about...” A precise conclusion unsupported by the evidence
Observational association “X was associated with Y” “X caused Y” without adequate causal support

Do not confuse evidence with recommendations

Even reasonably certain evidence about an effect does not automatically determine what people should do. Decisions may depend on benefits, harms, costs, feasibility, acceptability, equity, and values or preferences.

Cochrane explicitly distinguishes review conclusions from recommendations, noting that recommendations generally require judgments and evidence beyond the effect estimates contained in a systematic review.

Moving from “the intervention probably improves X” to “therefore all institutions should adopt it” introduces another inferential step that needs its own justification.

04 · A Practical Example

How Certainty Can Disappear Between the Evidence and the Sentence

Hypothetical Example

Does generative AI feedback improve student writing?

Imagine six hypothetical studies. Four report better writing scores among students receiving AI-generated feedback. Two report little difference. Most studies are small, follow students for only one assignment, use different writing rubrics, and compare different AI systems. Only one uses random assignment, and none examines whether improvements persist after students stop using the tool.

The overstated version

“Research demonstrates that generative AI feedback improves students' writing ability.”

The sentence compresses several uncertain steps. It treats short-term scores as writing ability, generalizes across different systems and measures, implies a causal conclusion stronger than most designs support, and says nothing about the inconsistency or short follow-up.

A calibrated synthesis

State the observed pattern Several hypothetical studies report higher writing scores when students receive AI-generated feedback.
Identify limitations The evidence is based largely on small, short-duration studies using heterogeneous systems and outcome measures.
Limit the inference The literature provides some evidence of short-term improvement in assessed writing performance but offers less certainty about durable writing development.
Preserve the unresolved question Whether improvements persist without AI assistance remains unestablished in this hypothetical evidence base.

The calibrated version is not weaker because it contains uncertainty. It is stronger because each claim corresponds more closely to what the evidence can actually support.

05 · What Researchers Often Get Wrong

Common Ways Literature Reviews Overstate Certainty

Misconception

“Many studies means high certainty”

A large evidence base can still contain substantial shared bias, indirectness, inconsistency, imprecision, or missing evidence. Quantity is relevant, but it does not substitute for appraisal.

Misconception

“Peer-reviewed evidence can be stated confidently”

Peer review is not a certainty assessment. Published studies vary in design, execution, precision, relevance, and risk of bias.

Misconception

“Statistically significant means established”

A statistically significant result may still be small, biased, indirect, or practically unimportant. Statistical significance does not determine certainty by itself.

Misconception

“Non-significant means there is no effect”

A non-significant estimate may simply be too imprecise to distinguish an important effect from little or no effect. Examine the estimate and confidence interval rather than turning a threshold into a substantive conclusion.

Misconception

“Using cautious words automatically prevents overclaiming”

Adding “may” to an unsupported causal claim does not repair the underlying inference. Certainty language needs to accompany accurate claims about design, population, outcome, and evidence strength.

Misconception

“A confident conclusion makes a literature review stronger”

A conclusion is strong when it is defensible, not when it is categorical. Explicitly identifying what remains uncertain may be one of the most important analytical contributions of a review.

06 · What This Means for You

Audit Every Strong Claim for the Uncertainty It Leaves Out

When revising your synthesis, highlight sentences containing words such as “demonstrates,” “proves,” “causes,” “consistently,” “clearly,” “effective,” “no effect,” and “established.” You do not need to remove them automatically. Ask what evidence licenses each one.

Then inspect the body of evidence supporting the claim rather than the sentence in isolation.

A simple decision framework

If evidence is consistent, precise, directly relevant, and at low risk of important bias
Use appropriately clear language without unnecessary hedging.
If estimates are imprecise
Make the range of plausible effects and resulting uncertainty visible.
If credible studies disagree
Reduce confidence in a broad conclusion unless the inconsistency can be explained defensibly.
If evidence is indirect for the population, intervention, context, or outcome of interest
Narrow the scope of the claim and identify the limitation in applicability.
If the design supports association but not a strong causal inference
Use associational language and keep causal explanations appropriately tentative.
If several major sources of uncertainty remain
Say explicitly that the evidence remains uncertain rather than hiding uncertainty behind polished prose.

Formal systematic reviews should follow their specified certainty-assessment framework. Cochrane reviews, for example, use GRADE to assess certainty for important outcomes and align conclusions with those judgments.

07 · A Quick Checklist

Before Making a Confident Claim About the Literature, Check:

For every important conclusion, check:
Does the study design support the type of inference I am making?
Have important risks of bias been reflected in the interpretation rather than merely reported elsewhere?
Are the relevant findings sufficiently consistent, or does unexplained variation reduce confidence?
Are estimates sufficiently precise for the conclusion I am drawing?
Is the evidence directly applicable to the population, context, intervention, and outcome named in my claim?
Could missing studies or selectively reported results make the evidence appear more certain than it is?
Am I distinguishing lack of evidence from evidence that an effect is absent?
Does the certainty of my wording match the certainty of the underlying evidence?
08 · Frequently Asked Questions

Frequently Asked Questions About Certainty in Literature Reviews

Should academic writing always use cautious language?

No. Excessive hedging can make strong evidence sound unnecessarily uncertain. The objective is calibration: use direct language when evidence supports it and qualified language when meaningful uncertainty remains.

When can I say research “shows” something?

There is no universal prohibited-word list. Ask whether the evidence supporting the statement is sufficiently credible, direct, consistent, and precise for the claim. More specific statements are usually easier to justify than broad declarations about what “research shows.”

Should I avoid the word “proves”?

In empirical research synthesis, “proves” is usually stronger than the evidence warrants because findings remain conditional on study designs, assumptions, measurements, populations, and uncertainty. More precise language normally communicates the evidential status better.

Does consistent evidence mean high certainty?

Not necessarily. Consistency is only one consideration. A consistent body of evidence may still have important risk of bias, indirectness, imprecision, or publication bias.

What is the difference between imprecision and inconsistency?

Imprecision concerns uncertainty around an estimate, often reflected in a wide confidence interval. Inconsistency concerns meaningful variation in findings across studies. Both can reduce confidence, but they represent different problems.

What does GRADE assess?

GRADE assesses certainty in a body of evidence for a particular outcome. In Cochrane's application, key domains include risk of bias, inconsistency, indirectness, imprecision, and publication bias, producing certainty judgments such as high, moderate, low, or very low.

Can certainty differ for different outcomes in the same review?

Yes. Evidence may be more certain for one outcome than another because the relevant studies, precision, consistency, directness, or risk of bias differ. Avoid giving an entire intervention or topic one blanket certainty label.

Is “more research is needed” enough when evidence is uncertain?

Usually not. Identify why uncertainty remains. Future research needs differ depending on whether the problem is risk of bias, inadequate sample size, inconsistent findings, indirect populations, missing outcomes, or another limitation.

09 · The Bottom Line

Make the Certainty of the Sentence Match the Certainty of the Evidence

The Bottom Line

Avoid making the literature look more certain than it is by allowing risk of bias, inconsistency, indirectness, imprecision, missing evidence, and the limits of study design to constrain both the scope and wording of your conclusions.

Uncertainty is not something to edit out of a literature review. When the evidence is strong, say so clearly. When it is limited, conflicting, indirect, or imprecise, make that uncertainty equally clear. The credibility of the synthesis depends on the match between what the evidence can support and what your sentences claim.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes