03 · What You Need to Know
Confidence Is a Judgment About the Claim, Not a Property of the Paper
Start With the Exact Conclusion You Want to Evaluate
Before asking how confident you should be, identify the claim.
“The intervention works” is usually too vague. Does it improve a particular outcome? By how much? Compared with what? Among which population? Over what period? Is the claim descriptive, predictive, causal, interpretive, or explanatory?
Evidence can support one version of a claim strongly and another only weakly.
A study might provide convincing evidence that two variables are associated while providing little basis for concluding that one causes the other. Confidence therefore cannot be assigned sensibly until the intended inference is clear.
Study Design Determines What Kinds of Inferences Are Available
Different designs provide different forms of evidential leverage.
A well-designed randomized experiment may provide strong evidence for certain causal comparisons. A representative survey may support population estimates. Longitudinal data can provide information about temporal patterns. Interviews may provide rich evidence about participants' experiences and interpretations. Historical research may draw on documentary and contextual evidence to explain past processes.
The question is not whether one design is universally superior. It is whether the design is appropriate for the conclusion.
This is why valid evidence must be evaluated in relation to the research question and claim.
Risk of Bias Affects Confidence
A result can be precise and still be systematically misleading.
Researchers therefore examine whether features of sampling, assignment, measurement, missing data, analysis, reporting, or other procedures could shift the findings away from the quantity or phenomenon of interest.
The relevant biases depend on the research design. Randomized trials, observational studies, diagnostic research, qualitative investigations, surveys, and other approaches face different methodological threats.
Confidence should decrease when a consequential source of bias remains plausible and insufficiently addressed.
Measurement Quality Matters
Researchers cannot be highly confident in a conclusion about a construct if the evidence does not represent that construct adequately.
A reliable-looking numerical score does not solve a mismatch between the measurement and the claim. If a study claims to improve “learning” but measures only immediate satisfaction, the problem is conceptual before it is statistical.
Researchers should therefore ask whether instruments, operational definitions, coding procedures, observations, or other measures provide defensible representations of the phenomena named in the conclusion.
Precision Affects Confidence in Quantitative Estimates
A point estimate provides one value, but researchers also need to know how uncertain that estimate is.
Cochrane guidance emphasizes interpreting estimates alongside confidence intervals. A relatively narrow interval can indicate that an effect is estimated with greater precision, while a very wide interval may leave substantially different substantive possibilities compatible with the data.
Suppose an estimated intervention effect is positive, but the interval is compatible with a substantial benefit, little meaningful effect, and possibly harm. The direction of the point estimate alone should not generate high confidence about what the intervention actually does.
Imprecision is therefore one important source of uncertainty, not the only one.
Statistical Significance Does Not Determine Confidence
A p-value is not a general-purpose confidence meter.
A statistically significant result can arise from a study with serious bias, poor measurement, an unimportant effect, or limited generalizability. Conversely, a study that does not cross a conventional significance threshold may still provide useful information about the plausible magnitude and direction of an effect.
Researchers should examine estimates, uncertainty, design, assumptions, and substantive relevance rather than translating “significant” into “trustworthy” and “non-significant” into “no evidence.”
Watch Out
Statistical significance and confidence in a scientific conclusion answer different questions. Crossing a statistical threshold does not certify the validity, importance, causality, or generalizability of the result.
Consistency Across Studies Can Increase Confidence
A conclusion becomes less dependent on one particular dataset when other relevant studies obtain compatible findings.
Repeated evidence can reduce concern that the original result arose from an unusual sample or study-specific circumstance. Studies conducted by independent teams or using different methods can provide particularly useful information about robustness.
But consistency should not be judged merely by counting how many papers reach the same headline conclusion.
Researchers should examine whether effect estimates are genuinely compatible, whether studies are sufficiently comparable, and whether they share limitations that could produce the same misleading result.
This is part of how replication and repeated evidence can strengthen confidence.
Inconsistency Can Lower Confidence or Reveal Heterogeneity
When studies produce substantially different results, researchers need to understand why.
Unexplained inconsistency can reduce confidence in a simple conclusion. But variation may also reveal genuine differences among populations, contexts, interventions, outcomes, or implementations.
The appropriate response is therefore not always to average everything together and move on. Researchers may need to refine the claim.
A conclusion such as “the intervention is effective” may become “the intervention appears effective under these conditions but not those.”
Directness and Relevance Matter
Evidence can be methodologically strong but only indirectly relevant to the question at hand.
Suppose high-quality studies evaluate an intervention among adults, but your conclusion concerns children. Or studies measure a surrogate outcome while the claim concerns a long-term outcome that matters directly to participants.
The inferential distance between the evidence and the intended conclusion introduces uncertainty.
In GRADE, this is addressed through the domain of indirectness. Other disciplines may describe the problem differently, but the principle is widely applicable: confidence should reflect how directly the evidence answers the actual question.
Publication and Availability of Evidence Can Affect Confidence
The visible literature may not represent all the research that has been conducted.
If studies with striking or positive findings are more likely to become available than studies with null, negative, or inconclusive results, the published evidence can create an exaggerated impression of consistency or effect magnitude.
Publication bias is therefore considered explicitly in frameworks such as GRADE.
Researchers should be especially cautious when a conclusion depends on a small literature dominated by unusually striking findings or when there are reasons to suspect selective availability of results.
Plausible Alternative Explanations Matter
A conclusion becomes stronger when evidence distinguishes it from credible alternatives.
Suppose students using a particular platform achieve higher grades. If students choose whether to use the platform, motivation, prior achievement, available study time, or instructor behavior may explain part of the association.
If the design cannot address these alternatives, confidence in a causal conclusion should remain lower even if the association itself is estimated precisely.
This connects confidence to how researchers move from observation to explanation. A persuasive explanation must do more than fit the observed pattern.
Confidence in One Study Is Different From Certainty in a Body of Evidence
A study can be rigorous and informative without settling the larger question.
Once multiple studies exist, researchers need to evaluate the evidence collectively. Systematic reviews can identify relevant studies using explicit procedures, while evidence-synthesis frameworks can structure judgments about the resulting body of evidence.
GRADE, for example, uses four levels of certainty in a body of evidence: high, moderate, low, and very low. Its assessment considers domains including risk of bias, inconsistency, indirectness, imprecision, and publication bias.
These categories belong to a particular evidence-assessment framework and should not be applied mechanically to every form of research. Qualitative evidence synthesis, for example, may use GRADE-CERQual, which evaluates methodological limitations, relevance, coherence, and adequacy when assessing confidence in synthesized qualitative findings.
The broader lesson is that confidence judgments should use criteria suited to the kind of evidence being evaluated.
High Confidence Does Not Mean Absolute Certainty
A conclusion can deserve high confidence without being logically impossible to revise.
The National Academies emphasizes that science aims for refined degrees of confidence rather than complete certainty. Claims become more credible when they survive repeated testing and scrutiny.
High confidence therefore means that the available evidence provides strong reasons to accept a conclusion within its stated scope, not that no conceivable future evidence could modify it.
This is why uncertainty can remain within reliable scientific knowledge.
Low Confidence Does Not Mean the Claim Is False
The opposite distinction is equally important.
Low confidence means that the evidence does not currently justify a strong conclusion. Perhaps studies are small, biased, inconsistent, indirect, or sparse. The claim might eventually turn out to be correct, but the present evidence cannot establish that with confidence.
Researchers should distinguish “we have evidence that this is false” from “the evidence is currently insufficient for us to know confidently whether it is true.”
Confidence Can Differ Across Outcomes in the Same Research Literature
A body of evidence may support one conclusion strongly and another weakly.
For example, researchers might have high confidence that an intervention improves a short-term outcome but low confidence about long-term effects because few studies followed participants for long enough.
Cochrane emphasizes assessing certainty for specific outcomes rather than assigning one blanket quality label to an entire review.
The same logic is useful more broadly. Avoid describing a field simply as having “strong evidence” without specifying strong evidence for what.
Confidence Should Be Updated When the Evidence Changes
Scientific confidence is not fixed permanently.
New studies may reduce imprecision, reveal bias, resolve inconsistency, improve measurement, demonstrate generalizability, or identify previously unknown boundary conditions.
Confidence can therefore increase or decrease as evidence accumulates.
This is part of why scientific knowledge can change when new evidence appears. Updating confidence is not indecision; it is the expected consequence of making conclusions answerable to evidence.