03 · What You Need to Know
Evidence Establishes Claims at Different Levels
Begin with a proposition, not a general impression
“The literature supports online learning” is difficult to evaluate because almost every important term is underspecified. Supports it for what? Compared with what? For which learners? Under which conditions? On which outcomes?
A more precise proposition might be: “Structured online retrieval practice improves short-term factual recall among undergraduate students compared with rereading.” That claim specifies a population, intervention, comparator, and outcome closely enough to evaluate what evidence bears on it.
The more ambitious the claim becomes, the more evidence it usually requires.
Separate observations from conclusions
A body of studies may establish that a pattern has been repeatedly observed without establishing the explanation for that pattern.
For example, several observational studies may consistently find that students who participate more frequently in online discussions obtain higher course grades. That may establish a recurring association under the studied conditions. It does not automatically establish that increasing discussion participation causes grades to improve.
Observed pattern
What the studies repeatedly measure, report, or document.
Supported inference
What can reasonably be concluded from that pattern given the designs, methods, uncertainty, and alternative explanations.
Ask what kind of claim the research designs can support
Different research designs answer different questions. Cross-sectional studies may establish associations at a particular time but usually provide limited evidence about temporal order or causation. Longitudinal designs can clarify temporal relationships but may still face confounding. Randomized experiments can provide stronger evidence for causal intervention effects when appropriately designed and conducted.
Qualitative studies can establish credible patterns in experiences, meanings, processes, and contexts when conducted rigorously, but their claims should be matched to their sampling and analytical foundations rather than translated automatically into population prevalence.
There is therefore no single hierarchy that answers every research question. Ask whether the method is capable of supporting the particular proposition you want to make.
Then ask how trustworthy the studies are
An appropriate design can still be poorly conducted. Attrition, confounding, measurement problems, selective reporting, deviations from intended interventions, or other sources of bias may weaken confidence in the findings.
This is why critical appraisal should affect the synthesis rather than remain a ceremonial table somewhere near the appendix. If the studies most directly supporting a claim have serious limitations, the certainty of the claim should change accordingly.
Look for consistency among sufficiently comparable evidence
Repeated findings can strengthen a conclusion when the studies are sufficiently comparable and their limitations do not all point toward the same distortion.
Consistency does not require identical results. Effects may vary in magnitude while still supporting the same general relationship. Conversely, studies can appear consistent only because unlike outcomes have been grouped under one broad label.
Before describing a finding as established, determine exactly what is recurring and whether the studies are addressing the same proposition.
Precision determines how specifically you can characterize an effect
Suppose a meta-analysis estimates a beneficial effect, but its confidence interval is wide enough to include both a trivial effect and a substantial benefit. The direction may appear encouraging, yet the magnitude remains uncertain.
Cochrane guidance emphasizes interpreting effect estimates together with their confidence intervals. Wide intervals indicate greater uncertainty, while narrow intervals can support more precise conclusions when other sources of uncertainty are limited. Cochrane also cautions against reducing interpretation to whether a result crosses a conventional threshold for statistical significance.
Directness determines how far the conclusion travels
A body of evidence can provide a credible answer for one population or context without establishing the same conclusion everywhere.
If every study involves first-year medical students, the literature may establish something reasonably well for that population while providing only indirect evidence for elementary pupils, working professionals, or university students generally.
Formal GRADE assessment treats indirectness as one domain affecting certainty. Cochrane describes this in terms of how well the available evidence matches the population, intervention, comparator, and outcome question to which the conclusion will be applied.
Certainty belongs to particular outcomes and claims
A research topic is rarely simply “established” or “not established.” Evidence can be strong for one conclusion and weak for another.
An intervention might have credible evidence for improving short-term test performance, limited evidence concerning long-term retention, and almost no evidence concerning transfer to authentic tasks. Saying that “the effectiveness of the intervention is established” would conceal these distinctions.
In Cochrane reviews, certainty is assessed for individual outcomes, with GRADE considering domains including risk of bias, inconsistency, indirectness, imprecision, and publication bias.
Distinguish evidence of little or no effect from insufficient evidence
This distinction is essential when deciding what has been established.
If precise evidence indicates little to no meaningful difference between two interventions, the literature may support that conclusion. If an estimate is so imprecise that substantial benefit and substantial harm both remain plausible, the literature has not established equivalence or absence of effect.
Cochrane explicitly warns against confusing lack of evidence of an effect with evidence that an effect is absent.
Watch Out
“No study found a statistically significant difference” does not by itself establish that no meaningful difference exists. Examine the estimates, their uncertainty, and the amount and quality of evidence available.
Ask whether plausible alternatives remain open
A conclusion becomes stronger when the evidence rules out important competing interpretations.
Suppose students who use a voluntary learning platform perform better academically. One explanation is that the platform improves learning. Another is that motivated students are more likely to use it. If the available studies cannot distinguish those possibilities, the association may be established while the causal explanation remains unresolved.
A synthesis should explicitly separate these levels of knowledge.
Look for convergence across appropriately different evidence
Evidence does not need to be identical to contribute to a conclusion. Different methods can sometimes strengthen understanding by addressing complementary parts of the same problem.
An experimental study might estimate whether an intervention changes an outcome, while qualitative studies illuminate how participants use it and longitudinal evidence examines whether benefits persist. Convergence across these forms of evidence can create a richer understanding, provided their distinct contributions are not treated as interchangeable.
Ask what remains true after the qualifications are added
This is one of the most practical tests for determining what the literature establishes.
Begin with your provisional conclusion. Then add the important qualifications: population, context, design limitations, inconsistent findings, measurement differences, follow-up duration, and uncertainty. What claim remains intact?
That residual claim is often much closer to what the literature actually establishes.
| Evidence situation |
Potential conclusion |
What remains unestablished |
| Repeated observational association |
An association is consistently observed under studied conditions |
That the exposure causes the outcome |
| Credible randomized evidence with precise estimates |
The intervention affects the measured outcome under the studied conditions |
Automatic generalization to different populations, implementations, or outcomes |
| Consistent qualitative findings |
A recurring experience or process is documented across the included contexts |
Its population prevalence unless appropriate evidence supports that inference |
| Imprecise evidence around little effect |
The effect remains uncertain |
That there is definitively no meaningful effect |
| Conflicting credible studies |
The evidence does not support one stable general conclusion, or the relationship may be conditional |
A universal effect unless the variation can be explained defensibly |
“Established” should not mean “permanently settled”
Empirical conclusions remain open to revision when stronger evidence, new populations, improved measurements, or better explanations become available.
This does not mean research can never establish anything. It means that empirical knowledge is bounded by the evidence on which it rests. A conclusion can be well supported while remaining conditional on its scope and assumptions. Such qualification is ordinary scholarship, even if it makes fewer appearances in press releases.
04 · A Practical Example
Moving From “The Studies Are Positive” to What Is Actually Established
Hypothetical Example
Does retrieval practice improve university students' learning?
Imagine ten hypothetical studies. Seven controlled experiments report better performance on later tests after retrieval practice than after rereading. Two report smaller and imprecise differences. One finds little difference when the final assessment requires transfer to substantially different problems. Most studies involve undergraduate students, use relatively short follow-up periods, and assess retention of material similar to what was practiced.
The overly broad conclusion
“Research has established that retrieval practice improves learning.”
The statement may be directionally plausible, but “learning” covers considerably more than the evidence described.
Work downward to the defensible claim
Identify the recurring finding Most hypothetical controlled studies report improved later performance relative to rereading.
Specify the outcome The strongest evidence concerns retention on assessments relatively similar to the practiced material.
Specify the population The evidence is concentrated among undergraduate students.
Identify the temporal boundary Most studies examine relatively short follow-up periods.
Preserve the unresolved issue Evidence for transfer to substantially different tasks and longer-term retention remains less certain.
A more defensible conclusion would therefore be: “The hypothetical evidence supports a beneficial effect of retrieval practice over rereading on relatively near-term retention among undergraduate students, while its effects on longer-term retention and transfer to substantially different tasks remain less well established.”
That sentence says more about the state of knowledge than the broader claim because it distinguishes what is supported from what remains open.