03 · What You Need to Know
Scientific Knowledge Is Cumulative, but Accumulation Is Not Simple Addition
Each Study Contributes a Particular Piece of Evidence
A research study investigates a question using particular observations, methods, measurements, assumptions, participants, settings, and analytical procedures.
Its findings therefore have a scope.
A study might establish that a relationship appears in one population. Another may estimate its magnitude more precisely. A third might test a possible mechanism. Another may examine whether the relationship persists under substantially different conditions.
These studies do not have to be duplicates to contribute to the same developing understanding.
This follows from how research produces knowledge: each investigation connects a question to evidence and an evidence-based conclusion, but the broader scientific claim must eventually be evaluated beyond any one study.
Repeated Studies Test Whether Findings Are Isolated
One of the first questions after an interesting result is whether something similar happens again when researchers obtain new data.
The National Academies defines replicability as obtaining consistent results across studies aimed at answering the same scientific question, with each study obtaining its own data.
Consistency across studies can increase confidence that a finding is not merely an isolated coincidence or an unusual feature of one dataset. This is one reason replication contributes to scientific reliability.
But replication is not binary. Results do not have to be numerically identical, and one replication does not certify a claim permanently. Researchers evaluate consistency in relation to the uncertainty inherent in the phenomenon and the studies.
Conceptual Replication Can Test Whether a Claim Survives Methodological Change
Researchers may also investigate the same underlying proposition using different operational definitions, instruments, populations, analytical approaches, or experimental procedures.
If a conclusion remains compatible across methods that have different weaknesses, confidence may increase because the result is less easily explained as an artifact of one particular procedure.
This is sometimes described as conceptual replication, although terminology varies among disciplines.
The underlying idea is broader than repeating a protocol: scientific claims become more informative when researchers learn whether they survive meaningful changes in how they are investigated.
Different Studies Can Reveal Generalizability
A finding observed in one context may or may not apply elsewhere.
Studies conducted across different populations, institutions, cultures, periods, environments, or implementation conditions can help researchers determine the boundaries of a claim.
The National Academies distinguishes this issue from replicability by defining generalizability as the extent to which results apply to contexts or populations different from those in the original study.
If findings remain compatible across diverse conditions, a broader conclusion may become defensible. If they differ systematically, researchers may discover important boundary conditions instead.
Disagreement Can Add Knowledge Rather Than Merely Subtract Confidence
Cumulative research is not a process in which every new study must agree with the previous ones.
When findings differ, researchers can ask whether the discrepancy arises from sampling variation, measurement, context, implementation, analytical decisions, methodological problems, or genuine differences in the underlying phenomenon.
This is why two well-conducted studies can reach different conclusions without making cumulative knowledge impossible.
Indeed, disagreement may transform a broad claim into a more precise one. “The intervention improves learning” might eventually become “The intervention tends to improve delayed retention when implemented with repeated retrieval opportunities, but benefits are smaller under other conditions.”
The second conclusion is less simple but potentially more informative.
Convergence Across Different Lines of Evidence Can Be Particularly Powerful
Research does not always accumulate through repeated versions of the same study.
Different methods may address complementary aspects of a question. An experiment might estimate an intervention's effect. Observational research might examine how it operates under routine conditions. Qualitative research might identify implementation processes or participant experiences. Longitudinal evidence might reveal whether effects persist.
These forms of evidence answer different questions, so they should not be treated as interchangeable. When their implications are compatible, however, they can contribute to a richer explanation than any one method provides.
The National Academies describes scientific confidence as emerging through a web of evidence developed across multiple lines of inquiry, not solely through one-to-one replication between individual studies.
Cumulative Knowledge Is Not a Vote Among Papers
Suppose 12 studies report results broadly supportive of a claim while four do not. It may be tempting to conclude that the claim wins 12 to 4.
That is not a defensible evidence synthesis.
The studies may differ greatly in sample size, risk of bias, relevance, measurement quality, precision, independence, population, design, and the questions they actually address. Several publications might even report different analyses of the same underlying dataset.
Watch Out
Scientific evidence should not be summarized by counting supportive and unsupportive papers. The weight of evidence depends on the characteristics and limitations of the studies, not merely on how many conclusions fall on each side.
Systematic Reviews Make the Accumulation Process Explicit
As a literature grows, informal reading becomes increasingly vulnerable to selective attention. Researchers may encounter the most famous studies, the newest papers, or results that happen to support their expectations while missing other relevant evidence.
Systematic reviews address this problem by using explicit methods to identify, select, appraise, and synthesize studies relevant to a defined question.
Cochrane describes evidence synthesis as bringing together data from included studies to draw conclusions about a body of evidence. The process includes examining study characteristics before statistical synthesis is considered.
This matters because meaningful synthesis requires researchers to understand what is being combined.
Meta-Analysis Can Combine Compatible Quantitative Evidence
When studies estimate sufficiently comparable quantities, meta-analysis can statistically combine their effect estimates.
The result may provide a summary estimate together with an indication of its uncertainty. Meta-analysis can also help researchers investigate variation among study results.
But statistical combination is not automatically appropriate simply because several studies exist. Cochrane emphasizes that researchers must first consider the review question, eligibility criteria, study characteristics, risk of bias, comparisons, and whether the available numerical results are meaningful to combine.
A polished forest plot cannot rescue an incoherent research question. Academics occasionally need reminders that diamonds at the bottom of figures are not epistemological gemstones.
Synthesis Can Reveal Heterogeneity
When studies produce different estimates, that variation may itself become an object of investigation.
Researchers may examine whether effects differ according to population characteristics, intervention intensity, measurement choices, research design, follow-up duration, or other factors.
This process can reveal heterogeneity, meaning that the phenomenon does not behave identically across all studies or conditions.
Rather than asking only for one average answer, cumulative research can therefore identify how and why the answer changes.
Accumulating Evidence Can Change the Original Question
Early research often begins with relatively broad questions: Does the intervention work? Are these variables related? Does this phenomenon occur?
As evidence accumulates, those questions can become more sophisticated.
Researchers may begin asking how large an effect is, which mechanism explains it, who benefits most, when the relationship disappears, what unintended consequences occur, or how long an outcome persists.
Knowledge has progressed not because researchers have stopped asking questions, but because earlier evidence has made more precise questions possible.
Contrary Evidence Can Revise the Body of Knowledge
Cumulative knowledge is not simply a one-way process in which confidence increases with every publication.
A strong new study may expose a weakness in earlier work. Improved measurement may reveal that an accepted effect was smaller than believed. Studies in new populations may show that a conclusion does not generalize as broadly as assumed. Reanalysis may identify an error.
New evidence can therefore increase, decrease, or redirect confidence.
This is part of why scientific knowledge can change when new evidence appears. A cumulative system must be capable of subtraction and revision as well as addition.
The Body of Evidence Becomes the Relevant Unit of Interpretation
Once substantial research exists, asking what “Study A says” is often less useful than asking what the relevant evidence collectively supports.
The National Academies argues that reviews of cumulative evidence can be more useful for assessing overall effects and generalizability than concentrating predominantly on the replicability of individual studies.
This is the transition from individual findings to a body of knowledge.
It also explains why one research study is rarely enough to give a definitive answer. Scientific understanding emerges from relationships among studies, not merely from the existence of individual publications.