03 · What You Need to Know
A Good Search Is Good for Something
The phrase “good literature search” is incomplete unless you specify what the search is supposed to accomplish.
If you want to understand what a new concept means, a search may have done its job once you can identify the important terminology, conceptual distinctions, foundational literature, and major scholarly conversations well enough to continue your research intelligently.
If you are developing an empirical study, the search needs to give you sufficient command of the relevant literature to formulate and justify the problem, position the study, inform appropriate methodological decisions, and interpret the eventual findings responsibly.
If you are conducting a systematic review, the standard changes again. The search helps determine which studies enter the evidence base. Search inadequacy can therefore affect the synthesis itself.
This is why the purpose of the literature search must be clear before its quality can be judged.
Start by Asking What Would Go Wrong If the Search Were Inadequate
A useful way to calibrate search quality is to consider the consequences of missing relevant literature.
Suppose you are searching to learn the terminology of an unfamiliar topic. Missing several papers may not prevent the search from accomplishing that purpose. You can continue exploring as your understanding develops.
Now suppose you are conducting a systematic review and several eligible studies are missed. Those omissions may alter the characteristics of the evidence base, affect estimates or interpretations, and potentially change the conclusions.
A 2024 case analysis compared two systematic reviews addressing a similar research question but using different limited database searches. One included 16 studies and the other 11, with only four studies overlapping. The example does not establish a universal database minimum, but it illustrates why apparently modest differences in evidence identification can produce substantially different review evidence sets.
The appropriate standard therefore rises with the consequences of omission.
Search Quality Begins With Conceptual Validity
Before asking how many databases you searched, ask whether the search actually represents the question you intend to investigate.
A technically flawless search for the wrong concept is still the wrong search.
Suppose your research concerns students' uncritical reliance on generative AI. You build an elaborate search around “technology dependence.” The syntax may be perfect, the databases appropriate, and the documentation immaculate. But if technology dependence represents a meaningfully different construct from the phenomenon you intend to study, retrieval quality is compromised at the conceptual level.
This is why search-strategy quality begins before Boolean operators.
You need to understand the research problem, identify the concepts that need to be searchable, and determine how those concepts are represented in the literature. If you initially know little about the topic, exploratory searching can help build that understanding before the more consequential search is finalized.
A Good Search Uses the Language of the Literature, Not Just Your Language
Relevant studies may describe the same concept using different terminology.
A developed search may therefore need synonyms, spelling variants, abbreviations, acronyms, older terminology, newer terminology, and controlled vocabulary where appropriate. At the same time, conceptually related terms should not be added indiscriminately merely to increase retrieval.
Quality lies in representing the concept adequately.
Known relevant papers are useful here. Examine how they describe the phenomenon in titles, abstracts, keywords, and indexing. If your strategy cannot retrieve several highly relevant papers you already know should be discoverable in the database, investigate why.
The explanation may reveal missing terminology, an overly restrictive search concept, unsuitable field searching, indexing differences, or another weakness in the strategy.
Retrieving Known Relevant Studies Is a Useful Test, but Not Proof of Completeness
Testing whether a strategy retrieves known relevant studies is sometimes called known-item testing.
It can answer a valuable diagnostic question: Can this search retrieve examples of the evidence it is supposed to find?
If the answer is no, something deserves investigation.
But retrieving every paper in your test set does not establish that the search will retrieve every other relevant study. The known papers may use unusually obvious terminology or come from journals well represented in the selected databases.
Known-item testing is therefore evidence about search performance, not certification of comprehensiveness.
The Sources You Search Must Fit the Evidence You Need
A search cannot retrieve records that are absent from the source being searched.
Database selection is therefore a substantive part of search quality. Relevant evidence may be distributed across disciplinary databases, multidisciplinary citation indexes, regional databases, trial registers, repositories, organizational websites, grey literature sources, or other information systems.
Searching more databases does not automatically solve this problem.
Five databases with heavily overlapping coverage may contribute less than three carefully chosen sources that cover different important parts of the evidence space. Conversely, one excellent database may still be insufficient for a question whose literature crosses disciplinary boundaries.
For systematic reviews, methodological guidance such as the Cochrane Handbook recommends selecting databases and supplementary sources according to the topic and the kinds of evidence being sought rather than relying on an arbitrary database count.
Watch Out
Rules such as “three databases are enough” or “five databases make a search comprehensive” are poor substitutes for source selection based on the actual evidence landscape. Ask what important literature each source can contribute and what may remain uncovered if it is omitted.
More Results Do Not Mean a Better Search
Result count is one of the easiest search characteristics to observe and one of the easiest to misinterpret.
A search retrieving 20,000 records is not automatically stronger than one retrieving 2,000. The larger search may simply contain far more irrelevant material.
Information retrieval provides two useful concepts for thinking about this problem: sensitivity and precision.
These measures involve a trade-off. Research evaluating literature searches has long recognized that increasing recall can reduce precision. A highly sensitive strategy may intentionally retrieve many irrelevant records so that fewer relevant studies are missed.
For Cochrane intervention reviews, the recommended objective is to maximize sensitivity while striving for reasonable precision. That balance is appropriate because missing eligible studies can be more consequential than screening additional irrelevant records.
Other search purposes may reasonably strike a different balance.
Precision Matters Because Researcher Time Is Not Infinite
It is easy to dismiss precision as mere convenience. It is more consequential than that.
Every irrelevant record consumes screening time. Extremely low precision can make a search impractical, particularly when researchers have limited resources. Screening fatigue may itself create opportunities for mistakes.
The goal is not maximum precision at any cost, however. An extremely precise search may achieve its efficiency by excluding terminology or concepts necessary to retrieve less obvious relevant studies.
A good search therefore manages the trade-off rather than optimizing one metric blindly.
If removing a term cuts 5,000 irrelevant records but also eliminates several known eligible studies, the apparent efficiency gain may not be defensible for a systematic review. If a term produces thousands of irrelevant records and no meaningful relevant retrieval, retaining it merely because “broader is better” may be equally difficult to justify.
Search Quality Cannot Usually Be Reduced to One Metric
Sensitivity and precision are useful concepts, but they do not capture the entire quality of a literature search.
A search may retrieve known relevant papers yet omit an important database. Another may use excellent databases but represent a central concept poorly. A third may have good retrieval performance but be reported so incompletely that readers cannot evaluate what was done.
The PRESS guideline reflects this multidimensional nature of search quality. It provides an evidence-based framework for peer review of electronic search strategies and considers elements such as translation of the research question, Boolean and proximity operators, subject headings, text words, spelling and syntax, and limits and filters.
These dimensions show why search quality cannot be judged from a keyword list alone.
Peer Review Can Improve Consequential Search Strategies
When the search is central to an evidence synthesis, another person can sometimes see problems that the original searcher has stopped noticing.
PRESS was developed specifically for peer review of electronic search strategies used in systematic reviews, health technology assessments, and other evidence syntheses. Its checklist provides a structured way to examine whether the question has been translated appropriately and whether the search contains problems in subject headings, free-text terms, operators, syntax, spelling, or limits.
Cochrane strongly recommends peer review of search strategies for intervention reviews before searches are run.
Peer review does not prove that a search will retrieve every eligible study. It is quality assurance: an opportunity to identify avoidable errors before they propagate through the review.
A Good Systematic Search Should Be Reproducible Enough to Inspect
Search quality is difficult to evaluate when nobody can tell what was actually done.
For systematic reviews, documentation should allow readers to identify the databases and platforms searched, understand the search strategies, see important limits or filters, know when searches were conducted, and understand relevant supplementary methods.
PRISMA-S was developed to improve reporting of literature searches in systematic reviews and provides specific reporting items for information sources, search strategies, limits, supplementary searches, peer review, search updates, and related procedures.
Reproducibility does not make a weak search good. A poor strategy can be documented perfectly. But without adequate documentation, methodological strengths and weaknesses become difficult to inspect.
This is why reproducibility, comprehensiveness, and systematicity should be evaluated as related but distinct properties.
A Search Can Be Good Enough Without Being Comprehensive
This is especially important outside formal evidence synthesis.
Suppose you are searching because you need to understand how researchers define “feedback literacy.” You identify several recent reviews, important conceptual papers, competing definitions, major empirical applications, and terminology that allows you to continue searching intelligently.
You probably have not identified every paper about feedback literacy.
That does not necessarily matter.
If your purpose was orientation, the search may be entirely adequate once additional searching stops materially changing your conceptual understanding.
The mistake would be to take that exploratory search and make a stronger claim than it can support, such as “no studies have examined feedback literacy in context X.” Establishing an absence requires greater confidence in evidence coverage than understanding a concept does.
Diminishing Returns Can Be Useful for Exploratory Searching
For open-ended exploratory searches, you may eventually notice that additional searching produces fewer meaningful conceptual discoveries.
The same terminology recurs. You recognize the major authors and debates. New papers largely reinforce categories you already understand. Citation trails lead back to familiar literature.
This can be a practical signal that the exploratory search has accomplished its immediate learning purpose.
It should not be mistaken for proof that every relevant paper has been identified.
Gusenbauer has argued that searching in an environment of abundant scholarly information requires clearer thresholds for deciding when searching is good enough, while recognizing that stopping rules for exploratory searching cannot be as concrete as those used in systematic evidence synthesis.
The useful question is whether further searching is still changing the decision or understanding that motivated the search.
“Saturation” Is Not a Universal Stopping Rule for Literature Searches
Researchers sometimes borrow the language of saturation and stop when they feel that no new papers or ideas are appearing.
That can be a reasonable informal observation during exploratory searching, but it should not automatically be treated as a validated stopping criterion for every literature search.
A new database, terminology variant, citation route, or disciplinary perspective may reveal literature that repetitive searches in the same source cannot.
For systematic reviews, stopping should follow the planned search methods rather than the researcher's subjective impression that enough literature has appeared.
In other words, repetition within one retrieval path does not prove that the wider evidence space has been adequately covered.
Search Adequacy for an Empirical Study Is Usually About Intellectual Coverage
When the literature search supports an empirical paper, thesis, or dissertation rather than constituting the primary research method, the central question is usually whether you understand the scholarship necessary to conduct and position the study responsibly.
You should know the important concepts and theories relevant to your problem. You should understand closely related empirical findings and meaningful disagreements. You should be aware of methodological approaches and limitations that affect your own design. You should be able to explain how your study relates to what has already been done.
This does not require citing every paper that exists.
Indeed, a manuscript can contain hundreds of references and still misunderstand the literature if it ignores a central theoretical distinction or major contradictory evidence.
Coverage should be intellectual as well as numerical.
Search Adequacy for a Review Study Is More Methodological
When the literature itself becomes the object of systematic evidence synthesis, search adequacy requires a more explicit methodological defense.
You need to ask whether the search concepts align with eligibility criteria, whether important terminology has been represented, whether appropriate databases and supplementary sources were searched, whether unjustified restrictions could introduce bias, whether the strategies were tested, and whether the search can be reported transparently.
The consequences of an inadequate search can be substantial.
Research comparing literature searches has shown that different search strategies can produce meaningfully different evidence sets. The quality of literature searching therefore contributes directly to the quality of the resulting systematic review.
This is why systematic searching is necessary for some studies but disproportionate for others.
The Number of Databases Is Evidence About the Search, Not a Quality Score
It is tempting to convert methodological judgment into a checklist:
Two databases: weak. Three databases: acceptable. Five databases: excellent.
That is too simple.
A study examining two systematic reviews with similar questions found that limited database searching contributed to substantially different sets of included studies. That finding provides a useful warning against overly narrow database coverage, but it does not establish that a particular number of databases guarantees adequacy.
The appropriate sources depend on the topic. A highly specialized question may be concentrated in a small number of databases. A multidisciplinary question may require considerably broader coverage.
The better question is: Which important part of the relevant evidence could I plausibly miss because of where I did or did not search?
Search Restrictions Need Justification
Date limits, language restrictions, publication-type restrictions, study-design filters, geographical restrictions, and other limits can sometimes be methodologically appropriate.
They can also remove relevant evidence.
A good search does not avoid all restrictions. It uses restrictions for reasons connected to the research question or methodology rather than simply because they make the result set easier to manage.
For example, restricting a review to studies published after a particular year may be defensible when an intervention or technology did not exist earlier. Applying the same restriction because screening older studies would take too long is a different justification and should be treated as a methodological constraint.
The search should make such decisions visible rather than allowing convenience to masquerade as conceptual relevance.
Quality Also Depends on Whether the Search Is Current Enough
A well-designed search can become outdated.
This matters particularly for systematic reviews and other evidence syntheses because new eligible studies may appear between the original search and publication.
Search updating is therefore part of the evidence-identification process. PRISMA-S includes reporting of search updates, and review methodologies may specify expectations concerning how recently searches should have been run.
For ordinary research projects, the same principle applies less formally. A literature review written from searches conducted several years earlier may no longer represent a rapidly developing field adequately.
Search quality is therefore partly temporal: the evidence base should be sufficiently current for the claim and research decision being made.
Good Enough Does Not Mean Perfect
Every literature search operates under constraints.
Databases have incomplete and overlapping coverage. Indexing is imperfect. Terminology varies. Search systems impose technical limits. Some evidence is difficult to access. Researchers have finite time, expertise, language capability, and resources.
A defensible search acknowledges those realities.
The goal is not to prove that no better search could possibly exist. It is to demonstrate that the search was sufficiently well designed and executed for its purpose, that important avoidable weaknesses were addressed, and that remaining limitations are proportionate to the conclusions being drawn.
That is a much more useful standard than perfection.