03 · What You Need to Know
Why a Search Can Return Far More Results Than You Expected
First, Ask Whether the Number Is Actually a Problem
There is no universally correct number of database results.
Two hundred results can be far too many for a highly specific question. Twenty thousand can be plausible for a broad systematic search in a large database. The appropriate number depends on the research question, database coverage, review type, number of concepts, terminology, and how much screening is feasible.
Result count is therefore a diagnostic signal, not a quality score.
Large result set
A numerical description of how many records satisfied the search.
Overly broad search
A methodological judgment that too many retrieved records fall outside the information need because the strategy represents the question poorly or too loosely.
Do not assume the second merely because you observe the first.
Look at the Results Before Editing the Query
A result count tells you how many records matched. It does not tell you why.
Inspect a sample of the results, including records near the top and, where the database allows useful sorting, records from different parts of the result set. Ask what makes the irrelevant records qualify.
Do they share one problematic keyword? Are they from another discipline? Is an acronym being interpreted differently? Is a broad term retrieving a neighboring concept? Does one concept appear only incidentally in abstracts?
Search refinement should respond to a pattern in the retrieval, not merely to the size displayed above the results.
Break the Search Into Concept Blocks
If your complete strategy contains several concepts, run them separately.
Suppose your search has this structure:
This is much more informative than modifying the complete query at random.
Concept blocks also make it easier to see whether parentheses are grouping your Boolean logic correctly.
One OR Term Can Dominate an Entire Concept Block
Suppose your concept is online learning:
("online learning" OR e-learning OR "online education" OR technology)
The first three expressions may be reasonably focused. Technology is much broader.
Because the terms are connected with OR, every record satisfying technology can enter the concept set even if it has nothing to do with online learning.
The problem is not OR. OR is doing exactly what it should: providing alternative routes into the concept block. The problem is that one alternative does not represent the same concept closely enough.
Watch Out
When an OR group is unexpectedly large, test its terms individually. A single broad or ambiguous alternative can contribute most of the irrelevant retrieval while hiding among several perfectly reasonable synonyms.
Test Individual Terms Within a Problematic OR Group
Run each important alternative separately and compare what it retrieves.
Term A Does it retrieve the intended concept consistently?
Term B Does it add useful terminology or mostly duplicate Term A?
Term C Does it have another disciplinary meaning?
Term D Is it actually broader than the concept rather than an alternative expression for it?
If one term is responsible for most of the noise, refine or remove that term rather than narrowing the entire search.
This is one reason the process of finding alternative terms for the same concept requires conceptual judgment rather than synonym accumulation.
Ambiguous Acronyms Can Produce Enormous Result Sets
Short acronyms are frequent offenders.
A two- or three-letter abbreviation may represent your concept in one discipline and several unrelated concepts elsewhere. Searching it across broad fields can introduce thousands of records.
Before deleting the acronym, test what it uniquely contributes. If it retrieves relevant articles that the expanded term misses, you may still want it. Consider whether it can be searched in a more informative field, combined with another positive term, or used in a proximity expression where supported.
If it contributes almost no unique relevant material, exclusion may be defensible.
The broader principles for handling acronyms and abbreviations in database searches are particularly important when a result set suddenly becomes enormous.
Check Whether Truncation Is Too Aggressive
Truncation deliberately expands retrieval. If the stem is too short, it can retrieve many unintended words.
Suppose you truncate a term at a point that seems efficient. Run the truncated expression separately and inspect the vocabulary in the resulting records. You may discover that the stem matches several unrelated word families.
If so, move the truncation point to the right or replace the truncated expression with a smaller set of explicit alternatives.
The shortest possible stem is not the most rigorous. It is merely the stem with the greatest capacity for mischief.
Wildcards Can Also Expand More Than Expected
A wildcard intended to capture one spelling difference may represent more characters or positions than you assumed.
This is particularly likely when a search has been copied from another database. Wildcard and truncation symbols are platform-specific, and identical-looking symbols can behave differently.
Check the database's current truncation and wildcard rules, then test the expression independently.
Check Whether Your Phrase Is Being Searched as a Phrase
A multiword concept searched as separate words can retrieve records in which the words occur independently.
Suppose you intend the established concept social media but the platform processes the words separately or expands them through its own query system. Records containing social in one context and media in another may become eligible depending on the database.
If the expression functions as a stable phrase in the literature, test phrase searching.
Do not automatically place every multiword term inside quotation marks. Exact phrases can also become too restrictive. The appropriate question is whether keeping the words together improves representation of the concept.
AND Can Be Too Loose When the Relationship Between Words Matters
Suppose you search:
teacher AND burnout
If both terms can appear anywhere in a long abstract or full-text field, you may retrieve articles in which they are unrelated.
If the intended concept depends on the terms occurring near one another, proximity searching may be more precise than AND.
A proximity condition can require the words to occur within a defined distance without forcing one exact phrase. This is especially useful in long text fields where accidental co-occurrence becomes more likely.
Check Which Fields the Database Is Actually Searching
A search may be broad because the interface is searching more text than you realized.
A field labeled Topic, Keyword, or Anywhere may search titles, abstracts, author keywords, subject terms, indexing fields, or even full text depending on the platform.
Search terms that perform well in titles and abstracts can become extremely noisy when expanded to full text.
Before changing the vocabulary, verify which fields the search is using.
Field Restriction Can Improve Precision Without Changing the Concept
If one term is useful but ambiguous in broad fields, consider restricting that particular term rather than every term.
For example:
The exact syntax depends on the database, and the restriction should be tested against known relevant records.
Do Not Automatically Restrict Everything to the Title
Title-only searching is a powerful way to reduce results. That does not make it the correct response to a large result set.
A relevant article may study your concept without naming it in the title. Restricting every concept to titles can substantially reduce recall, especially when terminology varies.
Use title restrictions when the concept must be prominent or when a particular term is otherwise too ambiguous. Do not use them merely because you want the database to display a smaller number.
Full-Text Searching Can Produce Large Amounts of Incidental Retrieval
If the database or discovery system searches article full text, every introduction, methods section, discussion, reference, and caption can become another place for your term to match.
This can be useful when you need information buried inside articles, but it is often unnecessarily broad when you are looking for papers centrally about a topic.
If full-text retrieval is causing the problem, consider whether full-text searching is actually appropriate for the information need or whether bibliographic fields would provide a better signal.
Check the Boolean Grouping Before Adding Another Concept
A missing pair of parentheses can make a search dramatically broader than intended.
Suppose you meant:
(adolescent OR teenager) AND depression
but entered:
adolescent OR teenager AND depression
Depending on the platform's precedence rules, records satisfying adolescent may enter the result set without satisfying the depression requirement.
Before assuming your terminology is too broad, check the Boolean structure.
Inspect the Database's Query Translation
Some platforms automatically modify what you type.
PubMed, for example, uses Automatic Term Mapping for many untagged searches. A seemingly simple query may be translated into MeSH concepts, synonyms, and other search expressions. Quotation marks, field tags, wildcards, and proximity syntax can alter that processing.
If the result set is unexpectedly large, inspect Search Details or the equivalent query translation where the database provides one.
You may discover that the database searched more than the visible query suggested.
Do Not Add Another Concept Solely to Reduce the Number
This is one of the most common ways a search quietly changes its question.
Suppose your actual question concerns online learning among university students. The search returns 12,000 records. You add academic performance because you are particularly interested in that outcome and the result count falls to 800.
That may look like an improvement. But studies of online learning among university students that investigate engagement, satisfaction, persistence, attendance, or other relevant outcomes are now excluded.
If academic performance was always an essential component of the research question, the new concept may be appropriate. If it was added only to reduce the result count, you have narrowed the question rather than merely improved the search.
Watch Out
Every new AND concept is an additional eligibility requirement at the retrieval stage. Do not add concepts merely because they reduce the result count. Add them only when the research question genuinely requires records to represent them.
More AND Is Not Automatically Better Precision
AND usually reduces results because records must satisfy additional conditions. The danger is that relevant records may not mention every concept in searchable metadata.
Cochrane's current Handbook cautions against searching every aspect of a review question when some elements, such as outcomes or comparators, may not be well represented in titles, abstracts, or indexing.
For high-sensitivity searches, a somewhat broader retrieval set may be methodologically preferable to an elegant but brittle strategy that requires too many concepts.
Use NOT With Particular Caution
NOT can make a large search dramatically smaller, which makes it tempting.
Suppose many irrelevant records concern children, so you add:
NOT child*
A relevant mixed-population study involving adults and children may now disappear.
The problem is not that NOT failed. It successfully removed records satisfying the excluded condition, including some that may also contain the evidence you need.
Before using exclusion to clean up a large result set, understand when Boolean NOT can remove relevant research.
Try Positive Refinement Before Negative Exclusion
Instead of defining what you do not want, first improve the representation of what you do want.
If a term has an unwanted meaning, possible responses include:
- replace it with a more specific synonym;
- restrict that term to a more informative field;
- search a stable multiword expression as a phrase;
- use proximity where textual closeness is meaningful;
- remove an overly broad OR alternative;
- correct truncation or wildcard behavior.
These changes can improve precision without subtracting an entire set of potentially relevant records.
Date Limits Should Follow the Question, Not the Result Count
Restricting a search to the last five or ten years is an efficient way to remove thousands of records. It is methodologically defensible only if the time restriction has a substantive rationale.
Examples might include:
- a technology that did not exist before a particular period;
- a policy introduced on a known date;
- a review intentionally updating an earlier review;
- a research question explicitly concerned with contemporary evidence.
"There were too many articles" is a practical concern, but it does not by itself justify declaring older research irrelevant.
Language Limits Can Also Exclude Relevant Evidence
Restricting results to English can reduce workload, but it can introduce language-related selection effects and exclude eligible studies.
Cochrane's current methodological standards state that searches for studies should not be restricted by language or publication status when conducting Cochrane intervention reviews.
Other review types may operate under different constraints, but the general principle remains useful: language limits should be justified transparently rather than used casually as a result-reduction tool.
Document-Type Limits Need the Same Scrutiny
A database may allow you to restrict results to journal articles, reviews, conference papers, clinical trials, dissertations, or other publication types.
Such limits can be appropriate when the research design or eligibility criteria require them. They can also exclude relevant studies because publication-type indexing is incomplete, inconsistent, or unavailable for newer records.
If study design matters, consider whether a validated search filter is more appropriate than an interface checkbox and whether the filter's sensitivity is suitable for the review.
Database Filters Are Not Automatically Methodologically Neutral
A filter button can feel safer than editing a Boolean search because it is built into the interface. It is still a retrieval decision.
Filters for date, language, age, study type, geography, species, or other characteristics depend on metadata and indexing. Records lacking the relevant metadata may be excluded even when they satisfy the substantive eligibility criteria.
Before applying a filter, determine:
- what metadata it uses;
- whether all relevant records receive that metadata;
- whether recent or unindexed records are affected;
- whether the filter corresponds exactly to your eligibility criterion.
Subject Headings Can Improve Precision When They Represent the Concept Well
If your free-text search is broad because terminology is ambiguous, a well-chosen subject heading may provide a more conceptually focused retrieval route.
That does not mean replacing keywords with subject headings. Relevant records may lack completed indexing or use emerging terminology. A combination of controlled vocabulary and free text can provide complementary access.
The usefulness of subject headings depends on whether the database has an appropriate controlled vocabulary and whether the heading accurately represents your concept.
Check Whether a Subject Heading Is Exploding Too Broadly
Controlled-vocabulary searches can also be unexpectedly broad.
A broader subject heading may automatically include narrower concepts beneath it. If those narrower concepts extend beyond your intended scope, explosion can add large amounts of irrelevant literature.
Inspect the hierarchy. Determine which narrower terms are being included and whether they belong to the research question.
Do not turn off explosion merely to reduce the result count. Turn it off when the narrower concepts are substantively inappropriate.
Search Each Concept Through Both Controlled Vocabulary and Free Text Before Diagnosing the Final Set
For a substantial literature search, it can be useful to compare the contribution of the two retrieval routes.
Controlled-vocabulary route What does the subject heading retrieve?
Free-text route What do the title, abstract, and keyword terms retrieve?
Combined concept What unique records does each route contribute when connected with OR?
Final search How does this concept behave once combined with the other concept blocks?
This can reveal whether the broad retrieval comes from the vocabulary term, the free-text terms, or both.
Do Not Judge Search Quality by Precision Alone
A search returning 100 highly relevant-looking results can feel better than one returning 5,000 mixed results. But if the 100-result search misses half of the eligible literature, its apparent precision comes at a substantial cost.
Search development often involves balancing precision and sensitivity or recall.
Higher precision
A larger proportion of the retrieved records are relevant, reducing screening burden.
Higher sensitivity
A larger proportion of the relevant records available in the searched resource are successfully retrieved.
Cochrane describes systematic-review searching as an attempt to maximize sensitivity while maintaining reasonable precision. That balance will differ for an exploratory search, scoping review, systematic review, or highly focused evidence query.
A Large Search May Be the Honest Representation of a Broad Question
Sometimes there is no technical problem to fix.
If your question concerns a heavily researched topic using broad concepts, thousands of records may genuinely satisfy the search. Narrowing further would require changing the conceptual scope or imposing additional eligibility criteria.
In that situation, the methodological response may be to manage screening efficiently rather than distort the search.
Screening tools, deduplication, prioritization workflows, and carefully designed eligibility criteria can reduce workload after retrieval. Search design should not be asked to perform every screening decision in advance.
Search Strategy and Screening Strategy Solve Different Problems
The search determines which records have the opportunity to be considered. Screening determines which retrieved records meet the eligibility criteria.
Trying to make the search perform all screening decisions can produce an excessively brittle query.
| Stage |
Main Question |
Typical Priority |
| Searching |
Which potentially relevant records should enter the candidate set? |
Adequate retrieval of relevant evidence |
| Deduplication |
Which records represent the same underlying citation? |
Remove duplicate workload without losing unique studies |
| Title/abstract screening |
Which records plausibly meet the eligibility criteria? |
Exclude clearly irrelevant records using human judgment |
| Full-text screening |
Which studies actually satisfy the detailed criteria? |
Make eligibility decisions using fuller information |
A large result set can therefore be inconvenient without being methodologically wrong.
Deduplicate Before Panicking About the Total
If you are searching multiple databases, the combined number of retrieved records can be much larger than the number of unique citations.
The same article may appear in MEDLINE, Embase, Scopus, Web of Science, CINAHL, PsycInfo, or other resources. Before interpreting the total screening burden, consider the likely effect of deduplication.
A search returning 15,000 records across several databases does not necessarily mean 15,000 unique articles require screening.
Do Not Optimize the Search to an Arbitrary Target Number
There is no methodological rule that says a database search should return 500 records, 1,000 records, or any other convenient number.
Optimizing toward a target count can create perverse decisions: adding a concept until the result count falls below 1,000, restricting to titles because 3,000 feels excessive, or removing older literature because the screening workload is uncomfortable.
The correct target is not a number. It is a defensible balance between finding relevant evidence and avoiding avoidable noise.
Make One Change at a Time
If you change the field, add a phrase, remove a synonym, add a concept, and apply a date filter simultaneously, you will not know which modification caused the improvement or damage.
Refine iteratively:
Baseline Save the original search and result count.
Change one element For example, remove one broad synonym.
Compare Inspect the new count and the records gained or lost.
Check relevance Test known relevant records and sample the affected results.
Keep or reverse Retain the change only if it improves the search for a defensible reason.
Continue Move to the next problem only after understanding the first modification.
This turns search refinement into an auditable process rather than a sequence of increasingly desperate clicks.
Known Relevant Records Are Useful Guardrails
If you know several articles that should be retrievable, check them after consequential narrowing decisions.
Did a title restriction remove one? Did eliminating an ambiguous acronym lose a paper that uses only the acronym? Did a date limit exclude foundational research that actually belongs to the question?
Known-item testing cannot establish complete sensitivity, but it can reveal obvious over-narrowing.
Inspect What a Narrowing Change Removes
Do not look only at the cleaner result set that remains.
If possible, construct or inspect the difference between the original and revised searches. What records disappeared? Are they genuinely irrelevant, or did the change remove a meaningful subset of the literature?
This is particularly important for NOT, field restrictions, date limits, and removal of high-yield synonyms.
Stopping Is Also Part of Search Refinement
Once the obvious sources of noise have been addressed, additional narrowing can produce diminishing returns.
You may reach a point where another restriction removes relatively few irrelevant records but begins to threaten relevant retrieval. At that stage, screening may be the better tool.
The same principle applies when deciding when to stop adding search terms: search strategies should stop growing or shrinking when additional complexity no longer improves the representation of the information need.