Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Keep Track of Which Studies Support, Contradict, or Complicate an Argument?

Studies rarely fit neatly into supporting and opposing camps. Track the specific claim each study bears on, the evidence it contributes, and the conditions that may explain agreement, contradiction, or qualification.

182
Tracking Evidence Across Studies Guide 182 of 247
01 · The Question

How Do You Know What the Literature Actually Says About an Argument?

After reading enough papers, you begin making mental classifications. This study supports the argument. That one contradicts it. Another seems mixed. A fourth somehow belongs in all three categories depending on which result you examine.

The difficulty is that research evidence rarely behaves like a vote. One paper may support a relationship in one population but not another. Two apparently contradictory studies may have measured different outcomes. A study that fails to reproduce an expected result may not necessarily demonstrate the opposite result. Another may support the general argument while revealing an important boundary condition.

If you record only whether each paper is "for" or "against" your emerging position, you can lose precisely the nuance that a literature review is supposed to explain. You need a system that tracks not merely the direction of evidence, but what claim the evidence bears on and under what conditions.

02 · The Short Answer

Track Evidence Against Specific Claims, Not Entire Papers

In Brief

To track whether studies support, contradict, or complicate an argument, define the specific claim first, then record what evidence each study contributes to that claim, the direction and strength of that contribution, and the methodological or contextual conditions needed to interpret it.

A whole paper should not automatically receive a single "support" or "contradict" label. Studies often contain several findings, and those findings may bear differently on different parts of your argument. Treat categories such as support, contradiction, and complication as analytical relationships between evidence and claims, not permanent properties of papers.

03 · What You Need to Know

The Unit You Need to Track Is the Study-to-Claim Relationship

Suppose your emerging argument is:

AI-generated feedback improves university students' academic writing.

It may be tempting to read each paper and assign one label:

  • supports;
  • contradicts; or
  • mixed.

But the apparent simplicity hides several questions. What does "improves" mean? Compared with what? Which students? What kind of writing? What kind of AI-generated feedback? Was improvement measured immediately or later? Does a study measuring student satisfaction provide evidence about writing improvement at all?

Before classifying evidence, you therefore need to know exactly what proposition you are asking the evidence to support.

Start by Breaking Broad Arguments Into Trackable Claims

Broad arguments often contain several propositions that should be evaluated separately.

For example, the broad statement:

AI feedback is effective in higher education.

could conceal very different claims:

  • AI feedback improves revision quality compared with no feedback;
  • AI feedback performs as well as instructor feedback;
  • students perceive AI feedback as useful;
  • students act on AI-generated feedback;
  • AI feedback reduces instructor workload; or
  • benefits persist beyond the immediate task.

A study may support one of these propositions while providing no evidence about the others. Tracking the specific claim prevents you from treating general relevance to a topic as evidence for every argument associated with that topic.

Distinguish Support From Mere Consistency

A finding may be consistent with an argument without providing strong evidence for it.

Suppose a small observational study reports that students who use a feedback tool more frequently tend to receive higher writing scores. That finding may be consistent with the proposition that the tool improves writing. But other explanations may remain plausible. More proficient or motivated students might simply use the tool more often.

How strongly a study bears on a claim depends partly on its design, measurement, analysis, and the alternative explanations it can reasonably address.

For that reason, do not use "support" as shorthand for "the result points in the same direction as my argument."

Consistent with the claim The finding points in a compatible direction, but the study may not provide a strong test of the proposition.
Provides evidence supporting the claim The study design and evidence bear meaningfully on the proposition, although the strength and scope of that support still require interpretation.

A Non-Significant Result Does Not Automatically Contradict an Argument

This distinction is especially important in quantitative research.

Suppose a study finds no statistically significant difference between two conditions. It may be tempting to record:

Contradicts.

That conclusion may be unwarranted. A non-significant result does not by itself establish equivalence, prove the absence of an effect, or demonstrate the opposite effect. Interpretation depends on matters such as the estimate, its uncertainty, study design, sample size, measurement, and the specific hypothesis being tested.

The American Statistical Association has cautioned against treating a threshold such as p >.05 as evidence that no effect exists. Statistical conclusions should instead consider the magnitude and uncertainty of estimates alongside the broader study context.

Your tracking system therefore needs more vocabulary than "supports" and "contradicts."

Use Categories That Preserve Different Relationships to the Claim

You do not need an elaborate taxonomy, but several distinctions can prevent premature simplification.

Relationship What It Means
Supports The evidence provides meaningful support for the specific claim under the conditions studied.
Consistent with The result aligns with the claim but does not provide a particularly strong or direct test of it.
Contradicts The evidence provides meaningful evidence inconsistent with the specific claim or supports an incompatible proposition.
Qualifies The general claim may hold, but only under narrower conditions than originally stated.
Complicates The evidence introduces distinctions, mechanisms, trade-offs, or mixed patterns that make a simple version of the claim inadequate.
No clear evidence The study does not provide sufficiently informative evidence for or against the claim.
Not applicable The study is relevant to the broader topic but does not actually test or address this particular claim.

These categories are analytical aids rather than universal methodological classifications. Adapt them to your project and define them consistently.

Record the Finding Before Recording Your Classification

If your matrix contains only the label supports, you may later forget why you assigned it.

A more defensible record includes at least:

  • the claim being evaluated;
  • the relevant finding;
  • your relationship classification;
  • important conditions or qualifications; and
  • a source locator that allows verification.

This preserves the evidence behind the label.

It also makes it easier to distinguish what the authors reported from your own interpretation. The authors may never have said that their study "contradicts" another paper. That may be a relationship you infer when synthesizing the literature.

Track Findings, Not Just Papers

One paper may contain several relevant results.

Imagine a study reporting that an intervention:

  • improved immediate performance;
  • did not clearly improve delayed retention;
  • increased student satisfaction; and
  • produced larger gains among one subgroup.

What is the paper's position?

There may be no useful single answer. It depends on the claim.

If your argument concerns immediate performance, the evidence may support it. If the argument concerns lasting learning, the evidence may be inconclusive. If your argument assumes comparable effects for all students, the subgroup result may complicate it.

When several findings matter, use separate entries or fields rather than forcing the entire paper into one category.

Record the Conditions Under Which the Finding Occurred

A qualified claim is often more informative than a broad one.

Instead of recording:

Study B contradicts Study A.

you might discover:

Study A reports improvement when the intervention is compared with no feedback, whereas Study B finds no clear advantage when it is compared with instructor feedback.

Now the disagreement looks different. The studies may not be testing the same comparative proposition.

Relevant conditions can include population, setting, intervention, comparison condition, duration, outcome definition, measurement instrument, study design, implementation, analytical approach, or theoretical framing.

This is one reason a well-designed literature matrix should contain analytically useful comparison fields rather than only bibliographic information and generic summaries.

Investigate Apparent Contradictions Before Treating Them as Genuine Disagreement

When two studies appear to conflict, compare them systematically.

Check the claim Are both studies actually addressing the same proposition?
Check the construct Are they defining and operationalizing the same concept?
Check the outcome Are they measuring the same thing?
Check the population and context Could the difference be conditional on who or where was studied?
Check the method Could design, measurement, implementation, or analysis contribute to the different findings?
Check uncertainty Are the estimates actually incompatible, or has one result simply crossed a statistical threshold while another has not?

The last point deserves particular attention. The American Statistical Association's guidance warns that scientific conclusions should not be based solely on whether a p-value passes a particular threshold. Two studies should not automatically be declared contradictory merely because one reports p <.05 and another reports p >.05.

Complication Is Often More Informative Than Contradiction

A study does not need to overturn an argument to change it substantially.

Suppose several studies suggest that collaborative learning improves student performance. A later study finds that the effect appears primarily when groups receive structured guidance.

That result may not simply "support" or "contradict" the earlier literature. It may qualify the proposition:

Collaborative learning may improve performance particularly when collaboration is appropriately structured.

This is intellectually valuable because the literature has moved from a broad claim toward a more conditional explanation.

Look for studies that reveal:

  • boundary conditions;
  • subgroup differences;
  • alternative mechanisms;
  • measurement dependencies;
  • contextual variation;
  • trade-offs; or
  • different interpretations of apparently similar evidence.

These studies often contribute more to a sophisticated synthesis than another paper that simply points in the expected direction.

Do Not Count Studies as If Every Study Casts One Equal Vote

Suppose seven studies appear to support an argument and three do not.

It is tempting to conclude that the literature is 7–3 in favor.

That arithmetic can be misleading. Studies may differ in design, sample, precision, risk of bias, measurement quality, relevance to the specific claim, and independence from one another. Several papers may even analyze overlapping datasets.

Watch Out

A literature review is not strengthened by treating the number of studies on each side as a substitute for evaluating the evidence. "Most studies found..." may sometimes be descriptively accurate, but the numerical majority alone does not establish which conclusion is best supported.

Formal systematic reviews and meta-analyses use explicit methods for evaluating and combining evidence. A narrative literature review should not imitate those methods informally by counting positive and negative studies and calling the result synthesis.

Let Your Categories Change as the Argument Becomes More Precise

Early in your review, you might track a broad proposition such as:

Technology X improves learning.

After comparing the literature, that proposition may become too crude. You may need to separate it into:

  • immediate achievement;
  • long-term retention;
  • student engagement;
  • particular learner groups;
  • specific implementation conditions; or
  • comparison with particular alternatives.

Your evidence-tracking system should evolve accordingly. Refining the claim is part of synthesis, not a failure to decide what you believed at the beginning.

04 · A Practical Example

What Looks Like Conflicting Evidence May Be a More Conditional Story

Hypothetical Example

Does AI-Generated Feedback Improve Student Writing?

Suppose you are reviewing five hypothetical studies relevant to the claim that AI-generated feedback improves university students' academic writing. Instead of assigning one label to each paper, you record the specific evidence and conditions.

Study Relevant Evidence Relationship to Claim Important Qualification
A Students receiving AI feedback produced stronger revisions than students receiving no feedback. Supports Comparison was no feedback, not instructor feedback.
B No clear difference in revision scores between AI and instructor feedback groups. Qualifies Does not establish that AI feedback is ineffective; it addresses comparative performance against instructor feedback.
C Students rated AI feedback highly but writing performance was not measured. Not applicable to performance claim Relevant to acceptability, not evidence of improved writing.
D Improvement appeared for lower-performing students but not clearly for higher-performing students. Complicates Possible subgroup difference.
E Immediate writing scores improved, but the difference was not evident at delayed assessment. Qualifies Potential distinction between immediate performance and retention.

A simple tally might have produced something like "two positive, one null, one mixed, one irrelevant." That tells you very little.

Initial argument AI-generated feedback improves university students' academic writing.
Compare the evidence Positive evidence is clearest against no feedback, while evidence against instructor feedback is less straightforward.
Notice the outcome distinction Student perceptions should not be treated as evidence of improved writing performance.
Notice the boundary conditions Effects may vary by learner characteristics and may differ between immediate performance and delayed outcomes.
Refine the synthesis The literature may support a narrower claim that AI-generated feedback can improve immediate writing performance under some conditions, while its comparative advantage over instructor feedback, durability, and consistency across learners remain less certain.

The final statement is more complicated than "the evidence supports AI feedback." That is precisely why it is more useful.

05 · What Researchers Often Get Wrong

Common Mistakes When Tracking Agreement and Disagreement

Misconception

Every Paper Is Either Supporting or Contradictory

Many studies do neither. A study may provide indirect evidence, address only one component of the claim, qualify it, reveal a boundary condition, or simply be relevant to the broader topic without testing the proposition you are evaluating.

Misconception

A Non-Significant Result Means the Study Found No Effect

Failure to cross a statistical significance threshold does not by itself demonstrate that an effect is absent. Examine the estimate, uncertainty, study design, sample, and the hypothesis being tested before deciding how the evidence bears on your claim.

Misconception

Studies With Opposite Conclusions Must Have Contradictory Evidence

Authors may frame conclusions differently even when underlying results are compatible, and apparently conflicting studies may examine different populations, outcomes, conditions, or constructs. Compare the evidence itself before classifying the relationship.

Misconception

The Side With More Studies Has Stronger Evidence

Study counts ignore differences in relevance, design, precision, quality, sample size, and independence. Five weak or only indirectly relevant studies do not automatically outweigh two studies that provide stronger evidence about the specific proposition.

Misconception

I Should Track Whether a Study Supports My Thesis

Framing the task this way can encourage confirmation bias. Track how the study bears on a defined claim regardless of whether the result helps the argument you expected to make. A literature review should allow the argument to change when the evidence requires it.

Misconception

Mixed Findings Are Just Messy Evidence

Mixed findings may reveal theoretically or practically important distinctions. Different outcomes, populations, contexts, measurements, or time points can expose the conditions under which a proposition does and does not hold. Investigate the mixture rather than treating it as an inconvenient residual category.

06 · What This Means for You

Build an Evidence Map Around Claims Rather Than Camps

When you encounter a study, resist asking only, Does this support my argument?

Ask instead:

What precise claim does this study provide evidence about, what does the evidence show, and what conditions limit what I can infer from it?

This change in question reduces the temptation to sort papers prematurely into favorable and unfavorable piles.

A simple decision framework

If the study directly addresses the claim and its evidence is consistent with it
Record the evidence and classify it as supporting, with appropriate methodological and contextual qualifications.
If the result points in the same direction but provides only indirect or limited evidence
Consider "consistent with" rather than overstating it as strong support.
If the evidence is genuinely incompatible with the claim
Record it as contradictory and examine why it differs from supporting evidence.
If the claim appears to hold only for particular populations, contexts, outcomes, or conditions
Record the study as qualifying the claim and revise the proposition accordingly.
If the study introduces several different patterns or mechanisms
Record how it complicates the argument rather than forcing it into one directional category.
If the study is about the general topic but does not provide evidence about the specific proposition
Mark it as not applicable to that claim rather than treating topical relevance as evidential support.

You can implement this system in a spreadsheet, database, or literature matrix designed for cross-source comparison. Whatever tool you use, preserve the finding behind the classification and maintain a route back to the original source.

As patterns emerge, you may discover that the strongest organizational structure for the written review follows the evidence itself. In that case, organizing parts of the literature around findings or the conditions explaining those findings may make the emerging argument easier for readers to follow.

07 · A Quick Checklist

Before Labeling a Study as Supporting or Contradictory, Check the Relationship

Before classifying the evidence, check:
Have I written the specific claim that I am evaluating rather than relying on a broad topic or thesis?
Does this study actually provide evidence about that claim?
Have I recorded the relevant finding before adding my support, contradiction, qualification, or complication label?
Am I distinguishing evidence consistent with a claim from evidence that provides a strong test of it?
If studies appear contradictory, have I checked whether they use comparable constructs, populations, methods, outcomes, and conditions?
Have I avoided interpreting statistical significance alone as proof that two studies agree or disagree?
Can one paper receive different classifications for different findings or claims when necessary?
Have I recorded boundary conditions or qualifications that could make the emerging argument more precise?
Am I evaluating evidence rather than simply counting how many studies appear on each side?
Can I trace each important classification back to the exact evidence in the original source?
08 · Frequently Asked Questions

Questions About Tracking Supporting and Contradictory Evidence

What should I do when a study partly supports and partly contradicts my argument?

Separate the relevant findings and connect each to the specific claim it addresses. A paper does not need one overall classification. One result may support a proposition while another qualifies or contradicts a different part of your argument.

Does a non-significant result contradict a study that found a significant effect?

Not necessarily. Statistical significance alone is not an appropriate test of whether two results are incompatible. Compare the effect estimates, uncertainty, designs, samples, measurements, and other relevant conditions before interpreting the studies as contradictory.

Should I count how many studies support each side?

A descriptive count may sometimes be informative, but it should not substitute for evaluating the evidence. Studies can differ substantially in design, relevance, precision, quality, and independence. A simple majority does not by itself establish the better-supported conclusion.

What is the difference between a study that contradicts and one that complicates an argument?

A contradictory study provides evidence that is meaningfully inconsistent with a particular claim. A complicating study shows that the claim is too simple, perhaps because outcomes vary across populations, conditions, measures, mechanisms, or dimensions of the phenomenon.

What does it mean for a study to qualify an argument?

Qualification narrows the conditions under which a claim appears defensible. Instead of concluding that an intervention works generally, for example, the evidence may suggest that it works primarily for a particular population, outcome, implementation approach, or comparison condition.

Should I include studies that contradict my expected argument?

Yes, when they are relevant and meet the criteria for the literature you are reviewing. Excluding inconvenient evidence can distort the synthesis. Contradictory studies may also reveal limitations, boundary conditions, methodological differences, or alternative explanations that substantially improve the eventual argument.

How should I record a study that is relevant but does not test my claim?

Mark it as relevant to the broader topic but not applicable to that particular claim. For example, a study of students' satisfaction with an intervention does not by itself provide evidence that the intervention improves learning outcomes.

Can my classification of a study change later?

Yes. As your claims become more precise and you understand the literature better, a study that initially looked contradictory may turn out to examine a different condition, while a supposedly supportive study may provide only indirect evidence. Revising classifications is a normal part of synthesis when the reason for the change is documented.

09 · The Bottom Line

Do Not Ask Which Side a Paper Is On

The Bottom Line

Track evidence by connecting specific findings to specific claims, then record whether those findings support, contradict, qualify, complicate, or provide no clear evidence about the proposition under the conditions actually studied.

The objective is not to accumulate papers on opposing sides of an argument. It is to understand why evidence converges or diverges and to let those relationships make your eventual claim more precise. A strong synthesis often replaces a simple yes-or-no conclusion with a better explanation of what appears to hold, for whom, under what conditions, and with what remaining uncertainty.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes