03 · What You Need to Know
The Unit You Need to Track Is the Study-to-Claim Relationship
Suppose your emerging argument is:
AI-generated feedback improves university students' academic writing.
It may be tempting to read each paper and assign one label:
- supports;
- contradicts; or
- mixed.
But the apparent simplicity hides several questions. What does "improves" mean? Compared with what? Which students? What kind of writing? What kind of AI-generated feedback? Was improvement measured immediately or later? Does a study measuring student satisfaction provide evidence about writing improvement at all?
Before classifying evidence, you therefore need to know exactly what proposition you are asking the evidence to support.
Start by Breaking Broad Arguments Into Trackable Claims
Broad arguments often contain several propositions that should be evaluated separately.
For example, the broad statement:
AI feedback is effective in higher education.
could conceal very different claims:
- AI feedback improves revision quality compared with no feedback;
- AI feedback performs as well as instructor feedback;
- students perceive AI feedback as useful;
- students act on AI-generated feedback;
- AI feedback reduces instructor workload; or
- benefits persist beyond the immediate task.
A study may support one of these propositions while providing no evidence about the others. Tracking the specific claim prevents you from treating general relevance to a topic as evidence for every argument associated with that topic.
Distinguish Support From Mere Consistency
A finding may be consistent with an argument without providing strong evidence for it.
Suppose a small observational study reports that students who use a feedback tool more frequently tend to receive higher writing scores. That finding may be consistent with the proposition that the tool improves writing. But other explanations may remain plausible. More proficient or motivated students might simply use the tool more often.
How strongly a study bears on a claim depends partly on its design, measurement, analysis, and the alternative explanations it can reasonably address.
For that reason, do not use "support" as shorthand for "the result points in the same direction as my argument."
Consistent with the claim
The finding points in a compatible direction, but the study may not provide a strong test of the proposition.
Provides evidence supporting the claim
The study design and evidence bear meaningfully on the proposition, although the strength and scope of that support still require interpretation.
A Non-Significant Result Does Not Automatically Contradict an Argument
This distinction is especially important in quantitative research.
Suppose a study finds no statistically significant difference between two conditions. It may be tempting to record:
Contradicts.
That conclusion may be unwarranted. A non-significant result does not by itself establish equivalence, prove the absence of an effect, or demonstrate the opposite effect. Interpretation depends on matters such as the estimate, its uncertainty, study design, sample size, measurement, and the specific hypothesis being tested.
The American Statistical Association has cautioned against treating a threshold such as p >.05 as evidence that no effect exists. Statistical conclusions should instead consider the magnitude and uncertainty of estimates alongside the broader study context.
Your tracking system therefore needs more vocabulary than "supports" and "contradicts."
Use Categories That Preserve Different Relationships to the Claim
You do not need an elaborate taxonomy, but several distinctions can prevent premature simplification.
| Relationship |
What It Means |
| Supports |
The evidence provides meaningful support for the specific claim under the conditions studied. |
| Consistent with |
The result aligns with the claim but does not provide a particularly strong or direct test of it. |
| Contradicts |
The evidence provides meaningful evidence inconsistent with the specific claim or supports an incompatible proposition. |
| Qualifies |
The general claim may hold, but only under narrower conditions than originally stated. |
| Complicates |
The evidence introduces distinctions, mechanisms, trade-offs, or mixed patterns that make a simple version of the claim inadequate. |
| No clear evidence |
The study does not provide sufficiently informative evidence for or against the claim. |
| Not applicable |
The study is relevant to the broader topic but does not actually test or address this particular claim. |
These categories are analytical aids rather than universal methodological classifications. Adapt them to your project and define them consistently.
Record the Finding Before Recording Your Classification
If your matrix contains only the label supports, you may later forget why you assigned it.
A more defensible record includes at least:
- the claim being evaluated;
- the relevant finding;
- your relationship classification;
- important conditions or qualifications; and
- a source locator that allows verification.
This preserves the evidence behind the label.
It also makes it easier to distinguish what the authors reported from your own interpretation. The authors may never have said that their study "contradicts" another paper. That may be a relationship you infer when synthesizing the literature.
Track Findings, Not Just Papers
One paper may contain several relevant results.
Imagine a study reporting that an intervention:
- improved immediate performance;
- did not clearly improve delayed retention;
- increased student satisfaction; and
- produced larger gains among one subgroup.
What is the paper's position?
There may be no useful single answer. It depends on the claim.
If your argument concerns immediate performance, the evidence may support it. If the argument concerns lasting learning, the evidence may be inconclusive. If your argument assumes comparable effects for all students, the subgroup result may complicate it.
When several findings matter, use separate entries or fields rather than forcing the entire paper into one category.
Record the Conditions Under Which the Finding Occurred
A qualified claim is often more informative than a broad one.
Instead of recording:
Study B contradicts Study A.
you might discover:
Study A reports improvement when the intervention is compared with no feedback, whereas Study B finds no clear advantage when it is compared with instructor feedback.
Now the disagreement looks different. The studies may not be testing the same comparative proposition.
Relevant conditions can include population, setting, intervention, comparison condition, duration, outcome definition, measurement instrument, study design, implementation, analytical approach, or theoretical framing.
This is one reason a well-designed literature matrix should contain analytically useful comparison fields rather than only bibliographic information and generic summaries.
Investigate Apparent Contradictions Before Treating Them as Genuine Disagreement
When two studies appear to conflict, compare them systematically.
Check the claim Are both studies actually addressing the same proposition?
Check the construct Are they defining and operationalizing the same concept?
Check the outcome Are they measuring the same thing?
Check the population and context Could the difference be conditional on who or where was studied?
Check the method Could design, measurement, implementation, or analysis contribute to the different findings?
Check uncertainty Are the estimates actually incompatible, or has one result simply crossed a statistical threshold while another has not?
The last point deserves particular attention. The American Statistical Association's guidance warns that scientific conclusions should not be based solely on whether a p-value passes a particular threshold. Two studies should not automatically be declared contradictory merely because one reports p <.05 and another reports p >.05.
Complication Is Often More Informative Than Contradiction
A study does not need to overturn an argument to change it substantially.
Suppose several studies suggest that collaborative learning improves student performance. A later study finds that the effect appears primarily when groups receive structured guidance.
That result may not simply "support" or "contradict" the earlier literature. It may qualify the proposition:
Collaborative learning may improve performance particularly when collaboration is appropriately structured.
This is intellectually valuable because the literature has moved from a broad claim toward a more conditional explanation.
Look for studies that reveal:
- boundary conditions;
- subgroup differences;
- alternative mechanisms;
- measurement dependencies;
- contextual variation;
- trade-offs; or
- different interpretations of apparently similar evidence.
These studies often contribute more to a sophisticated synthesis than another paper that simply points in the expected direction.
Do Not Count Studies as If Every Study Casts One Equal Vote
Suppose seven studies appear to support an argument and three do not.
It is tempting to conclude that the literature is 7–3 in favor.
That arithmetic can be misleading. Studies may differ in design, sample, precision, risk of bias, measurement quality, relevance to the specific claim, and independence from one another. Several papers may even analyze overlapping datasets.
Watch Out
A literature review is not strengthened by treating the number of studies on each side as a substitute for evaluating the evidence. "Most studies found..." may sometimes be descriptively accurate, but the numerical majority alone does not establish which conclusion is best supported.
Formal systematic reviews and meta-analyses use explicit methods for evaluating and combining evidence. A narrative literature review should not imitate those methods informally by counting positive and negative studies and calling the result synthesis.
Let Your Categories Change as the Argument Becomes More Precise
Early in your review, you might track a broad proposition such as:
Technology X improves learning.
After comparing the literature, that proposition may become too crude. You may need to separate it into:
- immediate achievement;
- long-term retention;
- student engagement;
- particular learner groups;
- specific implementation conditions; or
- comparison with particular alternatives.
Your evidence-tracking system should evolve accordingly. Refining the claim is part of synthesis, not a failure to decide what you believed at the beginning.