03 · What You Need to Know
Extract Enough to Reconstruct the Evidence, Not the Entire Paper
Start with the question your extraction needs to answer
Data extraction is selective by design. You are converting a much larger research report into a structured record containing the information needed for a particular purpose.
That purpose matters. In systematic reviews, Cochrane guidance emphasizes that reviewers collect information about study design, risk of bias, and results so that evidence can be presented and synthesized accurately. PRISMA 2020 likewise requires reviewers to define the outcomes and other variables for which data were sought and to explain how they selected results when several compatible measures, time points, or analyses were available.
The same logic is useful outside formal evidence synthesis. Before deciding what belongs in your notes or literature matrix, ask:
What will I need to know later in order to compare, interpret, cite, or use this study correctly?
Your answer determines the extraction fields.
A useful core record has several recurring elements
Although the exact fields should vary, most empirical studies can be reconstructed from a relatively compact set of information.
| Information to extract |
What it tells you |
Why you may need it later |
| Research question or objective |
What the study was trying to find out |
Prevents you from using the paper for a question it did not actually address |
| Context or setting |
Where, when, or under what conditions the research occurred |
Helps you judge applicability and compare contexts |
| Study design |
How the research was structured |
Shapes what kinds of inference the evidence can support |
| Sample or data source |
Who or what contributed the evidence |
Helps assess relevance, comparability, and scope |
| Key variables, constructs, phenomena, or intervention |
What was actually examined |
Prevents superficially similar studies from being treated as identical |
| Measures, procedures, or data collection |
How the central concepts became evidence |
Helps interpret what the findings really represent |
| Analytical approach |
How the researchers turned the data into findings |
Helps you understand what the reported result means |
| Relevant findings |
What the study actually reported |
Provides the evidence you may compare, synthesize, or cite |
| Important limitations |
What constrains interpretation |
Prevents conclusions from becoming stronger or broader than the evidence |
| Relevance to your research |
Why you saved the paper |
Lets your future self recover its role without rereading it |
This is a working framework rather than a universal reporting standard. Some studies require additional information, while several fields may be unnecessary for a particular purpose.
Record the research question in your own words
Start with what the study was designed to investigate.
A title can be broader than the actual study. An abstract can compress several objectives into one paragraph. Your extraction should make the target of the research explicit.
For example:
Research question: Does repeated automated formative feedback improve undergraduate students' writing performance compared with conventional instructor feedback?
That statement gives the rest of the extraction record a reference point. You can now ask whether the design, measures, analysis, and conclusions actually address that question.
If the paper contains several research questions, extract the ones relevant to your purpose rather than automatically copying every objective.
Record enough context to know where the evidence came from
Research findings occur somewhere.
Depending on your topic, useful contextual information may include country, institution type, educational level, clinical setting, industry, laboratory conditions, online environment, policy context, historical period, or another feature that affects interpretation.
Cochrane guidance for qualitative evidence synthesis explicitly recommends preserving contextual information during extraction because removing findings from the setting in which they were produced can lead to misinterpretation.
The same caution applies more broadly. A finding from first-year engineering students at one institution should not quietly become evidence about "all university students" simply because the contextual details disappeared from your notes.
Extract context when it could affect comparison, transferability, interpretation, or the scope of the claim you intend to make.
Record the study design before recording what the study supposedly proves
Study design shapes inference.
Record whether the study is, for example, a randomized experiment, quasi-experiment, cohort study, cross-sectional survey, case-control study, qualitative interview study, ethnography, case study, mixed-methods study, secondary-data analysis, systematic review, or another design appropriate to the field.
You do not need to force every study into a familiar label when its design does not fit neatly. Record enough methodological description to remember how the evidence was generated.
This becomes especially important when you later synthesize findings. A randomized comparison and a cross-sectional association may address related topics, but they do not automatically carry the same inferential meaning.
Record who or what produced the data
Sample information is easy to reduce to one number, such as N = 312. That is rarely enough.
Depending on the study, record characteristics that matter to your research question: age or educational level, profession, diagnosis, geographic setting, relevant demographic characteristics, eligibility criteria, sampling strategy, organizational characteristics, document type, dataset, or other features.
For example:
Sample: 312 first-year undergraduates from two universities; voluntary participation; 61% women.
Whether each of those details belongs in your extraction depends on your research question. If gender is irrelevant to your synthesis, collecting it from every study merely because it is reported may add work without improving interpretation.
Formal evidence synthesis makes this principle explicit. PRISMA 2020 asks review authors to define the variables they sought, including relevant participant and intervention characteristics, rather than treating every reported variable as automatically necessary.
Record how the central concept was actually measured
This is one of the most important extraction fields because papers can use the same conceptual label for different operational definitions.
Suppose several studies investigate "academic performance." One uses final course grades, another a standardized test, another an instructor-scored writing task, and another students' self-reported perceptions of learning.
If your extraction sheet records only "academic performance," those studies may look more comparable than they really are.
Concept
Academic performance
Operational measure
Score on a 40-item standardized achievement test administered immediately after the intervention
Record the second when measurement differences matter.
The same principle applies to exposures, interventions, predictors, qualitative phenomena, and other constructs. What did the researchers actually observe, ask, manipulate, code, or calculate?
For interventions, extract what participants actually received
Labels such as "blended learning," "AI feedback," "simulation," or "peer instruction" can conceal considerable variation.
If the intervention matters to your review or research, record enough detail to distinguish implementations. This may include duration, frequency, intensity, delivery mode, provider, components, comparison condition, and other features relevant to your question.
A study in which students use an AI feedback system once for 15 minutes is not automatically testing the same intervention as a semester-long system integrated into every writing assignment.
If you later observe conflicting findings, these details may help explain why.
Record the analytical approach at the level you will need later
You do not necessarily need to copy the entire statistical analysis plan or qualitative coding procedure.
Record the analysis that produced the finding you care about.
For a quantitative study, that may mean noting a group comparison, regression model, multilevel model, survival analysis, structural equation model, or another relevant procedure. If adjustment matters, record the key covariates or the fact that the result is adjusted.
For qualitative research, extraction may involve the analytical approach, how coding or interpretation proceeded, theoretical frameworks used, and the relationship between participant data and author-generated themes. Cochrane's qualitative evidence guidance notes that extracted findings can include participant quotations, author-developed themes, explanations, hypotheses, theory, and interpretations, depending on the synthesis method.
If you do not understand an analysis that is central to the finding, do not merely copy its name into your extraction sheet. Resolve enough of the methodological uncertainty to know what the resulting evidence means.
Extract the actual result, not merely “significant” or “no effect”
A useful extraction preserves the substance of the finding.
For quantitative research, record the effect estimate or relevant numerical result when it matters, together with an appropriate measure of uncertainty such as a confidence interval when reported and relevant to your use.
For example:
Weak extraction: AI feedback significantly improved writing.
Better extraction: Intervention group scored an estimated 4.2 points higher than comparison group on the writing assessment; 95% CI [1.3, 7.1].
The second record preserves magnitude and uncertainty rather than only a significance label.
Depending on the analysis, you might extract a mean difference, standardized effect, correlation, regression coefficient, risk ratio, odds ratio, prevalence estimate, or another quantity. If the statistics are unfamiliar, focus on what was estimated, its magnitude and direction, and its uncertainty.
Define which result you want before a paper gives you several choices
One study may report the same broad outcome in several ways: multiple measurement instruments, time points, adjusted and unadjusted analyses, subgroup estimates, alternative model specifications, or several related outcomes.
If you simply extract whichever result looks most interesting, your decisions can become inconsistent across studies.
This problem is sufficiently important that PRISMA 2020 asks systematic reviewers to report how outcomes were defined and how they selected which results to collect when multiple compatible results were available.
Even in a less formal literature review, decide what matters to your question. Are you interested in immediate post-intervention performance or long-term retention? Adjusted or unadjusted estimates? Total scale score or a particular subscale?
The decision should come from your research purpose rather than from which result happens to produce the most convenient conclusion.
Watch Out
If a study reports many outcomes, time points, subgroups, or analyses, do not silently select the result that best fits your developing argument. Define what evidence is relevant to your question and preserve important results that complicate the interpretation.
Record null and contradictory findings when they matter
An extraction record can become biased even when every number in it is technically correct.
Suppose a paper reports improvements on two outcomes but no clear difference on three others. If you extract only the positive results, your later synthesis may make the study appear more consistently favorable than it actually was.
Record findings according to the outcomes and questions relevant to your project, not according to whether the results are exciting.
The same applies to subgroup analyses, sensitivity analyses, qualitative exceptions, negative cases, and findings that complicate the authors' preferred interpretation when those results are consequential to your synthesis.
Keep the authors’ interpretation separate from the finding
It can be useful to extract what the authors believe their results mean, but store that separately from the evidence itself.
| Extraction field |
Example |
| Reported finding |
Students receiving AI-assisted feedback revised their drafts more frequently and achieved higher final writing scores. |
| Authors' interpretation |
Authors propose that increased revision behavior explains the improvement. |
| Your appraisal |
Revision frequency was observed, but the analysis did not directly establish it as the mechanism producing the score difference. |
This prevents the discussion from quietly becoming the result. When the interpretation is consequential, compare it with the study design and evidence before adopting it.
Extract limitations that affect your use, not every caveat in the discussion
A paper may list numerous limitations. Your extraction should prioritize those that change how you interpret, compare, or apply the study.
Examples might include a highly selected sample, substantial attrition, reliance on self-report for a key construct, short follow-up, lack of an appropriate comparison group, important missing data, limited contextual information, or analytical uncertainty.
Record limitations identified by the authors and consequential limitations you identify through your own reading.
A useful field is:
Limitation for my use: Outcome measured immediately after intervention; cannot support claims about long-term retention.
That is more actionable than copying a paragraph headed "Limitations."
Keep risk of bias or methodological appraisal separate from basic study characteristics
Study extraction and critical appraisal are related but not identical.
Recording that a study used a randomized design describes a study characteristic. Determining whether the randomization process creates a low or high risk of bias requires appraisal using appropriate criteria.
Extraction
What was done, who or what was studied, and what was found?
Appraisal
How much confidence should I place in the evidence, given potential methodological limitations or biases?
Formal reviews often maintain both. PRISMA 2020 treats data items and study risk-of-bias assessment as distinct methodological elements, while Cochrane provides separate guidance for collecting study data and evaluating bias.
You may store both in the same spreadsheet or review system, but keep the concepts distinct.
Always extract why the paper matters to your project
Formal systematic-review extraction forms are designed around predefined review variables. Ordinary scholarly reading needs an additional field that is easy to overlook:
Why did I save this?
Your answer might be:
- central evidence supporting my second argument;
- direct contradiction of the dominant finding;
- potential instrument for my own study;
- useful methodological precedent;
- foundational theoretical source;
- same intervention but different population;
- possible explanation for heterogeneity across studies;
- important limitation to the claim I am developing.
This is not a characteristic of the study itself. It is metadata about the relationship between the study and your research.
That relationship is often what you most need when you return months later.
Record where important information came from
An extraction sheet should help you return to the source quickly.
For consequential information, record a page, table, figure, supplementary file, section, or other locator where practical.
For example:
Primary outcome: Adjusted mean difference 4.2, 95% CI [1.3, 7.1] (Table 3).
Attrition: 18% intervention, 7% comparison (participant flow diagram).
Mechanism claim: Authors propose increased revision frequency as explanation (Discussion, p. 12).
These locators make later verification considerably easier and reduce the temptation to cite from your extraction sheet without returning to the original source when precision matters.
Do not extract information from memory
Populate the extraction record while the paper and relevant supplementary materials are available.
Memory tends to preserve the gist while losing exactly the details that distinguish studies: sample restrictions, measurement differences, comparison conditions, time points, analytical adjustments, and uncertainty.
If you later discover that a field was not extracted, return to the source rather than filling it from recollection.
For formal evidence synthesis, accuracy is sufficiently consequential that review guidance recommends processes such as standardized forms, piloting, independent extraction, checking by another reviewer, and procedures for resolving disagreements. PRISMA 2020 asks review authors to report how many reviewers collected data and whether they worked independently.
Use a consistent extraction structure when you need to compare studies
Free-form notes work well for exploratory reading. They become less useful when you need to compare twenty studies on the same dimensions.
At that point, a structured literature matrix or extraction form can help.
Rows might represent studies and columns the variables you need to compare. The fields should follow your research question rather than an enormous generic template.
For example:
| Study |
Design |
Sample |
Intervention or exposure |
Outcome measure |
Relevant finding |
Key limitation |
| Study A |
Randomized experiment |
First-year undergraduates |
AI feedback after each draft |
Rubric-scored writing task |
Higher post-test score |
No delayed follow-up |
| Study B |
Cross-sectional survey |
University students |
Self-reported AI use frequency |
Self-reported learning |
Positive association |
Cannot establish causal effect |
Placed side by side, the studies are no longer simply "two papers showing positive effects." Their evidence differs in ways that matter.
Pilot your extraction fields before committing to them
If you are extracting information from many studies, test your form on a small number first.
You may discover that a field is consistently ambiguous, that two fields should be separated, that an important variable is missing, or that you are collecting information you never use.
Guidance for systematic reviews has long recommended developing and pilot-testing data extraction forms before full extraction begins. The purpose is not bureaucratic neatness. It is to improve consistency and identify problems while they are still inexpensive to fix.
Draft Define the fields you think your research question requires.
Pilot Extract several studies that differ in design or reporting.
Revise Clarify ambiguous definitions, add missing fields, and remove information that does not serve the synthesis.
Apply consistently Use the revised structure across the remaining studies.
Define fields precisely enough that future-you knows what belongs there
A column labeled "Results" invites inconsistency. One day you may enter a p-value, another day a narrative conclusion, and later an effect estimate.
More specific fields can help:
- Primary outcome definition
- Outcome time point
- Effect estimate
- Uncertainty
- Authors' interpretation
- Limitation relevant to review question
You do not need all of these for every project. The principle is to define what a field means before you populate it repeatedly.
PRISMA 2020 similarly requires systematic reviewers to list and define the outcomes and other variables for which data were sought and to describe assumptions made about unclear or missing information.
Do not silently guess when information is missing
Sometimes a paper simply does not report what your extraction form asks for.
Do not convert "not reported" into your best guess.
Use a consistent marker such as "NR" for not reported, "unclear," or another code whose meaning you have defined. If you infer something from surrounding information, label it as an inference rather than as reported fact.
In formal evidence synthesis, missing or unclear information may warrant checking other reports of the same study or contacting study investigators. PRISMA specifically asks reviewers to report processes used to obtain or confirm relevant data from investigators.
For ordinary reading, contacting authors will rarely be necessary. The principle remains useful: distinguish what the study reports from what you think probably happened.
One study can have more than one report
This matters particularly in formal reviews.
A journal article, conference abstract, protocol, trial registry entry, dissertation, follow-up article, or secondary analysis may all refer to the same underlying study. Treating each report as an independent study can duplicate participants or evidence.
PRISMA 2020 explicitly distinguishes studies from reports. When your project requires rigorous evidence synthesis, identify whether multiple publications describe the same underlying research and decide how information across those reports will be handled.
For ordinary literature notes, a simpler practice may be enough: record related publications when you notice them so that you do not later mistake them for completely independent evidence.
Qualitative studies may require a different extraction logic
A matrix designed around sample size, effect estimate, and confidence interval will not adequately capture many qualitative studies.
Depending on the synthesis purpose, relevant extraction may include context, participants, methodological orientation, data collection, analytical approach, author-developed themes or findings, participant quotations, theoretical interpretations, and other information necessary to preserve meaning.
Cochrane's guidance for qualitative evidence synthesis stresses that contextual and methodological information should be extracted and that findings may occur throughout the report rather than in one neatly bounded results section.
Do not force every research design into a quantitative extraction template simply because spreadsheets prefer obedient columns.
Extraction should make synthesis easier
The point of extraction is not to build an impressive spreadsheet. It is to make later reasoning possible.
Once several studies are extracted consistently, you can ask questions that are difficult to answer while flipping among PDFs:
- Are studies actually measuring the same outcome?
- Do findings differ according to population or context?
- Are contradictory results associated with different research designs?
- Which methods recur across the literature?
- Where is evidence strong, sparse, or methodologically limited?
- Which papers provide genuinely independent evidence?
Those questions move you from collecting papers toward understanding a body of literature.
This is also where good extraction differs from simply taking extensive notes on each paper. Extraction creates comparable information across studies; notes can preserve broader ideas, reactions, questions, and connections that do not fit neatly into standardized fields.