03 · What You Need to Know
Separate the Empirical Question From the Decision It Is Supposed to Inform
The distinction between evidence and recommendation is especially visible in formal evidence-to-decision frameworks. Approaches such as GRADE do not move directly from “the research found an effect” to “therefore this option should be recommended.” They consider the evidence alongside additional criteria such as the balance of desirable and undesirable consequences, certainty of evidence, values and preferences, resources, equity, acceptability, and feasibility.
These frameworks were developed primarily for health and policy decision-making, so their specific procedures should not be imported mechanically into every discipline. The underlying lesson is broader: a recommendation is a judgment about what should be done, whereas an empirical finding describes, estimates, explains, interprets, or evaluates something about the world.
Look for Normative Language in the Question
Words such as “should,” “ought,” “best,” “appropriate,” “recommended,” and “preferable” often indicate that the question is asking for a decision rather than merely evidence.
For example:
- Should universities prohibit generative AI in examinations?
- What teaching strategy should instructors use?
- Which learning-management system should the university adopt?
- What policy should schools implement to regulate student AI use?
- Should researchers use generative AI when writing manuscripts?
These are legitimate questions. The issue is that the word “should” cannot usually be resolved by measuring one outcome and allowing that outcome to dictate the answer.
Ask What Empirical Claim Is Hidden Inside the Recommendation Question
Suppose the question is: “Should universities use AI tutors to improve student learning?”
Several empirical questions may sit underneath it. Do AI tutors improve specified learning outcomes relative to a relevant alternative? For which students? Under what conditions? What unintended consequences occur? How much does implementation cost? Are students and instructors willing and able to use the system?
Before trying to answer the recommendation, identify which of these uncertainties your study is actually capable of investigating.
Evidence About Effectiveness Is Not the Same as a Recommendation
Suppose a rigorous study finds that an intervention improves examination performance. It is tempting to conclude: “Therefore, universities should adopt it.”
That conclusion may be reasonable, but it does not follow from effectiveness alone. The improvement might be very small. The intervention could be expensive. It might create accessibility problems, introduce privacy risks, increase instructor workload, or produce other outcomes that matter to the decision.
Evidence-to-decision frameworks formalize this point by treating the balance of consequences and other decision criteria separately from evidence that an intervention produces a particular effect.
A Statistically Significant Difference Does Not Answer “Which Is Better?”
Suppose students using Platform A score significantly higher than students using Platform B. Can the researcher conclude that universities should choose Platform A?
Not necessarily. The magnitude of the difference matters, as do uncertainty, cost, usability, accessibility, implementation requirements, privacy, technical infrastructure, and the outcomes that stakeholders consider important.
“Which option produces a higher score on this outcome?” and “Which option should we choose?” are different questions.
Recommendations Require Values Because Outcomes Have to Be Weighted
Decisions often involve competing outcomes. An intervention might improve learning slightly while increasing student workload. A technology might save instructors time while introducing privacy concerns. A policy might reduce one form of misconduct while restricting legitimate educational uses.
Empirical research can estimate or describe these consequences. It cannot, by itself, determine how much importance should be assigned to each one.
Formal evidence-to-decision approaches therefore consider values and preferences explicitly. The eventual recommendation depends partly on how relevant stakeholders value different outcomes and trade-offs.
The “Best” Option Depends on the Criterion for Best
“Which teaching method is best?” appears empirical until you ask what “best” means.
Highest examination scores? Strongest long-term retention? Lowest cost? Greatest student satisfaction? Least instructor workload? Best accessibility? Most equitable outcomes?
Different criteria can produce different answers. Unless one criterion has been justified as decisive, “best” compresses several evaluative judgments into a single word.
A more researchable question may compare clearly specified outcomes rather than assuming that one outcome determines overall superiority.
Feasibility Can Change a Recommendation Without Changing the Evidence of Effect
An intervention may work under controlled conditions yet be difficult to implement in the intended setting. Required expertise may be unavailable. Infrastructure may be inadequate. Costs may exceed available resources. Faculty or students may find the approach unacceptable.
These conditions do not necessarily change whether the intervention can produce an effect. They can change whether adopting it is a sensible recommendation in that context.
This is one reason evidence-to-decision frameworks explicitly include feasibility and acceptability among their considerations.
Equity Can Matter Even When Average Outcomes Improve
An intervention can improve the average outcome while distributing benefits and burdens unevenly.
For example, an AI-based educational system might benefit students with reliable devices and high-speed internet while creating additional barriers for students with limited connectivity or accessibility needs. An average positive effect does not automatically resolve whether implementation would worsen or reduce inequities.
If the ultimate question is what an institution should do, distributional consequences may be relevant alongside the average effect.
The Recommendation May Change Across Contexts
The same evidence can support different decisions in different settings because resources, alternatives, infrastructure, stakeholder values, institutional missions, legal constraints, and implementation conditions differ.
A recommendation that makes sense for a well-resourced research university may not be appropriate for a small institution with limited technical support. The empirical evidence need not be contradictory for the decisions to differ.
This is why recommendations should usually make their context and assumptions visible rather than presenting one decision as universally compelled by the evidence.
Some Research Questions Are Intentionally Decision-Oriented
Not every “should” question needs to be rewritten into a narrow empirical question. Policy research, health technology assessment, implementation research, program evaluation, operations research, and other decision-oriented traditions may explicitly aim to support recommendations.
In such cases, the study should acknowledge that the recommendation requires an evidence-to-decision process rather than pretending that one empirical estimate supplies the answer automatically.
The methodological task then becomes broader: identify the decision, alternatives, relevant outcomes, evidence sources, stakeholders, values, constraints, and criteria by which the options will be judged.
A Single Study Rarely Supplies Every Input to a Major Recommendation
A researcher may investigate one important part of a decision without claiming to settle the whole decision.
For example, a randomized study might estimate the effect of AI tutoring on learning. A qualitative study might examine student acceptability. An economic analysis might estimate costs. Accessibility research might identify barriers for particular learners. Policy analysis might examine legal and governance implications.
A recommendation can synthesize these forms of evidence, but each study should remain clear about what it contributes.
Do Not Smuggle a Recommendation Into an Empirical Conclusion
A common progression looks like this:
“Students who used the intervention achieved higher scores. Therefore, universities should implement the intervention.”
The first sentence reports an empirical finding. The second introduces a recommendation. Between them lies an unstated decision process.
That process may ultimately support the recommendation, but the researcher should make the reasoning visible: How large was the benefit? How certain is the evidence? What disadvantages were observed? What alternatives exist? What resources are required? Who benefits or bears the burden?
Recommendation Questions Can Also Conceal Several Questions at Once
“What AI policy should universities adopt?” may require evidence about learning, academic integrity, privacy, accessibility, assessment, faculty workload, student practices, costs, governance, and implementation.
If one study attempts to answer all of those issues under a single recommendation question, it may be combining several distinct research questions into one.
The better approach may be to identify the specific empirical uncertainty the present study can address while treating the broader recommendation as the decision that the evidence will eventually inform.
Ask Whether Your Study Could Produce the Evidence Needed for the Recommendation
Even after the decision criteria are identified, the proposed study may address only some of them.
A survey of faculty attitudes can provide evidence about reported acceptability. It cannot by itself establish effectiveness, cost-effectiveness, student outcomes, or institutional feasibility. An experiment estimating learning outcomes cannot automatically answer questions about long-term implementation or stakeholder values.
This makes it essential to check whether the recommendation requires evidence that the proposed study could never produce.