03 · What You Need to Know
Different Designs Can Approach the Same Question Through Different Evidence
A research question specifies a problem, not always one unique route to an answer
Consider this question:
Does participation in an AI-supported tutoring program improve university students' academic performance?
There are several conceivable ways to investigate it.
Researchers could randomly assign eligible students to intervention and comparison conditions. They could compare naturally occurring groups when randomization is unavailable. They could follow students prospectively according to whether they use the program. They could analyze an existing policy change that created plausibly comparable groups under particular assumptions.
These designs address a related causal problem, but they do not create identical evidence.
Study-design literature acknowledges both that the question should drive design selection and that more than one design may sometimes be applicable.
The question may be the same while the identification strategy changes
For causal questions, one major difference among designs is how they create or approximate the comparison needed to estimate what would have happened in the absence of the exposure or intervention.
Randomization uses chance assignment to create groups that are expected, with adequate implementation and sample size, to be comparable with respect to measured and unmeasured baseline characteristics.
Quasi-experimental designs seek credible comparisons without random assignment, often exploiting intervention timing, eligibility rules, policy changes, discontinuities, or other structures. Observational studies may instead rely on measured covariates, temporal information, matching, weighting, statistical adjustment, or other assumptions to address confounding.
Each route can contribute evidence, but the assumptions required for interpretation differ.
Different designs expose the study to different biases
A design is partly a way of deciding which problems you are willing and able to manage.
Consider a prospective cohort study. It can establish temporal information and observe multiple outcomes, but loss to follow-up and confounding may become important concerns. Retrospective designs can be faster and less expensive but may be limited by variables that were not originally collected for the current research purpose.
A randomized experiment can substantially strengthen causal inference for an intervention under appropriate conditions, yet nonadherence, attrition, contamination, implementation failure, restricted eligibility, or artificial study conditions can still affect interpretation.
A qualitative design may provide rich evidence about mechanisms, experiences, or contextual processes that are largely invisible in outcome comparisons, but it generally answers a different inferential dimension from population prevalence or average causal-effect estimation.
The question “Which design has less bias?” therefore needs a qualifier: less vulnerable to which bias, for which inference?
Different designs may answer different versions of the same broad question
Researchers often say two studies address “the same question” when their operational questions are not actually identical.
Consider the broad problem of whether remote work affects employee productivity.
| Design |
Possible operational question |
What the evidence emphasizes |
| Cross-sectional observational study |
Are current levels of remote work associated with current productivity? |
Contemporaneous association |
| Prospective cohort study |
Do workers with different remote-work patterns subsequently show different productivity trajectories? |
Temporal patterns and longitudinal association |
| Randomized intervention study |
Does assignment to a particular remote-work arrangement change productivity compared with another arrangement? |
Causal effect of the assigned intervention under the study conditions |
| Qualitative case study |
How and under what organizational conditions does remote work shape employees' productive practices? |
Processes, mechanisms, experiences, and context |
All address the relationship between remote work and productivity, but the precise questions and answers differ.
Before declaring that two designs answer exactly the same research question, compare the population, exposure or intervention, comparison, outcome, time frame, context, and intended inference.
Sometimes two designs genuinely can target the same estimand or outcome
The distinction should not be pushed too far. Different designs can sometimes be constructed to estimate closely related quantities.
For example, a randomized trial and a carefully designed observational study might both attempt to estimate the effect of the same intervention on the same outcome in similar populations. Researchers can then compare how the estimates differ and whether design-related biases might account for discrepancies.
Such comparisons are methodologically informative precisely because different designs rely on different mechanisms for controlling bias and confounding.
However, observational and randomized studies that appear to investigate the same intervention may still differ in populations, treatment implementation, follow-up, outcome definitions, or other features. Apparent design comparisons therefore require careful examination before differences in findings are attributed solely to design.
A stronger design may answer the question with fewer assumptions
One way designs differ is in how much must be assumed before the evidence supports the intended conclusion.
Suppose researchers want to estimate the causal effect of an educational intervention.
In a well-conducted randomized experiment, random assignment can make baseline comparability less dependent on measuring every possible confounder. An observational comparison may require stronger assumptions that all important confounding has been adequately measured and addressed, among other conditions.
This does not mean observational evidence is useless. It means the inferential route is different.
When comparing designs, ask not only what data they produce but what must be believed for those data to answer the question.
A less controlled design may sometimes answer a more relevant version of the question
Greater internal control is not the only dimension of usefulness.
An intervention trial conducted under tightly controlled conditions might provide a strong estimate of efficacy for selected participants. A large observational study in routine practice might provide weaker causal identification but reveal how the intervention performs across populations, settings, implementation conditions, or time periods not represented in the trial.
The two designs may therefore contribute complementary evidence.
This is one reason the question should specify whether the researcher wants to know whether something can work under particular conditions, whether it does work under routine conditions, for whom it works, or through what processes.
Different designs may be appropriate because the apparently simple question contains several inferential dimensions.
Qualitative and quantitative designs can address the same phenomenon without being substitutes
Suppose researchers ask whether a new feedback system helps students learn.
A quantitative experiment might estimate changes in measured learning outcomes. A qualitative study might investigate how students interpret the feedback, how it changes their revision practices, and why it appears helpful in some circumstances but not others.
Both studies concern whether and how the system helps learning, but they produce fundamentally different forms of evidence.
Neither should be treated as a cheaper substitute for the other when the research question genuinely requires the kind of evidence the other approach produces.
If both forms of evidence are necessary and intentionally integrated, a mixed methods approach may be justified.
Triangulation across designs can strengthen a research program
Using different designs across several studies can be particularly valuable when their weaknesses differ.
If a similar conclusion appears across designs that are vulnerable to different sources of bias, confidence in the broader finding may increase. Conversely, disagreement across designs can reveal that context, measurement, selection, implementation, or assumptions matter more than initially recognized.
For some questions, methodological guidance has explicitly recommended examining evidence across multiple study designs because different designs have different patterns of bias and confounding.
This does not mean agreement among studies proves that the conclusion is correct. Shared measurement problems, common unmeasured confounding, publication bias, or similar assumptions can still produce convergence. But methodological diversity can provide useful complementary evidence.
The “same question” may need refinement before designs can be compared
Suppose one researcher proposes a randomized experiment and another proposes a cross-sectional survey for the question:
“Does social media affect academic performance?”
Before comparing designs, the question itself needs work.
What aspect of social media? Which population? What exposure or intervention? Compared with what? Which academic outcome? Over what period? Is the question asking about association or causation?
Research-question frameworks such as PICO are useful in applicable fields because specifying population, intervention or exposure, comparison, and outcome makes the evidentiary requirements more explicit. Methodological literature emphasizes that well-formulated questions guide appropriate design selection.
Sometimes the apparent existence of many equally suitable designs reflects an underspecified question rather than genuine methodological equivalence.
Compare designs by consequences, not labels
If several designs appear defensible, compare what changes when you choose one over another.
| Comparison question |
Why it matters |
| What evidence becomes observable? |
A design may provide temporal, contextual, comparative, or mechanistic information that another does not. |
| What inference can be supported? |
Some designs support description or association; others may provide a stronger basis for particular causal or explanatory claims. |
| Which assumptions are required? |
Different designs rely on different assumptions about confounding, measurement, selection, missing data, or interpretation. |
| Which biases are most threatening? |
A design may reduce one threat while increasing another. |
| Who or what can realistically be studied? |
Eligibility, recruitment, setting, and access can alter the population represented by the evidence. |
| What is ethically permissible? |
The strongest theoretical comparison may involve an intervention or exposure that cannot ethically be assigned. |
| What can be completed well? |
Resources, expertise, participant availability, and time affect whether the proposed design can actually be executed with adequate quality. |
This comparison shifts the question from “Which label is better?” to “Which evidentiary trade-offs best serve the research question?”
Different designs do not mean that all designs are equally appropriate
The fact that more than one design can address a question should not be interpreted as methodological relativism.
Some designs may be substantially better suited to the intended inference.
If a feasible randomized experiment can directly address an intervention-effect question, a one-time convenience survey may provide a much weaker answer to that causal question. Both might produce information about the topic, but that does not make them equivalent options.
The appropriate comparison is therefore among designs that can plausibly answer the question, followed by an assessment of their relative strengths, assumptions, limitations, ethics, and feasibility.
This is why research-design appropriateness remains meaningful even when several designs are possible.
More than one defensible design does not necessarily mean one “best” design
Sometimes one option dominates because it provides substantially stronger evidence with acceptable cost, ethics, and feasibility.
In other situations, designs involve genuine trade-offs.
One may provide stronger causal identification but require a narrow population. Another may offer broader real-world representation but depend on stronger assumptions. A third may illuminate mechanisms that neither of the first two can observe directly.
Whether one of these should be called the single best research design depends on which dimensions of the question and evidence are prioritized.