03 · What You Need to Know
Research Design Determines What the Evidence Can Tell You
Start with the research question, not the name of the design
When critically evaluating a paper, it is tempting to begin by identifying whether the study is randomized, cross-sectional, longitudinal, qualitative, mixed-methods, or something else.
That information matters, but the design cannot be judged in isolation.
A cross-sectional survey is not inherently weak. A randomized trial is not inherently appropriate. A qualitative interview study is not inherently less rigorous than an experiment. Each design is suited to particular kinds of questions and inferences.
Young and Solomon's guidance on critical appraisal places this relationship near the beginning of article evaluation: readers should identify the research question and determine whether the study design is appropriate for answering it.
So begin with:
What exactly are the researchers trying to find out?
Then ask:
What kind of evidence would be required to answer that question?
Translate the research question into an inferential task
Research questions often contain clues about the evidence they require.
| If the question asks... |
The study needs evidence about... |
Possible design families |
| How common is X? |
Prevalence or frequency in a defined population |
Cross-sectional or population survey designs with appropriate sampling |
| Are X and Y related? |
Association between variables |
Cross-sectional, cohort, correlational, or other observational designs |
| Does X predict Y? |
Prediction of an outcome from specified information |
Longitudinal or appropriately constructed predictive modeling designs |
| Does X cause Y? |
A causal contrast between what happens with and without X |
Randomized experiments where feasible, or appropriately designed causal observational studies under explicit assumptions |
| How does X change over time? |
Temporal development or within-unit change |
Longitudinal or repeated-measures designs |
| How do people experience X? |
Meaning, perception, interpretation, or lived experience |
Qualitative designs appropriate to the specific question and methodology |
| How accurately does X identify Y? |
Performance against an appropriate reference standard |
Diagnostic or classification accuracy designs |
| How and why does an intervention work? |
Effectiveness plus processes, mechanisms, context, or implementation |
Experimental, process-evaluation, qualitative, mixed-methods, or other complementary designs |
These are broad examples rather than rigid one-to-one rules. The important principle is that the wording of the question should imply an evidential requirement.
Descriptive questions need descriptive evidence
Suppose researchers ask:
How common is generative AI use among university students?
This is fundamentally a descriptive question. The researchers need evidence capable of estimating the frequency or prevalence of AI use in a defined population.
A cross-sectional survey can be entirely appropriate.
The critical issues then include how the target population was defined, how students were sampled, whether participation could create systematic differences, how AI use was measured, and whether the estimate is being generalized beyond the population represented by the data.
An experiment would not automatically be "better" simply because experiments are often placed higher in generic evidence hierarchies. Random assignment is unnecessary for estimating how prevalent something currently is.
Watch Out
Do not rank research designs without first asking what question they are supposed to answer. A design that is powerful for causal inference can be unnecessary or inappropriate for a descriptive, interpretive, diagnostic, or exploratory question.
Association questions do not automatically require experiments
Suppose the question is:
Is frequency of AI-tool use associated with students' academic engagement?
An observational design may be appropriate because the question asks whether variables are related rather than whether one causes the other.
A cross-sectional study could estimate whether higher AI use and engagement occur together at one time point. A longitudinal design could provide information about whether earlier AI use is associated with later engagement.
Those designs answer somewhat different questions, but neither is automatically invalid merely because it is observational.
The problem arises if the conclusion changes from:
AI use was associated with engagement
to:
AI use increased engagement.
The second statement asks the design to support a causal inference.
Causal questions require more than finding a statistically significant association
Causal questions are particularly vulnerable to design mismatch.
If the question is:
Does AI-assisted formative feedback improve students' writing performance?
the researchers want to know what would happen to writing performance under the intervention compared with what would happen under an appropriate alternative condition.
Random assignment is valuable because, when implemented successfully and combined with appropriate conduct and analysis, it can help create comparable groups with respect to both measured and unmeasured baseline factors. This strengthens causal interpretation of between-group differences.
Observational research can also address causal questions, but causal inference then depends on design features, measurements, assumptions, and analytical strategies capable of addressing alternative explanations such as confounding and selection.
The mere presence of regression adjustment does not automatically make an observational study causal.
When a paper makes an important causal claim, examine what the design actually does to distinguish causation from association.
Temporal questions require temporal information
If researchers want to know whether something changes, develops, precedes something else, or predicts a later outcome, time matters.
A single cross-sectional measurement can reveal differences among people at one moment. It cannot directly observe within-person change across time.
For example:
Cross-sectional question: Do first-year and fourth-year students differ in research self-efficacy?
A one-time comparison of the two groups can address that question.
But consider:
Developmental question: How does research self-efficacy develop as students progress through university?
Comparing first-year and fourth-year students once does not directly observe development within the same individuals. Cohort differences may reflect factors other than progression through university.
A longitudinal design following students over time would more directly address within-person change.
Difference at one time point
Group A currently differs from Group B.
Change over time
The same units changed as time or exposure progressed.
Those are not interchangeable inferences.
Prediction and explanation are different research tasks
A model can predict an outcome accurately without identifying its causes.
Suppose researchers develop a model that predicts student dropout from attendance, grades, learning-management-system activity, and demographic variables.
If the research question is:
Can we predict which students are at elevated risk of dropping out?
predictive performance is central. The design should support development and evaluation of a model that performs adequately on data not simply used to fit it, ideally with appropriate validation.
If the researchers instead ask:
What causes students to drop out?
a predictive model is not sufficient merely because it predicts well. A variable can be useful for prediction without being a causal determinant, and a causal factor need not be the strongest predictor.
Confusing explanation with prediction can lead researchers to assign causal meaning to coefficients that were generated for predictive purposes.
Questions about experience require evidence about experience
Suppose researchers want to understand:
How do doctoral students experience the use of generative AI during dissertation writing?
A well-designed qualitative study may be more appropriate than a large numerical survey if the aim is to understand meanings, tensions, interpretations, and experiences in depth.
Interviews, focus groups, observations, diaries, documents, or other qualitative materials may provide evidence suited to that question, depending on the methodological approach.
Evaluating such a study by asking why the researchers did not randomly assign participants or calculate statistical power would misunderstand the research question.
The appropriate critical question is whether the qualitative design, sampling, data generation, analysis, reflexivity, and interpretation are coherent with the type of knowledge the study seeks to produce.
This becomes especially important when learning how to evaluate qualitative research without imposing inappropriate quantitative standards.
Mixed questions may require more than one form of evidence
Some research questions contain multiple components.
Consider:
Does an AI-supported tutoring system improve mathematics performance, and how do students experience its use?
The first part concerns an intervention effect. The second concerns experience.
One dataset may not adequately answer both.
A mixed-methods design could combine quantitative evidence about performance with qualitative evidence about students' experiences. But merely including both numerical and textual data does not automatically create a strong mixed-methods study. The components need to address meaningful parts of the research question and be integrated appropriately.
When both forms of evidence are central, evaluate both methodological components and what their integration contributes.
The study population is part of the design-question match
A design can be appropriate in principle but still fail to answer the intended question because the wrong population was studied.
Suppose the research question asks whether an intervention helps novice teachers. The researchers test it exclusively among experienced teachers with more than ten years of practice.
The experiment itself might be rigorously conducted. It does not directly answer the stated question about novices.
Likewise, a study asking about "university students" but sampling only postgraduate engineering students from one specialized institution may need a narrower interpretation.
This is why evaluating study design cannot be separated entirely from judging whether the sample was appropriate.
The outcome must correspond to the question
A design-question mismatch can also occur through measurement.
Suppose researchers ask:
Does an educational intervention improve learning?
But their only outcome is student satisfaction.
The study may be perfectly capable of answering:
Were students more satisfied with the intervention?
It does not thereby answer whether they learned more.
Similarly, intention is not behavior, perceived usefulness is not objective effectiveness, and immediate performance is not necessarily long-term retention.
The design is not merely a label such as "randomized trial." It includes what was measured and when.
A randomized experiment using an outcome unrelated to the research question can still answer the wrong question very rigorously.
The comparison condition determines what the effect means
In intervention research, the question is not simply whether outcomes improved. It is usually whether outcomes differed relative to some comparison.
Consider an educational intervention in which student scores increase from 70 before the intervention to 80 afterward.
Can you conclude that the intervention caused the ten-point improvement?
Not necessarily. Students may improve because of ordinary instruction, repeated testing, maturation, external study, regression to the mean, or other factors.
An appropriate comparison condition helps establish what would plausibly have happened without the intervention.
The choice of comparator also changes the question being answered:
| Comparison |
Question it helps answer |
| Intervention vs. no intervention |
Does the intervention outperform receiving nothing additional? |
| Intervention vs. usual practice |
Does it outperform what participants would ordinarily receive? |
| Intervention vs. active alternative |
Does it outperform another plausible intervention? |
| Intervention vs. attention-matched condition |
Does the specific intervention contribute beyond attention or engagement associated with participating? |
The phrase "the intervention was effective" is incomplete until you know: effective compared with what?
Timing can change the research question
When an outcome is measured matters.
An intervention may improve performance immediately but have no detectable effect several months later. A treatment may produce short-term benefits accompanied by delayed harms. An attitude may change before behavior changes.
If the research question concerns durable learning, measuring the outcome immediately after instruction may not be enough.
If the question concerns immediate performance, a delayed assessment alone may miss the relevant short-term effect.
So ask:
Was the outcome measured at a time point capable of answering the temporal version of the question?
Retrospective designs can answer some questions well and others less well
Retrospective research uses information about events that have already occurred. Such designs can be valuable, particularly when prospective data collection would be impractical, expensive, slow, or unethical.
The important question is whether the available records contain sufficiently accurate information about the variables needed for the research question and whether selection or missingness creates serious bias.
A retrospective cohort can sometimes provide strong evidence about associations between past exposures and later outcomes. It may be less suitable when key confounders were never recorded or when exposure classification is unreliable.
Do not reject retrospective designs merely because they look backward. Evaluate whether the available data can reconstruct the relevant comparison credibly.
Randomization strengthens causal inference, but the word “randomized” is not enough
A randomized controlled trial is often an appropriate design for intervention-effect questions because random allocation can help create comparable groups before treatment.
But identifying a paper as randomized does not complete the appraisal.
You still need to ask:
- Was the allocation sequence genuinely random?
- Was allocation concealed appropriately?
- Did important deviations occur after randomization?
- Was attrition substantial or unequal?
- Were outcomes measured similarly across groups?
- Were analyses consistent with the design?
- Was the reported outcome prespecified?
Cochrane's Risk of Bias 2 framework addresses precisely these kinds of domains when evaluating randomized trials.
A design can be appropriate in principle and compromised in execution.
Design appropriateness and design execution are separate judgments
This distinction helps prevent several appraisal errors.
Design appropriateness
If implemented well, is this type of design capable of answering the research question?
Design execution
Was this particular study implemented well enough for the design's intended strengths to apply?
A randomized trial may be an appropriate design but poorly executed. A cross-sectional survey may be excellently executed but inappropriate for the causal claim attached to it.
You need both judgments.
Analysis cannot create a design that was never there
Sophisticated analysis can improve what researchers learn from a study. It cannot automatically manufacture missing design features.
Regression adjustment cannot retroactively randomize participants. Propensity scores do not guarantee that unmeasured confounding disappears. A machine-learning algorithm cannot make a convenience sample representative merely by processing it more intensively.
This does not mean those methods lack value. It means their inferential contribution depends on assumptions and the data available.
When a paper relies heavily on analytical sophistication, ask whether the analysis appropriately complements the design rather than being asked to rescue it.
Sometimes the paper answers a narrower question than its title suggests
Titles and abstracts often compress research into broad language.
A paper titled around "the effects of AI on student learning" might actually be a survey of self-reported AI use and perceived learning among students at one institution.
The study may still be informative. But its evidence more directly addresses:
How are reported AI use and perceived learning related in this sample?
That is a narrower question than whether AI affects learning generally.
Critical reading often involves reconstructing the question that the design actually answered and comparing it with the question the paper appears to claim it answered.
A mismatch does not always make the entire study useless
Suppose authors intended to demonstrate causality but used a design that supports only association.
You do not necessarily have to discard the paper.
The causal conclusion may fail while the associational evidence remains useful.
Claimed question Does X improve Y?
Available design Cross-sectional observational comparison.
Unsupported inference X causes improvement in Y.
Potentially supported inference X and Y were associated in the observed sample.
This is an example of why a design mismatch can become fatal to one inference without necessarily making every aspect of the paper worthless.
Use design-specific appraisal rather than one universal hierarchy
Generic evidence hierarchies can be useful for particular questions, especially intervention effectiveness, but they become misleading when applied mechanically to all research.
You would not criticize an ethnography because it lacks randomization if its purpose is to understand social practices within a particular setting. Nor would you use an ethnography alone to estimate population prevalence with a known margin of error.
Formal critical-appraisal tools reflect this diversity. Organizations such as JBI provide different appraisal instruments for randomized trials, cohort studies, cross-sectional studies, case-control studies, qualitative research, systematic reviews, diagnostic accuracy studies, and other designs.
The existence of multiple tools reflects a fundamental point: methodological quality must be judged relative to the kind of evidence a design is intended to produce.