Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Makes a Research Design Valid and Defensible?

A defensible research design is not simply one that follows a recognized method. It is one in which the research question, evidence, procedures, analysis, and conclusions are coherently aligned and the important methodological choices can be justified.

126
Valid and Defensible Research Design Guide 126 of 217
01 · The Question

When Can You Actually Defend Your Research Design?

Researchers are often told to choose the “right” research design. That advice sounds straightforward until you have to explain why a particular experiment, survey, case study, cohort design, interview study, or mixed-methods design can actually answer your research question.

A design can look methodologically sophisticated and still produce weak evidence. A familiar design label does not guarantee appropriate sampling, meaningful measurement, adequate control of alternative explanations, or conclusions that stay within what the data can support. Conversely, a study may have unavoidable limitations and still be methodologically defensible if its choices are appropriate, transparent, and interpreted within those limitations.

The more useful question, then, is not simply whether you used an accepted research design. It is whether the chain connecting your question to your conclusions is strong enough to justify the claims you intend to make.

02 · The Short Answer

A Valid Design Supports the Inference You Want to Make

In Brief

A research design is valid and defensible when its methodological choices are appropriate for the research question and provide sufficiently credible evidence for the conclusions being drawn.

This does not require a flawless study. It requires alignment among the question, design, participants or cases, measurements or evidence, procedures, analysis, and interpretation, together with serious attention to plausible sources of bias and uncertainty. What counts as convincing evidence also depends on the type of claim and the methodological tradition in which the study is conducted.

03 · What You Need to Know

Validity Is Built Across the Research Design

It is tempting to treat validity as a property that a study either possesses or lacks. In practice, validity is better considered in relation to the inferences a researcher wants to make and the evidence supporting those inferences. A design may support one conclusion reasonably well while providing little basis for another.

For example, a carefully conducted descriptive survey might provide useful evidence about the reported experiences of the sampled participants. The same study may provide much weaker grounds for claiming that one variable caused another. The problem is not that surveys are inherently invalid. The causal claim simply asks the design to establish something it may not have been designed to establish.

Start With the Claim, Not the Design Label

Before asking whether a design is valid, ask what the study is trying to establish. Is the objective to describe a phenomenon, estimate prevalence, compare groups, examine an association, explain a process, understand experiences, evaluate an intervention, predict an outcome, or support a causal inference?

Different questions require different forms of evidence. APA's research-methods guidance similarly frames design selection around the fit between the research question, the data or measurement techniques used to capture the phenomenon, and the subsequent analytic strategy.

This means there is no universally strongest design. A randomized experiment may offer substantial advantages for some causal questions, yet be inappropriate, unethical, or impractical for other questions. An interview study would not normally establish an intervention's average causal effect, but it could provide evidence about how participants experience that intervention that an experiment alone would not reveal.

Design label The methodological category used to describe the study, such as experimental, cross-sectional, longitudinal, case study, ethnographic, or mixed methods.
Design logic The reasoning that explains why the chosen participants, evidence, procedures, comparisons, timing, and analysis can address the research question.

A defensible study needs more than the correct label. It needs a convincing design logic.

The Research Question and Design Must Be Aligned

One of the first tests of defensibility is alignment. The design should generate the kind of evidence needed to answer the question actually being asked.

Suppose a researcher asks, “Does a new teaching strategy improve students' critical-thinking performance?” but conducts a one-time survey asking students whether they believe the strategy improved their thinking. Those responses may provide evidence about students' perceptions. They do not, by themselves, demonstrate that critical-thinking performance actually improved because of the strategy.

The mismatch occurs because the claim concerns change and potentially causation, whereas the evidence primarily concerns self-reported perceptions. A more defensible design would either narrow the research question to match the evidence or strengthen the design so that the evidence better addresses the original question.

Your Participants or Cases Must Fit the Population and Question

Who or what enters the study affects what can reasonably be concluded from it. Researchers therefore need to justify the target population, sampling frame, inclusion and exclusion criteria, recruitment procedures, sample size, and relevant characteristics of the resulting sample.

Selection processes can introduce systematic differences between the people included in a study and those who are not. CDC guidance, for example, identifies selection problems, variable participation, measurement error, confounding, and inadequate sample size among important sources of error that researchers should consider when planning data collection.

Representativeness is not equally necessary for every research purpose. Some experiments prioritize control and causal inference over population representativeness. Qualitative studies may deliberately select information-rich participants rather than statistically representative samples. The defensibility question is therefore not simply, “Is the sample representative?” It is, “Does this sampling strategy make sense for the inference this study intends to make?”

Your Measures Must Support the Interpretations You Make From Them

A sound design can still be undermined by poor measurement. Researchers need evidence that their operational definitions, instruments, observations, coding procedures, or other forms of data adequately represent the concepts they intend to investigate.

This is where reliability and validity of measurement become particularly important. Reliability concerns consistency, whereas validity concerns whether available evidence supports the intended interpretation and use of the resulting scores or measurements.

The Standards for Educational and Psychological Testing, developed jointly by the American Educational Research Association, American Psychological Association, and National Council on Measurement in Education, treat validity in terms of evidence and theory supporting interpretations of test scores for their intended uses. This is an important distinction because an instrument is not simply “valid” in every context and for every purpose.

Researchers should therefore resist assuming that adopting a published scale automatically resolves measurement validity. Even a previously studied instrument must be appropriate for the construct, purpose, population, language, context, and interpretation involved in the new study. The question of whether using a validated instrument makes the study itself valid is broader still, because measurement is only one part of the design.

The Design Should Address Plausible Alternative Explanations

Whenever researchers compare groups, estimate relationships, evaluate interventions, or make causal claims, alternative explanations become central to validity.

Could differences have existed before the intervention? Could participant characteristics explain the association? Could attrition have changed the composition of the groups? Could the way outcomes were measured have favored one condition? Could another variable be associated with both the exposure and the outcome?

These possibilities are among the biases, confounding factors, and other threats to valid research that design decisions should anticipate rather than leave entirely to post hoc explanation.

Different designs address these threats differently. Randomization, comparison groups, matching, restriction, repeated measurement, blinding where feasible, standardized procedures, careful sampling, and analytic adjustment may each strengthen particular inferences. None should be treated as a ritual requirement independent of the research question.

Watch Out

Statistical sophistication cannot automatically rescue a weak design. If an important variable was never measured, groups were selected in systematically different ways, an outcome was poorly operationalized, or the temporal sequence needed for a claim was never established, a more complicated statistical model does not necessarily repair the underlying evidential problem.

The Analysis Must Match the Design and the Data

Defensibility continues after data collection. The analytic strategy should reflect how the study was designed, how variables were measured, how observations are structured, what assumptions the method requires, and what question the analysis is intended to answer.

Researchers should be able to explain why an analysis was selected rather than merely report that a familiar test or model was used. Relevant issues may include independence of observations, clustering, repeated measurements, missing data, model assumptions, multiple comparisons, influential observations, coding decisions, and the distinction between planned and exploratory analyses.

Transparency matters here. APA reporting standards, for example, call for reporting methodological and analytic information that allows readers to evaluate research quality, including participant characteristics, measures, analytic strategies, exclusions, and data-dependent decisions where applicable.

Your Conclusions Must Not Outrun Your Evidence

Even a carefully executed study becomes difficult to defend when its conclusions make claims beyond what the design establishes.

A correlation does not automatically establish causation. A statistically significant difference does not necessarily establish practical importance. Findings from a narrowly defined sample do not automatically apply to every population. A null result does not necessarily prove that no meaningful effect exists.

This is why internal and external validity address different questions. A design may provide relatively strong evidence that an observed effect is attributable to the studied intervention under the study conditions while offering more limited evidence about whether the effect will occur elsewhere.

The strongest interpretation is usually not the boldest one. It is the one calibrated to what the design and evidence can reasonably support.

Defensibility Also Requires Transparency

Two researchers may make different but reasonable methodological choices when faced with the same research problem. Defensibility does not mean proving that your choice was the only possible choice. It means showing why it was appropriate given the question, theoretical assumptions, practical constraints, ethical considerations, and intended inference.

A reader should be able to reconstruct the logic of the study: what was done, why it was done, what assumptions were involved, what important threats were addressed, what limitations remain, and how those limitations affect interpretation.

This is also why reproducibility and transparent reporting matter. The National Academies distinguishes computational reproducibility from replicability and emphasizes the importance of sufficient methodological transparency for scientific scrutiny. A method that cannot be understood from the report is difficult for other researchers to evaluate, reproduce where applicable, or build upon.

Validity Does Not Look Identical Across Methodological Traditions

The language used to evaluate rigor varies across research traditions. Quantitative research commonly discusses issues such as measurement validity, internal validity, external validity, bias, precision, and confounding. Qualitative researchers may instead evaluate credibility, dependability, confirmability, and transferability, depending on the methodological tradition being used.

Mixed-methods research adds another layer because researchers must consider not only the quality of the quantitative and qualitative components but also whether combining them is methodologically coherent and contributes something meaningful to the research question.

These traditions should not be forced into a single checklist. What they share is a more fundamental expectation: the evidence, procedures, analytic reasoning, and conclusions should be coherent with the question and with the epistemological and methodological claims being made.

04 · A Practical Example

How a Plausible Study Becomes More Defensible

Hypothetical Example

Does an AI-supported feedback tool improve student writing?

A university researcher wants to determine whether students who use an AI-supported formative feedback tool improve their academic writing. The researcher initially plans to survey one class at the end of the semester and ask whether students think the tool improved their writing.

Question The intended claim concerns improvement in writing performance, not merely students' perceptions of the tool.
Design The researcher collects writing performance data before and after the intervention and includes an appropriate comparison condition rather than relying only on a post-use opinion survey.
Measurement Writing performance is assessed using clearly specified criteria, with procedures for applying those criteria consistently. Student perceptions are retained as a separate outcome rather than treated as evidence of actual performance improvement.
Threats The researcher considers whether the groups differ at baseline, whether students use the tool as intended, whether attrition differs between conditions, and whether assessors' knowledge of the condition could influence scoring.
Analysis The analysis accounts for the structure of the data and addresses the comparison relevant to the research question rather than simply testing every available variable for statistical significance.
Interpretation The conclusions are limited to what the design supports. If the participants came from one course at one institution, the researcher does not assume that the findings automatically generalize to every discipline, institution, or student population.

The revised study is more defensible not because it has become perfect, but because the evidence now corresponds more closely to the intended claim. Remaining limitations can still be acknowledged and considered when interpreting the findings.

05 · What Researchers Often Get Wrong

Common Mistakes When Judging Research Design Quality

Misconception

Is a More Complex Design Automatically More Rigorous?

No. Complexity and rigor are not synonyms. Adding variables, sophisticated statistical models, multiple data sources, or additional methodological stages can create complexity without improving the inference. A simpler design that directly addresses the research question may be considerably easier to defend.

Misconception

Does Following a Textbook Design Make the Study Valid?

A recognized design provides a methodological framework, not a guarantee of validity. Sampling, measurement, implementation, missing data, bias, analysis, and interpretation can still weaken the evidence. Researchers need to justify how the design was actually implemented in their particular study.

Misconception

Does a Validated Instrument Validate the Entire Study?

No. Evidence supporting an instrument's use addresses a measurement problem, not every aspect of study validity. Sampling bias, confounding, inappropriate comparisons, implementation problems, analytic errors, and overgeneralization can remain even when measurement is strong.

Misconception

Does a Large Sample Fix a Weak Design?

A larger sample can improve precision and statistical power under appropriate conditions, but it does not automatically remove systematic bias. A very large biased sample can estimate the wrong quantity with considerable precision. Sample size should therefore be considered alongside sampling quality, measurement, design, and the intended inference.

Misconception

Does Every Limitation Mean the Design Is Flawed?

No. Research involves constraints and trade-offs. A narrow sample, limited observation period, reliance on self-report, or reduced experimental control may be reasonable for a particular question. The important distinction is whether a limitation represents a justified trade-off or whether it undermines a central inference the study nevertheless claims to support. That distinction is explored more directly when considering whether a limitation is a design flaw or an acceptable trade-off.

06 · What This Means for You

Build a Chain of Justification From Question to Conclusion

When designing or reviewing a study, do not ask only whether each individual methodological choice appears acceptable. Examine how those choices work together.

A useful way to test your design is to imagine a skeptical but reasonable reader asking, “Why?” after every major decision. Why this design? Why these participants? Why this measure? Why this comparison? Why this timing? Why this analysis? Why does this evidence justify that conclusion?

You do not need an elaborate defense for every routine procedural detail. You should, however, be able to justify choices that materially affect what the findings mean.

A simple decision framework

If your evidence does not directly address the research question
Revise the question, collect more appropriate evidence, or change the design.
If an alternative explanation could plausibly account for the finding
Determine whether it can be addressed through design, measurement, analysis, or appropriately cautious interpretation.
If a methodological choice is unavoidable but introduces a limitation
Explain the trade-off and restrict the conclusions accordingly rather than pretending the limitation does not matter.
If your conclusion requires a stronger inference than your design supports
Narrow the claim. A defensible modest conclusion is methodologically stronger than an ambitious unsupported one.

The objective is not methodological invulnerability. Few studies could survive that standard. The objective is a coherent argument that the evidence generated by the design is appropriate for the conclusion, with uncertainty and limitations represented accurately.

07 · A Quick Checklist

Can You Defend the Logic of Your Research Design?

Before finalizing your research design, check:
Can you state precisely what your research question requires you to describe, compare, explain, predict, understand, or estimate?
Does the selected design generate evidence that can actually address that question?
Can you justify who or what is included in the study and how participants, cases, records, or observations are selected?
Do your measures, observations, or other evidence adequately represent the concepts you intend to investigate?
Have you identified the most plausible sources of bias, confounding, measurement error, or alternative explanation for your particular design?
Does your analysis correspond to the design, structure of the data, and research question?
Are important methodological decisions and deviations documented transparently enough for readers to evaluate the study?
Do your conclusions stay within the population, conditions, constructs, and kinds of inference that the evidence can support?
Can you explain the important limitations and how they change, rather than merely weaken, the interpretation of the findings?
08 · Frequently Asked Questions

Questions Researchers Ask About Design Validity

Can a research design ever be completely valid?

It is usually more useful to ask how strongly a design supports particular inferences than to classify an entire study as completely valid or invalid. Research designs operate under assumptions and constraints, and different validity concerns may matter for different conclusions.

What is the most valid research design?

There is no single most valid design for every research question. The appropriate design depends on what you want to know, what kind of inference you intend to make, the evidence that can answer the question, ethical considerations, feasibility, and the methodological tradition of the study.

Does a randomized controlled trial guarantee validity?

No. Randomization can strengthen causal inference by helping address confounding, but randomized studies can still be affected by attrition, nonadherence, measurement problems, compromised allocation procedures, missing data, inappropriate analysis, or limited applicability beyond the study conditions.

Can an observational study be valid and defensible?

Yes. Observational designs can provide strong and useful evidence for many research questions. Their defensibility depends on the question, sampling, measurement, temporal structure, handling of bias and confounding, analysis, and the claims made from the findings. They should not be judged simply by whether randomization was possible.

Can a study be valid even if it is not generalizable?

Potentially. A study may support a credible inference within the participants and conditions studied while providing limited evidence about other populations or settings. This is why internal validity and generalizability should not be treated as the same property.

Does acknowledging limitations make a study less defensible?

Usually the opposite. Transparent limitations help readers determine the boundaries of the evidence. A study becomes harder to defend when consequential limitations are ignored or when conclusions exceed what those limitations permit.

Can a study with important limitations still be useful?

Yes. Limitations do not automatically erase the evidential value of a study. Their importance depends on which inference they threaten, how severely they affect it, and whether the conclusions are appropriately qualified. A study can therefore produce useful evidence despite meaningful limitations.

09 · The Bottom Line

A Defensible Design Makes Its Reasoning Visible

The Bottom Line

A valid and defensible research design creates a coherent connection between the question being asked, the evidence collected, the way that evidence is analyzed, and the conclusion the researcher ultimately makes.

You do not need to eliminate every limitation. You need to anticipate the threats that matter for your intended inference, make methodological choices that address them where reasonably possible, explain consequential trade-offs, and keep your claims within the boundaries of the evidence.

10 · Sources and Further Reading

Authoritative Resources on Research Design and Validity

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes