03 · What You Need to Know
Validity Is Built Across the Research Design
It is tempting to treat validity as a property that a study either possesses or lacks. In practice, validity is better considered in relation to the inferences a researcher wants to make and the evidence supporting those inferences. A design may support one conclusion reasonably well while providing little basis for another.
For example, a carefully conducted descriptive survey might provide useful evidence about the reported experiences of the sampled participants. The same study may provide much weaker grounds for claiming that one variable caused another. The problem is not that surveys are inherently invalid. The causal claim simply asks the design to establish something it may not have been designed to establish.
Start With the Claim, Not the Design Label
Before asking whether a design is valid, ask what the study is trying to establish. Is the objective to describe a phenomenon, estimate prevalence, compare groups, examine an association, explain a process, understand experiences, evaluate an intervention, predict an outcome, or support a causal inference?
Different questions require different forms of evidence. APA's research-methods guidance similarly frames design selection around the fit between the research question, the data or measurement techniques used to capture the phenomenon, and the subsequent analytic strategy.
This means there is no universally strongest design. A randomized experiment may offer substantial advantages for some causal questions, yet be inappropriate, unethical, or impractical for other questions. An interview study would not normally establish an intervention's average causal effect, but it could provide evidence about how participants experience that intervention that an experiment alone would not reveal.
Design label
The methodological category used to describe the study, such as experimental, cross-sectional, longitudinal, case study, ethnographic, or mixed methods.
Design logic
The reasoning that explains why the chosen participants, evidence, procedures, comparisons, timing, and analysis can address the research question.
A defensible study needs more than the correct label. It needs a convincing design logic.
The Research Question and Design Must Be Aligned
One of the first tests of defensibility is alignment. The design should generate the kind of evidence needed to answer the question actually being asked.
Suppose a researcher asks, “Does a new teaching strategy improve students' critical-thinking performance?” but conducts a one-time survey asking students whether they believe the strategy improved their thinking. Those responses may provide evidence about students' perceptions. They do not, by themselves, demonstrate that critical-thinking performance actually improved because of the strategy.
The mismatch occurs because the claim concerns change and potentially causation, whereas the evidence primarily concerns self-reported perceptions. A more defensible design would either narrow the research question to match the evidence or strengthen the design so that the evidence better addresses the original question.
Your Participants or Cases Must Fit the Population and Question
Who or what enters the study affects what can reasonably be concluded from it. Researchers therefore need to justify the target population, sampling frame, inclusion and exclusion criteria, recruitment procedures, sample size, and relevant characteristics of the resulting sample.
Selection processes can introduce systematic differences between the people included in a study and those who are not. CDC guidance, for example, identifies selection problems, variable participation, measurement error, confounding, and inadequate sample size among important sources of error that researchers should consider when planning data collection.
Representativeness is not equally necessary for every research purpose. Some experiments prioritize control and causal inference over population representativeness. Qualitative studies may deliberately select information-rich participants rather than statistically representative samples. The defensibility question is therefore not simply, “Is the sample representative?” It is, “Does this sampling strategy make sense for the inference this study intends to make?”
Your Measures Must Support the Interpretations You Make From Them
A sound design can still be undermined by poor measurement. Researchers need evidence that their operational definitions, instruments, observations, coding procedures, or other forms of data adequately represent the concepts they intend to investigate.
This is where reliability and validity of measurement become particularly important. Reliability concerns consistency, whereas validity concerns whether available evidence supports the intended interpretation and use of the resulting scores or measurements.
The Standards for Educational and Psychological Testing, developed jointly by the American Educational Research Association, American Psychological Association, and National Council on Measurement in Education, treat validity in terms of evidence and theory supporting interpretations of test scores for their intended uses. This is an important distinction because an instrument is not simply “valid” in every context and for every purpose.
Researchers should therefore resist assuming that adopting a published scale automatically resolves measurement validity. Even a previously studied instrument must be appropriate for the construct, purpose, population, language, context, and interpretation involved in the new study. The question of whether using a validated instrument makes the study itself valid is broader still, because measurement is only one part of the design.
The Design Should Address Plausible Alternative Explanations
Whenever researchers compare groups, estimate relationships, evaluate interventions, or make causal claims, alternative explanations become central to validity.
Could differences have existed before the intervention? Could participant characteristics explain the association? Could attrition have changed the composition of the groups? Could the way outcomes were measured have favored one condition? Could another variable be associated with both the exposure and the outcome?
These possibilities are among the biases, confounding factors, and other threats to valid research that design decisions should anticipate rather than leave entirely to post hoc explanation.
Different designs address these threats differently. Randomization, comparison groups, matching, restriction, repeated measurement, blinding where feasible, standardized procedures, careful sampling, and analytic adjustment may each strengthen particular inferences. None should be treated as a ritual requirement independent of the research question.
Watch Out
Statistical sophistication cannot automatically rescue a weak design. If an important variable was never measured, groups were selected in systematically different ways, an outcome was poorly operationalized, or the temporal sequence needed for a claim was never established, a more complicated statistical model does not necessarily repair the underlying evidential problem.
The Analysis Must Match the Design and the Data
Defensibility continues after data collection. The analytic strategy should reflect how the study was designed, how variables were measured, how observations are structured, what assumptions the method requires, and what question the analysis is intended to answer.
Researchers should be able to explain why an analysis was selected rather than merely report that a familiar test or model was used. Relevant issues may include independence of observations, clustering, repeated measurements, missing data, model assumptions, multiple comparisons, influential observations, coding decisions, and the distinction between planned and exploratory analyses.
Transparency matters here. APA reporting standards, for example, call for reporting methodological and analytic information that allows readers to evaluate research quality, including participant characteristics, measures, analytic strategies, exclusions, and data-dependent decisions where applicable.
Your Conclusions Must Not Outrun Your Evidence
Even a carefully executed study becomes difficult to defend when its conclusions make claims beyond what the design establishes.
A correlation does not automatically establish causation. A statistically significant difference does not necessarily establish practical importance. Findings from a narrowly defined sample do not automatically apply to every population. A null result does not necessarily prove that no meaningful effect exists.
This is why internal and external validity address different questions. A design may provide relatively strong evidence that an observed effect is attributable to the studied intervention under the study conditions while offering more limited evidence about whether the effect will occur elsewhere.
The strongest interpretation is usually not the boldest one. It is the one calibrated to what the design and evidence can reasonably support.
Defensibility Also Requires Transparency
Two researchers may make different but reasonable methodological choices when faced with the same research problem. Defensibility does not mean proving that your choice was the only possible choice. It means showing why it was appropriate given the question, theoretical assumptions, practical constraints, ethical considerations, and intended inference.
A reader should be able to reconstruct the logic of the study: what was done, why it was done, what assumptions were involved, what important threats were addressed, what limitations remain, and how those limitations affect interpretation.
This is also why reproducibility and transparent reporting matter. The National Academies distinguishes computational reproducibility from replicability and emphasizes the importance of sufficient methodological transparency for scientific scrutiny. A method that cannot be understood from the report is difficult for other researchers to evaluate, reproduce where applicable, or build upon.
Validity Does Not Look Identical Across Methodological Traditions
The language used to evaluate rigor varies across research traditions. Quantitative research commonly discusses issues such as measurement validity, internal validity, external validity, bias, precision, and confounding. Qualitative researchers may instead evaluate credibility, dependability, confirmability, and transferability, depending on the methodological tradition being used.
Mixed-methods research adds another layer because researchers must consider not only the quality of the quantitative and qualitative components but also whether combining them is methodologically coherent and contributes something meaningful to the research question.
These traditions should not be forced into a single checklist. What they share is a more fundamental expectation: the evidence, procedures, analytic reasoning, and conclusions should be coherent with the question and with the epistemological and methodological claims being made.