03 · What You Need to Know
The Real Question Is How Qualitative Claims Become Trustworthy
The debate about reliability and validity in qualitative research is partly terminological and partly philosophical.
Quantitative measurement commonly uses reliability to address consistency and measurement error, while validity concerns whether evidence supports intended interpretations and uses of scores. Those concepts are essential when researchers use tests, questionnaires, ratings, or other quantitative measures.
Qualitative inquiry often works differently. Researchers may be examining meanings, experiences, practices, narratives, language, social processes, or contextualized interpretations rather than estimating a stable numerical attribute.
Applying a psychometric reliability coefficient to such work may therefore answer no meaningful methodological question.
Qualitative Research Does Not Get a Free Pass From Rigor
Rejecting inappropriate quantitative criteria does not mean rejecting standards of quality.
Qualitative researchers still need to explain why their design fits the question, how participants or data sources were selected, how data were generated, how analysis developed, how researcher perspectives affected the inquiry, and why the interpretation is sufficiently supported by the evidence.
APA's JARS-Qual reporting standards emphasize precisely these issues. They call for transparent reporting of the qualitative approach, researcher characteristics and perspectives, recruitment, data collection, analysis, methodological integrity, and the contextual conditions relevant to the findings.
The issue is therefore not rigor versus no rigor. It is whether the criteria used to judge rigor correspond to the logic of the inquiry.
Why Did Trustworthiness Become an Alternative Framework?
One influential response came from naturalistic inquiry, particularly the work associated with Lincoln and Guba.
Instead of directly applying conventional quantitative criteria, the trustworthiness framework proposed credibility, transferability, dependability, and confirmability. These concepts are often presented as parallels to concerns about truth value, applicability, consistency, and neutrality.
They remain widely used in qualitative health, education, social-science, and applied research.
The framework is influential, not compulsory. Qualitative methodology has continued to develop, and researchers working within different paradigms may use different quality criteria.
The detailed meaning of credibility, dependability, confirmability, and transferability therefore provides one important framework for rigor rather than the only legitimate vocabulary available.
Can Qualitative Researchers Still Use the Word Validity?
Yes.
Some qualitative scholars use validity while explicitly redefining it for interpretive inquiry. Others prefer trustworthiness, credibility, authenticity, methodological integrity, or methodology-specific criteria.
The problem is not the word itself. The problem arises when researchers use “validity” without explaining what it means in their methodological framework or import quantitative validation procedures that do not fit the study.
For example, saying that “the interview questions were validated by three experts” does not establish that the eventual qualitative interpretations are valid. Expert review of an interview guide may improve clarity, relevance, or coverage, but qualitative rigor extends far beyond the wording of interview questions.
Can Qualitative Researchers Use the Word Reliability?
Again, sometimes.
Some qualitative approaches value coding consistency, structured codebooks, reproducible categorization, or agreement among analysts. In those approaches, inter-coder or inter-rater reliability may be relevant.
Other qualitative methodologies treat analysis as interpretive rather than as a classification task in which coders are expected to reproduce the same answer independently.
In reflexive thematic analysis, for example, different researchers may notice different patterns because interpretation is understood as situated and theoretically informed. Treating disagreement as measurement error requiring elimination would impose an analytic philosophy that the method does not assume.
The appropriate question is therefore not “Did two coders agree?” but “Does coder agreement represent quality according to the analytic approach being used?”
Inter-Coder Reliability Is Not a Universal Requirement
Imagine two researchers independently coding interview transcripts.
In a structured qualitative content analysis, the objective may be to apply a defined coding framework consistently. Measuring agreement can then provide relevant information about whether coding categories are being applied as intended.
In a deeply interpretive analysis, researchers may instead discuss different readings of the data, use those differences reflexively, and develop an interpretation through dialogue. Calculating a kappa coefficient might add little because identical coding is not the methodological objective.
Coding consistency as a methodological goal
Agreement or reliability procedures may be relevant when the analysis treats coding categories as rules intended to be applied consistently.
Interpretive engagement as a methodological goal
Differences among researchers may be analytically productive when interpretation is understood as situated rather than as reproduction of one objectively correct coding scheme.
Neither position should be imposed automatically on every qualitative study.
An Interview Guide Is Not a Psychometric Scale
Researchers sometimes ask how to “validate” a semi-structured interview guide in the same way they would validate a questionnaire.
The two instruments serve different purposes.
A questionnaire scale may be designed to produce numerical scores interpreted as measurements of a construct. Its validity argument can therefore involve evidence concerning content, internal structure, relationships with other variables, and reliability.
A semi-structured interview guide generally serves as a tool for facilitating conversation about topics relevant to the research question. Its quality may depend on whether questions are understandable, open enough to elicit meaningful accounts, sensitive to context, aligned with the methodological approach, and adaptable to what participants say.
Expert review or pilot interviews may improve an interview guide. That does not make psychometric construct, criterion, or other measurement-validity procedures automatically appropriate.
Pilot Interviews Do Not “Validate” the Entire Qualitative Study
Piloting can help researchers determine whether questions are understandable, whether the interview sequence works, whether important topics are missing, or whether questions inadvertently constrain participants' responses.
That is useful methodological work.
But a successful pilot does not establish the credibility of findings that have not yet been generated. The quality of the eventual study also depends on sampling, researcher-participant relationships, data richness, analysis, reflexivity, methodological coherence, interpretation, and reporting.
Member Checking Is Not the Qualitative Equivalent of a Validity Test
Member checking is often presented in methods courses as the definitive way to establish validity in qualitative research: return the findings to participants, obtain agreement, and the study is validated.
That formulation is too mechanical.
Participants can help researchers identify factual errors, overlooked perspectives, or interpretations that do not resonate with their experiences. But participant agreement is not necessarily the final criterion for interpretive accuracy.
Participants may disagree with one another. They may reinterpret earlier experiences. They may reject an interpretation for social or political reasons. Some methodologies intentionally examine assumptions that participants themselves may not articulate.
Member checking should therefore be used because it serves the study's epistemological and methodological purpose, not because qualitative research supposedly requires one standard “validity test.”
Triangulation Is Also Not Automatic Validation
Triangulation can involve multiple data sources, researchers, methods, or theoretical perspectives.
When different sources converge, researchers may gain additional confidence in an interpretation. When they diverge, the disagreement may reveal complexity that would otherwise remain hidden.
For example, teachers may report that a new institutional policy is implemented consistently, while classroom observations reveal substantial variation. The disagreement does not mean triangulation failed. It may be the most interesting finding.
Triangulation therefore contributes to rigor when it deepens or challenges interpretation. It should not be treated as a ritual in which three sources must produce the same answer before a theme becomes valid.
Methodological Integrity Offers Another Way to Think About Quality
APA's qualitative reporting standards use the concept of methodological integrity, including attention to the fidelity and utility of the research.
This shifts emphasis toward whether research procedures support the study's goals and whether the resulting claims are grounded, contextualized, and useful for the inquiry being undertaken.
JARS-Qual also emphasizes reporting researchers' perspectives, changes to methodological procedures, participant recruitment, data-collection processes, analytic strategies, and the ways methodological integrity was enhanced or challenged.
That orientation is particularly useful because it avoids requiring every qualitative study to demonstrate rigor using exactly the same technique.
Transparency Is Necessary but Not Sufficient
A researcher can describe a weak study very transparently.
Detailed reporting helps readers evaluate what was done, but transparency cannot transform superficial interviews, inappropriate sampling, incoherent analysis, or unsupported claims into rigorous research.
Reporting guidelines such as SRQR, JARS-Qual, and COREQ should therefore be used to support complete reporting rather than as quality scores.
COREQ, for example, is specifically designed for reporting interview and focus-group research. Applying it to every qualitative methodology merely because the study contains words is not quite the methodological triumph one might hope for.
Qualitative Rigor Is Methodology-Specific
Different approaches ask different kinds of questions and make different kinds of claims.
A phenomenological study may emphasize careful engagement with lived experience. Grounded theory may focus on developing explanatory theoretical categories through iterative data generation and analysis. Ethnography may depend on sustained contextual engagement. Critical discourse research may examine how language constructs social realities and power relations.
These differences affect what counts as appropriate evidence of rigor.
Researchers should therefore identify the quality criteria recognized within their methodological tradition and explain how their procedures satisfy those criteria. The practical strategies for establishing trustworthiness should follow from that methodological logic.
Do Journal Reviewers Expect Reliability and Validity Anyway?
Sometimes they do, particularly in interdisciplinary fields where reviewers may come from quantitative traditions.
The solution is not necessarily to add inappropriate procedures merely to satisfy familiar terminology.
Instead, researchers can explain the quality framework they used, why it fits the methodology, and what evidence supports rigor. Where concepts such as reliability, validity, or inter-coder agreement genuinely apply, report them. Where they do not, explain the alternative criteria and procedures clearly enough that reviewers can evaluate the methodological reasoning.
Authoritative reporting standards can help. APA's JARS-Qual explicitly accommodates multiple qualitative approaches, while SRQR applies broadly to qualitative reporting and COREQ addresses interview and focus-group studies.