03 · What You Need to Know
Data Quality Is About Fitness for the Research Question
Researchers sometimes describe a dataset as simply “good” or “bad.” In practice, data quality is better evaluated in relation to intended use.
A dataset may be excellent for administrative reporting but unsuitable for causal inference. Records collected for billing may accurately document transactions while representing a psychological construct poorly. A survey dataset may contain complete responses but use an instrument inappropriate for the population being studied.
UK Research and Innovation guidance treats research data management and quality assurance as matters that should be considered throughout the research lifecycle. Data-quality guidance from the UK Natural Environment Research Council similarly emphasizes quality control during collection, entry or digitization, and subsequent checking so that errors and limitations are visible to users assessing whether data are suitable for a particular purpose.
The practical question is therefore:
“Are these data sufficiently trustworthy and appropriate for the claim I want to make?”
Poor Data Quality Can Mean Several Different Things
“Poor quality” is not a diagnosis. Before deciding what to do, identify the actual problem.
| Data-quality problem |
What it might look like |
Why it matters |
| Missingness |
Important variables are absent for many cases |
Loss of information and possible bias depending on why values are missing |
| Measurement error |
Observed values differ substantially from what they are intended to measure |
Associations, group differences, classifications, or estimates may be distorted |
| Misclassification |
Cases are assigned incorrectly to categories |
Comparisons and estimated relationships may become biased |
| Inconsistency |
Definitions, coding, units, or procedures differ across sites or time |
Values that appear comparable may represent different things |
| Invalid values |
Impossible dates, duplicate identifiers, values outside permitted ranges |
May indicate entry, transfer, coding, or processing errors |
| Insufficient granularity |
Broad categories replace the distinctions required by the question |
The necessary comparison or construct may not be recoverable |
| Unclear provenance |
How variables were generated, transformed, or collected is undocumented |
The researcher may be unable to judge what the values actually represent |
| Incomplete coverage |
Relevant people, periods, events, or settings are systematically absent |
The dataset may not represent the population or phenomenon required by the question |
These problems require different responses. There is no universal “data cleaning” procedure that solves all of them.
Missing Data Are Not Just Empty Cells
When values are missing, begin by determining what is missing, how much is missing, where the missingness occurs, and what processes may have produced it.
A dataset with 10% missing values does not tell you enough. If almost all missingness occurs in an optional background variable, the consequences may be modest. If 10% of participants lack the primary outcome, the issue is more consequential. If missing outcomes occur disproportionately among participants with particular characteristics, the problem may affect more than sample size.
Methodological guidance on missing data emphasizes that the appropriate analysis depends partly on assumptions about the mechanism generating missingness. There is no single analytical technique that automatically removes the problem.
This is why replacing every missing value with a mean, deleting every incomplete case, or applying multiple imputation by default can all be inappropriate depending on the situation.
More Complete Data Are Not Necessarily Better-Measured Data
A variable can have no missing values and still be poor evidence.
Suppose every student in a dataset has a recorded value for “engagement.” That looks reassuring until you discover that engagement is operationalized solely as the number of learning-management-system logins.
The problem is no longer completeness. It is measurement.
A login count may capture one observable behavior associated with engagement, but it does not automatically represent behavioral, emotional, or cognitive engagement more broadly. Whether it is an adequate measure depends on the construct, intended interpretation, and supporting evidence.
Measurement error and misclassification can bias estimated relationships. Methodological guidance from the STRATOS initiative emphasizes that errors in exposures, covariates, and outcomes can materially affect statistical analyses and should be considered during study design as well as analysis.
A Large Dataset Does Not Compensate for Systematic Error
Large datasets can produce precise estimates of the wrong quantity.
If a measurement is systematically biased, increasing the number of observations may reduce sampling uncertainty without correcting that bias. Similarly, thousands of records do not solve a problem in which the variable required by the research question was never collected.
This distinction between quantity and quality matters particularly when researchers use administrative, platform, registry, sensor, or other routinely collected data. Such sources can be enormously valuable, but they were often created for purposes different from the research question now being asked.
Before being impressed by the number of rows, inspect what the columns actually mean.
Inconsistent Definitions Can Create False Comparability
Longitudinal and multisite datasets deserve particular scrutiny because the same variable name may conceal changes in how information was collected.
Suppose an institution records “student withdrawal reason” for ten academic years. During the first five years, staff select from six categories. A new information system later introduces twelve categories and changes several definitions. Combining all ten years without reconciling those changes could create apparent trends that partly reflect the coding system rather than actual changes in student behavior.
The same problem can occur across institutions, laboratories, devices, assessors, versions of an instrument, or data-processing pipelines.
Check data dictionaries, codebooks, metadata, collection protocols, instrument versions, system documentation, and change histories where available. Consistency should be demonstrated rather than inferred from identical column headings.
Data Provenance Matters
Provenance concerns where data came from and what happened to them before they reached you.
Who collected the information? For what original purpose? Which instrument or system generated it? Were values transformed? Were records filtered? Were variables derived from other fields? Were duplicate cases removed? Did definitions change? Were manual corrections made?
Without adequate documentation, a plausible-looking variable can become surprisingly difficult to interpret.
Research-data integrity guidance emphasizes documentation, validation, consistency, and preservation of information necessary to understand and assess data. Reproducibility likewise depends on researchers being able to understand the data and procedures underlying reported findings.
Data Cleaning Can Fix Errors, but It Cannot Fix Every Data Problem
Cleaning is essential when datasets contain identifiable errors or inconsistencies. Typical tasks may include checking duplicates, validating ranges, standardizing formats, reconciling coding, inspecting impossible combinations, documenting missing values, and verifying transformations.
But cleaning has limits.
Repairable data problem
An error or inconsistency can be identified and corrected using defensible information, such as an incorrectly formatted date whose correct value can be verified against the source record.
Irrecoverable information problem
The information required for the research question was never collected, was measured inadequately, or cannot be reconstructed reliably from available evidence.
Suppose age is recorded as “230” because an extra zero was entered. If the source record confirms that the intended value was 23, correction is straightforward.
Now suppose a database contains only students' final course grades, but your research question requires their scores on a validated measure of critical thinking. No cleaning procedure can reconstruct the missing construct from the final grade merely because both are educational outcomes.
Knowing when not to “fix” data is part of data quality management.
Missing Data Methods Depend on Assumptions
Modern statistical methods can sometimes make better use of incomplete data than simply discarding every case with a missing value. Multiple imputation, likelihood-based methods, weighting approaches, and other techniques may be appropriate in particular settings.
They are not magic.
The validity of such methods depends on assumptions about how the data became missing, the variables included in the analysis, model specification, and other features of the research design. National Academies guidance on missing data in clinical trials has emphasized that substantial missing data can threaten scientific credibility and that analytical methods cannot simply be assumed to compensate for information lost through poor study design or conduct.
When different plausible assumptions about missingness lead to meaningfully different conclusions, sensitivity analyses may be important.
Measurement Error May Require More Than an Acknowledgment
Researchers sometimes recognize that a measure is imperfect and then proceed as though mentioning the limitation resolves it.
It does not.
The consequences of measurement error depend on which variable is affected, the nature of the error, whether the error differs among groups or according to other variables, and the analytical method being used. Error can bias estimates, affect classification, alter precision, and sometimes change the apparent direction or magnitude of relationships.
Where measurement quality is central to the study, consider whether validation information, repeated measurements, reference measures, calibration data, or appropriate analytical adjustments are available.
If none are available, the claim may need to become narrower.
Poor Data Can Change the Research Question
Suppose you planned to investigate whether participation in a mentoring program reduces student dropout. The administrative dataset contains mentoring participation and enrollment status but no reliable information about when students entered the mentoring program.
Without temporal information, the planned question may become difficult to answer because you cannot establish whether mentoring preceded the outcome in the way the proposed analysis requires.
You have several choices. Obtain the missing temporal information from another source. Ask a different question that the available data can support. Use a different study design. Or decide that the current dataset cannot answer the question.
The wrong choice is to preserve the original wording while quietly conducting an analysis that answers something weaker.
Sometimes Poor Data Make the Study No Longer Viable
Researchers understandably resist this conclusion, particularly after spending substantial time obtaining access, cleaning records, and learning the dataset.
But some data problems are fatal to a particular research question.
The primary outcome may be missing for most eligible cases. A central construct may have been represented by an indefensible proxy. Records from comparison groups may have been generated using incompatible procedures. The required dates may not exist. Data provenance may be so uncertain that key variables cannot be interpreted reliably.
If no defensible analytical method, supplementary source, alternative measure, or revised research question can address the problem, continuing with the original study may produce a manuscript without producing credible evidence.
That possibility is one reason data availability should be treated as one of the assumptions on which a research idea depends, rather than something to verify after the study has already been designed.
06 · What This Means for You
Audit the Data Before You Build the Study Around Them
If existing data are central to your research idea, do not wait until formal analysis to discover what they contain. Obtain enough information early to evaluate whether the study is feasible.
That may involve inspecting a data dictionary, metadata, blank data-collection forms, coding manuals, sample records, summary statistics, missingness reports, instrument documentation, system change histories, or a preliminary extract. What is possible depends on the dataset and applicable access, privacy, governance, and ethical requirements.
The purpose is not to perform the entire analysis prematurely. It is to test the assumption that the necessary evidence actually exists.
A simple decision framework
If the problem is a correctable coding or processing error
Correct it using a documented, reproducible rule supported by the source information.
If some data are missing but sufficient information remains
Characterize the missingness and use an analytical approach appropriate to the design, variables, and plausible missing-data mechanism.
If variable definitions differ across sites or periods
Determine whether they can be defensibly harmonized. If not, restrict comparisons or analyses rather than pretending the values are equivalent.
If a measure is an imperfect but defensible representation of the intended construct
Align the interpretation with what the measure can support and consider validation, sensitivity, or supplementary evidence where appropriate.
If the available variable does not adequately represent the construct required by the question
Obtain a better measure or revise the research question. Do not repair the mismatch by changing the label.
If a central variable or necessary temporal, contextual, or comparison information does not exist
Consider another data source, another method, or another research question.
If the remaining data cannot produce credible evidence for the central question
Do not continue with the original study solely because substantial work has already been invested in obtaining or cleaning the dataset.
Separate Data Problems From Research-Question Problems
Sometimes the data are adequate for a narrower question.
If your database cannot measure academic engagement comprehensively but reliably records LMS activity, you may still have a worthwhile study about digital activity patterns. That is not necessarily a lesser project. It is a different claim.
The important point is to make the change explicitly.
Do not write a research question about one construct and quietly substitute another because the latter happens to be available.
Consider Whether Another Method Can Produce Better Evidence
If the existing dataset cannot answer the question, ask whether the question is still worth pursuing through another approach.
You might collect primary data, combine sources, validate a subset of records, use qualitative evidence to address a different aspect of the question, or redesign the study entirely. The appropriate response depends on what information is missing and why it matters.
If the original method depended on data that simply do not exist, consider what happens if your preferred research method cannot be used. Protect the research question where possible rather than protecting the method or dataset merely because you chose it first.
Define Your Minimum Data Requirements in Advance
Before obtaining the full dataset, specify what would make it usable for the study.
Which variables are indispensable? What population and period must be covered? Which definitions must remain comparable? What degree of missingness would require reassessment? What documentation is necessary to interpret derived variables? Which measurement limitations would change the research question?
Not every threshold can be reduced to a universal percentage. A variable can tolerate substantial missingness in one study and become unusable with much less in another, depending on the pattern of missingness, analytical purpose, available auxiliary information, and assumptions required.
Define decision criteria in relation to the actual study rather than adopting arbitrary rules.
Document What Happened to the Data
Maintain a reproducible record of exclusions, corrections, recoding, transformations, merges, derived variables, handling of missing values, and other consequential processing decisions.
UKRI research-data guidance emphasizes data management across the research lifecycle, while data-quality frameworks stress documentation and quality-control procedures that allow users to understand errors and assess suitability.
This record is valuable for your future self, collaborators, reviewers, and anyone attempting to reproduce or reuse the research. Six months later, “I think we removed those cases because something looked strange” is not quite the audit trail one hopes for.
Watch Out
Never change, delete, recode, impute, or exclude observations merely because they make the findings inconvenient. Data-quality decisions should follow documented rules and defensible methodological reasoning. When uncertainty about a data problem cannot be resolved, preserve that uncertainty in the analysis and reporting rather than quietly converting it into certainty.
If poor data remove the study's ability to answer its central question and no realistic alternative exists, the problem may become the strongest argument against proceeding with the proposed study. At that point, reconsidering the project is a research decision, not merely a data-management decision.