Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Happens If the Data You Need Turn Out to Be Poor Quality?

Having data is not the same as having data capable of answering your research question. If the data turn out to be incomplete, inaccurate, inconsistent, poorly measured, or otherwise unsuitable, determine what can be repaired, what requires a narrower claim, and what makes the study no longer viable.

509
What If Your Research Data Are Poor Quality? Guide 509 of 533
01 · The Question

You Have the Data. What If They Are Not Good Enough?

Your research plan depends on a dataset that appears ideal. It contains thousands of records, covers several years, and includes variables that seem closely related to your research question.

Then you inspect the data.

The primary outcome is missing for many cases. Categories change from one year to another. Several variables have unclear definitions. Dates contradict one another. Values appear outside plausible ranges. A variable described as “engagement” turns out to be nothing more than the number of times a student logged into a platform.

You still have a large dataset. But you may no longer have the evidence you thought you had.

Poor-quality research data are not simply untidy data waiting to be cleaned. Some problems can be corrected or accommodated. Others introduce uncertainty or bias that no amount of statistical sophistication can fully recover. The important question is whether the available data remain fit for the specific research purpose you have in mind.

02 · The Short Answer

Data Quality Determines What the Evidence Can Support

In Brief

If the data you need turn out to be poor quality, identify exactly what is wrong, determine how the problem affects the variables and cases essential to your research question, and reassess what conclusions the data can still support.

Some problems can be corrected, modeled, documented, or accommodated through appropriate analysis. Others require a narrower question, additional data, different measures, or abandonment of the planned study. Data cleaning can correct certain errors, but it cannot manufacture information that was never collected or turn an invalid measure into a valid one.

03 · What You Need to Know

Data Quality Is About Fitness for the Research Question

Researchers sometimes describe a dataset as simply “good” or “bad.” In practice, data quality is better evaluated in relation to intended use.

A dataset may be excellent for administrative reporting but unsuitable for causal inference. Records collected for billing may accurately document transactions while representing a psychological construct poorly. A survey dataset may contain complete responses but use an instrument inappropriate for the population being studied.

UK Research and Innovation guidance treats research data management and quality assurance as matters that should be considered throughout the research lifecycle. Data-quality guidance from the UK Natural Environment Research Council similarly emphasizes quality control during collection, entry or digitization, and subsequent checking so that errors and limitations are visible to users assessing whether data are suitable for a particular purpose.

The practical question is therefore:

“Are these data sufficiently trustworthy and appropriate for the claim I want to make?”

Poor Data Quality Can Mean Several Different Things

“Poor quality” is not a diagnosis. Before deciding what to do, identify the actual problem.

Data-quality problem What it might look like Why it matters
Missingness Important variables are absent for many cases Loss of information and possible bias depending on why values are missing
Measurement error Observed values differ substantially from what they are intended to measure Associations, group differences, classifications, or estimates may be distorted
Misclassification Cases are assigned incorrectly to categories Comparisons and estimated relationships may become biased
Inconsistency Definitions, coding, units, or procedures differ across sites or time Values that appear comparable may represent different things
Invalid values Impossible dates, duplicate identifiers, values outside permitted ranges May indicate entry, transfer, coding, or processing errors
Insufficient granularity Broad categories replace the distinctions required by the question The necessary comparison or construct may not be recoverable
Unclear provenance How variables were generated, transformed, or collected is undocumented The researcher may be unable to judge what the values actually represent
Incomplete coverage Relevant people, periods, events, or settings are systematically absent The dataset may not represent the population or phenomenon required by the question

These problems require different responses. There is no universal “data cleaning” procedure that solves all of them.

Missing Data Are Not Just Empty Cells

When values are missing, begin by determining what is missing, how much is missing, where the missingness occurs, and what processes may have produced it.

A dataset with 10% missing values does not tell you enough. If almost all missingness occurs in an optional background variable, the consequences may be modest. If 10% of participants lack the primary outcome, the issue is more consequential. If missing outcomes occur disproportionately among participants with particular characteristics, the problem may affect more than sample size.

Methodological guidance on missing data emphasizes that the appropriate analysis depends partly on assumptions about the mechanism generating missingness. There is no single analytical technique that automatically removes the problem.

This is why replacing every missing value with a mean, deleting every incomplete case, or applying multiple imputation by default can all be inappropriate depending on the situation.

More Complete Data Are Not Necessarily Better-Measured Data

A variable can have no missing values and still be poor evidence.

Suppose every student in a dataset has a recorded value for “engagement.” That looks reassuring until you discover that engagement is operationalized solely as the number of learning-management-system logins.

The problem is no longer completeness. It is measurement.

A login count may capture one observable behavior associated with engagement, but it does not automatically represent behavioral, emotional, or cognitive engagement more broadly. Whether it is an adequate measure depends on the construct, intended interpretation, and supporting evidence.

Measurement error and misclassification can bias estimated relationships. Methodological guidance from the STRATOS initiative emphasizes that errors in exposures, covariates, and outcomes can materially affect statistical analyses and should be considered during study design as well as analysis.

A Large Dataset Does Not Compensate for Systematic Error

Large datasets can produce precise estimates of the wrong quantity.

If a measurement is systematically biased, increasing the number of observations may reduce sampling uncertainty without correcting that bias. Similarly, thousands of records do not solve a problem in which the variable required by the research question was never collected.

This distinction between quantity and quality matters particularly when researchers use administrative, platform, registry, sensor, or other routinely collected data. Such sources can be enormously valuable, but they were often created for purposes different from the research question now being asked.

Before being impressed by the number of rows, inspect what the columns actually mean.

Inconsistent Definitions Can Create False Comparability

Longitudinal and multisite datasets deserve particular scrutiny because the same variable name may conceal changes in how information was collected.

Suppose an institution records “student withdrawal reason” for ten academic years. During the first five years, staff select from six categories. A new information system later introduces twelve categories and changes several definitions. Combining all ten years without reconciling those changes could create apparent trends that partly reflect the coding system rather than actual changes in student behavior.

The same problem can occur across institutions, laboratories, devices, assessors, versions of an instrument, or data-processing pipelines.

Check data dictionaries, codebooks, metadata, collection protocols, instrument versions, system documentation, and change histories where available. Consistency should be demonstrated rather than inferred from identical column headings.

Data Provenance Matters

Provenance concerns where data came from and what happened to them before they reached you.

Who collected the information? For what original purpose? Which instrument or system generated it? Were values transformed? Were records filtered? Were variables derived from other fields? Were duplicate cases removed? Did definitions change? Were manual corrections made?

Without adequate documentation, a plausible-looking variable can become surprisingly difficult to interpret.

Research-data integrity guidance emphasizes documentation, validation, consistency, and preservation of information necessary to understand and assess data. Reproducibility likewise depends on researchers being able to understand the data and procedures underlying reported findings.

Data Cleaning Can Fix Errors, but It Cannot Fix Every Data Problem

Cleaning is essential when datasets contain identifiable errors or inconsistencies. Typical tasks may include checking duplicates, validating ranges, standardizing formats, reconciling coding, inspecting impossible combinations, documenting missing values, and verifying transformations.

But cleaning has limits.

Repairable data problem An error or inconsistency can be identified and corrected using defensible information, such as an incorrectly formatted date whose correct value can be verified against the source record.
Irrecoverable information problem The information required for the research question was never collected, was measured inadequately, or cannot be reconstructed reliably from available evidence.

Suppose age is recorded as “230” because an extra zero was entered. If the source record confirms that the intended value was 23, correction is straightforward.

Now suppose a database contains only students' final course grades, but your research question requires their scores on a validated measure of critical thinking. No cleaning procedure can reconstruct the missing construct from the final grade merely because both are educational outcomes.

Knowing when not to “fix” data is part of data quality management.

Missing Data Methods Depend on Assumptions

Modern statistical methods can sometimes make better use of incomplete data than simply discarding every case with a missing value. Multiple imputation, likelihood-based methods, weighting approaches, and other techniques may be appropriate in particular settings.

They are not magic.

The validity of such methods depends on assumptions about how the data became missing, the variables included in the analysis, model specification, and other features of the research design. National Academies guidance on missing data in clinical trials has emphasized that substantial missing data can threaten scientific credibility and that analytical methods cannot simply be assumed to compensate for information lost through poor study design or conduct.

When different plausible assumptions about missingness lead to meaningfully different conclusions, sensitivity analyses may be important.

Measurement Error May Require More Than an Acknowledgment

Researchers sometimes recognize that a measure is imperfect and then proceed as though mentioning the limitation resolves it.

It does not.

The consequences of measurement error depend on which variable is affected, the nature of the error, whether the error differs among groups or according to other variables, and the analytical method being used. Error can bias estimates, affect classification, alter precision, and sometimes change the apparent direction or magnitude of relationships.

Where measurement quality is central to the study, consider whether validation information, repeated measurements, reference measures, calibration data, or appropriate analytical adjustments are available.

If none are available, the claim may need to become narrower.

Poor Data Can Change the Research Question

Suppose you planned to investigate whether participation in a mentoring program reduces student dropout. The administrative dataset contains mentoring participation and enrollment status but no reliable information about when students entered the mentoring program.

Without temporal information, the planned question may become difficult to answer because you cannot establish whether mentoring preceded the outcome in the way the proposed analysis requires.

You have several choices. Obtain the missing temporal information from another source. Ask a different question that the available data can support. Use a different study design. Or decide that the current dataset cannot answer the question.

The wrong choice is to preserve the original wording while quietly conducting an analysis that answers something weaker.

Sometimes Poor Data Make the Study No Longer Viable

Researchers understandably resist this conclusion, particularly after spending substantial time obtaining access, cleaning records, and learning the dataset.

But some data problems are fatal to a particular research question.

The primary outcome may be missing for most eligible cases. A central construct may have been represented by an indefensible proxy. Records from comparison groups may have been generated using incompatible procedures. The required dates may not exist. Data provenance may be so uncertain that key variables cannot be interpreted reliably.

If no defensible analytical method, supplementary source, alternative measure, or revised research question can address the problem, continuing with the original study may produce a manuscript without producing credible evidence.

That possibility is one reason data availability should be treated as one of the assumptions on which a research idea depends, rather than something to verify after the study has already been designed.

04 · A Practical Example

When a Large Institutional Dataset Is Less Useful Than It Looks

Hypothetical Example

Predicting Student Dropout From Learning Analytics

Suppose a researcher obtains five years of institutional data covering 30,000 university students. The proposed study will investigate whether online engagement predicts subsequent dropout. The dataset contains demographics, course outcomes, enrollment status, and learning-management-system activity.

At first glance, the project appears exceptionally well resourced.

Problem 1: The outcome definition changed For the first three years, “dropout” includes students who temporarily stopped enrollment. In later years, those students are classified separately. The outcome is therefore not directly comparable across the entire dataset.
Problem 2: The exposure is a weak proxy The variable labeled “online engagement” is simply the number of LMS logins. It does not capture time on task, learning activity, participation quality, or whether a login involved meaningful academic engagement.
Problem 3: Activity data are incomplete Some courses used external learning platforms, so their students appear to have unusually low LMS activity even when they were actively participating elsewhere.
Problem 4: Missingness is structured Several background variables are disproportionately missing among students from particular programs because those programs used a different admissions system.
Assessment The researcher separates problems that can potentially be harmonized from those that change the interpretation. The dropout definition may be reconciled using underlying enrollment records. Missingness can be characterized and addressed using methods appropriate to its likely causes. But login frequency cannot simply be relabeled as a comprehensive measure of engagement.
Decision The research question is narrowed. Instead of claiming to study “student engagement,” the analysis examines whether recorded LMS activity is associated with subsequent enrollment outcomes, with appropriate attention to courses using external platforms. If the broader construct of engagement remains central, additional data or a different design will be required.

The dataset did not become useless. What changed was the claim it could legitimately support.

05 · What Researchers Often Get Wrong

Common Mistakes When Research Data Are Poor Quality

Misconception

A Large Sample Makes Data-Quality Problems Less Important

A larger sample can reduce some forms of sampling uncertainty, but it does not automatically correct systematic measurement error, misclassification, inconsistent definitions, selection problems, or an invalid proxy. More observations can make a biased estimate more precise without making it more accurate.

Misconception

Data Cleaning Can Fix Poor Data

Cleaning can identify and correct some errors, standardize formats, resolve defensible inconsistencies, and document anomalies. It cannot recover information that was never collected or transform a measure that does not represent the intended construct into one that does. Distinguish correctable data errors from fundamental limitations in what the dataset contains.

Misconception

You Should Delete Every Case With Missing Values

Complete-case analysis may be appropriate in some circumstances, but automatic deletion can waste information and may introduce bias depending on why data are missing. The amount, pattern, variables affected, and plausible missingness mechanisms should inform the analytical strategy.

Misconception

Multiple Imputation Solves Missing Data

Multiple imputation can be a valuable method, but its validity depends on assumptions and model specification. It does not recreate information with certainty, and it cannot rescue a study simply because a central variable was barely measured. Appropriate diagnostics, justification, and sometimes sensitivity analyses remain necessary.

Misconception

If the Database Calls a Variable “Engagement,” It Measures Engagement

A variable name is not evidence of construct validity. Inspect how the value was generated, what behavior or response it records, and whether that operationalization supports the interpretation required by your research question. Administrative labels are especially dangerous when they arrive looking suspiciously like theoretical constructs.

Misconception

You Can Decide How to Handle Data Problems After Seeing Which Method Gives the Best Result

Data-processing and analytical decisions should be driven by the nature of the problem and defensible methodological reasoning, not by which option produces the preferred finding. Document cleaning rules, exclusions, transformations, missing-data procedures, and sensitivity analyses transparently.

06 · What This Means for You

Audit the Data Before You Build the Study Around Them

If existing data are central to your research idea, do not wait until formal analysis to discover what they contain. Obtain enough information early to evaluate whether the study is feasible.

That may involve inspecting a data dictionary, metadata, blank data-collection forms, coding manuals, sample records, summary statistics, missingness reports, instrument documentation, system change histories, or a preliminary extract. What is possible depends on the dataset and applicable access, privacy, governance, and ethical requirements.

The purpose is not to perform the entire analysis prematurely. It is to test the assumption that the necessary evidence actually exists.

A simple decision framework

If the problem is a correctable coding or processing error
Correct it using a documented, reproducible rule supported by the source information.
If some data are missing but sufficient information remains
Characterize the missingness and use an analytical approach appropriate to the design, variables, and plausible missing-data mechanism.
If variable definitions differ across sites or periods
Determine whether they can be defensibly harmonized. If not, restrict comparisons or analyses rather than pretending the values are equivalent.
If a measure is an imperfect but defensible representation of the intended construct
Align the interpretation with what the measure can support and consider validation, sensitivity, or supplementary evidence where appropriate.
If the available variable does not adequately represent the construct required by the question
Obtain a better measure or revise the research question. Do not repair the mismatch by changing the label.
If a central variable or necessary temporal, contextual, or comparison information does not exist
Consider another data source, another method, or another research question.
If the remaining data cannot produce credible evidence for the central question
Do not continue with the original study solely because substantial work has already been invested in obtaining or cleaning the dataset.

Separate Data Problems From Research-Question Problems

Sometimes the data are adequate for a narrower question.

If your database cannot measure academic engagement comprehensively but reliably records LMS activity, you may still have a worthwhile study about digital activity patterns. That is not necessarily a lesser project. It is a different claim.

The important point is to make the change explicitly.

Do not write a research question about one construct and quietly substitute another because the latter happens to be available.

Consider Whether Another Method Can Produce Better Evidence

If the existing dataset cannot answer the question, ask whether the question is still worth pursuing through another approach.

You might collect primary data, combine sources, validate a subset of records, use qualitative evidence to address a different aspect of the question, or redesign the study entirely. The appropriate response depends on what information is missing and why it matters.

If the original method depended on data that simply do not exist, consider what happens if your preferred research method cannot be used. Protect the research question where possible rather than protecting the method or dataset merely because you chose it first.

Define Your Minimum Data Requirements in Advance

Before obtaining the full dataset, specify what would make it usable for the study.

Which variables are indispensable? What population and period must be covered? Which definitions must remain comparable? What degree of missingness would require reassessment? What documentation is necessary to interpret derived variables? Which measurement limitations would change the research question?

Not every threshold can be reduced to a universal percentage. A variable can tolerate substantial missingness in one study and become unusable with much less in another, depending on the pattern of missingness, analytical purpose, available auxiliary information, and assumptions required.

Define decision criteria in relation to the actual study rather than adopting arbitrary rules.

Document What Happened to the Data

Maintain a reproducible record of exclusions, corrections, recoding, transformations, merges, derived variables, handling of missing values, and other consequential processing decisions.

UKRI research-data guidance emphasizes data management across the research lifecycle, while data-quality frameworks stress documentation and quality-control procedures that allow users to understand errors and assess suitability.

This record is valuable for your future self, collaborators, reviewers, and anyone attempting to reproduce or reuse the research. Six months later, “I think we removed those cases because something looked strange” is not quite the audit trail one hopes for.

Watch Out

Never change, delete, recode, impute, or exclude observations merely because they make the findings inconvenient. Data-quality decisions should follow documented rules and defensible methodological reasoning. When uncertainty about a data problem cannot be resolved, preserve that uncertainty in the analysis and reporting rather than quietly converting it into certainty.

If poor data remove the study's ability to answer its central question and no realistic alternative exists, the problem may become the strongest argument against proceeding with the proposed study. At that point, reconsidering the project is a research decision, not merely a data-management decision.

07 · A Quick Checklist

Check Whether Your Data Are Fit for the Study

Before relying on a dataset, check:
Confirm that every variable essential to the primary research question actually exists in a usable form.
Inspect how central variables were defined, measured, coded, derived, and transformed rather than relying on their labels.
Quantify missingness for important variables and investigate where and among whom values are missing.
Check for impossible values, duplicates, contradictory records, coding errors, and unexpected anomalies.
Verify whether definitions, instruments, units, systems, and collection procedures changed across sites or time periods.
Assess whether proxies and operational measures support the constructs and claims required by the research question.
Review data provenance, documentation, codebooks, metadata, and collection procedures sufficiently to understand what the dataset represents.
Distinguish errors that can be defensibly corrected from information that cannot be recovered.
Plan missing-data, measurement-error, and sensitivity analyses according to the specific problem rather than applying a generic repair.
Decide whether the data still answer the original research question, require a narrower question, or make the proposed study infeasible.
08 · Frequently Asked Questions

Questions About Poor-Quality Research Data

What makes research data poor quality?

Data may be problematic because they are incomplete, inaccurate, inconsistently defined, measured with substantial error, misclassified, insufficiently documented, unrepresentative of the required population or period, or unsuitable for the construct and inference required by the research question. The relevant dimensions depend on how the data will be used.

Can data cleaning fix poor-quality data?

It can correct or manage some identifiable errors and inconsistencies, but not every problem. Cleaning cannot recreate information that was never collected, establish the validity of an inappropriate measure, or reliably reconstruct unknown values without assumptions. The appropriate response depends on the source and consequences of the problem.

How much missing data is too much?

There is no universal percentage that determines whether a dataset is acceptable. The consequences depend on which variables are missing, why they may be missing, how missingness relates to other variables and outcomes, the analytical method, available auxiliary information, and the research question. A small amount of strategically important missingness can sometimes matter more than a larger amount elsewhere.

Should I delete participants who have missing data?

Not automatically. Complete-case analysis may be reasonable under particular conditions, but deletion can reduce information and may introduce bias depending on the missingness process. Evaluate the pattern and likely causes of missingness and choose a method appropriate to the design and assumptions.

Can I use multiple imputation to solve missing data?

Multiple imputation can be appropriate for some missing-data problems, but it depends on assumptions and an appropriately specified imputation model. It should not be treated as a universal repair, particularly when a variable is almost entirely absent, was never collected for relevant cases, or is itself an inadequate measure of the intended construct.

Can I still use a secondary dataset if some variables are poor quality?

Possibly. Determine whether the problematic variables are necessary for the primary research question and whether their limitations can be addressed or reflected in a narrower interpretation. A dataset can remain useful for some questions even when it is unsuitable for others.

What should I do if the dataset does not contain the variable I need?

First determine whether another defensible measure or data source exists. If not, reconsider the research question or method. Avoid substituting a convenient proxy unless there is a sound conceptual and empirical basis for interpreting it as the construct required by the study.

When should poor data make me abandon a research study?

Consider substantial redesign or abandonment when data problems affect information essential to the primary question and cannot be corrected, adequately modeled, supplemented, or accommodated through a defensible change in scope. If the remaining evidence cannot support the intended inference, continuing because you have already invested heavily in the dataset risks turning past investment into the reason for preserving the study.

09 · The Bottom Line

Do Not Ask More of the Data Than They Can Tell You

The Bottom Line

If your research data turn out to be poor quality, diagnose the specific problem and determine whether the remaining evidence is still fit for your research question before deciding how to analyze it.

Correct what can be defensibly corrected, model or acknowledge uncertainty where appropriate, and narrow or redesign the study when the data cannot support the original claim. If essential information was never collected or cannot be interpreted reliably, sophisticated analysis cannot manufacture it after the fact.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes