Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Happens If You Collect Data You Don’t Know How to Analyze?

Collecting data without knowing how to analyze them does not automatically ruin a study, but it can reveal serious mismatches among the research question, design, measurements, and analysis. The solution is to diagnose the problem before choosing a convenient statistical test.

154
Data You Don't Know How to Analyze Guide 154 of 217
01 · The Question

You Have the Dataset. Now You Are Not Sure What to Do With It.

Data collection is finished. The spreadsheet is complete. Then comes an uncomfortable realization: you do not know which analysis will answer your research question.

This situation is more common than researchers sometimes admit. Perhaps the design is more complicated than expected. The outcome does not behave as anticipated. Several variables were measured without clear analytical roles. Observations are nested or repeated. Or the methods section simply promised that “appropriate statistical analyses” would be performed later.

The immediate temptation is to search for a statistical test that seems compatible with the available columns. Resist that shortcut. Not knowing how to analyze the data may be a statistical problem, but it may also reveal a deeper problem with the question, design, measurement, or data collection.

02 · The Short Answer

Do Not Force the Dataset Into a Convenient Test

In Brief

If you have already collected data that you do not know how to analyze, return to the research question and study design before choosing a statistical technique. Determine what the question requires, what evidence the design actually produced, and which analyses are defensible for the data you really have.

Sometimes the problem can be solved with an appropriate analytical method or methodological consultation. In other cases, the dataset cannot answer the original question as intended. When that happens, narrowing the claim, reframing the question transparently, collecting additional data if feasible, or acknowledging the limitation is more defensible than forcing an unsuitable analysis.

03 · What You Need to Know

Diagnose Why the Data Seem Difficult to Analyze

First Ask Whether the Problem Is Statistical or Structural

“I don't know how to analyze this” can describe several very different problems.

You may have a perfectly coherent design but lack experience with the required statistical model. That is primarily an expertise problem. Alternatively, the data may be missing a variable required to answer the question, lack an appropriate comparison, contain only one measurement when change was the intended outcome, or have a sampling structure that cannot support the intended generalization. Those are not problems that a more advanced statistical test necessarily solves.

Analysis problem The data and design can address the question, but you need an appropriate method or expertise to analyze them.
Design or data problem The evidence required by the question was not generated or collected, so changing the statistical method cannot fully repair the mismatch.

Distinguishing those situations should be your first task.

Return to the Original Research Question

Put the statistical software aside temporarily and write the research question in plain language. What exactly are you trying to describe, compare, associate, predict, explain, or estimate?

Then identify what quantity or pattern would constitute an answer. If the question asks whether two groups changed differently over time, you need evidence about change and the difference in that change. If it asks whether an exposure predicts an outcome, you need a clearly defined outcome and predictor, plus an analytical strategy appropriate to the design and predictive objective.

This prevents the available dataset from quietly redefining the research question. The goal is to determine whether the analysis can still match the question and design.

Reconstruct How the Data Were Generated

Before selecting an analysis, document the design as it actually occurred, not merely as it was intended.

Who or what was sampled? How were participants or cases selected? Were groups assigned or naturally occurring? How many groups or conditions exist? Were the same units measured repeatedly? Are observations nested within classrooms, hospitals, organizations, families, or other clusters? When were variables measured? Were any observations matched or paired?

These details determine which analytical methods are plausible and which interpretations are defensible. They can also reveal why a familiar test does not fit.

Create a Variable Map

Next, inventory the variables required for the research question and assign their analytical roles.

Identify the outcome, principal exposure or intervention, predictors, grouping variables, repeated-measure identifiers, potential adjustment variables, and other variables relevant to the intended analysis. Record their measurement scales and timing.

Do not assume that every variable collected belongs in the model. The role of each variable should follow from the question and design, not from the fact that a column exists in the spreadsheet.

Check Whether the Dataset Contains the Evidence the Question Requires

This is the uncomfortable but essential step.

Suppose your question asks whether an intervention improved participants' scores, but you measured scores only after the intervention and have no credible baseline or comparator relevant to the intended inference. You may still be able to describe the observed scores. You may even be able to compare naturally occurring groups if the design permits. But those analyses do not automatically answer whether the intervention caused improvement.

Similarly, if your question concerns long-term outcomes but the study collected only immediate post-intervention measurements, no statistical procedure can create the missing follow-up period.

Watch Out

Do not mistake a computable analysis for an answerable research question. Statistical software can produce output for many combinations of variables even when the design does not support the interpretation you want to make.

Do Not Search for a Test by Variable Type Alone

Typing “test for one categorical and one continuous variable” into a search engine may produce a familiar statistical procedure. It does not establish that the procedure answers your research question.

Statistical choice also depends on how observations are related, how groups were formed, what quantity you intend to estimate, how variables were measured, relevant assumptions, and what inferential claim the design can support.

Decision trees can be useful educational tools, but they should narrow analytical possibilities rather than replace methodological reasoning.

Sometimes You Need a Different Model, Not Different Data

Not every unfamiliar dataset indicates a failed design. You may simply have collected data with a structure that requires a method beyond the techniques you routinely use.

Repeated observations may call for longitudinal methods. Students nested within classes may require an approach that recognizes clustering. Binary, count, ordinal, censored, or time-to-event outcomes may require models suited to those outcome structures. Complex survey designs can require analyses incorporating sampling features.

If the question, design, and measurements are coherent, this is a good reason to seek appropriate expertise rather than simplify the data until they fit a familiar test.

Sometimes the Original Question Cannot Be Answered

There are limits to statistical rescue.

If an essential variable was never measured, you cannot recover its actual values through a more sophisticated model. If the design never included the comparison required by the question, an analysis cannot manufacture that comparison. If the temporal sequence needed for an inference was not observed, statistical adjustment does not create it.

At that point, ask what the dataset can answer legitimately. The defensible research question may be narrower, more descriptive, or more exploratory than originally intended.

That is disappointing, but it is preferable to overstating what the evidence supports.

Be Careful About Rewriting the Question After Seeing the Results

There is an important difference between recognizing that the original question cannot be answered and searching through the dataset for a new question that produces an attractive result.

Secondary and exploratory analyses can be scientifically useful. If a new question emerges from the collected data, label it appropriately and preserve the chronology of the research. Do not imply that a data-driven question or hypothesis was specified before collection if it was not.

This distinction becomes particularly important when the same data are used both to generate and evaluate a new hypothesis.

Seek Statistical Help With the Question, Not Just the Spreadsheet

If you consult a statistician or methodologist, provide more than the dataset.

Bring the research question, protocol or proposal, sampling strategy, instruments, variable definitions, timing of measurements, inclusion and exclusion rules, and any original analysis plan. Explain how data collection actually proceeded and identify deviations from the intended design.

A statistician cannot determine what a column means scientifically merely from its values. Nor can a methodologist infer the intended causal or substantive question from variable names alone.

The ideal time to obtain this input would have been earlier, which is why there are situations where a statistician or methodologist should be involved during study design. But late consultation can still prevent a difficult dataset from becoming a misleading analysis.

Do Not Modify the Data Merely to Make a Preferred Test Possible

Researchers sometimes respond to analytical difficulty by transforming variables, combining categories, deleting observations, splitting continuous measures into groups, or excluding inconvenient cases until a familiar procedure becomes available.

Some transformations and exclusions can be methodologically justified. The problem is making them primarily because they produce a preferred analysis or result.

Every consequential modification should have a defensible rationale. Record what was done, why it was done, and whether the decision occurred before or after examining the relevant results.

Use the Experience to Improve Future Study Planning

A dataset that is difficult to analyze can be an expensive lesson in why data collection and analysis should not be planned separately.

The CDC notes that analysis planning and data collection are interdependent and recommends deciding what to collect and how it will be analyzed before designing the questionnaire. Other methodological guidance similarly emphasizes planning the main analyses at the design stage because doing so clarifies the questions and identifies the data that need to be recorded.

For the next study, plan how the data will be analyzed before collection begins. Even a provisional analysis plan can expose missing variables, inappropriate measurement timing, insufficient sampling, or incompatible data structures before they become permanent features of the dataset.

04 · A Practical Example

When the Dataset Cannot Answer the Question as Originally Written

Hypothetical Example

Did a Workshop Improve Research Skills?

A researcher conducts a research-writing workshop for 80 participants. At the end of the workshop, everyone completes a 20-item research-skills assessment. The stated research question is: “Did participation in the workshop improve researchers' skills?” No pre-workshop assessment was administered, and there is no comparison group.

Original intention The researcher wants to estimate improvement attributable to the workshop.
Available evidence The dataset contains only post-workshop assessment scores from participants who received the workshop.
Diagnosis The data can describe participants' post-workshop performance, but they contain no direct information about how those same participants performed before the workshop and no comparison condition representing what might have occurred without it.
What not to do The researcher should not select an arbitrary benchmark, find that the post-workshop mean differs significantly from it, and present that result as evidence that the workshop caused improvement unless that benchmark has a defensible role in the design and question.
Defensible response The researcher reports what the available data can support, acknowledges that the original improvement question cannot be answered as intended, and redesigns future evaluation to collect evidence appropriate to change and the desired causal inference.

The obstacle was not failure to discover the correct statistical test. The study never generated the evidence needed for the original question. Recognizing that distinction prevents statistical analysis from becoming a substitute for research design.

05 · What Researchers Often Get Wrong

What Not to Do When You Are Unsure How to Analyze the Data

Misconception

“There Must Be a Statistical Test for Whatever Data I Collected”

You can usually calculate something from a dataset, but that does not mean the result answers the original research question. Statistical feasibility and scientific relevance are different requirements.

Misconception

“A Statistician Can Fix the Study After Data Collection”

A statistician may identify an appropriate method for a complex but coherent dataset and can help prevent further analytical errors. Statistical expertise cannot recreate variables, comparison groups, measurement occasions, randomization, or other design features that never existed.

Misconception

“I Should Try Different Tests Until One Is Significant”

Choosing among analyses according to the resulting p-values makes the observed results part of the method-selection process. Alternative analyses should have methodological purposes, such as sensitivity, robustness, or exploration, rather than serving as repeated opportunities to obtain a preferred conclusion.

Misconception

“If the Original Question Cannot Be Answered, the Dataset Is Useless”

Not necessarily. The data may support narrower descriptive, associational, methodological, or exploratory questions. Those analyses can still be valuable if their scope and post hoc status are reported transparently and the conclusions remain within what the design supports.

Misconception

“Changing the Research Question Solves the Problem”

Reframing a question may accurately reflect what the data can address, but it does not erase how or when the new question arose. A question developed after examining the data should not be represented as though it necessarily motivated the original design.

06 · What This Means for You

Work From the Question Back to the Dataset

If you already have difficult-to-analyze data, do not begin by scrolling through a list of statistical tests. Diagnose the study systematically.

A simple decision framework

If the data contain what the research question requires but the analysis is technically unfamiliar
Identify an appropriate method and obtain statistical or methodological assistance if needed.
If the data structure is more complex than expected
Use a method appropriate to that structure rather than altering the data merely to fit a familiar test.
If an essential variable or measurement occasion is missing
Determine whether additional data collection is feasible; otherwise narrow the question and acknowledge the limitation.
If the design cannot support the intended causal or comparative conclusion
Moderate the claim rather than relying on statistical complexity to imply evidence the design did not generate.
If a new question emerges from the dataset
Investigate it transparently as exploratory or post hoc where appropriate and avoid rewriting the history of the study.

The goal is not to salvage the original conclusion at any cost. It is to identify the strongest conclusion that the available evidence can legitimately support.

07 · A Quick Checklist

What to Do Before Running Another Statistical Test

If you already collected the data, check:
I can state the original research question precisely without referring to a statistical test.
I have reconstructed the design as it actually occurred, including sampling, grouping, timing, repeated measurements, and clustering.
I have identified the outcome, principal predictors or exposures, grouping variables, and other analytically relevant variables.
The dataset contains the measurements and comparisons required to answer the original question.
I am choosing the analysis from the question and design rather than from whichever test accepts the available variable types.
Any data transformations, exclusions, recoding, or derived variables have a methodological rationale.
If the required analysis exceeds my expertise, I have sought help and provided the methodological context rather than only the spreadsheet.
If the original question cannot be answered, I have narrowed the interpretation instead of forcing the data to support it.
New analyses or questions developed after examining the data will be identified transparently where that distinction matters.
08 · Frequently Asked Questions

Frequently Asked Questions When You Do Not Know How to Analyze Your Data

Is my study ruined if I did not plan the statistical analysis beforehand?

Not necessarily. First determine whether the design and collected data can answer the research question. A defensible analysis may still be possible. The absence of advance planning becomes more serious when essential measurements or design features are missing or when post hoc analytical flexibility affects how strongly the findings can be interpreted.

Can I ask a statistician to choose the test after I collect the data?

Yes, and appropriate consultation may be very useful. Give the statistician the research question, design, protocol, variable definitions, instruments, sampling information, and original plans as well as the data. Choosing a method requires understanding how the observations were generated and what the study is trying to estimate.

What if no statistical test seems to fit my data?

Determine why. The data may require a model outside the methods you know, the research question may not be sufficiently defined, or the design may not contain the evidence required to answer it. Those situations require different solutions.

Can I change my research question after collecting data?

You can investigate questions that the available data can legitimately address. If the question was developed after collection or examination of the data, preserve that chronology in reporting rather than presenting it as though it necessarily guided the original study.

Can I collect additional data if I discover something is missing?

Sometimes, depending on the study, ethics approval, protocol, sampling framework, resources, and whether additional collection would introduce comparability or selection problems. Do not simply append new observations without considering how the change affects the design and analysis.

Should I use the simplest statistical test available?

Prefer an analysis no more complicated than necessary, but simplicity does not justify ignoring clustering, repeated measurements, outcome structure, confounding, or other features essential to the question. The simplest appropriate method is preferable to either an inappropriate simple test or unnecessary complexity.

What should I do differently in my next study?

Before collecting data, map each research question to the evidence, variables, design features, and analytical approach required to answer it. Where several approaches remain possible, determine how you will choose among defensible analyses before the results themselves can drive that choice.

09 · The Bottom Line

Analyze the Study You Actually Conducted, Not the Study You Wish You Had Conducted

The Bottom Line

If you have collected data that you do not know how to analyze, return to the research question, reconstruct the design, identify what evidence the dataset actually contains, and then choose an analysis appropriate to that evidence.

Sometimes expert statistical help will solve the problem. Sometimes the original question cannot be answered as intended. In that case, a narrower and transparent conclusion is methodologically stronger than forcing the available data through a convenient test and claiming more than the study can support.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes