03 · What You Need to Know
Diagnose Why the Data Seem Difficult to Analyze
First Ask Whether the Problem Is Statistical or Structural
“I don't know how to analyze this” can describe several very different problems.
You may have a perfectly coherent design but lack experience with the required statistical model. That is primarily an expertise problem. Alternatively, the data may be missing a variable required to answer the question, lack an appropriate comparison, contain only one measurement when change was the intended outcome, or have a sampling structure that cannot support the intended generalization. Those are not problems that a more advanced statistical test necessarily solves.
Analysis problem
The data and design can address the question, but you need an appropriate method or expertise to analyze them.
Design or data problem
The evidence required by the question was not generated or collected, so changing the statistical method cannot fully repair the mismatch.
Distinguishing those situations should be your first task.
Return to the Original Research Question
Put the statistical software aside temporarily and write the research question in plain language. What exactly are you trying to describe, compare, associate, predict, explain, or estimate?
Then identify what quantity or pattern would constitute an answer. If the question asks whether two groups changed differently over time, you need evidence about change and the difference in that change. If it asks whether an exposure predicts an outcome, you need a clearly defined outcome and predictor, plus an analytical strategy appropriate to the design and predictive objective.
This prevents the available dataset from quietly redefining the research question. The goal is to determine whether the analysis can still match the question and design.
Reconstruct How the Data Were Generated
Before selecting an analysis, document the design as it actually occurred, not merely as it was intended.
Who or what was sampled? How were participants or cases selected? Were groups assigned or naturally occurring? How many groups or conditions exist? Were the same units measured repeatedly? Are observations nested within classrooms, hospitals, organizations, families, or other clusters? When were variables measured? Were any observations matched or paired?
These details determine which analytical methods are plausible and which interpretations are defensible. They can also reveal why a familiar test does not fit.
Create a Variable Map
Next, inventory the variables required for the research question and assign their analytical roles.
Identify the outcome, principal exposure or intervention, predictors, grouping variables, repeated-measure identifiers, potential adjustment variables, and other variables relevant to the intended analysis. Record their measurement scales and timing.
Do not assume that every variable collected belongs in the model. The role of each variable should follow from the question and design, not from the fact that a column exists in the spreadsheet.
Check Whether the Dataset Contains the Evidence the Question Requires
This is the uncomfortable but essential step.
Suppose your question asks whether an intervention improved participants' scores, but you measured scores only after the intervention and have no credible baseline or comparator relevant to the intended inference. You may still be able to describe the observed scores. You may even be able to compare naturally occurring groups if the design permits. But those analyses do not automatically answer whether the intervention caused improvement.
Similarly, if your question concerns long-term outcomes but the study collected only immediate post-intervention measurements, no statistical procedure can create the missing follow-up period.
Watch Out
Do not mistake a computable analysis for an answerable research question. Statistical software can produce output for many combinations of variables even when the design does not support the interpretation you want to make.
Do Not Search for a Test by Variable Type Alone
Typing “test for one categorical and one continuous variable” into a search engine may produce a familiar statistical procedure. It does not establish that the procedure answers your research question.
Statistical choice also depends on how observations are related, how groups were formed, what quantity you intend to estimate, how variables were measured, relevant assumptions, and what inferential claim the design can support.
Decision trees can be useful educational tools, but they should narrow analytical possibilities rather than replace methodological reasoning.
Sometimes You Need a Different Model, Not Different Data
Not every unfamiliar dataset indicates a failed design. You may simply have collected data with a structure that requires a method beyond the techniques you routinely use.
Repeated observations may call for longitudinal methods. Students nested within classes may require an approach that recognizes clustering. Binary, count, ordinal, censored, or time-to-event outcomes may require models suited to those outcome structures. Complex survey designs can require analyses incorporating sampling features.
If the question, design, and measurements are coherent, this is a good reason to seek appropriate expertise rather than simplify the data until they fit a familiar test.
Sometimes the Original Question Cannot Be Answered
There are limits to statistical rescue.
If an essential variable was never measured, you cannot recover its actual values through a more sophisticated model. If the design never included the comparison required by the question, an analysis cannot manufacture that comparison. If the temporal sequence needed for an inference was not observed, statistical adjustment does not create it.
At that point, ask what the dataset can answer legitimately. The defensible research question may be narrower, more descriptive, or more exploratory than originally intended.
That is disappointing, but it is preferable to overstating what the evidence supports.
Be Careful About Rewriting the Question After Seeing the Results
There is an important difference between recognizing that the original question cannot be answered and searching through the dataset for a new question that produces an attractive result.
Secondary and exploratory analyses can be scientifically useful. If a new question emerges from the collected data, label it appropriately and preserve the chronology of the research. Do not imply that a data-driven question or hypothesis was specified before collection if it was not.
This distinction becomes particularly important when the same data are used both to generate and evaluate a new hypothesis.
Seek Statistical Help With the Question, Not Just the Spreadsheet
If you consult a statistician or methodologist, provide more than the dataset.
Bring the research question, protocol or proposal, sampling strategy, instruments, variable definitions, timing of measurements, inclusion and exclusion rules, and any original analysis plan. Explain how data collection actually proceeded and identify deviations from the intended design.
A statistician cannot determine what a column means scientifically merely from its values. Nor can a methodologist infer the intended causal or substantive question from variable names alone.
The ideal time to obtain this input would have been earlier, which is why there are situations where a statistician or methodologist should be involved during study design. But late consultation can still prevent a difficult dataset from becoming a misleading analysis.
Do Not Modify the Data Merely to Make a Preferred Test Possible
Researchers sometimes respond to analytical difficulty by transforming variables, combining categories, deleting observations, splitting continuous measures into groups, or excluding inconvenient cases until a familiar procedure becomes available.
Some transformations and exclusions can be methodologically justified. The problem is making them primarily because they produce a preferred analysis or result.
Every consequential modification should have a defensible rationale. Record what was done, why it was done, and whether the decision occurred before or after examining the relevant results.
Use the Experience to Improve Future Study Planning
A dataset that is difficult to analyze can be an expensive lesson in why data collection and analysis should not be planned separately.
The CDC notes that analysis planning and data collection are interdependent and recommends deciding what to collect and how it will be analyzed before designing the questionnaire. Other methodological guidance similarly emphasizes planning the main analyses at the design stage because doing so clarifies the questions and identifies the data that need to be recorded.
For the next study, plan how the data will be analyzed before collection begins. Even a provisional analysis plan can expose missing variables, inappropriate measurement timing, insufficient sampling, or incompatible data structures before they become permanent features of the dataset.