03 · What You Need to Know
Alignment Runs From the Research Question All the Way to the Conclusion
Start with the relationship in the research question
Consider:
Does academic self-efficacy predict student engagement?
The focal variables are:
X = academic self-efficacy
Y = engagement
Now ask:
To whom does this X–Y relationship apply?
If the answer is individual students, then:
Unit of analysis = student
This simple logic is one of the clearest ways to identify what your study is actually studying.
The subject of the question often reveals the unit
Consider these questions:
| Research question |
Likely unit of analysis |
| Which students are more likely to persist? |
Student |
| Which classrooms show greater student engagement? |
Classroom |
| Do universities with formal AI policies have higher adoption rates? |
University |
| Do countries with greater research investment produce more publications? |
Country |
| Are internationally coauthored articles cited more frequently? |
Research article |
The unit is not chosen because it is convenient to analyze. It follows from what kind of entity the research question wants to compare.
A useful diagnostic is: “Who or what can have both X and Y?”
Suppose your question is:
Does faculty AI self-efficacy predict faculty AI adoption?
An individual faculty member can have both:
- a self-efficacy score;
- an adoption score.
Faculty member is therefore a coherent unit.
Now suppose the question is:
Does university AI policy predict institutional adoption rate?
A university can have both:
- a policy status;
- an institutional adoption rate.
University is the coherent unit.
The question should not assign variables to an entity that cannot possess them
Consider:
Does university AI policy increase faculty self-efficacy?
The university can possess the policy.
The faculty member possesses the self-efficacy score.
No single entity possesses both variables at the same level.
That does not make the question invalid.
It makes it cross-level.
The correct structure is:
University-level predictor → Individual-level outcome
This is the kind of question examined in a cross-level relationship.
Research-question alignment does not require every variable to be at the same level
A multilevel question can be perfectly coherent.
The requirement is that the levels are explicit.
For example:
Do faculty members working in universities with clearer AI policies report greater individual AI self-efficacy?
Now the structure is visible:
University policy = organizational level
Faculty self-efficacy = individual level
The question tells you that faculty are nested within universities.
An ambiguous noun can hide a mismatch
Consider:
Does institutional support improve adoption?
What is “institutional support”?
What is “adoption”?
The question could mean:
Individual level: Does perceived support predict one faculty member's adoption?
Group level: Does university support climate predict institutional adoption rate?
Cross level: Does university support climate predict individual faculty adoption?
Those are three different studies.
The unit should therefore be visible in the wording
Instead of:
Does support influence adoption?
write:
Do faculty members who perceive greater institutional support report greater individual AI adoption?
or:
Do universities with stronger AI-support climates have higher institutional AI adoption rates?
or:
Do faculty members in universities with stronger AI-support climates report greater individual AI adoption?
The nouns make the analytical level difficult to miss.
Next, check whether the variables actually match the unit
Suppose your unit is the university.
Your variables should meaningfully characterize universities.
Examples include:
- formal policy status;
- university budget;
- institutional adoption rate;
- organizational size;
- validated university-level climate.
Individual self-efficacy or personal attitudes do not automatically become university-level variables merely because faculty work at universities.
The measurement referent is a useful clue
Compare:
“I have adequate access to AI tools.”
with:
“Faculty in this university have adequate access to AI tools.”
The first refers to the individual.
The second refers to the university context.
Neither wording is automatically superior. The correct choice depends on the construct your question requires.
The unit of observation should also fit the inferential plan
Your unit of observation can differ from your unit of analysis.
Faculty members may provide observations used to characterize universities.
Documents may provide observations used to characterize organizations.
Repeated measurements may provide observations used to characterize people.
The distinction is legitimate when researchers explain the bridge between what is observed and what is analyzed.
If individuals provide group-level evidence, the composition process must fit
Suppose your unit is the university and institutional climate is measured from faculty responses.
You should ask:
- Is climate theoretically a university-level construct?
- Do the items use an appropriate institutional referent?
- Are faculty knowledgeable about the target?
- Is aggregation justified?
- Is agreement relevant and adequate?
- Are group scores reliable?
A research question and unit cannot be considered aligned if the variables themselves do not validly represent that unit.
Then check whether the sampling matches the unit
Suppose your question concerns universities.
You survey 3,000 faculty members from four universities.
The individual sample is large.
The university sample is four.
If you intend to generalize about variation among universities, the design provides very limited higher-level information.
Watch Out
The sample size that matters for a relationship depends on the level where that relationship varies. Thousands of lower-level observations do not create thousands of higher-level units.
A matching question and unit still need enough cases
You may correctly identify universities as the unit but still have too few universities for the intended model.
Unit alignment and adequate sampling are related but distinct requirements.
A conceptually correct design can still be empirically weak.
Check whether one row in the analytical dataset represents the right entity
A practical diagnostic is:
What does one row mean at the stage where the focal relationship is estimated?
If the research question compares universities and your fully aggregated analysis has one row per university, the structure is straightforward.
If the question concerns individuals nested in universities, one row may remain one individual while a multilevel model recognizes the shared university context.
The important point is not forcing one row to equal the unit in every design. It is understanding what the cases and grouping structure represent.
Repeated observations provide a useful example
Suppose 100 people report stress for 30 days.
There are 3,000 person-day records.
Question A:
Are generally more stressed people less satisfied with life?
This is a between-person question.
Question B:
On days when a person is more stressed than usual, does that person sleep worse?
This is a within-person question using repeated occasions nested within individuals.
The same dataset supports different analytical levels because the questions specify different sources of variation.
The number of rows does not define the unit
Having 3,000 rows does not mean the study has 3,000 independent people.
Likewise, duplicating a university-level policy variable across 100 faculty rows does not create 100 university observations.
Research-question alignment therefore requires understanding the hierarchy rather than reading N directly from the spreadsheet.
Next, check whether the statistical model matches the unit
Suppose schools are randomized to an intervention but outcomes are measured on students.
The design contains:
Experimental assignment = school level
Outcome observation = student level
If the analysis treats every student as though treatment had been independently randomized to them, the statistical model does not match the design.
Unit-of-analysis errors can therefore invalidate inference even when the substantive question itself is sensible.
Ignoring dependence is a common sign of mismatch
If students are nested within classrooms, employees within organizations, or repeated measures within individuals, observations from the same higher-level unit may be correlated.
Ordinary methods that assume independence may misstate uncertainty.
Recognizing nested or hierarchical data is therefore part of matching the analysis to the research question.
But using multilevel modeling does not automatically fix conceptual mismatch
Suppose the research question claims to study organizational climate, but the questionnaire measures only individual satisfaction.
Adding a random intercept for organization does not transform satisfaction into climate.
Statistical sophistication cannot repair a construct that exists at the wrong level.
Check whether the predictor and outcome are interpreted at their actual levels
Suppose a multilevel model finds:
University infrastructure predicts individual faculty adoption.
The correct conclusion is cross-level:
Faculty in universities with stronger infrastructure tend to report greater adoption.
It is not equivalent to:
Faculty with stronger infrastructure adopt more.
Faculty members do not individually possess the university's institutional infrastructure score in the same sense.
Next, inspect the conceptual framework
Every arrow should answer:
Between what kinds of entities does this relationship operate?
If a diagram connects:
Institutional support → Faculty self-efficacy → University adoption
the model moves:
university → individual → university
That is a complex multilevel pathway.
It may be defensible, but the theory must explain how the process moves across levels.
Simply drawing arrows does not solve the cross-level composition problem.
A mediator must also have a coherent level
Suppose university policy affects individual self-efficacy, which is proposed to mediate an effect on institutional adoption.
The mediator is individual-level while the final outcome is organizational-level.
This requires a theory of how changes among individuals aggregate into institutional outcomes.
Ordinary single-level mediation logic cannot simply be copied into the model without considering the multilevel structure.
The same applies to moderators
A university-level variable can moderate an individual-level relationship.
For example:
Faculty self-efficacy → Faculty adoption
may depend on:
University infrastructure.
This is coherent cross-level moderation when explicitly theorized.
Check whether the level changes when the research question changes
Compare:
Why do some faculty adopt AI more than others?
with:
Why do some universities have higher AI adoption than others?
Changing one noun fundamentally changes the explanatory target.
The predictors that belong in the study may change accordingly.
This is why “same topic” does not mean “same study”
Both questions may concern generative AI adoption.
One investigates individuals.
The other investigates organizations.
Topic is not unit of analysis.
The theoretical literature should be relevant to the level too
If your outcome is organizational adoption, literature explaining individual technology acceptance may provide useful mechanisms but may not be sufficient as the entire theoretical basis.
You may also need literature on organizational change, innovation diffusion, governance, institutional capability, or other higher-level processes.
Theoretical alignment is therefore partly a level-of-analysis issue.
Check whether the hypotheses name the correct entities
An ambiguous hypothesis:
Institutional support positively influences AI adoption.
A clearer individual hypothesis:
Faculty members who perceive greater institutional support will report greater individual AI adoption.
A clearer organizational hypothesis:
Universities with stronger AI-support climates will have higher institutional adoption rates.
A clearer cross-level hypothesis:
Faculty members in universities with stronger AI-support climates will report greater individual AI adoption.
Each hypothesis now identifies its unit structure.
Check whether your sample description matches the claims
If the primary analysis concerns universities, reporting only:
N = 2,000 faculty participants
is incomplete.
Readers need to know:
2,000 faculty nested in how many universities?
The number of higher-level units is essential for understanding the evidence behind university-level relationships.
Check whether your result statement changes nouns
Suppose the analysis compares universities and the results say:
Universities with higher average faculty self-efficacy showed greater adoption.
Then the discussion says:
Therefore, highly self-efficacious faculty are more likely to adopt AI.
The noun changed from universities to faculty.
That should trigger an immediate level-of-inference check.
Noun drift is a practical warning sign
Watch for shifts such as:
schools → students
organizations → employees
neighborhoods → residents
countries → citizens
teams → team members
Moving downward without supporting evidence risks the ecological fallacy.
Noun drift can also move upward
If an individual study concludes:
Faculty with greater self-efficacy adopt AI more frequently.
and then recommends:
Universities with greater faculty self-efficacy will therefore achieve institutional AI transformation.
the inference has moved upward.
This risks the atomistic fallacy.
Check whether aggregation changes the research question
Suppose your individual question is not supported, so you aggregate by university.
A significant university-level result does not rescue the original individual hypothesis.
You have answered another question.
The distinction matters because relationships can legitimately differ across levels.
There is no universal “correct” unit independent of the question
Is faculty member or university the correct unit for AI adoption research?
Either can be correct.
It depends on whether you want to understand:
individual behavior
or:
institutional differences.
A research topic does not come with one permanent unit of analysis attached to it.
The same phenomenon can support several coherent units
A study of educational technology could examine:
- students;
- teachers;
- classrooms;
- schools;
- universities;
- countries.
Each unit produces a different scientific question.
A study can even legitimately contain multiple units of analysis if the relationships are clearly separated or explicitly connected.
Unit alignment should be checked before choosing statistical tests
Researchers sometimes jump from:
“I have survey data.”
to:
“I will use regression.”
The better sequence is:
Research question → unit → variables and levels → sampling → data structure → analysis.
The statistical method comes after defining what kind of cases and relationships need to be analyzed.
The research design should be traceable as one chain
A coherent study should let you trace:
Research question What relationship am I trying to understand?
Unit of analysis To whom or to what does that relationship apply?
Variables Do the measurements validly describe that entity or the relevant cross-level entities?
Sampling Did I observe enough appropriate cases at the relevant level?
Analysis Does the statistical or qualitative method preserve the actual data structure?
Conclusion Am I making claims about the same entities the analysis supports?
If one step changes levels unexpectedly, investigate the mismatch.
The strongest diagnostic is to write four sentences
Complete these before collecting data:
My research question asks about ______.
My primary unit of analysis is ______.
My data will provide valid information about that unit because ______.
My conclusions will therefore be about ______.
If the first, second, and fourth blanks describe different entities, you either have a deliberate multilevel design or an alignment problem.
A fifth sentence helps in multilevel studies
Add:
Variables operating at a different level are ______, and they connect to the primary unit through ______.
This forces the cross-level mechanism into the open.
A simple unit-of-analysis matrix can prevent many errors
| Design element |
Question to answer |
| Research question |
Who or what is the claim about? |
| Predictor |
What entity possesses X? |
| Outcome |
What entity possesses Y? |
| Observations |
Where does the information come from? |
| Sampling |
What entities were actually sampled? |
| Data structure |
Are observations independent, nested, repeated, or cross-classified? |
| Analysis |
At what level is each coefficient or comparison estimated? |
| Conclusion |
What entity does the final claim describe? |
A mismatch does not always require abandoning the study
Sometimes the research question can be rewritten to match what the data genuinely support.
Suppose you wanted to ask:
Do universities with stronger institutional support have higher AI adoption?
but your data contain only individual perceptions and individual adoption.
A defensible revision may be:
Are faculty members who perceive greater institutional support more likely to report personal AI adoption?
The revised study is narrower, but the question now matches the evidence.
Alternatively, the design can be strengthened rather than the question weakened
If the organizational question is scientifically important, you can redesign the study to include:
- more universities;
- institution-level sampling;
- organizational measures;
- multiple respondents per institution;
- validated aggregation procedures;
- multilevel modeling.
The correct response to a mismatch depends on which component reflects the real scientific objective.
The question should drive the design, not the available spreadsheet
If the data you happen to possess can answer a useful but different question, ask that question honestly.
Do not force the dataset to support a grander level of inference simply because the original title sounds more ambitious.
Alignment matters in qualitative research too
Suppose researchers interview employees but claim to explain organizational culture.
The employees can provide evidence about the organization, but the researchers should show how multiple accounts, observations, and documents support the organizational interpretation.
If each interview is analyzed only as one person's experience, an organization-level conclusion may exceed the analytical evidence.
Case-study boundaries function similarly
If the case is a university, researchers need evidence capable of characterizing the university as a case.
If the case is a faculty member, the evidence should support conclusions about that person's experiences or practices.
The question “What is the case?” is therefore closely related to “What is the unit of analysis?”
Unit alignment is ultimately about inferential discipline
The purpose is not to police terminology.
It is to prevent a common methodological failure:
collecting evidence about one kind of thing and making conclusions about another without explaining the bridge.
When the bridge is theoretically and methodologically explicit, multilevel research can be powerful.
When it is hidden, even sophisticated analysis can answer a different question from the one printed at the top of the manuscript.