Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Know Whether Your Research Question and Unit of Analysis Actually Match?

Your research question and unit of analysis match when the entities named in the question are the entities your data and analysis can legitimately compare. Checking that alignment early prevents mismatched variables, inflated sample sizes, inappropriate aggregation, and conclusions made at the wrong level.

131
Matching the Research Question and Unit of Analysis Guide 131 of 223
01 · The Question

Your Question Says “Universities,” but Your Dataset Contains Faculty Members. Do They Match?

Consider this research question:

How does institutional support influence university adoption of generative AI?

Now imagine the study surveys individual faculty members and measures:

  • their personal perception of support;
  • their own AI self-efficacy;
  • their personal classroom AI use.

Nothing is necessarily wrong with those measurements.

But do they answer a question about universities?

Perhaps not.

The data may actually support a different question:

Are faculty members who perceive greater institutional support more likely to use generative AI in their own teaching?

The first question sounds organizational. The second is individual.

This mismatch is easy to miss because both studies use words such as “institutional support” and “AI adoption.” Yet they make claims about different entities.

Checking whether the research question and unit of analysis match means asking whether the people, groups, organizations, documents, events, or other entities named in the question are actually the entities your study is capable of comparing and making conclusions about.

02 · The Short Answer

The Noun in Your Question Should Match the Entity Your Analysis Can Support

In Brief

Your research question and unit of analysis match when the relationship in the question applies to the same kind of entity represented by the analytical cases and supported by the measurements, sampling, and data structure.

Identify the predictor and outcome, then ask: “To whom or to what does this relationship apply?” If the answer is students, your data and conclusions should support student-level inference. If it is schools, organizations, countries, documents, or another entity, the design must provide valid information at that level. In multilevel questions, each variable and pathway should be assigned explicitly to its level.

03 · What You Need to Know

Alignment Runs From the Research Question All the Way to the Conclusion

Start with the relationship in the research question

Consider:

Does academic self-efficacy predict student engagement?

The focal variables are:

X = academic self-efficacy

Y = engagement

Now ask:

To whom does this X–Y relationship apply?

If the answer is individual students, then:

Unit of analysis = student

This simple logic is one of the clearest ways to identify what your study is actually studying.

The subject of the question often reveals the unit

Consider these questions:

Research question Likely unit of analysis
Which students are more likely to persist? Student
Which classrooms show greater student engagement? Classroom
Do universities with formal AI policies have higher adoption rates? University
Do countries with greater research investment produce more publications? Country
Are internationally coauthored articles cited more frequently? Research article

The unit is not chosen because it is convenient to analyze. It follows from what kind of entity the research question wants to compare.

A useful diagnostic is: “Who or what can have both X and Y?”

Suppose your question is:

Does faculty AI self-efficacy predict faculty AI adoption?

An individual faculty member can have both:

  • a self-efficacy score;
  • an adoption score.

Faculty member is therefore a coherent unit.

Now suppose the question is:

Does university AI policy predict institutional adoption rate?

A university can have both:

  • a policy status;
  • an institutional adoption rate.

University is the coherent unit.

The question should not assign variables to an entity that cannot possess them

Consider:

Does university AI policy increase faculty self-efficacy?

The university can possess the policy.

The faculty member possesses the self-efficacy score.

No single entity possesses both variables at the same level.

That does not make the question invalid.

It makes it cross-level.

The correct structure is:

University-level predictor → Individual-level outcome

This is the kind of question examined in a cross-level relationship.

Research-question alignment does not require every variable to be at the same level

A multilevel question can be perfectly coherent.

The requirement is that the levels are explicit.

For example:

Do faculty members working in universities with clearer AI policies report greater individual AI self-efficacy?

Now the structure is visible:

University policy = organizational level

Faculty self-efficacy = individual level

The question tells you that faculty are nested within universities.

An ambiguous noun can hide a mismatch

Consider:

Does institutional support improve adoption?

What is “institutional support”?

What is “adoption”?

The question could mean:

Individual level: Does perceived support predict one faculty member's adoption?

Group level: Does university support climate predict institutional adoption rate?

Cross level: Does university support climate predict individual faculty adoption?

Those are three different studies.

The unit should therefore be visible in the wording

Instead of:

Does support influence adoption?

write:

Do faculty members who perceive greater institutional support report greater individual AI adoption?

or:

Do universities with stronger AI-support climates have higher institutional AI adoption rates?

or:

Do faculty members in universities with stronger AI-support climates report greater individual AI adoption?

The nouns make the analytical level difficult to miss.

Next, check whether the variables actually match the unit

Suppose your unit is the university.

Your variables should meaningfully characterize universities.

Examples include:

  • formal policy status;
  • university budget;
  • institutional adoption rate;
  • organizational size;
  • validated university-level climate.

Individual self-efficacy or personal attitudes do not automatically become university-level variables merely because faculty work at universities.

The measurement referent is a useful clue

Compare:

“I have adequate access to AI tools.”

with:

“Faculty in this university have adequate access to AI tools.”

The first refers to the individual.

The second refers to the university context.

Neither wording is automatically superior. The correct choice depends on the construct your question requires.

The unit of observation should also fit the inferential plan

Your unit of observation can differ from your unit of analysis.

Faculty members may provide observations used to characterize universities.

Documents may provide observations used to characterize organizations.

Repeated measurements may provide observations used to characterize people.

The distinction is legitimate when researchers explain the bridge between what is observed and what is analyzed.

If individuals provide group-level evidence, the composition process must fit

Suppose your unit is the university and institutional climate is measured from faculty responses.

You should ask:

  • Is climate theoretically a university-level construct?
  • Do the items use an appropriate institutional referent?
  • Are faculty knowledgeable about the target?
  • Is aggregation justified?
  • Is agreement relevant and adequate?
  • Are group scores reliable?

A research question and unit cannot be considered aligned if the variables themselves do not validly represent that unit.

Then check whether the sampling matches the unit

Suppose your question concerns universities.

You survey 3,000 faculty members from four universities.

The individual sample is large.

The university sample is four.

If you intend to generalize about variation among universities, the design provides very limited higher-level information.

Watch Out

The sample size that matters for a relationship depends on the level where that relationship varies. Thousands of lower-level observations do not create thousands of higher-level units.

A matching question and unit still need enough cases

You may correctly identify universities as the unit but still have too few universities for the intended model.

Unit alignment and adequate sampling are related but distinct requirements.

A conceptually correct design can still be empirically weak.

Check whether one row in the analytical dataset represents the right entity

A practical diagnostic is:

What does one row mean at the stage where the focal relationship is estimated?

If the research question compares universities and your fully aggregated analysis has one row per university, the structure is straightforward.

If the question concerns individuals nested in universities, one row may remain one individual while a multilevel model recognizes the shared university context.

The important point is not forcing one row to equal the unit in every design. It is understanding what the cases and grouping structure represent.

Repeated observations provide a useful example

Suppose 100 people report stress for 30 days.

There are 3,000 person-day records.

Question A:

Are generally more stressed people less satisfied with life?

This is a between-person question.

Question B:

On days when a person is more stressed than usual, does that person sleep worse?

This is a within-person question using repeated occasions nested within individuals.

The same dataset supports different analytical levels because the questions specify different sources of variation.

The number of rows does not define the unit

Having 3,000 rows does not mean the study has 3,000 independent people.

Likewise, duplicating a university-level policy variable across 100 faculty rows does not create 100 university observations.

Research-question alignment therefore requires understanding the hierarchy rather than reading N directly from the spreadsheet.

Next, check whether the statistical model matches the unit

Suppose schools are randomized to an intervention but outcomes are measured on students.

The design contains:

Experimental assignment = school level

Outcome observation = student level

If the analysis treats every student as though treatment had been independently randomized to them, the statistical model does not match the design.

Unit-of-analysis errors can therefore invalidate inference even when the substantive question itself is sensible.

Ignoring dependence is a common sign of mismatch

If students are nested within classrooms, employees within organizations, or repeated measures within individuals, observations from the same higher-level unit may be correlated.

Ordinary methods that assume independence may misstate uncertainty.

Recognizing nested or hierarchical data is therefore part of matching the analysis to the research question.

But using multilevel modeling does not automatically fix conceptual mismatch

Suppose the research question claims to study organizational climate, but the questionnaire measures only individual satisfaction.

Adding a random intercept for organization does not transform satisfaction into climate.

Statistical sophistication cannot repair a construct that exists at the wrong level.

Check whether the predictor and outcome are interpreted at their actual levels

Suppose a multilevel model finds:

University infrastructure predicts individual faculty adoption.

The correct conclusion is cross-level:

Faculty in universities with stronger infrastructure tend to report greater adoption.

It is not equivalent to:

Faculty with stronger infrastructure adopt more.

Faculty members do not individually possess the university's institutional infrastructure score in the same sense.

Next, inspect the conceptual framework

Every arrow should answer:

Between what kinds of entities does this relationship operate?

If a diagram connects:

Institutional support → Faculty self-efficacy → University adoption

the model moves:

university → individual → university

That is a complex multilevel pathway.

It may be defensible, but the theory must explain how the process moves across levels.

Simply drawing arrows does not solve the cross-level composition problem.

A mediator must also have a coherent level

Suppose university policy affects individual self-efficacy, which is proposed to mediate an effect on institutional adoption.

The mediator is individual-level while the final outcome is organizational-level.

This requires a theory of how changes among individuals aggregate into institutional outcomes.

Ordinary single-level mediation logic cannot simply be copied into the model without considering the multilevel structure.

The same applies to moderators

A university-level variable can moderate an individual-level relationship.

For example:

Faculty self-efficacy → Faculty adoption

may depend on:

University infrastructure.

This is coherent cross-level moderation when explicitly theorized.

Check whether the level changes when the research question changes

Compare:

Why do some faculty adopt AI more than others?

with:

Why do some universities have higher AI adoption than others?

Changing one noun fundamentally changes the explanatory target.

The predictors that belong in the study may change accordingly.

This is why “same topic” does not mean “same study”

Both questions may concern generative AI adoption.

One investigates individuals.

The other investigates organizations.

Topic is not unit of analysis.

The theoretical literature should be relevant to the level too

If your outcome is organizational adoption, literature explaining individual technology acceptance may provide useful mechanisms but may not be sufficient as the entire theoretical basis.

You may also need literature on organizational change, innovation diffusion, governance, institutional capability, or other higher-level processes.

Theoretical alignment is therefore partly a level-of-analysis issue.

Check whether the hypotheses name the correct entities

An ambiguous hypothesis:

Institutional support positively influences AI adoption.

A clearer individual hypothesis:

Faculty members who perceive greater institutional support will report greater individual AI adoption.

A clearer organizational hypothesis:

Universities with stronger AI-support climates will have higher institutional adoption rates.

A clearer cross-level hypothesis:

Faculty members in universities with stronger AI-support climates will report greater individual AI adoption.

Each hypothesis now identifies its unit structure.

Check whether your sample description matches the claims

If the primary analysis concerns universities, reporting only:

N = 2,000 faculty participants

is incomplete.

Readers need to know:

2,000 faculty nested in how many universities?

The number of higher-level units is essential for understanding the evidence behind university-level relationships.

Check whether your result statement changes nouns

Suppose the analysis compares universities and the results say:

Universities with higher average faculty self-efficacy showed greater adoption.

Then the discussion says:

Therefore, highly self-efficacious faculty are more likely to adopt AI.

The noun changed from universities to faculty.

That should trigger an immediate level-of-inference check.

Noun drift is a practical warning sign

Watch for shifts such as:

schools → students

organizations → employees

neighborhoods → residents

countries → citizens

teams → team members

Moving downward without supporting evidence risks the ecological fallacy.

Noun drift can also move upward

If an individual study concludes:

Faculty with greater self-efficacy adopt AI more frequently.

and then recommends:

Universities with greater faculty self-efficacy will therefore achieve institutional AI transformation.

the inference has moved upward.

This risks the atomistic fallacy.

Check whether aggregation changes the research question

Suppose your individual question is not supported, so you aggregate by university.

A significant university-level result does not rescue the original individual hypothesis.

You have answered another question.

The distinction matters because relationships can legitimately differ across levels.

There is no universal “correct” unit independent of the question

Is faculty member or university the correct unit for AI adoption research?

Either can be correct.

It depends on whether you want to understand:

individual behavior

or:

institutional differences.

A research topic does not come with one permanent unit of analysis attached to it.

The same phenomenon can support several coherent units

A study of educational technology could examine:

  • students;
  • teachers;
  • classrooms;
  • schools;
  • universities;
  • countries.

Each unit produces a different scientific question.

A study can even legitimately contain multiple units of analysis if the relationships are clearly separated or explicitly connected.

Unit alignment should be checked before choosing statistical tests

Researchers sometimes jump from:

“I have survey data.”

to:

“I will use regression.”

The better sequence is:

Research question → unit → variables and levels → sampling → data structure → analysis.

The statistical method comes after defining what kind of cases and relationships need to be analyzed.

The research design should be traceable as one chain

A coherent study should let you trace:

Research question What relationship am I trying to understand?
Unit of analysis To whom or to what does that relationship apply?
Variables Do the measurements validly describe that entity or the relevant cross-level entities?
Sampling Did I observe enough appropriate cases at the relevant level?
Analysis Does the statistical or qualitative method preserve the actual data structure?
Conclusion Am I making claims about the same entities the analysis supports?

If one step changes levels unexpectedly, investigate the mismatch.

The strongest diagnostic is to write four sentences

Complete these before collecting data:

My research question asks about ______.

My primary unit of analysis is ______.

My data will provide valid information about that unit because ______.

My conclusions will therefore be about ______.

If the first, second, and fourth blanks describe different entities, you either have a deliberate multilevel design or an alignment problem.

A fifth sentence helps in multilevel studies

Add:

Variables operating at a different level are ______, and they connect to the primary unit through ______.

This forces the cross-level mechanism into the open.

A simple unit-of-analysis matrix can prevent many errors

Design element Question to answer
Research question Who or what is the claim about?
Predictor What entity possesses X?
Outcome What entity possesses Y?
Observations Where does the information come from?
Sampling What entities were actually sampled?
Data structure Are observations independent, nested, repeated, or cross-classified?
Analysis At what level is each coefficient or comparison estimated?
Conclusion What entity does the final claim describe?

A mismatch does not always require abandoning the study

Sometimes the research question can be rewritten to match what the data genuinely support.

Suppose you wanted to ask:

Do universities with stronger institutional support have higher AI adoption?

but your data contain only individual perceptions and individual adoption.

A defensible revision may be:

Are faculty members who perceive greater institutional support more likely to report personal AI adoption?

The revised study is narrower, but the question now matches the evidence.

Alternatively, the design can be strengthened rather than the question weakened

If the organizational question is scientifically important, you can redesign the study to include:

  • more universities;
  • institution-level sampling;
  • organizational measures;
  • multiple respondents per institution;
  • validated aggregation procedures;
  • multilevel modeling.

The correct response to a mismatch depends on which component reflects the real scientific objective.

The question should drive the design, not the available spreadsheet

If the data you happen to possess can answer a useful but different question, ask that question honestly.

Do not force the dataset to support a grander level of inference simply because the original title sounds more ambitious.

Alignment matters in qualitative research too

Suppose researchers interview employees but claim to explain organizational culture.

The employees can provide evidence about the organization, but the researchers should show how multiple accounts, observations, and documents support the organizational interpretation.

If each interview is analyzed only as one person's experience, an organization-level conclusion may exceed the analytical evidence.

Case-study boundaries function similarly

If the case is a university, researchers need evidence capable of characterizing the university as a case.

If the case is a faculty member, the evidence should support conclusions about that person's experiences or practices.

The question “What is the case?” is therefore closely related to “What is the unit of analysis?”

Unit alignment is ultimately about inferential discipline

The purpose is not to police terminology.

It is to prevent a common methodological failure:

collecting evidence about one kind of thing and making conclusions about another without explaining the bridge.

When the bridge is theoretically and methodologically explicit, multilevel research can be powerful.

When it is hidden, even sophisticated analysis can answer a different question from the one printed at the top of the manuscript.

04 · A Practical Example

Diagnosing a Research Question That Sounds Institutional but Uses Individual Data

Hypothetical Example

University support and generative AI adoption

A researcher proposes: “How does institutional support influence university adoption of generative AI?” The survey is administered to individual faculty members and asks about personal perceptions of support and personal classroom AI use.

Question-level diagnosis The phrases “institutional support” and “university adoption” suggest an organizational-level study.
Measurement diagnosis The actual variables describe individual faculty perceptions and individual faculty behavior.
Mismatch The study is collecting individual-level evidence but framing the outcome as though universities themselves were the analytical cases.
Option 1: revise the question “Are faculty members who perceive greater institutional support more likely to report personal classroom AI adoption?”
Option 2: revise the design Sample multiple universities deliberately, measure institutional support at the university level or justify its aggregation, define an institutional adoption outcome, and use an organizational or multilevel analysis.

Neither solution is inherently superior.

The correct choice depends on whether the scientific target is faculty behavior or institutional variation.

05 · What Researchers Often Get Wrong

Common Signs That the Research Question and Unit Do Not Match

Misconception

The unit of analysis is simply whoever answers my survey

No. Respondents may provide observations used to study classrooms, organizations, households, communities, or other entities. The unit follows the substantive question, not recruitment alone.

Misconception

If a variable has “institutional” in its name, it is institutional-level

No. An individual's personal perception of institutional support can remain an individual-level variable. Construct definition and referent determine the level.

Misconception

If my dataset has one row per person, my research question must be individual-level

No. Multilevel analyses can retain individual rows while including higher-level contexts. However, the hierarchy must be identifiable and modeled appropriately.

Misconception

If my question names universities, thousands of faculty respondents give me thousands of university observations

No. University-level inference depends on the number of universities represented, not the number of people nested within them.

Misconception

Using a multilevel model automatically makes the question and variables multilevel

No. Statistical structure cannot repair constructs whose theoretical definitions and measurements operate at the wrong level.

Misconception

If the analysis produces a significant result, the unit must have been appropriate

No. Statistical significance does not validate the research question, construct level, sampling unit, or inferential target.

06 · What This Means for You

Audit the Question, Variables, Sample, Analysis, and Conclusion as One System

A simple decision framework

If the predictor and outcome both belong to individual people
Frame the question and conclusion at the individual level, while accounting for clustering where necessary.
If the predictor and outcome both belong to groups or organizations
Ensure that those groups are adequately sampled and that the variables validly characterize them.
If the predictor and outcome belong to different levels
Write the question explicitly as cross-level and use a design capable of connecting those levels.
If the data only support a lower-level question than originally planned
Narrow the research question rather than making a higher-level claim the evidence cannot support.
If the higher-level question is essential
Redesign the sampling, measurement, and analysis so the intended unit is genuinely represented.

The best unit is not the most impressive-sounding one. It is the one that matches the relationship you genuinely want to understand and the evidence you can legitimately obtain.

07 · A Quick Checklist

Does Your Research Question Actually Match Your Unit of Analysis?

Before finalizing the study, check:
What relationship is my research question asking about?
To whom or to what does that relationship apply?
Can the proposed unit meaningfully possess the predictor and outcome?
If the variables occupy different levels, have I stated the cross-level structure explicitly?
Do my measurements validly describe the entities named in the research question?
Does my unit of observation provide defensible information about the intended unit of analysis?
Have I sampled enough units at the level where the focal relationship varies?
Does the analytical method respect nesting, repeated observations, clustering, or other dependence?
Does my results statement refer to the same kind of entity represented by the coefficient or comparison?
Does the final conclusion avoid silently moving from individuals to groups or groups to individuals?
08 · Frequently Asked Questions

Frequently Asked Questions About Research Questions and Units of Analysis

How do I identify the unit of analysis from a research question?

Identify the focal relationship and ask, “To whom or to what does this relationship apply?” If the answer is students, schools, organizations, countries, documents, or another entity, that entity is usually the relevant unit of analysis for the question.

Does the unit of analysis have to be named explicitly in the research question?

Not grammatically in every case, but the entity should be unambiguous. Explicit nouns are especially helpful in multilevel research because they prevent confusion about whether a construct refers to individuals, groups, or organizations.

Can my predictor and outcome be at different levels?

Yes. That creates a cross-level question, such as university policy predicting individual faculty behavior. The data structure and analysis should explicitly connect the two levels.

If individuals answer my questionnaire, can organizations still be my unit of analysis?

Yes, if individuals serve as valid informants about the organization and their responses can defensibly be combined or modeled to characterize the organizational construct. The number and sampling of organizations also matter.

What if my question is about universities but I only have individual variables?

You may need to narrow the research question to an individual-level claim or redesign the study to collect valid university-level information. Simply calling individual variables “institutional” does not make them organization-level measures.

Does one row in the dataset always equal one unit of analysis?

No. Repeated observations can create many rows per individual, and group-level characteristics can be repeated across many individual rows. What a row represents should be interpreted within the full data structure and analytical model.

How do I know whether I need multilevel analysis?

Multilevel analysis becomes particularly relevant when lower-level observations are nested within higher-level units and the research question concerns variation or relationships at multiple levels. Other methods can also address clustering depending on the inferential goal.

What is the biggest warning sign of a unit-of-analysis mismatch?

One strong warning sign is noun drift: the analysis compares one kind of entity, such as universities, but the conclusion suddenly describes another, such as individual faculty members, without an explicit cross-level analysis supporting that move.

09 · The Bottom Line

Your Research Question, Data, Analysis, and Conclusion Should Be About the Same Thing

The Bottom Line

Your research question and unit of analysis match when the entities to which the proposed relationship applies are the same entities your measurements, sampling, and analytical structure can legitimately compare, or when any cross-level connections are made explicit.

Trace the study as one chain from question to unit, variables, observations, sample, analysis, and conclusion. If the entity changes unexpectedly anywhere along that chain, determine whether you have a deliberate multilevel design or a mismatch. Good unit-of-analysis decisions are ultimately about making sure the study answers the question it claims to answer.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes