03 · What You Need to Know
Multiple Units Are Legitimate When the Research Problem Is Genuinely Multilevel
Start with the simplest case: one unit of analysis
Many studies focus on one type of entity.
For example:
Students: Is academic self-efficacy associated with persistence?
Schools: Are school resources associated with graduation rates?
Countries: Is national research investment associated with scientific output?
Articles: Is international collaboration associated with citation impact?
In these studies, one main unit can adequately describe the substantive question.
Understanding what the unit of analysis is remains the first step before deciding whether the study actually needs more than one.
Multiple units arise when the questions refer to different kinds of entities
Consider a study involving faculty and universities.
Question 1:
Do faculty members with greater AI self-efficacy report greater AI adoption?
Unit of analysis: faculty member.
Question 2:
Do universities with formal AI governance frameworks have higher institutional adoption rates?
Unit of analysis: university.
The study now contains at least two meaningful analytical units because it asks distinct substantive questions about individuals and organizations.
Multiple units by design
Different research questions deliberately concern different entities.
Unit confusion
The analysis unintentionally mixes entities without specifying which claims belong to which level.
Having several data sources does not automatically mean having several units of analysis
Suppose a university case study uses faculty interviews, administrator interviews, policy documents, meeting minutes, and website materials.
That is multiple data sources.
But if all sources are integrated to understand one university as a case, the university may remain the single primary unit of analysis.
Participants, documents, and observations provide evidence about that analytical unit.
This is why participants are not automatically units of analysis.
Likewise, many observations do not automatically create many kinds of units
If 500 students each provide one survey response and the study examines individual-level relationships, there may still be only one unit of analysis: student.
If those students are nested within 20 classrooms but classroom differences are irrelevant to the substantive questions, the classrooms are nevertheless part of the data structure, though they may not be substantive analytical targets.
The distinction is important: hierarchy in the data does not automatically mean every level is a substantive unit of analysis.
Nested data frequently create the possibility of multiple analytical levels
Common hierarchies include repeated observations within people, students within classrooms, classrooms within schools, employees within departments, departments within organizations, and patients within hospitals.
These structures are examples of nested or hierarchical data.
The researcher may ask questions at one level, multiple levels, or across levels.
Multiple levels of data are not identical to multiple units of substantive inference
Suppose students are nested within schools.
The outcome is student achievement, and the question asks whether student study habits predict achievement.
School clustering may need to be accounted for statistically because students within the same school are not completely independent.
But the substantive unit of analysis may still be the student.
By contrast, if the study also asks whether school-level leadership explains differences among schools, schools become an explicit analytical level in the scientific argument.
Different levels can contain different variables
| Individual-level variables |
Organization-level variables |
| Employee self-efficacy |
Organizational size |
| Job satisfaction |
Formal AI policy |
| Age |
Leadership structure |
| AI adoption behavior |
Technology infrastructure |
A multilevel study can examine relationships among individual variables, relationships among organization variables, and relationships connecting the two levels.
This is why distinguishing individual-level and group-level variables becomes important once multiple analytical units are involved.
Within-individual and between-individual questions can also represent multiple levels
Suppose participants report stress and sleep every day.
Researchers could ask whether generally more stressed people sleep less than generally less stressed people. That is a between-person question.
They could also ask whether, on days when a person is more stressed than usual, that person sleeps less than usual. That is a within-person question.
The observations are measurement occasions nested within people. The same variables appear in both questions, but the source of variation differs.
A cross-level relationship connects different analytical levels
Suppose:
University AI policy → Faculty AI adoption
The predictor belongs to the university level. The outcome belongs to the faculty level.
The relationship therefore crosses levels.
Recognizing cross-level relationships prevents researchers from treating variables defined at different levels as ordinary interchangeable columns.
Cross-level moderation adds another layer
Suppose the relationship between individual AI self-efficacy and AI adoption depends on university support.
The proposed relationship is:
Faculty self-efficacy → Faculty adoption
but its strength varies according to university-level support.
This is a cross-level interaction.
Aggregating individuals can create a higher-level analytical unit
Suppose 40 employees within each organization rate organizational climate.
If the theoretical construct is organizational-level and aggregation is justified, researchers may calculate an organization-level score.
This is the logic behind aggregating individual-level data to the group level.
Aggregation does not eliminate the individual level from reality
If individual scores are averaged, variation among individuals within each group disappears from the aggregated dataset.
That may be appropriate for a purely group-level question.
But if researchers also care about individual-level relationships, collapsing everything to group means throws away information needed to answer those questions.
Multilevel models are often designed precisely for this structure
Multilevel, hierarchical, or mixed-effects models allow researchers to represent observations clustered within higher-level units.
A simplified two-level model might distinguish:
Level 1: faculty members
Level 2: universities
Individual outcomes can then be modeled using both faculty-level and university-level characteristics while accounting for dependence among faculty from the same university.
Multiple units require enough information at each relevant level
Suppose 2,000 faculty members are sampled from only eight universities.
The study has extensive individual-level information but very limited between-university information.
Estimating complex university-level effects or cross-level interactions may therefore be difficult even with a large overall participant count.
Watch Out
Lower-level sample size cannot fully compensate for a very small number of higher-level units. If the study asks university-level questions, the number and diversity of universities matter independently of how many faculty members are observed within each university.
The unit relevant to each parameter matters for precision
An individual-level coefficient draws heavily on variation among individuals. A university-level coefficient depends on variation among universities. A cross-level interaction requires information about both.
One headline sample size therefore does not describe the information available for every part of a multilevel study.
Separate analyses at each level can sometimes be appropriate
Not every multiple-unit study requires one enormous multilevel model.
A project might deliberately conduct one faculty-level analysis, one university-level analysis, and one qualitative institution-level comparison.
If these address distinct research questions, separate analyses may be clearer than forcing all questions into one model.
But separate analyses cannot answer every cross-level question
Suppose researchers separately find that faculty self-efficacy predicts individual adoption and that university support predicts institutional adoption rate.
Those two findings do not establish that university support changes the individual self-efficacy–adoption relationship.
That cross-level moderation question requires an analysis capable of linking the levels appropriately.
The same construct can exist at more than one level
Consider support.
At the individual level, perceived personal support may describe one employee's experience.
At the group level, support climate may represent a shared property of an organization.
These constructs are related but should not automatically be treated as identical simply because both use the word “support.”
Group means can separate within-group and between-group information
Suppose individual workload predicts burnout and employees are nested within organizations.
An employee's workload score contains at least two kinds of information: how overloaded that employee is relative to colleagues in the same organization, and whether the organization itself tends to have high average workload.
These can have different relationships with burnout.
A multilevel approach can separate within-group and between-group components rather than assuming one coefficient captures both.
Relationships can differ depending on the unit being compared
Suppose individual employees with greater autonomy report greater satisfaction.
That does not guarantee that organizations with greater average autonomy have greater average satisfaction by exactly the same amount.
The individual-level and organization-level relationships can differ because they represent different sources of variation.
This is why a relationship at one level can differ from the relationship at another.
Multiple units help avoid ecological and atomistic reasoning
If researchers observe a relationship among organizations and assume the same relationship holds among individuals, they risk an ecological fallacy.
If they observe an individual-level relationship and automatically infer an organization-level relationship, they risk an atomistic fallacy.
A multilevel framework can preserve both forms of variation and make clear which conclusion belongs to which analytical unit.
Multiple units should be visible in the research questions
A study claiming to examine multiple levels should not hide that structure until the methods section.
Compare:
RQ1: How is faculty AI self-efficacy related to individual classroom AI adoption?
RQ2: How is university AI infrastructure related to institutional adoption rates?
RQ3: Does university AI infrastructure moderate the relationship between faculty self-efficacy and individual adoption?
The unit and level of each question are visible immediately.
This clarity helps prevent what happens when a research question mixes individual-level and group-level explanations without specifying the intended structure.
Sampling should reflect every substantive unit
If universities are an important analytical unit, the study needs a defensible sample of universities.
If departments are also substantively important, the sampling strategy should provide meaningful departmental variation.
Recruiting hundreds of participants from one or two higher-level units cannot provide broad evidence about variation among higher-level entities.
Do not call every layer a unit of analysis merely because it exists
A dataset may contain items within scales within respondents within teams within organizations.
Not every layer is necessarily a substantive unit of analysis.
Questionnaire items may be measurement indicators rather than substantive analytical entities. Teams may be clustering units without being objects of substantive inference.
Use the term unit of analysis for entities relevant to the actual analytical or inferential question rather than every structural component in the data file.