03 · What You Need to Know
The Key Problem Is Moving From Individual Observations to Group-Level Meaning
Start by separating the unit of observation from the unit of analysis
Suppose 20 teachers in each school complete a questionnaire about school climate.
The teachers provide the information:
Unit of observation = teacher
If the study then compares schools based on aggregated climate scores:
Unit of analysis = school
This is exactly why the unit of observation can differ from the unit of analysis.
The arrangement is not inherently problematic. The methodological question is whether the individual observations legitimately support the group-level construct.
Individuals often serve as informants about groups
Many collective properties are difficult to measure without asking group members.
Examples include:
- organizational climate;
- team cohesion;
- school climate;
- classroom climate;
- shared leadership;
- collective efficacy;
- safety climate;
- innovation climate.
Researchers therefore ask individuals about experiences or perceptions that are theorized to reflect the collective environment.
The individuals are not necessarily the final target of inference. They function as informants about the group.
But not every individual question measures a group property
Compare:
“I personally feel confident using generative AI.”
This is an individual-level construct.
Now compare:
“Faculty members in this university receive adequate support for using generative AI.”
This item asks the respondent to characterize the university environment.
The second wording is more naturally aligned with a group-level construct because the referent is the collective.
The distinction matters because researchers should not simply average any personal attribute and relabel it as an organizational climate.
The theoretical construct should be group-level before the statistics are calculated
A group-level measure should begin with a group-level definition.
For example:
University AI-support climate is the shared perception among faculty that university policies, leadership, infrastructure, and practices support responsible AI adoption.
This definition explains why individual perceptions may be combined: the target construct is the shared environment.
Without such a definition, aggregation can become a purely numerical operation searching for a conceptual interpretation afterward.
Composition models explain how individual information becomes a higher-level construct
Multilevel theory distinguishes several ways lower-level data can compose into higher-level constructs.
In some cases, the group construct is essentially the average or consensus of members' individual perceptions.
In others, the pattern or dispersion of individual responses matters.
For example:
- average expertise may characterize a team;
- diversity depends on heterogeneity rather than consensus;
- team climate may require shared perceptions;
- minimum competence may matter more than the mean for some safety processes.
This is why aggregation should reflect theory rather than defaulting automatically to the arithmetic mean.
A simple mean is only one possible aggregation rule
The most familiar aggregation is:
The mean is appropriate when theory says the average level of the individual responses represents the higher-level property.
But other constructs may require proportions, dispersion indices, maxima, minima, or other composition rules.
Averaging does not automatically create consensus
Suppose ten employees rate organizational support as:
1, 1, 1, 1, 1, 5, 5, 5, 5, 5.
The mean is 3.
But does this organization have a shared “moderate support climate”?
Probably not. Half the group perceives almost no support and half perceives very strong support.
The mean hides disagreement.
This is why shared group-level constructs often require evidence about within-group agreement.
Within-group agreement asks whether members perceive the group similarly
If the construct is supposed to represent a shared climate, researchers need some basis for claiming that group members converge in their perceptions.
Agreement measures can help evaluate whether responses within each group are sufficiently similar relative to a specified reference.
The exact choice of index and threshold should reflect the measurement model and field-specific conventions.
Watch Out
A single cutoff should not replace substantive judgment. Strong aggregation arguments combine theoretical justification, appropriate item referents, within-group agreement, between-group differentiation, reliability, and transparent reporting.
Reliability of the group mean is a different question from agreement
Researchers sometimes treat agreement and reliability as interchangeable.
They are not.
Agreement asks whether members give similar ratings.
Reliability asks how dependably groups can be distinguished from one another given the observed responses and sampling of members.
A group mean can sometimes be estimated reliably even when individual agreement is imperfect, and high agreement does not by itself guarantee strong differentiation among groups.
The two concepts contribute different evidence.
Between-group variation matters because the study wants to compare groups
Suppose faculty perceptions of support are almost identical across every university.
Even if each university shows strong internal consensus, there is little between-university variation to explain.
A group-level predictor is most useful when groups differ meaningfully on the construct.
Multilevel variance estimates and intraclass correlations can help describe this differentiation.
ICC(1) is often used to describe group-related clustering
Conceptually, ICC(1) reflects the proportion of individual-response variance associated with group membership under a random-intercept model.
ICC(1) should not be treated as a universal pass-fail test for aggregation, but it can contribute evidence that groups matter empirically.
ICC(2) is often used to describe reliability of group means
ICC(2) is commonly interpreted as the reliability of aggregated group means under a particular sampling and random-effects framework.
Its magnitude depends partly on group size: group means based on more respondents can generally be estimated more reliably than means based on very few respondents, all else equal.
This matters because a university climate score based on two faculty members carries different measurement uncertainty from one based on fifty.
Group size therefore matters twice
The number of respondents per group affects:
- how well the group construct is measured;
- how stable the group mean or other aggregate is.
But the number of groups affects something different:
- how much information is available for estimating relationships among groups.
Researchers need to distinguish these two sample-size questions.
Many respondents in a few groups do not create a strong group-level study
Suppose researchers survey 5,000 employees from five organizations.
They may estimate each organization's average climate with considerable precision.
But there are still only five organizations available for examining whether organizational climate predicts organizational performance.
This is a central limitation.
Respondents per group
Help estimate the group's characteristics.
Number of groups
Provides information for comparing higher-level entities and estimating higher-level relationships.
The research question should use group-level nouns if the claim is group-level
Compare:
Do faculty who perceive greater institutional support adopt AI more frequently?
This is an individual-level question.
Now compare:
Do universities with stronger AI-support climates have higher institutional adoption rates?
This is a university-level question.
The wording reveals what kind of variation the study intends to explain.
This distinction should be clear before aggregation begins.
An individual-level relationship and a group-level relationship are not equivalent
Suppose faculty members who perceive stronger support are more likely to adopt AI.
This does not automatically mean universities with stronger average support have higher average adoption.
The first relationship compares individuals.
The second compares universities.
They may both be true, but one does not logically establish the other.
This is why relationships may differ across levels of analysis.
Aggregating both predictor and outcome moves the entire question upward
Suppose researchers average support and adoption within each university.
The dataset now has one support score and one adoption score per university.
The relationship between those variables is a between-university relationship.
Individual variation within universities has been discarded from that analysis.
The resulting coefficient cannot directly answer whether individual faculty members who perceive more support than their colleagues adopt more AI.
Keeping individual data allows multilevel analysis instead of full aggregation
If researchers care about both individuals and groups, they do not necessarily need to collapse everything to university means.
A multilevel model can retain faculty-level observations while including university-level variables.
This allows questions such as:
Do faculty members with greater self-efficacy adopt more AI?
Do universities differ in average adoption?
Does university support predict faculty adoption?
Does university support modify the individual self-efficacy–adoption relationship?
This is one reason a study can legitimately have more than one unit of analysis.
Multilevel modeling preserves the nested structure
Suppose faculty are nested in universities.
A simplified multilevel model recognizes that two faculty members from the same university share contextual influences that faculty from different universities do not.
The model can estimate variation within universities and variation between universities separately.
This avoids treating every individual as completely independent while also avoiding the information loss created by reducing each university to one row.
Contextual effects can differ from individual effects
Suppose individual faculty members with high AI competence are more likely to adopt AI.
Researchers may also ask whether working in a university where faculty generally have high AI competence predicts adoption beyond one's personal competence.
The second is a contextual or group-composition question.
It asks whether the environment created by being surrounded by more competent colleagues matters independently of the individual's own competence.
Such questions require separating individual and group-level components rather than using one undifferentiated score.
Group-mean centering can help distinguish individual and contextual components
Suppose individual competence is X.
Researchers can calculate:
Individual deviation from university mean:
and include:
University mean: X̄ⱼ
as a separate group-level predictor.
The model can then distinguish whether personal competence and average university competence have different associations with adoption.
Group-level inference requires enough groups
Suppose 2,000 teachers are sampled from six schools.
The teachers may provide excellent information about individual-level relationships.
But six schools provide very limited information about school-level predictors or school-to-school relationships.
The relevant N for a group-level hypothesis is not simply the number of individual respondents.
Sampling should therefore start at the group level when the question is about groups
If the primary research question concerns universities, the study needs a sampling strategy for universities.
Researchers should ask:
- How many universities are needed?
- How will universities be selected?
- Is there enough variation across universities?
- How many respondents per university are needed to measure the group construct well?
Recruiting many individuals from a convenient handful of institutions is not equivalent to sampling many institutions.
Unequal group sizes create additional measurement considerations
Suppose one university contributes 80 respondents and another contributes 4.
The two group means do not have the same measurement precision.
Researchers may need to consider whether unequal group sizes affect reliability, weighting, uncertainty, and interpretation.
Simply assigning equal confidence to every aggregate can obscure large differences in how well each group was measured.
Missing groups can threaten group-level generalization
If organizations with weak climates are less willing to participate, the resulting group sample may overrepresent high-performing institutions.
This is a group-level selection problem.
Individual response rates within participating groups and selection of groups into the study are distinct issues and should both be considered.
Individual nonresponse can also distort the group measure
Suppose only highly enthusiastic faculty complete the AI-support survey within each university.
The resulting university mean may not represent the faculty population accurately even if many universities participate.
Group-level measurement therefore depends on who responds within each group as well as which groups enter the sample.
Measurement invariance can matter across groups
If a scale is used across organizations, classrooms, countries, or other groups, researchers should consider whether respondents interpret and respond to the items comparably.
If the meaning of a scale differs systematically across groups, differences in group means may partly reflect measurement differences rather than genuine construct differences.
This becomes particularly important in cross-cultural and multi-institutional research.
Qualitative research can use individuals as informants about groups too
Suppose researchers interview faculty and administrators to understand university-level AI governance.
The interviewees are individuals, but the analytic case may be the university.
The researchers may triangulate multiple accounts, documents, and observations to construct a case-level interpretation.
The same principle applies: individual testimony is evidence about the group-level phenomenon, not automatically the final unit of analysis.
Multiple informants can strengthen group-level measurement
Using several respondents from one group can provide different perspectives and reduce dependence on one person's idiosyncratic view.
However, more informants do not automatically guarantee validity.
Researchers should still consider:
- who is knowledgeable about the target construct;
- whether respondents share the same referent;
- whether roles produce systematically different perspectives;
- whether disagreement itself is theoretically meaningful.
Disagreement is not always measurement error
If faculty and administrators strongly disagree about institutional support, researchers should not automatically force the responses into one mean and describe disagreement as noise.
The disagreement may reveal:
- unequal access to resources;
- role-based differences;
- fragmented implementation;
- subcultures within the organization;
- lack of a genuinely shared climate.
Sometimes variation within groups is itself the phenomenon of interest.
Not every group-level question requires aggregation
Some higher-level variables can be measured directly.
University size, budget, policy presence, accreditation category, infrastructure investment, or adoption rate may be available from administrative data.
In those cases, individual aggregation may be unnecessary for those variables.
Researchers can still combine directly measured group variables with individually measured outcomes in a multilevel design.
The ecological fallacy remains a risk after aggregation
Suppose universities with stronger support climates have higher adoption rates.
The finding concerns universities.
It does not necessarily mean that within each university, faculty members who perceive greater support are more likely to adopt AI.
Inferring the individual relationship from the aggregate relationship risks ecological fallacy.
The atomistic fallacy is the reverse risk
Suppose individual faculty who feel more supported are more likely to adopt AI.
That does not automatically establish that universities with stronger support climates have higher institutional adoption.
Group-level dynamics may differ from individual-level dynamics.
Your conclusion should reveal the level of the claim
Compare:
Individual-level conclusion: Faculty members who perceived more support reported greater AI adoption.
Group-level conclusion: Universities with stronger AI-support climates had higher institutional adoption rates.
Cross-level conclusion: Faculty members working in universities with stronger AI-support climates reported greater personal adoption.
These statements are not stylistic variants. They refer to different analytical structures.
The transition from individual observations to group claims should be documented explicitly
A strong methods section should explain:
- why the construct is conceptualized at the group level;
- who provided the individual observations;
- how the group referent was specified;
- how responses were combined;
- what evidence supported aggregation;
- how many respondents contributed per group;
- how many groups were analyzed;
- how group-level uncertainty or nesting was handled.
This makes the inferential bridge visible instead of leaving readers to assume that averaging responses was sufficient.