03 · What You Need to Know
Aggregation Changes Both the Data Structure and the Meaning of the Comparison
Start with individual-level data
Imagine a study involving 1,500 faculty members across 50 universities.
Each faculty member provides:
- AI teaching self-efficacy;
- perceived institutional support;
- frequency of AI use;
- academic rank.
At this stage, each faculty member has their own values. If the analysis compares faculty members, those variables operate at the individual level.
For example:
Do faculty members who perceive more support use AI more frequently?
That is an individual-level relationship.
Aggregation groups individuals by a higher-level entity
Suppose the researcher now wants to compare universities.
Faculty members are grouped according to university, and their support ratings are combined.
University A now has one summary value representing the average of those observed faculty responses.
If every university receives one such score, the researcher can compare universities rather than individual faculty members.
The unit represented by one row may change
Before full aggregation:
1 row = 1 faculty member
After full aggregation:
1 row = 1 university
This reflects the distinction between units of observation and units of analysis.
Faculty members provided the observations. Universities may become the units being compared.
Aggregation is therefore a substantive transformation
The numerical calculation may take seconds, but the inferential change is substantial.
At the individual level, the study asks:
How do people differ from one another?
At the aggregate level, it asks:
How do groups differ from one another?
Those are not the same statistical comparison.
Individual-level analysis
Compares variation among individual people.
Group-level analysis
Compares variation among collective entities such as classrooms, teams, schools, or organizations.
The arithmetic mean is only one form of aggregation
Researchers sometimes speak of aggregation as though it always means calculating an average.
But the correct group summary depends on what the group-level construct represents.
| Group construct |
Possible aggregation |
| Average faculty experience |
Mean |
| Proportion of faculty adopting AI |
Proportion or percentage |
| Total research output |
Sum |
| Team diversity |
Dispersion or diversity index |
| Minimum competence needed for a task |
Minimum |
| Highest available expertise |
Maximum |
| Variation in member attitudes |
Variance or standard deviation |
The aggregation operator should follow the theory of how individual characteristics combine into a higher-level property.
Composition models explain how lower-level information becomes a higher-level construct
Multilevel theory uses the idea of composition to describe how constructs at one level relate to constructs at another. Chan's influential framework emphasized that constructs referring to similar content can relate across levels in different ways, so researchers need to specify the functional relationship rather than simply assume that a group construct is an average individual construct.
For some constructs, the group mean may be theoretically appropriate.
For others, agreement among members matters.
For still others, disagreement, dispersion, minimum values, or patterns of configuration may be the defining feature.
Direct-consensus constructs are a common aggregation case
Suppose researchers define university AI-support climate as a shared faculty perception that the institution provides resources and encouragement for responsible AI use.
Faculty members answer items referring to their university.
If the theory treats consensus among members as defining the group property, researchers may aggregate those ratings when sufficient evidence supports the shared interpretation.
The logic is:
Individual perceptions → consensus within university → university-level climate score
Referent-shift constructs work somewhat differently
Instead of asking:
“I am confident using generative AI.”
researchers might ask:
“Faculty in my department are capable of using generative AI effectively.”
The respondent is still an individual, but the item referent has shifted from I to the group.
Aggregation can then be used to construct a group-level collective efficacy measure when theory and empirical evidence justify doing so.
The wording of the items therefore matters before any averages are computed.
Not every group variable originates through aggregation
Some variables exist directly at the higher level.
Examples include:
- number of students in a classroom;
- university budget;
- presence of a formal policy;
- organization size;
- school type;
- country-level legislation.
These variables do not require individual responses to construct them.
Aggregation is specifically relevant when lower-level observations are being transformed into higher-level information.
A group mean can represent composition without representing consensus
Suppose researchers calculate the mean age of employees in an organization.
No agreement is required. Employees do not need to be similar in age for the average age to be meaningful.
The construct is simply the group's average composition.
By contrast, if researchers claim that employees share a common safety climate, agreement becomes much more relevant because sharedness is part of the construct.
This distinction is essential.
Agreement is therefore construct-dependent
Whether within-group agreement is required depends on what the aggregate is supposed to mean.
| Aggregate |
Is consensus central? |
| Average employee age |
No |
| Percentage of teachers using AI |
No |
| Team diversity |
No; disagreement or heterogeneity defines the construct |
| Shared safety climate |
Usually yes |
| Shared organizational support climate |
Usually yes |
Researchers should therefore avoid applying agreement statistics mechanically to every aggregate variable.
Aggregation removes information about individual differences
Suppose five departments have 100 faculty members each.
After averaging all relevant variables within department, the analysis contains only five department-level cases.
Information about differences among the 500 faculty members within departments is no longer represented in the fully aggregated dataset.
This is not necessarily wrong. It may be exactly what a department-level question requires.
But the cost should be understood.
Very different groups can have the same mean
Consider two teams.
Team A ratings:
4, 4, 4, 4, 4
Team B ratings:
1, 2, 4, 6, 7
Both have a mean of 4.
Yet Team A shows complete agreement while Team B contains substantial disagreement.
If the construct is supposed to represent shared climate, treating the two groups as equivalent may conceal important information.
This is why dispersion may matter
Some theories treat variability within a group as substantively meaningful.
A team with sharply divided views about leadership may operate differently from a team whose members uniformly hold moderate views, even if both groups have identical means.
Researchers should therefore ask whether the group mean alone captures the higher-level phenomenon.
Aggregation changes the effective sample for group-level relationships
Suppose 2,400 employees are drawn from 30 organizations.
After full aggregation to the organization level, the group-level analysis contains:
N = 30 organizations
not:
N = 2,400 organizations.
Watch Out
Large numbers of individual respondents can improve estimation of each group's characteristics, but they do not create additional independent groups. Group-level relationships depend on variation across the groups themselves.
More people per group and more groups solve different problems
Suppose you want to estimate university climate and then relate climate to university performance.
More faculty respondents per university can improve the measurement of each university's climate.
More universities provide more information about the relationship between climate and university performance.
These are distinct design considerations.
More respondents per group
Can improve estimation and reliability of the group characteristic.
More groups
Increase information available for estimating between-group relationships.
Aggregation can improve signal when individual responses contain idiosyncratic noise
If individuals provide imperfect observations of a shared group environment, combining several informants can reduce the influence of one person's unusual response.
A university climate score based on 30 appropriately sampled faculty members may therefore provide a more stable estimate of the shared environment than the rating of a single faculty member.
This is one reason multiple informants can be valuable.
But aggregation can also hide meaningful subgroups
Suppose faculty from engineering report strong institutional support while faculty from the humanities report very weak support.
A university-wide average may obscure that systematic difference.
If the subgroups experience genuinely different environments, the university may not possess one homogeneous support climate.
Researchers should therefore inspect the possibility that the assumed group is too broad.
The group boundary itself requires justification
Why aggregate by university rather than department?
Why by school rather than classroom?
Why by country rather than region?
The correct grouping variable should correspond to where the proposed shared process exists.
If AI policy and support are implemented mainly at department level, university-wide aggregation may conceal the relevant context.
Aggregation can produce contextual variables
Suppose individual faculty digital competence is averaged within universities.
The university mean represents the typical competence level of faculty in that university.
Researchers can then investigate whether an individual's adoption behavior is related not only to their own competence but also to the average competence of colleagues.
This creates a contextual variable derived from individual-level information.
The individual score and the group mean should not be treated as interchangeable
Suppose Faculty Member A has competence = 5.
That score describes the individual.
Suppose their university mean is 3.5.
That score describes the aggregate context.
The two values answer different questions even though they were derived from the same original variable.
Multilevel modeling can retain both levels instead of fully aggregating
If researchers care about both individuals and groups, complete aggregation is not the only option.
A multilevel model can retain individual observations while introducing higher-level variables.
This allows researchers to investigate:
- individual-level relationships;
- group-level differences;
- contextual effects;
- cross-level relationships;
- cross-level moderation.
This is particularly useful when the study has nested or hierarchical data.
Ignoring grouping is not the opposite of aggregation
Researchers sometimes imagine only two choices:
Option A: average everything by group
Option B: analyze every individual as though grouping did not exist
Multilevel methods provide a third option: preserve the individual observations while explicitly modeling the group structure.
Aggregation can change a coefficient dramatically
Suppose individual support and individual adoption are positively related.
After aggregation, university-average support and university-average adoption might show:
- a stronger positive relationship;
- a weaker relationship;
- no relationship;
- even a relationship in the opposite direction.
This is possible because the aggregate analysis compares different variation.
The fact that relationships can differ across levels is one of the most important reasons to distinguish individual and aggregate analyses.
Aggregation creates a risk of ecological inference
Suppose universities with higher average AI competence show higher institutional adoption.
That university-level result does not prove that within universities, individual faculty members with higher competence are necessarily more likely to adopt AI.
Inferring the individual relationship from the aggregate relationship risks ecological fallacy.
Individual findings cannot simply be pushed upward either
If individual faculty competence predicts individual adoption, it does not automatically follow that universities with higher average competence will show greater institutional adoption.
That opposite movement risks atomistic inference.
Aggregation therefore changes both what is estimated and what conclusions are defensible.
Weighting may matter when group sizes differ
Suppose one school has 10 respondents and another has 200.
In a group-level analysis, should both group means receive equal statistical weight?
The answer depends on the estimand, sampling design, measurement precision, and analytical method.
Larger groups may yield more precise estimates of the group mean, but automatically weighting groups by size can also cause larger groups to dominate a question that conceptually concerns differences among groups.
Weighting should therefore follow the design rather than habit.
Missing responses can affect group aggregates
A university average is only as representative as the observed faculty contributing to it.
If respondents are systematically more enthusiastic about AI than nonrespondents, the group mean may overstate the university's actual support or adoption climate.
Aggregation does not remove individual-level nonresponse bias.
Unequal response patterns can also make groups differently reliable
A group score based on 40 representative respondents is not measured with the same certainty as a group score based on three respondents.
Researchers should report group sample sizes and consider how varying precision affects the analysis.
The appropriate aggregation should be decided before looking for favorable results
Researchers should avoid trying the mean, median, maximum, proportion, and several subgroup definitions and retaining whichever produces the strongest association.
The composition rule should be driven primarily by theory and design.
Exploratory alternatives can still be informative, but they should be identified transparently as exploratory.
Aggregation is easiest to defend when the inferential chain is explicit
A strong study can explain:
1. What lower-level variable was measured?
2. What higher-level construct is intended?
3. Why should lower-level responses compose into that construct?
4. What aggregation rule represents the composition process?
5. What empirical evidence supports the aggregation?
6. What higher-level relationship will then be analyzed?
This reasoning is more informative than simply stating that “responses were averaged by institution.”
Aggregation should match the research question
If the question asks about individual faculty behavior, aggregating everything to university means may destroy necessary information.
If the question asks about universities, analyzing every faculty response independently may fail to represent the intended unit.
The correct strategy follows from the alignment between question and unit.
This is why aggregation is closely connected to the broader issue of collecting individual data for a group-level research question.