03 · What You Need to Know
A Group-Level Measure Needs More Than a Group Mean
First ask whether the construct is actually group-level
The most important question is theoretical:
What exactly is the variable supposed to describe?
Consider:
Faculty AI self-efficacy
This naturally describes an individual faculty member's belief in their own capability.
Now consider:
University AI-support climate
This is intended to describe a shared institutional environment.
If the construct is inherently individual, averaging individual scores may produce an average but not necessarily a new collective construct.
A group mean and a group-level construct are not synonymous
Suppose the average faculty self-efficacy score in University A is 4.6.
That value can legitimately describe:
the average observed self-efficacy of faculty in University A.
But it does not automatically establish a collective construct called:
University A's collective AI efficacy.
Collective efficacy has a distinct theoretical meaning concerning the group's shared belief in collective capability.
The label should therefore follow the theory, not the spreadsheet calculation.
The composition model should be specified explicitly
Chan's composition framework distinguishes several ways constructs can relate across levels, emphasizing that higher-level constructs are not produced by one universal averaging rule.
The researcher should therefore state something like:
“University AI-support climate is conceptualized as a shared faculty perception of institutional practices and resources. Individual faculty ratings are treated as lower-level observations of that common institutional property.”
That statement provides the theoretical bridge.
Shared constructs require evidence of sharedness
If a group-level construct is defined by consensus, researchers need evidence that members perceive the target sufficiently similarly.
Suppose faculty within one university provide ratings:
4, 4, 5, 4, 4, 5, 4, 5
The responses appear relatively consistent.
Now suppose another university produces:
1, 1, 2, 2, 6, 6, 7, 7
The mean may fall near the middle of the scale, but describing the university as having one shared moderate climate would hide substantial disagreement.
Within-group agreement and interrater reliability are not the same thing
This distinction is fundamental.
Agreement asks:
Do members give similar ratings?
Reliability asks:
Can groups be differentiated consistently based on the ratings?
LeBreton and Senter emphasize the conceptual distinction between interrater agreement and interrater reliability and warn against treating the associated indices as interchangeable.
Agreement
How close are the ratings provided by members of the same group?
Reliability
How reliably can differences among groups or targets be distinguished using those ratings?
rwg is commonly used as an agreement index
The rwg
family of statistics was developed to evaluate within-group agreement relative to a specified null distribution of responses. James, Demaree, and Wolf's work provided an influential basis for these indices.
Conceptually, higher rwg values indicate that observed responses are more concentrated than expected under the chosen null distribution.
But interpretation depends on assumptions about that null distribution and the response scale.
Researchers should therefore be cautious about treating one conventional threshold as automatic proof that aggregation is justified.
rwg should not be the sole justification for aggregation
Methodological work has specifically warned that rwg and related agreement indices can be misused when treated as the only empirical criterion for moving data to a higher level.
A strong aggregation argument considers:
- theory;
- construct referent;
- agreement;
- nonindependence;
- between-group variation;
- reliability of group means;
- group size;
- sampling quality.
ICC(1) addresses a different question
In a simple random-intercept setting, ICC(1) is commonly used to describe the degree to which individual responses are associated with group membership.
An ICC(1) of.20 in this simplified illustration indicates that approximately 20% of the modeled variation is associated with differences between groups.
It does not mean that respondents agree perfectly or that aggregation is automatically valid.
Nonindependence itself matters
Individuals within the same group often resemble one another more than randomly selected individuals from different groups because they share environments, leadership, policies, experiences, or selection processes.
Bliese distinguishes agreement, reliability, and nonindependence as separate properties relevant to aggregation and multilevel analysis.
Recognizing nonindependence is also important even when researchers ultimately retain the individual-level outcome rather than aggregate it.
ICC(2) is often used to assess reliability of group means
ICC(2) is commonly interpreted as the reliability of the group mean under an appropriate model.
Conceptually, it asks whether the aggregate group scores can be estimated reliably enough to distinguish groups.
ICC(2) generally increases as the number of respondents per group increases, assuming the underlying ICC(1) and other conditions remain similar.
This means that a university mean based on 30 respondents may be substantially more reliable than one based on three.
High agreement and high ICC(2) answer different questions
Imagine every member within each university gives very similar ratings, but all universities have nearly the same mean.
Agreement may be strong.
Between-university differentiation may still be weak.
Alternatively, group means may differ reliably even though members within groups do not show perfect consensus.
This is why researchers often examine several pieces of evidence instead of searching for one decisive statistic.
There is no universal magic cutoff
Rules such as:
rwg must exceed.70
or:
ICC(2) must exceed.70
are frequently repeated.
Such conventions can provide reference points, but they should not be treated as universal laws.
The acceptable evidence depends on:
- construct definition;
- measurement reliability;
- group sizes;
- number of groups;
- purpose of aggregation;
- expected degree of consensus;
- research context.
LeBreton and Senter's review is useful precisely because it emphasizes how different agreement and reliability indices answer different methodological questions rather than supporting one mechanical criterion.
Item referent matters before agreement is calculated
Suppose the goal is to measure school climate.
Compare:
“I receive useful feedback from my principal.”
with:
“Teachers in this school receive useful feedback from school leadership.”
The first asks about a personal experience.
The second asks respondents to characterize a shared school environment.
Even strong agreement on the first item does not automatically make the underlying construct equivalent to shared school climate.
Respondents also need to be able to observe the target
Suppose junior faculty are asked to rate:
“The university governing board has a coherent AI investment strategy.”
Many respondents may lack sufficient knowledge to make that judgment.
Agreement among uninformed raters does not necessarily create valid organizational measurement.
Researchers should ask whether respondents are appropriately positioned to observe the phenomenon.
Multiple informants are most useful when they are informative
More respondents can improve reliability, but quantity does not compensate for poor informational access.
For a department-level teaching climate, ordinary faculty may be appropriate informants.
For confidential institutional investment decisions, senior administrators may be more appropriate sources.
Sometimes combining multiple roles provides a more complete view, but systematic role differences may also need to be modeled rather than averaged away.
Respondent sampling within groups matters
Suppose only faculty already enthusiastic about AI respond to the survey.
Their ratings may systematically overstate institutional AI support relative to the broader faculty population.
A large number of biased respondents can still produce a biased group mean.
Group-level measurement therefore requires attention to representativeness within each group.
Small group samples can create unstable aggregates
Suppose one university contributes two faculty respondents while another contributes forty.
Both universities can technically receive a mean score.
But the uncertainty surrounding those means differs substantially.
Researchers should establish minimum sampling strategies appropriate to the construct and acknowledge groups that are poorly represented.
More respondents per group can improve ICC(2), but not the number of groups
This distinction is essential.
If you have five organizations and survey 200 employees in each one, the organization means may be estimated very reliably.
But you still have only five organizations for estimating relationships among organizations.
Watch Out
Excellent measurement of a few groups does not create a large higher-level sample. Aggregation quality and higher-level statistical power are separate design problems.
Groups also need to differ meaningfully
If every university has nearly identical climate scores, there is little higher-level variation available to explain.
A group-level construct can be theoretically valid yet empirically unsuitable as a predictor in a particular sample because the observed groups are too homogeneous.
The group boundary must match the construct
Suppose support practices are determined mainly by departments rather than universities.
Faculty within the same department may share similar experiences, while departments within one university differ sharply.
Aggregating everyone to the university level could conceal the actual level at which the climate exists.
This is a theoretical question about where the social process operates.
Agreement can reveal whether the proposed grouping level is too broad
If agreement is consistently poor at the university level but substantially stronger within departments, that pattern may suggest that the relevant shared environment is departmental rather than university-wide.
Statistics do not make that theoretical decision alone, but they can reveal a mismatch worth investigating.
Disagreement can itself be substantively important
Suppose faculty are polarized about AI policy.
Some perceive excellent support while others perceive severe institutional barriers.
The variance within the university may indicate:
- unequal implementation;
- disciplinary subcultures;
- role-based inequalities;
- fragmented leadership;
- policy-practice gaps.
Collapsing the disagreement into one mean may erase the phenomenon the study should explain.
Not every aggregate needs shared agreement
This bears repeating because aggregation decisions differ by construct.
If the variable is:
percentage of faculty using AI
there is no requirement that faculty “agree” about whether they personally use AI.
The heterogeneity is exactly what generates the proportion.
If the variable is:
disciplinary diversity
lack of similarity is central to the construct.
Agreement criteria make sense primarily when sharedness is part of the theoretical definition.
Internal consistency does not justify aggregation
A high Cronbach's alpha or another internal-consistency coefficient may indicate that items on a scale hang together within respondents.
It does not establish that members of the same group agree with one another.
Internal consistency
Do items within a scale show coherence?
Within-group agreement
Do different people within the same group give similar ratings?
Group-mean reliability
Can aggregate group differences be estimated reliably?
These statistics answer different questions.
Factorial validity does not establish the level either
A scale can have an excellent individual-level factor structure while remaining inappropriate as a group-level measure.
Moving a construct upward may require considering whether the measurement model behaves appropriately at the between-group level as well.
In multilevel factor analysis, within-group and between-group structures can even differ.
Measurement invariance across groups can matter
If a scale is used across universities, cultures, occupations, or countries, researchers may need to consider whether the items have comparable meaning across those groups.
Apparent differences in group means are difficult to interpret if groups use the measurement scale differently.
Aggregation cannot repair noncomparability in the underlying measurement.
The group-level construct should predict something at the appropriate level
One source of construct validity can come from how the aggregate behaves in relation to theoretically relevant variables.
For example, if university AI-support climate is genuinely institutional, it might plausibly relate to:
- university-wide AI adoption;
- institutional investment;
- faculty development participation;
- cross-level faculty adoption outcomes.
Such relationships do not by themselves prove the composition model, but they can contribute to a broader validity argument.
A group-level measure can predict an individual-level outcome
Suppose validated university support climate predicts individual faculty AI adoption.
This is a cross-level relationship:
University climate → Faculty behavior
The study should use an analytical approach that recognizes faculty are nested within universities.
This is precisely the kind of question covered by a cross-level relationship.
A group-level variable should not be duplicated across individuals and treated as independent
Suppose every faculty member in University A receives the same university climate score of 4.3.
The researcher cannot pretend those repeated values represent independent observations of 4.3.
They all originate from the same higher-level entity.
Ignoring this clustering can distort standard errors and inference. Grouped organizational data are typically nonindependent, and methodological work has shown that treating such observations as independent can adversely affect statistical inference.
Aggregation is easier to defend when planned before data collection
Planning allows researchers to:
- define the higher-level construct explicitly;
- write items with the appropriate referent;
- sample enough groups;
- sample enough knowledgeable respondents per group;
- predefine the aggregation rule;
- identify appropriate agreement and reliability evidence.
Deciding after data collection that an individual survey “would be more interesting” as an organizational study can create problems that no statistical transformation can fully solve.
Changing the intended level after data collection requires caution
Suppose a questionnaire was designed entirely around personal experiences:
“I have adequate resources.”
“My supervisor supports me.”
“I am comfortable experimenting with AI.”
After data collection, the researcher decides to average the scores and call the result “organizational AI readiness.”
That interpretation may be unjustified because the original items were not designed to measure a shared organizational construct.
The possibility of changing the unit of analysis after data collection therefore needs to be evaluated substantively, not merely computationally.
Aggregation evidence should be reported transparently
A methods section should not simply say:
“Employee responses were aggregated to the organizational level.”
A stronger report explains:
- why the construct is group-level;
- how items referred to the group;
- which respondents contributed;
- how many respondents contributed per group;
- which aggregation operator was used;
- which agreement and reliability evidence was examined;
- how many higher-level groups entered the analysis.
Use multiple pieces of evidence rather than one ritual statistic
The strongest justification rarely takes the form:
“rwg was above.70, therefore aggregation was justified.”
Methodological literature distinguishes agreement, reliability, and nonindependence because they answer different questions about grouped data.
A more defensible conclusion considers theory and several empirical indicators together.
The relevant evidence depends on the composition model
For a shared climate, agreement may be central.
For average age, agreement is irrelevant.
For diversity, dispersion is the variable.
For the percentage adopting AI, the proportion itself is the group property.
There is therefore no universal aggregation checklist that can be applied without first defining what kind of higher-level construct is being created.
The final question is whether the group-level interpretation is scientifically coherent
After all calculations are complete, ask:
What exactly does a one-unit increase in this group score mean?
If the answer is:
“Faculty members in this university collectively perceive a more supportive institutional environment,”
the construct is clearly group-oriented.
If the answer remains:
“The numerical average happened to be higher,”
the higher-level interpretation may need more conceptual work.
This is the essential difference between aggregating individual data and validating a genuinely group-level measure.