Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

When Is It Justified to Treat Individual Responses as a Group-Level Measure?

Individual responses can support a group-level measure when the construct is genuinely collective and the measurement strategy shows that members' responses can meaningfully represent the group. A high average or a convenient ICC alone is not enough.

123
When Individual Responses Can Represent a Group Guide 123 of 223
01 · The Question

When Does Averaging People's Answers Actually Tell You Something About Their Group?

Suppose 30 faculty members at each university rate institutional support for generative AI.

You average their responses and call the result university AI-support climate.

Is that justified?

Perhaps. But the fact that the software can calculate a university mean does not establish that the mean validly represents a university-level construct.

Faculty members may be describing private experiences rather than a shared environment. They may disagree substantially. Some universities may have too few respondents. The questionnaire may use an individual referent even though the interpretation is organizational. Groups may barely differ from one another.

To make a credible group-level claim, researchers need a composition argument connecting individual observations to the collective construct and empirical evidence appropriate to that argument. Composition theory specifically addresses how constructs at different levels are functionally related rather than assuming that lower-level and higher-level constructs are automatically equivalent.

02 · The Short Answer

Aggregation Is Justified When Theory and Evidence Support the Group Interpretation

In Brief

It is justified to treat individual responses as a group-level measure when the construct is theoretically defined at the group level, the respondents can reasonably report on that shared target, the measurement referent matches the intended level, and the observed response pattern supports the proposed composition process.

For shared constructs such as climate or collective perceptions, researchers commonly examine within-group agreement, between-group differentiation, and reliability of group scores. These statistics provide different kinds of evidence and should not be reduced to one universal cutoff. The substantive definition of the group construct comes first.

03 · What You Need to Know

A Group-Level Measure Needs More Than a Group Mean

First ask whether the construct is actually group-level

The most important question is theoretical:

What exactly is the variable supposed to describe?

Consider:

Faculty AI self-efficacy

This naturally describes an individual faculty member's belief in their own capability.

Now consider:

University AI-support climate

This is intended to describe a shared institutional environment.

If the construct is inherently individual, averaging individual scores may produce an average but not necessarily a new collective construct.

A group mean and a group-level construct are not synonymous

Suppose the average faculty self-efficacy score in University A is 4.6.

That value can legitimately describe:

the average observed self-efficacy of faculty in University A.

But it does not automatically establish a collective construct called:

University A's collective AI efficacy.

Collective efficacy has a distinct theoretical meaning concerning the group's shared belief in collective capability.

The label should therefore follow the theory, not the spreadsheet calculation.

The composition model should be specified explicitly

Chan's composition framework distinguishes several ways constructs can relate across levels, emphasizing that higher-level constructs are not produced by one universal averaging rule.

The researcher should therefore state something like:

“University AI-support climate is conceptualized as a shared faculty perception of institutional practices and resources. Individual faculty ratings are treated as lower-level observations of that common institutional property.”

That statement provides the theoretical bridge.

Shared constructs require evidence of sharedness

If a group-level construct is defined by consensus, researchers need evidence that members perceive the target sufficiently similarly.

Suppose faculty within one university provide ratings:

4, 4, 5, 4, 4, 5, 4, 5

The responses appear relatively consistent.

Now suppose another university produces:

1, 1, 2, 2, 6, 6, 7, 7

The mean may fall near the middle of the scale, but describing the university as having one shared moderate climate would hide substantial disagreement.

Within-group agreement and interrater reliability are not the same thing

This distinction is fundamental.

Agreement asks:

Do members give similar ratings?

Reliability asks:

Can groups be differentiated consistently based on the ratings?

LeBreton and Senter emphasize the conceptual distinction between interrater agreement and interrater reliability and warn against treating the associated indices as interchangeable.

Agreement How close are the ratings provided by members of the same group?
Reliability How reliably can differences among groups or targets be distinguished using those ratings?

rwg is commonly used as an agreement index

The rwg family of statistics was developed to evaluate within-group agreement relative to a specified null distribution of responses. James, Demaree, and Wolf's work provided an influential basis for these indices.

Conceptually, higher rwg values indicate that observed responses are more concentrated than expected under the chosen null distribution.

But interpretation depends on assumptions about that null distribution and the response scale.

Researchers should therefore be cautious about treating one conventional threshold as automatic proof that aggregation is justified.

rwg should not be the sole justification for aggregation

Methodological work has specifically warned that rwg and related agreement indices can be misused when treated as the only empirical criterion for moving data to a higher level.

A strong aggregation argument considers:

  • theory;
  • construct referent;
  • agreement;
  • nonindependence;
  • between-group variation;
  • reliability of group means;
  • group size;
  • sampling quality.

ICC(1) addresses a different question

In a simple random-intercept setting, ICC(1) is commonly used to describe the degree to which individual responses are associated with group membership.

Conceptual ICC(1)
ICC(1) = σ²between / (σ²between + σ²within)
σ²between represents modeled between-group variance and σ²within represents modeled within-group variance.
If between-university variance is 0.20 and within-university variance is 0.80, ICC(1) = 0.20 / 1.00 = 0.20.

An ICC(1) of.20 in this simplified illustration indicates that approximately 20% of the modeled variation is associated with differences between groups.

It does not mean that respondents agree perfectly or that aggregation is automatically valid.

Nonindependence itself matters

Individuals within the same group often resemble one another more than randomly selected individuals from different groups because they share environments, leadership, policies, experiences, or selection processes.

Bliese distinguishes agreement, reliability, and nonindependence as separate properties relevant to aggregation and multilevel analysis.

Recognizing nonindependence is also important even when researchers ultimately retain the individual-level outcome rather than aggregate it.

ICC(2) is often used to assess reliability of group means

ICC(2) is commonly interpreted as the reliability of the group mean under an appropriate model.

Conceptually, it asks whether the aggregate group scores can be estimated reliably enough to distinguish groups.

ICC(2) generally increases as the number of respondents per group increases, assuming the underlying ICC(1) and other conditions remain similar.

This means that a university mean based on 30 respondents may be substantially more reliable than one based on three.

High agreement and high ICC(2) answer different questions

Imagine every member within each university gives very similar ratings, but all universities have nearly the same mean.

Agreement may be strong.

Between-university differentiation may still be weak.

Alternatively, group means may differ reliably even though members within groups do not show perfect consensus.

This is why researchers often examine several pieces of evidence instead of searching for one decisive statistic.

There is no universal magic cutoff

Rules such as:

rwg must exceed.70

or:

ICC(2) must exceed.70

are frequently repeated.

Such conventions can provide reference points, but they should not be treated as universal laws.

The acceptable evidence depends on:

  • construct definition;
  • measurement reliability;
  • group sizes;
  • number of groups;
  • purpose of aggregation;
  • expected degree of consensus;
  • research context.

LeBreton and Senter's review is useful precisely because it emphasizes how different agreement and reliability indices answer different methodological questions rather than supporting one mechanical criterion.

Item referent matters before agreement is calculated

Suppose the goal is to measure school climate.

Compare:

“I receive useful feedback from my principal.”

with:

“Teachers in this school receive useful feedback from school leadership.”

The first asks about a personal experience.

The second asks respondents to characterize a shared school environment.

Even strong agreement on the first item does not automatically make the underlying construct equivalent to shared school climate.

Respondents also need to be able to observe the target

Suppose junior faculty are asked to rate:

“The university governing board has a coherent AI investment strategy.”

Many respondents may lack sufficient knowledge to make that judgment.

Agreement among uninformed raters does not necessarily create valid organizational measurement.

Researchers should ask whether respondents are appropriately positioned to observe the phenomenon.

Multiple informants are most useful when they are informative

More respondents can improve reliability, but quantity does not compensate for poor informational access.

For a department-level teaching climate, ordinary faculty may be appropriate informants.

For confidential institutional investment decisions, senior administrators may be more appropriate sources.

Sometimes combining multiple roles provides a more complete view, but systematic role differences may also need to be modeled rather than averaged away.

Respondent sampling within groups matters

Suppose only faculty already enthusiastic about AI respond to the survey.

Their ratings may systematically overstate institutional AI support relative to the broader faculty population.

A large number of biased respondents can still produce a biased group mean.

Group-level measurement therefore requires attention to representativeness within each group.

Small group samples can create unstable aggregates

Suppose one university contributes two faculty respondents while another contributes forty.

Both universities can technically receive a mean score.

But the uncertainty surrounding those means differs substantially.

Researchers should establish minimum sampling strategies appropriate to the construct and acknowledge groups that are poorly represented.

More respondents per group can improve ICC(2), but not the number of groups

This distinction is essential.

If you have five organizations and survey 200 employees in each one, the organization means may be estimated very reliably.

But you still have only five organizations for estimating relationships among organizations.

Watch Out

Excellent measurement of a few groups does not create a large higher-level sample. Aggregation quality and higher-level statistical power are separate design problems.

Groups also need to differ meaningfully

If every university has nearly identical climate scores, there is little higher-level variation available to explain.

A group-level construct can be theoretically valid yet empirically unsuitable as a predictor in a particular sample because the observed groups are too homogeneous.

The group boundary must match the construct

Suppose support practices are determined mainly by departments rather than universities.

Faculty within the same department may share similar experiences, while departments within one university differ sharply.

Aggregating everyone to the university level could conceal the actual level at which the climate exists.

This is a theoretical question about where the social process operates.

Agreement can reveal whether the proposed grouping level is too broad

If agreement is consistently poor at the university level but substantially stronger within departments, that pattern may suggest that the relevant shared environment is departmental rather than university-wide.

Statistics do not make that theoretical decision alone, but they can reveal a mismatch worth investigating.

Disagreement can itself be substantively important

Suppose faculty are polarized about AI policy.

Some perceive excellent support while others perceive severe institutional barriers.

The variance within the university may indicate:

  • unequal implementation;
  • disciplinary subcultures;
  • role-based inequalities;
  • fragmented leadership;
  • policy-practice gaps.

Collapsing the disagreement into one mean may erase the phenomenon the study should explain.

Not every aggregate needs shared agreement

This bears repeating because aggregation decisions differ by construct.

If the variable is:

percentage of faculty using AI

there is no requirement that faculty “agree” about whether they personally use AI.

The heterogeneity is exactly what generates the proportion.

If the variable is:

disciplinary diversity

lack of similarity is central to the construct.

Agreement criteria make sense primarily when sharedness is part of the theoretical definition.

Internal consistency does not justify aggregation

A high Cronbach's alpha or another internal-consistency coefficient may indicate that items on a scale hang together within respondents.

It does not establish that members of the same group agree with one another.

Internal consistency Do items within a scale show coherence?
Within-group agreement Do different people within the same group give similar ratings?
Group-mean reliability Can aggregate group differences be estimated reliably?

These statistics answer different questions.

Factorial validity does not establish the level either

A scale can have an excellent individual-level factor structure while remaining inappropriate as a group-level measure.

Moving a construct upward may require considering whether the measurement model behaves appropriately at the between-group level as well.

In multilevel factor analysis, within-group and between-group structures can even differ.

Measurement invariance across groups can matter

If a scale is used across universities, cultures, occupations, or countries, researchers may need to consider whether the items have comparable meaning across those groups.

Apparent differences in group means are difficult to interpret if groups use the measurement scale differently.

Aggregation cannot repair noncomparability in the underlying measurement.

The group-level construct should predict something at the appropriate level

One source of construct validity can come from how the aggregate behaves in relation to theoretically relevant variables.

For example, if university AI-support climate is genuinely institutional, it might plausibly relate to:

  • university-wide AI adoption;
  • institutional investment;
  • faculty development participation;
  • cross-level faculty adoption outcomes.

Such relationships do not by themselves prove the composition model, but they can contribute to a broader validity argument.

A group-level measure can predict an individual-level outcome

Suppose validated university support climate predicts individual faculty AI adoption.

This is a cross-level relationship:

University climate → Faculty behavior

The study should use an analytical approach that recognizes faculty are nested within universities.

This is precisely the kind of question covered by a cross-level relationship.

A group-level variable should not be duplicated across individuals and treated as independent

Suppose every faculty member in University A receives the same university climate score of 4.3.

The researcher cannot pretend those repeated values represent independent observations of 4.3.

They all originate from the same higher-level entity.

Ignoring this clustering can distort standard errors and inference. Grouped organizational data are typically nonindependent, and methodological work has shown that treating such observations as independent can adversely affect statistical inference.

Aggregation is easier to defend when planned before data collection

Planning allows researchers to:

  • define the higher-level construct explicitly;
  • write items with the appropriate referent;
  • sample enough groups;
  • sample enough knowledgeable respondents per group;
  • predefine the aggregation rule;
  • identify appropriate agreement and reliability evidence.

Deciding after data collection that an individual survey “would be more interesting” as an organizational study can create problems that no statistical transformation can fully solve.

Changing the intended level after data collection requires caution

Suppose a questionnaire was designed entirely around personal experiences:

“I have adequate resources.”

“My supervisor supports me.”

“I am comfortable experimenting with AI.”

After data collection, the researcher decides to average the scores and call the result “organizational AI readiness.”

That interpretation may be unjustified because the original items were not designed to measure a shared organizational construct.

The possibility of changing the unit of analysis after data collection therefore needs to be evaluated substantively, not merely computationally.

Aggregation evidence should be reported transparently

A methods section should not simply say:

“Employee responses were aggregated to the organizational level.”

A stronger report explains:

  • why the construct is group-level;
  • how items referred to the group;
  • which respondents contributed;
  • how many respondents contributed per group;
  • which aggregation operator was used;
  • which agreement and reliability evidence was examined;
  • how many higher-level groups entered the analysis.

Use multiple pieces of evidence rather than one ritual statistic

The strongest justification rarely takes the form:

“rwg was above.70, therefore aggregation was justified.”

Methodological literature distinguishes agreement, reliability, and nonindependence because they answer different questions about grouped data.

A more defensible conclusion considers theory and several empirical indicators together.

The relevant evidence depends on the composition model

For a shared climate, agreement may be central.

For average age, agreement is irrelevant.

For diversity, dispersion is the variable.

For the percentage adopting AI, the proportion itself is the group property.

There is therefore no universal aggregation checklist that can be applied without first defining what kind of higher-level construct is being created.

The final question is whether the group-level interpretation is scientifically coherent

After all calculations are complete, ask:

What exactly does a one-unit increase in this group score mean?

If the answer is:

“Faculty members in this university collectively perceive a more supportive institutional environment,”

the construct is clearly group-oriented.

If the answer remains:

“The numerical average happened to be higher,”

the higher-level interpretation may need more conceptual work.

This is the essential difference between aggregating individual data and validating a genuinely group-level measure.

04 · A Practical Example

When Faculty Ratings Can Represent University AI-Support Climate

Hypothetical Example

Measuring a shared institutional environment

Researchers want to compare 70 universities on AI-support climate. They survey approximately 25 faculty members per institution using items about institutional policies, leadership practices, infrastructure, and professional development.

Theoretical definition The construct is explicitly defined as a shared faculty perception of the university environment rather than a private attitude.
Appropriate referent Items use wording such as “faculty in this university” and “this university provides,” directing respondents toward the shared institutional target.
Informant suitability Faculty are selected because they are exposed to the policies and support practices the construct is intended to capture.
Empirical evidence Researchers examine within-university agreement, between-university variation, and reliability of aggregate university scores rather than relying on one statistic alone.
Aggregation When the theoretical and empirical evidence is adequate, faculty responses are combined into one support-climate score per university.

The resulting university measure is defensible because the aggregation was built into the theory and measurement strategy rather than imposed after the survey was completed.

05 · What Researchers Often Get Wrong

Common Mistakes When Justifying Group-Level Measures

Misconception

A high group mean proves the group shares the perception

No. A high mean tells you the average response. It does not reveal whether members agree or whether some members hold dramatically different views.

Misconception

rwg above one conventional cutoff automatically validates aggregation

No. Agreement indices depend on assumptions and address only one part of the aggregation argument. Theory, reliability, between-group variation, group size, and measurement design also matter.

Misconception

A high Cronbach's alpha means members agree

No. Internal consistency concerns relationships among scale items, not agreement among different respondents within the same group.

Misconception

Every aggregate requires consensus

No. Consensus is relevant to shared constructs such as climate. It is not required for aggregates such as average age, proportions, totals, or diversity measures in the same way.

Misconception

More respondents per organization solve a small number of organizations

No. More respondents can improve measurement of each organization, but higher-level relationships still depend on the number of organizations observed.

Misconception

If a scale works at the individual level, it automatically works at the group level

No. Moving constructs across levels requires a composition argument. The group-level meaning, referent, agreement, reliability, and measurement structure may differ from the individual-level version.

06 · What This Means for You

Justify the Group Construct Before You Justify the Statistic

A simple decision framework

If the construct describes an inherently individual attribute
Do not relabel its average as a shared group construct without a separate theoretical composition argument.
If the construct is explicitly defined as a shared group perception
Use appropriate group-referent measurement and examine agreement, reliability, and between-group differentiation.
If the group property is compositional rather than consensual
Use the theoretically appropriate mean, proportion, dispersion, diversity index, or other summary without requiring irrelevant agreement criteria.
If substantial disagreement exists
Investigate whether the group is heterogeneous, the grouping level is wrong, or disagreement itself is the construct of interest before averaging it away.
If evidence for the group-level interpretation is weak
Retain the variable at the individual level or describe the aggregate more modestly rather than claiming a shared group construct.

The goal is not to pass a statistical ritual. It is to show that the group-level number means what the study says it means.

07 · A Quick Checklist

Before Treating Individual Responses as a Group-Level Measure, Check This

Before claiming a group-level construct, check:
Is the construct theoretically defined as belonging to the group?
Have I specified how lower-level responses compose into the higher-level construct?
Do the items use a referent appropriate to the intended group-level interpretation?
Are respondents knowledgeable about the group property they are rating?
Is the respondent sample within each group sufficiently representative?
If sharedness is part of the construct, is within-group agreement adequate?
Is there meaningful between-group variation or nonindependence?
Are the aggregate group scores estimated with appropriate reliability?
Are there enough higher-level groups for the planned group-level relationships?
Am I using several relevant pieces of evidence rather than one automatic cutoff?
08 · Frequently Asked Questions

Frequently Asked Questions About Group-Level Measurement

Can I simply average survey responses to create a group-level measure?

You can calculate the mean, but interpreting it as a group-level construct requires a theoretical composition argument. Shared constructs may also require evidence of agreement, between-group differentiation, and reliability.

Does every group-level measure require high within-group agreement?

No. Agreement is relevant when sharedness or consensus defines the construct. Group averages, proportions, diversity measures, and other compositional variables can be meaningful without member consensus.

What is rwg used for?

rwg and related indices are commonly used to assess within-group agreement relative to a specified null response distribution. They contribute evidence about sharedness but should not normally serve as the sole justification for aggregation.

What does ICC(1) tell me?

In common random-intercept applications, ICC(1) describes the proportion of modeled individual-response variation associated with group differences and therefore indicates the degree of clustering or nonindependence by group.

What does ICC(2) tell me?

ICC(2) is commonly used to describe the reliability of group mean scores. It is influenced by the underlying group-related variance and by how many respondents contribute to each group mean.

What cutoff should I use for rwg or ICC?

No single cutoff is universally appropriate. Conventional benchmarks can be useful descriptive references, but aggregation should be justified using construct theory, measurement design, agreement, reliability, group sizes, between-group variation, and the intended use of the aggregate.

Does high Cronbach's alpha justify aggregation?

No. Cronbach's alpha addresses internal consistency among scale items. It does not establish agreement among respondents within a group or reliability of aggregate group scores.

What should I do if group members strongly disagree?

Examine whether the construct is genuinely shared, whether the grouping level is appropriate, whether meaningful subgroups exist, and whether disagreement itself should be modeled. Averaging substantial heterogeneity may hide rather than represent the group phenomenon.

09 · The Bottom Line

A Group Mean Becomes a Group-Level Measure Only When the Interpretation Is Defensible

The Bottom Line

Individual responses can be treated as a group-level measure when theory specifies how those responses represent a collective property and the measurement evidence supports the intended composition process.

For shared constructs, examine the referent, respondent suitability, within-group agreement, between-group differentiation, and reliability of aggregate scores rather than relying on one convenient threshold. For other group constructs, use the composition rule the theory actually requires. Aggregation is justified by the meaning of the group construct first and by appropriate statistics second.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes