Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Happens When Your Data Are Collected From Individuals but Your Question Is About Groups?

It is common to collect data from individuals while asking questions about classrooms, teams, schools, or organizations. This design can be valid, but researchers must explain how individual observations represent the group-level construct and why the resulting group-level inference is justified.

121
Individual Data for Group-Level Questions Guide 121 of 223
01 · The Question

What If People Give You the Data, but the Group Is What You Actually Want to Study?

Suppose you want to compare universities on their institutional support for generative AI.

A university cannot complete a questionnaire. So you survey faculty members and ask them about policy clarity, available resources, technical support, leadership encouragement, and professional development.

The respondents are individuals.

But your research question is:

Do universities with stronger AI-support climates have higher institutional adoption rates?

You are therefore collecting information from individuals to make a claim about groups.

This is common in organizational research, education, public health, sociology, and multilevel research. It can be entirely defensible, but it creates an important methodological bridge that must be justified.

You need to explain why individual responses can be combined to represent the group, how the aggregation will be performed, whether members show enough agreement, whether groups differ meaningfully, and whether the number of groups is sufficient for the intended group-level inference.

02 · The Short Answer

Individuals Can Provide Group-Level Evidence, but the Link Must Be Justified

In Brief

When data are collected from individuals but the research question concerns groups, individuals function as observational sources while the groups become the substantive units of analysis, provided the individual responses can validly represent the intended group-level construct.

This usually requires more than averaging responses. Researchers should define the group-level construct theoretically, use an appropriate group referent when needed, evaluate whether within-group responses show meaningful agreement, assess between-group variation and reliability, and ensure that the analysis and conclusions remain at the group level or explicitly model multiple levels when individual and group questions coexist.

03 · What You Need to Know

The Key Problem Is Moving From Individual Observations to Group-Level Meaning

Start by separating the unit of observation from the unit of analysis

Suppose 20 teachers in each school complete a questionnaire about school climate.

The teachers provide the information:

Unit of observation = teacher

If the study then compares schools based on aggregated climate scores:

Unit of analysis = school

This is exactly why the unit of observation can differ from the unit of analysis.

The arrangement is not inherently problematic. The methodological question is whether the individual observations legitimately support the group-level construct.

Individuals often serve as informants about groups

Many collective properties are difficult to measure without asking group members.

Examples include:

  • organizational climate;
  • team cohesion;
  • school climate;
  • classroom climate;
  • shared leadership;
  • collective efficacy;
  • safety climate;
  • innovation climate.

Researchers therefore ask individuals about experiences or perceptions that are theorized to reflect the collective environment.

The individuals are not necessarily the final target of inference. They function as informants about the group.

But not every individual question measures a group property

Compare:

“I personally feel confident using generative AI.”

This is an individual-level construct.

Now compare:

“Faculty members in this university receive adequate support for using generative AI.”

This item asks the respondent to characterize the university environment.

The second wording is more naturally aligned with a group-level construct because the referent is the collective.

The distinction matters because researchers should not simply average any personal attribute and relabel it as an organizational climate.

The theoretical construct should be group-level before the statistics are calculated

A group-level measure should begin with a group-level definition.

For example:

University AI-support climate is the shared perception among faculty that university policies, leadership, infrastructure, and practices support responsible AI adoption.

This definition explains why individual perceptions may be combined: the target construct is the shared environment.

Without such a definition, aggregation can become a purely numerical operation searching for a conceptual interpretation afterward.

Composition models explain how individual information becomes a higher-level construct

Multilevel theory distinguishes several ways lower-level data can compose into higher-level constructs.

In some cases, the group construct is essentially the average or consensus of members' individual perceptions.

In others, the pattern or dispersion of individual responses matters.

For example:

  • average expertise may characterize a team;
  • diversity depends on heterogeneity rather than consensus;
  • team climate may require shared perceptions;
  • minimum competence may matter more than the mean for some safety processes.

This is why aggregation should reflect theory rather than defaulting automatically to the arithmetic mean.

A simple mean is only one possible aggregation rule

The most familiar aggregation is:

Group Mean
X̄ⱼ = ΣXᵢⱼ / nⱼ
Xᵢⱼ is individual i's response within group j, and X̄ⱼ is the resulting group mean.
If 25 faculty members in one university have support scores totaling 100, the university mean is 100 / 25 = 4.0.

The mean is appropriate when theory says the average level of the individual responses represents the higher-level property.

But other constructs may require proportions, dispersion indices, maxima, minima, or other composition rules.

Averaging does not automatically create consensus

Suppose ten employees rate organizational support as:

1, 1, 1, 1, 1, 5, 5, 5, 5, 5.

The mean is 3.

But does this organization have a shared “moderate support climate”?

Probably not. Half the group perceives almost no support and half perceives very strong support.

The mean hides disagreement.

This is why shared group-level constructs often require evidence about within-group agreement.

Within-group agreement asks whether members perceive the group similarly

If the construct is supposed to represent a shared climate, researchers need some basis for claiming that group members converge in their perceptions.

Agreement measures can help evaluate whether responses within each group are sufficiently similar relative to a specified reference.

The exact choice of index and threshold should reflect the measurement model and field-specific conventions.

Watch Out

A single cutoff should not replace substantive judgment. Strong aggregation arguments combine theoretical justification, appropriate item referents, within-group agreement, between-group differentiation, reliability, and transparent reporting.

Reliability of the group mean is a different question from agreement

Researchers sometimes treat agreement and reliability as interchangeable.

They are not.

Agreement asks whether members give similar ratings.

Reliability asks how dependably groups can be distinguished from one another given the observed responses and sampling of members.

A group mean can sometimes be estimated reliably even when individual agreement is imperfect, and high agreement does not by itself guarantee strong differentiation among groups.

The two concepts contribute different evidence.

Between-group variation matters because the study wants to compare groups

Suppose faculty perceptions of support are almost identical across every university.

Even if each university shows strong internal consensus, there is little between-university variation to explain.

A group-level predictor is most useful when groups differ meaningfully on the construct.

Multilevel variance estimates and intraclass correlations can help describe this differentiation.

ICC(1) is often used to describe group-related clustering

Conceptually, ICC(1) reflects the proportion of individual-response variance associated with group membership under a random-intercept model.

Conceptual ICC(1)
ICC(1) = between-group variance / total variance
Higher values indicate that responses from members of the same group tend to resemble one another more strongly relative to responses from different groups.
If 12% of variation in faculty support ratings lies between universities, group membership accounts for a nontrivial portion of the observed variability under the fitted model.

ICC(1) should not be treated as a universal pass-fail test for aggregation, but it can contribute evidence that groups matter empirically.

ICC(2) is often used to describe reliability of group means

ICC(2) is commonly interpreted as the reliability of aggregated group means under a particular sampling and random-effects framework.

Its magnitude depends partly on group size: group means based on more respondents can generally be estimated more reliably than means based on very few respondents, all else equal.

This matters because a university climate score based on two faculty members carries different measurement uncertainty from one based on fifty.

Group size therefore matters twice

The number of respondents per group affects:

  • how well the group construct is measured;
  • how stable the group mean or other aggregate is.

But the number of groups affects something different:

  • how much information is available for estimating relationships among groups.

Researchers need to distinguish these two sample-size questions.

Many respondents in a few groups do not create a strong group-level study

Suppose researchers survey 5,000 employees from five organizations.

They may estimate each organization's average climate with considerable precision.

But there are still only five organizations available for examining whether organizational climate predicts organizational performance.

This is a central limitation.

Respondents per group Help estimate the group's characteristics.
Number of groups Provides information for comparing higher-level entities and estimating higher-level relationships.

The research question should use group-level nouns if the claim is group-level

Compare:

Do faculty who perceive greater institutional support adopt AI more frequently?

This is an individual-level question.

Now compare:

Do universities with stronger AI-support climates have higher institutional adoption rates?

This is a university-level question.

The wording reveals what kind of variation the study intends to explain.

This distinction should be clear before aggregation begins.

An individual-level relationship and a group-level relationship are not equivalent

Suppose faculty members who perceive stronger support are more likely to adopt AI.

This does not automatically mean universities with stronger average support have higher average adoption.

The first relationship compares individuals.

The second compares universities.

They may both be true, but one does not logically establish the other.

This is why relationships may differ across levels of analysis.

Aggregating both predictor and outcome moves the entire question upward

Suppose researchers average support and adoption within each university.

The dataset now has one support score and one adoption score per university.

The relationship between those variables is a between-university relationship.

Individual variation within universities has been discarded from that analysis.

The resulting coefficient cannot directly answer whether individual faculty members who perceive more support than their colleagues adopt more AI.

Keeping individual data allows multilevel analysis instead of full aggregation

If researchers care about both individuals and groups, they do not necessarily need to collapse everything to university means.

A multilevel model can retain faculty-level observations while including university-level variables.

This allows questions such as:

Do faculty members with greater self-efficacy adopt more AI?

Do universities differ in average adoption?

Does university support predict faculty adoption?

Does university support modify the individual self-efficacy–adoption relationship?

This is one reason a study can legitimately have more than one unit of analysis.

Multilevel modeling preserves the nested structure

Suppose faculty are nested in universities.

A simplified multilevel model recognizes that two faculty members from the same university share contextual influences that faculty from different universities do not.

The model can estimate variation within universities and variation between universities separately.

This avoids treating every individual as completely independent while also avoiding the information loss created by reducing each university to one row.

Contextual effects can differ from individual effects

Suppose individual faculty members with high AI competence are more likely to adopt AI.

Researchers may also ask whether working in a university where faculty generally have high AI competence predicts adoption beyond one's personal competence.

The second is a contextual or group-composition question.

It asks whether the environment created by being surrounded by more competent colleagues matters independently of the individual's own competence.

Such questions require separating individual and group-level components rather than using one undifferentiated score.

Group-mean centering can help distinguish individual and contextual components

Suppose individual competence is X.

Researchers can calculate:

Individual deviation from university mean:

Within-Group Component
Xᵢⱼ - X̄ⱼ
This expresses how much an individual's score differs from the average score in that person's group.
If a faculty member scores 5.5 and the university mean is 4.5, the within-university deviation is +1.0.

and include:

University mean: X̄ⱼ

as a separate group-level predictor.

The model can then distinguish whether personal competence and average university competence have different associations with adoption.

Group-level inference requires enough groups

Suppose 2,000 teachers are sampled from six schools.

The teachers may provide excellent information about individual-level relationships.

But six schools provide very limited information about school-level predictors or school-to-school relationships.

The relevant N for a group-level hypothesis is not simply the number of individual respondents.

Sampling should therefore start at the group level when the question is about groups

If the primary research question concerns universities, the study needs a sampling strategy for universities.

Researchers should ask:

  • How many universities are needed?
  • How will universities be selected?
  • Is there enough variation across universities?
  • How many respondents per university are needed to measure the group construct well?

Recruiting many individuals from a convenient handful of institutions is not equivalent to sampling many institutions.

Unequal group sizes create additional measurement considerations

Suppose one university contributes 80 respondents and another contributes 4.

The two group means do not have the same measurement precision.

Researchers may need to consider whether unequal group sizes affect reliability, weighting, uncertainty, and interpretation.

Simply assigning equal confidence to every aggregate can obscure large differences in how well each group was measured.

Missing groups can threaten group-level generalization

If organizations with weak climates are less willing to participate, the resulting group sample may overrepresent high-performing institutions.

This is a group-level selection problem.

Individual response rates within participating groups and selection of groups into the study are distinct issues and should both be considered.

Individual nonresponse can also distort the group measure

Suppose only highly enthusiastic faculty complete the AI-support survey within each university.

The resulting university mean may not represent the faculty population accurately even if many universities participate.

Group-level measurement therefore depends on who responds within each group as well as which groups enter the sample.

Measurement invariance can matter across groups

If a scale is used across organizations, classrooms, countries, or other groups, researchers should consider whether respondents interpret and respond to the items comparably.

If the meaning of a scale differs systematically across groups, differences in group means may partly reflect measurement differences rather than genuine construct differences.

This becomes particularly important in cross-cultural and multi-institutional research.

Qualitative research can use individuals as informants about groups too

Suppose researchers interview faculty and administrators to understand university-level AI governance.

The interviewees are individuals, but the analytic case may be the university.

The researchers may triangulate multiple accounts, documents, and observations to construct a case-level interpretation.

The same principle applies: individual testimony is evidence about the group-level phenomenon, not automatically the final unit of analysis.

Multiple informants can strengthen group-level measurement

Using several respondents from one group can provide different perspectives and reduce dependence on one person's idiosyncratic view.

However, more informants do not automatically guarantee validity.

Researchers should still consider:

  • who is knowledgeable about the target construct;
  • whether respondents share the same referent;
  • whether roles produce systematically different perspectives;
  • whether disagreement itself is theoretically meaningful.

Disagreement is not always measurement error

If faculty and administrators strongly disagree about institutional support, researchers should not automatically force the responses into one mean and describe disagreement as noise.

The disagreement may reveal:

  • unequal access to resources;
  • role-based differences;
  • fragmented implementation;
  • subcultures within the organization;
  • lack of a genuinely shared climate.

Sometimes variation within groups is itself the phenomenon of interest.

Not every group-level question requires aggregation

Some higher-level variables can be measured directly.

University size, budget, policy presence, accreditation category, infrastructure investment, or adoption rate may be available from administrative data.

In those cases, individual aggregation may be unnecessary for those variables.

Researchers can still combine directly measured group variables with individually measured outcomes in a multilevel design.

The ecological fallacy remains a risk after aggregation

Suppose universities with stronger support climates have higher adoption rates.

The finding concerns universities.

It does not necessarily mean that within each university, faculty members who perceive greater support are more likely to adopt AI.

Inferring the individual relationship from the aggregate relationship risks ecological fallacy.

The atomistic fallacy is the reverse risk

Suppose individual faculty who feel more supported are more likely to adopt AI.

That does not automatically establish that universities with stronger support climates have higher institutional adoption.

Group-level dynamics may differ from individual-level dynamics.

Your conclusion should reveal the level of the claim

Compare:

Individual-level conclusion: Faculty members who perceived more support reported greater AI adoption.

Group-level conclusion: Universities with stronger AI-support climates had higher institutional adoption rates.

Cross-level conclusion: Faculty members working in universities with stronger AI-support climates reported greater personal adoption.

These statements are not stylistic variants. They refer to different analytical structures.

The transition from individual observations to group claims should be documented explicitly

A strong methods section should explain:

  • why the construct is conceptualized at the group level;
  • who provided the individual observations;
  • how the group referent was specified;
  • how responses were combined;
  • what evidence supported aggregation;
  • how many respondents contributed per group;
  • how many groups were analyzed;
  • how group-level uncertainty or nesting was handled.

This makes the inferential bridge visible instead of leaving readers to assume that averaging responses was sufficient.

04 · A Practical Example

From Faculty Responses to a University-Level AI-Support Climate

Hypothetical Example

Measuring institutional support through faculty reports

Researchers want to test whether universities with stronger AI-support climates have higher institutional rates of responsible AI adoption. They recruit 25 to 40 faculty members from each of 60 universities.

Define the group-level construct AI-support climate is defined as the shared faculty perception that institutional leadership, policy, infrastructure, and professional development support responsible AI use.
Use a group referent Items ask about what “faculty in this university” experience rather than only what the individual respondent personally experiences.
Evaluate aggregation Researchers examine within-university agreement, between-university variation, and reliability of university mean scores.
Create the higher-level measure When the theoretical and empirical evidence supports aggregation, faculty responses are combined to estimate each university's AI-support climate.
Analyze universities The university-level climate score is related to university-level adoption outcomes across the 60 institutions.

Faculty members supplied the observations, but the main analytical comparison is among universities.

The validity of the university-level conclusion depends on the entire composition process, not merely on the fact that a mean was calculated.

05 · What Researchers Often Get Wrong

Common Mistakes When Individual Data Are Used for Group-Level Questions

Misconception

If I average individual responses, I automatically have a group-level variable

No. Aggregation creates a group summary, but the group-level interpretation requires a theoretical composition argument and suitable measurement evidence.

Misconception

A high Cronbach's alpha means aggregation is justified

No. Internal consistency addresses whether items within a scale relate to one another. It does not establish that members of the same group agree or that the construct legitimately exists at the group level.

Misconception

More respondents always compensate for fewer groups

No. More respondents can improve measurement of each group, but group-level relationships still depend on the number and variation of groups themselves.

Misconception

If group members disagree, I should simply average them anyway

Not automatically. Strong disagreement may indicate that the construct is not shared, that subgroups exist, or that dispersion itself is theoretically important.

Misconception

A group-level association tells me the individual-level association

No. Aggregate and individual relationships can differ. Inferring the individual pattern from group-level evidence risks ecological fallacy.

Misconception

If individuals are my respondents, the study must be individual-level

No. Individuals can serve as observational sources for group-level constructs. The target of inference is determined by the research question and analytical structure, not simply by who completes the questionnaire.

06 · What This Means for You

Build an Explicit Bridge From Individual Reports to Group-Level Inference

A simple decision framework

If your question is genuinely about groups
Define the construct and outcome at the group level before deciding how individual data will contribute.
If individuals are being used as informants about a shared group property
Use item wording, sampling, and theory that make the group referent explicit.
If responses will be aggregated
Evaluate the composition rule, agreement, between-group variation, and reliability rather than relying on the arithmetic mean alone.
If individual-level and group-level questions both matter
Retain the hierarchical structure and consider a multilevel analysis rather than collapsing all individual information.
If you have many respondents but few groups
Be cautious about group-level inference because the higher-level sample size remains limited.

A useful planning sentence is:

“We collect responses from [individuals] because they provide information about [group-level construct], which will be represented at the [group] level using [composition method] and analyzed across [number/type of groups].”

If that sentence cannot be completed clearly, the design probably needs further work.

07 · A Quick Checklist

Before Using Individual Responses for a Group-Level Question, Check This

Before moving from individuals to groups, check:
Is the research question explicitly about a group-level entity?
Is the construct theoretically defined at the group level?
Do the survey items use an appropriate individual or group referent?
Is the chosen composition rule appropriate for the construct?
Is there sufficient within-group agreement when a shared construct is claimed?
Is there enough between-group variation to support meaningful group comparison?
Are group-level scores estimated reliably given the number of respondents per group?
Are there enough groups for the intended higher-level analysis?
Does the conclusion remain at the group level unless a separate individual-level analysis supports an individual claim?
08 · Frequently Asked Questions

Frequently Asked Questions About Individual Data and Group-Level Research

Can I study organizations using surveys completed by employees?

Yes. Employees can serve as informants about organizational constructs, but the study should justify why their responses represent the organization and how those responses are combined or modeled.

Do I always need to aggregate individual responses?

No. If the study contains both individual- and group-level questions, multilevel modeling may allow you to retain individual responses while including higher-level variables. Full aggregation is only one analytical strategy.

Is taking the group mean enough to create a group-level measure?

No. A mean is a mathematical summary. Group-level interpretation also requires a defensible construct definition, composition model, and evidence appropriate to the type of group construct being claimed.

Why does within-group agreement matter?

If a construct is supposed to represent a shared climate or perception, substantial disagreement among members can undermine the claim that the group has one common standing on that construct. The appropriate interpretation depends on theory and the specific agreement evidence.

What is the difference between ICC(1) and ICC(2)?

In common multilevel applications, ICC(1) describes the extent to which individual responses are clustered by group, while ICC(2) is often used to describe the reliability of group mean scores. Exact interpretation depends on the model and design.

How many respondents do I need per group?

There is no universal number. Requirements depend on the construct, expected agreement, group size, reliability, sampling design, measurement quality, and intended analysis. More respondents per group can improve estimation of group characteristics, but they do not replace the need for an adequate number of groups.

If I have 1,000 respondents from 10 organizations, is my sample size 1,000 or 10?

Both counts describe different parts of the design. There are 1,000 individual respondents nested within 10 organizations. Individual-level analyses use the lower-level information while respecting clustering, whereas organization-level relationships rely on variation across the 10 organizations.

Can group-level and individual-level results disagree?

Yes. They represent different sources of variation and can differ in magnitude or direction. Researchers should therefore avoid inferring one level directly from findings at another.

09 · The Bottom Line

Individual Respondents Can Tell You About Groups, but Averaging Alone Does Not Create Group-Level Evidence

The Bottom Line

When individuals provide data for a group-level research question, they serve as observational sources while the group becomes the target of analysis, and the validity of that transition depends on a defensible theoretical and measurement link between individual responses and the group construct.

Define the group-level construct first, collect individual reports with an appropriate referent, justify the composition or aggregation process, evaluate agreement and reliability where relevant, sample enough groups for the intended inference, and keep group-level conclusions separate from individual-level claims unless both have been analyzed directly.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes