03 · What You Need to Know
The Unit of Analysis Is Usually a Design Decision, but Existing Data Can Sometimes Support New Analyses
Ideally, the unit is identified before data collection
The unit of analysis should normally follow from the research question.
If your question is:
Which faculty members are more likely to adopt generative AI?
the faculty member is the natural unit.
If your question is:
Which universities have higher institutional adoption?
the university is the natural unit.
Because the question determines what kind of entity the study needs to compare, identifying the unit early affects sampling, measurement, data collection, and analysis.
This is why the unit of analysis should generally be established during research design rather than discovered only after the dataset has been assembled.
But research questions can evolve
Research is not always perfectly linear.
Existing datasets may contain information capable of answering questions that were not part of the original primary analysis.
Secondary analysis, exploratory analysis, multilevel analysis, and reanalysis of archived data routinely ask new questions of previously collected information.
The existence of a new question is not itself a methodological problem.
The question is whether the dataset is fit for purpose for that new analytical unit.
A new unit usually means a new research question
Suppose the original question is:
Do faculty members who perceive greater support use AI more often?
Unit of analysis: faculty member.
You now aggregate responses by university and ask:
Do universities with higher average perceived support have higher average AI use?
Unit of analysis: university.
This is not simply another statistical test of the same proposition.
The new analysis compares different entities and a different source of variation.
Original analysis
Compares people with other people.
New analysis
Compares groups with other groups.
The new unit must actually exist in the data
Suppose faculty records include a university identifier.
That allows respondents to be linked to their institutions.
If no university identifier was collected and institutions cannot be reconstructed reliably, a university-level analysis may simply be impossible.
Likewise, if repeated measurements were collected but no participant identifier exists, researchers cannot reliably recognize which observations belong to the same person.
The desired hierarchy cannot be invented after the fact if the linkage information was never retained.
Changing units does not necessarily mean changing units of observation
You may still have collected observations from individuals.
For example:
Unit of observation = faculty member
while a new aggregate analysis uses:
Unit of analysis = university.
The distinction between unit of observation and unit of analysis becomes particularly useful here.
The biggest temptation is simple aggregation
Suppose the dataset contains 40 faculty respondents per university.
You can calculate:
Computationally, the transformation is straightforward.
Methodologically, it may or may not be justified.
A group mean can always be calculated, but it may not represent the group construct you want
Suppose the original survey asked:
“I personally feel confident using AI.”
Averaging that item across faculty can legitimately produce:
the average individual AI confidence among observed faculty in each university.
But it does not automatically produce:
university collective AI efficacy.
The latter is a different theoretical construct.
Watch Out
Post hoc aggregation can change the numerical level of a score without changing what the original items actually measured. Averages cannot manufacture a shared group construct from measurements designed only to capture private individual attributes.
The original item referent matters
Compare:
“I receive adequate support for using AI.”
with:
“Faculty in this university receive adequate support for using AI.”
The first emphasizes an individual's experience.
The second asks the respondent to characterize a shared institutional environment.
If the original questionnaire used only individual referents, converting the scale into a shared group climate after data collection may require a much stronger justification.
Some variables can be aggregated more easily than others
Composition variables may permit straightforward summaries.
For example:
- average faculty age;
- percentage of faculty using AI;
- mean publication count;
- proportion holding doctoral degrees.
These group summaries do not require members to agree.
They describe the composition of the observed group.
By contrast, shared constructs such as organizational climate, collective efficacy, or team cohesion usually require a more explicit composition argument.
Shared group constructs may require evidence that was not planned originally
Researchers may need to consider:
- whether respondents refer to the same collective target;
- within-group agreement;
- between-group variation;
- reliability of aggregate group scores;
- number of respondents per group;
- whether respondents were appropriate informants.
This is why deciding whether individual responses can represent a group-level measure is a substantive measurement decision.
The original sampling strategy may be unsuitable for the new unit
Suppose your original study recruited 1,500 faculty members using convenience sampling.
You happen to obtain respondents from 30 universities, but:
- one university contributes 400 respondents;
- several contribute only one or two;
- most universities come from one geographic region;
- institutions were never deliberately sampled.
The dataset may be useful for individual-level questions.
It may provide weak support for broad university-level comparisons.
Sampling people from groups is not the same as sampling groups
If your new unit is the university, the university sample becomes substantively important.
You should ask:
Which universities entered the dataset, and why?
A large individual N cannot compensate for a highly selective or very small set of higher-level entities.
The effective sample size can shrink dramatically
Suppose you have:
2,000 faculty from 20 universities.
An individual-level analysis has information from 2,000 faculty observations, while accounting for their clustering if needed.
A fully aggregated university-level analysis contains:
N = 20 universities.
The shift from 2,000 to 20 is not a minor technicality.
It changes precision, model complexity, power, and the number of relationships that can be estimated credibly.
More respondents within a group do not create more groups
Suppose 1,000 participants come from four universities.
You may estimate the characteristics of those four universities quite precisely.
You still have only four universities for studying relationships among universities.
No aggregation formula can create the missing higher-level replication.
The new unit may introduce dependence that the original analysis ignored
Suppose the original data were analyzed as individual observations.
Once you recognize that faculty are nested within universities, you may discover that the individual observations are not fully independent.
This can mean the original individual-level analysis itself should account for the nested or hierarchical data structure.
Changing the research question can therefore reveal information about the original analysis as well.
You do not always need to fully aggregate when adding a higher-level question
Suppose you want to study university context but still care about individual faculty outcomes.
Instead of collapsing all faculty data into university means, you might retain faculty-level observations and use a multilevel approach.
This can preserve:
- individual variation;
- university variation;
- individual predictors;
- university predictors;
- cross-level relationships.
In other words, adding a university-level question does not always require abandoning the faculty-level unit.
The study may legitimately gain an additional unit of analysis
If the data support both levels, the revised project can ask:
Individual question: Do faculty with greater self-efficacy adopt AI more?
University question: Do universities with stronger infrastructure have higher average adoption?
Cross-level question: Does university infrastructure predict individual faculty adoption?
This is a legitimate case where a study can have more than one unit of analysis.
Changing from groups to individuals is also possible only if individual information exists
Suppose your original dataset contains one row per school with:
- average achievement;
- average parental income;
- school size.
You cannot later ask:
Are individual students from wealthier households more likely to achieve highly?
The individual information has already been collapsed.
Once lower-level detail is absent, it usually cannot be recovered from group averages.
Aggregate values cannot reconstruct individual distributions uniquely
Suppose a classroom mean achievement score is 80.
Many different sets of student scores can produce that same mean.
For example:
80, 80, 80, 80
and:
60, 70, 90, 100
both average 80.
The group mean therefore cannot reveal the individual distribution from which it came.
This is one reason ecological data cannot simply be converted into individual data
If only university-level percentages are available, individual-level relationships generally cannot be recovered without additional information and strong assumptions.
Attempting to infer individual patterns directly from aggregate patterns risks the ecological fallacy.
Changing the unit can alter variable meanings
Consider AI adoption.
At the individual level:
Does this faculty member use AI?
At the university level:
What proportion of faculty use AI?
or perhaps:
Has AI become formally institutionalized?
These are not identical outcomes.
A change in unit can therefore require new operational definitions rather than simply new aggregation.
The outcome may not aggregate naturally
Suppose the original outcome is individual intention to use AI.
What is the corresponding university-level outcome?
Average intention?
Percentage above some threshold?
Actual institutional adoption?
Those choices represent different constructs.
The data may support only some of them.
Changing the unit can alter the theoretical model
A theory explaining individual adoption might emphasize:
- self-efficacy;
- perceived usefulness;
- personal anxiety;
- experience.
A theory explaining institutional adoption might instead emphasize:
- governance;
- budgets;
- leadership;
- technical infrastructure;
- regulatory conditions.
A new unit therefore may require a different conceptual framework rather than simply recycling the individual model at a higher level.
Do not assume individual predictors become organizational predictors when averaged
If individual self-efficacy predicts individual adoption, it does not follow that average university self-efficacy explains institutional adoption.
That upward inference risks the atomistic fallacy.
The higher-level relationship must be estimated and theoretically justified separately.
The reverse relationship may differ too
A group-level association can be stronger, weaker, absent, or reversed compared with the individual association.
This is not unusual because relationships can differ at different levels.
Changing units after seeing the results creates additional analytic-flexibility concerns
Suppose the original individual-level hypothesis is not supported.
You then try:
- department averages;
- university averages;
- regional averages;
- several alternative aggregation rules;
- different subgroup definitions.
If only the favorable result is reported, readers cannot distinguish a prespecified higher-level hypothesis from a post hoc search.
Exploratory analysis is legitimate, but it should be labeled as exploratory rather than presented as though the higher-level model had been planned from the beginning.
Post hoc does not mean worthless
Unexpected patterns can generate valuable new questions.
A transparent report can state that:
the original study was designed at the individual level, and a secondary exploratory analysis examined university-level aggregates because the data contained multiple respondents per institution.
The key is to separate discovery from confirmation.
Changing units may require a new power analysis or precision assessment
A study powered for 1,000 individual observations may be severely underpowered for 15 organization-level observations.
The original sample-size calculation cannot simply be carried over.
The revised analysis should be evaluated using the actual number of higher-level units and the proposed model.
A new unit may require different missing-data handling
Suppose some universities have many respondents and others have very few.
Missing individual responses can affect the precision and representativeness of university-level aggregates.
Completely missing universities create a different problem: no individual-level technique can recover institutional variation for universities never observed.
A new unit may require different weighting
Suppose University A has 500 respondents and University B has 10.
If the scientific question treats universities as equally important analytical entities, simply pooling all individuals gives University A far more influence.
A group-level or multilevel analysis needs to reflect the actual estimand and sampling design rather than letting cluster size determine the question accidentally.
Secondary data may legitimately support units the original investigators did not analyze
Imagine a national survey containing individuals nested within provinces, schools, or households.
The original publication may focus on individuals.
A secondary researcher could ask province-level or multilevel questions if:
- the grouping identifiers are available;
- the relevant constructs can be measured validly;
- sampling supports the new inference;
- enough groups are observed;
- the analysis respects the design.
The fact that the original authors did not analyze that level does not automatically prohibit a defensible secondary analysis.
Qualitative studies can also refine the unit during inquiry
In some qualitative and interpretive research, the final analytical unit can emerge or be refined as researchers learn more about the phenomenon.
For example, a study initially focused on individual teacher practices may discover that instructional teams are the meaningful social entities shaping those practices.
Such development can be legitimate when documented transparently and when the collected evidence adequately represents the revised unit.
Methodological discussion of units of analysis has explicitly recognized the tension between determining units in advance and refining them during inquiry, while emphasizing fit between the unit and the phenomenon being studied.
Changing the unit is more defensible when the new unit was already represented structurally
Suppose your dataset already contains:
30 respondents × 50 universities
with institutional identifiers and group-referent climate items.
A university-level secondary analysis may be quite plausible.
Compare that with:
1,500 respondents from unknown institutions
with no identifiers and purely individual-referent items.
The same post hoc university question is far less defensible.
A practical feasibility test uses five questions
Before changing the unit, ask:
1. Can I identify the proposed new entities?
2. Do I have valid variables describing those entities?
3. Do I have enough independently observed entities?
4. Does the sampling design support conclusions about them?
5. Does the new analysis answer a coherent research question at that level?
If the answer to one or more is no, changing the unit may produce a dataset transformation without producing a defensible study.
Sometimes the correct answer is to redesign the next study
Your existing data may suggest an excellent organizational hypothesis while being inadequate to test it.
That is still scientifically useful.
Instead of overstating the secondary analysis, use it to design a follow-up study that deliberately samples organizations, measures organizational constructs, and gathers enough respondents within each organization.
A limitation can therefore become the design logic for the next project.