03 · What You Need to Know
“Who Participated?” and “What Did You Analyze?” Are Different Questions
What is a research participant?
A research participant is generally a person who takes part in a research study.
Depending on the design, participants may:
- complete questionnaires;
- participate in interviews or focus groups;
- undergo an intervention;
- perform experimental tasks;
- provide biological samples;
- allow researchers to observe their behavior;
- contribute repeated measurements over time.
The term therefore refers primarily to participation in the research process.
If 400 undergraduate students answer a questionnaire, those 400 students are participants.
If 30 school principals are interviewed, those principals are participants.
If 60 patients are enrolled in a randomized trial, those patients are participants.
What is the unit of analysis?
The unit of analysis identifies what kind of entity the primary analytical cases represent and what kind of entity the substantive conclusions describe.
If a study concludes:
Students with stronger academic self-efficacy report greater persistence
then students are the units of analysis.
If it concludes:
Universities with more supportive AI policies have higher institutional adoption rates
then universities are the units of analysis.
If it concludes:
Research articles with international collaboration receive more citations
then research articles are the units of analysis.
This distinction is central to understanding what your study is actually studying.
The participant and unit of analysis are often the same
In a conventional individual-level survey, the distinction may barely be noticeable.
| Study |
Participant |
Unit of analysis |
| Student self-efficacy and engagement survey |
Student |
Student |
| Employee burnout and turnover-intention survey |
Employee |
Employee |
| Patient quality-of-life study |
Patient |
Patient |
| Teacher attitudes toward AI |
Teacher |
Teacher |
In each example, individuals provide data about themselves and the study compares those same individuals.
This common arrangement may explain why researchers sometimes treat “participant,” “respondent,” “unit of observation,” and “unit of analysis” as though they were interchangeable.
They are not always interchangeable.
A participant is usually also a unit of observation
When people complete surveys, interviews, or measurements, they are usually units of observation because information is collected from them.
The distinction between unit of analysis and unit of observation helps clarify the logic:
Participant: who takes part in the study.
Unit of observation: where or from whom a particular piece of information is obtained.
Unit of analysis: what kind of entity the substantive analytical cases and conclusions concern.
In many individual surveys, one person occupies all three roles. In more complex designs, the roles separate.
Participants can act as informants about a larger entity
Suppose researchers study organizational climate across 80 companies.
Twenty employees from each company complete a survey describing leadership support, communication, collaboration, and innovation.
The employees are participants.
They are also units of observation because their responses are directly recorded.
But if responses are appropriately combined to characterize each company's organizational climate and companies are then compared, the company becomes the unit of analysis.
Participant
Employee who contributes information.
Unit of analysis
Company whose organizational characteristics are being compared.
Participants do not become organizations simply because their scores are averaged
The transition from individual responses to organizational measurement requires justification.
Suppose employees answer:
“I personally receive enough support from my supervisor.”
An average of these responses tells researchers something about the average employee experience within an organization.
That is not automatically identical to a shared organizational property called “support climate.”
To interpret individual reports as a group-level construct, researchers may need evidence that respondents are describing a common target and that meaningful within-group agreement and between-group variation exist.
This becomes important when deciding when individual responses can legitimately represent a group-level measure.
The wording of questions can reveal whether participants are reporting about themselves or their group
Compare two questionnaire items:
“I feel confident using generative AI in teaching.”
This clearly measures an individual attribute.
Now compare:
“Faculty members in this university receive adequate institutional support for using generative AI.”
The respondent is still an individual participant, but the referent is the university environment.
Item wording, theoretical definition, and aggregation strategy should therefore align with the unit at which the variable will eventually be interpreted.
The person interviewed may be an informant rather than the analytical case
This distinction is common in qualitative case studies.
Suppose researchers conduct interviews with:
- the university president;
- the CIO;
- the dean;
- faculty members;
- instructional designers.
The study investigates how a university developed its institutional AI governance system.
The interviewees are participants.
But the university may be the primary case or unit of analysis. Multiple participants provide different perspectives on that case.
The analysis may triangulate their accounts with institutional policies, meeting documents, and observations.
The resulting conclusion concerns the university-level governance process rather than treating each interviewee as an independent analytical unit.
Multiple participants can therefore contribute to one unit of analysis
This structure occurs in many designs.
| Participants |
Possible unit of analysis |
| Students |
Classroom |
| Teachers |
School |
| Employees |
Department or organization |
| Household members |
Household |
| Community residents |
Neighborhood |
| Executives and employees |
Company |
The presence of many participants per analytical case is often a clue that the design has multiple levels.
One participant can also generate many observational units
The reverse structure is equally important.
Suppose 100 students complete a short survey every evening for 30 days.
There are 100 participants, but potentially 3,000 person-day observations.
If the question asks:
On days when students experience unusually high stress, do they sleep worse that night?
daily occasions are lower-level observations nested within participants.
If the question instead asks:
Are generally more stressed students poorer sleepers than generally less stressed students?
the relevant comparison is between participants.
Thus, participant count and observation count can be very different.
A repeated-measures study does not suddenly have thousands of independent participants
Collecting 20 observations from each of 50 people yields 1,000 observations, not 1,000 independent people.
Observations from the same person are related because they share the same individual.
An analysis that treats all 1,000 rows as independent can underestimate uncertainty because it ignores this dependence.
This is one example of why identifying nested or hierarchical data matters before choosing an analytical method.
Participants can be nested inside higher-level units
Common research structures include:
Students within classrooms
Teachers within schools
Employees within departments
Patients within hospitals
Citizens within countries
Participants within the same higher-level unit may share environments, policies, teachers, leaders, resources, or other contextual characteristics.
That means their observations may be more similar than observations from participants in different groups.
Recognizing this hierarchy becomes essential when the research question contains variables at more than one level.
A participant-level outcome can be explained by a group-level predictor
Suppose researchers ask:
Does university-level AI policy predict individual faculty adoption of generative AI?
Faculty members are participants.
The outcome, faculty adoption, is individual-level.
The predictor, university AI policy, is organizational-level.
The relationship therefore crosses levels.
The research question involves a cross-level relationship, and a simple statement that “faculty were the participants” does not adequately describe the analytical structure.
Participant-level data can coexist with group-level outcomes
Suppose faculty members report their perceptions of institutional support, and researchers calculate an appropriately justified university-level support score.
The outcome is university-wide AI adoption rate.
Faculty members supplied data for one variable, but universities are the entities compared in the main analysis.
Again:
Participant ≠ necessarily unit of analysis.
Experimental assignment can create another important distinction
Suppose 30 schools participate in an educational trial. Fifteen schools receive an intervention and fifteen receive usual practice. Researchers then measure 5,000 students.
The students are participants if they directly participate in study procedures.
But treatment was assigned at the school level.
The school is therefore the experimental unit for the intervention assignment.
Student outcomes remain lower-level observations nested within schools.
The study cannot pretend that 5,000 independent treatment assignments occurred simply because 5,000 students were measured.
The number of participants is not always the relevant sample size for every question
Consider 2,000 faculty participants from only 10 universities.
If the primary question concerns variation among individual faculty members, the study contains substantial lower-level information, although clustering remains relevant.
If the question concerns university-level characteristics, however, only 10 universities provide independent between-university information.
Reporting only “N = 2,000 participants” can make the study appear to contain much more information about organizations than it actually does.
Watch Out
A large participant count does not automatically provide a large sample at every level of analysis. If your substantive question concerns 10 universities, observing thousands of people within those universities does not create thousands of independent universities.
The unit of analysis determines what your conclusion can safely describe
Suppose the analysis finds that universities with higher average AI readiness have greater institutional AI adoption.
The finding describes universities.
It does not automatically establish that individual faculty members with greater AI readiness are more likely to adopt AI within those universities.
Moving from a group-level finding to an individual conclusion risks the ecological fallacy.
Conversely, an individual-level relationship does not automatically establish a corresponding organizational relationship.
The nouns used in the conclusion should match the level represented by the analysis.
Participants can have different roles within the same study
A complex study may include several categories of participants.
For example, a study of university AI transformation might include:
- faculty survey participants;
- student survey participants;
- administrator interview participants.
The study could still use universities as its primary case-level unit of analysis if all of these data sources are integrated to compare institutional transformation across universities.
Alternatively, the study could conduct separate faculty-level, student-level, and institutional-level analyses.
The participant categories alone do not determine the analytical unit.
Not every study involving people has people as its unit of analysis
Researchers may use people to obtain information about:
- families;
- teams;
- communities;
- institutions;
- policies;
- events;
- decision processes;
- social interactions.
This is particularly common in organizational case studies, ethnography, policy research, and multilevel survey research.
Human participation and individual-level analysis therefore should not be treated as synonyms.
Some studies have units of analysis but no conventional participants
A bibliometric study may analyze journal articles.
A content analysis may analyze policy documents.
A historical study may analyze speeches or archival records.
A platform study may analyze posts, transactions, or interactions.
These studies have units of analysis even though they may have no participants in the conventional human-subject sense.
This demonstrates that “participant” and “unit of analysis” belong to different conceptual categories.
Secondary-data studies make this distinction especially visible
Suppose researchers analyze a national dataset originally collected from individuals.
The present researchers may never interact with the people who originally supplied the information.
Yet the analytical units may still be individuals if individual records are being analyzed.
Alternatively, the researchers could aggregate those records by province and analyze provinces.
Participation in original data collection and unit of analysis in a secondary study are therefore separate matters.
Qualitative studies can define analytical units at several scales
In qualitative research, researchers may refer to individuals as cases, but the analytical focus can also be:
- a family;
- an organization;
- a classroom episode;
- a decision process;
- a conversation;
- a policy implementation case.
Several participant interviews may contribute evidence to one case.
Researchers should therefore state explicitly what constitutes the case or analytical unit rather than assuming the interviewee defines it automatically.
Participant sampling and unit-of-analysis sampling may occur at different stages
Suppose researchers:
1. sample 30 universities;
2. sample 10 departments within each university;
3. recruit faculty members within those departments.
The study contains sampling decisions at multiple levels.
The primary unit of analysis depends on the research question, not merely on the final recruitment stage.
If the question concerns faculty attitudes, faculty members may be primary analytical units.
If the question concerns differences among universities, universities are the relevant higher-level units.
The key is alignment across question, data, analysis, and conclusion
Before collecting data, researchers should be able to distinguish four questions:
| Question |
What it identifies |
|
Who takes part?
|
Participants |
|
Where does the information come from?
|
Units of observation |
|
What entities are compared?
|
Units of analysis |
|
What entities do the conclusions describe?
|
Target of substantive inference |
If all four answers are “individual students,” the design is straightforward.
If the answers differ, the study is not necessarily wrong. It simply requires more explicit methodological reasoning.
03 · What You Need to Know
Multiple Units Are Legitimate When the Research Problem Is Genuinely Multilevel
Start with the simplest case: one unit of analysis
Many studies focus on one type of entity.
For example:
Students: Is academic self-efficacy associated with persistence?
Schools: Are school resources associated with graduation rates?
Countries: Is national research investment associated with scientific output?
Articles: Is international collaboration associated with citation impact?
In these studies, one main unit can adequately describe the substantive question.
Understanding what the unit of analysis is remains the first step before deciding whether the study actually needs more than one.
Multiple units arise when the questions refer to different kinds of entities
Consider a study involving faculty and universities.
Question 1:
Do faculty members with greater AI self-efficacy report greater AI adoption?
Unit of analysis: faculty member.
Question 2:
Do universities with formal AI governance frameworks have higher institutional adoption rates?
Unit of analysis: university.
The study now contains at least two meaningful analytical units because it asks distinct substantive questions about individuals and organizations.
Multiple units by design
Different research questions deliberately concern different entities.
Unit confusion
The analysis unintentionally mixes entities without specifying which claims belong to which level.
Having several data sources does not automatically mean having several units of analysis
Suppose a university case study uses:
- faculty interviews;
- administrator interviews;
- policy documents;
- meeting minutes;
- website materials.
That is multiple data sources.
But if all sources are integrated to understand one university as a case, the university may remain the single primary unit of analysis.
Participants, documents, and observations provide evidence about that analytical unit.
This is why participants are not automatically units of analysis.
Likewise, many observations do not automatically create many kinds of units
If 500 students each provide one survey response and the study examines individual-level relationships, there may still be only one unit of analysis: student.
If those students are nested within 20 classrooms but classroom differences are irrelevant to the substantive questions, the classrooms are nevertheless part of the data structure, though they may not be substantive analytical targets.
The distinction is important:
Hierarchy in the data does not automatically mean every level is a substantive unit of analysis.
Nested data frequently create the possibility of multiple analytical levels
Common hierarchies include:
Repeated observations within people
Students within classrooms
Classrooms within schools
Employees within departments
Departments within organizations
Patients within physicians within hospitals
These structures are examples of nested or hierarchical data.
The researcher may ask questions at one level, multiple levels, or across levels.
Multiple levels of data are not identical to multiple units of substantive inference
Suppose students are nested within schools.
The outcome is student achievement, and the question asks whether student study habits predict achievement.
School clustering may need to be accounted for statistically because students within the same school are not completely independent.
But the substantive unit of analysis may still be the student.
Schools matter statistically without necessarily becoming a separate substantive target.
By contrast, if the study also asks whether school-level leadership explains differences among schools, schools become an explicit analytical level in the scientific argument.
Different levels can contain different variables
Suppose employees are nested in organizations.
| Individual-level variables |
Organization-level variables |
| Employee self-efficacy |
Organizational size |
| Job satisfaction |
Formal AI policy |
| Age |
Leadership structure |
| AI adoption behavior |
Technology infrastructure |
A multilevel study can examine relationships among individual variables, relationships among organization variables, and relationships connecting the two levels.
This is why distinguishing individual-level and group-level variables becomes important once multiple analytical units are involved.
Within-individual and between-individual questions can also represent multiple levels
Multiple analytical levels do not require organizations or groups.
Suppose participants report stress and sleep every day.
Researchers could ask:
Between-person question: Do people who are generally more stressed sleep less than people who are generally less stressed?
and:
Within-person question: On days when a person is more stressed than usual, does that person sleep less than usual?
The observations are measurement occasions nested within people.
The same variables appear in both questions, but the source of variation differs.
Within-person and between-person relationships need not be identical.
One study can therefore ask separate questions at different levels
Consider faculty AI adoption.
| Question |
Primary level |
| Does faculty AI self-efficacy predict personal AI adoption? |
Individual faculty |
| Do universities with stronger AI infrastructure have higher adoption rates? |
University |
| Does university infrastructure predict individual faculty adoption? |
Cross-level |
| Does infrastructure change the relationship between self-efficacy and adoption? |
Cross-level interaction |
A coherent multilevel study can examine all four if theory, sample structure, measurement, and analytical capacity support them.
A cross-level relationship connects different analytical levels
Suppose:
University AI policy → Faculty AI adoption
The predictor belongs to the university level.
The outcome belongs to the faculty level.
The relationship crosses levels.
This is different from simply correlating two faculty-level variables or two university-level variables.
Recognizing cross-level relationships prevents researchers from pretending that variables defined at different levels are ordinary interchangeable columns.
Cross-level moderation adds another layer
Suppose the relationship between individual AI self-efficacy and AI adoption depends on university support.
The proposed relationship is:
Faculty self-efficacy → Faculty adoption
but its strength varies according to:
University-level support.
This is a cross-level interaction.
For example, self-efficacy may translate more strongly into adoption in universities that provide infrastructure and policy support than in institutions where confident faculty still face organizational barriers.
The research question therefore involves both individual and organizational units.
Aggregating individuals can create a higher-level analytical unit
Suppose 40 employees within each organization rate organizational climate.
If the theoretical construct is organizational-level and aggregation is justified, researchers may calculate an organization-level score.
The individual respondents remain sources of observation, while the organization-level score enters a higher-level analysis.
This is the logic behind aggregating individual-level data to the group level.
Aggregation does not eliminate the individual level from reality
If individual scores are averaged, variation among individuals within each group disappears from the aggregated dataset.
That may be appropriate for a purely group-level question.
But if researchers also care about individual-level relationships, collapsing everything to group means throws away information needed to answer those questions.
A multilevel analysis can often retain both sources of variation rather than forcing researchers to choose one and discard the other.
Multilevel models are often designed precisely for this structure
Multilevel, hierarchical, or mixed-effects models allow researchers to represent observations clustered within higher-level units.
A simplified two-level model might distinguish:
Level 1: faculty members
Level 2: universities
Individual outcomes can then be modeled using both faculty-level characteristics and university-level characteristics while accounting for dependence among faculty from the same university.
This is often preferable to pretending all faculty observations are independent or averaging everything into one university-level row when individual variation is substantively important.
Multiple units require enough information at each relevant level
Suppose 2,000 faculty members are sampled from only eight universities.
The study has extensive individual-level information but very limited between-university information.
Estimating complex university-level effects or cross-level interactions may therefore be difficult even with a large overall participant count.
Watch Out
Lower-level sample size cannot fully compensate for a very small number of higher-level units. If the study asks university-level questions, the number and diversity of universities matter independently of how many faculty members are observed within each university.
The unit relevant to each parameter matters for statistical power and precision
An individual-level coefficient draws heavily on variation among individuals.
A university-level coefficient depends on variation among universities.
A cross-level interaction requires information about both.
Thus, a study containing 5,000 students but only five schools may estimate some student-level relationships precisely while providing extremely limited evidence about school-level predictors.
One headline sample size does not describe the information available for every part of a multilevel model.
Multiple units require clarity about independence
Standard regression methods often assume that observations are independent conditional on included predictors.
But students in the same classroom share teachers and environments. Employees in the same organization share policies and leadership. Repeated measurements from one individual share the same person.
Ignoring these dependencies can underestimate standard errors and produce misleading inference.
This is one reason multilevel structure matters even when higher-level units are not themselves the main substantive focus.
Separate analyses at each level can sometimes be appropriate
Not every multiple-unit study requires one enormous multilevel model.
A project might deliberately conduct:
- one faculty-level analysis;
- one university-level analysis;
- a qualitative institution-level comparison.
If these are distinct research questions with distinct datasets or inferential targets, separate analyses may be clearer than forcing all questions into one model.
The important requirement is transparency about which unit each analysis addresses.
But separate analyses cannot answer every cross-level question
Suppose researchers separately find:
1. faculty self-efficacy predicts individual adoption;
2. university support predicts institutional adoption rate.
Those two findings do not establish that university support changes the individual self-efficacy–adoption relationship.
That cross-level moderation question requires an analysis capable of linking levels appropriately.
The same construct can exist at more than one level
Consider “support.”
At the individual level:
Perceived personal support may describe one employee's experience.
At the group level:
Support climate may represent a shared property of an organization.
These constructs are related but should not automatically be treated as identical simply because both are called support.
Likewise:
- individual efficacy differs from collective efficacy;
- personal trust differs from team trust climate;
- individual adoption differs from organizational adoption;
- personal socioeconomic status differs from neighborhood socioeconomic context.
Multiple levels require conceptual distinctions as well as statistical ones.
Group means can separate within-group and between-group information
Suppose individual workload predicts burnout and employees are nested within organizations.
An employee's workload score contains at least two conceptually different pieces of information:
1. how overloaded the employee is relative to colleagues in the same organization;
2. whether the organization itself tends to have high average workload.
These can have different relationships with burnout.
A multilevel approach can separate within-group and between-group components rather than assuming one coefficient captures both.
Relationships can differ depending on the unit being compared
Suppose individual employees with greater autonomy report greater satisfaction.
That does not guarantee that organizations with greater average autonomy have greater average satisfaction by exactly the same amount.
The individual-level and organization-level relationships can differ in magnitude or direction because they represent different variation.
This is why a relationship at one level can differ from the relationship at another.
Multiple units help avoid ecological and atomistic reasoning
If researchers observe a relationship among organizations and assume the same relationship holds among individuals, they risk an ecological fallacy.
If they observe an individual-level relationship and automatically infer an organization-level relationship, they risk an atomistic or individualistic fallacy.
A multilevel framework can preserve both forms of variation and make clear which conclusion belongs to which analytical unit.
One variable can be defined at one level and measured from another
Suppose organizational climate is a university-level construct measured through responses from individual faculty members.
The variable's conceptual level is university.
The source of the observations is individual.
Researchers therefore need a valid measurement bridge from individual reports to the higher-level construct.
This is exactly why collecting data from individuals for a group-level question requires explicit justification.
Qualitative research can also contain several units of analysis
A comparative case study may examine universities while also analyzing departments or teams nested within those universities.
For example, researchers might ask:
Institutional question: How do universities differ in AI governance?
Departmental question: How do departments interpret and implement university AI policies?
Individual question: How do faculty members respond to those departmental practices?
A qualitative multilevel case design can address all three if sampling, data collection, coding, and interpretation preserve the distinctions among levels.
Mixed-methods research does not automatically imply multiple units
A study can use surveys and interviews yet still have one unit of analysis.
For example, both methods may focus entirely on individual teachers.
Conversely, a purely quantitative study can have multiple units when students are nested in schools and the study analyzes both student- and school-level relationships.
Number of methods and number of analytical units are separate design dimensions.
Multiple units should be visible in the research questions
A study claiming to examine multiple levels should not hide that structure until the methods section.
Compare:
RQ1: How is faculty AI self-efficacy related to individual classroom AI adoption?
RQ2: How is university AI infrastructure related to institutional adoption rates?
RQ3: Does university AI infrastructure moderate the relationship between faculty self-efficacy and individual adoption?
The unit and level of each question are visible immediately.
This clarity helps prevent what happens when a research question mixes individual-level and group-level explanations without specifying the intended structure.
The hypotheses should preserve those levels too
A hypothesis such as:
Institutional support positively affects adoption
is ambiguous if “institutional support” is university-level but “adoption” could mean either individual faculty behavior or organizational adoption rate.
A clearer hypothesis is:
Universities with stronger AI-support climates will have higher institutional adoption rates.
or:
Faculty members working in universities with stronger AI-support climates will report greater individual AI adoption.
The entities in the sentence reveal the level of inference.
Sampling should also reflect every substantive unit
If universities are an important analytical unit, the study needs a defensible sample of universities.
If departments are also substantively important, the sampling strategy should provide adequate departmental variation.
Recruiting hundreds of participants from one or two higher-level units cannot provide broad evidence about variation among higher-level entities.
Multilevel questions therefore need multilevel sampling logic.
Measurement should match each unit
Individual constructs should be measured as individual constructs.
Group-level constructs need group-level theory and suitable operationalization.
For example:
Individual: “I have access to adequate AI support.”
Group referent: “Faculty members in this university have access to adequate AI support.”
Neither wording is automatically superior. They answer different measurement questions.
Analysis should not erase the hierarchy merely for convenience
Researchers sometimes simplify multilevel data by either:
- ignoring the groups and analyzing all individuals as independent; or
- averaging all individuals within groups and analyzing only the group means.
Either strategy may be reasonable for some specific questions, but both discard information.
Ignoring grouping loses the dependence structure and contextual variation.
Complete aggregation eliminates within-group variation.
If both levels matter substantively, a multilevel model often provides a more faithful representation.
A study can also have sequential units of analysis
Some research designs move from one unit to another across phases.
For example:
Phase 1: Survey individual faculty members.
Phase 2: Identify universities with unusually high or low adoption rates.
Phase 3: Conduct institutional case studies of selected universities.
The project uses individual-level analysis to inform selection of organization-level cases.
This is legitimate if the transition between analytical units is planned and explained.
Do not call every layer a unit of analysis merely because it exists
A dataset may contain:
items within scales within respondents within teams within organizations.
Not every layer is necessarily a substantive unit of analysis.
Questionnaire items may be measurement indicators rather than substantive analytical entities. Teams may be clustering units without being objects of substantive inference.
Use the term unit of analysis for entities relevant to the actual analytical or inferential question rather than every structural component in the data file.
Multiple units increase the importance of interpretation discipline
Suppose a model contains faculty-level and university-level effects.
The discussion should preserve those differences:
Individual-level finding: faculty with greater self-efficacy report greater adoption.
University-level finding: universities with stronger infrastructure have higher average adoption.
Cross-level finding: the self-efficacy–adoption relationship is stronger in universities with greater infrastructure.
These are three different findings.
Collapsing them into “self-efficacy and infrastructure increase AI adoption” loses the multilevel meaning of the study.