Beyond the Average Student: Latent Profile Analysis for Course-Evaluation Segments
A course mean of 3.6 can hide two entirely different student experiences averaged into one number. Latent profile analysis (LPA) recovers those hidden subgroups from the evaluation data itself, so you can see the delighted minority and the alienated cohort that the average erased.
Koji Education Team
Product
In brief
Latent profile analysis (LPA) is a model-based clustering method that identifies unobserved subgroups of students who share a similar pattern of course-evaluation responses, rather than assuming every respondent is a minor variation on one "average student." For course evaluation it answers a question the mean cannot: is a 3.6 the settled view of a homogeneous cohort, or the arithmetic midpoint of a satisfied cluster and a disengaged cluster that need opposite responses? Applied studies in higher education have used mixture models to classify university courses and cohorts by satisfaction pattern and to flag courses that are consistently underperforming for a subgroup.
What the research says
LPA belongs to the family of finite mixture models: it assumes the observed responses were generated by a mixture of a small number of latent classes, each with its own response pattern, and it estimates both how many classes there are and each student's probability of belonging to each. Spurk, Hirschi, Valero, Wang & Kauffeld (2020), in a widely cited review and "how to" guide in the Journal of Vocational Behavior, describe LPA as a categorical latent-variable approach for identifying latent subpopulations from a set of continuous indicators, and set out best-practice procedures for fitting and reporting it. The central methodological challenge — deciding how many profiles the data support — was addressed by Nylund, Asparouhov & Muthén (2007) in Structural Equation Modeling, whose Monte Carlo study found the Bayesian Information Criterion (BIC) and the bootstrap likelihood-ratio test (BLRT) to be the most reliable indicators of the correct number of classes, and cautioned against relying on cruder tests.
The approach has been applied directly to student satisfaction. Guerra, Bassi & Dias (2020), in Social Indicators Research, built a multiple-indicator latent growth mixture model on data from a large Italian university to track courses over time, identifying a clustering structure that isolated courses underperforming in student satisfaction across successive cohorts — a "low-quality teaching" profile that a simple year-on-year average would have smoothed away. Related work has used latent-class and two-level mixture item-response models to classify university courses by satisfaction while accounting for students' prior interest and perceptions of course organisation, showing that ostensibly similar mean scores can conceal qualitatively different satisfaction structures.
The unifying finding is that student evaluation data is frequently heterogeneous: it contains distinct response types (for example, uniformly enthusiastic, uniformly critical, and "content-but-for-assessment" patterns) whose existence is invisible in the aggregate. Person-centred methods like LPA recover those types; variable-centred methods (means, correlations, factor analysis) assume they do not exist.
Why it matters for course evaluation in practice
The mean, and even the full distribution of a single item, treats the cohort as one population. But two courses can post an identical 3.6 overall for opposite reasons: Course A because almost every student rates it a lukewarm 3-4, and Course B because half the class rates it 4.5 and half rates it 2.5. These demand different interventions — A needs a broad lift in a mediocre-but-uniform experience; B needs to understand why it is polarising a cohort, perhaps splitting along prior preparation, discipline background, or mode of attendance. A histogram hints at bimodality, but LPA does more: it uses the pattern across several items simultaneously to define the subgroups and can attach covariates (year of study, entry route, programme) to describe who falls into each profile. That turns "the class is split" into "there is a well-prepared, highly satisfied profile and a less-prepared profile that is struggling with the assessment load," which is directly actionable.
For programme-level quality assurance, the growth-mixture variant used by Guerra and colleagues is especially useful: it separates courses whose satisfaction is stably low (a structural problem) from those experiencing transient dips (noise or a one-off cohort), reducing the risk of over-reacting to regression-to-the-mean fluctuations while still catching genuinely persistent underperformance. In practice this changes how a QA committee triages its portfolio: instead of ranking every course by last term's mean and investigating the bottom decile — a list dominated by small classes and one-off noise — the committee can prioritise the courses that a mixture model places in a stable low-satisfaction trajectory, and set aside those whose apparent dip is a single anomalous cohort. The result is fewer wasted enhancement reviews and a defensible, evidence-based rationale for which courses receive scarce quality-improvement attention.
Limitations and honest caveats
LPA is powerful but easy to over-interpret, and a critical reviewer will press on several points. Profiles are model artefacts, not natural kinds. The classes are estimated, not observed; a different set of indicators, a different sample, or a different estimator can yield different profiles, and the solution should be cross-validated before it drives decisions. Class enumeration is genuinely hard. As Nylund et al. (2007) show, fit indices disagree, and the "best" number of classes is a judgement informed by BIC/BLRT and interpretability and theory — not a purely statistical verdict. Small samples and small classes are unstable; a profile containing a handful of students may not replicate, and per-course evaluation data is often too thin to support more than two or three classes reliably. Local maxima in estimation mean the model must be run from many random starts, or it can converge to a spurious solution. There is also a reification risk: labelling a cluster "the disengaged students" can harden a statistical convenience into an identity and invite exactly the kind of stereotyping that undermines fair evaluation. Finally, LPA describes structure; it does not explain it. Knowing there are two profiles does not tell you why — that still requires qualitative follow-up.
How Koji incorporates this
Koji is designed to make the inputs to a sound person-centred analysis available and to surface subgroup structure in reporting, without overclaiming that a dashboard cluster is a validated latent class.
- Multi-item structured data. LPA needs several indicators per respondent. Koji's structured questions (scale, single_choice, multiple_choice, ranking, yes_no) produce the multivariate response vectors that mixture models require, rather than a single overall score that cannot reveal profiles.
- Covariates for describing profiles. Koji captures cohort and context metadata, so an emergent profile can be described by who is in it (year, programme, attendance mode) — the step that turns a cluster into an actionable finding.
- Segment-aware reporting and triangulation. Koji's reporting is built to compare responses across cohorts and subgroups rather than collapsing everything to one mean, helping evaluators notice polarisation and triangulate a suspected split across cohorts before committing to it.
- Qualitative explanation of the segments. Because LPA describes but does not explain, Koji's AI-moderated conversational interviews and automatic thematic analysis of open-text responses can probe why a subgroup diverges — attaching the "why" to the "what" that a mixture model surfaces.
- Honest framing. Koji presents subgroup patterns as signals to investigate, not as certified latent classes; a formal LPA with proper class enumeration and cross-validation remains a dedicated statistical exercise best run on exported data.
Koji is intended to help evaluators see and interrogate heterogeneity, not to replace a statistician's model-selection judgement. For teams that run broader segmentation work, Koji's core research platform at koji.so applies the same AI-moderated interview engine to customer and product research, where person-centred segmentation is a routine goal.
Frequently asked questions
How is LPA different from just looking at the histogram of an item? A histogram shows the distribution of one item. LPA defines subgroups using the pattern across several items at once, so it can distinguish students who are uniformly critical from those who are satisfied on everything except assessment — a difference invisible in any single histogram.
How many profiles should I extract? There is no fixed answer. Nylund, Asparouhov & Muthén (2007) recommend leaning on the BIC and the bootstrap likelihood-ratio test, combined with interpretability and theory. More classes always fit better statistically, so parsimony and replicability must constrain the choice.
Is LPA the same as k-means clustering? No. k-means is a distance-based heuristic that forces every case into one cluster. LPA is model-based: it estimates a probability of membership in each class and provides fit statistics for choosing the number of classes, making it more principled and more honest about uncertainty.
Do I need a huge sample? More than for a simple mean. Small samples and small emergent classes are unstable and may not replicate, so per-course LPA is often underpowered; the method is most dependable at programme or institution scale, or pooled across cohorts.
Does finding profiles tell me what to do about them? Not on its own. LPA identifies structure; the reasons behind a profile require qualitative follow-up. Pair the segmentation with targeted open-text probing before acting.
Related resources
- Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages
- Q-Methodology: Surfacing the Distinct Viewpoints Students Hold About a Course
- Why Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores
- Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores
- Importance-Performance Analysis: Turning Course-Evaluation Scores into a Priority Map
- When Almost Everyone Scores 4.5: Ceiling Effects and Skew
References
- Spurk, D., Hirschi, A., Wang, M., Valero, D., & Kauffeld, S. (2020). Latent profile analysis: A review and "how to" guide of its application within vocational behavior research. Journal of Vocational Behavior, 120, 103445. https://doi.org/10.1016/j.jvb.2020.103445
- Nylund, K. L., Asparouhov, T., & Muthén, B. O. (2007). Deciding on the Number of Classes in Latent Class Analysis and Growth Mixture Modeling: A Monte Carlo Simulation Study. Structural Equation Modeling, 14(4), 535-569. https://doi.org/10.1080/10705510701575396
- Guerra, M., Bassi, F., & Dias, J. G. (2020). A Multiple-Indicator Latent Growth Mixture Model to Track Courses with Low-Quality Teaching. Social Indicators Research, 147(2), 361-381. https://doi.org/10.1007/s11205-019-02169-x
Related articles
Some Raters Are Harsh, Some Items Are Hard: Why Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores
Averaging Likert scores treats every student as an equally calibrated measuring instrument and every item as equally hard. Rasch and Many-Facet Rasch models do not. We explain the method, the evidence that rater severity and item difficulty distort raw means, the limitations, and how Koji applies the same logic.
When Almost Everyone Scores 4.5: Ceiling Effects, Skew, and What They Do to Course-Evaluation Statistics
Course-evaluation ratings pile up at the top of the scale, producing a strong ceiling effect and negative skew that breaks the statistics most universities still report. Here is what the evidence shows and how to report ratings honestly.
Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores
Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.
Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages for Course Evaluation
Course-evaluation data has a nested structure — students within sections within instructors — and a flat average ignores it. This guide explains how multilevel (hierarchical linear) models partition variance, why ignoring clustering understates uncertainty, and what it means for fair reporting.