New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Not Every Cohort Follows the Same Path: Latent Growth Curve and Growth Mixture Models for Course Evaluation

Latent growth curve and growth mixture models describe how student trajectories change across evaluation waves — and uncover subgroups that a single average path hides — while guarding against the subgroups that are only statistical artefacts.

Koji Education Team

Product

In brief

When you evaluate the same cohort at several points — early, mid, and end of a module, or term after term — the useful question is not only "did the average move?" but "did everyone move the same way?" Latent growth curve modelling (LGCM) treats each student's or cohort's trajectory as having a latent starting point and a latent rate of change, and estimates how much those differ across individuals. Growth mixture modelling (GMM) goes further, uncovering unobserved subgroups that follow qualitatively distinct paths — one improving, one flat-and-high, one declining — that a single average trajectory conceals. These methods are powerful for longitudinal evaluation, but GMM in particular can invent subgroups that are artefacts of non-normal data, so any class solution must be treated as a hypothesis, not a fact.

What the research says

Latent curve analysis was formalised by Meredith and Tisak (1990), who showed that repeated measures can be represented as a structural model with latent factors for the intercept (initial level) and slope (rate of change), imposing a structure on both the means and the covariances of the observed variables. Crucially, individuals are allowed to differ in both their starting point and their rate of change, and those individual differences are modelled as variances that can, in turn, be predicted by covariates. This reframed growth as an alternative to repeated-measures ANOVA and first-order autoregressive methods.

Curran, Obeidat and Losardo (2010), in an accessible "twelve frequently asked questions" primer, make an important honesty point: under many parameterisations LGCM is statistically equivalent to a multilevel model for change. What the latent-variable tradition adds is convenience — it is straightforward to add predictors of the slope, to model nonlinear trajectory shapes, to include several outcomes at once, and, by moving to a mixture, to search for distinct subgroups.

That mixture step is where growth mixture modelling and latent class growth analysis enter. Jung and Wickrama (2008) distinguish the two: latent class growth analysis (associated with Nagin) fixes the within-class variance to zero, so every member of a class shares one trajectory, while growth mixture modelling (associated with Muthén) allows individual variation around each class's mean trajectory. Both use finite-mixture estimation to recover latent trajectory classes, with the number of classes chosen by information criteria (BIC), classification quality (entropy), the bootstrap likelihood-ratio test, and — decisively — substantive interpretability.

Why it matters for course evaluation in practice

Course evaluation is increasingly longitudinal: a mid-term pulse plus an end-of-term survey, or a series across successive deliveries of a module. A raw average trajectory — or even a multilevel growth model that assumes a single population — can hide compositional heterogeneity. A redesign that lifts most students but alienates a minority can register as a flat average, when in fact two groups are moving in opposite directions.

LGCM answers who started where and who changed how fast, and lets you test whether a teaching change predicts the slope (the rate of improvement) rather than merely the level. GMM answers whether there are distinct response trajectories at all — for example, a group whose satisfaction climbs as a demanding course pays off, versus a group whose satisfaction erodes as workload bites. Framing an evaluation around trajectories rather than snapshots is often what turns "the numbers dipped in week six" into an actionable finding.

These models are distinct from their neighbours in this knowledge base. Latent profile analysis segments students at a single time point; latent transition analysis tracks movement between categorical states wave to wave; multilevel models model a single population's nesting; and the Reliable Change Index judges one student's two-point change against measurement error. LGCM and GMM model the shape of change across three or more waves and, in GMM, uncover subgroups following different shapes.

Limitations and honest caveats

The headline caveat comes from Bauer and Curran (2003), who demonstrated that finite normal mixtures can estimate multiple trajectory classes even when the population is a single homogeneous group, provided the repeated measures are non-normal. Multiple classes can appear optimal for skewed data even when only one group truly exists, and the within-class parameter estimates then become largely uninterpretable. Because course-evaluation scores are typically skewed and ceilinged, this risk is acute: a "three-class" solution may simply be describing the shape of one skewed distribution.

Other cautions follow. You need at least three waves to estimate a linear slope, and more for curved shapes; most end-of-term evaluation supplies too few. Class enumeration is unstable — BIC, entropy, and the bootstrap likelihood-ratio test can disagree, and solutions are sensitive to starting values and to how within-class variance and residuals are specified. Classes are latent, not observed, so naming one "the disengaged group" reifies a statistical construct. Attrition between waves biases trajectories if dropout is related to the outcome. The disciplined workflow is to pre-specify how many trajectories you expect and why, replicate across cohorts, report entropy and classification uncertainty, and treat classes as hypotheses to interrogate against the qualitative record — not as verdicts.

How Koji incorporates this

Koji's mid-cycle and always-on collection produces exactly the repeated measures these models require, rather than the single end-of-term snapshot that makes them impossible; its experience-sampling-style in-semester feedback is designed to build a longitudinal record keyed to student, cohort, and wave. Structured scale items supply the numeric trajectory, while the AI-moderated conversational open text supplies the reasons that let you interpret — or refute — a putative class.

Concretely, Koji is designed to hold responses in a longitudinal structure, to flag when an average masks divergent movement, and to pair any emergent trajectory class with the thematic evidence that either substantiates or dissolves it — a direct hedge against the over-extraction risk Bauer and Curran identified. Consistent with Koji's bias-aware reporting, trajectory classes are presented as hypotheses for a human reader to probe, not as settled segments. Teams running longitudinal product or customer cohorts can apply the same engine through Koji's core research platform at koji.so, where distinguishing real subgroups from artefacts of skew is an equally common problem.

A worked example

Suppose a redesigned second-year statistics module is evaluated in weeks 3, 7, and 11 across three consecutive cohorts. A single average trajectory shows satisfaction rising gently over the term. Fitting a latent growth curve first confirms meaningful variation in both the intercept and the slope: students do not start in the same place, and they do not improve at the same rate. Adding growth mixture modelling then suggests three trajectory classes — a large "steady climbers" group whose ratings rise as the material clicks, a smaller "early enthusiasts" group that starts high and drifts down as difficulty mounts, and a "persistently frustrated" minority that never recovers from a hard week-four assessment. Before believing the three-class solution, you check entropy (is classification clean?), confirm the result replicates across the three cohorts, and read the open-text comments for each putative class to see whether coherent, nameable experiences underlie them. Only the frustrated-minority class survives all three checks; the split between climbers and enthusiasts turns out to reflect skew rather than two real groups — exactly the over-extraction Bauer and Curran warned about. The actionable finding is the week-four assessment, not the invented enthusiast group.

Frequently asked questions

Is growth curve modelling just a multilevel model with a new name?

For a basic random-intercept-and-slope model, LGCM and a multilevel growth model are statistically equivalent. The latent-variable framing makes it easier to add predictors of the slope, model nonlinear shapes, include several outcomes, and — via mixtures — search for distinct trajectory subgroups, which is where it goes beyond a standard multilevel model.

How is GMM different from latent profile or latent transition analysis?

Latent profile analysis segments students at a single time point; latent transition analysis tracks movement between categorical classes wave to wave; growth mixture modelling classifies by the shape of a whole trajectory over three or more waves. GMM asks "who follows which path," not "who is in which state now."

How many time points do I need?

At least three to estimate a linear trajectory — two points give only a difference, which the Reliable Change Index handles better — and more for curved shapes. Most end-of-term evaluation gives too few waves, which is why mid-cycle collection is what makes these models feasible.

Could the subgroups be an artefact?

Yes. Bauer and Curran (2003) showed that non-normal, skewed data can produce multiple apparent classes even when the population is homogeneous. Because course-evaluation scores are usually skewed and ceilinged, treat any class solution as a hypothesis, report entropy and classification uncertainty, and replicate before believing it.

Can I use these models to compare instructors?

Indirectly. They describe how cohorts change and whether subgroups diverge, and a teaching change can be tested as a predictor of the slope or of class membership. For direct fair comparison of instructor levels, pair them with shrinkage or funnel-plot methods.

Related Resources

References