New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

Measuring Consensus, Not Just the Average: Mixed-Effects Location-Scale Models for Course Evaluation

Mixed-effects location-scale models model the spread of course-evaluation ratings — consensus versus polarisation — as an outcome in its own right, not just noise around the mean.

Koji Education Team

Product

In brief

Two instructors can share a mean of 4.0 out of 5 and be teaching completely differently: one earns near-unanimous 4s, the other splits the room into 5s and 3s. The average treats them as identical; a mixed-effects location-scale model does not. Developed for intensive longitudinal data in health research, this method models the mean (the "location") and the variability (the "scale") at the same time, lets each depend on covariates, and — decisively — allows a random scale effect, so that some instructors reliably produce more consensus and others more polarisation, over and above their average. For course evaluation, it turns disagreement from noise you average away into a signal you can measure, predict, and report.

What the research says

The mixed-effects location-scale model was introduced by Hedeker, Mermelstein and Demirtas (2008) for ecological momentary assessment data, where each subject supplies many repeated reports. Beyond the usual random subject effect on the mean, they added submodels for the within-subject variance — the scale — allowing it to depend on covariates and, importantly, to carry its own random subject effect. The model therefore estimates not only how high a person tends to score but how consistent they are, and whether that consistency itself varies from person to person. In their application, within-subject mood variability differed systematically by covariates, independently of average mood.

Rast, Hofer and Sparks (2012) applied the approach to individual differences in the within-person variability of affect, using Bayesian estimation, and showed that people differ reliably in their volatility and that these differences can be estimated and related to stable traits. Their contribution matters because it establishes within-unit variability as a meaningful, estimable quantity — not merely error to be pooled away.

Li and Hedeker (2012) extended the model to three levels — observations nested within days within subjects — adding random scale effects at the intermediate level. This nested structure is directly relevant to course evaluation, where responses are naturally nested (students within courses within instructors), so variability can and should be modelled at more than one level. Finally, Hedeker and Nordgren (2013) released MIXREGLS, software that makes these models estimable through a three-stage procedure with interfaces to SAS and R — establishing that this is implemented, usable methodology rather than theory alone.

Why it matters for course evaluation in practice

The mean is the default course-evaluation statistic, and it discards the very thing quality processes often care about most: agreement. A polarised 4.0 — some students love the course, some dislike it — and a consensual 4.0 mean different things for a redesign, for equity, and for how much weight a committee should place on the number. Existing tools touch parts of this: ceiling effects and skew describe the shape of a distribution, quantile regression models its tails, and robust estimators protect the mean from a few extreme scores. None of them models dispersion as an outcome with its own predictors and its own random effects.

A location-scale model does. It lets you ask whether variability shrinks as a course matures, whether large classes produce more polarised ratings than small ones, and whether a particular instructor is reliably more divisive than their mean suggests — a random scale effect. Because dispersion is estimated jointly with the mean rather than as an afterthought, the mean estimates themselves receive more honest standard errors when variance is heterogeneous — an advantage over a plain multilevel model or a GEE that assumes one residual variance for everyone.

Reporting a consensus measure alongside the mean changes conversations. An instructor with a middling average but exceptional consensus may be doing something that works for everyone; a high average with high polarisation may conceal a subgroup being underserved — an equity signal that a mean, or even a shrunken mean, would hide.

Limitations and honest caveats

The most important limit is data appetite. Modelling variance reliably needs many observations per unit — these methods come from studies with dozens of reports per person — and a class of twelve gives thin information about within-class spread, so scale estimates will be uncertain. Small classes are precisely where dispersion is least estimable, which is inconvenient because they are also where a mean is least stable. Interpretation demands care too: a large within-instructor variance can reflect genuine polarisation or simply a heterogeneous student intake, a confound rather than a property of the teaching; modelling covariates on the scale helps but does not settle causation.

There are technical caveats. Standard location-scale models assume conditional normality, which bounded, skewed, ceilinged Likert data violate, so ordinal or bounded variants — or at least sensitivity checks — may be needed. The variance submodel, random scale effects, and frequently Bayesian estimation are hard to explain to a committee, and a mis-explained "polarisation score" is easy to weaponise. Scanning many covariates for scale effects invites false positives. The disciplined approach is to use the model only where responses per unit are sufficient, pre-specify the scale predictors, check distributional assumptions, report scale effects with their uncertainty, and treat a polarisation flag as a prompt to read the open text rather than a verdict.

How Koji incorporates this

Koji's structured scale items provide the numeric distribution a location-scale model needs, and its AI-moderated open text supplies the reason behind the dispersion: when a model flags an instructor as unusually polarising, the thematic record can show whether two coherent camps exist — and what divides them — or whether the spread is simply noise. Koji is designed to report consensus alongside the average rather than the average alone, to attach uncertainty to any dispersion estimate so the small-class trap does not produce false polarisation flags, and to treat a polarisation signal as a trigger for qualitative follow-up rather than a standalone score.

This is Koji's bias-aware reporting applied to spread, not just level: a 4.0 is never presented as self-explanatory, and a divisive course is surfaced as a question to investigate, not a number to punish. As with every comparative signal in Koji, the platform frames dispersion as decision support for a human reviewer who owns the judgement. For product and customer teams, Koji's core research platform at koji.so applies the same AI-moderated interview engine, where "average satisfaction hides two segments" is an everyday finding.

Frequently asked questions

What does a location-scale model add over a normal multilevel model?

A standard multilevel model estimates the mean and assumes one residual variance. A location-scale model additionally models that variance — letting it depend on covariates and carry its own random effect — so you can measure and predict consistency, not just level, and obtain better standard errors when variability is uneven.

How is this different from quantile regression or robust estimators?

Quantile regression models specific percentiles (the tails) of the outcome; robust estimators protect the mean from outliers. Neither treats dispersion as a quantity to be modelled with predictors and random effects. A location-scale model makes the spread itself the outcome.

What is a random scale effect?

It is the finding that some units are reliably more variable than others, beyond what their mean or covariates explain — for course evaluation, an instructor who is consistently more polarising (or more consensual) than average. It is the dispersion analogue of a random intercept on the mean.

Do I need a big class for this to work?

Yes, relatively. Reliably estimating variance needs many responses per unit; these models originate in studies with dozens of reports per person. For small classes, scale estimates are uncertain, so report them with wide intervals or fall back to describing the distribution directly.

Is a high-variance instructor a bad instructor?

Not necessarily. High variability can mean genuine polarisation, a mixed student population, or a course that suits some students far better than others. It is a signal to investigate — ideally with the open-text record — not a judgement in itself.

Related Resources

References

  • Hedeker, D., Mermelstein, R. J., & Demirtas, H. (2008). An application of a mixed-effects location scale model for analysis of ecological momentary assessment (EMA) data. Biometrics, 64(2), 627–634. https://doi.org/10.1111/j.1541-0420.2007.00924.x
  • Rast, P., Hofer, S. M., & Sparks, C. (2012). Modeling individual differences in within-person variation of negative and positive affect in a mixed effects location scale model using BUGS/JAGS. Multivariate Behavioral Research, 47(2), 177–200. https://doi.org/10.1080/00273171.2012.658328
  • Li, X., & Hedeker, D. (2012). A three-level mixed-effects location scale model with an application to ecological momentary assessment data. Statistics in Medicine, 31(26), 3192–3210. https://doi.org/10.1002/sim.5393
  • Hedeker, D., & Nordgren, R. (2013). MIXREGLS: A program for mixed-effects location scale analysis. Journal of Statistical Software, 52(12), 1–38. https://doi.org/10.18637/jss.v052.i12

Related articles