New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

One Student, Many Teachers: Cross-Classified and Multiple-Membership Models for Fair Course Evaluation

Standard multilevel models assume a clean hierarchy, but students are taught by several instructors and instructors teach across programmes. Cross-classified and multiple-membership models partition that tangled variance honestly — and change which instructors look unusual.

Koji Education Team

Product

In brief

A plain multilevel model assumes course evaluation responses nest cleanly inside one course inside one instructor. Real teaching is not that tidy: a student is often taught by several instructors within a single module, and instructors reappear across programmes and cohorts. When the data structure is genuinely crossed or overlapping, forcing it into a strict hierarchy misattributes variance and can distort the very instructor comparisons a quality-assurance office cares about. Cross-classified random-effects models and multiple-membership models — introduced to education research by Goldstein and colleagues — represent that structure faithfully, giving each instructor a variance component that is not contaminated by the courses, cohorts, or co-teachers they happen to share.

This matters because unfair instructor rankings are frequently an artefact of a mis-specified model, not real differences in teaching. If you are already using multilevel models to avoid comparing raw averages, cross-classified and multiple-membership extensions are the next step when your teaching arrangements break the neat nesting assumption.

What the research says

Classical multilevel (hierarchical linear) models assume a strict tree: level-1 units (responses) nest in level-2 units (courses), which nest in level-3 units (departments), and so on. Each lower unit belongs to exactly one higher unit. Educational data routinely violate this. A widely cited structural review by Fielding and Goldstein (2006), Cross-classified and Multiple Membership Structures in Multilevel Models: An Introduction and Review, catalogues two common departures. In a cross-classified structure, two higher-level classifications are crossed rather than nested — for example, students are classified simultaneously by the course they took and by the instructor who taught them, and the same instructor teaches students who also appear under other courses. Neither instructor nor course nests inside the other; they cross.

In a multiple-membership structure, a single lower-level unit belongs to more than one higher-level unit at the same time. A student on a team-taught module is a member of two or three instructors simultaneously; a student who transfers between seminar groups is a member of both. Browne, Goldstein and Rasbash (2001) formalised the general multiple-membership multiple-classification (MMMC) framework and showed how membership weights (for example, the proportion of contact hours each instructor delivered) enter the model so that a student's outcome is a weighted function of every instructor who taught them.

The empirical payoff was demonstrated early. Hill and Goldstein (1998), Multilevel Modelling of Educational Data with Cross-Classification and Missing Identification for Units (Journal of Educational and Behavioral Statistics, 23(2), 117-146), fitted cross-classified models to pupil data where students were cross-classified by school and by neighbourhood, and showed that ignoring the crossed classification biases the estimated variance components and their standard errors. More recent simulation work in other applied fields (for example, cross-classified models in clinical outcomes research) confirms the general lesson: when a crossed structure is collapsed into a false hierarchy, the standard errors of covariate effects are estimated incorrectly, so significance tests about who differs from whom become untrustworthy.

The estimation machinery is now standard. Cross-classified and multiple-membership models are fitted by maximum likelihood, Markov chain Monte Carlo (as in the MLwiN implementations Goldstein's group developed), or integrated nested Laplace approximation. What was once a specialist technique is now available in mainstream mixed-model software, so the barrier is conceptual — recognising the structure — rather than computational.

Why it matters for course evaluation in practice

Three features of modern higher education make crossed and overlapping structures the rule, not the exception:

  1. Team teaching and modular delivery. A 15-credit module may be delivered by four lecturers. If you assign the whole module's evaluation score to one "module leader", you either over-credit or over-blame that person for teaching they did not do alone. A multiple-membership model spreads each response across the instructors who actually taught it, weighted by contact hours.

  2. Instructors who span courses and years. An instructor teaches Statistics I this term and Research Methods next term, to overlapping cohorts. A strict hierarchy that nests instructors within a single course cannot separate a durable instructor effect from a one-off course effect. A cross-classified model estimates an instructor variance component and a course variance component simultaneously, so you can ask the fair question: how much do evaluation scores vary between instructors once the course they taught is accounted for?

  3. Fair flagging. Quality committees increasingly flag instructors whose scores sit outside an expected band — the logic behind funnel plots and empirical-Bayes shrinkage. Those methods rest on a variance partition. If the partition is wrong because the model ignored crossing, the flags inherit the error. Getting the structure right is a precondition for defensible comparison, not a refinement after the fact.

The team-taught attribution problem is often discussed as a conceptual dilemma; multiple-membership modelling is the concrete statistical answer to it. Similarly, the Rasch and many-facet tradition separates rater severity from item difficulty; cross-classified models do the analogous job for the design facets (course, instructor, cohort) that produce a rating.

Limitations and honest caveats

These models are not a licence to over-interpret small data. Several cautions apply:

  • Variance components need data to estimate. An instructor who appears in only one course, taught alone, contributes almost nothing to separating the instructor classification from the course classification. If most of your instructors teach exactly one course solo, the crossed structure is weakly identified and a cross-classified model will not rescue you — you simply lack the cross-links that make the decomposition possible.

  • Membership weights are an assumption, not a measurement. Weighting each instructor by contact hours presumes influence on the evaluation is proportional to time. That is a modelling choice; a charismatic guest lecturer who taught two hours may shape a student's overall impression out of proportion to their share. Report the weighting scheme and test sensitivity to alternatives.

  • Crossed does not mean causal. Partitioning variance tells you how much scores differ between instructors; it does not tell you the difference is caused by teaching quality rather than by the confounds this corpus documents at length — course difficulty, discipline, class size, timing. Variance decomposition and confounding adjustment (for example, propensity-score methods or the E-value) answer different questions and should be combined, not substituted.

  • Complexity has a communication cost. A dean who struggles with a simple mean will not intuit a four-classification MMMC variance partition. The model can be correct and still fail if its output cannot be explained. Translate the result into a plain statement — "once we account for which course and which co-teachers were involved, genuine differences between instructors explain about X% of the variation" — before it reaches a committee.

  • Small higher-level counts bias variance estimates. With few instructors or few courses, maximum-likelihood variance estimates are biased downward and interval coverage is poor; Bayesian estimation with weakly informative priors is generally the more honest route, and its credible intervals should be reported rather than hidden.

How Koji incorporates this

Koji for Education is built so the analysis can respect the real teaching structure rather than a convenient fiction. Concretely:

  • Structured metadata on every response. Each conversational evaluation Koji collects is tagged with the module, the instructors who delivered it, the cohort, and the delivery mode. That metadata is exactly what a cross-classified or multiple-membership model needs; without it, the tidy-hierarchy shortcut is the only option available. Koji captures the crossing at source so it does not have to be reconstructed later.

  • Contact-share weighting for team-taught modules. Where a module records more than one instructor, Koji can attach membership weights (for example, by scheduled contact hours) so that reporting does not dump a shared score onto a single "module leader". This directly operationalises the multiple-membership logic for the team-taught attribution problem.

  • Variance-aware comparison, not raw league tables. Koji's instructor and programme reporting is designed to compare adjusted, shrunken estimates rather than raw averages, and to present the share of variation that is genuinely between-instructor after course and cohort are accounted for. The aim is to flag what is unusual without punishing an instructor for the difficulty of the course they were assigned.

  • Bias-aware, cautious flagging. Because small cohorts produce unstable estimates, Koji frames outlier flags as prompts for a conversation, not verdicts, and surfaces the uncertainty around each estimate. This is designed to mitigate — not eliminate — the risk that a mis-specified structure produces a false signal.

Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where crossed structures (customers who use several features, evaluated by several account teams) raise identical modelling questions.

Related resources

References

Frequently asked questions

What is the difference between a cross-classified model and a multiple-membership model?

In a cross-classified model each response still belongs to exactly one unit of each classification (one course and one instructor), but the two classifications cross rather than nest. In a multiple-membership model a single response belongs to several units of the same classification at once — a student taught by three instructors is a member of all three, usually with weights. Team teaching produces multiple membership; instructors reappearing across courses produce cross-classification. Real datasets often have both, which is why Browne, Goldstein and Rasbash (2001) combined them into one MMMC framework.

Why not just assign each module one "module leader" and use a normal hierarchy?

Because it manufactures a fiction. Assigning a team-taught module's whole score to one person over-credits or over-blames them for teaching several people delivered, and it hides the co-teaching structure from the variance decomposition. The estimate of "how much instructors differ" then absorbs course and co-teacher effects it should have separated out, so the resulting comparisons and any outlier flags are biased.

Do I need a huge dataset to fit these models?

You need cross-links, not just volume. The instructor and course classifications can only be separated if enough instructors teach more than one course and enough courses are taught by more than one instructor. If almost every instructor teaches a single course alone, the structure is weakly identified regardless of how many student responses you have, and the variance components will be unstable.

Does getting the structure right change who gets flagged as an outlier?

It can, materially. Outlier flagging rests on the estimated between-instructor variance and each instructor's standard error. Hill and Goldstein (1998) and later simulation studies show these are biased when a crossed structure is collapsed into a hierarchy. Since funnel plots and shrinkage flags are built directly on those quantities, correcting the structure can move instructors into or out of the "unusual" band.

Is this a causal method that proves an instructor teaches better or worse?

No. Cross-classified and multiple-membership models partition variance and estimate adjusted effects; they describe how much scores differ once design facets are accounted for. They do not remove confounding from course difficulty, discipline, or class size. For causal questions you still need a confounding-adjustment strategy such as propensity-score matching or an E-value sensitivity analysis, used alongside — not instead of — the correct multilevel structure.

How does Koji use membership weights?

Where a module records several instructors, Koji can weight each instructor's membership by a defensible share such as scheduled contact hours, so reporting reflects who actually taught. The weighting scheme is explicit and adjustable, because proportional-to-time is an assumption rather than a measured fact, and its sensitivity should be checked.

Related articles

analysis-reporting

Some Raters Are Harsh, Some Items Are Hard: Why Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores

Averaging Likert scores treats every student as an equally calibrated measuring instrument and every item as equally hard. Rasch and Many-Facet Rasch models do not. We explain the method, the evidence that rater severity and item difficulty distort raw means, the limitations, and how Koji applies the same logic.

analysis-reporting

Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages for Course Evaluation

Course-evaluation data has a nested structure — students within sections within instructors — and a flat average ignores it. This guide explains how multilevel (hierarchical linear) models partition variance, why ignoring clustering understates uncertainty, and what it means for fair reporting.

analysis-reporting

Which Instructor Scores Are Genuinely Unusual? Funnel Plots for Fair Course-Evaluation Comparison

Ranking instructors by mean evaluation score turns sampling noise into a league table. Funnel plots — Spiegelhalter's method for institutional comparison — show which scores are genuinely unusual and which are just small-sample wobble.

analysis-reporting

How Strong Would the Hidden Confounder Have to Be? The E-Value for Adjusted Course-Evaluation Comparisons

Every adjusted course-evaluation claim invites the objection "but you did not control for X". The E-value, from epidemiology, quantifies exactly how strong that unmeasured X would have to be to explain away your finding — turning a vague worry into a number.