Stop Ranking Lecturers Against Each Other: The Case for Criterion-Referenced Course Evaluation
Reporting that a lecturer is "in the 40th percentile of the department" feels rigorous and is anything but. Norm-referencing manufactures a bottom half by mathematical necessity. Criterion-referenced standards ask a better question: did the teaching meet defined good-practice criteria?
Koji Education Team
Product ·
The single most damaging thing universities do with course-evaluation data is rank instructors against one another. "Dr A is in the 40th percentile of the school." "This module is below the faculty median." It sounds rigorous, comparative, evidence-based. It is, in fact, a category error. Norm-referenced reporting guarantees that half of any group falls "below average" no matter how good the teaching is, amplifies tiny and statistically meaningless differences into rankings, and imports a competitive logic that the underlying data cannot support. The defensible alternative is criterion-referenced evaluation: judging teaching against explicit, pre-defined standards of good practice rather than against the performance of colleagues.
The mathematical trap of norm-referencing
Start with the arithmetic, because it is decisive. A norm-referenced system reports where each instructor sits relative to a distribution of other instructors — percentiles, ranks, "above/below the median." In any such system it is mathematically impossible for everyone to be above average. Half of every department sits below the median by construction, including departments where every member is an excellent teacher. You have built a machine that manufactures a failing bottom half regardless of the absolute quality of teaching. As the assessment literature puts it, criterion-referenced systems make it "entirely possible — and desirable — for every student to meet the standard," whereas norm-referenced systems cannot (Teachers Institute, norm- vs criterion-referenced).
Now layer on measurement error. Course-evaluation means are noisy, especially in the small classes typical of seminars and final-year options (see why small classes make course-evaluation statistics so uncertain). The differences between adjacent ranks are routinely smaller than the confidence interval around any single mean. Ranking on those differences is ranking on noise. The American Statistical Association's 2019 Statement on Using Student Evaluations of Teaching warned explicitly that student-rating differences should not be over-interpreted and that scores should not be used to make fine comparative distinctions between instructors. Greg Boysen's research showed the problem is not merely theoretical: evaluators read large, meaningful differences into trivial gaps between mean ratings even when explicitly warned not to (Boysen, 2015). Percentile reporting actively invites that error.
There is also a representativeness problem borrowed from the testing world: if the comparison group does not adequately represent the thing being measured, the ranking "can be misleading and unfair" and can "perpetuate systemic inequalities" (norm-referenced fairness concerns). A lecturer teaching a hard, required, quantitative course is benchmarked against colleagues teaching small popular electives — and the known biases of student ratings (difficulty, electivity, discipline) load the comparison before a single student responds.
What criterion-referenced evaluation looks like
Criterion-referencing flips the question from "How does this teaching compare to other teaching?" to "Did this teaching meet defined standards of good practice?" The standard is set in advance, is the same for everyone, and can — desirably — be met by all.
In practice this means:
- Define the criteria first. Drawn from evidence on effective teaching: clarity of learning outcomes and constructive alignment, timely and useful feedback, inclusive practice, appropriate challenge, responsiveness to students. These become the evaluation framework, not a ranking table.
- Set thresholds against the criterion, not the cohort. "At least 80% of students agree feedback helped them improve" is a criterion. "Above the school median on feedback" is a norm. The first is achievable by every course; the second cannot be.
- Report met / developing / not-yet-met, with the qualitative evidence behind each judgement, rather than a percentile.
- Treat the mean as a flag, not a verdict. A low criterion result triggers a conversation and a closer look, not an automatic ranking penalty.
This is not a lowering of standards — it is a clarifying of them. Criterion-referencing makes explicit what "good teaching" means at your institution, which a percentile never does. A 60th-percentile score tells a lecturer nothing about what to change.
But doesn't comparison drive improvement? And don't we need to make decisions?
The strongest objection is managerial. Deans and quality committees need to make decisions — about probation, promotion, resource allocation — and comparison feels indispensable. "If we can't rank, how do we identify who needs support? Doesn't a little competitive pressure raise the game?"
Three responses. First, criterion-referencing does identify who needs support — more precisely than ranking. A lecturer who falls short on the "useful feedback" criterion has a named, actionable gap; a lecturer in the 35th percentile has a number. The criterion tells you what to develop. Second, the evidence that ranking improves teaching is thin, while the evidence that it distorts behaviour is not: norm-referenced stakes incentivise grade leniency, entertainment over rigour, and gaming, because the goal becomes out-scoring colleagues rather than meeting a standard. Third, for genuinely high-stakes personnel decisions, the consensus of the measurement community — including the ASA — is that student ratings should be one input among several (peer observation, teaching portfolios, evidence of student learning) and never the sole or dominant basis for comparison (see student evaluations in tenure and promotion decisions). Norm-referenced rankings give a false precision that crowds out that richer judgement.
A fair limitation: setting good criteria is harder than sorting a spreadsheet by mean score. It requires the institution to articulate what it values in teaching and to commit to standards — work that ranking lets you avoid. But that work is the point. And benchmarking against disciplinary or sector norms is not worthless — it has a legitimate, diagnostic role in spotting systemic patterns (we cover its careful use in benchmarking course-evaluation scores across departments). The error is turning a diagnostic comparison into a high-stakes ranking of individuals.
Where Koji fits
A criterion-referenced approach needs evidence against criteria, not just a number to sort. This is where conversational, AI-native evaluation has the advantage over a static Likert form. Koji for Education lets you build evaluations around your defined good-practice criteria using six structured question types — scale items for the threshold measures, open-ended and probing items for the evidence behind each criterion. Its AI-moderated interviews gather the qualitative substance a met/developing/not-yet judgement actually requires, with standardized, bias-aware moderation so the evidence is consistent across hundreds of courses rather than dependent on which human happened to run the focus group. Automatic thematic analysis and quality scoring map open feedback to your criteria, and programme- and institution-level reporting presents results as criterion attainment with the supporting voice of students attached — not a percentile league table. Because the same engine powers the main koji.so research platform, the underlying method — structured questions plus AI-probed open dialogue — is the same one teams use to evaluate any experience against defined standards.
Koji does not eliminate the judgement involved in setting criteria; nothing can. What it changes is the evidence available to make that judgement honest.
A worked example
Make it concrete. Suppose a school wants to know whether feedback practice is sound. The norm-referenced version produces a table: Dr A is 12th of 20 on the "feedback was useful" item, Dr B is 4th, and a committee duly worries about the bottom third — even though all twenty scored above 4.0 and the gaps between them are within the margin of error. Nobody learns what to change; several good teachers are made to feel they are failing. The criterion-referenced version asks instead: did at least 80% of students agree the feedback helped them improve, and was it returned in time to use? On that standard, eighteen courses clear the bar and two do not — and the two that fall short come with open-text evidence showing why (marks returned after the next assignment was due). The first approach sorts twenty people into a hierarchy that the data cannot support. The second identifies two specific, fixable problems and leaves the other eighteen colleagues to get on with teaching. One manufactures anxiety; the other manufactures improvement.
The bottom line
Ranking lecturers by percentile is a statistical artefact dressed as accountability. It guarantees a bottom half, ranks on noise, and bakes in known biases. Define what good teaching means, measure each course against that standard, report attainment against criteria with the evidence attached, and reserve comparison for diagnosis rather than judgement. The question is not "who is below average?" — that question always has the same depressing answer. The question is "did the teaching meet the standard, and if not, what specifically should change?"