New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting11 min read

Should You Report an Instructor''s Percentile? Norm-Referenced vs Criterion-Referenced Course-Evaluation Scores

Telling a lecturer they are "in the 40th percentile of the department" is norm-referenced reporting — and it manufactures losers by construction, no matter how good everyone is. Criterion-referenced reporting asks instead whether teaching met a defined standard. Here is the evidence on why the choice matters and how to report responsibly.

Koji Education Team

Product

In brief

Reporting a course-evaluation score as a rank or percentile against a comparison group ("you are below the departmental mean," "you are in the 40th percentile") is norm-referenced — and it guarantees that half of any faculty will look "below average" even if every one of them is teaching well. Criterion-referenced reporting instead compares each course against a defined, absolute standard of acceptable teaching. The research on how evaluation data is actually interpreted shows that norm-referenced comparison invites systematic over-interpretation: reviewers treat tiny gaps between a score and the comparison mean as meaningful differences in teaching quality, when those gaps are usually within the noise. Boysen (2015) demonstrated this over-interpretation directly; Theall and Franklin (2001) warned that misused comparison data is one of the commonest sources of unfairness in student-ratings systems. The practical recommendation: report scores against criteria and with their uncertainty, and use comparison data only as cautious context, never as a ranking.

What the research says

The distinction is borrowed from educational measurement. A norm-referenced interpretation locates a score relative to a reference population — percentile ranks, stanines, "above/below the mean." A criterion-referenced interpretation compares a score against an absolute standard of what counts as adequate, independent of how others did. The two answer different questions: norm-referencing asks "how does this instructor compare to others?"; criterion-referencing asks "did this teaching meet our standard?"

Most institutional course-evaluation reports are implicitly norm-referenced. They print the instructor''s mean alongside a department, faculty, or institutional mean, and readers — promotion committees, chairs, the instructors themselves — instinctively read the comparison as a verdict. The research on what happens next is sobering.

Boysen (2015), in Teaching of Psychology ("Uses and Misuses of Student Evaluations of Teaching"), and in related experimental work, showed that faculty and administrators over-interpret small mean differences. In one study (N = 121), faculty interpreted small differences between course means as meaningful even when confidence intervals and statistical tests indicated no reliable difference. In a second (N = 183), differences explicitly labelled non-significant still shifted perceptions of teaching ability. Reviewers, Boysen argued, fall for the "belief in the law of small numbers" — treating every numeric gap as a true difference. Norm-referenced reporting is the perfect vehicle for this error, because it foregrounds exactly the small instructor-versus-mean gaps that are least reliable.

Theall and Franklin (2001), in New Directions for Institutional Research ("Looking for Bias in All the Wrong Places"), argued that the misuse of comparison and normative data is itself a leading source of unfairness in ratings systems — more so than many of the demographic biases that get more attention. When comparison groups are small, non-equivalent (different class sizes, disciplines, course levels), or unstable year to year, ranking an instructor against them imports all of those confounds into what looks like an objective benchmark.

Two further strands reinforce the point. The Berkeley statisticians Stark and Freishtat (2014), in their open review An Evaluation of Course Evaluations, showed that comparing averages of small, skewed, ordinal rating distributions to a comparison mean is statistically indefensible and recommended reporting distributions rather than rank-ordered averages. And Linse (2017) documented that the committees doing the comparing frequently lack the statistical literacy to know that a 0.2 gap on a 5-point scale from twelve respondents is noise. The convergent conclusion across all four sources: the reference frame you choose changes the decision, and the norm-referenced frame is the one most likely to produce confident, unfair conclusions.

Why it matters for course evaluation in practice

The choice of reference frame is not a presentational detail — it determines who gets flagged, promoted, or remediated. Three consequences matter most:

  • Norm-referencing manufactures failure by construction. In any distribution, half the values sit below the median. If your report frames "below the department mean" as a problem, you have guaranteed that roughly half your faculty are "problems" every cycle, regardless of absolute quality. A department of uniformly excellent teachers still produces a bottom half. This is a mathematical artefact, not a finding about teaching.
  • It amplifies the small-difference fallacy. Because comparison reporting puts the instructor''s mean next to the group mean, it directs attention to precisely the gap that is least likely to be real (see our companion piece on why a 4.2 vs 4.4 difference is usually noise). It also interacts with regression to the mean: a course that ranked "below average" one year will usually "improve" the next even if nothing changed, generating false credit and false alarm.
  • It misclassifies instructors. Even when ratings are unbiased, ranking on noisy scores reliably misassigns people near the middle of the distribution (see why course-evaluation scores misclassify instructors and Esarey & Valdes on fair ranking). Criterion-referenced reporting sidesteps this by not asking the data to do the one thing — fine-grained ranking — it is least able to do.

Criterion-referenced reporting is not free of difficulty, but it is more honest about the question. "Did at least 80% of students agree the learning objectives were clear?" is a defensible, actionable standard. It can be failed by everyone or met by everyone, which is exactly the point: it measures teaching against a goal, not against colleagues. Statistical remedies such as empirical-Bayes shrinkage further temper unstable small-sample comparisons by pulling noisy course means toward the overall average.

Limitations and honest caveats

A rigorous reader should not over-correct into the opposite error:

  • Comparison data is not worthless. Some context is legitimate and useful: a course rated far below both the criterion and every comparison group, consistently, across years, is a real signal. The problem is not comparison per se but ranking on small, noisy gaps and treating the comparison mean as a pass/fail line.
  • Criteria are value judgements, and can be set badly. A criterion-referenced standard is only as good as the threshold chosen. Set it too low and everyone passes meaninglessly; set it arbitrarily high and you have recreated the unfairness from the other side. Thresholds need justification, ideally tied to learning outcomes rather than to a round number.
  • Boysen''s and Theall & Franklin''s evidence is largely North American. The over-interpretation finding is robust and theoretically general (it is a cognitive bias, not a local custom), but the institutional reporting practices that trigger it vary, and European QA frameworks differ in how they mandate comparison.
  • Distributions help but ask more of readers. Stark and Freishtat''s recommendation to report full distributions rather than averages is sound, but it demands more statistical literacy from committees than a single comparison number — and the same literacy gap Linse documents is what makes distributions get ignored. Better reporting only works if paired with reader guidance.
  • No reference frame fixes a bad instrument. If the underlying ratings are contaminated by the biases covered elsewhere in this documentation, neither norm- nor criterion-referencing rescues them. Reference frame is a reporting decision downstream of measurement quality, not a substitute for it.

How Koji incorporates this

Koji for Education is designed to report against standards and uncertainty rather than to manufacture rankings:

  • Criterion-style reporting by default. Koji reports how a course performed against defined, outcome-linked questions — what proportion of students agreed objectives were clear, felt appropriately challenged, or would change a specific element — rather than leading with the instructor''s rank against a comparison mean. The headline answers "did we meet the standard?", not "who lost?".
  • Uncertainty made visible. For any aggregate, Koji can surface the response count and the spread behind a mean, so a committee sees that a 4.2-from-eleven-responses figure carries wide uncertainty — directly countering the small-difference over-interpretation Boysen documents.
  • Distributions and verbatims over rank-ordered averages. Consistent with Stark and Freishtat, Koji foregrounds the distribution of responses and representative open-text quotes rather than a single league-table number, giving committees the shape of the data, not just its midpoint.
  • Bias-aware, cohort-triangulated context. Where comparison is genuinely useful, Koji frames it as cautious context across comparable cohorts and over time rather than a single-cycle ranking, and flags small-sample and skewed distributions instead of averaging over them.
  • Designed to inform judgement, not automate it. Koji explicitly positions evaluation output as support for human, standards-referenced review — never as an automated ranking for personnel decisions — consistent with the documented unfairness of norm-referenced misuse.

The same reporting philosophy underpins Koji''s core research platform at [koji.so], which reports product and customer-research findings against goals and with their uncertainty rather than as decontextualised rankings.

Related Resources

References

  • Boysen, G. A. (2015). Uses and Misuses of Student Evaluations of Teaching: The Interpretation of Differences in Teaching Evaluation Means Irrespective of Statistical Information. Teaching of Psychology, 42(2), 109–118. https://doi.org/10.1177/0098628315569922
  • Theall, M., & Franklin, J. (2001). Looking for Bias in All the Wrong Places: A Search for Truth or a Witch Hunt in Student Ratings of Instruction? New Directions for Institutional Research, 2001(109), 45–56. https://doi.org/10.1002/ir.2
  • Stark, P. B., & Freishtat, R. (2014). An Evaluation of Course Evaluations. ScienceOpen Research. https://doi.org/10.14293/S2199-1006.1.SOR-EDU.AOFRQA.v1
  • Linse, A. R. (2017). Interpreting and using student ratings data: Guidance for faculty serving as administrators and on evaluation committees. Studies in Educational Evaluation, 54, 94–106. https://doi.org/10.1016/j.stueduc.2016.12.004
  • Glaser, R. (1963). Instructional technology and the measurement of learning outcomes: Some questions. American Psychologist, 18(8), 519–521. https://doi.org/10.1037/h0049294

Related articles

analysis-reporting

Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved

Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

analysis-reporting

The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise

Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.

analysis-reporting

An Evaluation of Course Evaluations: What the Berkeley Statisticians Concluded About Reporting Student Ratings

Two Berkeley statisticians argued that averaging student-rating scores to rank instructors for promotion and tenure should be abandoned for both statistical and substantive reasons. What Stark & Freishtat (2014) actually recommend — and how to report ratings responsibly instead.