New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

Common-Language Effect Size: Reporting Course-Evaluation Differences People Actually Understand

A 0.2-point gap in mean ratings means nothing to a committee. The common-language effect size (probability of superiority) restates a difference as a probability anyone can interpret. What the research says and how to report course evaluations honestly.

Koji Education Team

Product

The short answer

Course-evaluation reports are full of differences no reader can interpret: Course A scored 4.3, Course B scored 4.1. Is that gap large, trivial, meaningful? A common-language effect size (CLES) — also called the probability of superiority — answers in plain terms: it is the probability that a randomly chosen student from one course rates it higher than a randomly chosen student from the other. Instead of "a 0.2-point difference," you report "if you pick one student from each course at random, there is a 55% chance the Course A student rates higher" — a statement a dean, a student rep, or an AI assistant can immediately grasp. For ordinal, skewed evaluation data it is also more robust than differences in means.

BLUF: The common-language effect size (McGraw & Wong, 1992), or probability of superiority, restates a group difference as the probability that a randomly sampled member of one group scores higher than one from another. It communicates course-evaluation differences far more intuitively than raw mean gaps or standardised effect sizes, and its non-parametric form (Vargha & Delaney''s A; Ruscio, 2008) is robust to the skew, ordinality, and outliers typical of rating data. Koji is designed to report differences in these interpretable, robust terms rather than leaning on fragile mean comparisons.

What the research says

Kenneth McGraw and S. P. Wong introduced the common-language effect size in Psychological Bulletin (1992) to solve a communication problem: standardised effect sizes like Cohen''s d are statistically sound but opaque to non-specialists. Their CLES converts a difference between two groups into "the probability that a score sampled at random from one distribution will be greater than a score sampled at random from the other." Their own example is memorable — a height difference between men and women that corresponds to a CLES of about 0.92, meaning that in a random man–woman pairing there is roughly a 92% chance the man is taller. The number requires no training to understand, yet it is a genuine effect-size measure, computable from means and variances (under a normality assumption) or from the raw data directly.

András Vargha and Harold Delaney generalised this in the Journal of Educational and Behavioral Statistics (2000). They showed McGraw and Wong''s formula was tied to normal-distribution and equal-variance assumptions, and proposed the A statistic (the "measure of stochastic superiority"), a fully non-parametric version computed directly from the ranks of the data. A is closely related to the area under the ROC curve and to the Mann–Whitney U statistic, and it handles ties — which matter enormously for Likert data, where ties are the rule. Crucially for evaluation work, A does not assume equal-interval scores, normality, or equal variances.

John Ruscio, in Psychological Methods (2008), stress-tested this probability-based measure and found it notably robust: it is insensitive to base rates and holds up well under skew, unequal variances, extreme scores, and non-linear (e.g. ordinal) transformations that badly distort d. He argues its combination of interpretability and robustness makes it an attractive default for reporting effects, and it works naturally with ordinal data — exactly the properties course-evaluation reporting needs. Together the three papers give both an intuitive statistic (McGraw & Wong) and a rigorous, assumption-light foundation for using it on real, messy rating data (Vargha & Delaney; Ruscio).

Why it matters for course evaluation in practice

Evaluation audiences are non-statisticians making consequential judgements. A programme committee deciding whether Course A''s "higher" rating is meaningful, a student-partnership board comparing seminar streams, or a dean reviewing a department cannot act sensibly on a third-decimal-place difference in means, and most cannot interpret a Cohen''s d of 0.18 either. The probability of superiority gives them a sentence with a clear operational meaning and an implicit sense of scale: 50% is no difference at all, 55% is small, 70% is substantial. It reframes the question from "is this difference statistically significant?" (which large evaluation samples make trivially easy to achieve) to "how much does being in one course actually change the rating you''d give?" — the question decisions should turn on.

It also protects against two common reporting failures. First, overstating tiny gaps: significance tests on big cohorts flag differences that a CLES immediately exposes as near-51% coin-flips. Second, misusing means on skewed data: because evaluation distributions pile up at the ceiling, mean differences and d can mislead, whereas the rank-based A statistic reports the honest probability that one group out-rates the other regardless of distribution shape. Reporting a CLES alongside a confidence interval turns a fragile point estimate into a statement about both size and uncertainty — the combination responsible reporting requires.

Limitations and honest caveats

The CLES is a communication and robustness tool, not a cure for weak evidence. First, an intuitive number is still only as good as the data behind it: a clear 58% probability of superiority built on a biased or low-response sample is confidently wrong, and the statistic does nothing to fix validity or non-response problems. Second, the parametric McGraw–Wong formula inherits normality and equal-variance assumptions, so for skewed Likert data the non-parametric A (Vargha & Delaney) should be preferred; reporting the parametric version on ceilinged data reintroduces the very problem it is meant to avoid. Third, probability of superiority ignores the magnitude of individual differences — it counts who is higher, not by how much — so a large but rare gap and a small but consistent one can yield similar values; it complements rather than replaces distributional reporting. Fourth, like any single number it can be gamed or over-trusted; it should sit beside the raw distribution, a confidence interval, and qualitative context, not replace them. Finally, thresholds for "small/medium/large" (roughly .56/.64/.71 in Vargha and Delaney''s guidance) are conventions, not natural law, and should be used as rough signposts rather than decision rules. The defensible claim is specific: for communicating and robustly quantifying a difference in ordinal evaluation data, the probability of superiority beats a bare mean gap — not that it settles what the difference means.

How Koji incorporates this

Koji is designed to report differences in terms decision-makers can actually use, and to prefer robust, distribution-appropriate statistics over fragile mean comparisons.

  • Plain-language difference reporting. When Koji compares courses, cohorts, sections, or time points, it can express the difference as a probability of superiority — "a randomly chosen respondent from this cohort rates the course higher X% of the time" — turning an abstract gap into a statement stakeholders understand without training.
  • Robust, rank-based computation. Because evaluation data is ordinal and ceiling-heavy, Koji''s comparison logic favours rank-based, distribution-free summaries of the kind Vargha and Delaney and Ruscio validate, rather than defaulting to mean differences that skew distorts.
  • Uncertainty shown, not hidden. Differences are presented with their uncertainty (confidence intervals / ranges) so a 51% "advantage" is visibly indistinguishable from noise, directly countering the over-reading of trivial gaps that significance-only reporting invites.
  • Distribution alongside the summary. Koji keeps the full response distribution visible next to any effect-size statement, so reviewers see both the intuitive probability and the shape it summarises — honouring the caveat that a single number should never stand alone.
  • Consistent, comparable cross-course statistics. Structured question types produce comparable scales across courses, letting the same interpretable effect-size framework apply programme-wide rather than course-by-course.

Framed honestly, the probability of superiority does not make a weak dataset trustworthy — it is designed to communicate real differences honestly and robustly. Koji''s core research platform at koji.so uses the same interpretable-difference reporting for product and customer research, where "version B scored 0.3 higher" is just as meaningless to a stakeholder as it is to a dean.

Frequently asked questions

What is the common-language effect size? It is a way of expressing a difference between two groups as the probability that a randomly chosen member of one group scores higher than a randomly chosen member of the other (McGraw & Wong, 1992). A value of 50% means no difference; higher values mean a stronger tendency for one group to score above the other.

How is it different from Cohen''s d or a difference in means? Cohen''s d is a standardised difference that non-specialists rarely interpret intuitively, and a raw mean gap has no built-in sense of scale. The common-language effect size is a probability anyone can understand, and its non-parametric form is robust to skew and ordinality that distort d.

Why is it well suited to course-evaluation data? Evaluation ratings are ordinal, skewed, and ceilinged. The rank-based A statistic (Vargha & Delaney, 2000; Ruscio, 2008) makes no assumption of normality, equal variance, or equal spacing, so it reports the honest probability that one group out-rates another regardless of distribution shape.

Does it fix bias or low response rates? No. It is a reporting and robustness tool. A clear probability built on a biased or low-response sample is still misleading; validity and non-response must be addressed separately.

Should it replace reporting the distribution and confidence intervals? No. It should sit beside the raw distribution and an uncertainty estimate. The probability of superiority summarises who scores higher, not by how much, so it complements distributional reporting rather than replacing it.

Related resources

References

Related articles

analysis-reporting

Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved

Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

analysis-reporting

The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise

Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.

analysis-reporting

Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores

Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.