New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods11 min read

Selection Bias in Course Evaluations: What Goos and Salomons Found

A research-grounded reading of Goos & Salomons (2017) on selection bias in online course evaluations, with practical implications for response-rate policy and reporting at European universities.

Koji Education Team

Product

Answer in brief

The best European evidence on selection bias in course evaluations — Goos and Salomons (2017), based on more than 3,000 courses at a large European university — finds that course evaluation scores are systematically upward-biased: the students who choose to respond are more positive than the average enrolled student. The total bias they estimate is around 28% of a standard deviation of the score, large enough to materially change course rankings. Quality-assurance teams that compare or rank courses on raw evaluation means should therefore not interpret small differences as real differences, and should publish response rates alongside scores so that consumers of the data can adjust.

What the research says

The anchor paper for this brief is Goos, M., & Salomons, A. (2017). "Measuring teaching quality in higher education: assessing selection bias in course evaluations." Research in Higher Education, 58(4), 341–364. Goos and Salomons examine over 3,000 courses taught at a large European university and decompose selection bias in the resulting student evaluations into two components: bias from observed student characteristics (e.g. expected grade, programme, prior performance) and bias from unobserved characteristics that drive both response and rating.

Using the Heckman-style correction familiar from labour economics, they estimate that the total selection bias in average course evaluations is around 0.28 standard deviations of the score distribution. Critically, this is not a uniform shift — correcting for selection bias changes the relative ranking of courses, not just the absolute level. Some courses move up substantially in the corrected ranking, others move down, because the type of student who chooses to respond varies across courses (e.g. students with higher expected grades are more likely to respond in some courses, less likely in others). This means that policies that compare courses on raw mean scores are operating on biased data, and the bias is not constant.

Goos and Salomons' finding sits inside a broader literature on response and non-response in course evaluations.

Adams and Umbach (2012), "Nonresponse and online student evaluations of teaching: Understanding the influence of salience, fatigue, and academic environments," in Research in Higher Education (53(5), 576–591), analysed roughly 135,000 evaluations across 22,000 undergraduates. Response was predicted by salience (the student's investment in the course), fatigue (number of evaluations the student had been asked to complete), and academic environment. Non-respondents were systematically different from respondents — non-random non-response is the technical condition under which sample means are biased estimates of the population mean.

The earlier Spooren and Van Loon (2012) work and the comprehensive review by Spooren, Brockx, and Mortelmans (2013), "On the validity of student evaluation of teaching: the state of the art," Review of Educational Research (83(4), 598–642), document that the shift from in-class paper administration to online administration has reduced response rates from typical highs of 70–90% to typical lows of 20–40%, increasing exposure to selection bias of the sort Goos and Salomons quantify. A related finding from the same literature is that very low response rates (below ~50%) make the bias non-trivial, and below ~30% make individual-course score interpretation unreliable for any high-stakes purpose.

There is, importantly, no consensus that every online evaluation system is materially biased — some studies of specific institutions find smaller effects. The Goos and Salomons result is best read as an empirically grounded upper-bound estimate of the kind of bias that is plausible at a large European university with ordinary online response rates, not as a universal constant.

Why it matters for course evaluation in practice

Four operational implications follow for European QA teams.

Response rate is a first-class signal. Score without response rate is uninterpretable. A 4.3/5 from a 70% response is a different artefact than a 4.3/5 from a 22% response, even if both round to the same point. Dashboards, programme reports, and any external reporting (NVAO, QAA, ENQA-aligned reports) should display response rate alongside score, ideally with confidence intervals and an explicit warning when the response rate falls below the threshold at which selection bias becomes large.

Ranking is unsafe in the low-response regime. Because the Goos–Salomons correction changes relative rankings, comparing courses on raw means in the low-response regime can produce ordered lists that are not robust. Programme committees that act on ranked lists (e.g. "the bottom five modules trigger review") risk acting on noise plus selection bias rather than on real signal.

Effort to lift response rates is research-supported. Reminders that target salience, reductions in survey fatigue (fewer items, fewer parallel surveys), and small commitment devices (course time set aside for evaluation, with privacy preserved) all have empirical support for increasing response rates without distorting who responds. Importantly, paying students or grading on completion is not a defensible mitigation; it changes the meaning of the response.

Mid-cycle or multi-touch designs reduce the selection problem. When the same student is invited to give feedback at multiple points across the course, the selection bias on any single touch is partially offset by participation across touches. This is one mechanism by which mid-semester conversational feedback (see Mid-Semester Feedback and the Power of Consultation) reduces vulnerability to end-of-term selection bias.

Under the European Standards and Guidelines (ESG 2015) standard 1.3 (student-centred learning, teaching and assessment), and standard 1.7 (information management), institutions are expected to collect and report data in ways that genuinely support decisions. A QA system that reports only the score, without the response rate and the population it was drawn from, fails this expectation.

Limitations and honest caveats

Goos and Salomons (2017) is a single-institution study, even if it is a large one. The 0.28 SD bias estimate should not be transported verbatim to other institutions; the size of the bias depends on the local response-rate regime, on which student characteristics drive response, and on the structure of the evaluation instrument. Replications at other European universities are limited and methodologically heterogeneous.

The Heckman-style correction is sensitive to the choice of exclusion restrictions — that is, variables that affect response but not the rating. If those restrictions are wrong, the bias estimate is off. Goos and Salomons argue carefully for their identifying assumptions, but the assumption is non-trivial.

More broadly, "bias" here is a statistical concept (the sample mean differs from the population mean). It does not, by itself, tell you which courses are over- or under-rated in absolute terms or which students are over- or under-served. Pairing selection-bias adjustment with qualitative evidence is essential for any judgement about teaching or course design.

Finally, the literature is clear that low response rates do produce material bias, but it is not the case that every high-response-rate course is bias-free. Other biases (gender, course difficulty, grading leniency) operate alongside selection bias and can be present even at high response. This article is about one component of total measurement error, not the whole story.

How Koji incorporates this

Koji for Education (edu.koji.so) treats selection bias as a first-order design problem, not an afterthought. Four mechanisms address it concretely.

Response-rate–aware reporting. Every Koji course report displays the response rate prominently next to the result, along with a population-margin indicator that flags when the response rate is in the regime where selection bias is likely large. Programme committees reading the report see, at a glance, when a score should be treated as suggestive rather than decisive.

Confidence-bounded summaries. Quantitative summaries are presented with uncertainty ranges (not just point estimates), so that a 4.3 with 22% response is not visually equivalent to a 4.3 with 75% response. This is designed to mitigate the human tendency to compare point estimates as if they were equally precise.

Conversational depth over single-touch scalars. Koji's AI-moderated interviews lower the burden of responding (a focused conversation is shorter than a 30-item form) and produce qualitative evidence that does not collapse to a scalar mean. The qualitative evidence is less vulnerable to the direction of selection bias because a small but engaged sample of substantive open responses is informative about the issues raised, even if the sample is non-representative.

Mid-cycle and multi-touch defaults. Koji is designed for short, formative check-ins through the term, not only a year-end pulse. Distributing evaluation moments across the term reduces dependence on any single end-of-term sample and partially offsets end-of-term selection. The same engine powers Koji's core research platform at koji.so for product and customer research where response selection is also a known problem.

The central design choice is not to deny selection bias — no system can eliminate it — but to make it visible and to design around it: prominent response rate, uncertainty intervals, qualitative depth, and multi-touch collection.

Related resources

References

  • Goos, M., & Salomons, A. (2017). Measuring teaching quality in higher education: Assessing selection bias in course evaluations. Research in Higher Education, 58(4), 341–364. https://doi.org/10.1007/s11162-016-9429-8
  • Adams, M. J. D., & Umbach, P. D. (2012). Nonresponse and online student evaluations of teaching: Understanding the influence of salience, fatigue, and academic environments. Research in Higher Education, 53(5), 576–591. https://doi.org/10.1007/s11162-011-9240-5
  • Spooren, P., Brockx, B., & Mortelmans, D. (2013). On the validity of student evaluation of teaching: The state of the art. Review of Educational Research, 83(4), 598–642. https://doi.org/10.3102/0034654313496870
  • Esarey, J., & Valdes, N. (2020). Unbiased, reliable, and valid student evaluations can still be unfair. Assessment & Evaluation in Higher Education, 45(8), 1106–1120. https://doi.org/10.1080/02602938.2020.1724875