New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Why "I Learned a Lot" Can't Be Compared Across Courses: Reference Bias

Self-rated learning items are the backbone of most course evaluations, yet reference bias means students judge themselves against different implicit standards. We review the evidence that this distorts cross-group comparisons and what it means for benchmarking courses and programmes.

Koji Education Team

Product

In brief: When students answer "I learned a great deal in this course," they grade themselves against an implicit standard set by their peers and their own expectations — and that standard differs across courses, cohorts, and cultures. This reference bias means two equally effective courses can produce different self-rated-learning scores simply because students hold different yardsticks. Large studies (N over 200,000) show reference bias can even reverse the direction of group-level associations. The practical lesson: never benchmark or rank courses on self-rated learning alone, and pair the item with anchored, behavioural evidence.

The question this answers

Almost every course-evaluation form contains an item like "I learned a lot in this course" or "This course improved my understanding of the subject." Administrators treat these self-rated-learning items as the closest thing to an outcome measure — a student-reported proxy for whether teaching worked. But there is a measurement problem hiding in plain sight: a Likert rating of one's own learning is only meaningful relative to a reference point, and students do not all use the same one. This article explains reference bias, reviews the strongest evidence for it, and draws out what it means for comparing courses and programmes.

What the research says

The anchor is a large, recent study: Lira et al. (2022), Large Studies Reveal How Reference Bias Limits Policy Applications of Self-Report Measures, Scientific Reports, 12, 19189. Reference bias is defined as systematic error arising from differences in the implicit standards by which individuals evaluate the same behaviour. Across three studies with 229,685 adolescents, the authors found that students whose peers performed better academically rated themselves lower in self-regulation, and held higher standards for what counts as self-regulated behaviour. Critically:

  • The effect appeared in self-report measures but not in objective task measures of the same construct — pinpointing the bias to the act of self-rating against a shifting yardstick.
  • It produced paradoxical group-level predictions: within a school, higher self-reported self-regulation predicted later college persistence, but across schools, higher average self-regulation predicted lower persistence — a sign reversal driven by reference bias, not by the underlying trait.
  • The distorting reference point was set by near-peers (immediate classmates), not distant comparison groups, which is exactly the structure of a course cohort.

The foundational demonstration in education is West et al. (2016), Promise and Paradox: Measuring Students' Non-Cognitive Skills and the Impact of Schooling (Educational Evaluation and Policy Analysis, 38(1), 148–170). Studying Boston charter schools, West and colleagues found that students at schools that raised achievement most reported lower conscientiousness, self-control, and grit — the opposite of what the schools' effectiveness would predict — because demanding environments raised students' standards for what counts as "working hard." Self-reports were biased by the local reference frame.

A third strand connects this to survey methodology more broadly. Heine et al. (2002), What's Wrong With Cross-Cultural Comparisons of Subjective Likert Scales? (Journal of Personality and Social Psychology, 82(6), 903–918), showed that the "reference-group effect" undermines cross-cultural Likert comparisons: groups anchor their ratings to local norms, so the same number means different things in different populations. The mechanism is identical to what afflicts "I learned a lot" across disciplines and countries.

Why it matters for course evaluation in practice

Reference bias attacks the single use administrators most want from evaluations: comparison. Three implications:

  1. Cross-course benchmarking on self-rated learning is unsafe. A rigorous course that stretches students may score lower on "I learned a lot" than an easy one — not because students learned less, but because the demanding course raised their standard for what "a lot" means. Ranking courses or programmes on this item can invert the truth, exactly as Lira et al. and West et al. document.
  2. Cross-cultural and cross-discipline aggregation is fragile. In a European institution with international cohorts and multiple languages, the reference frame varies systematically. Pooling self-rated-learning scores across a multinational programme risks comparing yardsticks, not learning. This compounds the response-style differences already known to affect Likert data.
  3. Within-course, longitudinal use is safer. Reference bias is most damaging for between-group comparison. Tracking the same course over time, or reading self-rated learning alongside what students say they can now do, keeps the reference frame roughly constant and the signal more trustworthy.

In short: self-rated learning is a legitimate formative signal and a dangerous comparative metric. It tells you something about a course on its own terms; it lies when you line courses up against each other.

Limitations and honest caveats

The argument needs its own caveats, and a careful reader should weigh them:

  • Reference bias is not total invalidity. Lira et al. are explicit that reference-biased measures can retain strong validity within a group and at the individual level. "I learned a lot" is not noise; it is a signal whose comparability across groups is compromised. Discarding it entirely would over-correct.
  • The flagship evidence is about non-cognitive skills, not course evaluation specifically. The construct studied is self-regulation/conscientiousness, transferred by analogy to self-rated learning. The analogy is strong — both are self-reports judged against a peer-set standard — but it is an inference, not a direct course-evaluation experiment. Treat it as a well-supported mechanism rather than a measured course-evaluation effect size.
  • Magnitude varies. The dramatic sign reversals occur at the group/aggregate level with large samples; for a single instructor reading their own course's trend, the bias may be modest. The risk scales with how aggregated and how cross-group the comparison is.
  • Mitigations are imperfect. Anchoring vignettes, behaviourally specific items, and objective measures reduce but do not eliminate reference bias, and each adds respondent burden.

The honest conclusion: reference bias is a real, well-evidenced threat to comparing self-rated learning across courses, cohorts, and cultures — not a reason to delete the item, but a strong reason to stop treating it as a clean outcome metric.

How Koji incorporates this

Koji for Education is built around the recognition that a single self-rated-learning number is hard to compare across groups — so it captures the content of learning, not just a rating against an invisible yardstick.

  • AI-moderated conversational interviews that elicit anchored, behavioural evidence. Instead of stopping at "I learned a lot (4/5)," Koji probes what specifically a student can now do, which activities produced that, and where understanding broke down. A concrete answer — "I can now derive the model and apply it to a new dataset" — is far less reference-dependent than a Likert self-rating, because it points to behaviour rather than to a self-comparison.
  • Structured question types that separate construct from self-comparison. Koji's open_ended, scale, single_choice, multiple_choice, ranking, and yes_no items let an evaluation pair a self-rated-learning scale with behaviourally specific items, so reporting never rests on the biased item alone.
  • Automatic thematic analysis at programme level. By clustering open-text evidence of what students learned, Koji gives committees a comparison basis grounded in described outcomes rather than in raw self-ratings that may encode different standards.
  • Within-course trend reporting and bias-aware comparison. Koji is designed to favour within-course longitudinal views — where the reference frame is roughly stable — and to flag cross-cohort and cross-language comparisons as the place where reference bias is most likely, discouraging naive league tables of self-rated learning.
  • Triangulation by design. Koji positions self-rated learning as one strand to be triangulated with described capabilities, assessment evidence, and programme outcomes, rather than as a stand-alone measure of teaching effectiveness.

Koji is designed to mitigate reference bias by shifting weight from yardstick-dependent self-ratings to anchored, described evidence; it cannot eliminate a bias that is intrinsic to self-report. The same conversational engine powers general user and customer research at koji.so, where the reference-group effect is an equally well-known hazard in cross-market satisfaction benchmarking.

The practical takeaway for a quality office

Reference bias does not demand a wholesale redesign — it demands a change in how one item is read. Keep "I learned a great deal" on the form, because it carries useful within-course information and students expect to be asked. But annotate it in every report with a standing caveat: this number is comparable down a column (the same course over time) and treacherous across a row (different courses, cohorts, or languages side by side). Pair it routinely with at least one behaviourally anchored item and with open-text evidence of described capability, and require that any cross-course or cross-programme claim rest on those anchored sources rather than on the self-rating. A committee that internalises "self-rated learning is a within-group thermometer, not a between-group ruler" has absorbed the entire practical lesson of the reference-bias literature.

Related resources

References

  • Lira, B., O'Brien, J. M., Peña, P. A., Galla, B. M., D'Mello, S., Yeager, D. S., Defnet, A., Kautz, T., Munkacsy, K., & Duckworth, A. L. (2022). Large studies reveal how reference bias limits policy applications of self-report measures. Scientific Reports, 12, 19189. https://doi.org/10.1038/s41598-022-23373-9
  • West, M. R., Kraft, M. A., Finn, A. S., Martin, R. E., Duckworth, A. L., Gabrieli, C. F. O., & Gabrieli, J. D. E. (2016). Promise and paradox: Measuring students' non-cognitive skills and the impact of schooling. Educational Evaluation and Policy Analysis, 38(1), 148–170. https://doi.org/10.3102/0162373715597298
  • Heine, S. J., Lehman, D. R., Peng, K., & Greenholtz, J. (2002). What's wrong with cross-cultural comparisons of subjective Likert scales? The reference-group effect. Journal of Personality and Social Psychology, 82(6), 903–918. https://doi.org/10.1037/0022-3514.82.6.903

Related articles

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

research-methods

Why Students Click Straight Down the Middle: Satisficing in Course Evaluations

A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.

research-methods

What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback

Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.

research-methods

Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers

Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.