The Big-Fish-Little-Pond Problem: Why Course Ratings Depend on the Company Students Keep
A "4 out of 5" is not an absolute judgement — it is a comparison against a local reference frame the student never states. Decades of research on the Big-Fish-Little-Pond Effect explain why raw evaluation means are not safe to compare across cohorts, departments or institutions.
Koji Education Team
Product ·
Short answer: When a student rates a course, they are not reporting an absolute measurement. They are answering "compared to what?" — and the comparison frame is their local peer group, their other modules, and their expectations, none of which appear in your data. Educational psychology has documented this for forty years under the name the Big-Fish-Little-Pond Effect (BFLPE), and its lesson for course evaluation is uncomfortable: a raw rating is frame-dependent, so comparing raw means across cohorts, departments or institutions can be invalid without evidence that the frames are equivalent. Here is what the research shows, where the analogy holds, and where it doesn't.
The effect that started it all
In 1984, Herbert Marsh and John Parker asked a deceptively simple question in the title of their paper: is it better to be a relatively large fish in a small pond, even if you don't learn to swim as well? (Marsh & Parker, Journal of Personality and Social Psychology, 1984). Their finding, developed into a formal frame-of-reference model by Marsh (1987), was that equally able students hold lower academic self-concept when surrounded by higher-achieving peers. Your sense of how good you are is not read off your absolute ability; it is computed against the pond you happen to swim in.
The effect is not a small-sample curiosity. Marsh and Hau tested it across 103,558 fifteen-year-olds in 26 countries using PISA 2000 data. Individual achievement raised academic self-concept, as you would expect — but school-average achievement lowered it, and the negative effect appeared in all 26 countries (mean standardised coefficient −.20). A later meta-analysis of 33 studies covering more than 1.2 million students put the pooled effect at −0.28: move a student into a class one standard deviation more able, and their academic self-concept falls by roughly a quarter of a standard deviation, with no change in their actual ability.
Why this matters for a survey you thought was objective
The BFLPE is about self-concept, but it is an instance of a much more general phenomenon that reaches straight into your evaluation data: the reference-group effect. People answer subjective rating scales relative to a comparison standard they supply themselves, and that standard varies between groups.
Steven Heine and colleagues demonstrated this vividly. Although experts agree East Asians are more collectivistic than North Americans, direct Likert self-reports failed to show the difference — until the researchers manipulated the reference group, at which point the expected pattern reappeared (Heine, Lehman, Peng & Greenholtz, JPSP, 2002). Their conclusion has teeth for anyone who benchmarks: cross-group comparisons of raw Likert scores can be actively misleading, because a "4" is anchored to a different internal standard in each group.
Now map that onto course evaluation. A student in a demanding, high-achieving cohort compares this module against a field of strong modules; a student in a weaker programme compares against a weaker field. Two genuinely identical teaching experiences can receive different ratings purely because the reference frames differ. The number on your dashboard is a comparison, not a constant.
The direct evidence in student evaluations of teaching
Is there evidence the reference frame moves teaching ratings specifically, not just self-concept? Yes. Clark Nowell's work on relative grade expectations is the cleanest example. Examining student evaluations of teaching, Nowell (2007) found that what predicts higher ratings is not a student's absolute expected grade but their expected grade relative to their own history — as students expect to do better than they usually do, they reward the instructor. A follow-up study estimating the causal effect confirmed that relative, not absolute, expected grade drives evaluations (Nowell, Gale & Kerkvliet, Economics of Education Review, 2012). The comparison frame is doing the work.
The measurement consequence: you cannot compare what is not invariant
This connects to a formal statistical requirement that is routinely ignored in quality offices. To compare mean ratings across groups and conclude something about teaching, the instrument must be measurement invariant across those groups — students in different cohorts must interpret the scale the same way. When researchers actually test this, invariance frequently fails. A six-country study of student perceptions of teaching (Maulana et al., Frontiers in Psychology, 2020) used multi-group confirmatory factor analysis precisely because you cannot compare group means without first establishing that the underlying measurement holds across them. Skip that step — as almost every league table of departmental averages does — and you may be ranking reference frames, not teaching.
There is a methodological fix. Anchoring vignettes, developed by Gary King and colleagues, ask every respondent to rate the same hypothetical scenarios, exposing how differently groups use identical scale points and allowing you to correct for it (King, Murray, Salomon & Tandon, American Political Science Review, 2003). The point is that comparability is something you have to establish, not assume — which is exactly why comparing a 4.1 in one department to a 4.4 in another is a trap.
But isn't this just an analogy? The strongest counterargument
An honest reader should push back, and the objection is a good one. The BFLPE is about a student's self-concept — an internal, self-referent judgement. Rating an instructor is an external, other-referent judgement. The mechanism (social comparison to the local pond) is firmly established for self-concept, but its transfer to how students rate teaching is an analogy, not a demonstrated pathway. The one strand of directly relevant SET evidence — the relative-grade-expectation studies — shows framing effects operating through grades, not through peer ability per se. And even the BFLPE meta-analytic effect, at −0.28, is moderate, not overwhelming.
A defender of student evaluations would add three fair points. First, within a single course and cohort the reference frame is largely shared, so many legitimate uses — same-instructor comparisons over time, pre/post designs, within-cohort item comparisons — are far less affected. Second, the reference-frame problem bites hardest on exactly the comparison good quality assurance already treats with suspicion: ranking different departments or institutions against each other on raw means. Third, the remedy is methodological — test for invariance, use anchoring, benchmark only within comparable frames — not abandoning evaluation.
We agree. The reference-frame critique is decisive against naïve cross-group mean-ranking and weak as a blanket dismissal of student feedback. That is precisely the distinction most dashboards fail to make. It sits alongside related distortions we have covered — central-tendency compression and contrast effects from the preceding course — all of which share a root: the number hides the comparison that produced it.
Surfacing the reference frame instead of averaging over it
If the problem is a hidden "compared to what?", the answer is an instrument that asks it. A static Likert grid can only record the number; it cannot recover the frame the student used to generate it. That is the structural limitation.
Koji for Education takes a different route. Its AI-moderated conversational interviews probe the comparison directly — following a rating with "compared to your other modules, where does this sit, and why?" — so the reference point becomes visible data rather than invisible noise. Standardised, bias-aware AI moderation means every student is asked in a consistent way, removing the human-moderator inconsistency that would otherwise add another frame. Automatic thematic analysis then surfaces whether a "4" in one cohort means the same thing as a "4" in another, and programme- and institution-level reporting is framed to flag when cross-group comparisons rest on non-equivalent scales rather than presenting a spurious league table. Koji does not eliminate reference-frame effects — no instrument can — but it makes them legible and mitigates naïve comparison, which is the honest goal. The same conversational engine powers general user and customer research at koji.so, where the "compared to what?" problem is just as pervasive.
The next time two departments post different averages, ask the question the scale never did: compared to what? Until you can answer it, the gap between them may be a gap between ponds, not between teaching.