New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Does the Professor Students Rate Highest Teach Them Most? Value-Added Evidence vs Student Ratings

The strongest causal evidence — random-assignment studies that track follow-on course performance — shows the instructors students rate highest are often not the ones whose students learn most over time. What Carrell & West (2010) and Braga et al. (2014) found, and what it means for course evaluation.

Koji Education Team

Product

In brief: The most rigorous causal evidence on student evaluations comes from studies that randomly assign students to instructors and then track how those same students perform in later, more advanced courses. Carrell & West (2010) and Braga, Paccagnella & Pellizzari (2014) both found that student evaluations positively predict current-course achievement but are zero-to-negatively correlated with follow-on achievement. The instructors who earn the highest ratings are frequently those who teach to the immediate test rather than those who build durable understanding. The lesson for quality assurance is not to discard student feedback, but to stop treating a high average score as proof of deep learning.

What the research says

For decades the central validity question about Student Evaluation of Teaching (SET) has been: do higher-rated instructors actually produce more learning? Most studies answer this with a correlation between end-of-course ratings and end-of-course exam grades. Two natural experiments turned that question on its head by exploiting random assignment and follow-on course performance, and their findings are among the most uncomfortable in the literature.

The anchor study: Carrell & West (2010)

Scott Carrell and James West studied the United States Air Force Academy, an unusually clean research setting. Students there are randomly assigned to course sections, take a common standardized exam, and are required to take a fixed sequence of follow-on courses (e.g. Calculus I then Calculus II, then engineering courses that depend on the maths). This design removes the selection problems that plague ordinary SET research: students cannot choose easier professors, and every section faces the same assessment.

Carrell and West found two things. First, instructors who raised contemporaneous achievement — their students did better on the common exam for that course — received higher student evaluations. Second, those same instructors were associated with lower follow-on achievement: their students did worse in the next course in the sequence. Student evaluations were a positive predictor of current-course grades but a poor (even negative) predictor of deep, transferable learning. Strikingly, observable markers of expertise — academic rank, teaching experience, and holding a terminal degree — were negatively related to contemporaneous achievement and student ratings, but positively related to follow-on achievement. The paper appeared in the Journal of Political Economy (Carrell & West, 2010).

Corroboration: Braga, Paccagnella & Pellizzari (2014)

A near-replication in a completely different setting reached the same conclusion. Using administrative data from Bocconi University in Italy, where students are again effectively randomly assigned to teaching groups, Braga, Paccagnella and Pellizzari estimated teacher effectiveness from how students later performed in follow-on subjects. Teacher effectiveness was negatively correlated with student evaluations: the teachers who improved subsequent performance tended to receive worse evaluations. The authors offered a clean theoretical mechanism — teachers can either do "real teaching", which demands more student effort, or "teaching to the test", which secures high current grades without building future capability. If students are short-sighted and reward the instructor who maximises their immediate, low-effort utility, good teachers will rationally receive bad ratings. The study was published in Economics of Education Review (Braga, Paccagnella & Pellizzari, 2014).

The synthesis: Kornell & Hausman (2016)

Reviewing this literature, Kornell and Hausman framed the apparent contradiction precisely: how you measure learning determines the answer. When learning is measured by an end-of-course test, highly-rated teachers look most effective. When learning is measured by performance in subsequent related courses, teachers with lower ratings often look most effective. Their explanation draws on the cognitive-psychology distinction between performance and learning (Soderstrom & Bjork): "desirable difficulties" — effortful retrieval, spacing, productive struggle — depress short-term fluency and momentary satisfaction while improving long-term retention and transfer. A lecturer who makes everything feel easy generates high ratings and high immediate confidence, but not necessarily durable learning (Kornell & Hausman, 2016).

Why it matters for course evaluation in practice

The practical implication is not that student evaluations are worthless — it is that the construct they best capture is student experience and immediate satisfaction, which is correlated with, but not identical to, durable learning. For a quality-assurance office, three consequences follow:

  1. Do not use a single high average as evidence of learning. A 4.6/5 mean tells you students enjoyed the course and likely felt they learned. It does not tell you they will retain or transfer the material. Treating the rating as a proxy for educational quality risks rewarding exactly the teaching-to-the-test behaviour Braga et al. modelled.
  2. Beware perverse incentives. If promotion, contract renewal, or league-table position hinges on SET means, rational instructors will optimise for the rating — lighten workload, narrow the syllabus to the exam, inflate grades. The course difficulty and grading-leniency literatures document this drift directly.
  3. Triangulate learning with independent evidence. Follow-on performance, capstone and exit assessments, programme-level learning-outcome attainment, and peer observation measure something the survey cannot. SET should be one strand of evidence, weighted for what it actually measures.

Limitations and honest caveats

A critical reader should hold several caveats firmly:

  • Generalisability of the settings. The Air Force Academy is highly structured — mandatory sequences, standardised exams, a selected student body, military discipline. Bocconi is an elite Italian business university. Neither resembles a typical mass-intake European programme, and the negative correlation may be weaker, absent, or differently shaped elsewhere.
  • Effect sizes are modest, not catastrophic. The headline is a reversal of sign between contemporaneous and follow-on outcomes, but the magnitudes are not enormous. These studies show SET is a poor proxy for deep learning, not that high-rated teachers are actively harmful in every case.
  • "Follow-on performance" is itself imperfect. It is confounded by the next instructor, by student motivation, and by curriculum coupling. It is a better proxy for transfer than an end-of-course exam, but it is not a clean measure of learning either.
  • The contemporaneous-validity tradition disagrees in emphasis. Cohen (1981) and Feldman (1989), synthesising multisection validity studies, reported moderate positive correlations (around r = .43–.47) between ratings and same-course achievement. The newer studies do not overturn that; they add a crucial qualifier — the positive correlation lives in the short run and can vanish or reverse in the long run. Both can be true.
  • Mechanism is inferred, not proven. The "teaching to the test" model fits the data elegantly but is one explanation among several (e.g. differential difficulty, regression artefacts).

Acknowledging these limits is what separates a defensible interpretation from an anti-SET polemic. The honest claim is narrow and robust: a high average student rating is not, by itself, evidence of durable learning.

How Koji incorporates this

Koji is built on the premise that a Likert mean is the start of an evaluation, not the end of it. Several mechanisms are designed to mitigate exactly the gap these studies expose — though none can eliminate the underlying short-sightedness of in-the-moment student judgement.

  • AI-moderated conversational interviews that probe beyond the number. Where a static form records "Workload — 2/5", Koji's interviewer follows up: what felt heavy, whether the effort was productive or busywork, whether the difficulty mapped to learning. This helps distinguish a course that is hard and effective from one that is hard and badly designed — a distinction a single satisfaction score collapses.
  • Structured question types that separate experience from learning. Using scale, single_choice, open_ended, ranking, and yes_no items, an instrument can ask separately about enjoyment, perceived learning, effort, and confidence — rather than letting a global "how good was this course" item absorb all of them under a halo.
  • Triangulation across cohorts and time. Koji supports mid-cycle and formative collection and comparison across cohorts, so a programme can watch whether a popular course actually feeds success in downstream modules, rather than judging it on end-of-term sentiment alone.
  • Bias-aware reporting. Koji's reporting is designed to surface distributions and qualitative themes, not just a mean, so committees are less tempted to read a high average as proof of educational quality.
  • Closing-the-loop action tracking keeps the focus on whether changes improved outcomes, not just whether next year's satisfaction ticked up.

Koji is designed to mitigate — not eliminate — the performance-versus-learning gap by enriching the signal around the score and encouraging triangulation. The core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where the identical lesson holds: a satisfaction number is not the same as a true measure of value.

Related Resources

References