The Follow-On Course Test: When Randomised Evidence Turns Course Evaluations Upside Down
Two studies that randomly assigned students to professors - one American, one Italian - found the same unsettling result: the instructors whose students did best in later courses were rated worst by those students. Here is what that does and does not prove about course evaluation.
Koji Education Team
Product ยท
Short answer: In two rare settings where students were randomly assigned to instructors and then tracked into compulsory follow-on courses, teaching effectiveness measured by later achievement was negatively correlated with student evaluation scores. This is the single most damaging empirical result for the use of SET as a proxy for teaching quality, because random assignment removes the selection problems that let most critics dismiss observational findings. It does not, however, show that student feedback is worthless - it shows that satisfaction and durable learning are different constructs that a single instrument cannot measure at once.
Almost every argument about whether course evaluations measure teaching quality runs into the same wall: students choose their courses. Motivated students pick demanding lecturers, weaker students avoid them, and any correlation between ratings and outcomes is hopelessly confounded. You cannot randomise students to professors in a normal university.
Except in two places, someone effectively did.
The two natural experiments
Carrell and West (2010), published in the Journal of Political Economy, used data from the United States Air Force Academy, where students are randomly assigned to course sections and must take a fixed sequence of standardised follow-on courses (Carrell & West, 2010, JPE 118(3): 409-432). Because the introductory syllabus, the exams and the grading were common across sections, and because students could not select their instructor, the design isolates the instructor's causal contribution.
Their finding: professors who raised student achievement in their own course received better student evaluations - and their students went on to do worse in the more advanced follow-on courses. Instructors whose students underperformed contemporaneously but excelled later were rated poorly.
Braga, Paccagnella and Pellizzari (2014), in Economics of Education Review, replicated the logic in a European setting using administrative data from Bocconi University, where first-year students are randomly allocated to class sections (Braga et al., 2014, EER 41: 71-88). They measured teacher effectiveness by students' performance in subsequent courses and compared it with the evaluations those teachers received.
Their conclusion is blunt: the effectiveness measure is negatively correlated with student evaluations. The same paper reports that evaluations respond to meteorological conditions - a detail worth holding onto, because a measure that moves with the weather is telling you something about its construct.
Two very different institutions, two different countries, two different disciplines mixes, one direction of result.
Why this is harder to dismiss than the bias literature
The literature on bias in SET is large and genuinely contested; we have argued elsewhere that it should be read honestly rather than cherry-picked. The follow-on course studies are different in kind, for three reasons.
- Random assignment. The standard rebuttal to observational SET research - that students sort themselves - simply does not apply.
- An objective outcome. Follow-on achievement is measured by common, externally graded assessments, not by self-reported learning. This sidesteps the validity problems of student self-assessment.
- A longer horizon. Most validity studies correlate ratings with achievement in the same course. These studies ask whether the learning survived into the next one, which is much closer to what a degree programme actually claims to deliver.
The mechanism both papers propose is the same, and it is not a story about lazy students or vindictive ratings. Braga and colleagues model it as a choice between real teaching and teaching to the test. An instructor who covers exactly what the exam will ask, in a well-signposted and cognitively comfortable way, produces students who feel well taught and score well immediately. An instructor who demands that students struggle with material, build transferable understanding and tolerate confusion produces worse immediate results, more student frustration - and better long-run learning.
This should sound familiar. It is the same construct that shows up in the active-learning penalty, where students in effective active classrooms learn more but feel they learned less. The Harvard experiment showed it in perceptions; the random-assignment studies show it in careers-worth of downstream grades.
It is also consistent with the broader meta-analytic picture. Uttl, White and Gonzalez, reviewing multisection validity studies in Studies in Educational Evaluation, concluded that once prior-ability and small-study effects are handled properly, SET ratings and student learning are essentially unrelated.
Critics argue: two unusual institutions do not generalise
This objection is correct and important, and anyone citing these studies without it is overselling.
The Air Force Academy is an extreme case: a compulsory core curriculum, standardised exams, military discipline, mandatory attendance, and a student body unlike any civilian intake. Bocconi is a selective private business university. Neither resembles a typical European public institution with elective pathways and heterogeneous cohorts. The compulsory follow-on structure that makes the design possible is also what makes it unusual.
There are further limits worth being precise about:
- Negative correlation is not zero information. A negative relationship still means SET data carries signal - it is just signal about something other than durable learning. Knowing that students found a course frustrating is genuinely useful.
- These designs measure instructor value-added in a narrow, examinable domain. Follow-on grades reward the kind of learning that later exams test. That is a better outcome than satisfaction, but it is not the whole of education, as the limits of measuring employability illustrate.
- The effect sizes are not enormous. The headline is the sign of the correlation, not its magnitude.
The reasonable inference is therefore narrower than "course evaluations are worthless". It is this: using average satisfaction ratings as a proxy for teaching effectiveness in high-stakes decisions is not supported by the best-identified evidence available, and may invert the intended incentive. That is precisely the concern behind using SET in tenure and promotion.
What follows practically
The instinctive response - abandon student feedback - is the wrong one, and it is the response this evidence least supports. Students remain the only people present at every teaching session, and their experience is a legitimate object of enquiry in its own right. The problem is not asking them. The problem is asking them one badly-specified question and treating the average as a quality verdict.
Three practical shifts follow.
Separate the constructs you are measuring. Satisfaction, perceived learning, workload, clarity, and belonging are different things. Collapsing them into an overall score guarantees the halo dynamic these studies expose. Ask about experience as experience, and get evidence about learning from direct measures.
Ask about difficulty in a way that distinguishes productive from unproductive struggle. The follow-on studies imply that the instructors most worth keeping may be generating discomfort. A scale item cannot tell you whether a student was confused because the teaching was unclear or because the material was genuinely hard and the instructor refused to trivialise it. Those two answers have opposite implications, and they produce the same rating.
Stop ranking, start diagnosing. If ratings are negatively correlated with long-run effectiveness in the cleanest available designs, league tables of instructors are indefensible. Formative use is not.
Where Koji fits
The distinction between productive and unproductive difficulty is the clearest example of something a Likert scale structurally cannot capture and a conversation can. Koji for Education runs AI-moderated conversational interviews that follow up on a low clarity rating with the obvious next question - what specifically was unclear, and what did you do about it? The difference between "the lecturer never explained the notation" and "it took me three attempts and then it clicked" is the difference between a teaching problem and effective, demanding instruction. Both currently arrive as a 2 out of 5.
Koji supports six structured question types alongside open probing, so workload, clarity and perceived learning can be measured as separate constructs rather than collapsed into one number. Automatic thematic analysis of open-text responses surfaces those distinctions at scale, and quality scoring flags low-information responses so a committee is not reading noise. Because moderation is standardised and bias-aware, the probing is consistent across every student and every course - unlike human-run focus groups, where interviewer variation is added to the problem.
Crucially, Koji supports formative, mid-cycle collection and closing-the-loop action tracking. The follow-on evidence argues for using student feedback to improve teaching during a course, not to rank teachers after it - and that is the workflow Koji is designed around. Programme- and institution-level reporting keeps the focus on curriculum-level patterns rather than individual scoreboards, and data handling is GDPR-compliant and EU-appropriate throughout.
Koji does not claim to measure teaching effectiveness. Given the evidence above, no student-feedback instrument should. What it does is surface what students actually experienced, in enough detail that a programme director can tell the difference between a course that is failing and a course that is hard.
Colleagues running general user or customer research use the same AI interview engine on the main Koji platform.
Measure experience honestly, and stop pretending it is a proxy for learning. See how Koji for Education separates the two.