New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Washback: How Course Evaluation Quietly Reshapes the Teaching It Claims to Measure

A course evaluation is not a neutral thermometer. The moment scores carry consequences, they change the teaching they measure — a feedback effect language-assessment researchers call washback. Here is the evidence it operates in student ratings, why it is a validity problem, and how to design evaluation whose washback rewards good teaching.

Koji Education Team

Product · July 15, 2026

BLUF: Course evaluations are usually treated as a measurement instrument — a thermometer you dip into a class to read its temperature. But a thermometer that students, teachers, and committees can all see, and that carries consequences, does not merely measure the class; it changes it. In language-assessment research this feedback effect has a name — washback — and once you see it in course evaluation, the central design question shifts. It is no longer "is our number accurate?" but "in which direction does our number push teaching?" This article explains what washback is, presents causal evidence that it operates in student evaluations of teaching (SET), locates it inside Messick's theory of consequential validity, and argues that the fix is not to stop evaluating but to engineer instruments whose washback rewards good teaching rather than good performance.

Washback: a concept borrowed from testing

The term washback (or backwash) comes from language testing. In their foundational 1993 paper "Does Washback Exist?", J. Charles Alderson and Dianne Wall argued that high-stakes tests exert a powerful influence on classrooms — teachers teach to the test, learners study to the test — and they laid out fifteen distinct washback hypotheses to be tested empirically (Alderson & Wall, 1993, Applied Linguistics 14(2)). Their central caution transfers directly to the evaluation debate: washback is frequently assumed rather than demonstrated, and much of the early evidence rested on teachers' accounts rather than observation.

Course evaluation is not a test of students, but it functions as a high-stakes test of teachers. When ratings feed into probation, promotion, contract renewal, or a department league table, they acquire exactly the property Alderson and Wall worried about: they become powerful enough to reshape the behaviour they are supposed to observe. The instructor who knows a single "overall satisfaction" mean will decide their contract renewal has every rational incentive to teach toward that mean — and not necessarily toward learning.

The evidence that washback operates in student ratings

The strongest evidence is quasi-experimental. In a study exploiting the random assignment of students to instructors at the US Air Force Academy — where a common syllabus and centrally graded exams remove most confounds — Scott Carrell and James West found a striking dissociation. Instructors who raised their students' contemporaneous achievement and earned higher student evaluations tended to harm those same students' achievement in the follow-on course. Student evaluations were positive predictors of current-course grades but poor predictors of deep, downstream learning (Carrell & West, 2010, Journal of Political Economy 118(3)). The most plausible reading is washback: teaching optimised for the next survey (narrow, exam-focused, fluent, reassuring) diverges from teaching optimised for durable understanding.

Grade leniency is the clearest channel. Decades of work associate more generous grading with higher ratings, and experimental and longitudinal studies suggest the relationship is at least partly causal rather than a benign reflection of good teaching. When lenient grading buys better evaluations, the instrument is quietly rewarding a behaviour — grade inflation — that undermines the very standards a programme exists to uphold. This is negative washback in its purest form.

Bias findings compound the problem. Boring, Ottoboni, and Stark's re-analysis of 23,001 evaluations from a French university concluded that ratings "are more sensitive to students' gender bias and grade expectations than they are to teaching effectiveness", biasing scores against female instructors on dimensions as ostensibly objective as promptness of grading (Boring, Ottoboni & Stark, 2016, ScienceOpen Research). If the number is partly a measure of grade expectations, then teaching toward the number means managing expectations — not improving instruction.

Why washback is a validity problem, not just an ethics problem

It is tempting to file washback under "unintended consequences" and move on. Measurement theory does not let us. In Samuel Messick's unified account, validity is not a property of a test but of the interpretations and uses of its scores — and the social consequences of use are an integral part of the validity argument. A measure that produces perverse washback is, by Messick's standard, less valid for the purpose to which it is put, even if the items look reasonable. The consequential aspect of validity is exactly where washback lives.

This is why washback belongs alongside the more familiar validity threats we have written about — construct-irrelevant variance (the number reflecting things other than teaching) and the reliability-versus-validity distinction. Washback is the dynamic cousin of these static problems: it is what construct-irrelevant variance does over time when the score has teeth. It is also the deep structure beneath Goodhart's Law: "when a measure becomes a target, it ceases to be a good measure" is simply washback stated as an aphorism.

The counterargument: "all measurement changes behaviour — so why bother?"

The strongest objection is a good one. Every observation instrument in a social system exerts some backwash; if measuring inevitably distorts, is the honest conclusion that we should stop measuring teaching altogether? No — for two reasons.

First, washback has a direction. Alderson and Wall themselves distinguished beneficial from harmful washback. A well-designed evaluation can push teachers toward exactly the behaviours we want more of: clearer articulation of learning outcomes, timely formative feedback, inclusive design, alignment between assessment and objectives. The task is not to eliminate washback (impossible) but to engineer its sign.

Second, the alternative to evaluated teaching is not un-influenced teaching; it is teaching influenced by whatever informal signals fill the vacuum — corridor reputation, student complaints that reach a dean, enrolment numbers. Those signals have washback too, and it is usually worse: noisier, more biased, and invisible to quality assurance. The choice is between designed washback and accidental washback.

Designing for beneficial washback

If washback is inevitable and directional, evaluation design becomes an exercise in incentive engineering:

  • Measure process and learning behaviours, not just satisfaction. A number that rewards "I enjoyed this" invites showmanship — the classic Dr Fox effect. Questions anchored to specific, observable teaching practices create washback toward those practices.
  • Separate improvement from judgement. When the same instrument drives both formative development and high-stakes personnel decisions, the high-stakes use dominates and poisons the formative one. Formative, mid-cycle evaluation that never reaches a promotion committee generates far healthier washback.
  • Make gaming expensive. A single Likert mean is trivially gameable by charm or leniency. Rich, probed, qualitative evidence about what students actually did and understood is much harder to fake, so the cheapest path to a good result becomes — actually teaching well.
  • Report distributions and uncertainty, not league-table means, so that the target teachers optimise toward is not a spuriously precise decimal.

Where Koji fits

Koji for Education is built around the premise that the shape of the instrument determines the sign of the washback. Instead of a static Likert form that rewards expressiveness and leniency, Koji runs AI-moderated conversational interviews that probe beyond a single rating — asking students what they did, what changed in their understanding, and why — using six structured question types (open-ended, scale, single-choice, multiple-choice, ranking, yes/no). Because the evidence is conversational and thematically analysed rather than reducible to one number, the cheapest way to "score well" shifts toward genuinely clearer teaching and better-aligned assessment. Standardised, bias-aware AI moderation removes the human-moderator inconsistency that adds noise, and formative mid-cycle collection lets teachers act during a course rather than perform for an end-of-term verdict. Koji is careful about the claim: this mitigates and redirects washback; it does not eliminate it. No instrument can. The point is to make the washback work for you.

Many of the same dynamics show up whenever an organisation measures people using self-report — which is why teams running general user and customer research use the main Koji platform and its shared AI interview engine to get past satisfaction theatre to what respondents actually experienced.

The takeaway

Stop asking only whether your course-evaluation number is accurate. Ask where it pushes. A measure with consequences is a lever on teaching whether or not you designed it to be — and the responsible move is to design it deliberately, so that the easiest way to earn a good evaluation is to teach in the ways your students, your programme, and your accreditors actually value.

Ready to turn course evaluation from a number teachers game into feedback they can act on? Explore Koji for Education.