New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

The Dual-Purpose Problem: Why One Course Evaluation Can't Serve Both Improvement and Personnel Decisions

Most universities run a single end-of-term survey and then ask it to do two contradictory jobs: help teachers get better, and decide who gets promoted. Those two jobs have opposite design requirements. Trying to serve both from one instrument quietly corrupts both — and the fix is not a better questionnaire, it is two separate feedback streams.

Koji Education Team

Product ·

The short answer: A course evaluation designed to help an instructor improve and one designed to judge an instructor for tenure, promotion, or renewal are not the same instrument tuned differently — they have contradictory requirements. Formative evaluation needs candour, granularity, timing early enough to act, and psychological safety. Summative evaluation needs standardisation, comparability, defensibility, and stakes. When a single end-of-term survey is asked to do both, the summative stakes poison the formative honesty, and the formative openness undermines the summative rigour. The defensible design is to run two separated streams — a low-stakes improvement loop and a carefully bounded accountability record — rather than to keep polishing one questionnaire and hoping it can be both.

Why this is a design conflict, not a wording problem

Quality-assurance offices tend to treat weak course evaluations as a wording problem: better items, a cleaner scale, a nudge to raise response rates. But the deepest problem is upstream of any question. It is that the same dataset is expected to answer two questions that pull in opposite directions.

Consider what each purpose actually demands.

Formative (improvement) evaluation is for the teacher. Its value is diagnostic: what should I change next time, and why? That requires students to be specific and honest about what confused them, timing early enough that the current cohort can benefit, and — crucially — an environment where naming a problem carries no threat. The teacher needs signal about their own practice, not a number to be ranked against colleagues.

Summative (accountability) evaluation is for the institution. Its value is comparative and defensible: can we stand behind this as evidence in a personnel decision or an accreditation file? That requires standardised items across very different courses, statistical comparability, and a paper trail that will survive an appeal.

You cannot maximise both from one survey administered once, at the end, to everyone. The moment students know their end-of-term ratings feed a promotion file, the incentives around candour change — for them and, more importantly, for the instructor, who now has every reason to teach toward the score rather than toward the learning. This is simply Goodhart's Law operating inside your quality system: once the measure becomes the target, it stops being a good measure.

The evidence that the summative use is the weaker one

If we are going to let one purpose dominate the instrument, it had better be the well-supported one. It is not.

The most-cited synthesis of the does-it-measure-learning question is Uttl, White and Gonzalez's 2017 meta-analysis in Studies in Educational Evaluation, which re-analysed decades of multisection studies — the cleanest design, where different sections of the same course are compared against a common final. Their conclusion: once you correct for small-study artefacts, student rating scores are essentially unrelated to how much students actually learn, with ratings explaining at most about 1% of the variance in learning in the multisection data (Uttl, White & Gonzalez, 2017). Students do not, on average, learn more from instructors they rate more highly.

On the statistical mechanics, Philip Stark and Richard Freishtat's An Evaluation of Course Evaluations is blunt: the common practice of comparing an instructor's average score to a department average should be abandoned. Ratings are ordinal categories — the numbers are labels, not quantities — so averaging them and treating a 3-versus-4 gap as identical to a 6-versus-7 gap is not sound statistics. They recommend reporting the distribution of responses and the response rate, and explicitly warn against using averages in high-stakes personnel decisions (Stark & Freishtat, 2014).

Put those together and the summative use is on thin ice: the number does not track learning, and the standard way of reducing it to a comparison is statistically indefensible. Yet it is the summative use — the stakes — that reshapes the whole instrument and suppresses the formative candour that would have been valuable.

What the dual-purpose survey actually produces

When one survey serves both masters, you get predictable pathologies:

  • Bland, safe feedback. Students who suspect their comments carry weight in someone's career, or who simply fill in a required form at the end of a busy term, default to the middle and the generic. The rich, specific, improvement-relevant detail is exactly what gets lost.
  • Teaching to the evaluation. Instructors under summative pressure optimise for satisfaction — lighter workloads, more lenient grading, more entertaining delivery — rather than for durable learning, which the active-learning penalty literature shows can actually depress ratings even as it raises learning.
  • Late data. An end-of-term survey is summatively convenient (it covers the whole course) but formatively useless (the cohort that gave it has already left). The people who could act on it are gone.
  • A false sense of rigour. Because the number looks objective, committees treat a 4.1-versus-3.8 gap as a finding, when the confidence intervals around small classes routinely swamp differences of that size.

But doesn't a well-designed single survey solve this?

This is the strongest objection, and it deserves a direct answer. Advocates of unified instruments argue that a carefully validated questionnaire — with distinct formative and summative sections, or with formative items flagged as "for your eyes only" — can serve both. Two responses.

First, stakes are not a property of the questionnaire; they are a property of the system. You can label a section "formative", but if the same administration, the same cohort, and the same visible dataset also feed the promotion committee, students and staff will read the whole exercise as high-stakes. Psychological safety is not restored by a subheading.

Second, even a perfect instrument cannot fix the timing contradiction. Summative evaluation must cover the whole course, so it comes at the end. Formative evaluation is only useful if it comes early enough to change the course the respondents are in. One administration cannot be both retrospective-and-complete and mid-course-and-actionable. This is why the serious methodological literature treats formative and summative evaluation as different activities with different cadences, not two flavours of one survey.

A fair concession: triangulated summative evaluation — where student ratings are one input among peer observation, teaching portfolios and direct evidence of learning — is far more defensible than ratings alone, and we have argued exactly that in our piece on triangulating teaching evaluation. The dual-purpose problem is not an argument against summative evaluation existing. It is an argument against loading it onto the same instrument you rely on for improvement.

The honest design: two separated streams

The resolution is structural, not cosmetic. Separate the streams:

  1. A formative improvement loop, run mid-cycle, low-stakes, owned by the teacher and the teaching-and-learning centre, optimised for specific and actionable feedback — and never routed into a personnel file.
  2. A summative accountability record, run at end of term, standardised, reported as distributions rather than averages, and — critically — treated as one strand of evidence within a triangulated review, never as a decisive single number.

How Koji supports the separation

Koji for Education was built around this distinction rather than against it. The platform runs AI-moderated conversational interviews that probe beyond a rating — when a student says a topic was confusing, the moderator asks which topic and why — so the formative stream captures the diagnostic detail a Likert grid cannot. Because that improvement loop can be run mid-cycle, the cohort that gives the feedback is still present to benefit from the changes, resolving the timing contradiction that no single end-of-term survey can escape.

The AI moderation is standardised and bias-aware, which is what a defensible summative record needs — every student gets a consistent, neutral interviewer rather than the variable quality of human focus-group facilitation. Automatic thematic analysis turns open-text at scale into structured themes, and closing-the-loop action tracking records what was changed in response — the part of formative evaluation that legacy tools abandon entirely. Programme- and institution-level reporting keeps the summative strand comparable without collapsing it into a single averaged score. All of it is GDPR/AVG-compliant and EU-appropriate by design.

Legacy SET platforms — EvaSys, Qualtrics survey forms, paper sheets — were built for the one-survey-does-everything world: administer once, average a Likert score, rank. That is precisely the design that creates the dual-purpose problem. Koji treats improvement and accountability as two jobs with two cadences, which is what the methodology has said they are all along.

Many teaching-and-learning teams also run wider student- and staff-research projects; the same conversational interview engine powers the main Koji platform for general user and customer research, so the skills transfer across both.

If your course evaluations are being asked to be a coaching tool and a courtroom exhibit at the same time, the problem is not the questions. It is that you are running one instrument where you need two streams. See how Koji for Education separates the improvement loop from the accountability record.