New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Does Your Course Evaluation Ask Whether Students Actually Met the Learning Outcomes?

Most evaluation instruments interrogate the teacher's performance and ignore the one thing the course was designed to deliver: the intended learning outcomes. Design-focused evaluation flips the question — and it is the more useful one.

Koji for Education

Research & Editorial Team ·

Bottom line: Standard course evaluations ask students to rate the instructor — clarity, enthusiasm, organisation. They rarely ask whether the course's intended learning outcomes were actually achieved, or whether students perceived the teaching and assessment as aligned to them. That omission matters: under John Biggs's principle of constructive alignment, a course is only coherent when its outcomes, activities and assessment point the same way — and "design-focused evaluation" asks students precisely about that alignment. Switching the question from "how good was the lecturer?" to "did the way this course was built actually help you achieve what it promised?" produces feedback that is more diagnostic, less personality-driven, and far more useful for improving the course rather than judging the person.

The question your instrument is not asking

Pull up almost any institutional evaluation form and read the items. "The lecturer explained concepts clearly." "The teacher was enthusiastic." "The course was well organised." These are instructor-performance items. They orbit the teacher's personality and delivery.

Now ask: where on that form does a student report whether they actually achieved the learning outcomes the course was designed around? Where do they say whether the assessment tested what the teaching prepared them for? On most instruments, nowhere. We have built evaluation systems that meticulously rate the performer and barely mention the performance's purpose.

This is not a small gap. It means our headline feedback measure is structurally disconnected from the design logic that every modern curriculum claims to follow.

Constructive alignment: the design logic we evaluate against — except we don't

The dominant framework in higher-education course design is John Biggs's constructive alignment, introduced in 1996 and now embedded in quality frameworks and validation processes across Europe. Its claim is simple and powerful: a course works when the intended learning outcomes, the teaching and learning activities, and the assessment tasks are all aligned — when what students are asked to do in class, and what they are assessed on, both directly serve what they are supposed to be able to do by the end.

Biggs's deeper insight is about student perception: students, he argued, learn what they think they will be tested on. The student's experience of the assessment effectively drives their learning. Which means the single most important thing to evaluate is not whether the lecturer was charismatic, but whether students perceived the activities and assessment as coherently pointing at the stated outcomes — because that perceived alignment is what actually shapes how they learn.

So here is the contradiction at the heart of much current practice: institutions design courses using constructive alignment, validate them on constructive alignment, and then evaluate them with instruments that ask about almost everything except alignment.

Design-focused evaluation: asking the aligned question

There is an established alternative, and it has a name. Design-focused evaluation, developed by the educational researcher Calvin Smith, deliberately shifts evaluation away from rating inputs (the lecturer) or raw outputs (a global grade) and toward students' perceptions of the effectiveness of the alignment in the course's design. It asks students to report on how well specific designed elements — a particular activity, a piece of formative feedback, an assessment structure — helped them achieve the intended outcomes.

The shift sounds subtle. Its consequences are not.

It depersonalises the feedback. "Did the weekly problem sets help you reach the stated outcome of applying the method independently?" is a question about a design choice, not about whether students liked the teacher. That structurally reduces the halo, attractiveness and stereotype effects that contaminate personality-rating items — a meaningful benefit given the documented bias in instructor-focused ratings.

It produces actionable information. "The lecturer was disorganised" tells you little you can fix. "The assessment tested material the seminars never practised" names a specific alignment failure and points straight at the remedy. Design-focused items generate improvement actions rather than verdicts on individuals.

It tracks what actually matters. There is even evidence that course design outperforms instructor charisma as a predictor of evaluation quality and engagement: a 2024 study in Higher Education found course design to be a stronger predictor of student evaluation of quality and student engagement than teacher ratings. If design predicts the outcomes we care about better than personality does, evaluating design rather than personality is simply better measurement.

"But students can't judge whether they met the learning outcomes"

This is the serious objection, and it has two strands worth separating.

The first: self-reported learning is unreliable — students routinely misjudge how much they have learned, and feel they learn less from the active methods that teach them most. This is true and well-evidenced, and it is a genuine caution. But design-focused evaluation does not ask students to certify that they met the outcomes — that is what assessment is for. It asks them to report their experience of the alignment: whether the activities felt connected to the goals, whether the assessment matched the teaching, where the coherence broke down. Students are reliable witnesses to their own experience of a course's structure even when they are unreliable judges of their own knowledge gains. The trick is asking the question they can actually answer.

The second strand: isn't this just shifting from one set of subjective ratings to another? Partly — all student feedback is perception data. But perception of alignment is perception of something the course designer directly controls and can directly change, whereas perception of a lecturer's personality is largely not actionable and heavily biased. Trading unactionable, bias-prone perception for actionable, design-relevant perception is a real gain even if both are subjective.

A third, honest limitation: design-focused items are harder to write and harder to standardise than "rate your lecturer 1 to 5." A vague alignment question is no better than a vague satisfaction question. The quality of the method depends entirely on the quality of the questioning — which is exactly where most static survey forms fall down.

Why the instrument, not just the items, has to change

You could bolt a few alignment questions onto an existing survey, and that would help. But a paper or web form hits a ceiling fast: it can ask "did the assessment match the teaching?" but it cannot ask the student to explain where it diverged, cannot follow a promising thread, and cannot distinguish a single frustrated voice from a structural problem half the cohort experienced.

This is where Koji for Education fits the method rather than fighting it. Its AI-moderated conversational interviews are built to probe beyond a rating — so when a student says the assessment felt misaligned with the seminars, the moderator asks which assessment and how, turning a flat score into a located, fixable finding. The six structured question types let you frame outcome- and alignment-specific items (a ranking of which activities best supported a stated outcome; an open-ended probe on where coherence broke), and automatic thematic analysis aggregates those conversations into the recurring alignment gaps across a cohort — exactly the evidence a programme team needs to redesign rather than to judge. Because the moderation is standardised and bias-aware for every student, it sidesteps the human-moderator inconsistency of focus groups while keeping the depth, and its formative, mid-cycle mode means you can surface a misalignment in week 5 and fix it before the cohort is assessed, not autopsy it afterwards. It does not turn perception into proof — assessment still certifies learning — but it captures the alignment experience with a fidelity no Likert grid can, under GDPR/AVG-compliant, EU-appropriate data handling. Closing-the-loop tracking then records whether the redesign actually addressed the gap.

The same conversational interview engine powers the main Koji platform for product and user research, where "did this feature actually help users do the job they hired it for?" is the commercial twin of the alignment question.

The reframe

If your institution designs courses on constructive alignment, it should evaluate them on alignment too. Stop asking students only to rate the performer; start asking whether the course, as built, helped them reach what it set out to deliver. That single change moves your evaluation from a personality referendum to a design diagnostic — feedback you can actually act on, aimed at the thing the course was for.

Want feedback that tells you what to redesign, not just whom to judge? See how Koji for Education works.