New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Programme-Level vs Course-Level Evaluation: What Your Module Surveys Can't See

Course-by-course surveys give you 40 fragmented snapshots and no picture of the degree. Why programme-level evaluation — not module averages — is what ESG 2015 and good curriculum design actually require.

Koji Education Team

Product ·

Bottom line up front: Evaluating teaching one module at a time produces dozens of low-powered, noisy snapshots and no coherent view of the thing students actually graduate from — the programme. The European Standards and Guidelines (ESG 2015) ask institutions to monitor and review programmes, not just modules, and the most important quality problems — incoherent progression, duplicated content, skills that no single course owns — are invisible at module level by construction. The fix is not to abandon course feedback but to aggregate and triangulate it into programme-level evidence.

The unit of analysis problem

Almost every European university runs end-of-semester surveys at the level of the individual course or module. Each instructor gets a report; each report is read (or not) in isolation. This feels rigorous — every course is covered — but it quietly commits a methodological error: it treats the module as the unit of analysis when the unit students experience, and the unit that accreditation cares about, is the programme.

A degree is more than the sum of its modules. Constructive alignment, the framework introduced by John Biggs in 1996, holds that teaching and assessment should be aligned to intended learning outcomes — and the most consequential outcomes (critical thinking, research literacy, professional judgement, ethical reasoning) are developed across many courses, owned by none, and therefore measured by no single module survey. An analysis of constructive alignment in curricula notes that outcomes such as ethical development, intercultural competence and social responsibility are frequently espoused at programme level yet rarely assessed anywhere — precisely the kind of gap a module-by-module lens cannot detect.

What the regulators actually ask for

The Standards and Guidelines for Quality Assurance in the European Higher Education Area (ESG 2015) are explicit about the level of analysis. Standard 1.9, Ongoing monitoring and periodic review of programmes, requires that institutions "monitor and periodically review their programmes to ensure that they achieve the objectives set for them and respond to the needs of students and society." The guideline names what should feed that review: workload, progression and completion rates, the effectiveness of assessment, student expectations and satisfaction, and the learning environment — all programme-level constructs. Standard 1.3 on student-centred learning and 1.7 on information management likewise frame data as something to be analysed and acted upon at the level of the student journey, not the individual class.

In other words: the dominant practice (module surveys, read in isolation) and the regulatory expectation (programme monitoring, triangulated and acted upon) are misaligned. Many institutions generate vast quantities of course-level data and almost no genuine programme-level evidence.

Why module-level numbers are statistically weak

There is also a measurement reason to be sceptical of single-module averages. Module surveys are frequently small-sample, low-response affairs. Dennis Nulty's widely cited 2008 paper, "The adequacy of response rates to online and paper surveys in higher education", showed that the response rates needed for a defensible class-level estimate are far higher than most online evaluations achieve — and that small classes need proportionally more responses, not fewer. A 30-student elective with a 40% response rate yields an estimate so wide that comparing its 3.9 to another module's 4.2 is noise dressed as signal.

Aggregate those same responses to the programme level, however, and you gain statistical power. Patterns that are invisible in any single noisy module — a cohort that consistently reports assessment bunching in week 11, a progression cliff between years one and two, a research-methods thread students never feel they were taught — emerge clearly when feedback is pooled and read longitudinally. The signal was always there; the module lens just could not resolve it.

But doesn't programme-level evaluation lose the individual instructor?

This is the strongest objection, and it deserves a direct answer. Critics argue that aggregating to the programme dilutes accountability: if you stop reporting per-instructor numbers, how does a head of department spot the one course that is genuinely failing students?

Three responses. First, per-instructor numbers were never trustworthy for that purpose anyway — the best evidence, including Uttl, White and Gonzalez's 2017 meta-analysis "Meta-analysis of faculty's teaching effectiveness", found the correlation between student ratings and actual learning to be essentially zero once prior ability is controlled. Ranking instructors on noisy module means was always a misuse. Second, programme-level evaluation does not delete course data; it contextualises it — a weak module reads very differently when you can see whether it is a local problem or a symptom of a structural gap. Third, genuinely struggling courses are better surfaced through formative, mid-cycle feedback and triangulation with peer observation and outcome data than through a single summative league table. Aggregation and accountability are not opposites; crude ranking is just a poor proxy for both.

What good programme-level evaluation looks like

Doing this well means changing both what you collect and how you read it:

  • Map questions to programme outcomes, not just course satisfaction. Ask students about the coherence of their journey, the spacing of assessment, and whether the skills the programme promises are actually being built — questions no single module owns.
  • Triangulate. Combine evaluation data with progression and completion statistics, employer and graduate feedback, and peer review, as ESG 1.9 envisages. No single source is sufficient.
  • Read qualitatively, at scale. The richest programme-level signal lives in open text — students describe the journey in comments far more than they do in scale items. The barrier has always been that nobody can read thousands of comments across forty modules consistently.

A worked example: the missing methods thread

Consider a research-methods competence that a master's programme promises its graduates. It is introduced in a first-semester core module, assumed in a second-semester one, and assessed properly only in the final dissertation. No single module owns it, so no single module survey asks about it. Each course can return a respectable 4.0 while students quietly report — in dissertation supervision, far too late to act — that they never felt taught to design a study. The module-level dashboard shows green across the board; the programme-level reality is a gap that costs students at the most consequential moment of their degree.

This is the signature failure of course-by-course evaluation: problems that live between modules are owned by none of them. A programme-level lens asks a different question — not "how was this class?" but "did your degree build what it promised?" — and reads the answers across the whole cohort and the whole journey. The spacing of assessment, redundancy between courses, prerequisites that do not actually prepare students, threshold concepts that fall through the cracks: all of these are programme properties, invisible to any instrument pointed at a single module. The data to see them usually already exists, scattered across forty separate reports nobody reads together. The task is not collecting more; it is aggregating, threading and reading what you already have at the level where the curriculum actually lives.

Where Koji fits

This last barrier is exactly what an AI-native evaluation platform can lower. Koji for Education runs AI-moderated conversational interviews rather than static forms: instead of a fixed Likert grid, the AI can probe why a student felt under-supported and connect it to the wider programme experience. Its six structured question types (open-ended, scale, single- and multiple-choice, ranking, yes/no) let you ask outcome-level questions alongside course ones, and its automatic thematic analysis reads open-text feedback across every module in a programme with a single, standardised method — surfacing the cross-cutting themes that human readers, working module by module, reliably miss. Results roll up into programme- and institution-level reporting with action tracking, so a curriculum committee can see the journey, not just forty disconnected averages. Because the same conversational interview engine also powers general research on the main Koji platform, teams already doing customer or stakeholder research will find the approach familiar.

Koji does not claim to eliminate the weaknesses of student feedback — no instrument does. What it changes is the unit of analysis: from the noisy single module to the coherent programme, which is where both your students and your regulators are actually looking.

See your programme, not just your modules. Explore Koji for Education to move from forty disconnected reports to one evidence-led view of the degree.