The Credit Doesn't Match the Clock: Using Course Evaluation to Audit ECTS Workload
An ECTS credit is defined in hours of student work, yet the evidence shows nominal credits and actual study time routinely diverge. Course evaluation is the instrument best placed to catch the mismatch — if it asks the right questions.
Koji Education Team
Product ·
Bottom line up front: Under the European Credit Transfer and Accumulation System, one ECTS credit is defined as 25–30 hours of student work, and a full academic year as 1,500–1,800 hours. But the research is consistent: the workload students actually experience routinely diverges from the credits a module carries. Because credit allocation underpins degree comparability across the entire European Higher Education Area, a systematic mismatch is not a minor inconvenience — it is a validity problem at the heart of the system. Course evaluation, asked properly, is the cheapest and most direct instrument for detecting it. The caveat: perceived workload and measured study time are not the same thing, and conflating them produces bad decisions.
A credit is a promise about hours — and the promise is often wrong
The ECTS Users' Guide (2015) is unambiguous: credits express the workload "students typically need to complete all learning activities" — lectures, seminars, independent study, assessment — required to achieve the learning outcomes. One credit equals 25–30 hours; 60 credits per year equals 1,500–1,800 hours. This is the load-bearing assumption that lets a semester in Lisbon count toward a degree in Helsinki.
The evidence that the promise is frequently broken is now substantial. Studies across the EHEA repeatedly find that ECTS credits track contact teaching time reasonably well but only loosely track total study time, and that nominal allocations are often poorly calibrated to what students actually do. Some researchers argue the 25–30-hour equivalence is itself oversized, with knock-on effects on both education quality and student health, and warn that persistent mismatch is "a threat to the credibility of the ECTS system itself." The recurring empirical pattern is not random noise but structured divergence: some modules systematically demand far more than their credits imply, others far less.
There is a second, related problem the same literature surfaces — the assessment load specifically. Harland and colleagues' influential analysis of the "assessment arms race" (Teaching in Higher Education, 2015) documented students facing an average of 0.68 to 1.44 graded assessments per week, a "pedagogy of control" that drives surface learning and crowds out the unassessed activities that credits are supposed to fund. Over-assessment is workload miscalibration with a particular cause, and it is invisible to evaluation that only asks "were you satisfied?".
Why this is a measurement problem, not just a grumble
For a quality-assurance officer, ECTS workload is not pastoral hand-wringing — it is the validity of the credit as a unit of measurement. If a "5 ECTS" module reliably consumes the hours of a 7.5 ECTS module, then:
- Degree comparability breaks. The premise of credit mobility is that a credit means the same thing everywhere. Systematic local miscalibration quietly violates it.
- Learning outcomes are at risk. Under-resourced credits mean students cannot do the work the outcomes require; over-loaded credits push them toward surface strategies. The same evidence base finds that students' workload experience — not raw hours — has the strongest association with how they approach learning.
- It distorts your other evaluation data. A poor satisfaction score on a "hard" module may be measuring workload miscalibration, not teaching quality — a confound that pollutes any inference you draw from the ratings.
"But isn't self-reported workload hopelessly unreliable?" — the counterargument
This is the serious objection, and it has real force.
Students are poor clocks. Retrospective estimates of hours-per-week are subject to recall bias, rounding, and the tendency to over-report effort on courses they disliked. A single end-of-term item — "How many hours per week did you spend on this module?" — produces numbers too noisy to recalibrate credits with confidence. Anyone who has tried knows the distributions are wide and the means unstable.
The response is not to abandon the question but to ask it better, and to separate two distinct constructs the research is careful to distinguish: time-on-task (objective hours) and workload experience (the subjective sense of being overloaded). They are conceptually different and predict different things. Robust workload evidence therefore (a) collects time-use closer to the event rather than months later, reducing recall bias; (b) measures experienced workload and its distribution across the term, not just a single hours figure; and (c) triangulates self-report with timetable data and assessment schedules rather than trusting any one source. Done this way, self-report becomes a legitimate signal — the retrospective-recall problem we have written about is a problem of method, not of the construct.
What good workload evaluation asks
Concretely, a workload-aware evaluation gathers more than an hours estimate:
- Distribution, not just total. When in the term did the load concentrate? Bunched assessment deadlines are a design failure invisible to a term-average figure.
- Experienced vs. nominal. Did the workload feel proportionate to the credits awarded? This captures the workload-experience construct the evidence flags as most consequential.
- Where the time went. Productive learning activity versus busywork — over-assessment specifically. This connects to our work on survey and over-evaluation fatigue.
- Open-ended detail. A number tells you that a module is miscalibrated; a student's description tells you why — which is what a programme team needs to fix it.
Where Koji fits
Workload auditing is exactly the kind of evaluation that legacy SET tools handle badly: a single retrospective Likert item cannot separate time-on-task from workload experience, cannot capture distribution across a term, and cannot tell you why a module overran.
Koji for Education is built for this. Its formative, mid-cycle collection lets you sample workload during the term rather than reconstructing it months later — directly attacking the recall-bias objection. Its AI-moderated conversational interviews can probe where the time went and why it felt heavy, separating "I spent 12 hours" from "those 12 hours were six low-value quizzes" — the time-on-task versus workload-experience distinction the research insists on. Six structured question types let you combine a quantitative hours/distribution scale with qualitative probes in one instrument, and automatic thematic analysis aggregates the open-text across a cohort so a programme team can see whether miscalibration is systematic or idiosyncratic. Programme-level reporting then lets quality-assurance staff compare experienced workload against nominal ECTS across every module — the institution-wide view that turns scattered complaints into recalibration evidence, all under GDPR/AVG-compliant handling.
The same conversational engine handles "how much effort did this actually take and where did it go" for customer and product research on the main Koji platform.
The honest boundary: Koji surfaces and structures workload signals far better than a static form, but it does not turn self-report into a stopwatch. Recalibrating credits is a judgement that belongs to the programme team, informed by triangulated evidence — Koji makes that evidence collectable and comparable; it does not make the decision for you.
Reading the distribution, not just the mean
A final methodological point that separates useful workload data from misleading workload data: the mean hours figure is almost always the least informative statistic you can report. Two modules can carry an identical average of, say, ten hours per week and represent completely different student experiences — one a steady, manageable rhythm, the other a calm first half followed by a brutal deadline pile-up in weeks ten to twelve. The credit allocation looks defensible on average and indefensible in lived reality.
This is why workload evaluation should report distribution and timing, not a single number. Where in the term did the load spike? Did several modules in the same programme schedule their major assessments in the same fortnight — a coordination failure no single module evaluation can detect, but which programme-level analysis exposes immediately? And how does experienced intensity vary across the cohort, given that the same nominal workload lands very differently on a commuting student with caring responsibilities than on a residential full-timer? Treating workload as a distribution across time and across students, rather than a scalar, is what turns complaint into actionable recalibration evidence — and it is precisely the kind of structured, multi-dimensional signal that a single end-of-term Likert item is structurally incapable of capturing.
The takeaway
The ECTS credit is a measurement instrument, and like any instrument it can fall out of calibration. The evidence says it frequently has — that nominal credits and lived workload diverge, and that over-assessment makes it worse. Course evaluation is the natural audit tool, but only if it stops asking "were you satisfied?" and starts asking, carefully, "what did this actually cost you, when, and on what?" Get that right and evaluation defends the integrity of the credit itself.
Want to audit whether your credits match the clock? See how Koji for Education captures workload evidence that legacy surveys miss.