Does the 8am Slot Cost You Half a Point? Timetable Bias in Course Evaluations
Early-morning classes measurably damage attendance, sleep and grades - that part of the evidence is strong. Whether they also depress teaching ratings is genuinely contested. Here is how to tell the difference, and why the timetable is the confound most quality offices never model.
Koji Education Team
Product ยท
Short answer: The evidence that early-morning teaching slots harm attendance, sleep and academic performance is strong and large-scale. The evidence that they directly depress student evaluations of teaching (SET) is mixed and much weaker. Those are two different claims, and conflating them is a mistake. The defensible position is that timetable slot is a plausible confound in your evaluation data that you should test for locally - not a settled bias you should quietly adjust away.
Most bias conversations in course evaluation focus on attributes of the instructor: gender, accent, age, attractiveness. Far less attention goes to the variable that is assigned to instructors by a scheduling office with no pedagogical intent whatsoever - the slot in the week where the course happens to land. If a 08:00 Monday lecture systematically scores lower than the same course taught at 14:00, and that difference feeds into a promotion file, an institution has converted a registrar decision into a judgement about teaching quality.
The strong evidence: what the timetable does to students
The most rigorous recent work here is not about ratings at all. It is about behaviour.
Yeo and colleagues, publishing in Nature Human Behaviour in 2023, combined several large digital-trace datasets from a single large university (Yeo et al., 2023). The scale is unusual for education research:
- Attendance, measured from Wi-Fi connection logs for 23,391 students across 337 lecture courses: attendance at 08:00 classes was roughly 10 percentage points lower than at classes starting at 10:00 or later.
- Sleep, measured by actigraphy in 181 undergraduates across 7,329 nocturnal sleep recordings: students slept about one hour less on days with early-morning classes. Close to a third of students did not wake in time for their 08:00 class.
- Grades, across 33,818 students over six semesters: students with more morning-class days per week had lower GPA, with medium effect sizes (Cohen's d between -0.37 and -0.40) for those with three to five morning days per week compared with none.
That is a coherent causal story: early slots shorten sleep, suppress attendance, and depress performance. Note what it is not, though. None of those outcomes is a teaching rating.
The weaker evidence: what the timetable does to ratings
Here the literature is genuinely unsettled, and it would be dishonest to present it otherwise.
The question has been asked for decades. Hinkin's 1991 study in Journal of Management Education was titled, tellingly, The Effects of Time of Day on Student Teaching Evaluations: Perception versus Reality - and found that student and faculty beliefs about the effect diverged sharply from what the data showed. Since then, results have gone both ways. Some within-instructor, within-semester designs, comparing the same lecturer teaching the same course in two different slots, find no reliable effect of time of day on ratings. Others, such as a case study in the Journal of Computing Sciences in Colleges, find that later sections scored lower - and, revealingly, scored lower even on items that should have nothing to do with the clock, such as fairness of assessment methods.
That last detail is the interesting one, because it is a halo signature: when a circumstantial factor moves ratings on items it logically cannot touch, you are looking at a general affective response rather than a considered judgement. It is the same pattern documented for transient mood, where Braga, Paccagnella and Pellizzari (2014) found in Bocconi University administrative data that student evaluations respond to meteorological conditions.
So: strong mechanism evidence, contested outcome evidence. What should a quality office actually conclude?
Mechanism is not the same as bias - and this distinction matters
Suppose your 08:00 sections do score lower. There are at least three different explanations, and they call for opposite responses.
- Genuine experience difference. Students in that slot really did attend less, sleep less and struggle more. Their lower ratings are an accurate report of a worse learning experience. The measurement is valid; the cause is the timetable.
- Attribution error. The experience was worse, and students attributed it to the instructor rather than to the schedule - a textbook case of the fundamental attribution error in course evaluations. The measurement is contaminated as a judgement of teaching.
- Composition. Early slots are not filled randomly. Timetabling interacts with programme, year group, commuting distance and part-time status. What looks like a slot effect may be a cohort effect.
Only the second is bias in the strict sense. But all three make raw cross-slot comparison indefensible, which is the practical conclusion that matters. This is the same logical structure as class-size effects and discipline baselines: a property of the delivery context, not of the teaching, ends up encoded in the number.
Critics argue: this is a rounding error, and you are inventing an excuse
The strongest objection is that slot effects on ratings, where they exist at all, are small - certainly smaller than the mechanism effects on grades - and that treating every contextual variable as a confound eventually leaves nothing comparable at all. If class size, discipline, level, modality, cohort composition and now the timetable all need adjusting for, the evaluation system collapses under its own caveats. There is also a real risk of motivated reasoning: an instructor with poor scores has an obvious incentive to blame the slot.
Both points land, and the honest response is threefold.
First, the appropriate unit of concern is not the average effect but the decision margin. A 0.1-point slot effect is irrelevant when reading feedback formatively and potentially decisive when a department applies a norm-referenced cut score. The question is never "is the effect big?" but "is the effect big relative to the difference this institution is prepared to act on?" - which is why effect sizes matter more than differences.
Second, the answer to "too many confounds" is not to ignore them but to stop using the instrument for the thing that requires them all to be controlled. Ranking instructors on a shared scale demands comparability that SET data cannot deliver. Reading feedback to improve a course does not.
Third - and this is the part institutions skip - you do not have to rely on the published literature at all. Slot effects are heterogeneous across institutions, cultures and timetabling regimes. Whether they exist in your data is an empirical question you can answer with data you already hold.
What to actually do
- Test locally before you assume. Fit a model of overall rating on timetable slot, with course and instructor fixed effects where you have repeated observations. If the same instructor teaches the same course in two slots, you have a near-natural experiment sitting in your student records system.
- Record the slot as evaluation metadata. Most institutions never join timetable data to evaluation data, which makes the question unanswerable by construction. This is a five-minute schema decision with years of analytical consequence.
- Never rank across slots without acknowledging it. If comparison is unavoidable, compare like with like, or use shrinkage estimators rather than raw means.
- Ask students directly. The decisive evidence is often qualitative: did the schedule shape their experience, and did they factor it into their rating? A Likert scale cannot tell you. A follow-up question can.
Where Koji fits
Koji for Education is built around the last point. A static form collects a number and leaves you guessing at its cause. Koji runs AI-moderated conversational interviews that probe why a rating was given - so when a student rates attendance or engagement poorly, the moderator can follow up on whether the timetable, the commute, or the teaching drove it. That distinction is exactly what separates explanation 1 from explanation 2 above, and it is invisible to a five-point scale.
Because the moderation is standardised and bias-aware, every student is probed consistently - there is no human-interviewer variation to add on top of the variation you are already trying to isolate. Automatic thematic analysis surfaces scheduling as a theme across a programme when it recurs, rather than leaving it buried in free text that nobody reads. And programme- and institution-level reporting lets a quality office see whether a pattern is one unlucky lecturer or a systematic timetabling problem that belongs on the registrar's desk, not in a promotion file.
Koji does not eliminate timetable effects - nothing collected from students can. It surfaces them, and it makes them attributable, which is the precondition for acting on them fairly. Koji also supports formative, mid-cycle collection, so a scheduling problem can be caught in week four rather than diagnosed in a report the following year. All of it is GDPR-compliant with EU-appropriate data handling.
Teams who run general user and customer research alongside their academic work often use the same underlying AI interview engine on the main Koji platform.
Stop guessing whether the slot moved the score. See how Koji for Education surfaces the context behind a rating - and turn a confound you cannot control into evidence you can act on.