New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Do Early-Morning Classes Get Lower Evaluations? Time-of-Day and Scheduling as Confounds

Evidence that the timetable slot - early morning, late afternoon, days per week - shifts student performance and mood, and therefore course evaluation scores, independent of teaching quality. What QA teams should record and adjust for.

Koji Education Team

Product

In brief: Probably, at least a little. There is strong causal evidence that the time of day a class meets changes how well students perform, and student performance and mood are among the most reliable predictors of the evaluation scores they give. Pope (2016) and Dills and Hernandez-Julian (2008) both show meaningful time-of-day effects on grades using large datasets and within-student comparisons. Because the timetable slot is assigned by administrators, not chosen by the instructor, it is a construct-irrelevant confound in student evaluation of teaching (SET).

What the research says

Two econometrically careful studies establish the underlying mechanism - that when a class meets affects outcomes.

Nolan Pope (2016), in The Review of Economics and Statistics, analysed a panel of nearly two million 6th-11th grade students in Los Angeles County. Holding the school day fixed and exploiting variation in the order in which students take subjects, he found students are more productive earlier in the day: having a maths class in the morning rather than the afternoon raised GPA by about 0.072 and English by about 0.032, and a morning maths class raised state test scores by an amount comparable to a one-quarter standard-deviation improvement in teacher quality. The key methodological strength is that these are within-student comparisons, so the effect is not driven by better students choosing morning slots.

Angela Dills and Rey Hernandez-Julian (2008), in Economics of Education Review, used over 105,000 grades from Clemson University and controlled for class size, semester, meetings per week and fixed student and class characteristics. They found students performed better in classes later in the day (roughly 0.024 grade points per hour) and in classes meeting fewer days per week (Tuesday/Thursday better than Monday/Wednesday/Friday). Note that the direction is the opposite of Pope for the morning-versus-afternoon question - a useful reminder that the effect depends on population, fatigue and cumulative schedule load, not a simple "mornings are best" law.

The bridge to course evaluations is the well-established link between how students do and feel and the score they give. Students who earn better grades, and who are in a better mood, systematically rate courses more highly; grade satisfaction is one of the strongest single correlates of SET. If the timetable slot shifts performance and alertness - a 7:30am section fighting fatigue, or a Friday-afternoon slot with thinned attendance and low energy - then part of the resulting evaluation reflects the clock, not the instructor. This is the same construct-irrelevant-variance logic that applies to class size, discipline and the physical room.

Why it matters for course evaluation in practice

For quality assurance the risk is again comparability and fairness:

  1. Slot-driven score gaps. An instructor repeatedly timetabled into 8am or late-Friday slots may accumulate lower scores than a colleague with prime mid-morning slots teaching the identical course. If SET informs probation or promotion, the timetable becomes a hidden career variable.
  2. Cohort fatigue effects. A class that is the students' fourth session of the day is being evaluated by tired respondents. The Dills-Hernandez-Julian result on cumulative schedule load suggests fatigue, not just clock time, is doing work - a signal that raw slot comparisons are crude.
  3. Spurious trend signals. If a module is moved from an early to a mid-morning slot between years, a score rise may be misattributed to a teaching change. Longitudinal narratives in ESG-aligned self-evaluation reports can be distorted.
  4. Response-rate interactions. Timetable slots also affect who responds (an 8am or Friday-afternoon session may see lower attendance and thus a more self-selected respondent pool), compounding the confound with a non-response problem.

The pragmatic response is to record the slot as a covariate and interpret scores in light of it - not to abandon SET, and not to over-correct with a precise "morning penalty".

Limitations and honest caveats

  • The direct SET evidence is indirect. The strongest causal studies (Pope; Dills and Hernandez-Julian) measure grades and test scores, not SET. The link from performance/mood to ratings is well supported but is an inferential step, so the size of any time-of-day effect on SET specifically is uncertain and likely smaller than the grade effect.
  • The direction is not universal. Pope finds mornings better for schoolchildren; Dills and Hernandez-Julian find afternoons better for university students once fatigue and schedule density are accounted for. Chronotype, age, subject and cumulative load all moderate the effect. There is no single adjustment that applies everywhere.
  • Selection is only partly handled. Even within-student designs cannot fully rule out that students exert different effort in slots they did or did not prefer. In real timetables, slot assignment is quasi-random at best.
  • Small magnitudes. Grade effects are on the order of a few hundredths of a grade point per hour; the knock-on effect on a 5-point SET scale is plausibly a fraction of a point. It matters for fine-grained ranking, less so for coarse "meets expectations" judgements.

Honest framing: time of day is a real but modest confound whose sign depends on context. Treat it as a reason to avoid over-precise cross-instructor comparison, not as a coefficient to subtract.

How Koji incorporates this

Koji, as an AI-native evaluation platform, addresses the timetable confound through context capture, conversational probing and fatigue-aware collection design.

Record the slot as structured metadata. Timetable slot, day pattern, and session position in the day can be attached to each Koji collection via single_choice fields or import metadata. QA officers can then segment - comparing an instructor to peers teaching the same slot type rather than to a pooled average that mixes prime and unpopular times. The confound becomes a visible column, not an invisible penalty.

Distinguish clock complaints from teaching complaints. Because Koji runs an AI-moderated conversational interview rather than a static form, a low rating can be gently interrogated: "What shaped your experience of this session most?" A tired 8am respondent who says "honestly, it is just too early and I am exhausted" is giving construct-irrelevant feedback, which Koji's automatic thematic analysis tags separately from teaching-quality themes. The report then shows a dean how much of a low score is schedule-driven.

Fatigue-aware, formative timing. Koji supports mid-cycle and formative collection, so feedback need not be gathered only at the exhausted end of a long day or term. Shorter, well-timed conversational check-ins reduce the mood and fatigue contamination that a single end-of-term survey concentrates. Triangulating across cohorts and terms also helps separate a persistent teaching signal from a one-off slot effect.

Koji is deliberately conservative here: it surfaces the timetable as a candidate explanation for a score gap and leaves the judgement to humans, rather than silently re-scoring anyone. The same conversational engine powers Koji's core research platform at koji.so for product and customer research, where controlling for context around a respondent is equally central.

A practical protocol for quality-assurance teams

The scheduling confound is best managed at the point of interpretation rather than through statistical correction:

  1. Attach the slot to every collection. Record start time, day pattern and the session's position in the student's day. A slot that is invisible on the report cannot be accounted for.
  2. Segment before you compare. When comparing colleagues who teach the same module, check whether one is systematically in early-morning or end-of-week slots. If so, compare within slot type or annotate the difference explicitly.
  3. Watch rescheduling in trend data. If a module moves slot between years, flag it so a score change is not misread as a teaching change in a self-evaluation narrative.
  4. Rotate unpopular slots. Where feasible, share the burden of 8am and Friday-afternoon teaching across a team, so no individual's evaluation record is permanently penalised by the timetable.

Moderators: chronotype, age and blended delivery

The size and even the direction of time-of-day effects depend on the population. Adolescents (the subjects in Pope, 2016) skew toward later chronotypes, which is part of why morning under-performance is pronounced in schools; university cohorts are more heterogeneous, and cumulative daily load - not clock time alone - drives the afternoon advantage Dills and Hernandez-Julian (2008) observed. Blended and recorded delivery complicate the picture further: when lectures are recorded, the live-session slot selects a particular sub-group of attenders, entangling the scheduling confound with an attendance and self-selection effect. The honest conclusion for a European QA office is not to hunt for a universal correction coefficient, but to keep the slot visible, avoid over-precise cross-instructor ranking, and let qualitative feedback reveal when the clock rather than the teaching is doing the work.

Related resources

References