Does Course Difficulty and Workload Lower Student Evaluations? What Centra Found
A research-grounded look at whether harder, heavier courses are punished in student evaluations of teaching. Centra (2003) and Marsh & Roche (2000) show the relationship is non-linear and weaker than faculty fear — with practical implications for fair course evaluation.
Koji Education Team
Product
In brief: The widespread belief that demanding courses are systematically punished in student evaluations is not well supported by the evidence. Centra (2003), analysing more than 50,000 courses, found that once perceived learning was accounted for, expected grades had little effect on ratings, and that courses students judged "just right" in difficulty and workload — neither too easy nor too hard — received the highest evaluations, not the easiest ones. The relationship between rigour and ratings is curvilinear, not a straight downhill slope. The practical lesson for quality assurance is to stop reading a low score as proof of either bias or weak teaching, and to collect data that reveals why a course was rated as it was.
The question, and why it matters
Few beliefs are more entrenched in faculty common rooms than the idea that the path to good evaluations runs through easy grading and light reading lists. The fear has real consequences: it is invoked to discredit student feedback entirely ("students just reward the soft markers"), and it is used to justify watering down demanding courses. If the belief were true, student evaluations of teaching (SETs) would be close to worthless as a quality signal, because they would simply track leniency. If it is false — or far more nuanced than the folklore suggests — then evaluations retain value, but only if institutions interpret them with care. This article anchors on the most-cited empirical test of the claim and situates it within the wider literature.
What the research says
The anchor study is John A. Centra's 2003 article "Will teachers receive higher student evaluations by giving higher grades and less course work?" (Research in Higher Education, 44(5), 495–518). Centra analysed over 50,000 college courses whose instructors used the Student Instructional Report II, with regression models spanning eight subject-area groupings and controlling for class size, teaching method and — critically — students' perceived learning. Three findings stand out:
- Perceived learning was the strongest predictor of evaluations by a wide margin. Courses where students felt they had learned a great deal were rated highly, whatever the grading.
- After controlling for perceived learning, expected grades had little independent effect. The simple "lenient grading buys good ratings" story did not survive the controls.
- The relationship between difficulty/workload and ratings was curvilinear. Courses students rated as "just right" in difficulty received the highest evaluations. Courses seen as too easy — including, in the natural sciences, those with expected A grades — were rated lower, not higher.
This last point is the crux: students do not uniformly reward ease. They appear to reward an appropriate level of challenge, and they penalise courses that feel trivial as well as those that feel overwhelming.
Centra's findings align closely with the second major source, Herbert W. Marsh and Lawrence A. Roche's 2000 study, "Effects of grading leniency and low workload on students' evaluations of teaching: Popular myth, bias, validity, or innocent bystanders?" (Journal of Educational Psychology, 92(1), 202–228). Across two large studies, Marsh and Roche found that the relationship between grades and SETs is non-linear and modest, that higher workload was often associated with higher (not lower) ratings, and that the patterns were stable across a 12-year window — contrary to predictions that grade inflation would steadily corrupt the ratings. They concluded that grading leniency and workload are better understood as minor influences than as pervasive biases that invalidate SETs.
A third strand of evidence complicates the picture and keeps us honest. Greenwald and Gillmore (1997) argued, using within-instructor data, that expected grades do exert an upward pull on ratings that simple cross-sectional studies under-estimate, and proposed statistically adjusting SETs for grade expectations. The debate between the Marsh/Centra "myth" camp and the Greenwald "real but partial bias" camp has never fully resolved — which is precisely why a careful evaluation programme should neither dismiss grading effects nor treat them as the whole story.
Why it matters for course evaluation in practice
For a quality-assurance officer, the headline is liberating but demanding. It is liberating because it removes the excuse to ignore evaluations: the data are not simply a popularity contest rigged toward soft markers. It is demanding because it means a single mean score is genuinely ambiguous. A high rating could reflect excellent teaching of an appropriately challenging course — exactly what you want — or it could, in a minority of cases, reflect a course that asks too little. A low rating could signal real teaching problems, or a course pitched too hard, or a cohort that resented a fair-but-heavy workload. The number alone cannot tell you which.
The Centra curvilinear finding has a direct operational implication: you must measure perceived difficulty and workload separately, and you must capture the "why." Knowing that a course scored 3.8 is far less useful than knowing whether students found it "about right," "too demanding given the credit weighting," or "too easy to take seriously." Programmes that collect only an overall satisfaction number throw away exactly the information needed to interpret it.
Limitations and honest caveats
A PhD reader will rightly push back, and the honest answer is that the evidence is suggestive, not conclusive.
- Correlational design. Centra's data are observational. "Perceived learning" is itself a student perception, not an objective measure, and it may be endogenous — students who like a course may report more learning. Controlling for it could partly control away the very leniency effect under investigation.
- Self-report of difficulty. Perceived workload is not actual workload; students systematically mis-estimate hours, and the perception may be coloured by how engaging the course felt.
- Generalisability. Both Centra and Marsh & Roche draw heavily on North American institutions and instruments. European credit systems (ECTS), grading cultures and student expectations differ, so effect sizes should not be assumed to transfer unchanged.
- The Greenwald counter-position. Within-instructor and experimental designs can surface grade effects that cross-sectional studies miss. The "myth" framing may understate a real, if modest, bias.
- Publication era. These are mature studies; grade inflation and student consumerism have intensified since, and replication in contemporary cohorts is warranted.
None of these caveats overturns the central, robust pattern — ease does not straightforwardly buy good ratings, and challenge is not straightforwardly punished — but they should temper any claim that difficulty is irrelevant.
How Koji incorporates this
Koji is designed to address the exact interpretive gap this research exposes: the ambiguity of a lone difficulty-blind average. Rather than asking only "how would you rate this course," Koji's AI-moderated conversational interview can probe the reason behind a rating in the moment — following up on a low score to distinguish "the workload was unreasonable for the credits" from "the lectures were hard to follow" from "the material was too easy to stay engaged." This is the difference between a number and a diagnosis.
Concretely, Koji supports:
- Structured difficulty and workload items (
scaleandsingle_choice) that capture perceived challenge separately from overall satisfaction, so the curvilinear "just right" signal Centra identified is actually measurable in your data rather than collapsed into one mean. - Adaptive open-ended probing, where the interview asks why a course felt too hard or too easy, surfacing whether a low score reflects pitch, pacing, assessment design or teaching — the attribution problem at the heart of the difficulty debate.
- Automatic thematic analysis of those open responses, clustering comments into themes such as "workload vs. credit mismatch" or "insufficient challenge," so a programme director sees patterns rather than anecdotes.
- Bias-aware reporting that presents difficulty and workload distributions alongside satisfaction, discouraging the naive reading of a low average as either bias or teaching failure.
We are careful not to overclaim: Koji is designed to mitigate the misinterpretation of difficulty-driven scores; it cannot eliminate the underlying tension between rigour and student comfort, nor settle the Greenwald–Marsh debate. What it can do is give your quality cycle the structured and qualitative evidence to interpret a score in context. Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where the same "what's behind this rating?" problem recurs.
A practical reading rule for difficulty-aware interpretation
A simple heuristic operationalises the Centra finding for everyday quality work. When you encounter a course with a below-benchmark overall score, do not treat the number as a verdict; treat it as a prompt to ask three diagnostic questions before any judgement is recorded. First, where on the difficulty curve does the cohort place this course — too easy, just right, or too demanding for its credit weighting? Second, what does perceived learning look like, since Centra showed it is the dominant driver and a course can be hard, lower-rated and still highly effective? Third, do the open-text reasons attribute the score to teaching or to structural pitch? Only when those three are answered does the score acquire meaning. A programme that institutionalises this three-question rule converts an ambiguous average into actionable evidence, and protects instructors of demanding but well-taught courses from being penalised for refusing to dilute their content — exactly the perverse incentive the difficulty literature warns against.
Related Resources
- Does Grading Leniency Inflate Student Evaluations?
- Class Size and Student Evaluations: What Bedard and Kuhn Found
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
- The Halo Effect in Course Evaluations
- Selection Bias in Course Evaluations
- Why Averaging Likert Scores Misleads Course Evaluation
References
- Centra, J. A. (2003). Will teachers receive higher student evaluations by giving higher grades and less course work? Research in Higher Education, 44(5), 495–518. https://doi.org/10.1023/A:1025492407752
- Marsh, H. W., & Roche, L. A. (2000). Effects of grading leniency and low workload on students' evaluations of teaching: Popular myth, bias, validity, or innocent bystanders? Journal of Educational Psychology, 92(1), 202–228. https://doi.org/10.1037/0022-0663.92.1.202
- Greenwald, A. G., & Gillmore, G. M. (1997). Grading leniency is a removable contaminant of student ratings. American Psychologist, 52(11), 1209–1217. https://doi.org/10.1037/0003-066X.52.11.1209
Related articles
The Halo Effect in Course Evaluations: When One Impression Colours Every Rating
When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.
Class Size and Student Evaluations: What Bedard and Kuhn Found
Does class size bias student evaluations of teaching? Bedard and Kuhn (2008) found a large, non-linear negative effect of enrolment on instructor ratings even after controlling for instructor and course. This article synthesises the evidence and explains how to stop class size from contaminating cross-instructor comparisons.
Does Grading Leniency Inflate Student Evaluations? What the Evidence Actually Shows
A research-grounded look at the grading-leniency hypothesis in student evaluations of teaching: what Marsh and Roche, Greenwald and Gillmore, and Stroebe found, why the grades-ratings correlation is genuinely ambiguous, and how to design evaluation so leniency cannot quietly buy better scores.