New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

Should You Statistically Adjust Course Evaluations for Class Size and Difficulty? The IDEA Approach to Adjusted Scores

Some course-evaluation systems report "adjusted" scores that statistically correct for class size, discipline difficulty and student motivation. We examine what the IDEA system actually adjusts for, whether the practice is defensible, and how to contextualise scores without over-correcting.

Koji Education Team

Product

In brief

Course-evaluation scores are systematically nudged up or down by factors an instructor does not control — most reliably class size, disciplinary difficulty, and students' prior motivation to take the course. The IDEA system (Benton & Cashin, 2012; Hoyt & Lee, 2003) responds by reporting adjusted scores that statistically partial out five such "extraneous" variables so that instructors are compared on a fairer basis. The evidence supports contextualising ratings before you compare them, but statistical adjustment is a modelling choice with real limitations: it can over-correct, it depends on the specific variables measured, and it should never convert a noisy rating into a false impression of precision. Adjust to inform a conversation, not to manufacture a ranking.


What the research says

For decades the dominant worry about Student Evaluation of Teaching (SET) has been that scores reflect things other than teaching. Herbert Marsh's landmark synthesis concluded that class-average ratings are multidimensional, reliable, and "relatively unaffected by various biases" — but "relatively" is doing heavy lifting (Marsh, 1987). A handful of course and student characteristics move ratings enough to matter when scores are used comparatively.

The IDEA Center built an entire reporting system around this problem. Rather than pretending the confounds do not exist, IDEA measures them and reports an adjusted score alongside the raw one. Benton and Cashin's summary of the literature (IDEA Paper No. 50) lays out the rationale: the purpose of adjustment is "to separate the contributions of the teacher from the contributions of extraneous factors to student learning" (Benton & Cashin, 2012). Hoyt and Lee's technical report specifies the five variables the IDEA Diagnostic Form adjusts for (Hoyt & Lee, 2003):

  1. Student motivation — how much students wanted to take the course irrespective of who taught it. Motivated students learn more and rate higher, and this has nothing to do with the instructor.
  2. Student work habits — a general disposition toward academic effort that students bring with them.
  3. Course difficulty of subject matter — the intrinsic difficulty of the material, distinct from the instructor's demands.
  4. Class size — larger classes tend to earn slightly lower ratings on many items.
  5. Student effort — the effort students report investing, which partly reflects the course but also the cohort.

The empirical justification for each is drawn from the wider SET literature. Feldman's meta-analytic work established that elective vs required status, prior subject interest, and class size all correlate modestly with ratings (Feldman, 1978). The class-size relationship is small but consistent — a slight tendency for large-lecture instructors to be marked down, which IDEA treats as a minor adjustment. The disciplinary difference is larger: quantitative and "hard science" courses systematically receive lower ratings than social-science and humanities courses (a pattern later reinforced by work such as Uttl & Smibert, 2017, on quantitative courses). IDEA therefore also reports discipline-referenced comparisons so a mathematics instructor is not measured against a seminar leader.

The key methodological move is a regression adjustment: IDEA models the raw progress rating as a function of the extraneous variables using a large multi-institution database, then reports what a class would have scored had it been of average motivation, difficulty and size. This is conceptually the same idea as the "value-added" and shrinkage approaches used elsewhere in the measurement literature — you remove predictable, non-teaching variance so the residual is a cleaner signal of the instructor.


Why it matters for course evaluation in practice

If your quality-assurance office compares raw means across a faculty, you are almost certainly penalising the people who teach large, required, quantitative modules and rewarding those who teach small, elective, discursive ones. That is not a hypothesis — it is the most replicated finding in the SET literature. Three practical consequences follow.

First, raw cross-instructor rankings are indefensible. A 3.9 in a 300-person required statistics course and a 4.4 in a 15-person elective seminar are not comparable numbers. Any personnel or programme-review process that treats them as such is, in effect, measuring the teaching assignment rather than the teaching.

Second, adjustment changes who looks "good." IDEA's own reports routinely show instructors whose raw scores are unremarkable but whose adjusted scores are strong once the difficulty of their assignment is accounted for — and vice versa. For a QA officer, the adjusted score is often the more decision-relevant number.

Third, adjustment is a communication tool, not just a statistic. Showing an instructor "your raw score is 3.8, but for a course of this size and difficulty the expected score is 3.6, so you are performing above expectation" reframes a demoralising number into an actionable one. This is exactly the kind of responsible, context-first reporting that Linse (2017) urges evaluation committees to adopt.


Limitations and honest caveats

A PhD reader will immediately raise objections, and they are right to.

Adjustment can over-correct. If a genuinely effective instructor produces high student motivation and effort through their teaching, then partialling out motivation and effort removes real teaching signal along with the noise. Marsh's construct-validity work warns that some "biasing" variables are partly outcomes of good teaching, not just confounds. The direction of causality is not clean.

The model is only as good as its variables. IDEA adjusts for five things. It does not adjust for instructor gender, race, accent, age, or the dozens of other documented influences on ratings. An adjusted score is not a bias-free score; it is a score corrected for five specific, measurable factors and nothing else.

Generalisability. The regression weights come from IDEA's institutional database, which is predominantly North American. Applying those weights to a European institution, a different disciplinary mix, or a different rating instrument imports assumptions that may not hold. Cross-cultural response styles alone can distort the calibration.

False precision. Adjustment produces a tidy decimal that invites over-interpretation. But the underlying ratings still carry sampling error, and the difference between an adjusted 3.8 and 4.0 is very often noise. Adjustment does not rescue a score built on a 25% response rate.

Transparency and trust. Instructors may reasonably distrust a "black box" that changes their number via an opaque regression. If you adjust, you must be able to explain why and how — otherwise adjustment corrodes the legitimacy of the whole exercise.

The honest position: adjustment is a defensible way to make comparisons less unfair, not a way to make them fair. It narrows the gap between the number and the truth; it does not close it.


How Koji incorporates this

Koji for Education is built on the premise that a single adjusted number, however well-modelled, cannot carry the weight institutions put on it — so the platform is designed to mitigate the confounding problem through triangulation and context rather than a lone statistic.

  • Context is captured, not assumed. Koji records the structural metadata that drives the confounds — class size, module type, level, discipline, and delivery mode — and surfaces them alongside every score, so a QA officer never compares a large required module to a small elective without seeing the difference. This is the same logic as IDEA's adjustment, applied at the reporting layer through bias-aware reporting rather than a hidden regression.
  • Segmentation over global averages. Because Koji supports structured questions (scale, single_choice, multiple_choice, ranking, yes_no) tied to respondent metadata, cohorts can be compared like-with-like — motivated vs less-motivated intakes, on-campus vs distance — instead of collapsing everything into one adjusted mean.
  • The "why" behind the number. Where adjustment tries to remove the effect of motivation or difficulty statistically, Koji's AI-moderated conversational interviews probe it directly: when a student rates a hard module lower, the follow-up asks whether the difficulty was appropriate, poorly scaffolded, or genuinely a teaching failure. That distinction — which no regression can recover — is exactly what separates a confound from a criticism.
  • Automatic thematic analysis of the open-text then aggregates those probes, so a low score on a difficult course arrives with evidence about whether the difficulty was the problem.

Koji does not claim to eliminate confounding — it is explicitly designed to contextualise scores and to add the qualitative signal that statistical adjustment throws away. Institutions that run broader customer or product research often pair this with Koji's core research platform at koji.so, which applies the same AI-moderated interview engine outside the classroom.


Related resources

References

Related articles

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

analysis-reporting

Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores

Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.

analysis-reporting

Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages for Course Evaluation

Course-evaluation data has a nested structure — students within sections within instructors — and a flat average ignores it. This guide explains how multilevel (hierarchical linear) models partition variance, why ignoring clustering understates uncertainty, and what it means for fair reporting.

analysis-reporting

Should You Report an Instructor''s Percentile? Norm-Referenced vs Criterion-Referenced Course-Evaluation Scores

Telling a lecturer they are "in the 40th percentile of the department" is norm-referenced reporting — and it manufactures losers by construction, no matter how good everyone is. Criterion-referenced reporting asks instead whether teaching met a defined standard. Here is the evidence on why the choice matters and how to report responsibly.