Does the Room Bias the Rating? Physical Classroom Environment as a Confound in Course Evaluations
Evidence that the physical classroom - lighting, seating, comfort, technology - shifts student evaluation of teaching scores independent of instructional quality, and how to stop the estate from contaminating the teaching signal.
Koji Education Team
Product
In brief: Yes - the physical room a class is taught in can measurably move student evaluation of teaching (SET) scores, independent of how well the instructor actually teaches. In a controlled university study, Hill and Epps (2010) found that students taught in upgraded classrooms rated their instructors significantly higher on organisation and on learning new things than comparable students taught in standard rooms. For quality assurance this is a textbook source of construct-irrelevant variance: part of the number on the report reflects the estate, not the educator.
What the research says
The clearest single-study evidence comes from Mary C. Hill and Kathryn K. Epps (2010), who surveyed 237 undergraduate business students and compared satisfaction and teaching-evaluation ratings across "standard" and "upgraded" classrooms at the same institution. Upgraded rooms differed on concrete physical attributes - tiered seating, better lighting, more comfortable and ergonomic furniture, and improved instructional technology. Students in the upgraded rooms reported higher satisfaction with the classroom itself, which was unsurprising, but they also rated their instructors more highly - notably on perceived organisation and on the sense that they had learned new things. Because students did not choose their rooms on the basis of the instructor, the design isolates a room effect that bleeds into the teaching score.
This does not sit in isolation. A broad literature on the built learning environment documents that physical conditions - temperature, air quality, acoustics, natural light, layout and seating flexibility - affect student attention, comfort, wellbeing and self-reported experience. Systematic reviews of indoor environmental conditions in higher education (for example, work synthesised in the classroom-environment and student-wellbeing literature) consistently link poor thermal comfort, noise and inadequate lighting to worse concentration and lower satisfaction. Since global SET items ("Overall, this was an excellent course") are strongly coloured by affect and satisfaction, any environmental factor that depresses mood or comfort has a plausible pathway into the rating.
The mechanism matters. SET scores are not a clean readout of instructional effectiveness in the first place: the meta-analysis by Uttl, White and Gonzalez (2017) found that SET ratings explain at most about 1% of the variance in objective measures of student learning across multisection studies. When the "signal" (teaching-to-learning) is that weak, "noise" from context - the room, the timetable slot, class size - occupies a proportionally larger share of the score. A room effect that adds even a fifth of a scale point is not trivial when instructors are ranked and compared to two-decimal precision.
Why it matters for course evaluation in practice
For a quality-assurance office, the practical problem is comparability. If Dr A teaches in a refurbished, tiered lecture theatre with good AV and Dr B teaches the parallel section in a cramped, poorly ventilated room in an older building, a raw comparison of their mean scores is partly a comparison of the estate. Three concrete consequences follow:
- Unfair personnel inferences. Where SET feeds into probation, promotion or workload allocation, systematic room advantages become systematic career advantages that have nothing to do with teaching. This is the same fairness problem that arises with class size and discipline, but it is easier to overlook because room allocation is usually invisible on the evaluation report.
- Misdirected improvement effort. An instructor scoring low partly because of a hot, noisy room may be pushed into teaching-development activity that cannot fix the underlying cause. Meanwhile the actionable fix - a better room - sits with facilities, not the lecturer.
- Confounded trend data. If a programme is relocated to a new building, a jump in evaluation scores may be read as a teaching improvement when it is really an estate improvement. Longitudinal quality narratives to accreditors (ENQA/ESG-aligned self-evaluation reports) can be quietly distorted.
The correct response is not to dismiss SET, but to treat the room as a covariate to be recorded and reported alongside the score - exactly as one would record class size or delivery mode.
Limitations and honest caveats
A critical reader should hold this evidence to account:
- Effect sizes are modest and study-specific. Hill and Epps (2010) used a single institution, a single discipline (business), and a modest sample. The direction of the effect is well motivated theoretically, but the magnitude will vary and should not be over-generalised.
- Self-selection and confounding are hard to rule out. Upgraded rooms are not randomly assigned in most institutions; they may house particular course types, cohorts or timetable slots. The cleanest inference would come from a design that randomises rooms holding instructor and cohort fixed - rare in practice.
- Reverse and indirect pathways. A better room may genuinely improve teaching (more flexible layouts enable active learning), in which case part of the "room effect" on ratings is a real teaching effect, not pure contamination. Disentangling construct-irrelevant variance from a legitimate environment-enabled pedagogy change requires care.
- Publication and vintage. Some of the corroborating environment-and-wellbeing literature measures satisfaction or achievement rather than SET specifically, so it supports the plausibility of the pathway more than the size of the SET effect.
None of these caveats overturn the core point - environment is a live confound - but they should temper any attempt to "adjust away" a precise number.
How Koji incorporates this
Koji is an AI-native course-evaluation platform, and its design gives quality teams two levers against the room-as-confound problem.
Capture the context as structured metadata. Koji evaluations can attach structured fields - room, building, delivery mode, timetable slot, class size - to each collection using single_choice and yes_no question types or import-time metadata. That lets a QA officer segment and compare like-for-like (upgraded vs standard rooms) instead of pooling scores that differ on the estate. It turns an invisible confound into a visible, reportable variable.
Probe the reason behind a low number. The core of Koji is an AI-moderated conversational interview, not a bare Likert form. When a student gives a low rating, the moderator can ask a neutral, non-leading follow-up - "What most affected your experience of this session?" - and the student is free to say "the room was freezing and I could not hear the lecturer." Koji's automatic thematic analysis then tags logistics and environment complaints separately from teaching-quality themes. The practical payoff: a dean reading the report sees that three of the five low scores were driven by a construct-irrelevant estate problem, and routes it to facilities rather than to the instructor. This is designed to mitigate - not eliminate - the contamination, by making its share of the score legible.
Bias-aware, triangulated reporting. Because Koji separates the numeric global rating from the thematic open-text signal, and can triangulate across cohorts and terms, a spurious "improvement" caused by a building move is easier to spot: the numbers rise but the teaching-related themes do not change. Koji is careful here - it surfaces the possible confound for human judgement rather than silently re-scoring anyone.
The same conversational engine underpins Koji's core research platform at koji.so, where AI-moderated interviews are used for product and customer research; in education the emphasis is on separating what the university can act on (teaching) from what it must escalate (the estate).
A practical protocol for quality-assurance teams
Turning the room-as-confound insight into routine practice does not require a research programme - it requires a few disciplined habits:
- Record the room on every collection. Attach building, room and a simple quality flag (standard / upgraded / problem-flagged by estates) as structured metadata. Without this, the confound is unmeasurable and therefore unmanageable.
- Compare like-for-like before ranking. When two sections of the same module diverge, check the estate first. If one section sat in a flagged room, annotate the comparison rather than presenting the gap as a teaching difference.
- Read the open text before reacting to the number. A dip accompanied by comments about heat, noise or broken AV is an estates ticket, not a teaching-development referral. Route it accordingly and log the action.
- Feed patterns back to timetabling and estates. If a particular room consistently depresses scores across different instructors, that is strong, triangulated evidence for a capital or maintenance case - a genuinely useful by-product of evaluation data.
Why this matters for European institutions specifically
Estate quality is unevenly distributed across and within European universities - historic buildings, rapid post-expansion capacity pressures, and mixed refurbishment cycles mean that two colleagues in the same department can teach in very different physical conditions. Under ESG/ENQA expectations, institutions are asked to assure teaching quality fairly and to base judgements on sound evidence. Presenting unadjusted, room-confounded scores in a self-evaluation report or a promotion case is difficult to defend once the confound is named. Treating the room as a recorded, reportable variable is both a fairness safeguard for staff and a more honest evidence base for reviewers - and it converts a nuisance confound into an actionable estates signal.
Related resources
- /docs/is-student-evaluation-of-teaching-valid-spooren-state-of-the-art
- /docs/what-student-evaluations-measure-marsh-multidimensionality
- /docs/class-size-effects-student-evaluations
- /docs/do-cookies-treats-mood-bias-course-evaluations
- /docs/online-vs-face-to-face-course-modality-student-evaluations
- /docs/interpreting-reporting-student-ratings-responsibly
References
- Hill, M. C., & Epps, K. K. (2010). The impact of physical classroom environment on student satisfaction and student evaluation of teaching in the university environment. Academy of Educational Leadership Journal, 14(4), 65-79. https://digitalcommons.kennesaw.edu/facpubs/1308/
- Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22-42. https://doi.org/10.1016/j.stueduc.2016.08.007
- Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. https://doi.org/10.1037/0003-066X.50.9.741
- Review evidence on the physical learning environment and student wellbeing in higher education (2024), Building and Environment/related syntheses. https://www.sciencedirect.com/science/article/pii/S036013232400800X
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.
Do Cookies, Treats, and Mood Bias Course Evaluations?
Two controlled studies show that giving students chocolate or cookies before an evaluation measurably raises teaching scores. What the affect heuristic means for the validity of course evaluations — and how to design around it.
Is Student Evaluation of Teaching Valid? What the State-of-the-Art Review Concludes
A close reading of Spooren, Brockx and Mortelmans'' 2013 Review of Educational Research synthesis of SET validity — what it actually concludes, the evidence behind it, and what it means for how universities use student ratings.