Evaluating Team-Taught Courses: The Attribution Problem in Student Evaluations
Why a single overall rating conflates co-teachers, and how to design evaluations that separate instructor-level from course-level signal in team-taught courses.
Koji Education Team
Product
BLUF — Bottom line up front In a team-taught or co-taught course, a single overall rating conflates the instructors: the halo from one strong or weak teacher bleeds across the whole team, and students cannot reliably attribute their experience to the right person versus the course design. The direct empirical literature on team-taught SET is comparatively thin, but the well-established halo, dimensionality, and cross-classified-variance literatures converge on one conclusion — attribution must be designed into the instrument (per-instructor blocks, segment-tagged open text), because it cannot be recovered from a global rating after the fact.
The problem: one number, several teachers
Team teaching is now routine in European higher education — interdisciplinary modules, clinical rotations, studio critiques co-led by two or three faculty, first-year seminars with rotating specialists, and graduate courses where a course leader delivers most sessions while co-instructors contribute a handful. Kathryn Plank's edited volume Team Teaching: Across the Disciplines, Across the Academy (2011) documents this diversity of models across the sciences, social sciences, humanities, and arts, from first-year seminars to graduate courses, and treats team teaching as a deliberate pedagogical choice rather than an administrative accident.
Yet the instruments used to evaluate these courses were, for the most part, designed for one instructor. A standard student evaluation of teaching (SET) form asks for a global impression — "Overall, this instructor was effective" — and a handful of dimension items. When three people share the teaching, that single overall number has to stand in for all of them at once. This is the attribution problem: the evaluation produces one signal where the decision it informs (feedback, renewal, promotion) needs several. Worse, the ways that signal gets corrupted are systematic and well documented, so the errors are not random noise that averages out. They are directional.
What the research says
The honest starting point is that direct empirical work isolating SET behaviour in team-taught courses is comparatively thin. There is rich work on team teaching as pedagogy and on co-teaching outcomes, and there is a deep, rigorous literature on SET psychometrics in general. The defensible case about attribution is built by combining them.
SET is multidimensional, and mostly about the instructor — in single-instructor settings. Herbert Marsh's landmark review, Students' Evaluations of University Teaching: Dimensionality, Reliability, Validity, Potential Biases, and Utility (Marsh, 1984, Journal of Educational Psychology, 76(5), 707–754), synthesised his and others' research to conclude that class-average student ratings are (a) multidimensional, (b) reliable and stable, (c) primarily a function of the instructor who teaches a course rather than the course that is taught, (d) relatively valid against indicators of effective teaching, and (e) relatively unaffected by many hypothesised biases. Point (c) is the crux for our problem. In a single-instructor course it is reassuring: the rating tracks the teacher. In a team-taught course it is the source of the trouble, because "the instructor" is now several people, and the instrument was built assuming there was one. The instructor-level signal Marsh identified is real, but on a shared form it is smeared across the whole team rather than cleanly assigned.
The halo effect contaminates ratings — and on a shared form it contaminates across people. Thomas Hugh Feeley's study Evidence of Halo Effects in Student Evaluations of Communication Instruction (Feeley, 2002, Communication Education, 51(3), 225–236) had 128 students rate an instructor on nonverbal immediacy, teaching effectiveness, and attitude toward content, plus two items chosen precisely because they are irrelevant to teaching effectiveness: vocal clarity and physical attractiveness. All five measures were significantly inter-correlated (roughly r = 0.28 to 0.72), including the irrelevant ones — the signature of a halo effect, where raters fail to discriminate among conceptually distinct aspects of a target. The classic halo operates across items for one instructor. The team-teaching hazard is that the same failure to discriminate operates across people on one form: a strong overall impression of the charismatic lead lecturer pulls the quieter co-instructor's ratings up (undeserved credit), and a single frustrating co-teacher can drag the whole team's numbers down (undeserved blame). Either way, the weaker signal is the co-teacher whose individual contribution the reader most needs to see clearly.
The data structure is cross-classified, and the variance you want is instructor-level. Pieter Spooren's On the credibility of the judge: A cross-classified multilevel analysis on students' evaluation of teaching (Spooren, 2010, Studies in Educational Evaluation, 36(4), 121–131) makes the structural point explicit. Students evaluate several courses and teachers, teachers teach several courses, and courses are taught by different teachers — a cross-classified structure, not a simple hierarchy. Modelling it properly partitions SET variance into course, teacher, student, and student-by-teacher interaction components. The methodological lesson transfers directly: the instructor-level variance is a distinct, estimable quantity — but only if the instrument records which instructor each response refers to. A global rating on a team-taught course collapses the teacher facet into the course facet and throws away exactly the component a personnel or QA decision depends on. You cannot un-mix it afterward.
Where direct team-taught evidence does exist, it is cautious. Zach and Avugos, Co-teaching in higher education: implications for teaching, learning, engagement, and satisfaction (2024, Frontiers in Sports and Active Living, DOI 10.3389/fspor.2024.1424101), ran an action-research study of 50 undergraduate student-teachers across two co-taught seminars. Students reported emotional, social, and cognitive gains and a marginal preference for co-teaching over traditional instruction (means of about 4.65 vs 4.25). Notably, the authors found "no indications of higher quality" against traditional instruction and flagged the heavy instructor time and resource cost. The relevance here is twofold: co-teaching is a genuine, non-trivial format worth evaluating well, and even sympathetic empirical work does not hand us a validated method for cleanly rating the individuals within the team. That gap is precisely why the design-side argument carries the weight.
Why it matters in practice
Personnel decisions. SET data feed renewal, tenure, and promotion. If a junior co-instructor's individual effectiveness is invisible inside a team's global rating, they are judged on a number that is mostly about someone else. A halo from a strong lead can flatter a weak contributor; a halo from a weak one can damage a strong contributor. Both are unfair, and both are hard to appeal because the number looks objective.
Fairness between co-teachers. Exposure is not effectiveness. A co-instructor who teaches three of thirty sessions will be less salient in students' memory than the leader who runs the rest — so a shared form quietly rewards contact hours over teaching quality. A common institutional guardrail is to only evaluate faculty who teach beyond a threshold of the course (for example, at least around 20% or a set number of hours), and to state each instructor's role on the form so students rate only those they meaningfully interacted with.
QA and course improvement. A single number tells a programme director the course "scored 4.1" but not whether the weak spot was the second module's instructor, the assessment design, or the hand-off between co-teachers. Attribution is what turns an evaluation from a grade into a diagnosis you can act on.
Limitations and honest caveats
Intellectual honesty here matters more than a clean story.
- The direct evidence base is thin. There is no large, canonical body of studies establishing how standard SET instruments misbehave specifically in team-taught settings. The argument rests on transferring robust general findings (halo, dimensionality, cross-classified variance) to the team-taught case. That transfer is well motivated but is an inference, not a replicated team-taught result.
- Confounds are real. Exposure, session order, topic difficulty, and which instructor handled assessment all correlate with ratings and with each other. Even a per-instructor form does not fully disentangle "this teacher" from "the hard module this teacher happened to own."
- Generalisability. Marsh's and Feeley's work is largely North American; co-teaching norms, class sizes, and evaluation regimes differ across European systems. Zach and Avugos studied a specific sports-science seminar with 50 students. None of these should be over-extrapolated.
- The core limit is conceptual. A global rating is a single holistic judgment. There is no statistical operation that recovers separable per-instructor components from a number that never contained them. This is not a gap that better modelling closes after collection — it is a gap that only instrument design closes before collection.
How Koji incorporates this
Koji for Education is designed to mitigate the attribution problem, not to eliminate it — the caveats above still bind, and no tool recovers attribution from a legacy single global rating. What Koji can do is build attribution into the instrument so the instructor-level signal exists in the first place.
- Structured per-instructor blocks. Koji supports
open_ended,scale,single_choice,multiple_choice,ranking, andyes_noquestion types. A team-taught evaluation can therefore repeat a compact block per named instructor or per teaching segment, so each co-teacher gets their ownscaleandopen_endedresponses rather than sharing one global item — the design-side fix the literature points to. - AI-moderated conversational interviews that probe who. Rather than a static form, Koji's AI-moderated interview can ask a follow-up when a comment is ambiguous — "Which part of the course, or which instructor, are you describing here?" — capturing the attribution at the moment the student still remembers it, instead of leaving it to be guessed later.
- Thematic analysis attributed to named segments. Koji's automatic thematic analysis can tag open-text themes to specific teaching segments or instructors, so "the second module felt rushed" is filed against the right co-teacher rather than absorbed into a course-level average.
- Bias-aware reporting and multilevel-friendly export. Reporting is framed to flag halo-prone patterns (for example, near-identical ratings across all co-teachers), and data export is structured so an institutional-research team can fit cross-classified or multilevel models in the spirit of Spooren (2010), keeping course-level and instructor-level variance distinct.
- Closing the loop. Action tracking lets a programme direct a specific finding to the specific instructor or design element it concerns — the diagnostic payoff that a single global number cannot provide.
As a secondary note, the same AI-moderated interview engine underpins Koji's core research platform at koji.so, where it runs product and customer-research interviews — the same technique of probing which thing a respondent means, applied outside education.
The honest summary: Koji does not make a hard attribution problem disappear. It changes when attribution happens — from an impossible after-the-fact disaggregation to a designed-in property of the instrument — which is the only point in the pipeline where the evidence says the problem can actually be addressed.
References
- Feeley, T. H. (2002). Evidence of halo effects in student evaluations of communication instruction. Communication Education, 51(3), 225–236. https://doi.org/10.1080/03634520216519
- Marsh, H. W. (1984). Students' evaluations of university teaching: Dimensionality, reliability, validity, potential biases, and utility. Journal of Educational Psychology, 76(5), 707–754. https://doi.org/10.1037/0022-0663.76.5.707
- Plank, K. M. (Ed.). (2011). Team teaching: Across the disciplines, across the academy. Sterling, VA: Stylus Publishing. ISBN 9781579224547.
- Spooren, P. (2010). On the credibility of the judge: A cross-classified multilevel analysis on students' evaluation of teaching. Studies in Educational Evaluation, 36(4), 121–131. https://doi.org/10.1016/j.stueduc.2011.02.001
- Zach, S., & Avugos, S. (2024). Co-teaching in higher education: Implications for teaching, learning, engagement, and satisfaction. Frontiers in Sports and Active Living, 6, 1424101. https://doi.org/10.3389/fspor.2024.1424101
Related Resources
- What student evaluations actually measure: Marsh and multidimensionality
- Instructor immediacy, warmth, and the halo in course evaluations
- The halo effect in student evaluations of teaching
- Multilevel models for course evaluation and nested data
- Dimensionality of student ratings: global versus profile (d'Apollonia and Abrami)
Related articles
Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages for Course Evaluation
Course-evaluation data has a nested structure — students within sections within instructors — and a flat average ignores it. This guide explains how multilevel (hierarchical linear) models partition variance, why ignoring clustering understates uncertainty, and what it means for fair reporting.
The Halo Effect in Course Evaluations: When One Impression Colours Every Rating
When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.
One Number or Many? The Dimensionality Debate and How to Use Student Ratings for Personnel Decisions
Should a promotion committee use a single global teaching score or a detailed profile of many dimensions? The 1997 d'Apollonia & Abrami vs Marsh & Roche debate set the terms — and the answer depends on whether the purpose is summative judgement or formative improvement.