New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology8 min read

How Do You Evaluate a Team-Taught Course? The Attribution Problem in Multi-Instructor Modules

Standard evaluation instruments assume one course, one instructor. Team-taught and co-taught modules break that assumption — and the single averaged score they produce is often uninterpretable. Here is why, and what a defensible alternative looks like.

Koji Education Team

Product ·

The short answer: Most course-evaluation instruments are built on a hidden assumption — one course, one instructor — and team-taught modules violate it. When eight academics share a module and students rate "the teaching" on a single five-point scale, the resulting average blends contributions that should be read separately, masks weak and strong segments alike, and can unfairly penalise junior staff. The fix is not a better average; it is evaluation designed to attribute feedback to the right component, person, or teaching block — something a conversational, adaptive interview does far more naturally than a static form.

The assumption nobody states

Open almost any standard student-evaluation-of-teaching (SET) form and you will find items like "The instructor explained concepts clearly" or "The lecturer was well prepared." Singular. The instrument, the reporting, and the personnel processes downstream all assume a single owner of the course. University systems are, as the co-teaching literature puts it, "structured around the assumption of single instructors."

Team-taught, co-taught, and integrated modules are now common — in medicine and the health sciences, in interdisciplinary and STEM programmes, in capstones and problem-based curricula. And on these courses the single-owner assumption quietly fails. The student is asked one question about "the teaching" when there were five teachers, three teaching styles, and two segments that worked beautifully against one that did not.

Why the averaged score misleads

There are at least four distinct problems with collapsing a multi-instructor module into one mean.

1. Aggregation hides the signal. If weeks one to four were excellent and weeks five to eight were a mess, a 3.4 average tells the programme director nothing actionable. It is the methodological cousin of averaging Likert scores in general — except here the averaging spans genuinely different teaching experiences, not just an ordinal scale. The number is not just imprecise; it is pointed at the wrong unit of analysis.

2. Comparative contamination. When students rate co-teachers side by side, ratings become relative. The evidence and the lived experience both suggest a novice instructor paired with a celebrated senior colleague can receive lower ratings than they would have teaching alone, simply because students grade on the contrast. The score measures the pairing, not the person.

3. Attribution error. Students frequently cannot reliably remember which instructor taught which segment, especially weeks later at end of term — so even a well-intentioned "rate Dr X" item collects noise. Memory and the halo effect blur the boundaries the report pretends are crisp.

4. High-stakes misuse. None of this would matter much if the scores stayed formative. But course and instructor ratings feed merit, promotion, and tenure decisions, and the impact of co-teaching on those ratings has barely been studied. A 2023 analysis of co-teaching in undergraduate STEM noted that the effect of co-teaching on student evaluations "has not been explored," even though those ratings "may impact reward decisions" (Haag et al., 2023, CBE—Life Sciences Education). We are making consequential judgments on an instrument we know is mis-specified for these courses.

What the research recommends

The teaching-and-learning literature has converged on a clear principle: multi-instructor courses need component-level, not just course-level, evaluation. Work on integrated medical curricula argues explicitly that individual class evaluation is required to manage multi-instructor courses, because a single global rating cannot capture effective-teaching characteristics that vary block by block (Hwang et al., 2017, BMC Medical Education). And qualitative work comparing team-taught and individually taught modules finds that students experience them differently — valuing the diversity of perspectives but also reporting fragmentation and inconsistency that a single satisfaction score erases (qualitative study of team-taught versus individually taught undergraduate modules, Higher Education, 2016).

This also connects to a foundational point about what evaluations measure at all. Marsh's decades of work on the multidimensionality of student ratings (the SEEQ instrument) established that "teaching" is not one thing but several distinguishable factors. Team teaching simply makes that multidimensionality impossible to ignore: different instructors load on different dimensions, and forcing them into one composite throws away the structure that makes the data useful.

"But can't we just add a per-instructor question?"

This is the strongest counterargument, and it is partly right. Many legacy platforms now offer a "team-taught configuration" that repeats the instructor block for each named teacher. That is better than a single global item — but it has real limits.

  • Respondent burden. Repeating a full instructor battery for six teachers turns a short survey into a marathon, and burden is the single strongest predictor of abandonment. You buy attribution at the cost of response rate and data quality — a trade explored in our piece on response rates and non-response bias.
  • Rigid structure. A fixed per-instructor grid assumes you know in advance who taught what and that students can map their experience onto that grid. Guest lectures, swapped sessions, and genuinely collaborative co-teaching (two people in the room at once) break it.
  • It still averages. A per-instructor mean is still a mean. It tells you Dr X scored 3.8, not why the integration between Dr X's segment and Dr Y's felt disjointed — which is usually the thing a team-taught module most needs to fix.

So the honest position is: per-instructor items are a real improvement over a single global score, but they inherit the deeper problems of static, fixed-form evaluation. They make the form longer without making it adaptive.

A better unit of analysis: the conversation

The underlying issue is that a static form has to decide its structure before it knows the student's experience. An adaptive, conversational evaluation does not. It can ask "Which parts of this module worked best for you?" and then probe the specific segment, instructor, or transition the student raises — attributing feedback to the right component because the student led it there, not because a grid forced it.

This is where Koji for Education fits the problem rather than fighting it. Koji runs an AI-moderated conversational interview that adapts to what each student actually remembers and cares about, so attribution emerges from the dialogue instead of a rigid per-instructor matrix. Its automatic thematic analysis then clusters open feedback across hundreds of students into themes — "the hand-off between the statistics and ethics blocks was confusing" — at the segment level, not just the course level. The six structured question types (open-ended, scale, single-choice, multiple-choice, ranking, yes/no) let you keep comparable per-block metrics where you need them, while programme- and institution-level reporting rolls component-level insight up to the people making curriculum decisions. And because moderation is standardized, you avoid the inconsistency of different human note-takers interpreting "the teaching" differently across sections.

To be precise about the claim: Koji does not magically solve attribution — students still misremember, and genuinely fused co-teaching is hard to disentangle for anyone. What it does is reduce the structural mismatch between a single-owner instrument and a many-owner course, and surface segment-level themes that a global average buries.

The teaching-assistant blind spot

The attribution problem is not confined to senior academics sharing a lecture series. The most common multi-instructor structure in higher education is a professor-led lecture paired with seminars, labs, or tutorials run by teaching assistants, demonstrators, or graduate students — and standard instruments routinely ignore the TA entirely, attributing the whole experience to the named professor. This is doubly unfair: students who learned most from a skilled tutorial leader cannot say so, and a professor can be penalised for tutorial quality they never controlled. For early-career staff, whose teaching record matters acutely for progression, this invisibility is consequential. Evaluation that lets students speak to the part of the course that actually shaped their learning — lecture, seminar, or lab — restores both fairness and useful signal, and gives developing teachers the specific feedback a single course mean can never provide.

The bottom line

If your programme runs team-taught modules and evaluates them with a single "rate the teaching" score, you are not measuring teaching quality — you are measuring an average of experiences that should be read apart. The remedy is to move the unit of analysis from the course to the component, and from a fixed grid to an adaptive conversation that attributes feedback where it belongs.

See how conversational, component-aware evaluation works in practice with Koji for Education, and pair this with our methodology pieces on programme-level vs course-level evaluation and triangulating teaching evidence. The same adaptive interview engine powers general research on the main Koji platform.