New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Do Adjuncts Get Worse Course Evaluations Than Tenured Faculty? Employment Status as a Confound

The causal evidence says contingent faculty do not teach worse — and often teach better. Why employment status is a poor proxy for teaching quality, and how to read evaluation scores fairly across the tenure divide.

Koji Education Team

Product

In brief: There is no reliable evidence that contingent (adjunct, sessional, teaching-only) instructors teach worse than tenured faculty — and the strongest causal study finds the reverse. Figlio, Schapiro and Soter (2015), using eight cohorts of first-year students at Northwestern University, found students learned more from contingent faculty in their first-term courses, a result driven by the weakest quartile of tenure-track staff. Employment status is therefore a poor proxy for teaching quality. Comparing an adjunct's course-evaluation scores against a tenured colleague's — without accounting for the very different courses, class sizes, timetables and support they work under — will systematically mislead any personnel or quality-assurance decision built on it.

What the research says

The most-cited causal evidence comes from Figlio, Schapiro and Soter (2015), Are Tenure Track Professors Better Teachers? (Review of Economics and Statistics, 97(4), 715–724). The authors linked eight entering cohorts at Northwestern University to detailed records of which instructor taught each first-term course, then measured a "downstream" outcome that is much harder to game than an end-of-term rating: whether taking a course with a given instructor predicted the student going on to take — and do well in — a later course in the same subject. On this value-added measure, students consistently learned relatively more from contingent faculty. The effect was concentrated at the bottom: the weakest quartile of tenure-track and tenured staff underperformed their contingent counterparts, and the gap was largest for Northwestern's less-well-prepared students. The authors are careful that Northwestern's contingent faculty are often long-serving, well-supported teaching specialists — not the precariously employed adjuncts common elsewhere — but the headline stands: being on the tenure track did not make someone the better teacher.

Corroborating and complicating evidence:

  • Bettinger and Long (2010), Does Cheaper Mean Better? The Impact of Using Adjunct Instructors on Student Outcomes (Review of Economics and Statistics, 92(3), 598–613), used an instrumental-variables design on more than 43,000 students and found that taking a course with adjunct faculty modestly increased the likelihood a student took a further course in, or majored in, that subject — especially in applied, occupationally-linked fields. Adjuncts were not uniformly worse; in some disciplines they sparked more subsequent interest.
  • Xu (2019), Academic Performance in Community Colleges: The Influences of Part-Time and Full-Time Instructors (American Educational Research Journal, 56(2)), found that part-time instructors were associated with somewhat weaker outcomes on average, but that the effect was largely explained by working conditions — last-minute hiring, no office, no curricular say — rather than any deficit in the individuals. When conditions were equalised, much of the gap closed.

Read together, these studies deliver a consistent methodological message: employment status is entangled with a dozen other variables (course level, class size, timetable slot, institutional support, whether the instructor chose the material), and once you account for them, the raw status difference is small, inconsistent in sign, and frequently favourable to contingent staff.

Why it matters for course evaluation in practice

Quality-assurance offices routinely benchmark instructors against departmental or institutional averages. If contingent and permanent staff are pooled into a single distribution — or worse, contrasted directly — three problems follow.

First, confounded comparisons. Contingent faculty disproportionately teach large first-year service courses, required "gateway" modules, unpopular timetable slots and quantitative subjects. Every one of those is independently associated with lower evaluations in the wider literature (class size, course difficulty, electivity, time-of-day). An adjunct scoring 4.0 on a required 300-student first-year statistics course may be teaching far better than a tenured professor scoring 4.5 on a 12-person final-year elective they designed themselves. The number does not know that; the reader must supply it.

Second, high-stakes misuse. Contingent staff are precisely the group whose contracts are most easily not renewed on the strength of a single soft number. Given that the causal evidence does not support status as a quality signal, using evaluation scores to preferentially cull contingent staff is both unfair and evidentially unsound — the very kind of over-interpretation the Ryerson arbitration warned against.

Third, a demoralised majority. Contingent staff now teach the majority of undergraduate contact hours in many European and North-American systems. An evaluation regime that reads their scores naively, and offers them none of the developmental follow-up their permanent colleagues receive, wastes the feedback and erodes the goodwill of the people doing most of the teaching.

Limitations and honest caveats

The evidence deserves the same scrutiny we ask of any single finding.

  • Generalisability. Northwestern is a highly selective private university; its "contingent" faculty are frequently career teaching specialists on multi-year contracts, not the semester-to-semester adjuncts of a hard-pressed public system. The finding that its contingent staff outperform may not transfer to institutions where adjuncts are hired a week before term with no office and no curricular control — which is exactly Xu's point.
  • Outcome measures differ. Figlio et al. measure subsequent performance and course-taking, not the same-term rating that a course evaluation captures. It is entirely possible for contingent faculty to produce better downstream learning and receive lower end-of-term ratings — the two measures are known to diverge (see the value-added literature). So "adjuncts teach at least as well" and "adjuncts sometimes score lower on SET" can both be true, which is itself the argument against trusting the SET number.
  • Selection. Which instructor teaches which course is not random. The IV and value-added designs above work hard to break that link, but no observational study eliminates it entirely.
  • Heterogeneity. "Contingent" spans a retired industry expert teaching one beloved seminar and an overloaded graduate teaching assistant. Averaging across that category hides more than it reveals — the honest unit of analysis is the individual and the conditions, not the contract type.

How Koji incorporates this

Koji for Education is built so that employment status never silently contaminates a comparison.

  • Context-aware reporting, not naive league tables. Koji reports a course's results against genuinely comparable references — same course level, similar enrolment, similar delivery mode — rather than dropping every instructor into one institutional ranking. A first-year 300-seat required module is benchmarked against its peers, so a contingent instructor is not implicitly measured against a small final-year elective.
  • Low-inference, behaviour-anchored questions. Instead of asking students to rate a global impression that a status halo can colour, Koji's structured items (scale, single_choice, yes_no) target concrete teaching behaviours — "The instructor explained difficult concepts in more than one way", "Feedback on my work arrived in time to use it." Behaviour-anchored evidence travels across employment categories in a way that a global "overall quality" score does not.
  • AI-moderated conversational interviews that probe the why. Koji's evaluations are not a static Likert grid; an AI moderator asks open-ended follow-ups and then applies automatic thematic analysis. When a large required course scores lower, the transcripts usually reveal whether the driver is the instructor or the conditions — the timetable slot, the class size, the room, the lack of a follow-up seminar — the confounds Xu identified. That lets a QA office separate the person from the post.
  • Within-instructor trends over cross-sectional ranking. Koji foregrounds an instructor's own trajectory across cycles rather than a snapshot rank against dissimilar colleagues — the fairer, more actionable signal for contingent staff who move between very different courses.
  • Bias-aware framing. Reports flag when a comparison is being drawn across strongly different course characteristics, prompting the reader to interpret rather than rank.

None of this eliminates the confound — no instrument can — but it is designed to stop employment status masquerading as teaching quality. For teams who also run non-teaching research, Koji's core platform at koji.so applies the same AI-moderated interview engine to product and customer studies, where the same principle holds: the raw score means little until you know who answered and under what conditions.

What a fair contingent-staff evaluation policy looks like

Translating the evidence into policy is straightforward once the confound is named. A defensible approach has four features. First, benchmark within like-for-like course profiles — contingent staff teaching large required modules should be compared against the same kind of module, never against boutique electives. Second, never trigger a non-renewal decision on evaluation scores alone; require corroborating evidence from peer observation, teaching materials or downstream student performance, exactly as you would for a permanent colleague. Third, give contingent staff the same developmental follow-up — the consultation and support that the feedback-intervention literature shows is what actually turns evaluation into improvement — rather than treating their data as a pass/fail gate. Fourth, report the conditions alongside the scores: class size, timetable slot, whether the instructor chose the material, and how much notice they had. Xu's finding that working conditions, not the individuals, drive most of the apparent adjunct penalty means that a report which hides those conditions is not just incomplete — it is actively misleading about where any problem actually lies.

Related resources

References

  • Figlio, D. N., Schapiro, M. O., & Soter, K. B. (2015). Are Tenure Track Professors Better Teachers? Review of Economics and Statistics, 97(4), 715–724. https://doi.org/10.1162/REST_a_00529
  • Bettinger, E. P., & Long, B. T. (2010). Does Cheaper Mean Better? The Impact of Using Adjunct Instructors on Student Outcomes. Review of Economics and Statistics, 92(3), 598–613. https://doi.org/10.1162/REST_a_00014
  • Xu, D. (2019). Academic Performance in Community Colleges: The Influences of Part-Time and Full-Time Instructors. American Educational Research Journal, 56(2), 368–406. https://doi.org/10.3102/0002831218796131