New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Do Instructors and Students Agree? Self-Evaluation vs Student Ratings

Feldman's synthesis found instructor self-ratings and student ratings correlate only moderately (around r ≈ 0.3). What weak self–student agreement means for triangulation, faculty trust, and how to use both sources without privileging either.

Koji Education Team

Product

In brief: Instructors and their students agree only moderately about how good the teaching was. Feldman's (1989) synthesis of multi-source studies put the average correlation between instructor self-ratings and student ratings at roughly r ≈ 0.3 — positive, but far from interchangeable. The right conclusion is not that one party is "right" and the other "wrong," but that self-assessment and student feedback measure overlapping yet distinct things. Each is a legitimate, partial source of evidence, which is exactly why teaching quality should be judged by triangulating multiple perspectives rather than by any single rater.

A recurring move in debates about course evaluations is to ask which judge to trust: the students, who experienced the teaching, or the instructor, who designed it. The evidence reframes the question. Students and instructors are not two attempts to measure the same quantity; they are two vantage points that partially overlap. Knowing how much they overlap — and where they diverge — is what lets a quality office use both without naïvely averaging them or treating disagreement as someone's error.

What the research says

The anchor is Feldman (1989), "Instructional effectiveness of college teachers as judged by teachers themselves, current and former students, colleagues, administrators, and external (neutral) observers," published in Research in Higher Education. Feldman synthesised studies that collected ratings of the same teaching from multiple sources and compared them. The headline pattern: teachers' self-ratings and their current students' ratings are, at best, moderately related — the average association clustered around the r ≈ 0.3 region. Agreement was higher between some source pairs (for example, current and former students tend to converge more strongly) and notably weak between teachers' self-ratings and their colleagues' ratings. In short, the more independent the vantage points, the lower the agreement, and self-assessment was one of the least convergent sources.

Two interpretive points from Feldman's broader programme matter. First, the modest self–student correlation is not evidence that students are unreliable. Feldman's parallel work on the dimensions of student ratings (and the multisection validity tradition) shows student ratings carry real, structured information about specific instructional behaviours — organisation, clarity, stimulation of interest — that relate to achievement. The low self–student correlation tells us instructors and students weight different things, not that either signal is noise.

Second, self-assessment has a well-documented optimism problem. A large psychological literature (the self-enhancement and "above-average effect" findings) shows people, including professionals, tend to rate their own performance more favourably and with less discrimination than independent observers do. Applied to teaching, instructors often see intentions and effort that students never observe, while students see experienced clarity and fairness that instructors cannot feel from the front of the room. The two sources are partly looking at different objects.

This dovetails with the convergent-validity evidence on other sources. Our review of peer observation versus student evaluations finds a similar story — moderate, imperfect overlap between trained observers and students — and the general principle is developed in our note on triangulating teaching evaluation across multiple evidence sources. No single source is a gold standard; each is a fallible indicator that becomes trustworthy mainly in combination.

Why it matters for course evaluation in practice

For quality assurance, weak self–student agreement carries several concrete lessons.

First, disagreement is data, not error. When an instructor's self-assessment and their students' ratings diverge sharply, the productive response is to ask why — not to declare one side correct. A lecturer who rates their feedback practices highly while students rate them poorly may be giving feedback that is timely from the desk but unclear or too late from the seat. That gap is one of the most useful diagnostic signals an evaluation system can surface, and it is invisible if you only collect one source. This is also why pairing self-reflection with student data is a more honest improvement conversation than handing back scores alone — the loop-closing point developed in our piece on whether feedback actually improves teaching.

Second, collecting instructor self-evaluation can repair trust. Much faculty resistance to student ratings stems from feeling judged by a single, contestable number (see why faculty distrust course evaluations). Inviting the instructor's own structured self-assessment alongside the student data reframes evaluation as a multi-perspective conversation in which the instructor is a participant, not a defendant. Marsh and Roche's intervention work used exactly this design — teachers completed self-evaluations and then worked on targeted dimensions — and it is associated with better engagement and improvement.

Third, never let one source make a high-stakes decision alone. If self-ratings and student ratings overlap only at r ≈ 0.3, then either source used in isolation captures a minority of the relevant picture. For probation, promotion, or programme review, the defensible practice is a portfolio: student feedback, self-reflection, peer observation, and outcomes, each weighted for what it can and cannot see. This is the operational meaning of triangulation and the expectation embedded in European quality frameworks, which ask for multiple lines of evidence rather than a single metric.

Limitations and honest caveats

Several cautions apply. Feldman's synthesis is a synthesis of older studies, many predating modern instruments, online delivery, and today's student populations; the specific r ≈ 0.3 figure should be read as an order-of-magnitude summary, not a precise constant. Correlations of this kind also vary with the dimension being rated — agreement on concrete, observable behaviours (Did the instructor return work promptly?) tends to be higher than on global, evaluative judgements (Was this excellent teaching?), so a single average masks real heterogeneity.

There are measurement subtleties too. Self-ratings suffer range restriction — instructors who agree to be studied may be unusually reflective, and self-ratings often bunch toward the favourable end, which mechanically lowers correlations with any other source. A modest correlation can therefore understate genuine agreement on the underlying construct. Conversely, shared method or shared halo can inflate agreement between sources that draw on the same impressions. Correlations alone do not tell you which is happening.

Finally, moderate agreement does not adjudicate validity. That students and instructors disagree does not, by itself, tell you whose view better predicts learning. The criterion question — which source tracks actual student learning — is separate, and the SET-and-learning literature shows even student ratings predict learning only modestly. The honest framing is humility: each source is partial, none is a gold standard, and the value lies in combining them intelligently, not in crowning a winner.

How Koji incorporates this

Koji is designed to treat teaching quality as a multi-source judgement rather than a single student number — directly reflecting the triangulation lesson of this research.

  • Structured instructor self-evaluation alongside student data. Koji can run a self-evaluation study for instructors using the same structured question types (scale, open_ended, ranking, single_choice) used for students, so a course's self-assessment and its student feedback are captured in comparable form and can be set side by side rather than compared apples-to-oranges.
  • Surfacing the divergence, not hiding it. Because Koji's reporting can present multiple sources together, the most diagnostically valuable signal — where the instructor's view and the students' view diverge — becomes visible instead of buried. The AI-moderated interview can probe students for the concrete behaviour behind a low rating, giving the instructor something specific to reconcile against their self-perception.
  • Dimension-level comparison. Rather than a single overall gap, Koji's structured, multidimensional design lets a department see where self and student views agree (often concrete logistics) and where they part (often clarity or feedback usefulness), matching Feldman's finding that agreement is dimension-dependent.
  • Triangulation by design. Koji is built to sit within a portfolio of evidence — student feedback, self-reflection, and other sources — so no single rater drives a high-stakes outcome, consistent with the r ≈ 0.3 reality that any one source is partial.
  • Trust-preserving framing. By making the instructor a contributing voice through self-evaluation rather than only the object of student scoring, Koji supports the kind of multi-perspective conversation associated with better engagement and improvement.

Koji does not claim its tooling makes self and student views converge — the research suggests they legitimately should not fully converge. The aim is to capture both faithfully, compare them at the dimension level, and make disagreement a starting point for inquiry. The same multi-source, AI-moderated approach underpins Koji's core research platform at koji.so, where reconciling self-reported and observed behaviour is a daily problem in customer and product research.

Related Resources

References

  1. Feldman, K. A. (1989). Instructional effectiveness of college teachers as judged by teachers themselves, current and former students, colleagues, administrators, and external (neutral) observers. Research in Higher Education, 30(2), 137–194. https://doi.org/10.1007/BF00992716
  2. Marsh, H. W., & Roche, L. (1993). The use of students' evaluations and an individually structured intervention to enhance university teaching effectiveness. American Educational Research Journal, 30(1), 217–251. https://doi.org/10.3102/00028312030001217
  3. Feldman, K. A. (1989). The association between student ratings of specific instructional dimensions and student achievement: Refining and extending the synthesis of data from multisection validity studies. Research in Higher Education, 30(6), 583–645. https://doi.org/10.1007/BF00992392