Peer Observation of Teaching Is Not the Objective Corrective You Think It Is
When student evaluations look biased, the reflex is to add peer observation as the "objective" counterweight. But a single colleague dropping into a single class is one of the least reliable measures in the toolkit — reliable only under conditions most institutions never meet.
Koji Education Team
Product ·
Bottom line up front: Peer observation of teaching is widely proposed as the objective corrective to biased, popularity-driven student ratings. The evidence is more sobering. A single colleague visiting a single class produces a judgement with poor reliability, vulnerable to halo, leniency, and reactivity. Peer observation can reach strong reliability — but only with structured instruments, multiple trained observers, and several visits, which is precisely the resource-heavy design most universities skip. Used well, it is a valuable strand of triangulation. Used as the casual, once-a-year drop-in it usually is, it swaps one flawed number for another and calls it rigour.
The seductive logic — and its flaw
The argument writes itself. Student evaluations are contaminated by bias and by the Dr Fox effect — charisma outscoring content. Students are not disciplinary experts and cannot judge whether the material is current or the pedagogy sound. So bring in a peer: a fellow academic who can judge those things, watching the teaching directly. Objective, expert, immune to the popularity contest.
The trouble is that "a colleague watched a class and formed a view" is not a measurement in any technical sense until you ask the questions we ask of every other instrument: is it reliable, and is it valid? For the informal version of peer observation that dominates practice, the answer to the first question is: not very.
What the reliability evidence actually says
The key finding in the peer-observation literature is conditional. When observation is done with a validated, structured instrument and multiple trained observers, reliability can be strong. In one validation of a Peer Observation and Evaluation Tool, intraclass correlation coefficients across lectures rated by two to three observers ranged from 0.66 to 0.97 — good agreement (Peer Observation and Evaluation Tool validation, PMC). More recent work has pushed the same approach to online teaching, validating an instrument with many-facet Rasch measurement to establish construct validity and reliability (Springer, ETR&D, 2024).
Read that carefully, because the conditions are the whole story. Those numbers come from structured tools, multiple observers per session, and explicit training and rating protocols. Strip those conditions away — as almost every routine peer-observation scheme does — and you are left with one untrained colleague, one visit, one holistic impression. That is not the design that produced 0.66-0.97. It is the design that produces the well-known threats:
- Single-observer subjectivity. With one rater there is no inter-rater agreement to speak of — you have that observer''s idiosyncratic standards and nothing to check them against. Establishing dependability is exactly the problem generalizability theory exists to formalise: one observer on one occasion is a single, noisy draw.
- Reactivity (the observed-lesson effect). People teach differently when watched. A pre-announced observation captures a performance staged for a colleague, not a typical class — the teaching equivalent of a Hawthorne effect.
- Halo and collegiality bias. Observers who know and like the person — or who will be observed by them next term — tend to rate generously. Reciprocal leniency is a structural feature of small departments.
- Tiny sample of behaviour. One 50-minute slice cannot represent a semester any more than one exam question can represent a syllabus. Generalisability improves only as you add occasions and observers, and each addition multiplies the scheme''s cost in scarce academic time. That economic reality is why most peer-observation schemes quietly settle for the single unreliable visit — and then, understandably but wrongly, report its verdict with more confidence than one occasion can bear.
The developmental-versus-judgemental tension
There is a deeper problem that reliability statistics do not capture: peer observation is asked to do two incompatible jobs, the same dual-purpose conflict that undermines student evaluations.
Most peer-observation schemes were designed as developmental — a supportive, formative conversation between colleagues to improve practice. That works precisely because it is low-stakes and non-judgemental; honesty flows when nobody''s promotion is on the line. The moment an institution repurposes the same visit as evidence for probation, promotion, or performance management, the candour evaporates. Observers soften; the observed perform; the developmental value and the evaluative value cannot both survive in one visit. Bolting a summative purpose onto a formative instrument breaks it — a lesson the student-evaluation world learned the hard way.
The counterargument: surely two flawed measures beat one?
The strongest defence of peer observation is that it does not need to be perfect — it only needs to be independent of student ratings. Triangulation gains its power from combining sources whose errors do not correlate; even a noisy peer measure adds information if its noise is different from the students''.
This is right, and it is the correct reason to keep peer observation in the mix — but it comes with two conditions people forget. First, triangulation assumes the sources are genuinely independent; if peer observers are quietly influenced by a lecturer''s reputation or by the very student scores they are meant to cross-check, the errors correlate and the apparent corroboration is an illusion. Second, combining measures only helps if each is interpreted with its uncertainty attached. A peer observation reported as a confident verdict — rather than as one noisy, single-occasion data point — does not strengthen a judgement; it launders subjectivity as expertise. The value of peer observation is real, but it is conditional on humility about what a single visit can support. It is a strand of evidence, not a trump card, exactly as triangulation across multiple sources requires.
What good peer observation requires
If you want peer observation to carry evaluative weight, the reliability evidence tells you the price:
- A structured, validated instrument — a defined rubric, not "jot down your impressions." This is the same argument for validated instruments over home-grown forms that applies to student surveys.
- Multiple trained observers per teacher, so inter-rater reliability can actually be estimated.
- Several observations across the term to average out the bad-day and staged-lesson effects.
- A firewall between developmental and judgemental use — separate schemes, separate records, so the formative conversation stays safe.
Most institutions run none of this. They run one drop-in, by one untrained colleague, once a year, and file the result as if it settled something.
Where student voice — and Koji — fit
None of this argues against peer observation. It argues for honesty about what each strand of evidence can bear, and for making the student-voice strand as rigorous as possible so it pulls its weight in the triangle rather than being dismissed as mere popularity.
That is where Koji for Education contributes. It does not observe teaching — that is the peer''s job. What it does is make the student contribution to triangulation far more informative than a Likert average: AI-moderated conversational interviews that probe why students experienced a class as they did, six structured question types for a comparable quantitative backbone, automatic thematic analysis of open text, and standardised, bias-aware moderation that removes the between-rater inconsistency a human process introduces. A robust, well-characterised student measure is a better triangulation partner for peer observation than a noisy mean — and, because Koji reports the substance behind the scores, it makes it harder to over-interpret any single source, student or peer. The same interview engine underpins general research on the main Koji platform.
The bottom line
Peer observation feels objective, and that feeling is most of its appeal. But the reliability that justifies the feeling only appears under conditions — structured tools, multiple trained raters, repeated visits — that the typical annual drop-in never meets. Keep peer observation; it adds a genuinely independent lens. Just stop treating a single colleague''s single visit as the rigorous corrective to student ratings. Both are strands. The discipline is in weighting each for what it can actually support.
Make the student-voice strand of your evaluation evidence rigorous enough to triangulate with confidence — explore Koji for Education.