Triangulation in Teaching Evaluation: Why Student Ratings Are Necessary but Not Sufficient
Student ratings are one valid window onto teaching — and only one. The methodological literature has been clear for two decades: sound evaluation triangulates multiple sources. Here is what that means in practice, and how to do it without drowning in instruments.
Koji for Education
Research & Editorial Team · June 6, 2026
Bottom line: No single source can validly measure something as multidimensional as teaching. Student ratings see things peers cannot; peers see things students cannot; instructors and learning evidence see different things again. The methodological consensus — articulated most clearly by Ronald Berk — is triangulation: combine multiple, complementary sources so each compensates for the others'' blind spots. Student feedback is necessary, but on its own it is not sufficient.
The single-source trap
Most institutions evaluate teaching with one instrument: the end-of-term student rating, averaged and compared. We have written at length about why that average is a poor high-stakes measure and why averaging Likert scores misleads. But there is a more basic problem that predates any debate about bias: teaching is not one thing, and no single observer sees all of it.
A student experiences whether a class was clear, well-paced, fairly assessed, and engaging. A student does not see whether the content was current and accurate, whether the assessment aligned with the discipline''s standards, whether the instructor was scaffolding toward a coherent curriculum, or how the course fit a programme. A peer reviewer sees some of those and misses what only a learner in the room can feel. An instructor''s own reflection captures intent and design rationale invisible to everyone else. Asking the student rating to carry the whole judgement is asking one instrument to measure dimensions it was never built to detect.
Berk''s framework: twelve-plus sources, one principle
The clearest statement of the alternative comes from Ronald A. Berk. In Survey of 12 Strategies to Measure Teaching Effectiveness (International Journal of Teaching and Learning in Higher Education, 2005) and his later work on using multiple sources of evidence, Berk catalogues the available sources: student ratings, peer ratings, self-evaluation, video, student interviews, alumni ratings, employer ratings, administrator ratings, teaching scholarship, teaching awards, learning-outcome measures, and teaching portfolios.
His argument is not that you should use all twelve. It is that multiple sources build on the strengths of each while compensating for the weaknesses of any one, and that, given the complexity of teaching, this triangulation is the only defensible basis for formative and summative decisions. His one-line summary has become a touchstone in the field: student ratings are a necessary, but not sufficient, source of evidence.
This matters because the bias and validity findings reviewed elsewhere on this blog — including the Uttl et al. (2017) meta-analysis showing SET ratings explain at most ~1% of variance in learning — are devastating only if you rely on a single source. Within a triangulated design, student ratings stop being a flawed verdict and become one valid, bounded signal among several. Triangulation is not a way of rescuing a bad metric; it is the methodologically correct response to the fact that every individual source is partial.
What triangulation looks like in practice
Triangulation has a reputation for being impractical — twelve instruments, twelve administrative burdens. It needn''t be. A workable model for most programmes uses three or four complementary strands:
- Student voice, read for substance. Not just the scale averages, but the open-ended accounts of what helped and hindered learning. This is the richest, most actionable strand — and the one most institutions under-use.
- Peer or expert review. A colleague observes design, currency and disciplinary rigour — the dimensions students cannot judge. Done formatively, with a rubric, it is one of the most developmental sources available.
- Instructor self-reflection / teaching portfolio. The instructor explains intent, design choices and how they responded to prior feedback — closing the loop and surfacing context.
- Evidence of learning, where feasible. Assessment patterns, capability growth, or — at programme level — alumni and employer signal connecting teaching to graduate outcomes.
The principle of convergent validity does the work: when independent sources point the same way, confidence rises; when they diverge, that divergence is itself diagnostic information, not noise to be averaged away.
But isn''t triangulation just more work for the same answer?
This is the practical counterargument and it deserves a candid response. Three risks are real. First, cost: more sources mean more time from already-stretched staff. Second, false rigour: bolting four biased instruments together does not magically produce an unbiased result — if all four share a bias (say, all favour confident, charismatic instructors), triangulation can launder that bias as consensus. Third, incommensurability: the sources don''t reduce to a common scale, so committees may quietly fall back on the one number that does — the student average — defeating the purpose.
These are genuine, but they are arguments for doing triangulation well, not for abandoning it. Cost is contained by keeping each strand lightweight and using automation for the heavy lifting (thematic analysis of open text, in particular). The shared-bias risk is mitigated by choosing sources with different vulnerabilities — students and peers are biased by different things, which is precisely why combining them helps. And incommensurability is a feature, not a bug, if committees are trained to weigh qualitative convergence rather than to demand a single composite score. The failure mode is not triangulation; it is triangulation collapsed back into a number.
Sequencing: formative first, summative with care
Triangulation is easier to adopt when you separate its two jobs. Formatively — to help an instructor improve — you can be generous and exploratory: gather student conversation, a peer''s developmental observation, and the instructor''s own reflection, and let the convergences and tensions between them drive a conversation about practice. Here, divergence between sources is the most interesting output, not a problem to resolve. Low stakes mean you can afford breadth.
Summatively — to inform a renewal, promotion or programme-review judgement — the bar is higher and the discipline tighter. Each source should be weighted for what it can validly speak to: student voice on experience and engagement; peers on rigour, currency and design; learning evidence on outcomes. No source should be allowed to dominate outside its competence, and small numerical differences should never be treated as rankings. The committee''s task is to assess whether independent sources converge, and to treat unexplained divergence as a prompt for more evidence rather than as something to average into a composite.
A simple rule of thumb keeps this honest: collect richly, judge conservatively. Most of the volume of evidence should feed improvement, where the cost of being wrong is low and the value of detail is high. Only a deliberately triangulated, conservatively weighted subset should feed high-stakes decisions. Institutions that invert this — collecting thinly and judging aggressively on a single average — get the worst of both worlds.
Where Koji fits
Triangulation has historically been expensive because the qualitative strands — student interviews, open-text analysis, self-reflection — were labour-intensive to collect and harder still to synthesise. That is exactly the constraint Koji for Education removes. Koji''s AI-moderated conversational interviews make the deepest, most valid student strand — the interview, not the form — feasible at scale, probing beyond a rating to why students experienced a course as they did. Its automatic thematic analysis turns hundreds of open responses into structured patterns, so the rich strand becomes usable evidence rather than an unread pile of comments. Its six structured question types let a single instrument carry both comparable items and deep open exploration, and bias-aware, standardised AI moderation ensures every student is asked comparable questions — removing one of the inconsistencies that undermines multi-source designs.
Because Koji also supports programme- and institution-level reporting and closing-the-loop action tracking, the student strand can be combined cleanly with peer review and learning evidence in a single quality picture — and the loop back to instructors and committees is documented. The same conversational engine powers general research on the main Koji platform, so teams running both teaching evaluation and wider stakeholder research use one consistent method.
Koji does not replace peer review or self-reflection — nor should any tool claim to be the whole of triangulation. It makes the student strand far richer and far cheaper to run well, which is usually the strand institutions most need to strengthen. If you are moving from single-source ratings to a defensible multi-source model, explore Koji for Education.