Teaching Portfolios Are Not the Objective Alternative to Student Evaluations You Think
When student evaluations are attacked as biased, the teaching portfolio is offered as the grown-up alternative. But portfolios have their own reliability and self-presentation problems. The honest case is not portfolio versus survey — it is portfolio plus survey plus more.
Koji Education Team
Product · July 31, 2026
Bottom line up front: Whenever student evaluations of teaching get criticised for bias, someone in the room proposes the teaching portfolio as the objective, holistic alternative. It is a valuable instrument — but it is not objective, and treating it as the fix simply moves the measurement problem somewhere less visible. Portfolios are curated self-presentations scored by human raters, which introduces impression management and often low inter-rater reliability. The defensible position, supported by the evaluation literature for two decades, is not portfolio instead of survey. It is multiple imperfect sources, each covering the others' blind spots.
What a teaching portfolio is — and why it appeals
The teaching portfolio, popularised in higher education largely through Peter Seldin's work from the early 1990s onward, is a structured, self-authored dossier: a statement of teaching philosophy, sample syllabi and assignments, evidence of student learning, peer-observation notes, and reflective commentary. Its appeal is obvious and genuine. Unlike a single end-of-term number, a portfolio shows teaching as a practice over time. It captures course design, intellectual ambition and reflective growth that a five-point scale never touches. When faculty argue that student evaluations should not decide tenure and promotion, the portfolio is their natural counter-proposal — and a reasonable one.
The trouble starts when the portfolio is asked to be not just richer but more objective. It is not.
The reliability problem nobody mentions
A portfolio only informs a personnel decision if different reviewers, reading the same dossier, reach similar conclusions. That is inter-rater reliability, and it is exactly where portfolio assessment is weakest. Research on high-stakes portfolio and performance assessments of teaching has repeatedly found reliability to be modest or low. A 2024 study of a high-stakes performance assessment of teacher candidates, for example, reported inter-rater reliability that was low across all content areas examined (Education Sciences, 2024). The mechanism is not incompetence; it is that rich, holistic evidence is interpreted, and interpretation varies between readers — the same reason two analysts rarely agree on open-text coding without a shared, enforced protocol. A rubric helps, but it does not make the judgement objective; it makes it a little more consistent.
So the portfolio does not escape the reliability critique aimed at student surveys. It relocates it — from "are students biased raters?" to "are committee members consistent raters?" — and the second question is often answered less favourably than the first.
The self-presentation problem
A student survey is an unsolicited signal: the instructor does not choose which responses arrive. A portfolio is the opposite — it is curated by the person being evaluated. That is a feature for formative reflection and a bug for summative judgement. Every portfolio is, by design, a best-foot-forward document. The syllabus that flopped, the semester the redesign backfired, the module with the awful results — none of these are obligated to appear. This is not dishonesty; it is impression management, the same self-favouring selection that makes self-evaluation a weak sole source of evidence. A skilled writer with a modest teaching record can assemble a compelling portfolio; a brilliant teacher who writes reluctantly can assemble a thin one. The document rewards narrative ability partly independently of teaching quality — a close cousin of the Dr. Fox effect, where expressiveness masquerades as substance.
The same trap as every other "objective corrective"
There is a pattern here worth naming. Each time student evaluations are criticised, a supposedly objective alternative is nominated — and on inspection, each carries its own systematic error. Peer observation of teaching is not the objective corrective it looks like: a one-off visit is a tiny, reactive sample scored by a colleague with their own norms. Pass rates are not an objective measure of teaching quality: they are confounded by intake, grading standards and leniency pressure. Teaching portfolios join the list. The recurring error is believing that qualitative richness or administrative distance from students equals objectivity. It does not. Every source of evidence about teaching is a measurement with its own bias profile.
What the evidence actually recommends
The most-cited synthesis here is Ronald Berk's Survey of 12 Strategies to Measure Teaching Effectiveness (2005), which reviews twelve possible sources — student ratings, peer ratings, self-evaluation, videos, student interviews, alumni ratings, employer ratings, administrator ratings, teaching scholarship, teaching awards, learning-outcome measures and teaching portfolios — and concludes that no single source is sufficient. Berk argues for a unified conceptualisation that combines multiple sources for both formative and summative decisions, because multiple sources "build on the strengths of all sources while compensating for the weaknesses in any single source" (Berk, 2005, IJTLHE). That is the same logic behind triangulation in teaching evaluation: student ratings are necessary but not sufficient, and so is everything else.
Crucially, Berk also warns against mixing purposes. A portfolio assembled for improvement and a portfolio assembled to win a promotion case are different documents, and asking one artefact to do both jobs corrupts both — the dual-purpose problem that quietly wrecks so many evaluation systems.
But isn't a portfolio still better than a biased survey?
The strongest version of the objection is fair: even if portfolios are unreliable, at least they capture teaching, whereas biased student ratings can penalise instructors for their gender, accent or the difficulty of their course. Two responses.
First, this is a false choice. The point is not to defend student surveys as sufficient — they are not — but to reject the idea that portfolios are their objective replacement. The right move is combination, not substitution.
Second, portfolios can be improved precisely where surveys are weakest, if we stop pretending they are objective and start engineering their reliability: shared rubrics, calibrated multiple raters, required inclusion of a genuine failure-and-response case, and — critically — better student-voice evidence inside the portfolio than a bare score. A teacher who can show what students actually said, thematically analysed with the raw quotes attached, has stronger portfolio evidence than one who pastes in a mean rating. That is where the source of student feedback matters.
Where Koji fits
If the honest model is "multiple imperfect sources", then the quality of each source is worth improving on its own terms — including the student-voice source that ends up inside every portfolio and every promotion case. A single satisfaction number is thin portfolio evidence and, as faculty rightly distrust it, unpersuasive to committees.
Koji for Education strengthens that source. Its AI-moderated conversational interviews probe why students experienced a course as they did, across six structured question types, and its automatic thematic analysis turns hundreds of open responses into named themes with the underlying quotes attached and quality-scored. For a portfolio, that converts "4.2/5" into defensible, evidence-linked claims: students consistently reported that the redesigned problem sets improved their confidence, in their own words. Because the moderation is standardised and bias-aware, the student-voice evidence a teacher carries into a portfolio is more consistent than a colleague's one-off observation — and because reporting is formative and mid-cycle as well as summative, the portfolio can honestly show change over time rather than a single snapshot. Koji does not make teaching evaluation objective; nothing does. It makes one of your evidence sources markedly stronger, so triangulation rests on better inputs.
Teams that also run general user and customer research will find the main Koji platform runs on the same AI interview engine, so the practice transfers beyond the classroom.
Stop asking which single instrument is the objective one. Build a portfolio of evidence — and make the student-voice piece of it something a committee can actually trust.
What a defensible portfolio process looks like
If portfolios are going to inform real decisions, engineer their weaknesses out rather than pretending they are absent. Four practices do most of the work. Calibrate raters against shared exemplars before they score anything, and use at least two, because a single reader's judgement is the least reliable configuration there is. Require a genuine failure-and-response case — a redesign that flopped and what the teacher did next — so the document tests reflective practice rather than rewarding only success stories. Separate the improvement portfolio from the promotion portfolio, because one artefact cannot serve both purposes without corrupting both. And strengthen the student-voice exhibit: a mean rating is weak evidence, whereas thematically analysed student feedback with the underlying quotes attached lets a committee see what students actually experienced, not just how they scored it. None of this makes a portfolio objective. It makes it accountable — a document whose claims a second reader can check against evidence, which is the most any teaching-evaluation instrument can honestly offer.
Want student-feedback evidence strong enough to put in a teaching portfolio? Explore Koji for Education.