The Teaching Portfolio: The Summative Evidence Student Ratings Cannot Provide
What a teaching portfolio documents that student evaluations miss, what Centra (1994) and Quinlan (2002) found about its reliability, and how to make portfolio evidence defensible for tenure, promotion and accreditation.
Koji Education Team
Product
The short answer
A teaching portfolio is a structured, instructor-curated dossier of evidence about teaching — a reflective statement, syllabi and materials, evidence of student learning, peer-observation notes, and student evaluation data. Its value is that it documents the parts of teaching that a student rating questionnaire cannot see: course design, currency of content, the scholarship behind teaching decisions, and how an instructor responds to feedback over time. But that same richness is its weakness for high-stakes decisions: unless multiple trained reviewers apply explicit criteria, portfolio judgements are idiosyncratic. Treat the portfolio as the narrative that gives meaning to your quantitative evidence — not as a replacement for it, and not as a stand-alone basis for ranking staff.
Bottom line for evaluation committees: Student ratings and teaching portfolios measure different things and should be combined, not substituted. Use student-experience data (ratings plus open-text) as one evidentiary strand inside the portfolio, and require multiple raters with a shared rubric before any portfolio informs a personnel decision.
What the research says
The foundational empirical treatment is Centra (1994), The use of the teaching portfolio and student evaluations for summative evaluation (Journal of Higher Education, 65(5), 555–570). Centra argued that portfolios and student ratings capture complementary, non-overlapping information. Student ratings are reliable for dimensions students are well placed to judge — clarity, organisation, workload, interaction — but they say nothing about whether the syllabus is current, whether the assessment aligns with the learning outcomes, or whether the instructor is reading and acting on the literature of their discipline. The portfolio is where that evidence lives. Centra's central caution, however, was about reliability: because a portfolio is self-selected and holistically judged, its dependability for summative use rests on having several trained raters applying explicit criteria. A single reader's holistic impression is not a defensible basis for tenure or promotion.
Quinlan (2002), Inside the peer review process: how academics review a colleague''s teaching portfolio (Teaching and Teacher Education, 18(8), 1035–1049), looked at exactly how academics reason when they read a colleague''s portfolio. Studying seven reviewers of a biochemistry course portfolio, she found they used normative, case-based reasoning — comparing the portfolio''s practices to their own experience, to prototypical departmental practice, and to their prior knowledge of the teacher. The reflective commentary, the student evaluations, and the syllabus were the documents that mattered most to their judgement. The finding is double-edged: reviewers are thoughtful and context-sensitive, but they anchor on their own teaching as the yardstick, which is precisely the source of inter-rater inconsistency that Centra warned about.
Berk (2005), Survey of 12 strategies to measure teaching effectiveness (International Journal of Teaching and Learning in Higher Education, 17(1), 48–62), places the portfolio in its proper context. Berk''s thesis is that no single source of evidence is adequate; student ratings, peer observation, self-evaluation, teaching scholarship, learning outcomes, and the portfolio are twelve strands, each with distinct strengths and blind spots. The portfolio''s job is to integrate several of those strands into one triangulated, instructor-authored case. Seldin''s long-running practitioner work (Seldin, Miller & Seldin, 2010) makes the same organisational argument: a good portfolio is selective (8–12 pages plus appendices), reflective rather than merely a scrapbook, and explicitly tied to the criteria the reader will use.
Taken together, the literature converges on a clear position: the portfolio is the most complete single artefact for representing teaching, and simultaneously the one most dependent on procedural rigour to be fair.
Why it matters for course evaluation in practice
Most quality-assurance systems in European higher education over-rely on the end-of-course student questionnaire because it is cheap, comparable, and numeric. But a mean of 4.2 on a five-point scale tells a promotion panel almost nothing about why a course works or whether it is well designed. The portfolio is the instrument that turns a number into an argument.
Three practical consequences follow. First, evidence that never reaches the committee otherwise — a redesigned assessment that improved pass rates, a response to last year''s feedback, a teaching innovation piloted mid-cycle — only becomes visible if the instructor is asked to curate and interpret it. Second, the portfolio rebalances the power of a single questionnaire: an instructor teaching a hard, quantitative, compulsory course (systematically rated lower, as the discipline-bias literature shows) can contextualise their ratings with evidence of learning gains rather than being ranked purely on a biased metric. Third, for accreditation (ESG/ENQA, NVAO, AACSB, and the like), reviewers increasingly want to see how feedback closes the loop — and the portfolio, with its reflective commentary and action log, is the natural place to document the improvement cycle.
The catch is procedural. A portfolio scheme that lacks a shared rubric, trained readers, and multiple raters simply moves subjectivity from the student to the committee. The research is unambiguous that fairness depends on the process around the portfolio, not the portfolio itself.
Limitations and honest caveats
A critical reader should hold several objections in view.
- Reliability is the Achilles heel. Centra''s and Quinlan''s findings both point to the same problem: holistic judgement of a curated document is only as consistent as the rubric and the number of raters behind it. Two readers can reach different conclusions about the same portfolio. Committees that use a single reader, or no rubric, are producing evidence no more defensible than the raw student mean they were trying to supplement.
- Curation is selection bias by design. The instructor chooses what to include. A portfolio documents an instructor''s best case, not a representative sample of their teaching. This is legitimate for a reflective, developmental purpose but is a genuine threat to validity when the same document is used for summative ranking.
- Effort and equity. Portfolios are labour-intensive to produce and to read. Staff with less time — often those with heavier teaching or caring loads — may produce thinner portfolios for reasons unrelated to teaching quality, importing a new inequity.
- Generalisability. Quinlan''s study followed only seven reviewers of a single portfolio; the peer-review-reasoning literature is thin and discipline-specific. Claims about how reviewers behave should be read as suggestive, not settled.
- It does not measure learning directly. A portfolio can present evidence of learning, but the evidence is only as good as the underlying measures. A polished narrative around weak outcome data is still weak evidence.
None of this argues against portfolios; it argues for surrounding them with rubrics, multiple trained raters, and a clear separation between formative (developmental) and summative (high-stakes) use.
How Koji incorporates this
Koji is not a portfolio system and does not pretend to be one — the reflective commentary, syllabi and peer-observation notes are the instructor''s to assemble. What Koji does is produce the student-experience strand of the portfolio to a far higher evidentiary standard than a Likert questionnaire, so the evidence an instructor curates is worth curating.
- Beyond the number, an interviewable record. Koji''s AI-moderated conversational interviews probe why a student gives the rating they give, generating open-text evidence that an instructor can quote and interpret in a reflective statement — exactly the material Quinlan found reviewers weighted most heavily.
- Structured question types for portfolio-ready evidence. Using
scale,single_choice,ranking,yes_noandopen_endeditems in one instrument lets an instructor document both a defensible quantitative baseline and the qualitative context around it, rather than a bare mean. - Automatic thematic analysis and quality scoring turn a term''s worth of comments into a small number of evidenced themes with representative quotations — the kind of synthesised, triangulated exhibit a portfolio needs, and one that a committee can audit rather than take on trust.
- Closing-the-loop / action tracking. Because Koji supports mid-cycle (formative) collection as well as end-of-course, an instructor can document a loop: what students said, what changed, and what the next cohort reported — the single most persuasive thing a teaching portfolio can contain, and increasingly what accreditation reviewers ask to see.
- Bias-aware reporting lets an instructor present ratings alongside the contextual factors (course difficulty, compulsory status, cohort size) that the confounds literature shows depress scores, so a portfolio argues from evidence rather than pleading against a number.
Used this way, Koji strengthens one strand of Berk''s twelve without overclaiming to be the others; the peer observation, the self-reflection, and the multi-rater rubric remain essential. Beyond the classroom, Koji''s core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where practitioners face the identical problem of turning satisfaction scores into an evidenced narrative.
Related resources
- Peer Observation vs Student Evaluations: What Each Actually Measures
- Do Instructors and Students Agree? Self-Evaluation vs Student Ratings
- Should Student Evaluations Decide Tenure? The Ryerson Arbitration
- Interpreting and Reporting Student Ratings Responsibly
- What Do Student Evaluations Actually Measure? Marsh and the SEEQ
- One Number or Many? The Dimensionality Debate
References
- Centra, J. A. (1994). The use of the teaching portfolio and student evaluations for summative evaluation. Journal of Higher Education, 65(5), 555–570. https://doi.org/10.2307/2943777
- Quinlan, K. M. (2002). Inside the peer review process: How academics review a colleague''s teaching portfolio. Teaching and Teacher Education, 18(8), 1035–1049. https://doi.org/10.1016/S0742-051X(02)00058-6
- Berk, R. A. (2005). Survey of 12 strategies to measure teaching effectiveness. International Journal of Teaching and Learning in Higher Education, 17(1), 48–62. https://www.isetl.org/ijtlhe/pdf/IJTLHE8.pdf
- Seldin, P., Miller, J. E., & Seldin, C. A. (2010). The Teaching Portfolio: A Practical Guide to Improved Performance and Promotion/Tenure Decisions (4th ed.). San Francisco: Jossey-Bass.
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Should Student Evaluations Decide Tenure? The Ryerson Arbitration and the Limits of High-Stakes SET
A landmark 2018 Canadian arbitration ruled that student evaluations of teaching should not be used to measure teaching effectiveness for promotion and tenure. This guide explains the decision, the evidence behind it, and what it means for governing the high-stakes use of course-evaluation data.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.
Do Instructors and Students Agree? Self-Evaluation vs Student Ratings
Feldman's synthesis found instructor self-ratings and student ratings correlate only moderately (around r ≈ 0.3). What weak self–student agreement means for triangulation, faculty trust, and how to use both sources without privileging either.