New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Educational Connoisseurship and Criticism: Eisner's Case for Expert Judgement in Teaching Evaluation

Elliot Eisner argued that evaluating teaching is more like art criticism than measurement: it needs a connoisseur's trained perception and a critic's public disclosure. Here is the model, its evidence, its limits, and how it complements student ratings.

Koji Education Team

Product

In brief

Elliot Eisner's model of educational connoisseurship and criticism holds that judging the quality of teaching is closer to what an art, wine or literary critic does than to what a psychometrician does: it requires a connoisseur's trained perception to notice what matters, and a critic's skill to disclose it publicly so others can see it too. Connoisseurship is the private art of appreciation; criticism is the public art of description, interpretation, evaluation and thematics. For course evaluation, Eisner's framework legitimises expert, qualitative judgement — from peers, programme directors, or trained reviewers — as evidence in its own right, and warns against reducing teaching quality to a number that any layperson could read but no expert would trust.

What the research says

Eisner, an art educator at Stanford, imported the vocabulary of the arts into educational evaluation across the 1970s and formalised it in works including The Use of Qualitative Forms of Evaluation for Improving Educational Practice (1979) and later The Enlightened Eye (1991). His starting premise is that teaching is in part an art, and that important qualities of a classroom — the pacing of a discussion, the intellectual generosity of feedback, the way a teacher reads and adjusts to a room — are perceptible to a trained eye but invisible to a standardised instrument.

Two paired concepts do the work. Connoisseurship is "the art of appreciation": like an experienced wine taster or music critic, the educational connoisseur has cultivated the perceptual acuity to distinguish subtle but consequential qualities that a novice would miss. Connoisseurship is private — it is enough to perceive. Criticism is "the art of disclosure": it renders those perceptions public and vivid so that others come to see what the connoisseur sees. Eisner analysed educational criticism into dimensions — description (what is happening), interpretation (what it means), evaluation (its educational worth), and thematics (the transferable lessons).

Eisner's model was part of the same 1970s qualitative turn as Parlett and Hamilton's illuminative evaluation and Robert Stake's responsive evaluation — a shared reaction against the assumption that quasi-experimental and psychometric methods are the only respectable basis for educational judgement. Where it is distinctive is its unapologetic reliance on expertise: quality is what a qualified judge, able to justify the verdict publicly, discerns.

The empirical backdrop makes the case sharper than it may sound. If the field had a valid quantitative measure of teaching quality, expert judgement would be a luxury. It does not. Uttl, White and Gonzalez (2017) found student ratings explain about 1% of the variance in learning; Spooren, Brockx and Mortelmans (2013) judged the validity evidence for SET fragmented at best. Meanwhile, peer observation of teaching — the practical vehicle for connoisseurship — shows only modest convergence with student ratings, implying the two capture different things rather than the same thing measured twice. Eisner would read that not as a failure of peer review but as evidence that a single number cannot substitute for a trained judge who can say why a class works.

Why it matters for course evaluation in practice

  1. Legitimise expert qualitative evidence. Under connoisseurship-and-criticism, a well-argued peer-observation narrative or a programme director's structured judgement is not "soft" data to be outranked by the mean — it is a distinct and valid form of evidence, provided the critic can disclose the grounds publicly.

  2. Demand disclosure, not just verdicts. The model's discipline is that criticism must be made vivid and defensible. A score of 3.4 discloses nothing; a critic's account that describes, interprets, evaluates and draws transferable themes tells a department what to change. This is why connoisseurship pairs naturally with triangulation across sources.

  3. Develop the judges. Connoisseurship is cultivated, not innate. Investing in trained reviewers and calibrated peer-observation raises the quality of the evidence — the human analogue of improving an instrument.

Limitations and honest caveats

  • Subjectivity and elitism. The most obvious objection: expert judgement can encode the expert's biases, tastes and blind spots, and "who counts as a connoisseur?" is a question of power as much as skill. Without calibration and multiple critics, connoisseurship can reproduce the very halo, similarity and status biases documented across the SET literature.
  • Reliability and auditability. A single critic's verdict is hard to audit and may not replicate across reviewers. Frameworks like the ESG expect systematic, defensible processes; connoisseurship must therefore be structured (shared criteria, multiple observers, documented reasoning) to be accreditation-grade.
  • It sidelines the student voice if used alone. A connoisseur observes teaching; students experience it over a whole term. The two are complementary, and privileging the expert can silence the learner. Eisner intended criticism to augment, not replace, other evidence.
  • Scale and cost. Trained expert observation of every course is unaffordable at scale, which historically confined connoisseurship to periodic review rather than routine monitoring.

How Koji incorporates this

Eisner's model is about the human expert — and Koji does not claim to be one. What Koji does is supply the connoisseur with far richer, better-organised material to perceive and disclose, and it does so at a scale human observation never could.

  • Rich qualitative evidence for the critic to interpret. Koji's AI-moderated conversational interviews generate detailed, probing accounts of the student experience across a whole cohort. Instead of a mean and a handful of comments, a programme director acting as connoisseur receives structured themes, representative verbatim quotes, and quality-scored open text — the raw material for description, interpretation and thematics.
  • Thematic synthesis that mirrors Eisner's dimensions. Automatic thematic analysis surfaces what is happening (description) and clusters it into transferable patterns (thematics), leaving the expert to supply the interpretation and evaluation Eisner reserved for trained human judgement.
  • Triangulation, not replacement. Koji is designed to sit alongside peer observation and self-evaluation so the connoisseur can cross-check a trained-eye verdict against the lived student account — exactly the multi-source stance the model implies.
  • Support for calibration. Consistent, structured reporting across courses helps reviewers develop and calibrate their perception over time, the practical route to Eisner's cultivated connoisseurship.

The honest framing matters here: Koji augments expert judgement; it does not automate it. The AI moderates and organises evidence; the connoisseurship — the trained perception of what is educationally worthwhile — remains human, and Koji's reporting is designed to make that human judgement better informed and more defensible, not to substitute a score for it. The same AI-moderated interview engine underpins Koji's core research platform at koji.so, where skilled researchers likewise rely on it to surface the evidence their own expertise then interprets.

Making connoisseurship accreditation-grade

The gap between a defensible connoisseurship process and an idiosyncratic one is procedural, and it is bridgeable. Three safeguards do most of the work. First, multiple critics: pairing or triangulating reviewers converts a private verdict into an inter-subjective one and exposes the individual blind spots that make single-observer judgement fragile. Second, shared descriptive criteria: giving reviewers a common vocabulary for description — the observable features of clarity, challenge, responsiveness, inclusivity — disciplines the leap from perception to evaluation and makes disagreements traceable to evidence rather than taste. Third, documented reasoning: Eisner's insistence that criticism be publicly disclosed maps directly onto the audit expectations of the ESG. A review that records what was observed, how it was interpreted, and why it was judged as it was, is one an external panel can scrutinise.

Structured this way, connoisseurship stops being the antithesis of quality assurance and becomes one of its more demanding forms. The reviewer is not asked to suppress expert judgement in favour of a checklist; they are asked to make that judgement visible, comparable and contestable. That is a higher bar than assigning a number, and — where teaching quality genuinely cannot be reduced to a valid scalar — a more honest one. It also reframes the perennial worry about reviewer subjectivity: the goal is not to pretend the judgement is objective, but to make its grounds so explicit that a disagreeing colleague can point to exactly where and why they see the class differently.

Related resources

References

  • Eisner, E. W. (1979). The use of qualitative forms of evaluation for improving educational practice. Educational Evaluation and Policy Analysis, 1(6), 11-19. https://doi.org/10.3102/01623737001006011
  • Eisner, E. W. (1991). The Enlightened Eye: Qualitative Inquiry and the Enhancement of Educational Practice. New York: Macmillan.
  • Eisner, E. W. (1976). Educational connoisseurship and criticism: Their form and functions in educational evaluation. Journal of Aesthetic Education, 10(3/4), 135-150. https://doi.org/10.2307/3332067
  • Uttl, B., White, C. A. & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22-42. https://doi.org/10.1016/j.stueduc.2016.08.007
  • Spooren, P., Brockx, B. & Mortelmans, D. (2013). On the validity of student evaluation of teaching: The state of the art. Review of Educational Research, 83(4), 598-642. https://doi.org/10.3102/0034654313496870

Related articles

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

analysis-reporting

How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis

When you turn thousands of free-text comments into themes and counts, how do you know the coding is trustworthy? O Connor and Joffe (2020) on intercoder reliability, Braun and Clarke on thematic analysis, and what rigorous qualitative QA looks like.

research-methods

What Can Open-Text Student Comments Tell You That Likert Scores Cannot?

A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.

research-methods

What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback

Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.