New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Your Lecture-Designed Evaluation Form Fails Lab, Studio, and Clinical Courses

The standard course evaluation was built for a person talking at the front of a room. Used on a lab, a design studio, or a clinical placement, it asks the wrong questions about the wrong teaching — and the resulting numbers are not just imprecise, they are measuring the wrong thing.

Koji Education Team

Product ·

Bottom line up front: Most institutions run every course through one evaluation instrument — and that instrument was designed, item by item, around the lecture: a knowledgeable person explaining content to a room. Apply the same form to a chemistry lab, an architecture studio, a clinical placement, or a music conservatoire, and its questions misfire. "The lectures were clear", "the instructor explained concepts well", "the pace was appropriate" describe a mode of teaching that practical courses barely use. The result is not merely noisy data; it is construct mismatch — the form measures the wrong teaching, penalises pedagogies it was never built to see, and produces scores that are unfair to compare against lecture courses. Different signature pedagogies need different evaluation questions.

Signature pedagogies: why the lecture is not the universal case

The conceptual frame comes from Lee Shulman's influential 2005 essay Signature Pedagogies in the Professions in Daedalus. Shulman observed that each profession has characteristic forms of teaching that train novices "to think, to perform, and to act with integrity" — the bedside round in medicine, the case dialogue in law, the design crit in architecture, the studio in the arts, the bench in the sciences. These are not lectures with a different label. They organise attention, feedback, error, and supervision in fundamentally different ways.

A design studio runs on the critique — public, iterative, often uncomfortable feedback on work in progress. A clinical placement runs on supervised practice with graduated responsibility and real consequences. A laboratory runs on procedure, safety, and the productive failure of experiments that do not work. A conservatoire runs on the master-apprentice lesson. None of these is well described by "the instructor explained the material clearly". In a good studio, the instructor may explain very little; the learning is in the making and the crit. In a good lab, "clarity" is about protocol and safety, not exposition.

This matters for evaluation because European higher education is full of these settings. The sector includes a large and growing population of professionally-oriented institutions — universities of applied sciences, polytechnics, conservatoires, and the practical faculties of research universities — tracked across roughly 3,500 institutions in the European Tertiary Education Register. A meaningful share of all teaching in Europe is not a lecture. Yet a meaningful share of all evaluation assumes it is.

Three ways the lecture-form misfires on practical courses

1. It asks about activities that did not happen. Items about lectures, slides, explanation, and pace are often inapplicable in a studio or placement. Students faced with inapplicable items either skip them (creating missing data) or answer anyway against some vague impression (creating meaningless data). Either way the response is uninformative. The "not applicable" option, where it exists at all, is usually treated as missing rather than as the signal it is.

2. It is silent on what actually matters. The lecture form rarely asks about the things that define quality in practical teaching: the usefulness and timeliness of feedback on iterative work, the quality of supervision, access to equipment and materials, psychological safety during public critique, the realism of clinical exposure, the balance between autonomy and support. A form that cannot ask about supervision quality is structurally blind to the single most important variable in a clinical or studio course. We make the general version of this point in evaluating work-integrated learning placements and internships and in evaluating doctoral supervision — both settings where the standard form has almost nothing relevant to ask.

3. It produces scores that are unfairly compared. Because practical courses are demanding, sometimes uncomfortable (a hard crit, a failed experiment, a stressful ward), and ask more of students than passive attendance, they can attract systematically different ratings than a polished lecture — not because the teaching is worse, but because the experience is more effortful. When these scores are then dropped into the same league table as lecture courses, the comparison is invalid. This is a form of construct-irrelevant variance, the validity threat we examine in construct-irrelevant variance in course evaluation: the score reflects the mode of teaching as much as its quality.

There is a parallel here to the multi-instructor problem: practical courses are often team-taught or jointly supervised, so even attributing a score is fraught, as we discuss in the multi-instructor attribution problem.

But isn't one common instrument necessary for comparability and fairness?

This is the strongest argument for the status quo, and administrators make it constantly: if every course uses a different form, we lose comparability, benchmarking becomes impossible, and we open the door to gaming — let unpopular courses pick easier questions. It is a real concern and deserves a real answer, not a wave of the hand.

But the argument rests on a false premise: that a single instrument delivers comparability. It does not — it delivers the appearance of comparability while quietly measuring different constructs in different settings. Scoring a lecture and a studio on "the instructor explained concepts clearly" does not make them comparable; it makes them incomparable in a way that is hidden, because the same item means different things in the two contexts. True comparability comes not from identical items but from a shared framework applied through context-appropriate questions — much as constructive alignment (which we cover in aligning evaluation to learning outcomes) holds the intent constant while the activities differ.

The defensible design is a common core plus a contextual module: a small set of cross-cutting items that genuinely apply everywhere (clarity of expectations, fairness of assessment, whether feedback was useful and actionable, overall) reported across all courses for legitimate high-level comparison, plus a mode-specific module (studio-crit items, clinical-supervision items, lab items) that captures what actually drives quality in that pedagogy and is compared only against like courses. That structure preserves the comparability administrators need and the validity the lecture-only form throws away. The choice is not "one form or chaos"; it is "honest partial comparability or dishonest total comparability". And the contextual modules are where programme-level versus course-level evaluation judgements should be grounded.

What good practice looks like

  • Map your evaluation to the pedagogy, not the timetable slot. Before evaluating a course, ask what its signature pedagogy is — exposition, critique, supervised practice, apprenticeship — and whether your instrument has anything relevant to ask about it.
  • Build a common core plus contextual modules. Keep a short genuinely-universal core for comparison; add mode-specific items for studios, labs, clinics, and placements.
  • Treat "not applicable" as information. A high N/A rate on lecture items in a studio course is not missing data — it is the instrument telling you it does not fit.
  • Lead with open questions in practical settings. The richest evidence about a crit, a placement, or a lab comes from students describing what happened, not rating a generic item. This is where qualitative depth earns its keep — and where the analysis cost has historically been prohibitive.

Where Koji fits

Adapting evaluation to many different pedagogies sounds operationally heavy — a different form for every mode, more open text than any committee can read. This is exactly the overhead Koji for Education is designed to remove. Because Koji runs AI-moderated conversational interviews rather than a fixed paper form, the evaluation can adapt to the course in real time: a studio student gets asked about the critique and their work-in-progress; a clinical student gets asked about supervision and patient exposure; a lab student about protocol, safety, and what they did when an experiment failed. The moderator follows the student's actual experience instead of forcing it through lecture-shaped items — while remaining standardized in how it probes, so the evidence stays comparable within each mode.

Koji's six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) make the common-core-plus-contextual-module design native rather than a spreadsheet workaround, and its automatic thematic analysis turns the heavy open-text load of practical-course feedback — exactly where the real signal lives — into structured, programme-level themes without a committee hand-coding crit transcripts. Its programme- and institution-level reporting can hold a shared core view across all courses while preserving mode-specific detail underneath. The same conversational interview engine powers the main Koji platform for organisations researching varied, hard-to-standardise human experiences.

Stated precisely: Koji does not make a studio and a lecture identical — they should not be — and it does not eliminate the difficulty of comparing across pedagogies. It removes the operational barrier to evaluating each pedagogy on its own terms, and surfaces the context-appropriate evidence that a one-size lecture form cannot.

The takeaway

The standard course evaluation is a lecture instrument wearing a universal badge. Used on the labs, studios, clinics, and placements that make up a large share of European higher education, it asks about activities that did not happen, stays silent on the supervision and feedback that actually define quality, and produces scores that are quietly unfair to compare against lectures. The fix is not a single perfect form and not evaluation anarchy, but a common core plus pedagogy-specific modules — holding the framework constant while matching the questions to how the teaching actually works. Signature pedagogies differ; the evaluation should too.

Running labs, studios, clinics, or placements through a form built for lectures? See how Koji for Education adapts the evaluation to each pedagogy while keeping the evidence comparable.