New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology10 min read

The Jingle-Jangle Trap: What Does "Teaching Effectiveness" Even Mean?

Two century-old measurement fallacies quietly undermine most course-evaluation scores. If your instrument does not define the construct it claims to measure, a single "overall effectiveness" number is measuring something — but not what you think.

Koji Education Team

Product ·

Bottom line up front: Before you can ask whether a course-evaluation score is biased, reliable or fair, you have to answer a prior question that most instruments skip: what construct is it measuring? "Teaching effectiveness" sounds self-evident, but it is a textbook site of two century-old measurement errors — the jingle fallacy (assuming one label names one thing) and the jangle fallacy (assuming two labels name two things). Ignore them, and your headline "overall effectiveness" item is measuring something real, but almost certainly not what its name promises.

Two fallacies, one measurement problem

The terms come from early psychometrics. The jingle-jangle fallacy pairs Edward Thorndike's 1904 warning about the jingle fallacy — believing that two things called by the same name are therefore the same — with Truman Kelley's 1927 jangle fallacy — believing that two things with different names are therefore different.

Applied to course evaluation, both are everywhere:

  • The jingle fallacy in action: Two universities each field an item called "overall teaching effectiveness." One instrument's respondents interpret it as "did I enjoy this?"; another's as "did I learn?"; a third's as "was the lecturer clear and organised?" The shared label hides three different constructs. Benchmarking one against another — a routine cross-institution comparison — silently assumes a sameness that does not exist.
  • The jangle fallacy in action: An instrument asks separately about "clarity," "organisation," "communication" and "explains concepts well," treats them as four dimensions, and reports four numbers — when factor analysis would show they are one underlying factor wearing four names. The dashboard implies a differentiated profile that the data cannot support.

Neither fallacy is exotic. They are the default failure mode of any instrument written by intuition rather than construct definition — which describes a large share of the student-evaluation-of-teaching (SET) forms in use across European higher education.

Why the construct question comes first

Validity, in modern measurement theory, is not a property of a test but of the interpretation of its scores. You cannot say a course-evaluation number is valid or invalid in the abstract; you can only ask whether the meaning you attach to it is defensible. If you have not specified what "effectiveness" denotes — satisfaction, perceived learning, actual learning, instructor behaviours, or some tangle of all four — then no downstream statistic can rescue you. This is the deeper point behind the well-worn debate about whether course evaluations are valid at all: much of the confusion is really a construct-definition failure masquerading as a reliability problem.

The most consequential jingle in the entire field is the equation of satisfaction with learning. Institutions collect satisfaction — how students felt about a course — and then interpret it as if it measured teaching effectiveness in the sense of learning caused. The evidence that these are different constructs is now strong. Uttl, White and Gonzalez's 2017 meta-analysis, Meta-analysis of faculty's teaching effectiveness: student evaluation of teaching ratings and student learning are not related, re-analysed the multisection-study literature and found that once you account for small-study effects and publication bias, there is essentially no correlation between SET ratings and actual learning. The label "effectiveness" jingles two constructs into one; the data pull them back apart.

This is not a fringe finding. It aligns with a broader pattern — the documented active-learning penalty, where courses that demonstrably improve learning can receive lower ratings because effortful learning feels worse in the moment. If satisfaction and learning were the same construct, that could not happen. That it happens routinely is the jingle fallacy caught in the act.

The jangle fallacy and the illusion of a profile

The mirror error is just as costly for reporting. Marsh's influential work on students' evaluations established that teaching, as students perceive it, is genuinely multidimensional — but only when items are written and validated to separate distinct factors. Most operational SET forms are not. They list a dozen positively-worded items that all load on a single "general impression" factor, then report them as if each were an independent dimension. The result is a spider chart or a table of sub-scores that looks diagnostic — "strong on clarity, weaker on feedback" — when the underlying data are one halo-driven number in four costumes. Committees make personnel and curriculum decisions on these phantom distinctions. The jangle fallacy manufactures precision that the instrument never possessed.

But isn't this just academic hair-splitting?

The fair objection: institutions need a workable number to run quality assurance at scale, and demanding a fully-specified construct model for every evaluation item is a counsel of perfection that would paralyse practice. Deans do not have time for factor analysis; they have committees to feed.

Two responses. First, the fallacies are not a reason to abandon measurement — they are a reason to name what you measure honestly. It costs nothing to relabel an "overall teaching effectiveness" item as an "overall satisfaction" item if satisfaction is what it captures; the paralysis comes from the pretence, not from the rigour. Honest, modest labels are more usable, not less, because they stop committees from over-reading. Second, the practical damage is real, not theoretical: when satisfaction is mislabelled as effectiveness, the active-learning penalty systematically punishes exactly the teaching institutions claim to want, and phantom sub-scores drive unfair comparisons between colleagues. Getting the construct right is not perfectionism; it is the difference between a number that helps and a number that misleads while looking authoritative.

Escaping the trap: measure the construct, not the label

The way out is not a better single number — it is a shift from scoring a label to understanding a construct:

  • Define before you measure. Decide whether a given item is about satisfaction, perceived learning, or specific teaching behaviours, and label it as such. Do not let one word carry three meanings.
  • Separate feeling from learning. Collect satisfaction if you want satisfaction, but do not interpret it as learning. Where you care about learning, seek evidence of it directly rather than inferring it from enjoyment.
  • Earn your dimensions. Only report sub-scores as distinct if they are shown to be distinct. A validated multidimensional instrument is defensible; a dozen halo items presented as a profile is not.
  • Prefer explanation over a rating. A number tells you how much; it cannot tell you of what. Qualitative depth is what disambiguates the construct — it tells you whether "effective" meant clear, engaging, demanding or merely pleasant.

How Koji addresses the construct problem

Koji for Education is built to attack the jingle-jangle trap at its root, because a conversational interview does not have to compress a complex construct into one ambiguous word. Its AI moderator can probe what a rating actually means — following an "it was effective" with "effective in what sense: did you understand the material better, or did you enjoy it more?" — so satisfaction and learning are pulled apart in the conversation rather than jingled together on a form. Koji's six structured question types let institutions ask precisely-scoped questions rather than one catch-all effectiveness item, and its automatic thematic analysis reads the open-text explanation that reveals which construct a student had in mind. The point is not a cleverer average; it is that Koji surfaces the meaning behind the number, which is exactly what the jingle and jangle fallacies destroy. The same interview engine powers construct-aware research on the main Koji platform, where "what do users actually mean by this rating?" is the central question.

The takeaway

"Teaching effectiveness" is not a measurement you can take; it is a construct you must define. Every course-evaluation number rests on an unstated answer to "effective at what?" — and when that answer is left to a single ambiguous word, the jingle fallacy fuses satisfaction with learning while the jangle fallacy splinters one halo into a fake profile. The remedy is not more decimal places. It is the discipline to name what you measure, and the depth to understand what students actually meant.

Want evaluation that captures meaning, not just a label? Explore Koji for Education.