Teacher Clarity Predicts Learning Better Than Charisma: What Course Evaluations Should Measure
A meta-analysis of 144 effects and 73,000+ students shows teacher clarity explains roughly 13% of the variance in student learning. Here is what that means for the items you put on a course evaluation.
Koji Education Team
Product
In brief
If you want a course evaluation to say something about learning rather than about likeability, ask students about teacher clarity — whether explanations were structured, examples were relevant, and objectives were signposted — not just whether the instructor was engaging or enjoyable. The largest meta-analysis on the topic (Titsworth, Mazer, Goodboy, Bolkan & Myers, 2015, Communication Education) synthesised 144 effect sizes from more than 73,000 students and found that teacher clarity accounts for approximately 13% of the variance in student learning — a substantial, replicable relationship. Clarity is a low-inference, teachable behaviour, which makes it far more actionable feedback for an instructor than a global "overall satisfaction" number.
What the research says
Titsworth and colleagues (2015) conducted two meta-analyses on the clarity–learning relationship. The first pooled 144 reported effects across a combined sample of N = 73,281 students and estimated that teacher clarity is associated with roughly 13% of the variance in student learning. The second analysis (46 studies, N = 13,501) examined moderators and found that study design — survey versus experiment — moderated clarity's effect on affective learning, and that the relationship was generally stronger for affective learning (attitudes toward the subject and motivation to continue) than for cognitive learning (tested knowledge). The authors define clarity as the degree to which a teacher's verbal and written messages are structured, coherent, and comprehensible: clear objectives, logical sequencing, relevant examples, previews and summaries, and explicit signposting of what matters.
Clarity does not sit in isolation. It belongs to a family of instructional communication behaviours — clarity, immediacy (verbal and nonverbal warmth), and credibility — that have each been meta-analysed and reviewed as predictors of motivation and engagement (see Bolkan and colleagues' functional review, Frontiers in Psychology, 2021). This matters because clarity is frequently confused with charisma. The classic Dr. Fox effect (Naftulin, Ware & Donnelly, 1973) demonstrated that an expressive, confident lecturer delivering content-free nonsense could earn high ratings — a cautionary tale about mistaking expressiveness for substance. More recent work on the fluency illusion (for example, Carpenter and colleagues) shows that a smooth, polished delivery inflates students' perceived learning without a matching gain in actual learning. Clarity, by contrast, is the part of "good teaching" that most consistently tracks real learning outcomes.
A second anchor point is the multisection validity literature. Cohen's (1981) meta-analysis and the multisection paradigm — where many sections of the same course, using a common final exam, are correlated with their student ratings — established that specific dimensions such as "skill" and "structure/clarity" correlate more strongly with achievement than a single global rating does. The convergence is clear: when researchers isolate which teaching behaviours move the needle on learning, clarity is reliably near the top.
Why it matters for course evaluation in practice
The practical implication is about item selection and construct validity. Many institutional instruments still lean on a single global item ("Overall, how would you rate this instructor?") or on affect-laden prompts ("The instructor was enthusiastic/engaging"). Those items are easy to complete but they measure a blend of enjoyment, personality, and prior interest — constructs that are only weakly tied to learning and that are contaminated by well-documented biases. If your quality-assurance goal is to identify teaching that helps students learn, a global affect item is the wrong instrument.
Clarity-focused items have three advantages for a QA process:
- They are diagnostic. "The learning objectives for each session were clear" or "The instructor explained difficult concepts with useful examples" points a teacher at a concrete, changeable behaviour. "Overall satisfaction: 3.8/5" points nowhere.
- They are low-inference. Students are asked to report on observable behaviours (Were objectives stated? Were examples given?) rather than to render a summary judgement. Low-inference items reduce halo contamination and are more reliable across raters.
- They are actionable in a development conversation. A dean or educational developer can build a concrete improvement plan around clarity — restructure the module, add worked examples, preview and summarise — in a way that is impossible around a bare satisfaction score.
For programme-level quality assurance, clarity items also aggregate more meaningfully. A department that scores consistently low on "explanations were clear" has a specific, addressable signal; a department that scores low on "overall satisfaction" has a mood, not a diagnosis.
Limitations and honest caveats
A critical reader should hold several caveats in view. First, 13% of variance is a correlation, not a guarantee of causation. Much of the clarity literature is cross-sectional and survey-based, so reverse causation and third variables (student ability, prior knowledge, course difficulty) are live threats. Titsworth and colleagues themselves flag that design moderated the effects, and that experimental estimates can diverge from survey estimates.
Second, "learning" in this literature is heterogeneous. Affective learning (self-reported attitudes and motivation) and cognitive learning (test performance) are different constructs, and clarity's relationship is stronger with the former. Because affective learning is often measured by student self-report, some of the clarity–affect correlation may reflect shared method variance — a student who found a course clear will also report enjoying it. Self-reported learning is a weak proxy for measured learning, a point reinforced by the reference-bias and fluency-illusion literatures.
Third, clarity is necessary but not sufficient. A perfectly clear presentation of a poorly designed curriculum still produces limited learning. Clarity items should complement — not replace — evidence on constructive alignment, assessment quality, and workload.
Fourth, most of this evidence base is Anglophone and communication-discipline-centred, and much of it predates large online and blended cohorts. Generalisability to European multilingual, multi-campus settings should be treated as a hypothesis to test locally, not an assumption. Measurement invariance across languages and cohorts cannot be presumed.
How Koji incorporates this
Koji for Education is built around the premise that a single Likert number rarely captures why a course did or did not work — which is exactly the gap the clarity research exposes. Concretely:
- AI-moderated conversational interviews probe beyond the number. When a student rates clarity low, Koji's interviewer follows up conversationally — "Which part was hardest to follow, and what would have helped?" — surfacing whether the problem was sequencing, missing examples, unclear objectives, or pacing. This turns a diffuse clarity score into a diagnosis an instructor can act on, addressing the "13% of variance is real but a number alone is not actionable" problem head-on.
- Structured question types map to low-inference behaviours. Using
scale,single_choice, andyes_noitems, an institution can ask about the specific, observable clarity behaviours the meta-analysis identifies (stated objectives, worked examples, previews/summaries) rather than a global affect item — improving construct validity by design. - Automatic thematic analysis of open text clusters free-response comments so that "explanations were confusing" or "I never knew what the objectives were" emerge as named themes across a cohort, triangulating the quantitative clarity items against students' own words.
- Bias-aware reporting is designed to help committees separate clarity signal from the charisma/likeability halo that the Dr. Fox and fluency-illusion literatures warn about — flagging when open-text sentiment is dominated by personality language rather than learning.
- Mid-cycle (formative) collection lets a teacher fix a clarity problem while the course is still running, when it can still benefit the current cohort, rather than only after grades are in.
These are framed as mechanisms designed to mitigate the measurement problems clarity research identifies; no evaluation instrument, Koji included, eliminates confounds or converts self-report into measured learning. Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where "why behind the score" has the same diagnostic value.
Example clarity items you can adapt
Because clarity is low-inference, a good item names an observable behaviour and avoids a summary judgement. Institutions translating the research into an instrument often field prompts such as: "The instructor stated the learning objectives for each session clearly"; "Difficult concepts were explained with useful examples"; "The instructor previewed what was coming and summarised what had been covered"; "The structure of each session was easy to follow"; and an open prompt, "Which topic was hardest to follow, and what would have made it clearer?" Each maps to a component of clarity that the meta-analytic literature associates with learning, and each hands a teacher a concrete lever rather than a verdict. Pairing the scale items with the open prompt ensures that a low rating always arrives with an explanation the instructor can act on, which is exactly where a bare global-satisfaction number leaves them stranded.
Related Resources
- Low-inference teaching behaviors on course evaluations
- The Dr. Fox effect: instructor expressiveness and evaluations
- Instructor immediacy, warmth and the halo effect
- The lecture fluency illusion: perceived vs actual learning
- Cohen 1981: multisection validity of student ratings
- What student evaluations measure: Marsh multidimensionality
References
- Titsworth, S., Mazer, J. P., Goodboy, A. K., Bolkan, S., & Myers, S. A. (2015). Two meta-analyses exploring the relationship between teacher clarity and student learning. Communication Education, 64(4), 385–418. https://doi.org/10.1080/03634523.2015.1041998
- Bolkan, S., Goodboy, A. K., & others (2021). A functional review of research on clarity, immediacy, and credibility of teachers and their impacts on motivation and engagement of students. Frontiers in Psychology, 12, 712419. https://doi.org/10.3389/fpsyg.2021.712419
- Cohen, P. A. (1981). Student ratings of instruction and student achievement: A meta-analysis of multisection validity studies. Review of Educational Research, 51(3), 281–309. https://doi.org/10.3102/00346543051003281
- Naftulin, D. H., Ware, J. E., & Donnelly, F. A. (1973). The Doctor Fox lecture: A paradigm of educational seduction. Journal of Medical Education, 48(7), 630–635. https://doi.org/10.1097/00001888-197307000-00003
- Feldman, K. A. (1989). The association between student ratings of specific instructional dimensions and student achievement. Research in Higher Education, 30(6), 583–645. https://doi.org/10.1007/BF00992392
Related articles
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.
Stop Asking "Was the Lecturer Clear?": The Case for Low-Inference Teaching-Behaviour Items
High-inference items like "Is the instructor clear?" tell teachers nothing about what to change. Murray's low-inference behaviour research shows how concrete, observable items make evaluations diagnostic.
The Fluency Illusion: Why a Polished Lecturer Earns Higher Course Evaluations Without Teaching More
A research-grounded look at the lecture fluency illusion — why a smooth, confident presentation inflates student ratings and perceived learning while leaving actual learning unchanged — and what it means for interpreting course evaluations.
The Warmth Halo: How Instructor Immediacy Shapes Course Evaluations More Than Learning Does
What the meta-analytic evidence on teacher nonverbal immediacy and warmth tells us about course evaluations — why warmth strongly predicts how much students like a course and think they learned, but only weakly predicts what they actually learn.