Stop Asking "Was the Lecturer Clear?": The Case for Low-Inference Teaching-Behaviour Items
High-inference items like "Is the instructor clear?" tell teachers nothing about what to change. Murray's low-inference behaviour research shows how concrete, observable items make evaluations diagnostic.
Koji Education Team
Product
In brief
Most course-evaluation items are high-inference: they ask students to summarise an abstract quality ("Was the instructor clear / enthusiastic / organised?"). Harry Murray's research showed that the same teaching can be captured far more usefully through low-inference items that name concrete, observable behaviours ("signals the start of each new topic", "gives examples of difficult concepts"). Murray (1983) found that 26 specific low-inference behaviours, loading onto factors like clarity, enthusiasm and rapport, reliably distinguished instructors rated low, medium and high. The practical payoff is diagnostic: low-inference feedback tells a teacher what to do differently, where a high-inference score only tells them that something was wrong. Koji's structured questions and conversational probing are designed to elicit exactly this behaviour-level detail.
What the research says
The anchor is Harry G. Murray (1983), "Low-Inference Classroom Teaching Behaviors and Student Ratings of College Teaching Effectiveness" (Journal of Educational Psychology, 75(1), 138–149). Trained observers visited classes of university lecturers who had received low, medium or high student ratings, and recorded the frequency of roughly 60 specific, concrete teaching behaviours — things an observer can count without interpretation, such as "uses concrete examples", "repeats difficult ideas", "speaks expressively", "moves around the room", "addresses students by name". Findings:
- 26 individual behaviours differed significantly across the low/medium/high rating groups.
- The behaviours clustered into interpretable factors — notably clarity, enthusiasm and rapport — that distinguished effective from less effective instructors.
- The link between concrete observed behaviour and global student ratings was strong enough to suggest that student ratings, for all their flaws, are partly tracking real, specifiable teaching actions rather than pure impression.
The distinction Murray drew is the key idea: a high-inference item ("Is the instructor clear?") requires the rater to perform a large inferential leap and yields a number that is hard to act on; a low-inference item ("The instructor defined new terms when they were introduced") names a behaviour close to the observable surface, so it both measures more reliably and tells the teacher what specifically to change. Murray developed this into the Teaching Behaviors Inventory, and revisited the literature in a 2007 review chapter ("Low-Inference Teaching Behaviors and College Teaching Effectiveness: Recent Developments and Controversies").
This connects to a broader measurement tradition. The idea that ratings should be anchored to concrete, described behaviours rather than abstract traits is the basis of Behaviorally Anchored Rating Scales (BARS), introduced by Smith and Kendall (1963) and widely used in performance appraisal precisely because behavioural anchors reduce rater idiosyncrasy. It also dovetails with Marsh's case that teaching is multidimensional — see our note on what student evaluations actually measure.
Why it matters for course evaluation in practice
The dominant survey format — "Rate your agreement: the instructor was well-organised" — is high-inference by default. Switching toward low-inference items changes what evaluation can do:
- Feedback becomes actionable. "Score 3.4 on clarity" tells a lecturer nothing operational. "Students reported you rarely summarised at the end of class and rarely gave examples of hard concepts" points to a concrete change. Diagnostic specificity is the entire value proposition of formative evaluation (see do student evaluations improve teaching).
- Reliability improves. Because low-inference items require less interpretation, different students are more likely to agree on whether a behaviour occurred, narrowing the rater-idiosyncrasy that inflates measurement error in generalizability terms.
- It partially decouples 'quality' from 'likeability.' High-inference global items are highly exposed to the halo effect and the Dr. Fox effect, where charisma inflates ratings of substance. Asking about specific behaviours makes it harder for a single warm impression to flood every item.
- It supports developmental, not just judgmental, use. Behaviour-level data is what an academic-development unit can coach against; a global score is not.
Limitations and honest caveats
A careful reader should weigh several genuine objections:
- Students are not trained observers. Murray's data came from trained observers counting behaviours; ordinary students completing a survey may still infer or misremember whether a behaviour occurred. The reliability advantage of low-inference items is relative, not absolute, when the rater is an untrained student weeks after the fact.
- Behaviour lists can be context-bound. Effective behaviours in a small seminar (probing dialogue) differ from those in a 300-seat lecture (clear signposting, audibility). A single fixed inventory may fit some teaching modes poorly; the items must match the format.
- Low-inference does not equal valid. Counting behaviours assumes the listed behaviours actually cause learning. Some behaviours that correlate with high ratings may reflect expressiveness rather than learning gains — the very confound the Dr. Fox studies raise. Behaviour items still need validation against learning outcomes.
- Length cost. A thorough behaviour inventory is long, and length degrades data quality through fatigue and breakoff. There is a real trade-off between behavioural granularity and respondent burden.
- Generalizability. Murray's studies are decades old and largely North American; the behaviour–rating link is robust, but specific inventories should be checked against the local teaching culture and discipline.
How Koji incorporates this
Koji is designed to mitigate the diagnostic poverty of high-inference items, while being honest that students are not trained observers:
- Structured behaviour-level questions. Koji's question types (open_ended, scale, single_choice, multiple_choice, yes_no, ranking) let you replace a single "Was the instructor clear?" with concrete behavioural items ("Did the instructor summarise key points at the end of class?" "Were difficult concepts illustrated with examples?"), echoing Murray's Teaching Behaviors Inventory and the BARS logic of anchoring ratings to described behaviour.
- Conversational elicitation of specifics. This is the distinctive mechanism. When a student gives a vague global judgment, Koji's AI-moderated interview can probe for the underlying behaviour ("You said the course felt disorganised — what specifically made it feel that way?"), converting a high-inference impression into low-inference, actionable detail that a static survey cannot extract.
- Automatic thematic analysis into behavioural categories. Koji's analysis of open text and transcripts can cluster comments into concrete behavioural themes (signposting, pacing, examples, feedback timeliness), reconstructing behaviour-level signal even from narrative feedback.
- Triangulation with peer observation. Because behaviour-level student data aligns with what peer observers record, Koji's reporting can be triangulated against peer observation, strengthening validity rather than relying on student ratings alone.
We frame all of this as mitigation: conversational probing and thematic analysis improve the specificity of feedback, but they do not turn students into trained observers, and behaviour items still require local validation. Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where the move from abstract satisfaction scores to concrete behavioural detail is equally valuable.
Practical checklist
- Audit your instrument for high-inference trait words (clear, organised, enthusiastic, fair) and ask what observable behaviours sit beneath each.
- Add a small set of low-inference, behaviour-anchored items matched to the teaching format.
- Keep global items if you need a summary number, but treat the behavioural items as the diagnostic layer.
- Pair behaviour items with open text or conversational probing to capture format-specific behaviours your fixed list misses.
- Validate which behaviours actually relate to learning in your context before treating them as targets.
Related Resources
- What Do Student Evaluations Actually Measure? (Marsh & the SEEQ)
- Peer Observation vs Student Evaluations
- The Dr. Fox Effect: Does Charisma Inflate Ratings?
- The Halo Effect in Course Evaluations
- Designing Open-Ended Questions That Get Useful Answers
- Cognitive Interviewing for Course-Evaluation Questions
From trait to behaviour: a translation table
The fastest way to see the difference is to translate the abstract trait items most instruments use into the observable behaviours that sit beneath them. A few illustrative pairs:
- "The instructor was clear" → "Defined new terms when introduced", "Summarised the main points at the end of class", "Gave examples of difficult concepts", "Signalled the move from one topic to the next."
- "The instructor was enthusiastic" → "Spoke expressively rather than in a monotone", "Used gestures and movement", "Showed visible interest in the subject."
- "The instructor was well-organised" → "Stated the goals of the session at the start", "Followed a logical sequence", "Kept to the stated plan for the class."
- "The instructor had good rapport" → "Addressed students by name", "Invited and responded to questions", "Was available and approachable after class."
These map directly onto Murray's clarity, enthusiasm and rapport factors. The behavioural version measures more reliably and hands the teacher a concrete to-do list. Notice, too, that a behaviour can be present without producing learning — a vivid, gesture-rich lecturer scores high on the enthusiasm behaviours yet may still leave students underprepared, which is precisely the expressiveness confound the Dr. Fox studies expose. That is why behaviour items belong alongside, not instead of, evidence of actual learning.
Putting it to work without bloating the survey
A practical instrument keeps a small number of global items for trend reporting and benchmarking, then adds a focused set of behaviour-anchored items in the areas an institution most wants to develop — typically clarity and feedback. The behavioural layer is what academic-development units coach against; the global layer is what committees track over time. Reserve the heaviest behavioural detail for formative, mid-cycle collection, where its diagnostic value is highest and respondent burden is most justified.
References
- Murray, H. G. (1983). Low-inference classroom teaching behaviors and student ratings of college teaching effectiveness. Journal of Educational Psychology, 75(1), 138–149. https://doi.org/10.1037/0022-0663.75.1.138
- Murray, H. G. (2007). Low-inference teaching behaviors and college teaching effectiveness: Recent developments and controversies. In R. P. Perry & J. C. Smart (Eds.), The Scholarship of Teaching and Learning in Higher Education (pp. 145–183). Springer. https://doi.org/10.1007/1-4020-5742-3_6
- Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47(2), 149–155. https://doi.org/10.1037/h0047060
- Marsh, H. W. (1987). Students' evaluations of university teaching: Research findings, methodological issues, and directions for future research. International Journal of Educational Research, 11(3), 253–388. https://doi.org/10.1016/0883-0355(87)90001-2
Related articles
How to Design Open-Ended Course Evaluation Questions That Get Useful Answers
Open-ended evaluation questions usually fail not because students have nothing to say, but because the prompt asks for too little. The survey-methodology evidence on answer-box size, verbal instructions, and specific framing — and what it means for the comment box.
The Dr. Fox Effect: Does Lecturer Charisma Inflate Course Evaluations?
The 1973 Dr. Fox experiment, its meta-analytic re-interpretation, and the 2014 re-revisitation — what the evidence really says about whether expressive delivery seduces students into rating empty teaching highly, and how to evaluate substance over showmanship.
The Halo Effect in Course Evaluations: When One Impression Colours Every Rating
When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.