Research methods
Course evaluation methodology, bias reduction, and quality assurance frameworks.
Racial and Ethnic Bias in Student Evaluations of Teaching: What the Evidence Shows
Peer-reviewed evidence shows students rate instructors of colour systematically lower than White peers on the same teaching. We review Reid (2010), Chávez & Mitchell (2020) and Bavishi et al. (2010), the limitations, and how to design evaluation so race-linked bias does not contaminate quality decisions.
Does the Room Bias the Rating? Physical Classroom Environment as a Confound in Course Evaluations
Evidence that the physical classroom - lighting, seating, comfort, technology - shifts student evaluation of teaching scores independent of instructional quality, and how to stop the estate from contaminating the teaching signal.
Is the Weak SET-Learning Correlation an Artefact of Unreliable Measures? What Correction for Attenuation Does — and Doesn't — Prove
Disattenuation lets you estimate what a correlation would be if both measures were perfectly reliable. It is a legitimate tool that both sides of the student-ratings debate have used — and abused — to argue the true SET-learning link is stronger or weaker than the raw number suggests.
Do Early-Morning Classes Get Lower Evaluations? Time-of-Day and Scheduling as Confounds
Evidence that the timetable slot - early morning, late afternoon, days per week - shifts student performance and mood, and therefore course evaluation scores, independent of teaching quality. What QA teams should record and adjust for.
Do Students Who Skip Class Rate Teaching Differently? Attendance as a Confound in Course Evaluations
Attendance and the perceived need to attend are entangled with course-evaluation scores - and with who fills the survey out. Grounded in Burns and Ludlow (2005), this explains the confound and how to keep it from distorting SET.
Evaluating Doctoral Supervision: What PRES and Supervision-Experience Measures Actually Capture
A research-grounded look at how the postgraduate research (PGR) supervision experience is measured, what the UK PRES and validated instruments like the QSDI capture, and where those measures fall short.
Evaluating Team-Taught Courses: The Attribution Problem in Student Evaluations
Why a single overall rating conflates co-teachers, and how to design evaluations that separate instructor-level from course-level signal in team-taught courses.
Why Was the Course Rated Well, Not Just Whether? Mediation Analysis and the Mechanism Behind an Evaluation Score
A high rating tells you students were satisfied; it doesn't tell you why. Mediation analysis tests the pathway — did clarity raise satisfaction by increasing engagement? — but drawing that causal chain from a single end-of-term survey is far harder than the classic recipe suggests.
Evaluating Clinical Placements and Work-Based Learning: Why the End-of-Course Questionnaire Is the Wrong Instrument
Why standard end-of-course student evaluation questionnaires fail to capture placement quality, and what validated clinical/practice-learning-environment instruments (MCPI, CLES+T, UCEEM, PET, CLEI, DREEM) measure instead.
Do Adjuncts Get Worse Course Evaluations Than Tenured Faculty? Employment Status as a Confound
The causal evidence says contingent faculty do not teach worse — and often teach better. Why employment status is a poor proxy for teaching quality, and how to read evaluation scores fairly across the tenure divide.
Do Older Professors Get Lower Course Evaluations? Age and Seniority as a Confound
Stonebraker and Stone found a small but robust negative effect of instructor age on student ratings, emerging after the mid-forties. What the evidence does and does not show, and how to stop age from contaminating comparisons.
Does a Better Researcher Make a Better Teacher? The Teaching-Research Nexus and Your Evaluations
Hattie and Marsh found the correlation between research productivity and teaching quality is essentially zero. What that means for reading course evaluations and for keeping teaching and research evidence separate.
The Critical Incident Technique for Course Feedback
How Flanagan's Critical Incident Technique collects concrete, behaviourally-anchored student feedback that global Likert ratings cannot capture, and how it applies to course evaluation.
Q-Methodology: Surfacing the Distinct Viewpoints Students Hold About a Course
Q-methodology uses forced-choice card sorts and by-person factor analysis to reveal the two-to-four genuinely distinct viewpoints students hold about a course, a rigorous complement to Likert SET averages.
Beauty Bias in Course Evaluations: Are Attractive Instructors Rated as Better Teachers?
What the research says about physical-attractiveness bias in student evaluations of teaching — from Hamermesh & Parker (2005) to German and laboratory replications — and how to design evaluation so appearance does not masquerade as teaching quality.
Why Students Click Straight Down the Middle: Satisficing in Course Evaluations
A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.
Why "The Pace Was About Right" Breaks Your Scale: Ideal-Point Unfolding Models for Course Evaluation
Most rating models assume more of a trait always means more agreement. For "appropriateness" items — workload, pace, difficulty — that assumption is false, and it quietly wrecks the scale. Ideal-point unfolding models fix it.
Contribution Analysis: Making Credible Causal Claims from Course Evaluation
You changed a course and outcomes improved — but did the teaching cause it? Contribution analysis is a theory-based method for making defensible causal claims when a randomised trial is impossible. Here is how it applies to course evaluation.
The Most Significant Change Technique: Story-Based Course Evaluation That Surfaces What Students Actually Value
Likert averages tell you where a course sits; they cannot tell you what transformed a student. The Most Significant Change technique collects and collectively selects stories of change to reveal unanticipated outcomes and shared values. Here is how it works and where it fits course evaluation.
SALG: Can Students Reliably Report Their Own Learning Gains?
The SALG instrument asks about learning gains, not satisfaction. What Seymour and colleagues built, what the validity evidence (and Porter''s and Bowman''s critiques) show about self-reported gains, and how to use gains-framed questions responsibly.
The One-Minute Paper: Formative Feedback That Also Improves Learning
The one-minute paper and the muddiest point are the best-known Classroom Assessment Techniques. What Chizmar & Ostrosky (1998) and Stead (2005) found about their effect on learning and feedback, and how to run continuous formative collection well.
The Success Case Method: Evaluating a Course Through Its Best and Worst Cases
Brinkerhoff''s Success Case Method evaluates a course by studying its most and least successful students, not its average. What the approach is, its evidence and biases, and how to run the two-stage screen-then-interview design.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.
Course Evaluations Measure Reaction, Not Learning: What Kirkpatrick's Four Levels Reveal
Kirkpatrick's four-level model explains why an end-of-term course evaluation is a Level-1 'reaction' measure — and why decades of meta-analytic evidence show reaction correlates almost nothing with actual learning.
Fixing Nonresponse While You Field, Not After: Responsive and Adaptive Survey Design for Course Evaluation
Responsive and adaptive survey design monitors who is answering while a course evaluation is still open and reallocates effort toward under-represented students to reduce nonresponse bias — a step earlier than post-hoc weighting.
Beyond the End-of-Term Survey: Stufflebeam's CIPP Model for Programme-Level Evaluation
Most course evaluation stops at student satisfaction. Stufflebeam's CIPP model (Context, Input, Process, Product) reframes evaluation as decision-support across the whole programme lifecycle — and maps cleanly onto ESG/ENQA quality-cycle thinking.
When, Not Just Whether, Students Leave: Cox Proportional-Hazards Models for Course Evaluation
Most evaluation analytics ask whether a student withdrew. Cox proportional-hazards regression asks when — modelling the risk of withdrawal over time as a function of course, cohort, and early-experience signals — and why that is a sharper quality question than a pass/fail dropout flag.
Do Your Evaluation Items Actually Cover Teaching? Content Validity, the CVR and the CVI
An evaluation form can be reliable and still measure the wrong things. Content validity, quantified with Lawshe's CVR and the Content Validity Index, tests whether your items cover the teaching domain.
Not All Course Attributes Are Equal: The Kano Model and the Asymmetry of Student Satisfaction
A course-evaluation mean assumes every attribute affects satisfaction the same way. The Kano model shows it does not: some attributes only cause dissatisfaction when they are missing, others only delight when present. We explain must-be, one-dimensional and attractive quality — and why averaging hides them.
Can You Train the Bias Out of Evaluation Reviewers? Frame-of-Reference Rater Training
The committees and peer observers who read evaluations carry halo, contrast and attribution biases. Frame-of-reference training is the evidence-based method for making their judgments more accurate.
Does Evaluating Course After Course Change How Students Answer? Panel Conditioning in Repeated Course Evaluations
Students at a European university complete dozens of course evaluations across a degree. Panel-conditioning research shows that the mere act of being surveyed repeatedly can change later answers — a threat to comparing scores across years and cohorts.
The Nominal Group Technique: Structured Group Feedback That Ranks What Students Actually Care About
A survey tells you how students rated a course; it rarely tells you what they would fix first. The Nominal Group Technique — a structured, facilitated method — generates and prioritises student feedback with the whole cohort in the room, and has a track record in higher-education course evaluation.
Stop Arguing About Wording — Test It: Split-Ballot Experiments for Course-Evaluation Questions
Committees spend hours debating whether to word an evaluation item one way or another. Schuman and Presser showed that small wording changes can move survey answers by double-digit margins — and that the way to settle the debate is a randomised split-ballot experiment, not opinion.
Is Your Course Pushing Students Toward Deep or Surface Learning? The R-SPQ-2F as an Evaluation Lens
Most course evaluations ask whether students liked the teaching. The deep/surface approaches tradition asks a more consequential question: did the course lead students to engage meaningfully or just memorise to pass? The R-SPQ-2F instrument makes that measurable.
Numbers First, Then the Why: Explanatory Sequential Mixed Methods for Course Evaluation
A Likert average tells you a course scored 3.4; it never tells you why. The explanatory sequential mixed methods design (Creswell & Plano Clark) fixes this by using a quantitative survey to decide exactly which students to follow up with qualitatively — turning an unexplained number into an evidenced account.
Your Course Evaluation Measures Satisfaction, Not Emotion. Pekrun Says That's a Problem
Standard course evaluations ask whether students were satisfied. Pekrun's control-value theory and the Achievement Emotions Questionnaire show that discrete emotions — enjoyment, boredom, anxiety, hope, hopelessness — drive learning and are absent from almost every institutional survey. Here is why that gap matters and how to close it.
Narrative Inquiry for Course Evaluation: Reading Student Stories as Wholes, Not Fragments
Thematic coding chops student feedback into fragments and loses the plot. Narrative inquiry keeps each student's course experience whole and temporal, surfacing turning points that a code frequency cannot.
Should Course Evaluations Ask Whether Teaching Matched a Student's Learning Style?
Learning styles are one of the most durable myths in education, and Pashler et al. (2008) found no credible evidence for tailoring instruction to them. Here is why a course-evaluation item that asks students whether teaching 'suited their learning style' quietly measures a debunked construct — and what to ask instead.
What Cognitive Load Theory Says Your Course Evaluation Should — and Shouldn't — Ask
Cognitive load theory (Sweller, van Merriënboer & Paas, 2019) distinguishes the unavoidable difficulty of content from difficulty caused by poor design. That distinction changes what a course evaluation should measure: not overall 'difficulty', but the design choices that impose or remove extraneous load.
Evaluating Active Learning: The ICAP Framework as a Course-Evaluation Lens
Most course evaluations ask whether a course was 'engaging' — a word that conflates enjoyment with learning. Chi and Wylie's (2014) ICAP framework replaces it with an observable ladder of cognitive engagement (Passive, Active, Constructive, Interactive), giving evaluation items that measure what students actually did.
Desirable Difficulties: Why the Teaching That Improves Learning Often Lowers Satisfaction
Spacing, interleaving, and retrieval practice are among the best-evidenced ways to make learning durable — and they make a course feel harder and less smooth in the moment. Bjork & Bjork's (2011) desirable-difficulties principle explains why end-of-term satisfaction ratings systematically penalise the most effective teaching.
Does Course Difficulty and Workload Lower Student Evaluations? What Centra Found
A research-grounded look at whether harder, heavier courses are punished in student evaluations of teaching. Centra (2003) and Marsh & Roche (2000) show the relationship is non-linear and weaker than faculty fear — with practical implications for fair course evaluation.
Is Your Course Actually Aligned? Using Biggs's Constructive Alignment as an Evaluation Lens
Biggs's constructive alignment says learning outcomes, teaching activities, and assessment must point the same way. Here is how to turn that theory into a course-evaluation instrument that diagnoses misalignment students can feel but rarely name.
Does Your Course Support Autonomy, Competence, and Relatedness? Self-Determination Theory as an Evaluation Lens
Self-Determination Theory says motivation depends on three basic needs — autonomy, competence, and relatedness. Here is how to evaluate a course by whether it feeds or starves those needs, instead of only whether students were satisfied.
Stop Waiting for the End of Term: Experience Sampling for In-the-Moment Course Feedback
Experience-sampling methods capture what students feel and think during a course, not their reconstructed memory of it months later. Here is why in-the-moment data can be more valid than the end-of-term survey — and how to use it responsibly.
Why Quantitative Courses Get Lower Evaluations: The Discipline-Bias Problem
Uttl & Smibert (2017) show that instructors of quantitative courses receive systematically lower student evaluations than those teaching qualitative subjects — a bias with real career consequences. What the evidence says and how to compare ratings fairly across disciplines.
Focus Groups for Course Evaluation: What They Reveal That Surveys Miss (and Where They Fail)
What the methodological literature — Stalmeijer et al.'s AMEE Guide No. 91, Kaplowitz & Hoehn's comparative study, and randomized method comparisons — says about using focus groups for course evaluation, their known failure modes, and how AI-moderated individual interviews capture the depth without the group-dynamics distortions.
Do University Teachers Get Better With Experience? A 13-Year Growth-Model Study Says: Not Automatically
Herbert Marsh's multilevel growth model of student evaluations over 13 years found teachers' ratings neither improve nor decline with experience — individual differences dominate. What this means for tenure assumptions, development policy, and how evaluation data should track trajectories.
The Delphi Method for Course and Programme Evaluation: Building Expert Consensus Without the Loudest Voice Winning
How the Delphi technique produces structured, anonymous expert consensus on course-evaluation criteria and programme standards — its evidence base, its limitations, and where it fits alongside student feedback.
Beyond the Student Survey: Structured Classroom-Observation Protocols (COPUS and TDOP) as Evaluation Evidence
What COPUS and the TDOP measure that student surveys and traditional peer visits cannot: low-inference, reliable records of what actually happens in a classroom. Their evidence base, their limits, and where they fit in a triangulated evaluation.
Which Item Is Biased? Differential Item Functioning Detection for Course-Evaluation Questions
Measurement invariance tests the whole scale; differential item functioning (DIF) pinpoints the single biased item. A practical guide to Mantel-Haenszel and logistic-regression DIF, uniform vs non-uniform bias, and what to do when a course-evaluation item behaves differently across groups.
Do the Clicks Confirm the Comments? Triangulating Course Evaluations with LMS Learning-Analytics Data
Learning-analytics trace data from your VLE looks like an objective check on what students say in course evaluations. The research says it is a weaker and more course-specific signal than most institutions assume — here is how to triangulate it honestly.
How Big a Difference Can You Actually Detect? A-Priori Power Analysis for Course Evaluation
Reliability tells you how stable a score is; statistical power tells you whether your sample can detect a real difference at all. A practical guide to a-priori power analysis, minimum detectable effects, and why most single-class course-evaluation comparisons are underpowered before they begin.
Should Your Course Evaluation Ask How Students Used AI? What the Learning Evidence Says
Students now study with generative AI in nearly every course, and it changes what a satisfaction rating means. The evidence on AI and learning explains why — and what a course evaluation should and should not try to ask about it.
Asking the Questions Students Won't Answer Honestly: The Randomized Response Technique for Sensitive Course-Evaluation Items
When a course or climate survey asks about harassment, discrimination, academic misconduct, or truancy, ordinary anonymity is not enough — students still under-report. The randomized response technique adds provable privacy so respondents can answer honestly. What Warner (1965) proposed, what the validation evidence shows, and where it fails.
Not Every Cohort Follows the Same Path: Latent Growth Curve and Growth Mixture Models for Course Evaluation
Latent growth curve and growth mixture models describe how student trajectories change across evaluation waves — and uncover subgroups that a single average path hides — while guarding against the subgroups that are only statistical artefacts.
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.
Will Students Open Up to an AI That Runs Their Course Evaluation? Algorithm Aversion, Appreciation, and Disclosure
Do students trust an AI to conduct their course evaluation, and does an AI moderator make them more or less candid? The evidence on algorithm aversion, appreciation, and disclosure-to-machines is more encouraging than the sceptics assume.
Does Your Course Build Self-Regulated Learners? Metacognition and SRL as an Evaluation Lens
Self-regulated learning — the cycle of planning, monitoring and reflecting — predicts academic achievement, yet standard course evaluations never ask whether a course developed it. Here is the SRL evidence and how to turn it into evaluation questions.
Should Course Evaluations Measure Whether Students Feel They Belong? The Sense-of-Belonging Evidence
Sense of belonging predicts persistence, engagement, and mental health in higher education — often more reliably than satisfaction. Here is what the evidence says about measuring belonging in course evaluation, and how to do it without turning a survey into a diagnostic instrument.
Should Course Evaluations Ask Whether a Course Fostered a Growth Mindset? What the Meta-Analytic Evidence Says
Growth-mindset interventions show weak average effects in two large meta-analyses. Here is what that means for whether — and how — course evaluations should ask about mindset.
Did the Course Build Students Who Can Use Feedback? Feedback Literacy as an Evaluation Lens
Carless and Boud's feedback-literacy framework reframes a course-quality question: not "was feedback given?" but "did the course develop students who can appreciate, judge, manage, and act on feedback?"
Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.
One Score, Two Things: Item Response Tree Models Separate Opinion from Response Style
Item response tree (IRTree) models split each Likert answer into latent decisions — respond or stay neutral, agree or disagree, moderate or extreme — so a student's real opinion of a course can be estimated apart from their personal scale-use style.
The Halo Effect in Course Evaluations: When One Impression Colours Every Rating
When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.
Stop Asking Students to Rate Everything Highly: Discrete Choice Experiments Reveal What They Will Trade Off
Discrete choice experiments (DCEs) ask students to choose between course scenarios rather than rate each feature, forcing real trade-offs and yielding a random-utility estimate of what actually drives their preferences.
Bifactor and ESEM Models: When a Multidimensional Course Evaluation Still Justifies One Overall Score
Bifactor and exploratory structural equation models let you estimate a general teaching-quality factor and specific factors at once, then test with omega-hierarchical and ECV whether reporting a single overall course-evaluation score is defensible.
Did the Course Spark Interest? The Four-Phase Model of Interest Development as an Evaluation Lens
Most evaluations ask whether students enjoyed a course. Hidi and Renninger's four-phase model distinguishes a momentary spark from durable individual interest — a sharper, more consequential thing to measure.
When Everything Scores 4/5: Best-Worst Scaling (MaxDiff) for Course-Evaluation Priorities
Likert ratings on course evaluations cluster near the top and cannot tell you what matters most to students. Best-worst scaling (Louviere, Flynn & Marley) forces trade-offs that reveal genuine priorities. What the method is, its limits, and how it fits a Koji study.
Let Students Name Their Own Criteria: The Repertory Grid Technique for Course Evaluation
Standard evaluation forms impose the institution's categories on students. The repertory grid technique, built on Kelly's personal construct theory, elicits the dimensions students themselves use to judge a course — surfacing criteria a fixed questionnaire never asked about.
Before You Field It, Test It: Cognitive Interviewing for Course-Evaluation Questions
Why the wording of a course-evaluation item should be cognitively pretested before it reaches students, what Beatty and Willis (2007) established about think-aloud and verbal probing, and how Koji operationalises probing at scale.
Do Cookies, Treats, and Mood Bias Course Evaluations?
Two controlled studies show that giving students chocolate or cookies before an evaluation measurably raises teaching scores. What the affect heuristic means for the validity of course evaluations — and how to design around it.
Do Instructors and Students Agree? Self-Evaluation vs Student Ratings
Feldman's synthesis found instructor self-ratings and student ratings correlate only moderately (around r ≈ 0.3). What weak self–student agreement means for triangulation, faculty trust, and how to use both sources without privileging either.
Does a Conversational Course Evaluation Make Students Less Honest? Mode Effects and Social Desirability
Survey mode shapes honesty: interviewer-administered surveys invite more social-desirability bias than self-administered ones. What Tourangeau and Yan (2007) and the mode-effects literature mean for anonymous, AI-moderated conversational course evaluations.
Measuring What Students Will Not Admit Directly: The List Experiment for Sensitive Course-Evaluation Questions
Some of the most important evaluation questions - Did you actually attend? Did you use AI on assessments? Did the grade you expected shape your rating? - are exactly the ones students answer dishonestly. The list experiment (item-count technique) estimates their prevalence without ever asking any student to admit anything.
Do the Numbers on Your Rating Scale Change the Score? The Evidence on Numeric Anchors
Schwarz and colleagues showed that the numeric values printed on a rating scale (0 to 10 vs minus 5 to plus 5) systematically shift responses even when the verbal labels are identical. Here is what that means for course-evaluation design, comparability, and reporting.
Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
Online course evaluations consistently draw lower response rates than in-class paper forms, but the research shows the resulting scores are largely equivalent. Here is what Dommeyer, Nulty, and Stowell actually found, and what an adequate response rate really requires.
Did the Course Produce Flow? Csikszentmihalyi's Flow Theory as a Course-Evaluation Lens
Csikszentmihalyi's flow theory reframes course evaluation around whether teaching created states of deep, balanced absorption. Here is what the evidence supports, where it breaks down, and how to measure it responsibly.
Did Students Leave More Confident They Can Succeed? Academic Self-Efficacy as a Course-Evaluation Outcome
Academic self-efficacy is among the strongest psychological correlates of student achievement. Here is what the evidence says, why it belongs in course evaluation, its limits, and how to measure it without kidding yourself.
The Last Week Counts Double: The Peak-End Rule and What Students Actually Remember About Your Course
Retrospective memory of an experience is dominated by its most intense moment and its ending, not its average or its length. Here is what the peak-end rule means for end-of-semester course evaluations.
Stop Asking "Was the Lecturer Clear?": The Case for Low-Inference Teaching-Behaviour Items
High-inference items like "Is the instructor clear?" tell teachers nothing about what to change. Murray's low-inference behaviour research shows how concrete, observable items make evaluations diagnostic.
How Often Is "Often"? Why Vague Quantifiers Quietly Distort Course-Evaluation Data
Words like "often", "usually" and "sometimes" mean different things to different students, which makes course-evaluation responses hard to compare. Here is the survey-methodology evidence and what to do about it.
Sliders, Visual-Analogue, or Radio Buttons? The Evidence on Course-Evaluation Response Formats
Slider widgets look modern, but the survey-methodology evidence says they cost you data. Funke (2016), Couper et al. (2006) and Bosch et al. (2019) on choosing a response widget for online course evaluations.
Does the Professor Students Rate Highest Teach Them Most? Value-Added Evidence vs Student Ratings
The strongest causal evidence — random-assignment studies that track follow-on course performance — shows the instructors students rate highest are often not the ones whose students learn most over time. What Carrell & West (2010) and Braga et al. (2014) found, and what it means for course evaluation.
Which Evaluation Questions Are Redundant? Mutual Information and mRMR for Trimming Your Questionnaire
Two evaluation items are redundant to the degree that one predicts the other — mutual information. The mRMR rule keeps items that are maximally relevant to your decision and minimally redundant with each other, yielding a shorter, sharper questionnaire.
Is a Low Response Rate Biasing Your Course Evaluations? MCAR, MAR, and Missing-Not-at-Random
A low response rate does not automatically bias a course-evaluation score — the missing-data mechanism (MCAR, MAR, or MNAR) does. How to diagnose nonresponse bias instead of chasing a rate threshold.
Do Elective Courses Get Higher Ratings Than Required Ones? Course Characteristics and the Feldman Evidence
Course features the instructor cannot control — whether a course is elective or required, its level, size, subject, and timing — systematically shift student ratings. What Kenneth Feldman's synthesis established, and why comparing raw ratings across different course types is unfair.
One Number or Many? The Dimensionality Debate and How to Use Student Ratings for Personnel Decisions
Should a promotion committee use a single global teaching score or a detailed profile of many dimensions? The 1997 d'Apollonia & Abrami vs Marsh & Roche debate set the terms — and the answer depends on whether the purpose is summative judgement or formative improvement.
Deleting Incomplete Surveys Throws Away Good Data: Multiple Imputation for Missing Course-Evaluation Responses
Listwise deletion of partially-complete course evaluations shrinks your sample and can bias results. How multiple imputation (Rubin 1987; Schafer & Graham 2002) handles item nonresponse honestly.
The Qualitatively Different Ways Students Experience Your Course: Phenomenography
A standard survey tells you how many students rated "feedback" a 3. Phenomenography, the higher-education research tradition founded by Ference Marton, instead maps the qualitatively different ways students experience a phenomenon — and explains why the same number means different things.
Goal-Free Evaluation: Why Scriven Says Course Reviews Should Ignore the Stated Objectives
Placeholder
Agree/Disagree or Item-Specific? The Scale Choice That Quietly Degrades Course-Evaluation Data
"The lecturer was well organised: Strongly disagree to Strongly agree" feels natural, but the agree/disagree format invites acquiescence and lower data quality. Saris et al. (2010) on why item-specific scales measure better.
Response-Order Effects: Does Where an Answer Sits Change How Students Rate Your Course?
The order in which answer options appear can shift course-evaluation responses independent of what students think. We unpack Krosnick & Alwin (1987) on primacy and recency, why it matters for instrument design, and how to limit it.
Is 4.2 Really Worse Than 4.4? Bayesian Estimation and Credible Intervals for Course Evaluations
Bayesian estimation answers the question institutions actually ask — how probable is it that this instructor is below standard? — with credible intervals and a region of practical equivalence instead of raw averages or p-values.
Can Instructors Buy Better Ratings With Easy Grades? Instrumental Variables and the Grade–Evaluation Endogeneity Problem
Grades and ratings are jointly determined, so a naive regression overstates the grade effect. How Krautmann & Sander (1999) used two-stage least squares to recover the causal effect of expected grades on evaluations.
What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.
The Student-as-Consumer Effect: What a Consumer Mindset Does to Course Evaluations
When students see themselves as paying customers, what happens to how they rate courses — and how they learn? A look at Bunce, Baird & Jones (2017) and why the consumer frame quietly distorts the meaning of satisfaction scores.
Is Student Evaluation of Teaching Valid? What the State-of-the-Art Review Concludes
A close reading of Spooren, Brockx and Mortelmans'' 2013 Review of Educational Research synthesis of SET validity — what it actually concludes, the evidence behind it, and what it means for how universities use student ratings.
How Many Scale Points Should a Course-Evaluation Question Have?
What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.
Is One "Overall" Question Enough? Single-Item vs Multi-Item Measures in Course Evaluation
Can a single global question replace a multi-item battery in course evaluation? Gogol et al. (2014) and the single-item-measure literature show when one item is defensible and when it is not.
Where Should the Concern Threshold Sit? Standard-Setting Methods (Angoff, Bookmark) for Course-Evaluation Triggers
Most institutions pick a review trigger - below 3.5, act - out of thin air. Standard-setting, the discipline that decides pass marks on high-stakes exams, offers a defensible, panel-based way to set a criterion-referenced course-evaluation threshold that will survive an appeal.
Which Variables Should You Actually Control For? Causal Diagrams and the Backdoor Criterion for Course Evaluation
Adjusting for more variables is not automatically more rigorous. Causal diagrams and Pearl's backdoor criterion tell you which covariates remove bias in a course-evaluation comparison and which ones create it.
Did the Course Make Students Both Confident and Convinced It Was Worth It? Expectancy-Value Theory as an Evaluation Lens
Most course evaluations measure satisfaction. Eccles and Wigfield's expectancy-value theory says the outcomes that actually predict effort, persistence, and course choice are whether students expect to succeed and whether they value the task. Here is how to evaluate both — without confusing them with satisfaction.
The Fluency Illusion: Why a Polished Lecturer Earns Higher Course Evaluations Without Teaching More
A research-grounded look at the lecture fluency illusion — why a smooth, confident presentation inflates student ratings and perceived learning while leaving actual learning unchanged — and what it means for interpreting course evaluations.
Grounded Theory for Open-Text Course Feedback: Building Explanation, Not Just Themes
How grounded theory's constant comparison, theoretical sampling, and saturation turn open-text course comments into an explanatory account of why students respond as they do — and how it differs from thematic analysis.
The Warmth Halo: How Instructor Immediacy Shapes Course Evaluations More Than Learning Does
What the meta-analytic evidence on teacher nonverbal immediacy and warmth tells us about course evaluations — why warmth strongly predicts how much students like a course and think they learned, but only weakly predicts what they actually learn.
How Do We Know It Is Bias? Natural-Experiment Evidence on Gender in Course Evaluations
Random-assignment and quasi-experimental studies — Maastricht, Sciences Po, and controlled online courses — provide causal evidence that gender bias in course evaluations is real and not explained by differences in teaching effectiveness or student learning.
Do Students Completing Evaluations on Their Phones Give Worse Data? The Device-Effects Evidence
Most students now answer course evaluations on a smartphone. Does the device degrade the data? The evidence says ratings stay stable across devices, but participation and the length of open-text comments differ — with clear design implications.
Did Your Adjustment Actually Work? Negative Controls for Detecting Hidden Bias in Course-Evaluation Comparisons
After adjusting for the confounders you could measure, a negative control tests for the ones you missed: a variable that should show no effect if your analysis is unbiased. If it lights up, your comparison is contaminated.
Beyond the Lecturer: What the Course Experience Questionnaire (CEQ) Measures and Why It Predicts Learning
Ramsden's Course Experience Questionnaire reframed evaluation around the learning environment, not the instructor's personality. We unpack what the CEQ measures, the evidence that its scales predict deep learning and outcomes, its limitations, and how Koji operationalises learning-environment evaluation.
Are Students Customers? The SERVQUAL Gap Model, HEdPERF, and What Service-Quality Thinking Adds to Course Evaluation
The SERVQUAL gap model and its higher-education variant HEdPERF measure the distance between what students expect and what they perceive they received. We assess the evidence, the sharp limitations of treating students as customers, and how Koji uses expectation framing without collapsing learning into satisfaction.
Which Combinations of Course Features Produce High Ratings? Qualitative Comparative Analysis for Programme Evaluation
Qualitative Comparative Analysis (QCA) finds which combinations of course conditions are consistently sufficient for a good outcome, capturing equifinality and conjunctural causation that net-effect regression misses — but it is sensitive to calibration choices, limited case diversity and contradictions, and is descriptive of your cases, not a causal proof.
Validity Is About the Use, Not the Instrument: Applying Kane's Argument-Based Framework to Course Evaluation
Asking whether course evaluations are valid is the wrong question. Kane's argument-based framework asks whether a specific interpretation and use of the scores is justified. We rebuild the SET debate as an interpretation-use argument, expose where each inference breaks, and show how Koji strengthens the weak links.
Should Every Point on a Course-Evaluation Scale Be Labelled? The Evidence on Verbal Anchors
Labelling every response option, not just the endpoints, tends to raise reliability and tame response styles — but it also interacts with the number of categories and the language of your students. What the rating-scale-format research says for course evaluation.
Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap
A randomised Harvard experiment found students learned more in active classrooms but rated their own learning lower. What the feeling-of-learning gap means for interpreting course-evaluation items that ask how much students learned.
Did Students Learn Something They Can Use Elsewhere? Transfer of Learning as a Course-Evaluation Lens
The ultimate test of a course is not whether students liked it or passed the exam, but whether the learning survives outside the classroom. Barnett and Ceci's taxonomy of transfer — and the sobering evidence that far transfer is hard — gives course evaluation a demanding, honest outcome to aim at.
Cognitive Apprenticeship and the Zone of Proximal Development as a Course-Evaluation Lens
Cognitive apprenticeship makes expert thinking visible through modelling, coaching and scaffolding. Here is what the theory implies for course-evaluation items — and why "was the lecturer clear?" misses most of what good teaching does.
It''s Not Just Quality — It''s Met Expectations: Expectancy-Disconfirmation and Course Evaluations
Satisfaction is not the same as quality. Oliver''s expectancy-disconfirmation model explains why an identical course earns different evaluation scores depending on what students expected, and what that means for interpreting and managing course feedback.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.
What Achievement Goal Theory Reveals About What Course Evaluations Should Measure
Achievement goal theory distinguishes mastery from performance orientations. Here is what the evidence says about using that lens to design and interpret course evaluations — and why "how motivated were you?" is the wrong question.
Total Survey Error: The Framework That Connects Every Course-Evaluation Quality Decision
The Total Survey Error framework organises the whole zoo of course-evaluation biases — coverage, sampling, nonresponse, measurement, processing — into one map, and tells you where to spend your limited effort for the biggest gain in accuracy.
How Long Should a Course Evaluation Be? Questionnaire Length, Breakoff, and Answer Quality
What the survey-methodology evidence says about questionnaire length: longer instruments depress participation and degrade answers to later questions, but ruthless shortening is not automatically the answer. A research-grounded guide for designing course evaluations.
The Dr. Fox Effect: Does Lecturer Charisma Inflate Course Evaluations?
The 1973 Dr. Fox experiment, its meta-analytic re-interpretation, and the 2014 re-revisitation — what the evidence really says about whether expressive delivery seduces students into rating empty teaching highly, and how to evaluate substance over showmanship.
Network Psychometrics: Treating Course-Evaluation Items as a Network, Not a Hidden Factor
Network psychometrics models course-evaluation items as a system of mutually reinforcing responses rather than reflections of a single hidden factor. This guide explains partial-correlation networks, centrality and bridge items, the evidence base, the honest caveats, and how Koji applies the idea.
Class Size and Student Evaluations: What Bedard and Kuhn Found
Does class size bias student evaluations of teaching? Bedard and Kuhn (2008) found a large, non-linear negative effect of enrolment on instructor ratings even after controlling for instructor and course. This article synthesises the evidence and explains how to stop class size from contaminating cross-instructor comparisons.
Likability and Prior Subject Interest: The Hidden Confounds in Student Evaluations
Two things a lecturer cannot control — how likeable students find them and how interested students already were in the subject — quietly move evaluation scores. The evidence on which is a real bias, and which is not.
When Students Use ChatGPT to Write Their Course Feedback: AI-Generated Open-Text and Data Integrity
Generative AI lets anyone produce fluent, coherent open-text survey answers, breaking the old assumption that a coherent response is a human response. What the evidence says about AI-written feedback, why detection is failing, and how to protect the integrity of qualitative course evaluation.
Can You Weight Your Way Out of a Low Response Rate? Post-Stratification and Nonresponse Adjustment in Course Evaluations
A high response rate does not guarantee unbiased course evaluations, and a low one is not automatically wrong. Here is what survey methodology says about post-stratification weighting, when it corrects nonresponse bias, and when it just adds noise.
Learn What Students Actually Weigh: The Factorial Survey (Vignette) Experiment for Course Evaluation
A factorial survey embeds a randomised experiment inside a survey: students rate realistic course vignettes whose features are varied independently, so the weight of each feature on their judgement can be recovered cleanly. This article explains the design, its distinction from conjoint and anchoring vignettes, and how to avoid its pitfalls.
The Hawthorne Effect and Observation Reactivity in Course Evaluation
Does being observed change teaching and student behaviour enough to bias evaluation evidence? What the research actually shows about the Hawthorne effect and reactivity, its contested size, and how to design evaluation that does not depend on the one week someone is watching.
Peer Observation vs Student Evaluations: What Each Actually Measures (and Why You Need Both)
Students and faculty observers see different things in the same classroom. The evidence on convergent validity shows why neither source alone can carry a high-stakes judgement of teaching.
Some of Your Course-Evaluation Responses Are Careless: Detecting Insufficient-Effort Responding
A minority of course-evaluation responses are produced without genuine attention. Here is what the careless-responding literature shows and how to screen for it before computing means.
Should a Course-Evaluation Scale Have a Neutral Midpoint? What "Neither Agree Nor Disagree" Really Captures
A neutral midpoint looks harmless, but the evidence shows it often functions as a hidden "don't know." Here is what the research says about including or omitting the middle option in course evaluations.
Translating a Course Evaluation Is Not Translation: Back-Translation, TRAPD, and Cross-Language Equivalence
Running the same course evaluation in several languages requires more than a good translator. Brislin's back-translation, the TRAPD model, and the ITC Guidelines explain how to keep items equivalent across languages.
Response-Shift Bias: Why Self-Reported Learning Gains in Course Evaluations Can Mislead
When you ask students how much they improved, the course itself has changed the yardstick they use to answer. Response-shift bias, and the retrospective pre-test that corrects it, explained for evaluation committees.
Latent Transition Analysis: Modelling How Student Segments Move Between Evaluation Waves
Latent transition analysis (LTA) extends latent profile analysis over time, estimating how students move between response segments from a mid-semester to an end-of-semester evaluation. This guide explains the method, the evidence, its limitations, and how Koji uses multi-wave data to support it.
Demand Characteristics: When Students Guess What Your Course Evaluation Wants to Hear
Orne''s concept of demand characteristics explains why students who infer what an evaluation is "for" answer to fit it. What the reactivity evidence says, and how to design evaluations that measure experience rather than compliance.
When Students Round to the Nearest Five: Digit Preference and Heaping in Numeric Course-Evaluation Answers
Typed numbers in evaluations pile up on multiples of five and ten. Whipple's and Myers' indices measure this heaping; it biases means and tails and flags data quality.
Your 200 Responses Are Worth 120: Design Effects, Effective Sample Size, and the Finite-Population Correction
Clustering and weighting inflate variance so the effective sample size is below the headcount; a near-census of a small class earns a finite-population correction. Kish's design effect makes precision honest.
Ask Fewer Questions, Measure More Precisely: Computerized Adaptive Testing for Course Evaluations
Computerized adaptive testing (CAT) uses item response theory to choose each next question based on a respondent's previous answers, reaching a target precision with far fewer items — but it needs a large, well-calibrated, unidimensional item bank, which most course-evaluation instruments do not yet have.
Engagement Surveys vs Course Evaluations: What NSSE Measures and Why Porter Says Be Careful
How student-engagement surveys such as NSSE differ from course evaluations, what Kuh (2009) argues they capture, why Porter (2011) questions their validity, and how to triangulate engagement and evaluation evidence without over-trusting either.
How Many Open-Text Comments Are Enough? Thematic Saturation in Course Evaluations
The qualitative-research evidence on data saturation — Guest et al. (2006) and Hennink et al. (2017) — applied to course-evaluation free text: how many comments you need before new themes stop appearing, and why richness per comment matters more than raw count.
Mid-Semester Feedback and the Power of Consultation: What the Meta-Analyses Show
Penny & Coe (2004) and Cohen (1980) on the formative effectiveness of mid-semester student feedback when combined with consultation, and the implications for European course-evaluation design.
Is RateMyProfessors a Valid Measure of Teaching? What the Evidence Says
A research-grounded assessment of whether public sites like RateMyProfessors.com measure teaching quality, anchored to Timmerman (2008), Felton, Mitchell and Stinson (2004), Rosen (2018) and Bleske-Rechek and Fritsch (2011) - and what it implies for institutional course evaluation.
Unipolar or Bipolar? The Course-Evaluation Scale Choice That Decides How Many Points You Need
A research-grounded guide to unipolar versus bipolar response scales in course evaluation: what each format measures, why bipolar scales skew positive, and why the optimal number of categories depends on which one you choose.
Should Course Evaluations Offer a "Don't Know" or "Not Applicable" Option?
The survey-methodology evidence on no-opinion and "not applicable" response options in course evaluations, anchored to Krosnick et al. (2002) and Krosnick (1991), with practical guidance on when an N/A option helps and when it invites satisficing.
Does Grading Leniency Inflate Student Evaluations? What the Evidence Actually Shows
A research-grounded look at the grading-leniency hypothesis in student evaluations of teaching: what Marsh and Roche, Greenwald and Gillmore, and Stroebe found, why the grades-ratings correlation is genuinely ambiguous, and how to design evaluation so leniency cannot quietly buy better scores.
Do Online Courses Get Lower Evaluations? Course Modality as a Confound
Students often rate the same instructor teaching the same content lower when the course is delivered online. This guide separates course modality (how the course is taught) from administration mode (how the survey is run), reviews the evidence, and explains why modality must be controlled before comparing scores.
Should You Run Course Evaluations Before or After the Final Exam? The Timing Evidence
A research-grounded look at whether course evaluations should be administered before or after the final examination, anchored to Vehovar and Strlekar (2024), Arnold (2009), and Overall and Marsh (1980), with practical guidance for European quality assurance.
Do Student Ratings Track Real Learning? What Cohen's 1981 Multisection Meta-Analysis Found — and Why the Answer Changed
Cohen (1981) found a moderate positive correlation between student ratings and student achievement in multisection courses — the foundational evidence that ratings have validity. Forty years of reanalysis has since shrunk that correlation toward zero. Here is what the multisection paradigm actually proves, and how to use ratings responsibly given both findings.
The Teachers Who Help You Most Get Rated Worst: What the Bocconi Natural Experiment Proved
Braga, Paccagnella and Pellizzari (2014) used near-random assignment of students to professors at Bocconi University to show that the teachers who most improved students' performance in later courses received the lowest student evaluations. Here is the design, the effect sizes, the caveats, and what it means for European quality assurance.
Should You Send More Students to the Better Question? Multi-Armed Bandits and Adaptive Experiments
Adaptive experiments reallocate respondents toward the better-performing survey variant while a study runs. Bandit algorithms help engagement but distort inference - here is the trade-off.
Do Student Evaluations Measure Teaching or the Instructor's Personality?
Research consistently finds that the Big Five personality traits of an instructor explain substantial variance in student evaluation of teaching (SET) scores, over and above grades and perceived learning. What that means for fair, valid course evaluation.
Where Did Students Get Stuck? Threshold Concepts as a Course-Evaluation Lens
Meyer and Land's threshold concepts reframe course evaluation from satisfaction to whether students crossed the transformative, troublesome ideas a discipline turns on — and got unstuck from liminality.
Beyond the Likert Scale: Adaptive Comparative Judgement in Evaluation
Humans are better at judging "which of these two is better" than at assigning an absolute number. Adaptive Comparative Judgement (Pollitt 2012), built on Thurstone's law of comparative judgement, turns pairwise choices into a reliable scale — with lessons for course evaluation.
It's Not What They Say, It's How: Discourse Analysis of Open-Text Course Feedback
Discourse analysis reads student comments as language doing work — positioning the writer and drawing on cultural repertoires — exposing framing and bias that theme counts and sentiment scores miss.
Evaluating Online and Blended Teaching: The Community of Inquiry Framework
A satisfaction mean tells you little about why an online course works. The Community of Inquiry framework (Garrison, Anderson & Archer) and its validated survey give course evaluation a research-grounded structure for teaching, social and cognitive presence.
Are Student Ratings Just a Mood? The Longitudinal Stability Evidence
A common objection to course evaluations is that students cannot judge teaching until years later. Overall and Marsh (1980) tested this directly by re-surveying the same students after at least a year. We review the stability evidence, its limits, and what it means for how feedback should be timed and used.
Can Forced-Choice Items Beat Response Bias in Course Evaluations? The Thurstonian IRT Evidence
Forced-choice (ipsative) formats were designed to suppress acquiescence, halo, and social-desirability response styles that contaminate ordinary Likert course evaluations. Brown and Maydeu-Olivares'' Thurstonian IRT model solves the classic ipsative-data problem — but the format is costly to build and not a free lunch. Here is the evidence and what it means for a university QA office.
Do Your Evaluation Items Actually Form a Scale? Mokken Analysis and Nonparametric Item Response Theory
Before you average a set of course-evaluation items into a subscale score, you should check that they form a scale at all. Mokken scale analysis tests that with far weaker assumptions than Rasch or factor analysis — and tells you whether the items even order students the same way.
Selection Bias in Course Evaluations: What Goos and Salomons Found
A research-grounded reading of Goos & Salomons (2017) on selection bias in online course evaluations, with practical implications for response-rate policy and reporting at European universities.
Why "I Learned a Lot" Can't Be Compared Across Courses: Reference Bias
Self-rated learning items are the backbone of most course evaluations, yet reference bias means students judge themselves against different implicit standards. We review the evidence that this distorts cross-group comparisons and what it means for benchmarking courses and programmes.
Do Student Evaluations Encourage Grade Inflation? The Incentive Problem
Wolfgang Stroebe argues that using student evaluations for high-stakes personnel decisions creates an incentive structure that rewards lenient grading and easy courses. We review the theory, the empirical evidence, the honest caveats, and what a defensible evaluation system should do instead.
Acquiescence Bias and Reverse-Worded Items: Should Course Evaluations Flip the Question?
Reverse-worded items are the classic survey-design fix for yea-saying. But Weijters and Baumgartner show negated and reversed items create their own measurement problems. What the evidence means for course-evaluation questionnaire design.
Teacher Clarity Predicts Learning Better Than Charisma: What Course Evaluations Should Measure
A meta-analysis of 144 effects and 73,000+ students shows teacher clarity explains roughly 13% of the variance in student learning. Here is what that means for the items you put on a course evaluation.
Do Student Ratings, Peer Observation and Learning Outcomes Agree? The Multitrait-Multimethod Test
If student surveys, peer observation, and learning outcomes all claim to measure teaching quality, do they converge — and do a course evaluation's sub-scores stay distinct or collapse into one halo? Campbell & Fiske's multitrait-multimethod matrix is the classic test, and it explains why triangulation, not any single method, is the honest standard.
Accent and Origin Bias in Course Evaluations: What Rubin (1992) Revealed
Donald Rubin's classic experiment showed students rated an identical recorded lecture as harder to understand when they believed the instructor was Asian — evidence that perceived accent and origin bias course evaluations. What it means for QA in multilingual European universities, and how to design evaluation that resists it.
Shorter Surveys Without Losing Coverage: Planned Missingness for Course Evaluations
Planned missing data designs let you cover more questions while each student answers fewer. Here is how the three-form design works, what the evidence says, and where it fits in course evaluation.
Illuminative Evaluation: What Parlett and Hamilton's 'Learning Milieu' Adds to Course Feedback
Parlett and Hamilton's 1972 illuminative evaluation reframes course evaluation as illuminating how a course works within its learning milieu, not just scoring objectives. Here is the evidence, its limits, and how it changes QA practice.
Realist Evaluation for Course Feedback: 'What Works, for Whom, in What Circumstances'
Pawson and Tilley's realist evaluation replaces 'did it work?' with 'what works, for whom, in what circumstances?' using context-mechanism-outcome configurations. Here is the theory, the evidence, the caveats, and how to apply it to course evaluation.
Educational Connoisseurship and Criticism: Eisner's Case for Expert Judgement in Teaching Evaluation
Elliot Eisner argued that evaluating teaching is more like art criticism than measurement: it needs a connoisseur's trained perception and a critic's public disclosure. Here is the model, its evidence, its limits, and how it complements student ratings.
Do First Impressions Decide Your Course Evaluation? The Thin-Slice Evidence
Ambady and Rosenthal showed that silent 30-second clips of an instructor predict end-of-term ratings. What thin-slice judgments mean for the validity of course evaluations and how to design feedback that probes substance, not first impressions.
Grid or One Question at a Time? Matrix Formats and Straightlining in Course Evaluations
Rendering rating items as a single grid instead of one question per screen quietly degrades course-evaluation data. The controlled evidence on item non-response, breakoff and straightlining — and what to do about it.
Did Crossing the Threshold Change the Rating? Regression Discontinuity for Course-Evaluation Cutoff Effects
When a rule assigns students, instructors, or courses on either side of a sharp cutoff, regression discontinuity turns that arbitrary boundary into near-experimental evidence about what actually moved an evaluation score.
Who Dropped Out of Your Longitudinal Evaluation — and Did It Break the Comparison? Differential Attrition and Panel Mortality
In any pre/post or multi-wave evaluation, the students who stop responding are rarely a random subset — and when they leave one condition faster than another, differential attrition can manufacture an effect that was never there.
Question Order and Context Effects: How the Sequence of Items Shapes Course-Evaluation Answers
The order in which you ask evaluation questions changes the answers you get. Drawing on Schwarz (1999), Strack, Martin & Schwarz (1988) and Tourangeau, Rips & Rasinski (2000), we explain part-whole and assimilation/contrast effects, what they do to your data, and how conversational evaluation reduces the damage.
Do Anti-Bias Statements Reduce Gender Bias in Student Evaluations of Teaching?
A randomised experiment (Peterson et al., 2019) found that a short anti-bias statement placed on the evaluation form significantly raised ratings of women instructors without changing ratings of men. We review the evidence, the caveats, and how to deploy warning-label language responsibly in course evaluation.
Responsive Evaluation (Stake): Letting Stakeholder Issues Drive Course Evaluation
Robert Stake''s responsive evaluation begins not with pre-set objectives but with the issues and concerns that stakeholders actually raise, using emergent, largely qualitative design. We explain the model, its evidence base, its limits, and how it maps onto conversational course evaluation.
Do Interactive Follow-Up Probes Improve Open-Text Course Feedback Quality?
Survey-methodology experiments (Holland & Christian, 2009; Smyth et al., 2009) show that an interactive follow-up probe after a student''s first open-text answer increases response length and the number of distinct themes without raising item non-response. We review the evidence and its implications for conversational, AI-moderated course evaluation.