New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs

Research methods

Course evaluation methodology, bias reduction, and quality assurance frameworks.

Racial and Ethnic Bias in Student Evaluations of Teaching: What the Evidence Shows

Peer-reviewed evidence shows students rate instructors of colour systematically lower than White peers on the same teaching. We review Reid (2010), Chávez & Mitchell (2020) and Bavishi et al. (2010), the limitations, and how to design evaluation so race-linked bias does not contaminate quality decisions.

10 min readadvanced

Does the Room Bias the Rating? Physical Classroom Environment as a Confound in Course Evaluations

Evidence that the physical classroom - lighting, seating, comfort, technology - shifts student evaluation of teaching scores independent of instructional quality, and how to stop the estate from contaminating the teaching signal.

9 min readadvanced

Is the Weak SET-Learning Correlation an Artefact of Unreliable Measures? What Correction for Attenuation Does — and Doesn't — Prove

Disattenuation lets you estimate what a correlation would be if both measures were perfectly reliable. It is a legitimate tool that both sides of the student-ratings debate have used — and abused — to argue the true SET-learning link is stronger or weaker than the raw number suggests.

11 min readadvanced

Do Early-Morning Classes Get Lower Evaluations? Time-of-Day and Scheduling as Confounds

Evidence that the timetable slot - early morning, late afternoon, days per week - shifts student performance and mood, and therefore course evaluation scores, independent of teaching quality. What QA teams should record and adjust for.

9 min readadvanced

Do Students Who Skip Class Rate Teaching Differently? Attendance as a Confound in Course Evaluations

Attendance and the perceived need to attend are entangled with course-evaluation scores - and with who fills the survey out. Grounded in Burns and Ludlow (2005), this explains the confound and how to keep it from distorting SET.

10 min readadvanced

Evaluating Doctoral Supervision: What PRES and Supervision-Experience Measures Actually Capture

A research-grounded look at how the postgraduate research (PGR) supervision experience is measured, what the UK PRES and validated instruments like the QSDI capture, and where those measures fall short.

10 min readadvanced

Evaluating Team-Taught Courses: The Attribution Problem in Student Evaluations

Why a single overall rating conflates co-teachers, and how to design evaluations that separate instructor-level from course-level signal in team-taught courses.

10 min readadvanced

Why Was the Course Rated Well, Not Just Whether? Mediation Analysis and the Mechanism Behind an Evaluation Score

A high rating tells you students were satisfied; it doesn't tell you why. Mediation analysis tests the pathway — did clarity raise satisfaction by increasing engagement? — but drawing that causal chain from a single end-of-term survey is far harder than the classic recipe suggests.

11 min readadvanced

Evaluating Clinical Placements and Work-Based Learning: Why the End-of-Course Questionnaire Is the Wrong Instrument

Why standard end-of-course student evaluation questionnaires fail to capture placement quality, and what validated clinical/practice-learning-environment instruments (MCPI, CLES+T, UCEEM, PET, CLEI, DREEM) measure instead.

11 min readadvanced

Do Adjuncts Get Worse Course Evaluations Than Tenured Faculty? Employment Status as a Confound

The causal evidence says contingent faculty do not teach worse — and often teach better. Why employment status is a poor proxy for teaching quality, and how to read evaluation scores fairly across the tenure divide.

9 min readadvanced

Do Older Professors Get Lower Course Evaluations? Age and Seniority as a Confound

Stonebraker and Stone found a small but robust negative effect of instructor age on student ratings, emerging after the mid-forties. What the evidence does and does not show, and how to stop age from contaminating comparisons.

8 min readintermediate

Does a Better Researcher Make a Better Teacher? The Teaching-Research Nexus and Your Evaluations

Hattie and Marsh found the correlation between research productivity and teaching quality is essentially zero. What that means for reading course evaluations and for keeping teaching and research evidence separate.

9 min readadvanced

The Critical Incident Technique for Course Feedback

How Flanagan's Critical Incident Technique collects concrete, behaviourally-anchored student feedback that global Likert ratings cannot capture, and how it applies to course evaluation.

9 min readadvanced

Q-Methodology: Surfacing the Distinct Viewpoints Students Hold About a Course

Q-methodology uses forced-choice card sorts and by-person factor analysis to reveal the two-to-four genuinely distinct viewpoints students hold about a course, a rigorous complement to Likert SET averages.

14 minadvanced

Beauty Bias in Course Evaluations: Are Attractive Instructors Rated as Better Teachers?

What the research says about physical-attractiveness bias in student evaluations of teaching — from Hamermesh & Parker (2005) to German and laboratory replications — and how to design evaluation so appearance does not masquerade as teaching quality.

9 min readadvanced

Why Students Click Straight Down the Middle: Satisficing in Course Evaluations

A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.

9 min readintermediate

Why "The Pace Was About Right" Breaks Your Scale: Ideal-Point Unfolding Models for Course Evaluation

Most rating models assume more of a trait always means more agreement. For "appropriateness" items — workload, pace, difficulty — that assumption is false, and it quietly wrecks the scale. Ideal-point unfolding models fix it.

10 min readadvanced

Contribution Analysis: Making Credible Causal Claims from Course Evaluation

You changed a course and outcomes improved — but did the teaching cause it? Contribution analysis is a theory-based method for making defensible causal claims when a randomised trial is impossible. Here is how it applies to course evaluation.

10 min readadvanced

The Most Significant Change Technique: Story-Based Course Evaluation That Surfaces What Students Actually Value

Likert averages tell you where a course sits; they cannot tell you what transformed a student. The Most Significant Change technique collects and collectively selects stories of change to reveal unanticipated outcomes and shared values. Here is how it works and where it fits course evaluation.

10 min readintermediate

SALG: Can Students Reliably Report Their Own Learning Gains?

The SALG instrument asks about learning gains, not satisfaction. What Seymour and colleagues built, what the validity evidence (and Porter''s and Bowman''s critiques) show about self-reported gains, and how to use gains-framed questions responsibly.

9 min readadvanced

The One-Minute Paper: Formative Feedback That Also Improves Learning

The one-minute paper and the muddiest point are the best-known Classroom Assessment Techniques. What Chizmar & Ostrosky (1998) and Stead (2005) found about their effect on learning and feedback, and how to run continuous formative collection well.

8 min readintermediate

The Success Case Method: Evaluating a Course Through Its Best and Worst Cases

Brinkerhoff''s Success Case Method evaluates a course by studying its most and least successful students, not its average. What the approach is, its evidence and biases, and how to run the two-stage screen-then-interview design.

9 min readadvanced

What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback

Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.

10 min readadvanced

Course Evaluations Measure Reaction, Not Learning: What Kirkpatrick's Four Levels Reveal

Kirkpatrick's four-level model explains why an end-of-term course evaluation is a Level-1 'reaction' measure — and why decades of meta-analytic evidence show reaction correlates almost nothing with actual learning.

9 min readintermediate

Fixing Nonresponse While You Field, Not After: Responsive and Adaptive Survey Design for Course Evaluation

Responsive and adaptive survey design monitors who is answering while a course evaluation is still open and reallocates effort toward under-represented students to reduce nonresponse bias — a step earlier than post-hoc weighting.

9 min readadvanced

Beyond the End-of-Term Survey: Stufflebeam's CIPP Model for Programme-Level Evaluation

Most course evaluation stops at student satisfaction. Stufflebeam's CIPP model (Context, Input, Process, Product) reframes evaluation as decision-support across the whole programme lifecycle — and maps cleanly onto ESG/ENQA quality-cycle thinking.

9 min readintermediate

When, Not Just Whether, Students Leave: Cox Proportional-Hazards Models for Course Evaluation

Most evaluation analytics ask whether a student withdrew. Cox proportional-hazards regression asks when — modelling the risk of withdrawal over time as a function of course, cohort, and early-experience signals — and why that is a sharper quality question than a pass/fail dropout flag.

10 min readadvanced

Do Your Evaluation Items Actually Cover Teaching? Content Validity, the CVR and the CVI

An evaluation form can be reliable and still measure the wrong things. Content validity, quantified with Lawshe's CVR and the Content Validity Index, tests whether your items cover the teaching domain.

10 min readintermediate

Not All Course Attributes Are Equal: The Kano Model and the Asymmetry of Student Satisfaction

A course-evaluation mean assumes every attribute affects satisfaction the same way. The Kano model shows it does not: some attributes only cause dissatisfaction when they are missing, others only delight when present. We explain must-be, one-dimensional and attractive quality — and why averaging hides them.

10 min readadvanced

Can You Train the Bias Out of Evaluation Reviewers? Frame-of-Reference Rater Training

The committees and peer observers who read evaluations carry halo, contrast and attribution biases. Frame-of-reference training is the evidence-based method for making their judgments more accurate.

10 min readadvanced

Does Evaluating Course After Course Change How Students Answer? Panel Conditioning in Repeated Course Evaluations

Students at a European university complete dozens of course evaluations across a degree. Panel-conditioning research shows that the mere act of being surveyed repeatedly can change later answers — a threat to comparing scores across years and cohorts.

9 min readadvanced

The Nominal Group Technique: Structured Group Feedback That Ranks What Students Actually Care About

A survey tells you how students rated a course; it rarely tells you what they would fix first. The Nominal Group Technique — a structured, facilitated method — generates and prioritises student feedback with the whole cohort in the room, and has a track record in higher-education course evaluation.

9 min readintermediate

Stop Arguing About Wording — Test It: Split-Ballot Experiments for Course-Evaluation Questions

Committees spend hours debating whether to word an evaluation item one way or another. Schuman and Presser showed that small wording changes can move survey answers by double-digit margins — and that the way to settle the debate is a randomised split-ballot experiment, not opinion.

9 min readadvanced

Is Your Course Pushing Students Toward Deep or Surface Learning? The R-SPQ-2F as an Evaluation Lens

Most course evaluations ask whether students liked the teaching. The deep/surface approaches tradition asks a more consequential question: did the course lead students to engage meaningfully or just memorise to pass? The R-SPQ-2F instrument makes that measurable.

10 min readadvanced

Numbers First, Then the Why: Explanatory Sequential Mixed Methods for Course Evaluation

A Likert average tells you a course scored 3.4; it never tells you why. The explanatory sequential mixed methods design (Creswell & Plano Clark) fixes this by using a quantitative survey to decide exactly which students to follow up with qualitatively — turning an unexplained number into an evidenced account.

9 min readadvanced

Your Course Evaluation Measures Satisfaction, Not Emotion. Pekrun Says That's a Problem

Standard course evaluations ask whether students were satisfied. Pekrun's control-value theory and the Achievement Emotions Questionnaire show that discrete emotions — enjoyment, boredom, anxiety, hope, hopelessness — drive learning and are absent from almost every institutional survey. Here is why that gap matters and how to close it.

9 min readintermediate

Narrative Inquiry for Course Evaluation: Reading Student Stories as Wholes, Not Fragments

Thematic coding chops student feedback into fragments and loses the plot. Narrative inquiry keeps each student's course experience whole and temporal, surfacing turning points that a code frequency cannot.

9 min readintermediate

Should Course Evaluations Ask Whether Teaching Matched a Student's Learning Style?

Learning styles are one of the most durable myths in education, and Pashler et al. (2008) found no credible evidence for tailoring instruction to them. Here is why a course-evaluation item that asks students whether teaching 'suited their learning style' quietly measures a debunked construct — and what to ask instead.

9 min readintermediate

What Cognitive Load Theory Says Your Course Evaluation Should — and Shouldn't — Ask

Cognitive load theory (Sweller, van Merriënboer & Paas, 2019) distinguishes the unavoidable difficulty of content from difficulty caused by poor design. That distinction changes what a course evaluation should measure: not overall 'difficulty', but the design choices that impose or remove extraneous load.

10 min readadvanced

Evaluating Active Learning: The ICAP Framework as a Course-Evaluation Lens

Most course evaluations ask whether a course was 'engaging' — a word that conflates enjoyment with learning. Chi and Wylie's (2014) ICAP framework replaces it with an observable ladder of cognitive engagement (Passive, Active, Constructive, Interactive), giving evaluation items that measure what students actually did.

10 min readadvanced

Desirable Difficulties: Why the Teaching That Improves Learning Often Lowers Satisfaction

Spacing, interleaving, and retrieval practice are among the best-evidenced ways to make learning durable — and they make a course feel harder and less smooth in the moment. Bjork & Bjork's (2011) desirable-difficulties principle explains why end-of-term satisfaction ratings systematically penalise the most effective teaching.

11 min readadvanced

Does Course Difficulty and Workload Lower Student Evaluations? What Centra Found

A research-grounded look at whether harder, heavier courses are punished in student evaluations of teaching. Centra (2003) and Marsh & Roche (2000) show the relationship is non-linear and weaker than faculty fear — with practical implications for fair course evaluation.

9 min readintermediate

Is Your Course Actually Aligned? Using Biggs's Constructive Alignment as an Evaluation Lens

Biggs's constructive alignment says learning outcomes, teaching activities, and assessment must point the same way. Here is how to turn that theory into a course-evaluation instrument that diagnoses misalignment students can feel but rarely name.

10 min readadvanced

Does Your Course Support Autonomy, Competence, and Relatedness? Self-Determination Theory as an Evaluation Lens

Self-Determination Theory says motivation depends on three basic needs — autonomy, competence, and relatedness. Here is how to evaluate a course by whether it feeds or starves those needs, instead of only whether students were satisfied.

10 min readadvanced

Stop Waiting for the End of Term: Experience Sampling for In-the-Moment Course Feedback

Experience-sampling methods capture what students feel and think during a course, not their reconstructed memory of it months later. Here is why in-the-moment data can be more valid than the end-of-term survey — and how to use it responsibly.

11 min readadvanced

Why Quantitative Courses Get Lower Evaluations: The Discipline-Bias Problem

Uttl & Smibert (2017) show that instructors of quantitative courses receive systematically lower student evaluations than those teaching qualitative subjects — a bias with real career consequences. What the evidence says and how to compare ratings fairly across disciplines.

9 min readadvanced

Focus Groups for Course Evaluation: What They Reveal That Surveys Miss (and Where They Fail)

What the methodological literature — Stalmeijer et al.'s AMEE Guide No. 91, Kaplowitz & Hoehn's comparative study, and randomized method comparisons — says about using focus groups for course evaluation, their known failure modes, and how AI-moderated individual interviews capture the depth without the group-dynamics distortions.

9 min readintermediate

Do University Teachers Get Better With Experience? A 13-Year Growth-Model Study Says: Not Automatically

Herbert Marsh's multilevel growth model of student evaluations over 13 years found teachers' ratings neither improve nor decline with experience — individual differences dominate. What this means for tenure assumptions, development policy, and how evaluation data should track trajectories.

9 min readadvanced

The Delphi Method for Course and Programme Evaluation: Building Expert Consensus Without the Loudest Voice Winning

How the Delphi technique produces structured, anonymous expert consensus on course-evaluation criteria and programme standards — its evidence base, its limitations, and where it fits alongside student feedback.

9 min readadvanced

Beyond the Student Survey: Structured Classroom-Observation Protocols (COPUS and TDOP) as Evaluation Evidence

What COPUS and the TDOP measure that student surveys and traditional peer visits cannot: low-inference, reliable records of what actually happens in a classroom. Their evidence base, their limits, and where they fit in a triangulated evaluation.

10 min readintermediate

Which Item Is Biased? Differential Item Functioning Detection for Course-Evaluation Questions

Measurement invariance tests the whole scale; differential item functioning (DIF) pinpoints the single biased item. A practical guide to Mantel-Haenszel and logistic-regression DIF, uniform vs non-uniform bias, and what to do when a course-evaluation item behaves differently across groups.

11 min readadvanced

Do the Clicks Confirm the Comments? Triangulating Course Evaluations with LMS Learning-Analytics Data

Learning-analytics trace data from your VLE looks like an objective check on what students say in course evaluations. The research says it is a weaker and more course-specific signal than most institutions assume — here is how to triangulate it honestly.

9 min readadvanced

How Big a Difference Can You Actually Detect? A-Priori Power Analysis for Course Evaluation

Reliability tells you how stable a score is; statistical power tells you whether your sample can detect a real difference at all. A practical guide to a-priori power analysis, minimum detectable effects, and why most single-class course-evaluation comparisons are underpowered before they begin.

11 min readadvanced

Should Your Course Evaluation Ask How Students Used AI? What the Learning Evidence Says

Students now study with generative AI in nearly every course, and it changes what a satisfaction rating means. The evidence on AI and learning explains why — and what a course evaluation should and should not try to ask about it.

10 min readintermediate

Asking the Questions Students Won't Answer Honestly: The Randomized Response Technique for Sensitive Course-Evaluation Items

When a course or climate survey asks about harassment, discrimination, academic misconduct, or truancy, ordinary anonymity is not enough — students still under-report. The randomized response technique adds provable privacy so respondents can answer honestly. What Warner (1965) proposed, what the validation evidence shows, and where it fails.

11 min readadvanced

Not Every Cohort Follows the Same Path: Latent Growth Curve and Growth Mixture Models for Course Evaluation

Latent growth curve and growth mixture models describe how student trajectories change across evaluation waves — and uncover subgroups that a single average path hides — while guarding against the subgroups that are only statistical artefacts.

9 min readadvanced

Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers

Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.

10 min readadvanced

Will Students Open Up to an AI That Runs Their Course Evaluation? Algorithm Aversion, Appreciation, and Disclosure

Do students trust an AI to conduct their course evaluation, and does an AI moderator make them more or less candid? The evidence on algorithm aversion, appreciation, and disclosure-to-machines is more encouraging than the sceptics assume.

10 min readadvanced

Does Your Course Build Self-Regulated Learners? Metacognition and SRL as an Evaluation Lens

Self-regulated learning — the cycle of planning, monitoring and reflecting — predicts academic achievement, yet standard course evaluations never ask whether a course developed it. Here is the SRL evidence and how to turn it into evaluation questions.

9 min readintermediate

Should Course Evaluations Measure Whether Students Feel They Belong? The Sense-of-Belonging Evidence

Sense of belonging predicts persistence, engagement, and mental health in higher education — often more reliably than satisfaction. Here is what the evidence says about measuring belonging in course evaluation, and how to do it without turning a survey into a diagnostic instrument.

9 min readintermediate

Should Course Evaluations Ask Whether a Course Fostered a Growth Mindset? What the Meta-Analytic Evidence Says

Growth-mindset interventions show weak average effects in two large meta-analyses. Here is what that means for whether — and how — course evaluations should ask about mindset.

9 min readintermediate

Did the Course Build Students Who Can Use Feedback? Feedback Literacy as an Evaluation Lens

Carless and Boud's feedback-literacy framework reframes a course-quality question: not "was feedback given?" but "did the course develop students who can appreciate, judge, manage, and act on feedback?"

10 min readintermediate

Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation

A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.

12 min readadvanced

One Score, Two Things: Item Response Tree Models Separate Opinion from Response Style

Item response tree (IRTree) models split each Likert answer into latent decisions — respond or stay neutral, agree or disagree, moderate or extreme — so a student's real opinion of a course can be estimated apart from their personal scale-use style.

9 min readadvanced

The Halo Effect in Course Evaluations: When One Impression Colours Every Rating

When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.

9 min readintermediate

Stop Asking Students to Rate Everything Highly: Discrete Choice Experiments Reveal What They Will Trade Off

Discrete choice experiments (DCEs) ask students to choose between course scenarios rather than rate each feature, forcing real trade-offs and yielding a random-utility estimate of what actually drives their preferences.

9 min readadvanced

Bifactor and ESEM Models: When a Multidimensional Course Evaluation Still Justifies One Overall Score

Bifactor and exploratory structural equation models let you estimate a general teaching-quality factor and specific factors at once, then test with omega-hierarchical and ECV whether reporting a single overall course-evaluation score is defensible.

10 min readadvanced

Did the Course Spark Interest? The Four-Phase Model of Interest Development as an Evaluation Lens

Most evaluations ask whether students enjoyed a course. Hidi and Renninger's four-phase model distinguishes a momentary spark from durable individual interest — a sharper, more consequential thing to measure.

10 min readintermediate

When Everything Scores 4/5: Best-Worst Scaling (MaxDiff) for Course-Evaluation Priorities

Likert ratings on course evaluations cluster near the top and cannot tell you what matters most to students. Best-worst scaling (Louviere, Flynn & Marley) forces trade-offs that reveal genuine priorities. What the method is, its limits, and how it fits a Koji study.

9 min readadvanced

Let Students Name Their Own Criteria: The Repertory Grid Technique for Course Evaluation

Standard evaluation forms impose the institution's categories on students. The repertory grid technique, built on Kelly's personal construct theory, elicits the dimensions students themselves use to judge a course — surfacing criteria a fixed questionnaire never asked about.

10 min readadvanced

Before You Field It, Test It: Cognitive Interviewing for Course-Evaluation Questions

Why the wording of a course-evaluation item should be cognitively pretested before it reaches students, what Beatty and Willis (2007) established about think-aloud and verbal probing, and how Koji operationalises probing at scale.

9 min readadvanced

Do Cookies, Treats, and Mood Bias Course Evaluations?

Two controlled studies show that giving students chocolate or cookies before an evaluation measurably raises teaching scores. What the affect heuristic means for the validity of course evaluations — and how to design around it.

8 min readintermediate

Do Instructors and Students Agree? Self-Evaluation vs Student Ratings

Feldman's synthesis found instructor self-ratings and student ratings correlate only moderately (around r ≈ 0.3). What weak self–student agreement means for triangulation, faculty trust, and how to use both sources without privileging either.

9 min readadvanced

Does a Conversational Course Evaluation Make Students Less Honest? Mode Effects and Social Desirability

Survey mode shapes honesty: interviewer-administered surveys invite more social-desirability bias than self-administered ones. What Tourangeau and Yan (2007) and the mode-effects literature mean for anonymous, AI-moderated conversational course evaluations.

10 min readadvanced

Measuring What Students Will Not Admit Directly: The List Experiment for Sensitive Course-Evaluation Questions

Some of the most important evaluation questions - Did you actually attend? Did you use AI on assessments? Did the grade you expected shape your rating? - are exactly the ones students answer dishonestly. The list experiment (item-count technique) estimates their prevalence without ever asking any student to admit anything.

11 min readadvanced

Do the Numbers on Your Rating Scale Change the Score? The Evidence on Numeric Anchors

Schwarz and colleagues showed that the numeric values printed on a rating scale (0 to 10 vs minus 5 to plus 5) systematically shift responses even when the verbal labels are identical. Here is what that means for course-evaluation design, comparability, and reporting.

9 min readadvanced

Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?

Online course evaluations consistently draw lower response rates than in-class paper forms, but the research shows the resulting scores are largely equivalent. Here is what Dommeyer, Nulty, and Stowell actually found, and what an adequate response rate really requires.

9 min readintermediate

Did the Course Produce Flow? Csikszentmihalyi's Flow Theory as a Course-Evaluation Lens

Csikszentmihalyi's flow theory reframes course evaluation around whether teaching created states of deep, balanced absorption. Here is what the evidence supports, where it breaks down, and how to measure it responsibly.

9 min readintermediate

Did Students Leave More Confident They Can Succeed? Academic Self-Efficacy as a Course-Evaluation Outcome

Academic self-efficacy is among the strongest psychological correlates of student achievement. Here is what the evidence says, why it belongs in course evaluation, its limits, and how to measure it without kidding yourself.

9 min readintermediate

The Last Week Counts Double: The Peak-End Rule and What Students Actually Remember About Your Course

Retrospective memory of an experience is dominated by its most intense moment and its ending, not its average or its length. Here is what the peak-end rule means for end-of-semester course evaluations.

10 min readintermediate

Stop Asking "Was the Lecturer Clear?": The Case for Low-Inference Teaching-Behaviour Items

High-inference items like "Is the instructor clear?" tell teachers nothing about what to change. Murray's low-inference behaviour research shows how concrete, observable items make evaluations diagnostic.

10 min readadvanced

How Often Is "Often"? Why Vague Quantifiers Quietly Distort Course-Evaluation Data

Words like "often", "usually" and "sometimes" mean different things to different students, which makes course-evaluation responses hard to compare. Here is the survey-methodology evidence and what to do about it.

9 min readintermediate

Sliders, Visual-Analogue, or Radio Buttons? The Evidence on Course-Evaluation Response Formats

Slider widgets look modern, but the survey-methodology evidence says they cost you data. Funke (2016), Couper et al. (2006) and Bosch et al. (2019) on choosing a response widget for online course evaluations.

8 min readintermediate

Does the Professor Students Rate Highest Teach Them Most? Value-Added Evidence vs Student Ratings

The strongest causal evidence — random-assignment studies that track follow-on course performance — shows the instructors students rate highest are often not the ones whose students learn most over time. What Carrell & West (2010) and Braga et al. (2014) found, and what it means for course evaluation.

10 min readadvanced

Which Evaluation Questions Are Redundant? Mutual Information and mRMR for Trimming Your Questionnaire

Two evaluation items are redundant to the degree that one predicts the other — mutual information. The mRMR rule keeps items that are maximally relevant to your decision and minimally redundant with each other, yielding a shorter, sharper questionnaire.

9 min readadvanced

Is a Low Response Rate Biasing Your Course Evaluations? MCAR, MAR, and Missing-Not-at-Random

A low response rate does not automatically bias a course-evaluation score — the missing-data mechanism (MCAR, MAR, or MNAR) does. How to diagnose nonresponse bias instead of chasing a rate threshold.

9 min readadvanced

Do Elective Courses Get Higher Ratings Than Required Ones? Course Characteristics and the Feldman Evidence

Course features the instructor cannot control — whether a course is elective or required, its level, size, subject, and timing — systematically shift student ratings. What Kenneth Feldman's synthesis established, and why comparing raw ratings across different course types is unfair.

9 min readintermediate

One Number or Many? The Dimensionality Debate and How to Use Student Ratings for Personnel Decisions

Should a promotion committee use a single global teaching score or a detailed profile of many dimensions? The 1997 d'Apollonia & Abrami vs Marsh & Roche debate set the terms — and the answer depends on whether the purpose is summative judgement or formative improvement.

10 min readadvanced

Deleting Incomplete Surveys Throws Away Good Data: Multiple Imputation for Missing Course-Evaluation Responses

Listwise deletion of partially-complete course evaluations shrinks your sample and can bias results. How multiple imputation (Rubin 1987; Schafer & Graham 2002) handles item nonresponse honestly.

9 min readadvanced

The Qualitatively Different Ways Students Experience Your Course: Phenomenography

A standard survey tells you how many students rated "feedback" a 3. Phenomenography, the higher-education research tradition founded by Ference Marton, instead maps the qualitatively different ways students experience a phenomenon — and explains why the same number means different things.

9 min readadvanced

Goal-Free Evaluation: Why Scriven Says Course Reviews Should Ignore the Stated Objectives

Placeholder

10 min readadvanced

Agree/Disagree or Item-Specific? The Scale Choice That Quietly Degrades Course-Evaluation Data

"The lecturer was well organised: Strongly disagree to Strongly agree" feels natural, but the agree/disagree format invites acquiescence and lower data quality. Saris et al. (2010) on why item-specific scales measure better.

9 min readadvanced

Response-Order Effects: Does Where an Answer Sits Change How Students Rate Your Course?

The order in which answer options appear can shift course-evaluation responses independent of what students think. We unpack Krosnick & Alwin (1987) on primacy and recency, why it matters for instrument design, and how to limit it.

9 min readadvanced

Is 4.2 Really Worse Than 4.4? Bayesian Estimation and Credible Intervals for Course Evaluations

Bayesian estimation answers the question institutions actually ask — how probable is it that this instructor is below standard? — with credible intervals and a region of practical equivalence instead of raw averages or p-values.

9 min readadvanced

Can Instructors Buy Better Ratings With Easy Grades? Instrumental Variables and the Grade–Evaluation Endogeneity Problem

Grades and ratings are jointly determined, so a naive regression overstates the grade effect. How Krautmann & Sander (1999) used two-stage least squares to recover the causal effect of expected grades on evaluations.

10 min readadvanced

What Can Open-Text Student Comments Tell You That Likert Scores Cannot?

A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.

9 min readintermediate

The Student-as-Consumer Effect: What a Consumer Mindset Does to Course Evaluations

When students see themselves as paying customers, what happens to how they rate courses — and how they learn? A look at Bunce, Baird & Jones (2017) and why the consumer frame quietly distorts the meaning of satisfaction scores.

9 min readintermediate

Is Student Evaluation of Teaching Valid? What the State-of-the-Art Review Concludes

A close reading of Spooren, Brockx and Mortelmans'' 2013 Review of Educational Research synthesis of SET validity — what it actually concludes, the evidence behind it, and what it means for how universities use student ratings.

10 min readadvanced

How Many Scale Points Should a Course-Evaluation Question Have?

What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.

9 min readintermediate

Is One "Overall" Question Enough? Single-Item vs Multi-Item Measures in Course Evaluation

Can a single global question replace a multi-item battery in course evaluation? Gogol et al. (2014) and the single-item-measure literature show when one item is defensible and when it is not.

9 min readintermediate

Where Should the Concern Threshold Sit? Standard-Setting Methods (Angoff, Bookmark) for Course-Evaluation Triggers

Most institutions pick a review trigger - below 3.5, act - out of thin air. Standard-setting, the discipline that decides pass marks on high-stakes exams, offers a defensible, panel-based way to set a criterion-referenced course-evaluation threshold that will survive an appeal.

11 min readadvanced

Which Variables Should You Actually Control For? Causal Diagrams and the Backdoor Criterion for Course Evaluation

Adjusting for more variables is not automatically more rigorous. Causal diagrams and Pearl's backdoor criterion tell you which covariates remove bias in a course-evaluation comparison and which ones create it.

10 min readadvanced

Did the Course Make Students Both Confident and Convinced It Was Worth It? Expectancy-Value Theory as an Evaluation Lens

Most course evaluations measure satisfaction. Eccles and Wigfield's expectancy-value theory says the outcomes that actually predict effort, persistence, and course choice are whether students expect to succeed and whether they value the task. Here is how to evaluate both — without confusing them with satisfaction.

9 min readintermediate

The Fluency Illusion: Why a Polished Lecturer Earns Higher Course Evaluations Without Teaching More

A research-grounded look at the lecture fluency illusion — why a smooth, confident presentation inflates student ratings and perceived learning while leaving actual learning unchanged — and what it means for interpreting course evaluations.

9 min readintermediate

Grounded Theory for Open-Text Course Feedback: Building Explanation, Not Just Themes

How grounded theory's constant comparison, theoretical sampling, and saturation turn open-text course comments into an explanatory account of why students respond as they do — and how it differs from thematic analysis.

9 min readadvanced

The Warmth Halo: How Instructor Immediacy Shapes Course Evaluations More Than Learning Does

What the meta-analytic evidence on teacher nonverbal immediacy and warmth tells us about course evaluations — why warmth strongly predicts how much students like a course and think they learned, but only weakly predicts what they actually learn.

9 min readintermediate

How Do We Know It Is Bias? Natural-Experiment Evidence on Gender in Course Evaluations

Random-assignment and quasi-experimental studies — Maastricht, Sciences Po, and controlled online courses — provide causal evidence that gender bias in course evaluations is real and not explained by differences in teaching effectiveness or student learning.

10 min readadvanced

Do Students Completing Evaluations on Their Phones Give Worse Data? The Device-Effects Evidence

Most students now answer course evaluations on a smartphone. Does the device degrade the data? The evidence says ratings stay stable across devices, but participation and the length of open-text comments differ — with clear design implications.

9 min readintermediate

Did Your Adjustment Actually Work? Negative Controls for Detecting Hidden Bias in Course-Evaluation Comparisons

After adjusting for the confounders you could measure, a negative control tests for the ones you missed: a variable that should show no effect if your analysis is unbiased. If it lights up, your comparison is contaminated.

9 min readadvanced

Beyond the Lecturer: What the Course Experience Questionnaire (CEQ) Measures and Why It Predicts Learning

Ramsden's Course Experience Questionnaire reframed evaluation around the learning environment, not the instructor's personality. We unpack what the CEQ measures, the evidence that its scales predict deep learning and outcomes, its limitations, and how Koji operationalises learning-environment evaluation.

10 min readadvanced

Are Students Customers? The SERVQUAL Gap Model, HEdPERF, and What Service-Quality Thinking Adds to Course Evaluation

The SERVQUAL gap model and its higher-education variant HEdPERF measure the distance between what students expect and what they perceive they received. We assess the evidence, the sharp limitations of treating students as customers, and how Koji uses expectation framing without collapsing learning into satisfaction.

11 min readadvanced

Which Combinations of Course Features Produce High Ratings? Qualitative Comparative Analysis for Programme Evaluation

Qualitative Comparative Analysis (QCA) finds which combinations of course conditions are consistently sufficient for a good outcome, capturing equifinality and conjunctural causation that net-effect regression misses — but it is sensitive to calibration choices, limited case diversity and contradictions, and is descriptive of your cases, not a causal proof.

10 min readadvanced

Validity Is About the Use, Not the Instrument: Applying Kane's Argument-Based Framework to Course Evaluation

Asking whether course evaluations are valid is the wrong question. Kane's argument-based framework asks whether a specific interpretation and use of the scores is justified. We rebuild the SET debate as an interpretation-use argument, expose where each inference breaks, and show how Koji strengthens the weak links.

12 min readadvanced

Should Every Point on a Course-Evaluation Scale Be Labelled? The Evidence on Verbal Anchors

Labelling every response option, not just the endpoints, tends to raise reliability and tame response styles — but it also interacts with the number of categories and the language of your students. What the rating-scale-format research says for course evaluation.

9 min readintermediate

Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap

A randomised Harvard experiment found students learned more in active classrooms but rated their own learning lower. What the feeling-of-learning gap means for interpreting course-evaluation items that ask how much students learned.

9 min readadvanced

Did Students Learn Something They Can Use Elsewhere? Transfer of Learning as a Course-Evaluation Lens

The ultimate test of a course is not whether students liked it or passed the exam, but whether the learning survives outside the classroom. Barnett and Ceci's taxonomy of transfer — and the sobering evidence that far transfer is hard — gives course evaluation a demanding, honest outcome to aim at.

9 min readadvanced

Cognitive Apprenticeship and the Zone of Proximal Development as a Course-Evaluation Lens

Cognitive apprenticeship makes expert thinking visible through modelling, coaching and scaffolding. Here is what the theory implies for course-evaluation items — and why "was the lecturer clear?" misses most of what good teaching does.

10 min readadvanced

It''s Not Just Quality — It''s Met Expectations: Expectancy-Disconfirmation and Course Evaluations

Satisfaction is not the same as quality. Oliver''s expectancy-disconfirmation model explains why an identical course earns different evaluation scores depending on what students expected, and what that means for interpreting and managing course feedback.

9 min readadvanced

Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited

A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.

12 min readadvanced

What Achievement Goal Theory Reveals About What Course Evaluations Should Measure

Achievement goal theory distinguishes mastery from performance orientations. Here is what the evidence says about using that lens to design and interpret course evaluations — and why "how motivated were you?" is the wrong question.

10 min readadvanced

Total Survey Error: The Framework That Connects Every Course-Evaluation Quality Decision

The Total Survey Error framework organises the whole zoo of course-evaluation biases — coverage, sampling, nonresponse, measurement, processing — into one map, and tells you where to spend your limited effort for the biggest gain in accuracy.

9 min readadvanced

How Long Should a Course Evaluation Be? Questionnaire Length, Breakoff, and Answer Quality

What the survey-methodology evidence says about questionnaire length: longer instruments depress participation and degrade answers to later questions, but ruthless shortening is not automatically the answer. A research-grounded guide for designing course evaluations.

9 min readintermediate

The Dr. Fox Effect: Does Lecturer Charisma Inflate Course Evaluations?

The 1973 Dr. Fox experiment, its meta-analytic re-interpretation, and the 2014 re-revisitation — what the evidence really says about whether expressive delivery seduces students into rating empty teaching highly, and how to evaluate substance over showmanship.

9 min readadvanced

Network Psychometrics: Treating Course-Evaluation Items as a Network, Not a Hidden Factor

Network psychometrics models course-evaluation items as a system of mutually reinforcing responses rather than reflections of a single hidden factor. This guide explains partial-correlation networks, centrality and bridge items, the evidence base, the honest caveats, and how Koji applies the idea.

9 min readadvanced

Class Size and Student Evaluations: What Bedard and Kuhn Found

Does class size bias student evaluations of teaching? Bedard and Kuhn (2008) found a large, non-linear negative effect of enrolment on instructor ratings even after controlling for instructor and course. This article synthesises the evidence and explains how to stop class size from contaminating cross-instructor comparisons.

9 min readadvanced

Likability and Prior Subject Interest: The Hidden Confounds in Student Evaluations

Two things a lecturer cannot control — how likeable students find them and how interested students already were in the subject — quietly move evaluation scores. The evidence on which is a real bias, and which is not.

9 min readintermediate

When Students Use ChatGPT to Write Their Course Feedback: AI-Generated Open-Text and Data Integrity

Generative AI lets anyone produce fluent, coherent open-text survey answers, breaking the old assumption that a coherent response is a human response. What the evidence says about AI-written feedback, why detection is failing, and how to protect the integrity of qualitative course evaluation.

11 min readadvanced

Can You Weight Your Way Out of a Low Response Rate? Post-Stratification and Nonresponse Adjustment in Course Evaluations

A high response rate does not guarantee unbiased course evaluations, and a low one is not automatically wrong. Here is what survey methodology says about post-stratification weighting, when it corrects nonresponse bias, and when it just adds noise.

10 min readadvanced

Learn What Students Actually Weigh: The Factorial Survey (Vignette) Experiment for Course Evaluation

A factorial survey embeds a randomised experiment inside a survey: students rate realistic course vignettes whose features are varied independently, so the weight of each feature on their judgement can be recovered cleanly. This article explains the design, its distinction from conjoint and anchoring vignettes, and how to avoid its pitfalls.

9 min readadvanced

The Hawthorne Effect and Observation Reactivity in Course Evaluation

Does being observed change teaching and student behaviour enough to bias evaluation evidence? What the research actually shows about the Hawthorne effect and reactivity, its contested size, and how to design evaluation that does not depend on the one week someone is watching.

9 min readintermediate

Peer Observation vs Student Evaluations: What Each Actually Measures (and Why You Need Both)

Students and faculty observers see different things in the same classroom. The evidence on convergent validity shows why neither source alone can carry a high-stakes judgement of teaching.

9 min readintermediate

Some of Your Course-Evaluation Responses Are Careless: Detecting Insufficient-Effort Responding

A minority of course-evaluation responses are produced without genuine attention. Here is what the careless-responding literature shows and how to screen for it before computing means.

9 min readadvanced

Should a Course-Evaluation Scale Have a Neutral Midpoint? What "Neither Agree Nor Disagree" Really Captures

A neutral midpoint looks harmless, but the evidence shows it often functions as a hidden "don't know." Here is what the research says about including or omitting the middle option in course evaluations.

9 min readintermediate

Translating a Course Evaluation Is Not Translation: Back-Translation, TRAPD, and Cross-Language Equivalence

Running the same course evaluation in several languages requires more than a good translator. Brislin's back-translation, the TRAPD model, and the ITC Guidelines explain how to keep items equivalent across languages.

10 min readadvanced

Response-Shift Bias: Why Self-Reported Learning Gains in Course Evaluations Can Mislead

When you ask students how much they improved, the course itself has changed the yardstick they use to answer. Response-shift bias, and the retrospective pre-test that corrects it, explained for evaluation committees.

9 min readadvanced

Latent Transition Analysis: Modelling How Student Segments Move Between Evaluation Waves

Latent transition analysis (LTA) extends latent profile analysis over time, estimating how students move between response segments from a mid-semester to an end-of-semester evaluation. This guide explains the method, the evidence, its limitations, and how Koji uses multi-wave data to support it.

9 min readadvanced

Demand Characteristics: When Students Guess What Your Course Evaluation Wants to Hear

Orne''s concept of demand characteristics explains why students who infer what an evaluation is "for" answer to fit it. What the reactivity evidence says, and how to design evaluations that measure experience rather than compliance.

9 min readadvanced

When Students Round to the Nearest Five: Digit Preference and Heaping in Numeric Course-Evaluation Answers

Typed numbers in evaluations pile up on multiples of five and ten. Whipple's and Myers' indices measure this heaping; it biases means and tails and flags data quality.

9 min readintermediate

Your 200 Responses Are Worth 120: Design Effects, Effective Sample Size, and the Finite-Population Correction

Clustering and weighting inflate variance so the effective sample size is below the headcount; a near-census of a small class earns a finite-population correction. Kish's design effect makes precision honest.

9 min readadvanced

Ask Fewer Questions, Measure More Precisely: Computerized Adaptive Testing for Course Evaluations

Computerized adaptive testing (CAT) uses item response theory to choose each next question based on a respondent's previous answers, reaching a target precision with far fewer items — but it needs a large, well-calibrated, unidimensional item bank, which most course-evaluation instruments do not yet have.

9 min readadvanced

Engagement Surveys vs Course Evaluations: What NSSE Measures and Why Porter Says Be Careful

How student-engagement surveys such as NSSE differ from course evaluations, what Kuh (2009) argues they capture, why Porter (2011) questions their validity, and how to triangulate engagement and evaluation evidence without over-trusting either.

9 min readadvanced

How Many Open-Text Comments Are Enough? Thematic Saturation in Course Evaluations

The qualitative-research evidence on data saturation — Guest et al. (2006) and Hennink et al. (2017) — applied to course-evaluation free text: how many comments you need before new themes stop appearing, and why richness per comment matters more than raw count.

8 min readintermediate

Mid-Semester Feedback and the Power of Consultation: What the Meta-Analyses Show

Penny & Coe (2004) and Cohen (1980) on the formative effectiveness of mid-semester student feedback when combined with consultation, and the implications for European course-evaluation design.

12 min readadvanced

Is RateMyProfessors a Valid Measure of Teaching? What the Evidence Says

A research-grounded assessment of whether public sites like RateMyProfessors.com measure teaching quality, anchored to Timmerman (2008), Felton, Mitchell and Stinson (2004), Rosen (2018) and Bleske-Rechek and Fritsch (2011) - and what it implies for institutional course evaluation.

10 min readintermediate

Unipolar or Bipolar? The Course-Evaluation Scale Choice That Decides How Many Points You Need

A research-grounded guide to unipolar versus bipolar response scales in course evaluation: what each format measures, why bipolar scales skew positive, and why the optimal number of categories depends on which one you choose.

9 min readadvanced

Should Course Evaluations Offer a "Don't Know" or "Not Applicable" Option?

The survey-methodology evidence on no-opinion and "not applicable" response options in course evaluations, anchored to Krosnick et al. (2002) and Krosnick (1991), with practical guidance on when an N/A option helps and when it invites satisficing.

9 min readintermediate

Does Grading Leniency Inflate Student Evaluations? What the Evidence Actually Shows

A research-grounded look at the grading-leniency hypothesis in student evaluations of teaching: what Marsh and Roche, Greenwald and Gillmore, and Stroebe found, why the grades-ratings correlation is genuinely ambiguous, and how to design evaluation so leniency cannot quietly buy better scores.

9 min readadvanced

Do Online Courses Get Lower Evaluations? Course Modality as a Confound

Students often rate the same instructor teaching the same content lower when the course is delivered online. This guide separates course modality (how the course is taught) from administration mode (how the survey is run), reviews the evidence, and explains why modality must be controlled before comparing scores.

9 min readintermediate

Should You Run Course Evaluations Before or After the Final Exam? The Timing Evidence

A research-grounded look at whether course evaluations should be administered before or after the final examination, anchored to Vehovar and Strlekar (2024), Arnold (2009), and Overall and Marsh (1980), with practical guidance for European quality assurance.

9 min readintermediate

Do Student Ratings Track Real Learning? What Cohen's 1981 Multisection Meta-Analysis Found — and Why the Answer Changed

Cohen (1981) found a moderate positive correlation between student ratings and student achievement in multisection courses — the foundational evidence that ratings have validity. Forty years of reanalysis has since shrunk that correlation toward zero. Here is what the multisection paradigm actually proves, and how to use ratings responsibly given both findings.

10 min readadvanced

The Teachers Who Help You Most Get Rated Worst: What the Bocconi Natural Experiment Proved

Braga, Paccagnella and Pellizzari (2014) used near-random assignment of students to professors at Bocconi University to show that the teachers who most improved students' performance in later courses received the lowest student evaluations. Here is the design, the effect sizes, the caveats, and what it means for European quality assurance.

10 min readadvanced

Should You Send More Students to the Better Question? Multi-Armed Bandits and Adaptive Experiments

Adaptive experiments reallocate respondents toward the better-performing survey variant while a study runs. Bandit algorithms help engagement but distort inference - here is the trade-off.

9 min readadvanced

Do Student Evaluations Measure Teaching or the Instructor's Personality?

Research consistently finds that the Big Five personality traits of an instructor explain substantial variance in student evaluation of teaching (SET) scores, over and above grades and perceived learning. What that means for fair, valid course evaluation.

9 min readadvanced

Where Did Students Get Stuck? Threshold Concepts as a Course-Evaluation Lens

Meyer and Land's threshold concepts reframe course evaluation from satisfaction to whether students crossed the transformative, troublesome ideas a discipline turns on — and got unstuck from liminality.

9 min readintermediate

Beyond the Likert Scale: Adaptive Comparative Judgement in Evaluation

Humans are better at judging "which of these two is better" than at assigning an absolute number. Adaptive Comparative Judgement (Pollitt 2012), built on Thurstone's law of comparative judgement, turns pairwise choices into a reliable scale — with lessons for course evaluation.

10 min readadvanced

It's Not What They Say, It's How: Discourse Analysis of Open-Text Course Feedback

Discourse analysis reads student comments as language doing work — positioning the writer and drawing on cultural repertoires — exposing framing and bias that theme counts and sentiment scores miss.

9 min readadvanced

Evaluating Online and Blended Teaching: The Community of Inquiry Framework

A satisfaction mean tells you little about why an online course works. The Community of Inquiry framework (Garrison, Anderson & Archer) and its validated survey give course evaluation a research-grounded structure for teaching, social and cognitive presence.

9 min readintermediate

Are Student Ratings Just a Mood? The Longitudinal Stability Evidence

A common objection to course evaluations is that students cannot judge teaching until years later. Overall and Marsh (1980) tested this directly by re-surveying the same students after at least a year. We review the stability evidence, its limits, and what it means for how feedback should be timed and used.

8 min readadvanced

Can Forced-Choice Items Beat Response Bias in Course Evaluations? The Thurstonian IRT Evidence

Forced-choice (ipsative) formats were designed to suppress acquiescence, halo, and social-desirability response styles that contaminate ordinary Likert course evaluations. Brown and Maydeu-Olivares'' Thurstonian IRT model solves the classic ipsative-data problem — but the format is costly to build and not a free lunch. Here is the evidence and what it means for a university QA office.

11 min readadvanced

Do Your Evaluation Items Actually Form a Scale? Mokken Analysis and Nonparametric Item Response Theory

Before you average a set of course-evaluation items into a subscale score, you should check that they form a scale at all. Mokken scale analysis tests that with far weaker assumptions than Rasch or factor analysis — and tells you whether the items even order students the same way.

10 min readadvanced

Selection Bias in Course Evaluations: What Goos and Salomons Found

A research-grounded reading of Goos & Salomons (2017) on selection bias in online course evaluations, with practical implications for response-rate policy and reporting at European universities.

11 min readadvanced

Why "I Learned a Lot" Can't Be Compared Across Courses: Reference Bias

Self-rated learning items are the backbone of most course evaluations, yet reference bias means students judge themselves against different implicit standards. We review the evidence that this distorts cross-group comparisons and what it means for benchmarking courses and programmes.

9 min readadvanced

Do Student Evaluations Encourage Grade Inflation? The Incentive Problem

Wolfgang Stroebe argues that using student evaluations for high-stakes personnel decisions creates an incentive structure that rewards lenient grading and easy courses. We review the theory, the empirical evidence, the honest caveats, and what a defensible evaluation system should do instead.

9 min readadvanced

Acquiescence Bias and Reverse-Worded Items: Should Course Evaluations Flip the Question?

Reverse-worded items are the classic survey-design fix for yea-saying. But Weijters and Baumgartner show negated and reversed items create their own measurement problems. What the evidence means for course-evaluation questionnaire design.

9 min readadvanced

Teacher Clarity Predicts Learning Better Than Charisma: What Course Evaluations Should Measure

A meta-analysis of 144 effects and 73,000+ students shows teacher clarity explains roughly 13% of the variance in student learning. Here is what that means for the items you put on a course evaluation.

9 min readintermediate

Do Student Ratings, Peer Observation and Learning Outcomes Agree? The Multitrait-Multimethod Test

If student surveys, peer observation, and learning outcomes all claim to measure teaching quality, do they converge — and do a course evaluation's sub-scores stay distinct or collapse into one halo? Campbell & Fiske's multitrait-multimethod matrix is the classic test, and it explains why triangulation, not any single method, is the honest standard.

12 min readadvanced

Accent and Origin Bias in Course Evaluations: What Rubin (1992) Revealed

Donald Rubin's classic experiment showed students rated an identical recorded lecture as harder to understand when they believed the instructor was Asian — evidence that perceived accent and origin bias course evaluations. What it means for QA in multilingual European universities, and how to design evaluation that resists it.

9 min readadvanced

Shorter Surveys Without Losing Coverage: Planned Missingness for Course Evaluations

Planned missing data designs let you cover more questions while each student answers fewer. Here is how the three-form design works, what the evidence says, and where it fits in course evaluation.

10 min readadvanced

Illuminative Evaluation: What Parlett and Hamilton's 'Learning Milieu' Adds to Course Feedback

Parlett and Hamilton's 1972 illuminative evaluation reframes course evaluation as illuminating how a course works within its learning milieu, not just scoring objectives. Here is the evidence, its limits, and how it changes QA practice.

10 min readadvanced

Realist Evaluation for Course Feedback: 'What Works, for Whom, in What Circumstances'

Pawson and Tilley's realist evaluation replaces 'did it work?' with 'what works, for whom, in what circumstances?' using context-mechanism-outcome configurations. Here is the theory, the evidence, the caveats, and how to apply it to course evaluation.

11 min readadvanced

Educational Connoisseurship and Criticism: Eisner's Case for Expert Judgement in Teaching Evaluation

Elliot Eisner argued that evaluating teaching is more like art criticism than measurement: it needs a connoisseur's trained perception and a critic's public disclosure. Here is the model, its evidence, its limits, and how it complements student ratings.

10 min readadvanced

Do First Impressions Decide Your Course Evaluation? The Thin-Slice Evidence

Ambady and Rosenthal showed that silent 30-second clips of an instructor predict end-of-term ratings. What thin-slice judgments mean for the validity of course evaluations and how to design feedback that probes substance, not first impressions.

9 min readadvanced

Grid or One Question at a Time? Matrix Formats and Straightlining in Course Evaluations

Rendering rating items as a single grid instead of one question per screen quietly degrades course-evaluation data. The controlled evidence on item non-response, breakoff and straightlining — and what to do about it.

9 min readintermediate

Did Crossing the Threshold Change the Rating? Regression Discontinuity for Course-Evaluation Cutoff Effects

When a rule assigns students, instructors, or courses on either side of a sharp cutoff, regression discontinuity turns that arbitrary boundary into near-experimental evidence about what actually moved an evaluation score.

11 min readadvanced

Who Dropped Out of Your Longitudinal Evaluation — and Did It Break the Comparison? Differential Attrition and Panel Mortality

In any pre/post or multi-wave evaluation, the students who stop responding are rarely a random subset — and when they leave one condition faster than another, differential attrition can manufacture an effect that was never there.

10 min readadvanced

Question Order and Context Effects: How the Sequence of Items Shapes Course-Evaluation Answers

The order in which you ask evaluation questions changes the answers you get. Drawing on Schwarz (1999), Strack, Martin & Schwarz (1988) and Tourangeau, Rips & Rasinski (2000), we explain part-whole and assimilation/contrast effects, what they do to your data, and how conversational evaluation reduces the damage.

10 min readadvanced

Do Anti-Bias Statements Reduce Gender Bias in Student Evaluations of Teaching?

A randomised experiment (Peterson et al., 2019) found that a short anti-bias statement placed on the evaluation form significantly raised ratings of women instructors without changing ratings of men. We review the evidence, the caveats, and how to deploy warning-label language responsibly in course evaluation.

9 min readintermediate

Responsive Evaluation (Stake): Letting Stakeholder Issues Drive Course Evaluation

Robert Stake''s responsive evaluation begins not with pre-set objectives but with the issues and concerns that stakeholders actually raise, using emergent, largely qualitative design. We explain the model, its evidence base, its limits, and how it maps onto conversational course evaluation.

10 min readadvanced

Do Interactive Follow-Up Probes Improve Open-Text Course Feedback Quality?

Survey-methodology experiments (Holland & Christian, 2009; Smyth et al., 2009) show that an interactive follow-up probe after a student''s first open-text answer increases response length and the number of distinct themes without raising item non-response. We review the evidence and its implications for conversational, AI-moderated course evaluation.

9 min readintermediate