Analysis & reporting
Quality scoring, thematic analysis, triangulation, and report generation.
Does Test Anxiety Distort Course Evaluations — and Should Evaluations Measure It?
A 30-year meta-analysis confirms test anxiety reliably depresses performance. That makes it both a hidden confound in course ratings and, arguably, a course-quality signal worth capturing.
Comparing Instructors on Many Criteria at Once: Data Envelopment Analysis for Course Evaluation
Data Envelopment Analysis (DEA) rates instructors, modules, or departments against the best observed performers on several inputs and outputs at once — a fair multi-criteria alternative to ranking on one mean.
Measuring Consensus, Not Just the Average: Mixed-Effects Location-Scale Models for Course Evaluation
Mixed-effects location-scale models model the spread of course-evaluation ratings — consensus versus polarisation — as an outcome in its own right, not just noise around the mean.
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You
Natural-language processing can turn thousands of free-text course-evaluation comments into themes and sentiment at scale — but the research (Cunningham-Nelson 2019; Sunar & Khalid 2023) shows where automated analysis is reliable and where human judgement is still required.
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
Is the Workload-Rating Relationship a Straight Line? Splines and Generalized Additive Models for Course Evaluation
When you regress a rating on workload, difficulty, or class size, a straight line assumes each extra unit matters the same everywhere. Restricted cubic splines and generalized additive models let the data reveal the real curve — often a U-shape or a plateau — without slicing a continuous predictor into arbitrary bins.
When Ratings Pile Up at 5: Tobit and Censored Regression for Ceiling-Bounded Course Evaluation
When a third of your class marks the top of the scale, the 5 is a floor on their true opinion, not its ceiling. Ordinary regression on such data underestimates real differences. The Tobit model treats the pile-up at the boundary as censoring and recovers the effect that a naive analysis flattens.
Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Pooling Your Own Sections Without Faking Precision: Random-Effects Meta-Analysis for Course Evaluation
When you combine ratings across many small sections or terms, averaging the averages pretends they all measure one fixed truth. Random-effects meta-analysis treats each section as a noisy estimate of a genuinely varying effect, weights them properly, and reports how much real spread remains — with a prediction interval, not just a mean.
How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis
When you turn thousands of free-text comments into themes and counts, how do you know the coding is trustworthy? O Connor and Joffe (2020) on intercoder reliability, Braun and Clarke on thematic analysis, and what rigorous qualitative QA looks like.
Why a Single Student Survey Can't Stand Alone: Common-Method Bias in Course Evaluation
Common-method bias (Podsakoff et al., 2003) explains why correlations inside a single end-of-term student survey are inflated by the shared method itself - and why triangulating teaching evidence matters. A research-grounded guide for quality assurance.
Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?
Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.
Cronbach's Alpha and Course-Evaluation Reliability: What It Does and Doesn't Tell You
A high Cronbach's alpha on your course-evaluation instrument is widely read as proof of a "reliable" survey. The psychometric literature says that reading is largely wrong. What alpha actually measures, where it misleads, and what to report instead.
How Much of a Course-Evaluation Change Actually Matters? The Minimal Important Difference
Statistical significance tells you a course-evaluation change is real; the Minimal Important Difference (MID) tells you whether it is big enough to matter. Here is how to set one, using anchor-based and distribution-based methods imported from health measurement.
Can You Compare Course Ratings Across Cultures? Anchoring Vignettes and the King Method
When students from different countries interpret the same rating scale differently, their scores are not comparable. King, Murray, Salomon and Tandon (2004) introduced anchoring vignettes to correct this. Here is how the technique works and what it means for international and multi-campus course evaluation.
Are Your Respondents Representative? Using Early-vs-Late Wave Analysis to Estimate Nonresponse Bias
A low response rate is not automatically biased — what matters is whether respondents differ from non-respondents. Armstrong and Overton (1977) gave us a cheap diagnostic: compare early and late responders. Here is how to use wave analysis on course-evaluation data and where it breaks down.
When the AI Summary Says Something No Student Did: Faithfulness and Hallucination in LLM Course-Feedback Summaries
A large language model that summarises hundreds of open-text comments can invent a theme, a sentiment, or even a quote that no student wrote. The abstractive-summarisation research explains why - and what a defensible AI feedback pipeline must do to stay faithful to the source.
Should You Use Net Promoter Score for Courses? The "Would You Recommend" Question in Higher Education
Net Promoter Score is migrating from customer experience into student feedback. What Reichheld (2003) actually claimed, why Keiningham et al. (2007) failed to replicate its superiority, and whether a single recommend-question belongs in course evaluation.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Can You Report a Class Mean? ICC(1), ICC(2), and r_wg for Aggregating Student Ratings
Before you average student ratings into a class score, three organisational-psychology indices decide whether you can: r_wg (within-class agreement), ICC(1) (how much variance is between classes), and ICC(2) (the reliability of the class mean).
You Changed the Questionnaire — Can You Still Compare Years? Test Equating and Linking for Course Evaluation
When you revise a course-evaluation form, scores before and after are not automatically comparable. Kolen and Brennan's equating/linking/prediction hierarchy, and the anchor-item designs behind it, tell you what a cross-revision trend can honestly claim.
Beyond Keyword Counts: Semantic Text Embeddings for Searching, Deduplicating, and Mapping Open-Text Feedback
Text embeddings turn each open-text comment into a point in a meaning-space where semantically similar feedback sits close together — enabling semantic search, near-duplicate detection, coverage measurement, and routing that keyword counts and topic models cannot deliver.
An Evaluation of Course Evaluations: What the Berkeley Statisticians Concluded About Reporting Student Ratings
Two Berkeley statisticians argued that averaging student-rating scores to rank instructors for promotion and tenure should be abandoned for both statistical and substantive reasons. What Stark & Freishtat (2014) actually recommend — and how to report ratings responsibly instead.
Gendered Language in Student Comments: Why Men Are ''Brilliant'' and Women Are ''Caring''
Bias in course evaluations is not only in the numbers — it is in the words. What Mitchell & Martin (2018) and Storage et al. (2016) found about systematically different language applied to men and women, and why open-text analysis must account for it.
Do Students'' Written Comments Match Their Ratings? What Concordance Tells You
Open-text comments and Likert scores usually agree — but the gaps are where the insight lives. What Alhija & Fresko (2009) and Brockx et al. (2012) found about the consistency between qualitative and quantitative course-evaluation data, and how to read it.
A Mismeasured Predictor Biases Its Own Slope: Errors-in-Variables and Regression Dilution in Course Evaluation
Self-reported workload, engagement and interest are all measured with error, so their regression slopes are attenuated toward zero. Why more responses do not fix it, and how reliability lets you correct it.
Even Fair Evaluations Can Rank Unfairly: Why Course-Evaluation Scores Misclassify Instructors
A simulation-based finding that even unbiased, reliable, and valid course-evaluation scores produce high misclassification rates when used to rank or compare instructors — and what it means for tenure, promotion, and merit decisions.
Some Raters Are Harsh, Some Items Are Hard: Why Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores
Averaging Likert scores treats every student as an equally calibrated measuring instrument and every item as equally hard. Rasch and Many-Facet Rasch models do not. We explain the method, the evidence that rater severity and item difficulty distort raw means, the limitations, and how Koji applies the same logic.
Should You Average a 1–5 Course-Evaluation Scale at All? The Ordinal-vs-Interval Debate
Is it legitimate to compute a mean from a five-point course-evaluation scale? The fifty-year ordinal-versus-interval debate, from Jamieson to Carifio and Perla to Norman, and what it means for reporting student ratings.
When You Cannot Assume the Non-Responders Are Like the Responders: Manski Bounds for Course Evaluation
Weighting and imputation assume the silent majority resemble the responders. Partial identification refuses that assumption and reports the interval the true mean could occupy, showing exactly how much your conclusion rests on belief.
Why One Cruel Comment Outweighs Twenty Kind Ones: Negativity Bias in Reading Course Evaluations
Instructors and committees fixate on the harshest open-text comment and discount the praise. Baumeister''s "bad is stronger than good" and the negativity-bias literature explain why, and how to read qualitative course feedback fairly.
Correcting for Who Chose to Respond: The Heckman Selection Model for Course-Evaluation Nonresponse
When students who respond to course evaluations differ from those who skip them on the very thing you are measuring, reweighting cannot help. The Heckman selection model models response and rating jointly to correct for selection on unobservables — under fragile assumptions this article makes explicit.
Control for Everything Stable About an Instructor: Fixed-Effects Panel Models for Course Evaluation
Fixed-effects panel models compare each instructor only to themselves across terms, silently controlling for every stable confounder — measured or not. This article explains the within-transformation, the FE-vs-random-effects choice and Hausman test, and what FE can and cannot tell a quality office.
Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations
Year-over-year movements in course-evaluation scores are mostly statistical noise regressing toward an instructor's true average. Here is why regression to the mean fools quality-assurance committees, and how to read evaluation trends responsibly.
When Almost Everyone Scores 4.5: Ceiling Effects, Skew, and What They Do to Course-Evaluation Statistics
Course-evaluation ratings pile up at the top of the scale, producing a strong ceiling effect and negative skew that breaks the statistics most universities still report. Here is what the evidence shows and how to report ratings honestly.
How Much of the Rating Gap Is Bias? The Oaxaca-Blinder Decomposition for Course Evaluation
When two groups of instructors get different average ratings, the Oaxaca-Blinder decomposition splits the gap into a part explained by measurable circumstances and an unexplained residual. This article shows how to use it responsibly — and why the unexplained part is a bound on bias, not a measurement of it.
Is Your Missing Course-Evaluation Data Random? MAR, MNAR, and What to Do About It
A low response rate is a missing-data problem. Rubin's MCAR/MAR/MNAR framework explains when a course-evaluation mean is biased and when multiple imputation or maximum likelihood can help.
You Flagged the Lowest Instructor, Then Tested If They Were Below Average: Selective Inference and the Winner's Curse
Ranking instructors and then testing the extreme one invalidates the p-value and biases the flagged score. Selective inference, FCR intervals, and empirical-Bayes shrinkage put the inference right.
Your Ratings Have a Term Rhythm: STL Seasonal-Trend Decomposition for Rolling Course Evaluation
Continuous feedback carries a semester rhythm on top of a trend on top of noise. STL decomposes the series so you read real change without being fooled by the calendar.
Can a Large Language Model Code Your Open-Text Course Feedback? What the Agreement Studies Show
Peer-reviewed evidence on how closely GPT-4-class models match human coders when categorising open-ended student comments — agreement statistics, where they fail, and how to use them responsibly in course-evaluation quality assurance.
Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores
Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.
Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages for Course Evaluation
Course-evaluation data has a nested structure — students within sections within instructors — and a flat average ignores it. This guide explains how multilevel (hierarchical linear) models partition variance, why ignoring clustering understates uncertainty, and what it means for fair reporting.
Why Is the Effect Bigger in Some Sections? Meta-Regression for Course Evaluation
Meta-regression explains between-section heterogeneity in pooled course-evaluation results by modelling section-level moderators — powerful for generating explanations, weak for confirming them.
Beyond "Positive or Negative": Aspect-Based Sentiment Analysis of Open-Text Course Feedback
Document-level sentiment scoring collapses a rich student comment into one number and hides the most useful signal. Aspect-based sentiment analysis (ABSA) separates *what* the comment is about from *how* the student feels about it. Here is the evidence on ABSA for course feedback, its accuracy, its failure modes, and how Koji applies it.
Should You Report an Instructor''s Percentile? Norm-Referenced vs Criterion-Referenced Course-Evaluation Scores
Telling a lecturer they are "in the 40th percentile of the department" is norm-referenced reporting — and it manufactures losers by construction, no matter how good everyone is. Criterion-referenced reporting asks instead whether teaching met a defined standard. Here is the evidence on why the choice matters and how to report responsibly.
Topic Modeling for Open-Text Course Evaluations: LDA, STM and BERTopic Explained
How LDA, structural topic models and BERTopic surface themes across tens of thousands of open-text student comments — what each method does well, where it fails, and how to keep a human in the loop.
Top-Box vs Mean: How to Report Course Evaluation Scores Without Throwing Away Information
Reporting the percentage of students who chose the top box feels intuitive, but collapsing a scale to favorable/unfavorable discards information. Here is what the measurement evidence says and how to report responsibly.
Did the Teaching Change Actually Work? Interrupted Time Series for Course-Evaluation Trends
Comparing this year''s evaluation mean to last year''s cannot tell you whether a curriculum redesign worked. Interrupted time series with segmented regression can — here is how to apply it, and where it breaks.
The Framework Method for Open-Text Course Feedback: A Structured Alternative to Thematic Coding
When a quality committee — not a lone qualitative researcher — has to make sense of thousands of student comments, the Framework Method offers a transparent, auditable matrix-based approach. What Gale et al. (2013) actually proposed, and how to use it.
Statistical Process Control for Course Evaluations: Using Control Charts to Separate Signal from Noise
How Shewhart control charts and statistical process control distinguish real changes in teaching quality from ordinary random variation in course evaluation scores - so QA teams stop reacting to noise. Grounded in Sivena and Nikolaidis (2019, 2022).
Why "Assessment and Feedback" Consistently Scores Lowest in Student Surveys
A research-grounded explanation of why assessment and feedback is the perennial weak spot in the UK NSS and comparable instruments, and what a low score actually tells course-evaluation and QA teams.
Propensity Score Matching for Course-Evaluation Confounds
How propensity score matching (PSM) adjusts student-evaluation-of-teaching comparisons for confounds like class size, course level, discipline, and grading — and where it honestly falls short.
Coefficient Omega vs Cronbach's Alpha: Reporting the Reliability of a Course-Evaluation Scale
Cronbach's alpha assumes every item measures the construct equally well — an assumption course-evaluation subscales rarely meet. Here is why McDonald's omega is the more defensible reliability coefficient, and how to report it.
Some of Your Course Evaluations Were Answered in Ten Seconds: Response Latency as a Data-Quality Signal
Response times ('paradata') are a free, objective quality signal in online course evaluations. Evidence from Zhang & Conrad and others shows speeders straightline more — and how to use latency to screen data without deleting honest fast responses.
Should You Statistically Adjust Course Evaluations for Class Size and Difficulty? The IDEA Approach to Adjusted Scores
Some course-evaluation systems report "adjusted" scores that statistically correct for class size, discipline difficulty and student motivation. We examine what the IDEA system actually adjusts for, whether the practice is defensible, and how to contextualise scores without over-correcting.
Importance-Performance Analysis: Turning Course-Evaluation Scores into a Priority Map
A course-evaluation report that lists twenty item means tells you nothing about where to act first. Importance-Performance Analysis (IPA) plots each attribute by how much it matters to students against how well you did, producing a four-quadrant map that separates urgent fixes from wasted effort.
Beyond the Average Student: Latent Profile Analysis for Course-Evaluation Segments
A course mean of 3.6 can hide two entirely different student experiences averaged into one number. Latent profile analysis (LPA) recovers those hidden subgroups from the evaluation data itself, so you can see the delighted minority and the alienated cohort that the average erased.
Did COVID Lower Course Evaluations? Emergency Remote Teaching as a Natural Experiment in Confounding
What student-evaluation data from the 2020 emergency shift to remote teaching reveals about how much SET scores reflect factors outside an instructor's control — and why pandemic-era ratings need a giant asterisk in any personnel decision.
When Combining Sections Reverses the Result: Simpson's Paradox in Course-Evaluation Data
Aggregating course-evaluation scores across sections, cohorts, or years can reverse the very conclusion you are trying to draw. What Simpson's paradox is, how it appears in evaluation data, and how to report so the reversal cannot bite you.
The Course Reviewed Straight After a Brilliant One Looks Worse: Contrast Effects and Narrow Bracketing in Evaluation Review
Course-evaluation bias research focuses on the student filling in the form. The evidence on sequential judgement says the committee reading twenty reports in an afternoon is biased too — by what it read immediately before.
National Student Surveys Are Not Course Evaluations: What NSS and Studiebarometeret Can and Cannot Tell You
National student surveys sit at the wrong level of analysis to diagnose a course. Cheng and Marsh showed that most apparent difference between UK universities is not reliable variance at all — here is how to use national data alongside your own instrument instead of in place of it.
Does the Number Anchor the Committee Before It Reads a Word? Anchoring Bias in Evaluation Review
When a review panel sees an instructor's 3.8 mean before reading the comments, that number quietly pulls every later judgement toward it. What the anchoring literature says, and how to sequence evaluation review to resist it.
How Padding a Report With Extra Data Weakens a Strong Signal: The Dilution Effect in Evaluation Review
A classic finding (Nisbett, Zukier & Lemley, 1981) shows that adding irrelevant, non-diagnostic information makes judgements less extreme. For evaluation committees, a padded report can dilute a genuinely strong signal.
Was It the Teacher or the Situation? The Fundamental Attribution Error in Reading Course Evaluations
Committees read a low evaluation score as evidence of a weak teacher. Social-psychology research on the fundamental attribution error shows why that inference is systematically biased — and how to read scores in situational context.
Which Instructor Scores Are Genuinely Unusual? Funnel Plots for Fair Course-Evaluation Comparison
Ranking instructors by mean evaluation score turns sampling noise into a league table. Funnel plots — Spiegelhalter's method for institutional comparison — show which scores are genuinely unusual and which are just small-sample wobble.
Reading Evaluations to Confirm What You Already Believe: Confirmation Bias in Interpreting Course Feedback
Whoever reads a course evaluation already has a hypothesis about the instructor. Confirmation bias shapes which comments they weight, how ambiguity is resolved, and what the "data" is taken to show. Here is the evidence and the guardrails.
Fair Confidence Intervals for Small Classes: The Bootstrap for Course-Evaluation Reporting
Small classes and skewed rating distributions break the textbook confidence interval. The bootstrap resamples the data you actually have to produce honest uncertainty bounds. Here is the method, its limits, and how to report it.
One Vivid Comment Is Not a Pattern: Base-Rate Neglect in Reading Course Evaluations
A single scathing open-text comment can outweigh forty neutral ones in a reviewer's mind. Bar-Hillel and Kahneman & Tversky showed why: people ignore base rates in favour of vivid, individuating detail. Here is how base-rate neglect distorts course-evaluation review — and how to design against it.
Is That Score Gap Real or Just Luck of the Draw? Permutation Tests for Course-Evaluation Comparisons
Comparing two instructors' evaluation averages with a t-test quietly assumes normal, equal-variance data you rarely have with small, skewed Likert samples. Permutation tests, formalised by Fisher and reviewed by Ernst, answer the comparison question by shuffling the data itself — with almost no distributional assumptions. Here is when and how to use them.
Equivalence Testing (TOST): How to Show Two Instructors Really Do Score the Same
A non-significant t-test never proves two course-evaluation means are equivalent. Equivalence testing (TOST) does — here is what the method is, how to set a smallest effect size of interest, and how to use it for defensible decisions.
Denominator Neglect: Why Raw Comment Counts and Percentages Mislead Course-Evaluation Readers
Denominator neglect and ratio bias make "five complaints" and "20% negative" feel worse than they are. Here is what the research says about this reasoning error and how to report course evaluations so it does not distort decisions.
The Availability Heuristic: Why a Few Vivid Comments Distort How You Read Course Evaluations
The availability heuristic means the most memorable, extreme open-text comment feels more frequent than it is. What the research says, why it corrupts qualitative course-evaluation review, and how to read comments by prevalence instead of vividness.
Ordinal Regression for Course-Evaluation Data: Why Cumulative-Link Models Beat Averaging Likert Scores
Averaging Likert-scale course-evaluation items treats ordered categories as if they were equal-interval numbers. What the research says about the errors this causes, and why cumulative-link (ordinal) regression is the defensible analysis method.
Common-Language Effect Size: Reporting Course-Evaluation Differences People Actually Understand
A 0.2-point gap in mean ratings means nothing to a committee. The common-language effect size (probability of superiority) restates a difference as a probability anyone can interpret. What the research says and how to report course evaluations honestly.
Compare 60 Instructors, Expect 3 False Alarms: Multiple Comparisons and the False Discovery Rate
When a QA office tests every instructor against a benchmark, chance alone produces "significant" outliers. What the multiple-comparisons literature — Bonferroni, and Benjamini & Hochberg''s false discovery rate — says about flagging course-evaluation scores fairly.
Where Should You Set the Review Trigger? Signal Detection Theory for Course-Evaluation Thresholds
Deciding which courses to flag for review is a signal-detection problem. What Swets, Dawes & Monahan and the ROC literature say about separating how well a score discriminates from where you set the cut — and how to choose a threshold that reflects the real cost of misses versus false alarms.
Collider Bias: How Selecting Who Gets Evaluated Invents Correlations That Are Not There
Berkson''s paradox and collider bias explain why analysing only the students who respond — or only the courses that survive — can manufacture correlations that do not exist in the population. What the causal-inference literature says, and why controlling for a collider makes things worse.
Did the Teaching Change Cause the Score to Move? Difference-in-Differences for Course Evaluation
Difference-in-differences lets you estimate whether a course redesign actually moved evaluation scores by comparing a treated course against a similar untouched one over time. Here is how the design works, when its parallel-trends assumption holds, and how to use it honestly in quality assurance.
One Instructor, Two Hundred Defensible Scores: Specification-Curve Analysis for Course Evaluation
Every course-evaluation "score" is the product of dozens of defensible analytic choices — which items count, how to weight, mean or median, whether to adjust for class size. Specification-curve and multiverse analysis compute the result across all of them, so a personnel decision rests on the pattern, not on one analyst's arbitrary path.
The Department Average Is Not the Student: The Ecological Fallacy in Course Evaluation
A correlation that holds between department averages need not hold — and can even reverse — for individual students. The ecological fallacy is inferring individual-level relationships from group-level data. Here is how it distorts course-evaluation analysis and how to reason at the level your decision is actually made.
Your 'Significant' Instructor Difference Might Point the Wrong Way: Type S and Type M Errors in Small-Cohort Evaluation
Statistical significance does not protect a small-cohort evaluation from pointing the wrong way or exaggerating the gap. Type S and Type M error analysis shows why, and what to report instead.
Does the Effect Depend on Who or What? Moderation and Interaction Effects in Course-Evaluation Analysis
A bias or a teaching effect on evaluations often holds only for some students, courses or conditions. Moderation analysis tests those "it depends" claims properly, and it is far harder than it looks.
Is There Really No Difference, or Just No Evidence? Bayes Factors for Course-Evaluation Comparisons
A non-significant p-value cannot confirm two instructors scored the same — it only fails to reject. Bayes factors quantify evidence FOR the null as well as against it, distinguishing absence of evidence from evidence of absence in course-evaluation comparisons.
Build a Synthetic Comparison Department: Synthetic Control Methods for a Teaching Change
When one programme redesigns its teaching, you rarely have a clean control group. Synthetic control methods build a weighted "synthetic" comparator from other programmes so you can estimate whether the change actually moved evaluation scores.
How Strong Would the Hidden Confounder Have to Be? The E-Value for Adjusted Course-Evaluation Comparisons
Every adjusted course-evaluation claim invites the objection "but you did not control for X". The E-value, from epidemiology, quantifies exactly how strong that unmeasured X would have to be to explain away your finding — turning a vague worry into a number.
When a Few Retaliatory 1s Sink the Average: Robust Estimators for Course Evaluation
In a class of twelve, two vindictive 1s can drag the mean by half a point. The arithmetic mean has a breakdown point of zero — a single extreme value can move it arbitrarily far. Trimmed means, Winsorizing, and M-estimators resist that without throwing away respondents.
How Many Dimensions Does Your Evaluation Really Measure? Parallel Analysis for Factor Retention
Deciding how many factors your evaluation instrument measures with the eigenvalue-greater-than-one rule or the scree plot routinely gets the wrong answer. Parallel analysis (Horn, 1965) compares your data against random noise — and is one of the most accurate methods available.
The Average Hides the Tails: Quantile Regression for Course-Evaluation Data
Ordinary regression models the mean, but a teaching change can lift the median while sinking the unhappiest students. Quantile regression (Koenker & Bassett, 1978) models the whole distribution, revealing effects the average conceals.
When Can You Stop Collecting? Optional Stopping and Sequential Analysis for Rolling Course Evaluation
Watching results accumulate in a live course-evaluation dashboard and stopping when a difference looks significant inflates false positives. This guide explains the optional-stopping problem, the sequential-analysis methods that fix it, the evidence, the caveats, and how Koji handles rolling collection responsibly.
Reporting a Class of Four Without Exposing Anyone: Cell Suppression, k-Anonymity, and Differential Privacy
You cannot safely report an average for a class of four students without risking that individuals are identified or their comments inferred. This guide covers minimum-cell suppression, k-anonymity, and differential privacy for course-evaluation reporting — the methods, the evidence, the privacy-utility trade-off, and how Koji protects small-cohort respondents.
One Student, Many Teachers: Cross-Classified and Multiple-Membership Models for Fair Course Evaluation
Standard multilevel models assume a clean hierarchy, but students are taught by several instructors and instructors teach across programmes. Cross-classified and multiple-membership models partition that tangled variance honestly — and change which instructors look unusual.
Where Exactly Do Students Abandon Your Evaluation? Discrete-Time Survival Analysis of Breakoff
A completion rate tells you how many students quit; it cannot tell you where or why. Discrete-time survival analysis models the hazard of breakoff question by question, turning a single number into an actionable map of where your evaluation loses people.
Exactly Balanced Comparison Groups Without the Guesswork: Entropy Balancing for Course Evaluation
Propensity-score matching throws data away and needs you to iterate a model until the groups look balanced. Entropy balancing reweights the data so the groups are exactly balanced on the moments you specify, in one step. Here is what it does for fair course-evaluation comparison.
A High Correlation Does Not Mean Two Evaluation Methods Agree: The Bland-Altman Limits of Agreement
Two evaluation methods can correlate strongly yet disagree by a full scale point. Bland-Altman limits of agreement plot the differences, not the correlation, to show whether student ratings, peer review or AI coding can actually be used interchangeably.
Mapping the Patterns You Cannot Average: Multiple Correspondence Analysis for Categorical Course-Evaluation Data
Much course-evaluation data is genuinely categorical — programme, mode, agree/disagree, chosen theme. Multiple correspondence analysis (MCA) maps how those categories cluster on a two-dimensional plane, revealing response patterns that averaging destroys.
High Agreement, Low Kappa: Choosing an Agreement Coefficient for Coding Course Feedback
Cohen's kappa can collapse to near zero even when two coders agree on 95 percent of comments. Here is why the kappa paradox happens and when to report Gwet's AC1 or Krippendorff's alpha instead.
Did This Student Really Change? The Reliable Change Index for Mid-to-End Course Evaluation
When a student's mid-semester and end-of-semester ratings differ, the Reliable Change Index tells you whether the shift is larger than measurement error alone would produce - a per-individual test borrowed from clinical psychology.
Combining Several Rankings Into One Fair Order: Rank Aggregation for Course Evaluation
When you have to merge several rankings - by different criteria, cohorts or panel members - into one, the method you pick changes the winner. Kemeny, Borda and Condorcet from social-choice theory show why, and how to do it defensibly.
Multidimensional Scaling for Course Evaluation: A Perceptual Map of How Students See Your Courses
Multidimensional scaling turns a table of similarities into a two-dimensional map. Here is how MDS reveals the hidden structure in course-evaluation data that averages and factor analysis both miss.
Count Models for Course Evaluation: Why You Should Not Average the Number of Comments
How many students wrote a comment? How many mentioned assessment? These are counts, and averaging them or running OLS gives biased answers. Poisson, negative-binomial, and zero-inflated models do it right.
Which Aspects of Teaching Actually Drive the Overall Score? Relative Weights and Dominance Analysis
When clarity, workload, support, and feedback are all correlated, regression coefficients cannot tell you which one drives the overall rating. Relative weights and dominance analysis partition the explained variance fairly.
Which Teaching Behaviours Drive the Overall Score — Nonlinearly? Random Forests and Variable Importance for Course Evaluation
Random forests and their variable-importance measures reveal which teaching items predict the overall rating when the relationships are nonlinear and interacting — but Gini importance is biased toward correlated and high-cardinality predictors, so use conditional permutation importance and read the output as association, not cause.
Turning Several Quality Signals Into One Defensible Decision: The Analytic Hierarchy Process for Course Evaluation
The Analytic Hierarchy Process derives criterion weights from pairwise comparisons and checks their internal consistency, giving a transparent, auditable way to combine student ratings, peer observation and learning outcomes into one decision — but it is subjective, vulnerable to rank reversal, and its 1–9 scale and 0.10 consistency threshold are conventions, not laws.
What Will This Instructor Score Next Term — With a Guarantee? Conformal Prediction for Course Evaluation
Conformal prediction wraps any scoring model in a distribution-free prediction interval that is guaranteed to cover the true next value at a rate you choose. Here is what it does, why it fits course evaluation, and how to use it honestly.
Your Outcome Is a Proportion, Not a Mean: Beta Regression for Course Evaluation
When the thing you are modelling is a proportion between 0 and 1 — the share recommending a course, the top-box rate — ordinary regression misbehaves at the boundaries. Beta regression is the purpose-built tool. Here is how it works and when to use it.
The Average Effect Across Everyone, Not Inside One Classroom: GEE for Course Evaluation
When responses are clustered in courses but you want the population-average effect of a change — not the effect for a specific classroom — generalized estimating equations are the right tool, and they are robust to getting the correlation structure wrong.
Relating a Battery of Teaching Behaviours to a Battery of Outcomes: Canonical Correlation Analysis
When you want to relate a whole set of teaching-behaviour items to a whole set of outcome items, running many pairwise correlations misleads. Canonical correlation analysis maps the joint structure.
When Did the Rating Actually Shift? Changepoint Detection for Course-Evaluation Time Series
A rating series drifts and the real question is when it shifted, without knowing the date in advance. Changepoint detection dates unknown structural breaks in term-by-term scores.
Your Comment Categories Sum to 100%: Compositional Data Analysis for Course-Evaluation Shares
When course-evaluation numbers are shares of a fixed whole, ordinary means and correlations mislead. Compositional data analysis and log-ratios fix the closure trap.