New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs

Analysis & reporting

Quality scoring, thematic analysis, triangulation, and report generation.

Does Test Anxiety Distort Course Evaluations — and Should Evaluations Measure It?

A 30-year meta-analysis confirms test anxiety reliably depresses performance. That makes it both a hidden confound in course ratings and, arguably, a course-quality signal worth capturing.

10 min readadvanced

Comparing Instructors on Many Criteria at Once: Data Envelopment Analysis for Course Evaluation

Data Envelopment Analysis (DEA) rates instructors, modules, or departments against the best observed performers on several inputs and outputs at once — a fair multi-criteria alternative to ranking on one mean.

9 min readadvanced

Measuring Consensus, Not Just the Average: Mixed-Effects Location-Scale Models for Course Evaluation

Mixed-effects location-scale models model the spread of course-evaluation ratings — consensus versus polarisation — as an outcome in its own right, not just noise around the mean.

9 min readadvanced

How Many Responses Do You Need for a Reliable Course Evaluation?

Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.

10 min readadvanced

Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You

Natural-language processing can turn thousands of free-text course-evaluation comments into themes and sentiment at scale — but the research (Cunningham-Nelson 2019; Sunar & Khalid 2023) shows where automated analysis is reliable and where human judgement is still required.

10 min readintermediate

Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved

Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.

9 min readadvanced

Is the Workload-Rating Relationship a Straight Line? Splines and Generalized Additive Models for Course Evaluation

When you regress a rating on workload, difficulty, or class size, a straight line assumes each extra unit matters the same everywhere. Restricted cubic splines and generalized additive models let the data reveal the real curve — often a U-shape or a plateau — without slicing a continuous predictor into arbitrary bins.

10 min readadvanced

When Ratings Pile Up at 5: Tobit and Censored Regression for Ceiling-Bounded Course Evaluation

When a third of your class marks the top of the scale, the 5 is a floor on their true opinion, not its ceiling. Ordinary regression on such data underestimates real differences. The Tobit model treats the pile-up at the boundary as censoring and recovers the effect that a naive analysis flattens.

10 min readadvanced

Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough

Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.

11 min readadvanced

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

10 min readintermediate

Pooling Your Own Sections Without Faking Precision: Random-Effects Meta-Analysis for Course Evaluation

When you combine ratings across many small sections or terms, averaging the averages pretends they all measure one fixed truth. Random-effects meta-analysis treats each section as a noisy estimate of a genuinely varying effect, weights them properly, and reports how much real spread remains — with a prediction interval, not just a mean.

10 min readadvanced

How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis

When you turn thousands of free-text comments into themes and counts, how do you know the coding is trustworthy? O Connor and Joffe (2020) on intercoder reliability, Braun and Clarke on thematic analysis, and what rigorous qualitative QA looks like.

10 min readadvanced

Why a Single Student Survey Can't Stand Alone: Common-Method Bias in Course Evaluation

Common-method bias (Podsakoff et al., 2003) explains why correlations inside a single end-of-term student survey are inflated by the shared method itself - and why triangulating teaching evidence matters. A research-grounded guide for quality assurance.

10 min readadvanced

Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?

Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.

10 min readadvanced

Cronbach's Alpha and Course-Evaluation Reliability: What It Does and Doesn't Tell You

A high Cronbach's alpha on your course-evaluation instrument is widely read as proof of a "reliable" survey. The psychometric literature says that reading is largely wrong. What alpha actually measures, where it misleads, and what to report instead.

10 min readadvanced

How Much of a Course-Evaluation Change Actually Matters? The Minimal Important Difference

Statistical significance tells you a course-evaluation change is real; the Minimal Important Difference (MID) tells you whether it is big enough to matter. Here is how to set one, using anchor-based and distribution-based methods imported from health measurement.

10 min readintermediate

Can You Compare Course Ratings Across Cultures? Anchoring Vignettes and the King Method

When students from different countries interpret the same rating scale differently, their scores are not comparable. King, Murray, Salomon and Tandon (2004) introduced anchoring vignettes to correct this. Here is how the technique works and what it means for international and multi-campus course evaluation.

10 min readadvanced

Are Your Respondents Representative? Using Early-vs-Late Wave Analysis to Estimate Nonresponse Bias

A low response rate is not automatically biased — what matters is whether respondents differ from non-respondents. Armstrong and Overton (1977) gave us a cheap diagnostic: compare early and late responders. Here is how to use wave analysis on course-evaluation data and where it breaks down.

9 min readintermediate

When the AI Summary Says Something No Student Did: Faithfulness and Hallucination in LLM Course-Feedback Summaries

A large language model that summarises hundreds of open-text comments can invent a theme, a sentiment, or even a quote that no student wrote. The abstractive-summarisation research explains why - and what a defensible AI feedback pipeline must do to stay faithful to the source.

11 min readintermediate

Should You Use Net Promoter Score for Courses? The "Would You Recommend" Question in Higher Education

Net Promoter Score is migrating from customer experience into student feedback. What Reichheld (2003) actually claimed, why Keiningham et al. (2007) failed to replicate its superiority, and whether a single recommend-question belongs in course evaluation.

8 min readintermediate

The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise

Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.

9 min readadvanced

Can You Report a Class Mean? ICC(1), ICC(2), and r_wg for Aggregating Student Ratings

Before you average student ratings into a class score, three organisational-psychology indices decide whether you can: r_wg (within-class agreement), ICC(1) (how much variance is between classes), and ICC(2) (the reliability of the class mean).

10 min readadvanced

You Changed the Questionnaire — Can You Still Compare Years? Test Equating and Linking for Course Evaluation

When you revise a course-evaluation form, scores before and after are not automatically comparable. Kolen and Brennan's equating/linking/prediction hierarchy, and the anchor-item designs behind it, tell you what a cross-revision trend can honestly claim.

10 min readadvanced

Beyond Keyword Counts: Semantic Text Embeddings for Searching, Deduplicating, and Mapping Open-Text Feedback

Text embeddings turn each open-text comment into a point in a meaning-space where semantically similar feedback sits close together — enabling semantic search, near-duplicate detection, coverage measurement, and routing that keyword counts and topic models cannot deliver.

9 min readintermediate

An Evaluation of Course Evaluations: What the Berkeley Statisticians Concluded About Reporting Student Ratings

Two Berkeley statisticians argued that averaging student-rating scores to rank instructors for promotion and tenure should be abandoned for both statistical and substantive reasons. What Stark & Freishtat (2014) actually recommend — and how to report ratings responsibly instead.

10 min readadvanced

Gendered Language in Student Comments: Why Men Are ''Brilliant'' and Women Are ''Caring''

Bias in course evaluations is not only in the numbers — it is in the words. What Mitchell & Martin (2018) and Storage et al. (2016) found about systematically different language applied to men and women, and why open-text analysis must account for it.

10 min readadvanced

Do Students'' Written Comments Match Their Ratings? What Concordance Tells You

Open-text comments and Likert scores usually agree — but the gaps are where the insight lives. What Alhija & Fresko (2009) and Brockx et al. (2012) found about the consistency between qualitative and quantitative course-evaluation data, and how to read it.

9 min readintermediate

A Mismeasured Predictor Biases Its Own Slope: Errors-in-Variables and Regression Dilution in Course Evaluation

Self-reported workload, engagement and interest are all measured with error, so their regression slopes are attenuated toward zero. Why more responses do not fix it, and how reliability lets you correct it.

10 min readadvanced

Even Fair Evaluations Can Rank Unfairly: Why Course-Evaluation Scores Misclassify Instructors

A simulation-based finding that even unbiased, reliable, and valid course-evaluation scores produce high misclassification rates when used to rank or compare instructors — and what it means for tenure, promotion, and merit decisions.

10 min readadvanced

Some Raters Are Harsh, Some Items Are Hard: Why Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores

Averaging Likert scores treats every student as an equally calibrated measuring instrument and every item as equally hard. Rasch and Many-Facet Rasch models do not. We explain the method, the evidence that rater severity and item difficulty distort raw means, the limitations, and how Koji applies the same logic.

11 min readadvanced

Should You Average a 1–5 Course-Evaluation Scale at All? The Ordinal-vs-Interval Debate

Is it legitimate to compute a mean from a five-point course-evaluation scale? The fifty-year ordinal-versus-interval debate, from Jamieson to Carifio and Perla to Norman, and what it means for reporting student ratings.

9 min readadvanced

When You Cannot Assume the Non-Responders Are Like the Responders: Manski Bounds for Course Evaluation

Weighting and imputation assume the silent majority resemble the responders. Partial identification refuses that assumption and reports the interval the true mean could occupy, showing exactly how much your conclusion rests on belief.

10 min readadvanced

Why One Cruel Comment Outweighs Twenty Kind Ones: Negativity Bias in Reading Course Evaluations

Instructors and committees fixate on the harshest open-text comment and discount the praise. Baumeister''s "bad is stronger than good" and the negativity-bias literature explain why, and how to read qualitative course feedback fairly.

8 min readintermediate

Correcting for Who Chose to Respond: The Heckman Selection Model for Course-Evaluation Nonresponse

When students who respond to course evaluations differ from those who skip them on the very thing you are measuring, reweighting cannot help. The Heckman selection model models response and rating jointly to correct for selection on unobservables — under fragile assumptions this article makes explicit.

9 min readadvanced

Control for Everything Stable About an Instructor: Fixed-Effects Panel Models for Course Evaluation

Fixed-effects panel models compare each instructor only to themselves across terms, silently controlling for every stable confounder — measured or not. This article explains the within-transformation, the FE-vs-random-effects choice and Hausman test, and what FE can and cannot tell a quality office.

9 min readadvanced

Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations

Year-over-year movements in course-evaluation scores are mostly statistical noise regressing toward an instructor's true average. Here is why regression to the mean fools quality-assurance committees, and how to read evaluation trends responsibly.

10 min readadvanced

When Almost Everyone Scores 4.5: Ceiling Effects, Skew, and What They Do to Course-Evaluation Statistics

Course-evaluation ratings pile up at the top of the scale, producing a strong ceiling effect and negative skew that breaks the statistics most universities still report. Here is what the evidence shows and how to report ratings honestly.

9 min readadvanced

How Much of the Rating Gap Is Bias? The Oaxaca-Blinder Decomposition for Course Evaluation

When two groups of instructors get different average ratings, the Oaxaca-Blinder decomposition splits the gap into a part explained by measurable circumstances and an unexplained residual. This article shows how to use it responsibly — and why the unexplained part is a bound on bias, not a measurement of it.

9 min readadvanced

Is Your Missing Course-Evaluation Data Random? MAR, MNAR, and What to Do About It

A low response rate is a missing-data problem. Rubin's MCAR/MAR/MNAR framework explains when a course-evaluation mean is biased and when multiple imputation or maximum likelihood can help.

10 min readadvanced

You Flagged the Lowest Instructor, Then Tested If They Were Below Average: Selective Inference and the Winner's Curse

Ranking instructors and then testing the extreme one invalidates the p-value and biases the flagged score. Selective inference, FCR intervals, and empirical-Bayes shrinkage put the inference right.

9 min readadvanced

Your Ratings Have a Term Rhythm: STL Seasonal-Trend Decomposition for Rolling Course Evaluation

Continuous feedback carries a semester rhythm on top of a trend on top of noise. STL decomposes the series so you read real change without being fooled by the calendar.

9 min readintermediate

Can a Large Language Model Code Your Open-Text Course Feedback? What the Agreement Studies Show

Peer-reviewed evidence on how closely GPT-4-class models match human coders when categorising open-ended student comments — agreement statistics, where they fail, and how to use them responsibly in course-evaluation quality assurance.

9 min readadvanced

Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores

Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.

10 min readadvanced

Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages for Course Evaluation

Course-evaluation data has a nested structure — students within sections within instructors — and a flat average ignores it. This guide explains how multilevel (hierarchical linear) models partition variance, why ignoring clustering understates uncertainty, and what it means for fair reporting.

10 min readadvanced

Why Is the Effect Bigger in Some Sections? Meta-Regression for Course Evaluation

Meta-regression explains between-section heterogeneity in pooled course-evaluation results by modelling section-level moderators — powerful for generating explanations, weak for confirming them.

9 min readadvanced

Beyond "Positive or Negative": Aspect-Based Sentiment Analysis of Open-Text Course Feedback

Document-level sentiment scoring collapses a rich student comment into one number and hides the most useful signal. Aspect-based sentiment analysis (ABSA) separates *what* the comment is about from *how* the student feels about it. Here is the evidence on ABSA for course feedback, its accuracy, its failure modes, and how Koji applies it.

11 min readadvanced

Should You Report an Instructor''s Percentile? Norm-Referenced vs Criterion-Referenced Course-Evaluation Scores

Telling a lecturer they are "in the 40th percentile of the department" is norm-referenced reporting — and it manufactures losers by construction, no matter how good everyone is. Criterion-referenced reporting asks instead whether teaching met a defined standard. Here is the evidence on why the choice matters and how to report responsibly.

11 min readadvanced

Topic Modeling for Open-Text Course Evaluations: LDA, STM and BERTopic Explained

How LDA, structural topic models and BERTopic surface themes across tens of thousands of open-text student comments — what each method does well, where it fails, and how to keep a human in the loop.

10 min readadvanced

Top-Box vs Mean: How to Report Course Evaluation Scores Without Throwing Away Information

Reporting the percentage of students who chose the top box feels intuitive, but collapsing a scale to favorable/unfavorable discards information. Here is what the measurement evidence says and how to report responsibly.

9 min readintermediate

Did the Teaching Change Actually Work? Interrupted Time Series for Course-Evaluation Trends

Comparing this year''s evaluation mean to last year''s cannot tell you whether a curriculum redesign worked. Interrupted time series with segmented regression can — here is how to apply it, and where it breaks.

10 min readadvanced

The Framework Method for Open-Text Course Feedback: A Structured Alternative to Thematic Coding

When a quality committee — not a lone qualitative researcher — has to make sense of thousands of student comments, the Framework Method offers a transparent, auditable matrix-based approach. What Gale et al. (2013) actually proposed, and how to use it.

10 min readintermediate

Statistical Process Control for Course Evaluations: Using Control Charts to Separate Signal from Noise

How Shewhart control charts and statistical process control distinguish real changes in teaching quality from ordinary random variation in course evaluation scores - so QA teams stop reacting to noise. Grounded in Sivena and Nikolaidis (2019, 2022).

10 min readadvanced

Why "Assessment and Feedback" Consistently Scores Lowest in Student Surveys

A research-grounded explanation of why assessment and feedback is the perennial weak spot in the UK NSS and comparable instruments, and what a low score actually tells course-evaluation and QA teams.

11 min readadvanced

Propensity Score Matching for Course-Evaluation Confounds

How propensity score matching (PSM) adjusts student-evaluation-of-teaching comparisons for confounds like class size, course level, discipline, and grading — and where it honestly falls short.

14 minadvanced

Coefficient Omega vs Cronbach's Alpha: Reporting the Reliability of a Course-Evaluation Scale

Cronbach's alpha assumes every item measures the construct equally well — an assumption course-evaluation subscales rarely meet. Here is why McDonald's omega is the more defensible reliability coefficient, and how to report it.

9 min readadvanced

Some of Your Course Evaluations Were Answered in Ten Seconds: Response Latency as a Data-Quality Signal

Response times ('paradata') are a free, objective quality signal in online course evaluations. Evidence from Zhang & Conrad and others shows speeders straightline more — and how to use latency to screen data without deleting honest fast responses.

9 min readadvanced

Should You Statistically Adjust Course Evaluations for Class Size and Difficulty? The IDEA Approach to Adjusted Scores

Some course-evaluation systems report "adjusted" scores that statistically correct for class size, discipline difficulty and student motivation. We examine what the IDEA system actually adjusts for, whether the practice is defensible, and how to contextualise scores without over-correcting.

10 min readadvanced

Importance-Performance Analysis: Turning Course-Evaluation Scores into a Priority Map

A course-evaluation report that lists twenty item means tells you nothing about where to act first. Importance-Performance Analysis (IPA) plots each attribute by how much it matters to students against how well you did, producing a four-quadrant map that separates urgent fixes from wasted effort.

9 min readintermediate

Beyond the Average Student: Latent Profile Analysis for Course-Evaluation Segments

A course mean of 3.6 can hide two entirely different student experiences averaged into one number. Latent profile analysis (LPA) recovers those hidden subgroups from the evaluation data itself, so you can see the delighted minority and the alienated cohort that the average erased.

9 min readadvanced

Did COVID Lower Course Evaluations? Emergency Remote Teaching as a Natural Experiment in Confounding

What student-evaluation data from the 2020 emergency shift to remote teaching reveals about how much SET scores reflect factors outside an instructor's control — and why pandemic-era ratings need a giant asterisk in any personnel decision.

9 min readintermediate

When Combining Sections Reverses the Result: Simpson's Paradox in Course-Evaluation Data

Aggregating course-evaluation scores across sections, cohorts, or years can reverse the very conclusion you are trying to draw. What Simpson's paradox is, how it appears in evaluation data, and how to report so the reversal cannot bite you.

9 min readadvanced

The Course Reviewed Straight After a Brilliant One Looks Worse: Contrast Effects and Narrow Bracketing in Evaluation Review

Course-evaluation bias research focuses on the student filling in the form. The evidence on sequential judgement says the committee reading twenty reports in an afternoon is biased too — by what it read immediately before.

10 min readadvanced

National Student Surveys Are Not Course Evaluations: What NSS and Studiebarometeret Can and Cannot Tell You

National student surveys sit at the wrong level of analysis to diagnose a course. Cheng and Marsh showed that most apparent difference between UK universities is not reliable variance at all — here is how to use national data alongside your own instrument instead of in place of it.

9 min readadvanced

Does the Number Anchor the Committee Before It Reads a Word? Anchoring Bias in Evaluation Review

When a review panel sees an instructor's 3.8 mean before reading the comments, that number quietly pulls every later judgement toward it. What the anchoring literature says, and how to sequence evaluation review to resist it.

9 min readadvanced

How Padding a Report With Extra Data Weakens a Strong Signal: The Dilution Effect in Evaluation Review

A classic finding (Nisbett, Zukier & Lemley, 1981) shows that adding irrelevant, non-diagnostic information makes judgements less extreme. For evaluation committees, a padded report can dilute a genuinely strong signal.

10 min readadvanced

Was It the Teacher or the Situation? The Fundamental Attribution Error in Reading Course Evaluations

Committees read a low evaluation score as evidence of a weak teacher. Social-psychology research on the fundamental attribution error shows why that inference is systematically biased — and how to read scores in situational context.

10 min readadvanced

Which Instructor Scores Are Genuinely Unusual? Funnel Plots for Fair Course-Evaluation Comparison

Ranking instructors by mean evaluation score turns sampling noise into a league table. Funnel plots — Spiegelhalter's method for institutional comparison — show which scores are genuinely unusual and which are just small-sample wobble.

10 min readadvanced

Reading Evaluations to Confirm What You Already Believe: Confirmation Bias in Interpreting Course Feedback

Whoever reads a course evaluation already has a hypothesis about the instructor. Confirmation bias shapes which comments they weight, how ambiguity is resolved, and what the "data" is taken to show. Here is the evidence and the guardrails.

9 min readintermediate

Fair Confidence Intervals for Small Classes: The Bootstrap for Course-Evaluation Reporting

Small classes and skewed rating distributions break the textbook confidence interval. The bootstrap resamples the data you actually have to produce honest uncertainty bounds. Here is the method, its limits, and how to report it.

10 min readadvanced

One Vivid Comment Is Not a Pattern: Base-Rate Neglect in Reading Course Evaluations

A single scathing open-text comment can outweigh forty neutral ones in a reviewer's mind. Bar-Hillel and Kahneman & Tversky showed why: people ignore base rates in favour of vivid, individuating detail. Here is how base-rate neglect distorts course-evaluation review — and how to design against it.

8 min readintermediate

Is That Score Gap Real or Just Luck of the Draw? Permutation Tests for Course-Evaluation Comparisons

Comparing two instructors' evaluation averages with a t-test quietly assumes normal, equal-variance data you rarely have with small, skewed Likert samples. Permutation tests, formalised by Fisher and reviewed by Ernst, answer the comparison question by shuffling the data itself — with almost no distributional assumptions. Here is when and how to use them.

9 min readadvanced

Equivalence Testing (TOST): How to Show Two Instructors Really Do Score the Same

A non-significant t-test never proves two course-evaluation means are equivalent. Equivalence testing (TOST) does — here is what the method is, how to set a smallest effect size of interest, and how to use it for defensible decisions.

10 min readadvanced

Denominator Neglect: Why Raw Comment Counts and Percentages Mislead Course-Evaluation Readers

Denominator neglect and ratio bias make "five complaints" and "20% negative" feel worse than they are. Here is what the research says about this reasoning error and how to report course evaluations so it does not distort decisions.

9 min readintermediate

The Availability Heuristic: Why a Few Vivid Comments Distort How You Read Course Evaluations

The availability heuristic means the most memorable, extreme open-text comment feels more frequent than it is. What the research says, why it corrupts qualitative course-evaluation review, and how to read comments by prevalence instead of vividness.

9 min readintermediate

Ordinal Regression for Course-Evaluation Data: Why Cumulative-Link Models Beat Averaging Likert Scores

Averaging Likert-scale course-evaluation items treats ordered categories as if they were equal-interval numbers. What the research says about the errors this causes, and why cumulative-link (ordinal) regression is the defensible analysis method.

10 min readadvanced

Common-Language Effect Size: Reporting Course-Evaluation Differences People Actually Understand

A 0.2-point gap in mean ratings means nothing to a committee. The common-language effect size (probability of superiority) restates a difference as a probability anyone can interpret. What the research says and how to report course evaluations honestly.

9 min readintermediate

Compare 60 Instructors, Expect 3 False Alarms: Multiple Comparisons and the False Discovery Rate

When a QA office tests every instructor against a benchmark, chance alone produces "significant" outliers. What the multiple-comparisons literature — Bonferroni, and Benjamini & Hochberg''s false discovery rate — says about flagging course-evaluation scores fairly.

10 min readadvanced

Where Should You Set the Review Trigger? Signal Detection Theory for Course-Evaluation Thresholds

Deciding which courses to flag for review is a signal-detection problem. What Swets, Dawes & Monahan and the ROC literature say about separating how well a score discriminates from where you set the cut — and how to choose a threshold that reflects the real cost of misses versus false alarms.

10 min readadvanced

Collider Bias: How Selecting Who Gets Evaluated Invents Correlations That Are Not There

Berkson''s paradox and collider bias explain why analysing only the students who respond — or only the courses that survive — can manufacture correlations that do not exist in the population. What the causal-inference literature says, and why controlling for a collider makes things worse.

11 min readadvanced

Did the Teaching Change Cause the Score to Move? Difference-in-Differences for Course Evaluation

Difference-in-differences lets you estimate whether a course redesign actually moved evaluation scores by comparing a treated course against a similar untouched one over time. Here is how the design works, when its parallel-trends assumption holds, and how to use it honestly in quality assurance.

11 min readadvanced

One Instructor, Two Hundred Defensible Scores: Specification-Curve Analysis for Course Evaluation

Every course-evaluation "score" is the product of dozens of defensible analytic choices — which items count, how to weight, mean or median, whether to adjust for class size. Specification-curve and multiverse analysis compute the result across all of them, so a personnel decision rests on the pattern, not on one analyst's arbitrary path.

11 min readadvanced

The Department Average Is Not the Student: The Ecological Fallacy in Course Evaluation

A correlation that holds between department averages need not hold — and can even reverse — for individual students. The ecological fallacy is inferring individual-level relationships from group-level data. Here is how it distorts course-evaluation analysis and how to reason at the level your decision is actually made.

10 min readintermediate

Your 'Significant' Instructor Difference Might Point the Wrong Way: Type S and Type M Errors in Small-Cohort Evaluation

Statistical significance does not protect a small-cohort evaluation from pointing the wrong way or exaggerating the gap. Type S and Type M error analysis shows why, and what to report instead.

10 min readadvanced

Does the Effect Depend on Who or What? Moderation and Interaction Effects in Course-Evaluation Analysis

A bias or a teaching effect on evaluations often holds only for some students, courses or conditions. Moderation analysis tests those "it depends" claims properly, and it is far harder than it looks.

11 min readadvanced

Is There Really No Difference, or Just No Evidence? Bayes Factors for Course-Evaluation Comparisons

A non-significant p-value cannot confirm two instructors scored the same — it only fails to reject. Bayes factors quantify evidence FOR the null as well as against it, distinguishing absence of evidence from evidence of absence in course-evaluation comparisons.

10 min readadvanced

Build a Synthetic Comparison Department: Synthetic Control Methods for a Teaching Change

When one programme redesigns its teaching, you rarely have a clean control group. Synthetic control methods build a weighted "synthetic" comparator from other programmes so you can estimate whether the change actually moved evaluation scores.

9 min readadvanced

How Strong Would the Hidden Confounder Have to Be? The E-Value for Adjusted Course-Evaluation Comparisons

Every adjusted course-evaluation claim invites the objection "but you did not control for X". The E-value, from epidemiology, quantifies exactly how strong that unmeasured X would have to be to explain away your finding — turning a vague worry into a number.

9 min readadvanced

When a Few Retaliatory 1s Sink the Average: Robust Estimators for Course Evaluation

In a class of twelve, two vindictive 1s can drag the mean by half a point. The arithmetic mean has a breakdown point of zero — a single extreme value can move it arbitrarily far. Trimmed means, Winsorizing, and M-estimators resist that without throwing away respondents.

9 min readadvanced

How Many Dimensions Does Your Evaluation Really Measure? Parallel Analysis for Factor Retention

Deciding how many factors your evaluation instrument measures with the eigenvalue-greater-than-one rule or the scree plot routinely gets the wrong answer. Parallel analysis (Horn, 1965) compares your data against random noise — and is one of the most accurate methods available.

9 min readadvanced

The Average Hides the Tails: Quantile Regression for Course-Evaluation Data

Ordinary regression models the mean, but a teaching change can lift the median while sinking the unhappiest students. Quantile regression (Koenker & Bassett, 1978) models the whole distribution, revealing effects the average conceals.

9 min readadvanced

When Can You Stop Collecting? Optional Stopping and Sequential Analysis for Rolling Course Evaluation

Watching results accumulate in a live course-evaluation dashboard and stopping when a difference looks significant inflates false positives. This guide explains the optional-stopping problem, the sequential-analysis methods that fix it, the evidence, the caveats, and how Koji handles rolling collection responsibly.

9 min readadvanced

Reporting a Class of Four Without Exposing Anyone: Cell Suppression, k-Anonymity, and Differential Privacy

You cannot safely report an average for a class of four students without risking that individuals are identified or their comments inferred. This guide covers minimum-cell suppression, k-anonymity, and differential privacy for course-evaluation reporting — the methods, the evidence, the privacy-utility trade-off, and how Koji protects small-cohort respondents.

9 min readadvanced

One Student, Many Teachers: Cross-Classified and Multiple-Membership Models for Fair Course Evaluation

Standard multilevel models assume a clean hierarchy, but students are taught by several instructors and instructors teach across programmes. Cross-classified and multiple-membership models partition that tangled variance honestly — and change which instructors look unusual.

10 min readadvanced

Where Exactly Do Students Abandon Your Evaluation? Discrete-Time Survival Analysis of Breakoff

A completion rate tells you how many students quit; it cannot tell you where or why. Discrete-time survival analysis models the hazard of breakoff question by question, turning a single number into an actionable map of where your evaluation loses people.

9 min readadvanced

Exactly Balanced Comparison Groups Without the Guesswork: Entropy Balancing for Course Evaluation

Propensity-score matching throws data away and needs you to iterate a model until the groups look balanced. Entropy balancing reweights the data so the groups are exactly balanced on the moments you specify, in one step. Here is what it does for fair course-evaluation comparison.

9 min readadvanced

A High Correlation Does Not Mean Two Evaluation Methods Agree: The Bland-Altman Limits of Agreement

Two evaluation methods can correlate strongly yet disagree by a full scale point. Bland-Altman limits of agreement plot the differences, not the correlation, to show whether student ratings, peer review or AI coding can actually be used interchangeably.

9 min readintermediate

Mapping the Patterns You Cannot Average: Multiple Correspondence Analysis for Categorical Course-Evaluation Data

Much course-evaluation data is genuinely categorical — programme, mode, agree/disagree, chosen theme. Multiple correspondence analysis (MCA) maps how those categories cluster on a two-dimensional plane, revealing response patterns that averaging destroys.

9 min readadvanced

High Agreement, Low Kappa: Choosing an Agreement Coefficient for Coding Course Feedback

Cohen's kappa can collapse to near zero even when two coders agree on 95 percent of comments. Here is why the kappa paradox happens and when to report Gwet's AC1 or Krippendorff's alpha instead.

9 min readintermediate

Did This Student Really Change? The Reliable Change Index for Mid-to-End Course Evaluation

When a student's mid-semester and end-of-semester ratings differ, the Reliable Change Index tells you whether the shift is larger than measurement error alone would produce - a per-individual test borrowed from clinical psychology.

9 min readintermediate

Combining Several Rankings Into One Fair Order: Rank Aggregation for Course Evaluation

When you have to merge several rankings - by different criteria, cohorts or panel members - into one, the method you pick changes the winner. Kemeny, Borda and Condorcet from social-choice theory show why, and how to do it defensibly.

10 min readadvanced

Multidimensional Scaling for Course Evaluation: A Perceptual Map of How Students See Your Courses

Multidimensional scaling turns a table of similarities into a two-dimensional map. Here is how MDS reveals the hidden structure in course-evaluation data that averages and factor analysis both miss.

9 min readadvanced

Count Models for Course Evaluation: Why You Should Not Average the Number of Comments

How many students wrote a comment? How many mentioned assessment? These are counts, and averaging them or running OLS gives biased answers. Poisson, negative-binomial, and zero-inflated models do it right.

9 min readadvanced

Which Aspects of Teaching Actually Drive the Overall Score? Relative Weights and Dominance Analysis

When clarity, workload, support, and feedback are all correlated, regression coefficients cannot tell you which one drives the overall rating. Relative weights and dominance analysis partition the explained variance fairly.

9 min readadvanced

Which Teaching Behaviours Drive the Overall Score — Nonlinearly? Random Forests and Variable Importance for Course Evaluation

Random forests and their variable-importance measures reveal which teaching items predict the overall rating when the relationships are nonlinear and interacting — but Gini importance is biased toward correlated and high-cardinality predictors, so use conditional permutation importance and read the output as association, not cause.

9 min readadvanced

Turning Several Quality Signals Into One Defensible Decision: The Analytic Hierarchy Process for Course Evaluation

The Analytic Hierarchy Process derives criterion weights from pairwise comparisons and checks their internal consistency, giving a transparent, auditable way to combine student ratings, peer observation and learning outcomes into one decision — but it is subjective, vulnerable to rank reversal, and its 1–9 scale and 0.10 consistency threshold are conventions, not laws.

9 min readintermediate

What Will This Instructor Score Next Term — With a Guarantee? Conformal Prediction for Course Evaluation

Conformal prediction wraps any scoring model in a distribution-free prediction interval that is guaranteed to cover the true next value at a rate you choose. Here is what it does, why it fits course evaluation, and how to use it honestly.

9 min readadvanced

Your Outcome Is a Proportion, Not a Mean: Beta Regression for Course Evaluation

When the thing you are modelling is a proportion between 0 and 1 — the share recommending a course, the top-box rate — ordinary regression misbehaves at the boundaries. Beta regression is the purpose-built tool. Here is how it works and when to use it.

9 min readadvanced

The Average Effect Across Everyone, Not Inside One Classroom: GEE for Course Evaluation

When responses are clustered in courses but you want the population-average effect of a change — not the effect for a specific classroom — generalized estimating equations are the right tool, and they are robust to getting the correlation structure wrong.

9 min readadvanced

Relating a Battery of Teaching Behaviours to a Battery of Outcomes: Canonical Correlation Analysis

When you want to relate a whole set of teaching-behaviour items to a whole set of outcome items, running many pairwise correlations misleads. Canonical correlation analysis maps the joint structure.

9 min readadvanced

When Did the Rating Actually Shift? Changepoint Detection for Course-Evaluation Time Series

A rating series drifts and the real question is when it shifted, without knowing the date in advance. Changepoint detection dates unknown structural breaks in term-by-term scores.

9 min readadvanced

Your Comment Categories Sum to 100%: Compositional Data Analysis for Course-Evaluation Shares

When course-evaluation numbers are shares of a fixed whole, ordinary means and correlations mislead. Compositional data analysis and log-ratios fix the closure trap.

9 min readadvanced