Cronbach's Alpha and Course-Evaluation Reliability: What It Does and Doesn't Tell You
A high Cronbach's alpha on your course-evaluation instrument is widely read as proof of a "reliable" survey. The psychometric literature says that reading is largely wrong. What alpha actually measures, where it misleads, and what to report instead.
Koji Education Team
Product
In brief: Cronbach's alpha is the most-cited "proof of reliability" in course-evaluation instrument validation — and the most misunderstood. Sijtsma (2009) shows alpha is neither a measure of internal consistency nor, in general, an accurate reliability estimate: it is a lower bound that almost always understates true reliability and says nothing about whether items measure one thing. McNeish (2018) and Cortina (1993) reinforce that a high alpha is driven heavily by item count and inter-item correlation, not unidimensionality. For course evaluation, alpha should be reported with its assumptions, never as a stand-alone certificate of quality.
Why this matters before a single number is quoted
Open almost any technical appendix for a teaching-evaluation instrument and you will find a sentence like "the scale demonstrated good reliability (alpha = 0.89)." It reads as a clean bill of psychometric health, and committees treat it as one. But three widely cited methodological papers show that this single number is doing far less work than its prominence implies — and is routinely asked to certify a property (that the items measure one coherent construct) that it cannot certify at all. For an evaluation programme that supplies evidence to accreditation bodies and informs module review, mis-citing alpha is not a pedantic statistical sin; it is a claim about measurement that may not hold.
What the research says
The sharpest critique is Sijtsma (2009), "On the Use, the Misuse, and the Very Limited Usefulness of Cronbach's Alpha," in Psychometrika. Sijtsma makes two points that overturn the everyday interpretation. First, alpha is a lower bound to the true reliability of a test score: under the usual assumptions it is almost always smaller than the actual reliability, sometimes substantially. So a "low" alpha does not necessarily mean an unreliable instrument — it may mean alpha is a loose bound. Second, and more damaging to common practice, alpha is widely reported as a measure of internal consistency or unidimensionality, and Sijtsma shows it is essentially unrelated to the internal structure of the test. You can have a high alpha with a multidimensional set of items and a lower alpha with a unidimensional one. The number tells you almost nothing about whether your "teaching quality" scale measures one thing or several.
Cortina (1993), "What Is Coefficient Alpha?" in the Journal of Applied Psychology — one of the most-cited papers in the field — anticipates this. Cortina demonstrates that alpha is a function of both the average inter-item correlation and the number of items. Add more items, even weakly related ones, and alpha rises mechanically. A long scale can post an impressive alpha while being heterogeneous; a short, genuinely unidimensional scale can look worse. Cortina's explicit conclusion: alpha is not an index of homogeneity or unidimensionality, and should not be read as one.
McNeish (2018), "Thanks Coefficient Alpha, We'll Take It From Here," in Psychological Methods, completes the picture by cataloguing alpha's restrictive assumptions — chiefly tau-equivalence (every item contributes equally to the true score) and uncorrelated errors. When these fail, as they routinely do in real evaluation data, alpha can understate reliability, making an instrument look worse than it is. McNeish recommends modern alternatives — McDonald's omega, the greatest lower bound, and related model-based coefficients — that relax tau-equivalence and typically give more defensible estimates. (The literature is not unanimous: replies such as Raykov and Marcoulides defend alpha under specified conditions, which is itself a reason to report assumptions rather than a bare number.)
The convergent message across all three: a high alpha is necessary-sounding but neither sufficient nor self-explanatory. It is inflated by item count, blind to dimensionality, and usually a conservative bound on the quantity people think it estimates.
Why it matters for course evaluation in practice
For institutional-research and QA staff, four practical consequences follow.
1. Stop equating high alpha with "the survey is good." A teaching-evaluation battery that bundles distinct constructs — organisation, enthusiasm, assessment fairness, workload — can return a high alpha precisely because it is long, while masking that it measures several different things. Reporting alpha alone invites the reader to infer a coherence the data may not have.
2. Test dimensionality separately. Because alpha cannot tell you whether your items are unidimensional, that question must be answered by something else — factor analysis or a clear theoretical structure (as in Marsh's multidimensional SEEQ). If you intend to report a single "overall teaching" score, you owe evidence that a single dimension underlies it; alpha is not that evidence.
3. Do not over-trim items chasing a higher alpha. Because alpha rises with item count, the "alpha if item deleted" ritual can push designers to drop content-valid items that happen to lower the coefficient, narrowing the construct to inflate a statistic. That is optimising the proxy at the expense of the thing it proxies for — a textbook Goodhart problem.
4. Report assumptions and alternatives. Where reliability evidence is required for accreditation, present omega alongside (or instead of) alpha, state the number of items, the average inter-item correlation, and the dimensional structure. A committee that sees "omega = 0.84 on a unidimensional five-item scale, average inter-item r = 0.51" learns far more than one shown "alpha = 0.89."
Limitations and honest caveats
In the interest of the same rigour the topic demands:
- Alpha is not useless. Sijtsma's title is provocative, but his point is "very limited usefulness," not "no usefulness." As a quick, assumption-light lower bound, alpha remains a reasonable first screen — a very low alpha is still a warning. The error is treating it as a final verdict.
- The alternatives have their own assumptions. Omega and the greatest lower bound require a fitted measurement model and adequate sample size; on small classes or short scales they can be unstable or, in the case of the GLB, upward-biased. "Use omega" is not a free lunch.
- Reliability is not validity. Even a correctly estimated reliability coefficient says nothing about whether the instrument measures teaching effectiveness as opposed to student satisfaction or grading leniency. The wider validity debate around student evaluations is logically prior to any reliability statistic.
- Context of estimation. Alpha (and omega) describe a score in a particular sample and administration. A coefficient from one cohort does not automatically license claims about another — reliability is a property of scores in a context, not of the questionnaire in the abstract.
- Live scholarly debate. The alpha-versus-omega question is genuinely contested; some methodologists defend alpha under tau-equivalence. The defensible stance is transparency about assumptions, not allegiance to one coefficient.
How Koji incorporates this
Koji is an AI-native, AI-moderated course-evaluation platform, and its approach to reliability is designed around the limits this literature exposes rather than leaning on a single headline coefficient.
- Reliability reporting with assumptions, not a bare number. Koji's analysis and reporting layer is designed to present reliability evidence in context — alongside item count, dimensional structure, and modern coefficients where appropriate — so that committees and accreditation reviewers are not handed an alpha to over-interpret. The goal is bias-aware, honest reporting rather than a reassuring statistic.
- Triangulation that does not depend on internal consistency at all. The deeper response to alpha's limits is to not rest the case for trustworthy feedback on a single Likert battery. Koji's AI-moderated conversational interview captures open-ended reasoning and probes beyond the number, and its automatic thematic analysis surfaces convergent themes across many students — a form of evidence whose credibility comes from cross-respondent agreement on substance, not from the inter-item correlation of one scale.
- Dimensional awareness in design. Because alpha cannot certify unidimensionality, Koji supports structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) mapped to distinct constructs, so that "organisation," "feedback," and "workload" are collected and reported as the separate dimensions they are rather than blended into one opaque index.
- Quality scoring at the response level. Rather than inferring data quality solely from a scale-level coefficient, Koji scores individual responses for substantiveness and flags low-effort or inattentive answers — addressing the data-quality question that alpha was never designed to answer.
The framing is deliberately careful: Koji is designed to mitigate over-reliance on a single, widely misread reliability statistic by reporting in context and triangulating across evidence types — not to replace psychometrics, which still belongs in any rigorous evaluation. The same AI-moderated interview engine powers product and customer research on the core platform at koji.so, where the temptation to summarise a construct in one number is just as strong.
Related Resources
- Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
- How Many Responses Do You Need for a Reliable Course Evaluation?
- What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and Multidimensional Feedback
- Interpreting and Reporting Student Ratings Responsibly
- Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?
- Can You Fairly Rank Instructors by Their Course-Evaluation Scores?
References
- Sijtsma, K. (2009). On the Use, the Misuse, and the Very Limited Usefulness of Cronbach's Alpha. Psychometrika, 74(1), 107–120. https://doi.org/10.1007/s11336-008-9101-0
- Cortina, J. M. (1993). What Is Coefficient Alpha? An Examination of Theory and Applications. Journal of Applied Psychology, 78(1), 98–104. https://doi.org/10.1037/0021-9010.78.1.98
- McNeish, D. (2018). Thanks Coefficient Alpha, We'll Take It From Here. Psychological Methods, 23(3), 412–433. https://doi.org/10.1037/met0000144
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.