The Reliability Number in Your Evaluation Report Is Probably Wrong: Alpha, Omega, and What Justifies a Composite Score
Most course-evaluation reports certify their scale with a single Cronbach's alpha. It is usually the wrong coefficient — and a high value is routinely misread as proof the items measure one thing. Here is what alpha actually assumes, why omega is the better default, and what it means for the averages you report.
Koji Education Team
Product ·
Bottom line up front: If your course-evaluation report certifies the instrument with a single Cronbach''s alpha — "the scale is reliable, α = 0.89" — you are almost certainly reporting the wrong number, and quite possibly reading it backwards. Alpha rests on assumptions that student-evaluation scales rarely meet, so it usually underestimates reliability; and a high alpha is not evidence that your items measure a single construct, even though that is how it is nearly always interpreted. For most evaluation instruments, McDonald''s omega is the more defensible default, and the deeper lesson is that no single coefficient licenses collapsing ten items into one average.
This is a quietly consequential point. The reliability statistic is the credential a questionnaire carries into every committee meeting. It is the sentence that ends the methodological conversation — "the instrument is validated, alpha is high" — and lets everyone move on to the averages. If that credential is the wrong number, the conversation ended too early.
What alpha actually assumes
Cronbach''s alpha (Cronbach, 1951) is not a general-purpose "reliability button." It is a specific estimate that is exactly correct only under a demanding set of conditions: the items must be tau-equivalent (every item measures the latent construct on the same scale, with equal factor loadings), the errors must be uncorrelated, and — critically — the scale must be unidimensional. When these hold, alpha equals reliability. When they do not, alpha is a lower bound at best, and its relationship to true reliability becomes unpredictable.
Klaas Sijtsma put this bluntly in a widely cited 2009 Psychometrika paper, "On the use, the misuse, and the very limited usefulness of Cronbach''s alpha," arguing that alpha is neither a measure of internal consistency in the way people assume nor a good estimate of reliability, and that its near-universal reporting is largely a ritual. Almost a decade later, Daniel McNeish''s 2018 Psychological Methods paper carried the deliberately provocative title "Thanks Coefficient Alpha, We''ll Take It From Here," documenting that alpha is "riddled with problems stemming from unrealistic assumptions" and that violating those assumptions frequently yields reliability estimates that are too small — making a scale look less reliable than it is.
Two facts follow that matter for anyone running a course-evaluation programme.
First, alpha usually understates reliability for real instruments. Tau-equivalence — equal loadings across items — is almost never true. A course-evaluation form asks about clarity, workload, feedback, organisation, and enthusiasm; these items do not load equally on whatever common factor they share. Under the more realistic congeneric model (unequal loadings), alpha is a strict lower bound. So the paradox is that the coefficient people report to reassure themselves is systematically pessimistic. If your alpha is 0.78, your actual reliability may be higher.
Second, and more dangerously, a high alpha is not evidence of unidimensionality. This is the misreading that does real damage. Alpha increases with the number of items and with average inter-item correlation. Add enough items and you can push alpha past 0.9 on a scale that is manifestly measuring two or three distinct things. A high alpha tells you the items hang together on average; it says nothing about whether they hang together because they reflect one construct or because a general "halo" impression bleeds across all of them. As we argued in The Halo Effect, high inter-item correlation on evaluation forms is frequently a symptom of exactly the problem you should be worried about — a single global impression contaminating ten supposedly independent questions — not proof of a clean, well-behaved scale.
Omega, and what it fixes
McDonald''s omega estimates reliability from a factor model rather than assuming tau-equivalence. Coefficient omega (specifically omega-total, or omega-hierarchical when you want to isolate the variance attributable to a general factor) allows items to load unequally and can be computed from the same data you already collect. Because it does not force the false assumption of equal loadings, omega typically gives a more accurate — and often slightly higher — reliability estimate for congeneric scales.
More importantly, moving to a factor-model framing forces the right question. To compute omega, you have to fit a model, which means you have to look at the dimensional structure. Is this a one-factor scale, or does "assessment and feedback" form a cluster distinct from "delivery and clarity"? (It usually does — see Why "Assessment and Feedback" Always Scores Lowest.) Omega-hierarchical then answers the question a composite score actually depends on: how much of the total variance is explained by the single general factor you are about to average across? If that proportion is low, your overall average is a blend of several constructs wearing one number''s clothing — the construct-irrelevant variance problem, dressed as a summary statistic.
None of this is exotic. Omega is available in standard, free software and has been the psychometricians'' recommended default for years. The barrier is not technical; it is institutional inertia. Alpha is what the last report used, so it is what this report uses.
But doesn''t this just make the number look better — isn''t alpha the safe, conservative choice?
This is the strongest objection, and it deserves a direct answer. The reasoning goes: alpha is a lower bound, so if alpha clears your threshold, true reliability is at least that high; reporting the conservative number is prudent. Why chase a higher omega that could look like putting a thumb on the scale?
Two problems. First, "conservative" is only a virtue when the direction of the error is harmless. Underestimating reliability is not harmless: it can sink a genuinely sound formative instrument, or trigger needless "revisions" that add redundant items purely to inflate alpha — which improves the statistic while degrading the questionnaire and lengthening it for students. Second, and decisively, the safety argument addresses only the magnitude of alpha, not its misinterpretation. The real damage is not that alpha is a touch low; it is that a high alpha is paraded as proof the scale measures one thing, justifying a single overall average. On that question alpha is not conservative — it is silent, and its silence is routinely read as assent. Omega, computed through a factor model, at least makes you confront the dimensionality you are averaging over. The conservative move is the one that stops you over-claiming.
What this means for the averages you publish
The practical implications are modest and concrete:
- Report omega, not alpha, as your default — and report it per subscale, not just for the whole form. A single reliability figure for a multidimensional instrument is a category error.
- Stop treating a reliability coefficient as evidence of unidimensionality. If you want to average items, justify it with a dimensionality analysis (a factor model), not with a high alpha.
- Remember what reliability is not. A perfectly reliable scale can be perfectly biased. Reliability is the consistency of the ruler, not the truth of the measurement — the reliability-versus-validity distinction that a good alpha can lull you into forgetting. Consistency in the presence of halo or central-tendency effects is consistency in being wrong the same way every time.
Where Koji fits
The reliability-coefficient debate is a symptom of a deeper design choice: building an evaluation system around a small set of Likert items whose whole job is to be averaged. Koji for Education starts from the other end. Instead of relying on a single composite to carry the meaning, its AI-moderated conversational interviews probe why behind each rating, and its automatic thematic analysis surfaces the distinct dimensions in students'' own words rather than assuming a clean one-factor structure that a coefficient can never actually guarantee. Where structured items are used, Koji supports six question types (open-ended, scale, single- and multiple-choice, ranking, yes/no) and quality-scores responses, so the burden of validity does not rest on one number. That does not exempt anyone from psychometrics — it reduces the stakes of getting a single coefficient wrong, because the composite average is no longer the only thing you have.
The same AI interview engine underpins the main Koji platform for general user and customer research, where the identical problem — a satisfaction average standing in for a multidimensional reality — shows up under different names.
If you take one thing from this piece: the reliability number in your report is a claim, not a certificate. Ask which coefficient it is, what it assumes, and whether it was ever entitled to justify the average printed beside it.