Is the Weak SET-Learning Correlation an Artefact of Unreliable Measures? What Correction for Attenuation Does — and Doesn't — Prove
Disattenuation lets you estimate what a correlation would be if both measures were perfectly reliable. It is a legitimate tool that both sides of the student-ratings debate have used — and abused — to argue the true SET-learning link is stronger or weaker than the raw number suggests.
Koji Education Team
Product
The short answer
Correction for attenuation ("disattenuation") estimates what the correlation between two variables would be if both were measured without error, by dividing the observed correlation by the square root of the product of the two reliabilities. It is a century-old, entirely legitimate psychometric technique — and it sits at the heart of the fight over whether student evaluations of teaching (SET) measure learning. Because unreliable measures attenuate (shrink) correlations toward zero, a defender of SET can argue the true link is stronger than the modest raw correlation implies. But disattenuation is fragile: applied to small samples or shaky reliability estimates it produces unstable, over-optimistic numbers — and can even yield "corrected" correlations above 1.0, a mathematical impossibility that signals the method has been pushed past its limits.
BLUF for busy readers: Observed correlations are dragged toward zero by measurement error. The disattenuation formula, r_true = r_observed / √(r_xx · r_yy), estimates the error-free correlation. Use it to understand why a validity coefficient looks weak, not to manufacture a strong one. Corrected values above 1.0, or corrections resting on tiny samples, are warnings — not evidence.
What the research says
The technique dates to Spearman (1904), "The proof and measurement of association between two things," in the American Journal of Psychology. Spearman's insight was that any observed correlation between two fallible measures understates the correlation between the true scores they imperfectly capture, and that the degree of understatement is governed by each measure's reliability. His correction formula recovers an estimate of that true-score correlation.
A century of use has made the technique both indispensable and controversial. Muchinsky (1996), writing in Educational and Psychological Measurement, reviews the correction and the debates around it: the difference between double correction (disattenuating for unreliability in both variables) and single correction (only one), the sensitivity of the result to the reliability estimates plugged in, and the notorious problem that the correction can produce validity coefficients greater than 1.00 when the reliabilities are underestimated or the sample is small. His verdict is nuanced: the correction is valuable — it is embedded in meta-analysis and validity-generalisation work — but it demands honest, well-estimated reliabilities and a clear statement of which correlation (observed or corrected) is being claimed.
The stakes for course evaluation are concrete. Cohen's (1981) landmark meta-analysis of multisection-validity studies reported a substantial correlation between student ratings and student achievement — a figure often cited as evidence that SET is valid. But part of that apparent strength came from correcting small-sample correlations for attenuation and other artefacts. Decades later, Uttl, White, and Gonzalez (2017) in Studies in Educational Evaluation re-analysed the multisection literature and reached the opposite conclusion: once you account for small sample sizes and stop over-crediting artefact corrections, "student evaluation of teaching ratings and student learning are not related." Their critique is, in large part, a critique of how attenuation and small-sample corrections were applied — a reminder that disattenuation is only as trustworthy as the reliabilities and sample sizes behind it.
Why it matters for course evaluation in practice
Disattenuation is not an abstraction; it changes what your validity evidence appears to say:
- Interpreting SET-learning validity claims. When a report cites a correlation between ratings and learning, the first question is: raw or corrected? A disattenuated coefficient can look impressively high while the observed relationship a student or dean actually experiences is modest. Both numbers are legitimate, but they answer different questions — "how well do the true constructs relate?" versus "how well do our actual, fallible instruments track each other?"
- Correlating a rating scale with an outcome. If you correlate an end-of-term teaching-quality scale (with, say, a Cronbach's alpha of 0.80) against a noisy learning measure like a short quiz (reliability 0.60), the raw correlation is suppressed by both unreliabilities. Disattenuation shows the ceiling the relationship could reach with perfect instruments — useful for deciding whether a weak result reflects a weak link or merely weak measures.
- Meta-analysis and cross-programme synthesis. Institutions increasingly pool evaluation evidence across courses and cohorts. Attenuation corrections are standard in meta-analysis, and getting them wrong — as the Cohen/Uttl exchange shows — can flip a headline conclusion.
- Not over-claiming from unreliable open text. If you quantify themes from a handful of open comments with modest inter-rater reliability, correlations with other measures are heavily attenuated; but "correcting" them on a tiny sample is exactly the abuse Uttl and colleagues warned about.
Limitations and honest caveats
This is a technique that punishes overreach, and a PhD reader will insist on every caveat:
- It amplifies error in the reliabilities, not just the correlation. You are dividing by a product of estimated reliabilities. If those estimates are too low (which happens routinely with short scales and small samples), the corrected correlation balloons — sometimes past 1.0. A corrected value above 1.0 is not a strong finding; it is proof the inputs were off.
- Small samples make it wildly unstable. The multisection-validity literature is full of studies with a handful of sections. Correcting a correlation estimated from n = 5 sections yields a number with enormous, usually unreported, uncertainty. This is the crux of the Uttl et al. critique of Cohen.
- Reliability is not one thing. Whether you use Cronbach's alpha, coefficient omega, test–retest, or a generalizability-theory coefficient changes the correction. Using an internal-consistency coefficient when the relevant error is between-occasions can badly misstate the true reliability, and hence the correction.
- The corrected correlation is hypothetical. It describes a world with perfect instruments that does not exist. Decisions are made with the fallible instruments you actually have, so the observed correlation is what governs real predictive use. Disattenuation informs measurement strategy; it does not license acting as if the error were gone.
- It corrects for random error only. Systematic biases — a warmth halo, leniency, reference bias — are not measurement noise and are not fixed by disattenuation. Correcting for attenuation on a biased measure gives you a more precise estimate of a biased thing.
The honest framing: disattenuation is a legitimate way to separate "the measures are weak" from "the relationship is weak," but it is not a lever for making inconvenient correlations look bigger. The Cohen-versus-Uttl history is a cautionary tale about exactly that temptation.
How Koji incorporates this
Koji's stance is that reliability should be measured and reported, not assumed — which is the precondition for using (or scrutinising) any attenuation correction responsibly:
- Reliability estimates as a standard output. Koji's structured
scaleitems produce the multi-item data needed to compute internal-consistency reliability (alpha and omega) for each scale, so an analyst has the honest reliability figures that a defensible correction requires — rather than plugging in a guessed or borrowed coefficient. - Surfacing sample size next to every coefficient. Because the platform tracks how many responses underlie each estimate, it can flag when a correlation (and any correction applied to it) rests on too few cases to be stable — the specific failure mode behind inflated disattenuated values. Making n visible is a direct guard against the small-sample abuse the Uttl reanalysis exposed.
- Reducing the error that causes attenuation in the first place. The deeper fix for attenuation is not to correct for unreliability but to reduce it. Koji's AI-moderated conversational interviews are designed to elicit clearer, more considered responses than a rushed Likert grid, and automatic thematic analysis of open text applies a consistent coding frame — both aimed at raising the reliability of the underlying measures so the raw correlation is less attenuated to begin with. Koji frames this as reducing measurement error, not eliminating it.
- Keeping raw and corrected estimates distinct. Koji's reporting philosophy is to present the observed relationship a decision-maker actually faces, and to treat any disattenuated figure as a clearly-labelled measurement diagnostic — never letting a hypothetical error-free coefficient masquerade as the operational one.
Koji does not, and should not, silently disattenuate correlations behind a dashboard number; doing so is exactly how the SET-learning debate went wrong. Its contribution is to make reliability and sample size transparent so that whoever does apply a correction does it with honest inputs. Koji's core research platform at koji.so brings the same reliability-aware measurement discipline to product and customer research, where noisy self-report measures attenuate correlations just as severely.
References
- Spearman, C. (1904). The proof and measurement of association between two things. American Journal of Psychology, 15(1), 72–101. https://www.jstor.org/stable/1412159
- Muchinsky, P. M. (1996). The correction for attenuation. Educational and Psychological Measurement, 56(1), 63–75. https://doi.org/10.1177/0013164496056001004
- Cohen, P. A. (1981). Student ratings of instruction and student achievement: A meta-analysis of multisection validity studies. Review of Educational Research, 51(3), 281–309. https://doi.org/10.3102/00346543051003281
- Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
Frequently asked questions
What does correction for attenuation actually do? It estimates what a correlation between two variables would be if each were measured with perfect reliability, by dividing the observed correlation by the square root of the product of the two reliabilities. Because measurement error drags correlations toward zero, the corrected value is always larger than the observed one.
Why can a corrected correlation exceed 1.0? Because the formula divides by estimated reliabilities. If those estimates are too low — common with short scales and small samples — the division inflates the result past 1.0, which is mathematically impossible for a real correlation. A value above 1.0 is a signal that the reliability inputs or the sample were inadequate, not a strong finding.
Does disattenuation prove student evaluations measure learning? No. Cohen's 1981 meta-analysis reported a strong corrected correlation, but Uttl and colleagues (2017) showed much of that strength came from over-crediting small-sample and attenuation corrections; their reanalysis found ratings and learning are essentially unrelated. Disattenuation clarifies the role of measurement error; it does not settle validity on its own.
Should I report the raw or the corrected correlation? Report both, clearly labelled. The observed correlation is what your fallible instruments actually deliver and what governs real decisions. The corrected correlation is a measurement diagnostic showing the ceiling with perfect instruments. Presenting only the corrected figure overstates the operational relationship.
Does Koji correct correlations for attenuation automatically? No — silently disattenuating a dashboard number is exactly how the SET-learning debate went astray. Koji instead computes honest reliability estimates and keeps sample sizes visible, so any correction an analyst applies uses trustworthy inputs, and it works to reduce measurement error at the source so the raw correlation is less attenuated.
Related resources
- Cronbach's Alpha and Course-Evaluation Reliability: What It Does and Doesn't Tell You
- Coefficient Omega vs Cronbach's Alpha: Reporting the Reliability of a Course-Evaluation Scale
- Do Student Ratings Track Real Learning? What Cohen's 1981 Multisection Meta-Analysis Found
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
- Generalizability Theory and the Reliability of Student Ratings
- Why "I Learned a Lot" Can't Be Compared Across Courses: Reference Bias
Related articles
Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
Cronbach's Alpha and Course-Evaluation Reliability: What It Does and Doesn't Tell You
A high Cronbach's alpha on your course-evaluation instrument is widely read as proof of a "reliable" survey. The psychometric literature says that reading is largely wrong. What alpha actually measures, where it misleads, and what to report instead.
Coefficient Omega vs Cronbach's Alpha: Reporting the Reliability of a Course-Evaluation Scale
Cronbach's alpha assumes every item measures the construct equally well — an assumption course-evaluation subscales rarely meet. Here is why McDonald's omega is the more defensible reliability coefficient, and how to report it.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.