Did This Student Really Change? The Reliable Change Index for Mid-to-End Course Evaluation
When a student's mid-semester and end-of-semester ratings differ, the Reliable Change Index tells you whether the shift is larger than measurement error alone would produce - a per-individual test borrowed from clinical psychology.
Koji Education Team
Product
In brief
Run a mid-semester survey, act on it, run an end-of-semester survey, and a student's overall rating moves from 3 to 4. Did your intervention work, or is that just noise? The Reliable Change Index (RCI) answers this at the level of the individual respondent. Borrowed from clinical psychology, it divides the change score by the standard error of the difference between two measurements; if the result exceeds 1.96, the change is larger than measurement error alone would plausibly produce (p < 0.05, two-tailed). The RCI stops you from celebrating — or mourning — a swing that a perfectly stable student could have produced just by answering on a slightly different day.
What the research says
Jacobson and Truax (1991), "Clinical significance: A statistical approach to defining meaningful change in psychotherapy research" (Journal of Consulting and Clinical Psychology, 59(1), 12-19, doi:10.1037/0022-006X.59.1.12), introduced the RCI to solve a problem identical to formative course evaluation: a group mean can improve while it stays unknowable whether any individual changed reliably. Their index standardises each person's change:
RC = (x_post − x_pre) / S_diff
where S_diff is the standard error of the difference, computed from the measure's reliability and standard deviation: S_diff = √(2 × S_E²), and S_E = SD × √(1 − reliability). An RC above +1.96 or below −1.96 marks a statistically reliable change for that individual.
The S_diff term has a careful history. Christensen and Mendoza (1986), "A method of assessing change in a single subject: An alteration of the RC index" (Behavior Therapy, 17(3), 305-308, doi:10.1016/S0005-7894(86)80060-0), corrected an error in the original 1984 formulation, giving the two-error-term version Jacobson and Truax adopted. Maassen (2004), "The standard error in the Jacobson and Truax Reliable Change Index" (Journal of the International Neuropsychological Society, 10(6), 888-893, doi:10.1017/S1355617704106097), showed that the "classical" S_diff conflates measurement error with practice effects and can give poor estimates, and clarified when a practice-effect correction (the Chelune modification) is warranted. The RCI has proved durable: Zahra and Hedge (2010) and a 2022 Cognitive Behaviour Therapist review, "Reliable change and the reliable change index: still useful after all these years," confirm it remains the standard first step in judging individual change.
Two design points matter for course evaluation:
- Reliability drives everything. A short, noisy instrument has a large S_E, so S_diff is large and almost no change qualifies as reliable. This is the same reliability logic behind how many responses you need and generalizability theory.
- RCI separates reliable change from clinical (practical) significance. Jacobson and Truax paired it with a cutoff for crossing from a "dysfunctional" to a "functional" range; in education the analogue is crossing an acceptability threshold, not merely moving.
Why it matters for course evaluation in practice
Mid-semester feedback is valuable precisely because you can still act on it, and the consultation meta-analyses show it can shift end-of-term ratings. But evaluating whether your specific change worked invites two mistakes the RCI prevents:
- Over-reading noise. A cohort's mean creeps from 3.9 to 4.1 and a committee declares victory. With a typical instrument reliability of 0.8 and SD near 1.0, S_diff is around 0.63, so an individual needs to move roughly 1.2 points before the change is reliable. Most of that 0.2-point cohort drift is students wobbling within measurement error.
- The regression-to-the-mean trap. The lowest-rated section "improves" next cycle even if nothing changed. The RCI does not by itself cure regression to the mean — that needs a comparison group — but by demanding change exceed measurement error it filters out the small rebounds that regression manufactures.
Used well, the RCI lets you report something honest and specific: "of 45 students who gave mid and end ratings, 11 showed reliable improvement, 2 reliable decline, and 32 no reliable change." That is far more defensible than "average went up." It complements the aggregate view from statistical process control, which asks whether the process has shifted, with a per-student question: did this person actually move?
Limitations and honest caveats
- It needs a reliability estimate, and the right one. RCI uses test-retest reliability ideally; plugging in Cronbach's alpha from a single administration mis-estimates S_diff. Get the reliability wrong and every individual verdict is wrong.
- Practice and response-shift effects. Answering the same items twice changes how students use the scale; the mid-point re-anchors their standard. This overlaps with response-shift bias and can masquerade as, or mask, real change. Maassen's practice-effect correction addresses part of it.
- Single-item overall ratings are too coarse. A 1-5 overall item has poor reliability and a step size that swamps the RCI; the index is far more useful on a multi-item scale score.
- Reliable is not meaningful. A reliable one-category move may be pedagogically trivial; pair the RCI with a substantive threshold, exactly as Jacobson and Truax paired it with a functional cutoff.
- Not a causal claim. RCI says the change is bigger than noise, not that you caused it. Attributing it to your intervention still needs a design — a comparison section, a difference-in-differences logic, or randomisation.
- Small classes. With few paired responses, the count of reliable changers is itself noisy; report it with caution and never single out identifiable individuals.
How Koji incorporates this
Koji is built around mid-cycle and end-of-cycle collection on the same cohort, which is exactly the paired design the RCI needs.
- Formative-then-summative on one cohort: Koji supports a mid-semester study and a matched end-of-semester study, producing the paired scores an RCI is computed from — while preserving anonymity thresholds so no individual is exposed.
- Multi-item scale scores, not single numbers: because Koji builds structured, dimension-tagged instruments rather than one overall star, the scale scores have the reliability that makes an RCI usable; noisy single-item change is not mistaken for signal.
- Reliability-aware reporting: rather than trumpeting a two-tenths mean shift, Koji reporting is designed to distinguish movement that exceeds measurement error from movement that does not, keeping "closing the loop" claims honest.
- Conversational probing of genuine movers: when a reliable change is detected in aggregate, the AI-moderated interview explores what actually changed, turning a number into an explanation. Koji's core research platform at koji.so applies the same before-and-after logic to product research, where teams routinely over-read small pre/post survey shifts.
Koji is designed to mitigate the over-interpretation of small pre-to-post changes, not to certify causation — that still requires a proper comparison design.
Frequently asked questions
What exactly does the Reliable Change Index tell me?
It tells you whether one respondent's change between two administrations is larger than what measurement error alone would plausibly produce. An RC value above +1.96 or below −1.96 means the change is statistically reliable at p < 0.05; values in between are within the noise band.
How is the standard error of the difference calculated?
S_diff = √(2 × S_E²), where S_E = SD × √(1 − reliability). It combines the measurement error of both the pre and post scores, following the Christensen and Mendoza (1986) correction to the original formula.
Why can't I just compare the two means?
Because a mean can drift while no individual changed reliably, and small mean shifts are usually within measurement error. The RCI works at the individual level and explicitly accounts for the instrument's reliability, so it separates real movers from noise.
Does a reliable change prove my intervention worked?
No. The RCI shows the change exceeds measurement error, not that you caused it. Attributing the change to your action needs a comparison group or randomisation, because regression to the mean and practice effects can move scores on their own.
Can I use the RCI on a single overall rating?
It is far weaker there. A single 1-5 item has low reliability and a coarse step size, so S_diff is large and the index rarely detects change. Use it on a multi-item scale score.
What reliability figure should I plug in?
Ideally a test-retest reliability for the same instrument and interval. Using a single-administration alpha misstates the standard error. If you only have alpha, treat the results as approximate and say so.
Related resources
- Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations
- Statistical Process Control for Course Evaluations: Using Control Charts to Separate Signal from Noise
- Mid-Semester Feedback and the Power of Consultation: What the Meta-Analyses Show
- Response-Shift Bias: Why Self-Reported Learning Gains in Course Evaluations Can Mislead
- The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
- How Many Responses Do You Need for a Reliable Course Evaluation?
References
- Jacobson, N. S., & Truax, P. (1991). Clinical significance: A statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59(1), 12-19. doi:10.1037/0022-006X.59.1.12
- Christensen, L., & Mendoza, J. L. (1986). A method of assessing change in a single subject: An alteration of the RC index. Behavior Therapy, 17(3), 305-308. doi:10.1016/S0005-7894(86)80060-0
- Maassen, G. H. (2004). The standard error in the Jacobson and Truax Reliable Change Index: The classical approach to the assessment of reliable change. Journal of the International Neuropsychological Society, 10(6), 888-893. doi:10.1017/S1355617704106097
- Zahra, D., & Hedge, C. (2010). The reliable change index: Why isn't it more popular in academic psychology? PsyPAG Quarterly, 76, 14-19.
- Evans, C., Margison, F., & Barkham, M. (1998). The contribution of reliable and clinically significant change methods to evidence-based mental health. Evidence-Based Mental Health, 1(3), 70-72. doi:10.1136/ebmh.1.3.70
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations
Year-over-year movements in course-evaluation scores are mostly statistical noise regressing toward an instructor's true average. Here is why regression to the mean fools quality-assurance committees, and how to read evaluation trends responsibly.
Statistical Process Control for Course Evaluations: Using Control Charts to Separate Signal from Noise
How Shewhart control charts and statistical process control distinguish real changes in teaching quality from ordinary random variation in course evaluation scores - so QA teams stop reacting to noise. Grounded in Sivena and Nikolaidis (2019, 2022).