New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

Did This Student Really Change? The Reliable Change Index for Mid-to-End Course Evaluation

When a student's mid-semester and end-of-semester ratings differ, the Reliable Change Index tells you whether the shift is larger than measurement error alone would produce - a per-individual test borrowed from clinical psychology.

Koji Education Team

Product

In brief

Run a mid-semester survey, act on it, run an end-of-semester survey, and a student's overall rating moves from 3 to 4. Did your intervention work, or is that just noise? The Reliable Change Index (RCI) answers this at the level of the individual respondent. Borrowed from clinical psychology, it divides the change score by the standard error of the difference between two measurements; if the result exceeds 1.96, the change is larger than measurement error alone would plausibly produce (p < 0.05, two-tailed). The RCI stops you from celebrating — or mourning — a swing that a perfectly stable student could have produced just by answering on a slightly different day.

What the research says

Jacobson and Truax (1991), "Clinical significance: A statistical approach to defining meaningful change in psychotherapy research" (Journal of Consulting and Clinical Psychology, 59(1), 12-19, doi:10.1037/0022-006X.59.1.12), introduced the RCI to solve a problem identical to formative course evaluation: a group mean can improve while it stays unknowable whether any individual changed reliably. Their index standardises each person's change:

RC = (x_post − x_pre) / S_diff

where S_diff is the standard error of the difference, computed from the measure's reliability and standard deviation: S_diff = √(2 × S_E²), and S_E = SD × √(1 − reliability). An RC above +1.96 or below −1.96 marks a statistically reliable change for that individual.

The S_diff term has a careful history. Christensen and Mendoza (1986), "A method of assessing change in a single subject: An alteration of the RC index" (Behavior Therapy, 17(3), 305-308, doi:10.1016/S0005-7894(86)80060-0), corrected an error in the original 1984 formulation, giving the two-error-term version Jacobson and Truax adopted. Maassen (2004), "The standard error in the Jacobson and Truax Reliable Change Index" (Journal of the International Neuropsychological Society, 10(6), 888-893, doi:10.1017/S1355617704106097), showed that the "classical" S_diff conflates measurement error with practice effects and can give poor estimates, and clarified when a practice-effect correction (the Chelune modification) is warranted. The RCI has proved durable: Zahra and Hedge (2010) and a 2022 Cognitive Behaviour Therapist review, "Reliable change and the reliable change index: still useful after all these years," confirm it remains the standard first step in judging individual change.

Two design points matter for course evaluation:

  • Reliability drives everything. A short, noisy instrument has a large S_E, so S_diff is large and almost no change qualifies as reliable. This is the same reliability logic behind how many responses you need and generalizability theory.
  • RCI separates reliable change from clinical (practical) significance. Jacobson and Truax paired it with a cutoff for crossing from a "dysfunctional" to a "functional" range; in education the analogue is crossing an acceptability threshold, not merely moving.

Why it matters for course evaluation in practice

Mid-semester feedback is valuable precisely because you can still act on it, and the consultation meta-analyses show it can shift end-of-term ratings. But evaluating whether your specific change worked invites two mistakes the RCI prevents:

  • Over-reading noise. A cohort's mean creeps from 3.9 to 4.1 and a committee declares victory. With a typical instrument reliability of 0.8 and SD near 1.0, S_diff is around 0.63, so an individual needs to move roughly 1.2 points before the change is reliable. Most of that 0.2-point cohort drift is students wobbling within measurement error.
  • The regression-to-the-mean trap. The lowest-rated section "improves" next cycle even if nothing changed. The RCI does not by itself cure regression to the mean — that needs a comparison group — but by demanding change exceed measurement error it filters out the small rebounds that regression manufactures.

Used well, the RCI lets you report something honest and specific: "of 45 students who gave mid and end ratings, 11 showed reliable improvement, 2 reliable decline, and 32 no reliable change." That is far more defensible than "average went up." It complements the aggregate view from statistical process control, which asks whether the process has shifted, with a per-student question: did this person actually move?

Limitations and honest caveats

  • It needs a reliability estimate, and the right one. RCI uses test-retest reliability ideally; plugging in Cronbach's alpha from a single administration mis-estimates S_diff. Get the reliability wrong and every individual verdict is wrong.
  • Practice and response-shift effects. Answering the same items twice changes how students use the scale; the mid-point re-anchors their standard. This overlaps with response-shift bias and can masquerade as, or mask, real change. Maassen's practice-effect correction addresses part of it.
  • Single-item overall ratings are too coarse. A 1-5 overall item has poor reliability and a step size that swamps the RCI; the index is far more useful on a multi-item scale score.
  • Reliable is not meaningful. A reliable one-category move may be pedagogically trivial; pair the RCI with a substantive threshold, exactly as Jacobson and Truax paired it with a functional cutoff.
  • Not a causal claim. RCI says the change is bigger than noise, not that you caused it. Attributing it to your intervention still needs a design — a comparison section, a difference-in-differences logic, or randomisation.
  • Small classes. With few paired responses, the count of reliable changers is itself noisy; report it with caution and never single out identifiable individuals.

How Koji incorporates this

Koji is built around mid-cycle and end-of-cycle collection on the same cohort, which is exactly the paired design the RCI needs.

  • Formative-then-summative on one cohort: Koji supports a mid-semester study and a matched end-of-semester study, producing the paired scores an RCI is computed from — while preserving anonymity thresholds so no individual is exposed.
  • Multi-item scale scores, not single numbers: because Koji builds structured, dimension-tagged instruments rather than one overall star, the scale scores have the reliability that makes an RCI usable; noisy single-item change is not mistaken for signal.
  • Reliability-aware reporting: rather than trumpeting a two-tenths mean shift, Koji reporting is designed to distinguish movement that exceeds measurement error from movement that does not, keeping "closing the loop" claims honest.
  • Conversational probing of genuine movers: when a reliable change is detected in aggregate, the AI-moderated interview explores what actually changed, turning a number into an explanation. Koji's core research platform at koji.so applies the same before-and-after logic to product research, where teams routinely over-read small pre/post survey shifts.

Koji is designed to mitigate the over-interpretation of small pre-to-post changes, not to certify causation — that still requires a proper comparison design.

Frequently asked questions

What exactly does the Reliable Change Index tell me?

It tells you whether one respondent's change between two administrations is larger than what measurement error alone would plausibly produce. An RC value above +1.96 or below −1.96 means the change is statistically reliable at p < 0.05; values in between are within the noise band.

How is the standard error of the difference calculated?

S_diff = √(2 × S_E²), where S_E = SD × √(1 − reliability). It combines the measurement error of both the pre and post scores, following the Christensen and Mendoza (1986) correction to the original formula.

Why can't I just compare the two means?

Because a mean can drift while no individual changed reliably, and small mean shifts are usually within measurement error. The RCI works at the individual level and explicitly accounts for the instrument's reliability, so it separates real movers from noise.

Does a reliable change prove my intervention worked?

No. The RCI shows the change exceeds measurement error, not that you caused it. Attributing the change to your action needs a comparison group or randomisation, because regression to the mean and practice effects can move scores on their own.

Can I use the RCI on a single overall rating?

It is far weaker there. A single 1-5 item has low reliability and a coarse step size, so S_diff is large and the index rarely detects change. Use it on a multi-item scale score.

What reliability figure should I plug in?

Ideally a test-retest reliability for the same instrument and interval. Using a single-administration alpha misstates the standard error. If you only have alpha, treat the results as approximate and say so.

Related resources

References

  • Jacobson, N. S., & Truax, P. (1991). Clinical significance: A statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59(1), 12-19. doi:10.1037/0022-006X.59.1.12
  • Christensen, L., & Mendoza, J. L. (1986). A method of assessing change in a single subject: An alteration of the RC index. Behavior Therapy, 17(3), 305-308. doi:10.1016/S0005-7894(86)80060-0
  • Maassen, G. H. (2004). The standard error in the Jacobson and Truax Reliable Change Index: The classical approach to the assessment of reliable change. Journal of the International Neuropsychological Society, 10(6), 888-893. doi:10.1017/S1355617704106097
  • Zahra, D., & Hedge, C. (2010). The reliable change index: Why isn't it more popular in academic psychology? PsyPAG Quarterly, 76, 14-19.
  • Evans, C., Margison, F., & Barkham, M. (1998). The contribution of reliable and clinically significant change methods to evidence-based mental health. Evidence-Based Mental Health, 1(3), 70-72. doi:10.1136/ebmh.1.3.70

Related articles