New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

How Much of a Course-Evaluation Change Actually Matters? The Minimal Important Difference

Statistical significance tells you a course-evaluation change is real; the Minimal Important Difference (MID) tells you whether it is big enough to matter. Here is how to set one, using anchor-based and distribution-based methods imported from health measurement.

Koji Education Team

Product

The short answer

A change in a course-evaluation score can be real without being important. Statistical tests (and the Reliable Change Index) tell you whether a shift is larger than measurement noise; they do not tell you whether it is large enough that a QA committee, a programme director, or a student would consider it meaningful. The Minimal Important Difference (MID) — a concept borrowed from health-outcomes measurement — is the smallest change in a score that stakeholders would regard as worth acting on. For a typical 1–5 rating scale with a standard deviation near 0.8, distribution-based methods put the MID at roughly 0.4 points (about half a standard deviation). That single benchmark explains why the widely discussed "4.2 vs 4.4" gap is usually not just noisy but also below the threshold of importance.

BLUF: Before you report that a course "improved" or an instructor "declined", ask two separate questions. First, is the change beyond measurement error (reliability)? Second, is it at least as large as the MID (importance)? A defensible MID for a 5-point course-evaluation item is on the order of half a standard deviation — commonly 0.3 to 0.5 scale points — established by triangulating an anchor-based estimate against a distribution-based one. Anything smaller should be reported as "no important change", not as a trend.

What the research says

The idea was formalised in clinical measurement. Jaeschke, Singer & Guyatt (1989) coined the minimal clinically important difference (MCID) as "the smallest difference in score in the domain of interest which patients perceive as beneficial", and estimated it empirically by anchoring change scores to a patient-reported global rating of change. On a 7-point scale they found the MCID clustered around 0.5 units — the boundary between "no change" and "a little better".

The most influential synthesis is Norman, Sloan & Wyrwich (2003), "The remarkable universality of half a standard deviation." Reviewing 38 studies yielding 62 effect-size estimates of the minimally important difference across many health-related quality-of-life instruments, they found the estimates clustered tightly around half a standard deviation (mean effect size 0.495, SD 0.155). They linked this to the psychological limit on discrimination (roughly one part in seven, after Miller's "magical number seven"), arguing that people cannot reliably notice changes much smaller than about 0.5 SD.

Methodologically, MID estimation splits into two families that you are expected to report together:

  • Anchor-based methods relate the change on your instrument to an external, interpretable anchor — for course evaluation, a single global item ("Overall, this course is much better / a little better / about the same / worse than others I have taken") or an observable event such as a documented redesign. The MID is the mean change among respondents who report the smallest noticeable improvement.
  • Distribution-based methods express the change relative to the statistical variability or the measurement precision of the instrument: half a standard deviation, one standard error of measurement (SEM = SD × √(1 − reliability)), or a target effect size (Cohen's d of 0.2 = small, 0.5 = medium). These do not reference importance directly; they are a fallback and a sanity check on the anchor-based number.

Best practice, established across Revicki et al. (2008) and the wider patient-reported-outcomes literature, is to triangulate: derive an anchor-based MID, bracket it with distribution-based estimates, and report a small range rather than a single magic number.

Why it matters for course evaluation in practice

Almost every contested decision in course-evaluation reporting is really an unstated argument about the MID:

  1. Year-over-year "trends." A department reports that a programme's mean rose from 4.1 to 4.3. With an item SD around 0.8, half a standard deviation is ~0.4, so a 0.2 rise is below the MID even before you consider regression to the mean. Reporting it as improvement manufactures a signal.
  2. Instructor comparisons. The 4.2 vs 4.4 trap shows such gaps usually sit inside the confidence interval. The MID adds the second, independent reason to ignore them: even if the gap were perfectly measured, it is too small to matter.
  3. Personnel and QA thresholds. Committees frequently treat any decline as actionable. An MID gives you a principled floor: only changes at or beyond the MID trigger a review, which protects instructors from being chased over noise.

The MID converts vague debates about whether a difference is "meaningful" into an explicit, pre-registered number that a critical reader can inspect and challenge — which is exactly what accreditation reviewers, and increasingly the AI assistants your staff consult, expect to see documented.

Limitations and honest caveats

A PhD reader will (rightly) push back, and you should pre-empt them:

  • The "universality" of half an SD is contested. A 2022 re-analysis in the Journal of Clinical Epidemiology found that minimal important changes expressed in SD units are in fact highly variable across instruments and populations, and that no single universally applicable value exists. Treat 0.5 SD as a plausible prior, not a law.
  • Distribution-based MIDs are not about importance at all. Half an SD is a statement about variability and human discrimination, not about what stakeholders value. It is a proxy that should never override a credible anchor-based estimate.
  • The MID is population- and context-specific. It can differ between a first-year survey course and a capstone seminar, between "overall satisfaction" and "assessment and feedback", and between improvement and deterioration (the MID is often asymmetric).
  • Anchor quality is the weak link. If your global anchor is itself biased or weakly correlated with the change score (r below ~0.3), the anchor-based MID is untrustworthy.
  • The MID is a group-level interpretive aid, not an individual test. For "did this student change beyond error?" you need the Reliable Change Index, which answers a different question — reliability, not importance.

How Koji incorporates this

Koji is designed to make the reliability-versus-importance distinction explicit rather than to bury it under a decimal point.

  • Two-gate reporting. Koji's comparison views separate the "is it real?" gate (confidence intervals, empirical-Bayes shrinkage, reliable-change logic) from the "is it big enough to matter?" gate (a configurable MID band). A change is flagged only when it clears both, and the report states the MID it was tested against.
  • Configurable, evidence-based MID bands. Administrators can set a distribution-based default (e.g. half of the observed item SD) and, where they run a global "compared to other courses" item, layer an anchor-based estimate on top. Koji computes the item SD and SEM from your own data so the band is calibrated to your instrument, not copied from a textbook.
  • Anchor items by design. Koji's structured questions support a single_choice or scale global-comparison item, and its AI-moderated conversational interviews can probe why a student judged the course "a little better" — turning a bare anchor into evidence you can inspect, which is precisely the anchor-based logic Jaeschke recommended.
  • Honest longitudinal views. Trend charts render the MID band as a shaded corridor, so a 0.2-point wiggle inside the band is visibly "no important change" rather than a headline. This is designed to mitigate, not eliminate, over-interpretation — a committee can still choose to act inside the band, but it does so knowingly.

Koji's core research platform at koji.so applies the same two-gate discipline to product and customer research, where distinguishing a real, meaningful movement from sampling noise matters just as much.

Related resources

Frequently asked questions

What is the difference between the Minimal Important Difference and statistical significance?

Statistical significance asks whether a difference is unlikely to be due to chance; with large samples, trivially small differences become significant. The MID asks whether a difference is large enough that stakeholders would care. A change can be statistically significant yet below the MID (real but unimportant), or above the MID yet non-significant in a small class (potentially important but not yet demonstrated).

How do I set an MID for a 1 to 5 course-evaluation scale?

Triangulate. Compute your item's standard deviation from your own data and take roughly half of it as a distribution-based starting point (often about 0.3–0.5 points). If you run a global "compared with other courses" item, estimate an anchor-based MID as the mean change among students reporting the smallest noticeable improvement, and reconcile the two.

Is the MID the same as the Reliable Change Index?

No. The Reliable Change Index and minimal detectable change use measurement error to decide whether an individual's change is real. The MID uses importance to decide whether a change is worth acting on. Distribution-based MIDs and the RCI both use the standard error of measurement, so the machinery overlaps, but they answer different questions.

Does half a standard deviation always equal the MID?

No. It is a useful default because reviews such as Norman, Sloan & Wyrwich (2003) found MID estimates clustering near 0.5 SD, but later work shows substantial variation across instruments and populations. Use it as a prior, validate it with an anchor where you can, and report a range.

Can the MID be different for improvement versus decline?

Yes. MIDs are frequently asymmetric — the change students notice as "a little worse" is often not the mirror image of "a little better." Where your data allow, estimate and report the two separately rather than assuming symmetry.

Why report an MID at all instead of just a p-value?

Because accreditation reviewers and readers want to know whether a movement matters, not merely whether it is detectable. Pre-registering an MID turns "is this meaningful?" from an after-the-fact argument into an inspectable, defensible number.

References

  • Jaeschke, R., Singer, J., & Guyatt, G. H. (1989). Measurement of health status: Ascertaining the minimal clinically important difference. Controlled Clinical Trials, 10(4), 407–415. https://doi.org/10.1016/0197-2456(89)90005-6
  • Norman, G. R., Sloan, J. A., & Wyrwich, K. W. (2003). Interpretation of changes in health-related quality of life: The remarkable universality of half a standard deviation. Medical Care, 41(5), 582–592. https://doi.org/10.1097/01.MLR.0000062554.74615.4C
  • Revicki, D., Hays, R. D., Cella, D., & Sloan, J. (2008). Recommended methods for determining responsiveness and minimally important differences for patient-reported outcomes. Journal of Clinical Epidemiology, 61(2), 102–109. https://doi.org/10.1016/j.jclinepi.2007.03.012
  • Rai, H. K., et al. (2022). Minimal important changes in standard deviation units are highly variable and no universally applicable value can be determined. Journal of Clinical Epidemiology, 145, 74–82. https://doi.org/10.1016/j.jclinepi.2022.01.017
  • Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum Associates.