You Changed the Questionnaire — Can You Still Compare Years? Test Equating and Linking for Course Evaluation
When you revise a course-evaluation form, scores before and after are not automatically comparable. Kolen and Brennan's equating/linking/prediction hierarchy, and the anchor-item designs behind it, tell you what a cross-revision trend can honestly claim.
Koji Education Team
Product
Every few years a university revises its course-evaluation questionnaire: an item is reworded, two are merged, a five-point scale becomes a seven-point scale, a "would you recommend" question is added. Then someone asks the natural question — "is teaching improving compared with three years ago?" — and pulls a trend line straight across the revision. That trend line is very often wrong, because the numbers on either side of the change were produced by different instruments and are not on the same scale. Psychometrics has a mature toolkit for exactly this problem, developed for standardised testing where scores from different test forms must be used interchangeably. It is called equating, scaling, and linking, and course-evaluation offices are quietly re-inventing a worse version of it every time they change a form.
The short answer (BLUF)
When you revise an evaluation instrument, scores before and after the change are not automatically comparable, and a raw year-over-year trend across the revision can manufacture or hide change that has nothing to do with teaching. Kolen and Brennan's framework distinguishes three levels of comparability: equating (strict interchangeability, only possible when two forms measure the same construct with near-equal difficulty and reliability), scale aligning/linking (weaker, approximate comparability), and prediction (no real comparability at all — you can only forecast one from the other). Most questionnaire revisions qualify for linking at best, not equating. The practical rules: keep a set of unchanged anchor items across the revision, use them to align the scales, and label any pre/post comparison with the strength of linking it actually rests on.
What the research says
Kolen and Brennan's Test Equating, Scaling, and Linking is the standard reference. Its central insight is a hierarchy of comparability claims. Equating is the strongest: two forms are equated only if they measure the same construct, are built to the same content and statistical specifications, and are close in difficulty, so that a score of 4.1 on Form A means the same thing as 4.1 on Form B for any purpose. Scaling to achieve comparability (linking) is weaker — the forms are aligned as well as possible but interchangeability is approximate. Predicting is weakest: you regress new-form scores on old-form scores and accept that the relationship is directional and imperfect. Dorans and Holland (2000) formalised how to evaluate whether a proposed link is strong enough to call equating, introducing criteria such as population invariance — a true equating relationship should hold regardless of which subgroup of respondents you compute it on. If the link between your old and new evaluation forms changes depending on whether you look at first-years or finalists, it is not an equating; it is at best a rough linking.
The mechanics depend on an anchor: items common to both the old and new forms, administered to comparable groups, that act as a bridge. Equating designs include the common-item nonequivalent groups design — precisely the situation of a questionnaire revision, where this year's students take the new form, last year's took the old, and a handful of retained items connect them. The authoritative Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014) devote an entire chapter to scaling, norming, and linking, and warn explicitly against treating linked scores from different instruments as interchangeable without evidence.
Why it matters for course evaluation in practice
The consequences are concrete and, for individual staff, high-stakes.
A reworded item can shift the mean by more than any real change in teaching. Splitting "the course was well organised and clearly explained" into two separate items, or moving from "agree/disagree" to an item-specific scale, systematically changes how students respond. If a department's average moves 0.2 across a revision year, the honest first hypothesis is the instrument, not the teaching.
Scale-length changes break naïve rescaling. Converting a seven-point score to a five-point score by simple linear proportion assumes the two scales are used the same way — they are not. Respondents cluster differently on scales of different lengths, so linear rescaling introduces bias precisely where you are trying to compare.
Trend lines and control charts inherit the error. Any longitudinal method — regression to the mean checks, interrupted time series, changepoint detection — assumes a stable measurement scale. Run one across an un-linked revision and the "changepoint" it detects may simply be the day the form changed.
The disciplined alternative is to plan revisions like a testing organisation: retain a defensible set of anchor items, collect at least one cycle where old and new coexist or where the anchor is present, estimate the linking, and then report post-revision scores on the pre-revision scale (or vice versa) with an explicit note on linking strength.
Limitations and honest caveats
Equating theory was built for long, high-stakes cognitive tests with large samples and carefully engineered parallel forms; course evaluations are short affective questionnaires with modest per-class samples and items that are not designed to be parallel. This mismatch matters. With few anchor items and small samples, linking estimates are themselves noisy, and a poorly chosen anchor (items that were also reworded, or that behave differently after the revision — anchor "drift") can bias the link more than doing nothing. Equating also cannot rescue a revision that changed the construct: if the new form measures something genuinely different — say, it adds inclusive-teaching items the old form lacked — then no statistical bridge makes the totals comparable, and pretending otherwise is worse than admitting the break. Finally, population invariance is an ideal rarely met perfectly; the practical question is whether the departure from invariance is small enough to tolerate for your specific use, which is a judgement, not a formula.
How Koji incorporates this
Koji treats questionnaire change as a measurement event, not a silent edit. Because Koji versions instruments and stores every item response in structured form (scale, single_choice, ranking, yes_no), it retains the raw material needed to build an anchor-based link when a form is revised, rather than discarding the old-scale detail and leaving only incomparable aggregates. When a programme revises its questionnaire, Koji is designed to preserve a set of unchanged anchor items and flag any longitudinal comparison that crosses a revision boundary, so a dashboard does not quietly draw a trend line across two different instruments. Its reporting is built to label comparisons by the strength of comparability they rest on — genuine like-for-like on retained items versus approximate linking on revised ones — mirroring the equating/linking/prediction hierarchy. And because Koji's AI-moderated conversational interview captures why students answer as they do, a revision-year shift can be interrogated qualitatively — did students react to the new wording or to the teaching? — rather than being mistaken for change. The same versioned-instrument discipline underpins Koji's core research platform at koji.so, where tracking studies over time face the identical "did the metric move or did the question change" problem.
As always, this mitigates rather than eliminates the risk: Koji can preserve anchors and flag boundaries, but it cannot equate two forms that no longer measure the same thing, and it will say so rather than fabricate a comparison.
Frequently asked questions
What is the difference between equating and linking?
Equating is strict interchangeability: a score on the new form means exactly what the same score meant on the old form, for any use. Linking is weaker, approximate alignment. Kolen and Brennan reserve "equating" for forms that measure the same construct with near-equal difficulty and reliability; most course-evaluation revisions qualify only for linking, and some only for prediction.
Why can't I just rescale a 7-point score to a 5-point score proportionally?
Because respondents use scales of different lengths differently — the clustering and endpoint avoidance change with the number of points. A linear proportion assumes identical scale use, which is exactly the assumption a scale-length change violates, so it introduces bias rather than removing it.
What is an anchor item and why do I need one?
An anchor item is a question kept identical across the old and new forms. Anchors act as a bridge that lets you estimate how the two scales relate even though different cohorts answered them. Without a stable anchor, a revision leaves no defensible way to place old and new scores on a common scale.
How many items should I keep unchanged when revising a questionnaire?
There is no universal number, but testing practice favours an anchor that is representative of the whole instrument's content and long enough to estimate the link stably given your sample sizes. For short course evaluations, retaining several core items untouched — and confirming they did not "drift" after the revision — is far safer than changing everything at once.
Can equating fix a comparison if the new form measures something new?
No. If the revision changed the construct — for example by adding a dimension the old form never covered — the totals are not measuring the same thing, and no statistical link makes them interchangeable. The honest move is to treat it as a fresh baseline, not to bridge across the break.
How does this affect trend charts and changepoint detection?
Every longitudinal method assumes a stable scale. Run a trend line, interrupted time series, or changepoint detector across an un-linked revision and it may flag the revision itself as the "change." Link the scales first, or annotate the break, before drawing any trend across it.
Related resources
- Generalizability Theory and the Reliability of Student Ratings
- Norm-Referenced vs Criterion-Referenced Course-Evaluation Scores
- Regression to the Mean in Year-over-Year Changes
- Interrupted Time Series for Course-Evaluation Trends
- Changepoint Detection for Course-Evaluation Time Series
- Agree/Disagree or Item-Specific Scales
References
- Kolen, M. J., & Brennan, R. L. (2014). Test Equating, Scaling, and Linking: Methods and Practices (3rd ed.). Springer. https://doi.org/10.1007/978-1-4939-0317-7
- Dorans, N. J., & Holland, P. W. (2000). Population invariance and the equatability of tests: Basic theory and the linear case. Journal of Educational Measurement, 37(4), 281–306. https://doi.org/10.1111/j.1745-3984.2000.tb01088.x
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. AERA.
- Holland, P. W., & Dorans, N. J. (2006). Linking and equating. In R. L. Brennan (Ed.), Educational Measurement (4th ed., pp. 187–220). American Council on Education/Praeger.
Related articles
Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations
Year-over-year movements in course-evaluation scores are mostly statistical noise regressing toward an instructor's true average. Here is why regression to the mean fools quality-assurance committees, and how to read evaluation trends responsibly.
Did the Teaching Change Actually Work? Interrupted Time Series for Course-Evaluation Trends
Comparing this year''s evaluation mean to last year''s cannot tell you whether a curriculum redesign worked. Interrupted time series with segmented regression can — here is how to apply it, and where it breaks.
When Did the Rating Actually Shift? Changepoint Detection for Course-Evaluation Time Series
A rating series drifts and the real question is when it shifted, without knowing the date in advance. Changepoint detection dates unknown structural breaks in term-by-term scores.