New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

You Changed the Questionnaire — Can You Still Compare Years? Test Equating and Linking for Course Evaluation

When you revise a course-evaluation form, scores before and after are not automatically comparable. Kolen and Brennan's equating/linking/prediction hierarchy, and the anchor-item designs behind it, tell you what a cross-revision trend can honestly claim.

Koji Education Team

Product

Every few years a university revises its course-evaluation questionnaire: an item is reworded, two are merged, a five-point scale becomes a seven-point scale, a "would you recommend" question is added. Then someone asks the natural question — "is teaching improving compared with three years ago?" — and pulls a trend line straight across the revision. That trend line is very often wrong, because the numbers on either side of the change were produced by different instruments and are not on the same scale. Psychometrics has a mature toolkit for exactly this problem, developed for standardised testing where scores from different test forms must be used interchangeably. It is called equating, scaling, and linking, and course-evaluation offices are quietly re-inventing a worse version of it every time they change a form.

The short answer (BLUF)

When you revise an evaluation instrument, scores before and after the change are not automatically comparable, and a raw year-over-year trend across the revision can manufacture or hide change that has nothing to do with teaching. Kolen and Brennan's framework distinguishes three levels of comparability: equating (strict interchangeability, only possible when two forms measure the same construct with near-equal difficulty and reliability), scale aligning/linking (weaker, approximate comparability), and prediction (no real comparability at all — you can only forecast one from the other). Most questionnaire revisions qualify for linking at best, not equating. The practical rules: keep a set of unchanged anchor items across the revision, use them to align the scales, and label any pre/post comparison with the strength of linking it actually rests on.

What the research says

Kolen and Brennan's Test Equating, Scaling, and Linking is the standard reference. Its central insight is a hierarchy of comparability claims. Equating is the strongest: two forms are equated only if they measure the same construct, are built to the same content and statistical specifications, and are close in difficulty, so that a score of 4.1 on Form A means the same thing as 4.1 on Form B for any purpose. Scaling to achieve comparability (linking) is weaker — the forms are aligned as well as possible but interchangeability is approximate. Predicting is weakest: you regress new-form scores on old-form scores and accept that the relationship is directional and imperfect. Dorans and Holland (2000) formalised how to evaluate whether a proposed link is strong enough to call equating, introducing criteria such as population invariance — a true equating relationship should hold regardless of which subgroup of respondents you compute it on. If the link between your old and new evaluation forms changes depending on whether you look at first-years or finalists, it is not an equating; it is at best a rough linking.

The mechanics depend on an anchor: items common to both the old and new forms, administered to comparable groups, that act as a bridge. Equating designs include the common-item nonequivalent groups design — precisely the situation of a questionnaire revision, where this year's students take the new form, last year's took the old, and a handful of retained items connect them. The authoritative Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014) devote an entire chapter to scaling, norming, and linking, and warn explicitly against treating linked scores from different instruments as interchangeable without evidence.

Why it matters for course evaluation in practice

The consequences are concrete and, for individual staff, high-stakes.

A reworded item can shift the mean by more than any real change in teaching. Splitting "the course was well organised and clearly explained" into two separate items, or moving from "agree/disagree" to an item-specific scale, systematically changes how students respond. If a department's average moves 0.2 across a revision year, the honest first hypothesis is the instrument, not the teaching.

Scale-length changes break naïve rescaling. Converting a seven-point score to a five-point score by simple linear proportion assumes the two scales are used the same way — they are not. Respondents cluster differently on scales of different lengths, so linear rescaling introduces bias precisely where you are trying to compare.

Trend lines and control charts inherit the error. Any longitudinal method — regression to the mean checks, interrupted time series, changepoint detection — assumes a stable measurement scale. Run one across an un-linked revision and the "changepoint" it detects may simply be the day the form changed.

The disciplined alternative is to plan revisions like a testing organisation: retain a defensible set of anchor items, collect at least one cycle where old and new coexist or where the anchor is present, estimate the linking, and then report post-revision scores on the pre-revision scale (or vice versa) with an explicit note on linking strength.

Limitations and honest caveats

Equating theory was built for long, high-stakes cognitive tests with large samples and carefully engineered parallel forms; course evaluations are short affective questionnaires with modest per-class samples and items that are not designed to be parallel. This mismatch matters. With few anchor items and small samples, linking estimates are themselves noisy, and a poorly chosen anchor (items that were also reworded, or that behave differently after the revision — anchor "drift") can bias the link more than doing nothing. Equating also cannot rescue a revision that changed the construct: if the new form measures something genuinely different — say, it adds inclusive-teaching items the old form lacked — then no statistical bridge makes the totals comparable, and pretending otherwise is worse than admitting the break. Finally, population invariance is an ideal rarely met perfectly; the practical question is whether the departure from invariance is small enough to tolerate for your specific use, which is a judgement, not a formula.

How Koji incorporates this

Koji treats questionnaire change as a measurement event, not a silent edit. Because Koji versions instruments and stores every item response in structured form (scale, single_choice, ranking, yes_no), it retains the raw material needed to build an anchor-based link when a form is revised, rather than discarding the old-scale detail and leaving only incomparable aggregates. When a programme revises its questionnaire, Koji is designed to preserve a set of unchanged anchor items and flag any longitudinal comparison that crosses a revision boundary, so a dashboard does not quietly draw a trend line across two different instruments. Its reporting is built to label comparisons by the strength of comparability they rest on — genuine like-for-like on retained items versus approximate linking on revised ones — mirroring the equating/linking/prediction hierarchy. And because Koji's AI-moderated conversational interview captures why students answer as they do, a revision-year shift can be interrogated qualitatively — did students react to the new wording or to the teaching? — rather than being mistaken for change. The same versioned-instrument discipline underpins Koji's core research platform at koji.so, where tracking studies over time face the identical "did the metric move or did the question change" problem.

As always, this mitigates rather than eliminates the risk: Koji can preserve anchors and flag boundaries, but it cannot equate two forms that no longer measure the same thing, and it will say so rather than fabricate a comparison.

Frequently asked questions

What is the difference between equating and linking?

Equating is strict interchangeability: a score on the new form means exactly what the same score meant on the old form, for any use. Linking is weaker, approximate alignment. Kolen and Brennan reserve "equating" for forms that measure the same construct with near-equal difficulty and reliability; most course-evaluation revisions qualify only for linking, and some only for prediction.

Why can't I just rescale a 7-point score to a 5-point score proportionally?

Because respondents use scales of different lengths differently — the clustering and endpoint avoidance change with the number of points. A linear proportion assumes identical scale use, which is exactly the assumption a scale-length change violates, so it introduces bias rather than removing it.

What is an anchor item and why do I need one?

An anchor item is a question kept identical across the old and new forms. Anchors act as a bridge that lets you estimate how the two scales relate even though different cohorts answered them. Without a stable anchor, a revision leaves no defensible way to place old and new scores on a common scale.

How many items should I keep unchanged when revising a questionnaire?

There is no universal number, but testing practice favours an anchor that is representative of the whole instrument's content and long enough to estimate the link stably given your sample sizes. For short course evaluations, retaining several core items untouched — and confirming they did not "drift" after the revision — is far safer than changing everything at once.

Can equating fix a comparison if the new form measures something new?

No. If the revision changed the construct — for example by adding a dimension the old form never covered — the totals are not measuring the same thing, and no statistical link makes them interchangeable. The honest move is to treat it as a fresh baseline, not to bridge across the break.

How does this affect trend charts and changepoint detection?

Every longitudinal method assumes a stable scale. Run a trend line, interrupted time series, or changepoint detector across an un-linked revision and it may flag the revision itself as the "change." Link the scales first, or annotate the break, before drawing any trend across it.

Related resources

References

  • Kolen, M. J., & Brennan, R. L. (2014). Test Equating, Scaling, and Linking: Methods and Practices (3rd ed.). Springer. https://doi.org/10.1007/978-1-4939-0317-7
  • Dorans, N. J., & Holland, P. W. (2000). Population invariance and the equatability of tests: Basic theory and the linear case. Journal of Educational Measurement, 37(4), 281–306. https://doi.org/10.1111/j.1745-3984.2000.tb01088.x
  • American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. AERA.
  • Holland, P. W., & Dorans, N. J. (2006). Linking and equating. In R. L. Brennan (Ed.), Educational Measurement (4th ed., pp. 187–220). American Council on Education/Praeger.