Paper, Web, or Phone: How the Mode of Your Course Evaluation Changes the Answer
Moving course evaluations from paper to online is treated as an IT decision. It is a measurement decision. The administration mode changes who responds — and can change what they say — in ways that quietly break comparisons across years, cohorts, and departments.
Koji Education Team
Product ·
Bottom line: How you administer a course evaluation — paper in the lecture hall, a web link by email, a prompt on a phone — is not a neutral logistics choice. It is a measurement decision that changes your response rate dramatically and can subtly change the responses themselves. The best-documented effect is on who answers: the shift from in-class paper to online routinely halves response rates. The reassuring news is that, on the evidence, mode does far less to the average score than most faculty fear. Knowing the difference between those two effects is what separates a defensible evaluation programme from one that compares last year's paper numbers to this year's online numbers as if they were the same measurement.
The decision everyone treats as plumbing
When an institution retires its paper forms and moves evaluations online, the change is usually owned by IT or the registry. It saves paper, saves class time, and produces tidy data. What almost nobody does is treat it as what it is: a change to the measurement instrument. In survey methodology, the channel through which you collect data — the mode — is a recognised source of error in its own right, sitting inside the measurement and nonresponse branches of the Total Survey Error framework. Changing mode mid-stream is like recalibrating a thermometer and then comparing this month's readings to last month's without noting the recalibration.
There are two distinct questions to keep separate: does mode change who responds (a nonresponse effect), and does mode change what a given respondent says (a measurement effect)? Conflating them produces most of the confused arguments about online evaluations.
Mode strongly changes who responds
The response-rate effect is large, consistent, and well-documented. In-class paper evaluations enjoy a captive audience: the form is handed out, everyone present fills it in, and response rates are high. Emailed online evaluations rely on voluntary action later, and a large share of students simply never click.
The canonical comparison comes from Dommeyer, Baum, Hanna and Chapman (2004, Assessment & Evaluation in Higher Education, 29, 611–623), who reported a 75% response rate for in-class paper evaluations versus 43% online for the same courses. That gap — on the order of 30 percentage points — is typical of the literature. Donald Nulty's widely cited review, "The adequacy of response rates to online and paper surveys: what can be done?" (2008, Assessment & Evaluation in Higher Education, 33, 301–314), documents the same pattern and works through what response rate you actually need before the data can bear weight for accountability versus improvement purposes.
Why this matters: a lower response rate is not automatically biased, but it widens the door for nonresponse bias, because the students who opt in online may differ from those who would have been captured in class. The mode change does not just add noise; it changes the composition of your sample. Comparing a paper-era 4.2 to an online-era 4.2 assumes the two 4.2s were produced by comparable populations. Often they were not.
But mode does surprisingly little to the score itself
Here is where faculty fear tends to outrun the evidence. The common worry is that online evaluations, completed at home by a self-selected and possibly disgruntled minority, will produce systematically harsher scores. The data are largely reassuring on this point.
The same Dommeyer et al. study found that online and in-class administration did not produce significantly different mean evaluation scores, even when incentives were used to lift online response. Multiple subsequent comparisons have reached broadly similar conclusions: the level of ratings is fairly robust to mode, even as the response rate is not. There is some evidence that online administration yields slightly longer and more numerous open-text comments, plausibly because typing at one's own pace is easier than scribbling in the last two minutes of class — a quality gain, not a loss.
So the honest summary is asymmetric: mode is a big deal for representation and a small deal for measurement of the mean. That asymmetry is the actionable insight. The thing to worry about when you go online is not "will my scores drop" but "will enough of the right students respond that the number means anything."
The mobile wrinkle nobody planned for
The newest layer is that "online" is no longer one thing. A growing share of students complete evaluations on a phone, often on a small screen, in fragments of attention. Long grids of Likert items designed for a full-page paper form render badly on mobile and invite satisficing — straight-lining down a column, skipping the open text, tapping to be done. Mode is now fragmenting further, and an instrument designed for paper and merely displayed on a phone is being answered under conditions its designers never imagined. If you migrated a paper form to the web a decade ago and have not revisited it, a large and rising share of your responses are being collected through a channel the instrument was never validated for.
But doesn't this mean we should just go back to paper?
The obvious counterargument: if paper gets 75% and online gets 43%, and scores are similar anyway, why not keep paper and enjoy the better response rate? Three reasons this is too quick.
First, the paper response rate is bought with class time and a captive audience — a coercive element that has its own costs and that the sector is moving away from for good reasons, including the pressure it puts on students and the missed responses from anyone absent that day. Second, paper collapses at scale and across modalities: it cannot evaluate online, hybrid, block, or work-integrated courses where there is no shared room to hand out forms. Third, paper's headline response rate hides its own coverage gap — only students physically present are sampled, which for many modern courses is a shrinking and non-random subset.
The right response to the mode problem is not nostalgia. It is to (a) hold mode constant when you compare across time, or explicitly flag the break when you cannot; (b) design the instrument for the mode students actually use; and (c) invest in the response-rate levers Nulty and others identify — in-class time to complete, reminders, clarity about anonymity, and visible evidence that feedback is acted on — rather than assuming the channel is neutral. A critic might fairly add that some mode differences are confounded with timing (online evaluations often stay open until after grades are released), and that is exactly the kind of variable a serious programme controls rather than ignores.
Where Koji fits: one standardised, conversational mode
Koji for Education approaches the mode problem from a different direction: rather than porting a paper grid onto a phone, it is built mobile-first as a conversational instrument, so the channel and the design match.
- Because Koji uses AI-moderated conversational interviews rather than a long static grid, it sidesteps the mobile satisficing that kills open-text quality — the interaction is a short, adaptive exchange rather than a wall of Likert rows.
- The AI moderation is standardised across every respondent and every device, which removes the human-moderator and layout inconsistencies that make mode comparisons fragile. Everyone gets the same well-calibrated conversation whether on a laptop or a phone.
- Six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) and automatic thematic analysis capture the richer open text that online modes make possible, at programme scale.
- Because the mode is consistent and the instrument is designed for it, year-on-year comparisons hold up better than a paper-to-web-to-mobile migration ever could — and Koji works equally for in-person, online, and work-integrated courses that paper cannot reach.
Koji cannot repeal nonresponse — no tool can make a student answer — but by making mode a designed constant rather than an accidental variable, it removes one of the quiet reasons evaluation trends are so often uninterpretable. The same conversational engine underpins the main Koji platform for customer and user research, where mode effects are just as real.
Stop comparing this year's mode to last year's. Explore Koji for Education for a standardised, mobile-native evaluation instrument built for the channel your students actually use.