New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends8 min read

The Future of Student Feedback: From Satisfaction Surveys to Conversational Interviews

Response rates are falling, the data is thin, and students suspect nothing changes. The satisfaction survey has reached the limits of its design. We look at the evidence for a different paradigm — conversational, AI-moderated feedback — and where it genuinely helps versus where the hype runs ahead.

Koji for Education

Research & Editorial Team · June 6, 2026

Bottom line: The end-of-term Likert survey is hitting structural limits — falling response rates, shallow data, and student cynicism that feedback changes nothing. The emerging alternative is conversational, AI-moderated feedback that asks fewer fixed questions and probes more deeply. The evidence is promising but still early; the honest framing is that conversational methods address the survey''s specific failure modes — depth, engagement, and closing the loop — not that they are a panacea. Institutions should pilot deliberately, keeping a lightweight trended instrument alongside.

The survey paradigm is running out of road

The standard student feedback survey was designed for a world of paper forms handed out in the last ten minutes of class. It has not aged well. Two problems have compounded.

The first is falling response rates. Moving evaluations online — necessary for scale and analysis — reliably depresses participation. In a widely cited review, Nulty (2008) found online response rates averaged roughly 23% lower than in-class collection, with typical online rates landing around 50–60% against 70–80% for paper. Individual institutions have seen sharper drops: one large public university fell from 73% on paper to as low as 43% after going online. We examine the consequences for representativeness in our piece on response rates and non-response bias — low, self-selected response makes the resulting averages unreliable in exactly the way a quality process can least afford.

The second is shallow data. A five-point scale on "overall satisfaction" compresses a semester of experience into a single ordinal tick. As we argue in why averaging Likert scores misleads, that number is both statistically fragile and substantively thin. It tells you that students were lukewarm, never why.

Underneath both is a third, human problem: students suspect nothing changes. When feedback does not visibly lead to action, motivation to provide it collapses — a self-reinforcing spiral of survey fatigue. The instrument asks more and more from students while showing them less and less in return.

What changes with conversation

The conversational paradigm inverts the design. Instead of a fixed battery of items, an AI moderator asks an opening question and then follows up — probing a vague answer, asking for an example, distinguishing a course-design problem from a one-off incident. The interaction feels like being heard rather than processed, which is precisely the engagement lever a static form lacks.

Three concrete advantages follow, each mapped to a specific survey failure mode:

  • Depth instead of compression. A conversation surfaces the reason behind a rating — the actionable layer that a scale discards.
  • Engagement instead of fatigue. A short, responsive conversation that adapts to the student is, for many, less tedious than a long matrix of Likert grids, which can lift willingness to participate meaningfully.
  • Closing the loop instead of the void. When the method is designed around action tracking, students can see that themes get addressed — directly countering the "nothing changes" driver of fatigue.

This is not science fiction; it is where careful institutions are already piloting. The broader regulatory and quality context for AI in this space — and the guardrails it demands — is something we cover in AI in European higher-education quality assurance.

But isn''t this just chatbot hype dressed up as research?

This is the right question to ask, and a skeptical audience should hold the claim to account. Several cautions are legitimate.

First, evidence maturity. The literature on conversational feedback is young. Claims that AI interviews "boost response rates and depth" are plausible and supported by early deployments, but the multi-year, multi-institution randomised evidence that exists for, say, the bias literature does not yet exist here at the same scale. Honesty requires saying so. The correct posture is pilot and measure, not rip and replace on faith.

Second, new biases for old. An AI moderator removes human-moderator inconsistency, but a model has its own tendencies — in how it phrases probes, what it follows up on, which answers it treats as complete. These must be audited, not assumed away. Replacing a known bias with an unexamined one is no progress.

Third, data protection. Conversational feedback collects richer, more identifiable free text. For European institutions, that raises real GDPR/AVG obligations around consent, minimisation, retention and anonymisation in reporting. A conversational tool that is not engineered for EU data handling is a liability, not an upgrade.

Fourth, the comparability problem. Institutions track satisfaction trends year over year. Abandoning the Likert instrument entirely breaks that time series. The pragmatic answer most thoughtful pilots adopt is a hybrid: keep a short, stable set of scale items for the trend line, and add a conversational layer for the depth. You lose nothing comparable and gain the qualitative signal.

None of these cautions argue for staying with the broken survey. They argue for adopting the new paradigm with rigour — which is exactly the standard a PhD-literate quality office should demand of any vendor, including us.

How to run a defensible pilot

For a skeptical quality office, the credible path is a structured pilot, not a leap of faith. A few principles keep it honest:

  • Run it head-to-head. Keep your existing Likert instrument for a subset of courses and add the conversational method alongside in others — or, better, run both on the same cohort. You want a real comparison, not a vibe.
  • Pre-register what "better" means. Decide in advance which outcomes count: response rate, completion depth, proportion of actionable comments, time-to-insight for the quality office, and student-reported experience of the process. Measuring after the fact invites motivated interpretation.
  • Audit the moderation. Sample the AI''s probes and follow-ups for leading questions, uneven depth, or topics it systematically over- or under-pursues. Treat the model as an instrument to be calibrated, not a black box to be trusted.
  • Get the data protection right first. Define consent, retention, anonymisation-in-reporting and access controls before a single interview runs. For EU institutions this is a precondition, not a clean-up task.
  • Protect the trend line. Keep a small, stable set of scale items so you do not sacrifice year-over-year comparability while you evaluate the richer layer.

A pilot run this way answers the only question that matters — does the new method produce better, fairer, more actionable evidence for your students — with data rather than enthusiasm. If it does not, you have lost little. If it does, you have an evidence-backed case for change that will survive scrutiny from exactly the audience that should be scrutinising it.

Where Koji fits

Koji for Education is built for this transition, with the cautions above designed in rather than waved away. Its AI-moderated conversational interviews deliver the depth-and-engagement advantages directly, while its six structured question types (open-ended, scale, single-choice, multiple-choice, ranking, yes/no) make the hybrid model native: keep your trended Likert items and add conversational probing in one instrument, so you never break the time series. Bias-aware, standardised moderation addresses the new-bias concern by asking every student comparable, neutral questions and making the moderation logic consistent rather than ad hoc. Automatic thematic analysis turns the richer data into patterns a quality office can act on, and closing-the-loop action tracking is the part that actually fights survey fatigue — students see that themes lead somewhere. Crucially, Koji is designed for GDPR/AVG-compliant, EU-appropriate data handling, which is non-negotiable for the more identifiable text conversations produce.

We are deliberately not claiming conversational feedback is a finished, proven replacement — the evidence is still maturing, and we would distrust any vendor who told you otherwise. What we claim is narrower and defensible: conversational, AI-moderated feedback targets the specific failure modes of the satisfaction survey, and a hybrid pilot lets you test that on your own students without losing your trend data. The same conversational engine underpins general user and market research on the main Koji platform, so the method is battle-tested well beyond education.

If your response rates are sliding and your evaluation data feels thin, the move is not another reminder email — it is to rethink the instrument. Explore Koji for Education and run a hybrid pilot with the rigour the question deserves.