Focus Groups vs Surveys for Course Evaluation: The Depth-Versus-Scale Trade-Off
Surveys scale but stay shallow; focus groups go deep but cannot scale and carry their own biases. For decades that trade-off has forced universities to choose. Conversational AI is the first method that credibly attacks both ends at once.
Koji Education Team
Product ·
Bottom line up front: The survey and the focus group have opposite strengths. Surveys reach every student but flatten their answers into ratings and short comments; focus groups reveal why students think what they think but reach a handful of people and introduce group-dynamic biases of their own. Course-evaluation practice has been stuck choosing between reach and depth. Understanding exactly where each method breaks is the key to seeing why AI-moderated conversational interviews are not just "nicer surveys" — they target the trade-off itself.
Two methods, opposite failure modes
Almost every institution runs the end-of-semester survey. A smaller number supplement it with focus groups — often only for programme review or when a course is already in trouble. The two are usually treated as a hierarchy (surveys for everyone, focus groups for special cases) rather than as what they actually are: instruments with mirror-image weaknesses.
The survey scales and flattens. You can send it to 10,000 students and get structured, comparable data back. But a Likert item cannot ask a follow-up. When a student rates "assessment and feedback" a 2, the survey cannot ask which assessment, what about the feedback, or what would have helped. You are left with a number and, if you are lucky, a one-line comment written by a tired student at 11pm. The result is broad but thin — which is why "assessment and feedback" scores lowest almost everywhere and almost nobody can say precisely why from the survey alone.
The focus group deepens and cannot scale. A skilled moderator can chase exactly those follow-ups, surface the reasoning behind a rating, and let students build on each other''s points. The cost is reach and effort: you can realistically run a few groups per programme, and each one demands a room, a scheduled time, a trained facilitator, and hours of transcription and analysis.
The focus group''s own biases
It is a mistake to treat the focus group as the "truth" that surveys only approximate. The qualitative-methods literature is candid about its distortions.
Group size is a genuine constraint. Krueger and Casey — the standard reference — recommend roughly 6 to 10 participants, dropping to 4 to 6 "mini-groups" for complex topics, precisely because the format degrades outside that band (Krueger & Casey, Focus Groups: A Practical Guide for Applied Research). Too few and participants feel pressure to talk more than they want to; too many and everyone competes for airtime and detail collapses.
Then there is the dominant participant. As practitioner guidance puts it, one dominating voice can influence how everyone else in the session responds — and with a group of six, it takes exactly one confident student to reshape the discussion. The related risk is groupthink: participants converging on a shared view out of politeness or deference to that dominant voice rather than genuine agreement. A quiet student who disliked the course may simply nod along.
Add self-selection (who volunteers for a focus group is not a random sample), moderator effects, and the social-desirability pressure of saying critical things about a lecturer out loud in front of peers, and the focus group''s "depth" turns out to be depth of a particular, skewed slice of the cohort. It is genuinely useful — and genuinely not representative. Reaching thematic saturation, moreover, typically requires several groups per student population, so even the depth a focus group offers is only trustworthy after multiple sessions — multiplying the cost at exactly the moment you most need breadth. Few evaluation budgets ever fund enough groups to get there, which means the "deep" data is often thin data wearing a qualitative badge.
So the honest scorecard is: surveys give you representative but shallow; focus groups give you deep but unrepresentative and labour-intensive. Neither is the answer on its own, which is why the methods literature has long argued for triangulation across multiple evidence sources.
The counterargument: can''t you just add open-text boxes and run more groups?
The obvious rebuttal is that the trade-off is soft, not hard: bolt more free-text questions onto the survey to add depth, and run more focus groups to add reach.
Both moves hit ceilings quickly. Adding open-text questions raises the burden on every respondent, which depresses response rates and answer quality — you get more boxes, more of them blank, and the ones that are filled still cannot be followed up. And "run more focus groups" runs straight into the resource wall that made them exceptional in the first place; there is no version of focus groups that covers every module every semester within a real quality-assurance budget. The trade-off is not an accident of tooling. It is baked into what a static form and a human-moderated room can each physically do.
Why conversational AI attacks the trade-off itself
The reason to care about the anatomy of these two methods is that a third option now exists that is not a compromise between them but an attempt to get the strengths of both.
An AI-moderated conversational interview keeps the survey''s reach — every student in a 10,000-person cohort can be interviewed simultaneously — while behaving like a good focus-group moderator with each of them individually. When a student says the feedback was unhelpful, the AI can ask which assignment, what was missing, and what would have helped — the follow-up a Likert item structurally cannot make. This is the depth-at-scale move, and it is why we argue the one-size-fits-all form is the wrong default.
It also sidesteps the focus group''s specific biases:
- No dominant participant, no groupthink. Each student is interviewed alone, so no confident peer reshapes the room and nobody self-censors in front of classmates. You get the reflective depth of a conversation without the social dynamics that distort a group.
- Representative, not self-selected. Because it is issued to the whole cohort like a survey rather than to a handful of volunteers, it does not depend on who is willing to give up an evening.
- Consistent moderation. A single AI moderator applies the same probing style to every student — removing the between-facilitator inconsistency that makes focus-group data hard to compare across sessions or campuses.
- Analysis that scales with the data. Depth is only useful if you can read it. Koji''s automatic thematic analysis clusters thousands of conversational transcripts into themes, so more depth does not mean an unmanageable transcription pile — a problem we unpack in why a sentiment score is not insight.
Koji supports six structured question types (open-ended, scale, single- and multiple-choice, ranking, and yes/no), so the conversational depth sits alongside the comparable, quantitative backbone a quality office still needs. The same AI interview engine powers general user and customer research on the main Koji platform, where the depth-versus-scale problem is identical.
We should be careful about the claim. A conversational AI does not reproduce a focus group — it deliberately removes the group, and with it the genuine value of students hearing and building on each other''s ideas. For co-creation and deliberative work, a real group still matters, and students-as-partners approaches remain irreplaceable. What the AI interview replaces is the focus group''s role as a depth instrument for evaluation — and there, doing it one-to-one at cohort scale is a strict improvement on doing it six-at-a-time for the few.
The bottom line
For thirty years, course evaluation has been a choice between a survey that reaches everyone and says little, and a focus group that says a lot about a few. The trade-off was real because forms and rooms have hard physical limits. Conversational AI is the first method that credibly refuses the choice: survey-scale reach, interview-depth probing, and none of the group-dynamic distortions. It will not make focus groups obsolete for co-design — but as the default way to ask an entire cohort why, it changes what "depth at scale" can mean.
Want interview-depth feedback from every student, not just a focus-group few? See Koji for Education.