Focus Groups for Course Evaluation: What They Reveal That Surveys Miss (and Where They Fail)
What the methodological literature — Stalmeijer et al.'s AMEE Guide No. 91, Kaplowitz & Hoehn's comparative study, and randomized method comparisons — says about using focus groups for course evaluation, their known failure modes, and how AI-moderated individual interviews capture the depth without the group-dynamics distortions.
Koji Education Team
Product
Answer first
Focus groups can surface why students rate a course the way they do — the reasoning, context, and shared experiences that a Likert item cannot capture — but the methodological literature is equally clear about their failure modes: dominant voices, conformity pressure, moderator effects, suppression of socially sensitive topics, and severe scalability limits. A university that relies on surveys alone misses depth; one that relies on focus groups alone misses representativeness and candour. The evidence supports treating group discussion as one qualitative source among several — and increasingly, AI-moderated one-to-one interviews as a way to get interview-grade depth at survey-grade scale.
What the research says
The most systematic treatment of focus groups in an educational context is Stalmeijer, McNaughton and Van Mook''s AMEE Guide No. 91 (Medical Teacher, 2014, 36(11), 923–939). The Guide defines focus groups as group discussions organised to explore a specific set of issues, distinguished from mere group interviews "by the explicit use of the group interaction as research data." That definition matters for course evaluation: the point of a focus group is not to collect eight individual opinions efficiently, but to observe how students negotiate, challenge, and build on each other''s accounts of a course. The Guide sets out the conditions under which that interaction produces trustworthy data: a clearly scoped question, a trained moderator, deliberate composition of groups (homogeneous enough for comfort, heterogeneous enough for productive disagreement), a structured question route, and systematic analysis of transcripts rather than impressionistic note-taking.
The comparative evidence shows the method is not interchangeable with individual interviews. Kaplowitz and Hoehn (Ecological Economics, 2001, 36(2), 237–247) ran a controlled comparison of focus groups and individual interviews on the same topic and found the two formats yield different, complementary information — and, strikingly, that individual interviews were about 18 times more likely to raise socially sensitive topics than focus groups. Although the study''s domain was natural-resource valuation, the mechanism is directly transferable to course evaluation: a student is far less likely to raise a struggling instructor''s behaviour, their own academic difficulties, or experiences of unfair treatment in front of peers than in a private conversation. Guest and colleagues'' randomized comparison of the two methods (International Journal of Social Research Methodology, 2017) similarly found that the choice of format materially shapes the data collected — overlapping but non-identical content — reinforcing that method choice should follow the research question rather than convenience.
For course evaluation specifically, this literature implies a division of labour. Focus groups excel at exploring shared, non-threatening experiences: how a cohort experienced a redesigned curriculum, what a group of students collectively understood a module''s aims to be, where timetabling or assessment structures created common friction. They are weak precisely where much evaluative signal lives: individual struggle, minority experiences, criticism of identifiable teachers, and anything students perceive as risky to say aloud.
Why it matters for course evaluation in practice
Many European quality-assurance offices run "student panels," "course committees," or end-of-module focus groups as a complement to standardized questionnaires — often because response rates on surveys are falling and committees want richer material. The research suggests three practical cautions.
First, group data is not a vote. Because the unit of analysis is the interaction, a focus group cannot tell you what proportion of students hold a view. Six students agreeing in a room may reflect genuine consensus — or one articulate student anchoring the discussion while others acquiesce. Treating focus-group themes as prevalence estimates ("students think the workload is too high") commits a category error the AMEE Guide explicitly warns against.
Second, selection compounds the problem. Students who volunteer for a 60-minute discussion are systematically unrepresentative: more engaged, more available, often more satisfied or more aggrieved than the median student. A survey with a 40% response rate has a documented non-response problem; a focus group of eight volunteers from a cohort of 300 has a far more severe one, usually undocumented.
Third, sensitive signal is structurally suppressed. If individual formats are an order of magnitude more likely to surface sensitive topics, then the issues QA most needs to hear about — supervision problems, assessment unfairness, discriminatory treatment — are the issues least likely to appear in a group transcript. A programme director who concludes "the focus group raised no concerns" may simply have chosen a format in which concerns are structurally unlikely to be raised.
None of this makes focus groups useless. Used as the AMEE Guide prescribes — scoped questions, trained moderation, systematic transcript analysis, findings triangulated against survey and outcome data — they add a dimension surveys cannot. The failure mode is using them casually as a cheap qualitative garnish on a quantitative report.
Limitations and honest caveats
The comparative evidence base has real limits. Kaplowitz and Hoehn''s study was conducted in natural-resource valuation, not higher education; the 18:1 sensitivity ratio is a single-study estimate from one cultural context and should be read as evidence of direction, not a universal constant. Guest et al.''s randomized comparison concerned health topics with a specific population; transportability to European student cohorts is plausible but not demonstrated. The AMEE Guide is a methodological synthesis, not an empirical validation study — it codifies expert consensus and accumulated practice rather than testing focus groups against a learning-outcome criterion. There is, to our knowledge, no randomized study comparing focus groups, surveys, and individual interviews head-to-head on course evaluation content specifically; institutions should treat the division-of-labour argument above as well-grounded inference, not settled fact. Finally, moderator skill is a large, poorly quantified source of variance: the same protocol run by different moderators can produce very different data, which complicates any institutional standardisation of focus-group programmes.
How Koji incorporates this
Koji''s design starts from the finding that individual, private formats surface what group formats suppress — and that the traditional barrier to individual interviews has always been cost, not value.
- AI-moderated one-to-one interviews at cohort scale. Koji replaces the eight-volunteer focus group with structured conversational interviews available to every student in a course. Each interview is private, which targets exactly the sensitivity gap Kaplowitz and Hoehn documented: there is no peer audience, no dominant voice to anchor the discussion, and no conformity pressure.
- Probing without moderator variance. The AI interviewer asks follow-up questions ("you said the feedback came too late — too late for what, specifically?") the way a trained focus-group moderator would, but identically and tirelessly across hundreds of conversations. This is designed to mitigate the moderator-skill variance the methodological literature flags, though it introduces its own design responsibilities (prompt neutrality, avoiding leading questions).
- Prevalence and depth in the same instrument. Because every student can participate, themes extracted from Koji interviews come with denominators. Where a focus group can only say "this came up," Koji''s automatic thematic analysis reports how many students raised a theme, unprompted or probed — closing the theme-versus-prevalence gap that group methods cannot.
- Structured questions where they belong. Koji studies combine open conversational segments with structured items (scale, single_choice, multiple_choice, ranking, yes_no), so quantitative comparability is preserved alongside qualitative depth rather than traded against it.
- Group methods still have a place. Nothing in Koji prevents a QA office from running curriculum-redesign focus groups; Koji''s interview data is designed to triangulate with them — flagging, for instance, when a privately raised theme never appears in the public panel discussion, which is itself diagnostic.
Institutions that also run user or customer research will recognise the same trade-offs; Koji''s core research platform at koji.so applies the same AI-moderated interview engine to product and market research, where the focus-group-versus-interview literature originated.
Practical checklist
If you retain focus groups in your evaluation mix:
- Scope each group to a question interaction can actually answer (shared experience, not prevalence or sensitive topics).
- Recruit deliberately; document who is in the room and who is not.
- Use a trained moderator and a written question route.
- Transcribe and analyse systematically; report themes with quotes, never percentages.
- Triangulate every focus-group theme against survey distributions and individual-format data before acting on it.
- Route sensitive domains (supervision, fairness, wellbeing) to private, individual channels by design.
Where this sits in European quality assurance
The ESG (Standards and Guidelines for Quality Assurance in the European Higher Education Area) require institutions to collect and act on student feedback as part of ongoing programme monitoring, but they are deliberately method-agnostic — which leaves QA offices to defend their methodological choices to review panels themselves. Focus groups appear frequently in self-evaluation reports precisely because they photograph well: a panel reads "we held focus groups with students" as evidence of dialogue. The literature reviewed above suggests review panels — and the institutions reporting to them — should ask harder questions of that evidence. Who was in the room, and how were they recruited? Was the moderator independent of the teaching team being discussed (students soften criticism when the moderator has a stake)? Were transcripts analysed systematically, or did minutes simply record the loudest consensus? And crucially: what individual-format channel existed for the topics the group format structurally suppresses?
A defensible qualitative evidence portfolio for accreditation therefore pairs formats deliberately. Group methods document shared curricular experience and give student representatives a deliberative role — consistent with the student-partnership ethos many European agencies now expect. Individual formats — private interviews, well-designed open-text questions with follow-up probing — carry the sensitive and prevalence-bearing load. When both point the same way, the institution has triangulated evidence; when they diverge, the divergence itself is a finding worth reporting, because it usually marks a topic students will only discuss privately. Institutions that can show panels this two-channel architecture, with the group/individual division of labour justified from the methodological literature rather than from convenience, convert a routine compliance exercise into a genuinely persuasive quality narrative — and, not incidentally, hear about problems while they are still small.
Related Resources
- Nominal group technique for structured student feedback
- Q methodology: mapping student viewpoints
- Critical incident technique for course feedback
- Mode effects and social desirability in conversational course evaluation
- Inter-rater reliability in thematic analysis of open-text feedback
- Most significant change technique for qualitative course evaluation
References
- Stalmeijer, R. E., McNaughton, N., & Van Mook, W. N. K. A. (2014). Using focus groups in medical education research: AMEE Guide No. 91. Medical Teacher, 36(11), 923–939. https://doi.org/10.3109/0142159X.2014.917165
- Kaplowitz, M. D., & Hoehn, J. P. (2001). Do focus groups and individual interviews reveal the same information for natural resource valuation? Ecological Economics, 36(2), 237–247. https://doi.org/10.1016/S0921-8009(00)00226-3
- Guest, G., Namey, E., Taylor, J., Eley, N., & McKenna, K. (2017). Comparing focus groups and individual interviews: findings from a randomized study. International Journal of Social Research Methodology, 20(6), 693–708. https://doi.org/10.1080/13645579.2017.1281601
Related articles
How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis
When you turn thousands of free-text comments into themes and counts, how do you know the coding is trustworthy? O Connor and Joffe (2020) on intercoder reliability, Braun and Clarke on thematic analysis, and what rigorous qualitative QA looks like.
The Critical Incident Technique for Course Feedback
How Flanagan's Critical Incident Technique collects concrete, behaviourally-anchored student feedback that global Likert ratings cannot capture, and how it applies to course evaluation.
Q-Methodology: Surfacing the Distinct Viewpoints Students Hold About a Course
Q-methodology uses forced-choice card sorts and by-person factor analysis to reveal the two-to-four genuinely distinct viewpoints students hold about a course, a rigorous complement to Likert SET averages.
The Most Significant Change Technique: Story-Based Course Evaluation That Surfaces What Students Actually Value
Likert averages tell you where a course sits; they cannot tell you what transformed a student. The Most Significant Change technique collects and collectively selects stories of change to reveal unanticipated outcomes and shared values. Here is how it works and where it fits course evaluation.