Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You
Natural-language processing can turn thousands of free-text course-evaluation comments into themes and sentiment at scale — but the research (Cunningham-Nelson 2019; Sunar & Khalid 2023) shows where automated analysis is reliable and where human judgement is still required.
Koji Education Team
Product
Answer
Natural-language processing (NLP) makes it feasible to read every open-ended course-evaluation comment instead of the handful a busy committee skims — surfacing recurring themes, sentiment, and outlier issues across thousands of responses in minutes. The research shows this is genuinely useful for discovery and visualisation at scale, and that modern systems classify student feedback with usable accuracy. But the same literature is candid about limits: sentiment labels miss sarcasm, mixed messages and context; theme models can fragment or blur categories; and the analysis can launder a thin or biased set of comments into an authoritative-looking dashboard. The defensible practice is to use NLP to route human attention, not to replace human reading — and to keep the qualitative evidence anchored to verbatim quotes.
What the research says
Cunningham-Nelson, Baktashmotlagh and Boles (2019), writing in IEEE Transactions on Education, demonstrated an automated methodology that visualises students'' free-text comments from course-satisfaction surveys to reveal which teaching-and-learning aspects are performing well and which need attention. Their contribution is practical: rather than asking staff to read hundreds of comments unaided, an NLP pipeline (with a pre-processing stage that normalises and structures the text) extracts the salient topics and presents them visually, turning an unmanageable corpus into an interpretable summary. In related work in the same research programme, NLP approaches reached accuracies around 85% when assessing conceptual content from textual responses — a useful benchmark for what automated text classification can achieve in an educational setting, while also implying a meaningful error rate that matters at the level of an individual comment.
Sunar and Khalid (2023), in a systematic review in IEEE Transactions on Learning Technologies, synthesised the broader field of NLP applied to students'' feedback to instructors. They map the pipeline end to end — the data sources, the labelling and translation effort required, the methods used, and the categories of prediction/analysis targets (sentiment, topic/theme, aspect-based opinion, actionable suggestions). Two themes stand out for practitioners. First, labelling is the bottleneck: supervised models need human-coded training data, and the quality of automated output is hostage to the quality and representativeness of that coding. Second, the field is methodologically fragmented — many bespoke pipelines, inconsistent evaluation, and limited cross-institutional generalisability — which counsels against treating any single tool''s output as a turnkey ground truth.
The honest synthesis: NLP reliably compresses large comment sets into navigable structure (Cunningham-Nelson et al.), but the maturity, transparency and generalisability of these methods vary widely (Sunar & Khalid), so automated themes and sentiment should be treated as hypotheses to verify against the verbatim text, not as findings in themselves.
Why it matters for course evaluation in practice
1. Comments are where the actionable signal lives — if you can read them all. Likert numbers tell you that satisfaction dipped; comments tell you why (the assessment brief was unclear, the lab sessions were too rushed). NLP is what makes "read everything" tractable for a programme with thousands of responses, so the rich qualitative layer is no longer sacrificed to time pressure.
2. Sentiment scores are a triage tool, not a verdict. A −0.3 average sentiment is a prompt to investigate, not a conclusion. Automated sentiment routinely mislabels "the course was hard but I learned a lot" (positive experience, negative-leaning words) and misses irony. Use it to rank where to look; confirm by reading.
3. Themes must stay traceable to quotes. For quality assurance and accreditation, a theme like "assessment clarity" is only credible if a reviewer can click through to the actual student sentences behind it. A model that outputs a bar chart with no audit trail invites the ecological version of cherry-picking.
4. Garbage in, authoritative-looking garbage out. If only the most aggrieved 20% of students comment, NLP will faithfully summarise their themes and present them with a veneer of completeness. The analysis cannot fix a biased or thin sample — it can only make it look polished. This is why representativeness (response rate and count) must be reported alongside any text analytics.
Limitations and honest caveats
- Accuracy is aggregate, not per-comment. An 85%-accurate classifier is genuinely useful across a corpus but wrong roughly one time in seven on individual comments — unacceptable if a single misclassified comment is used against an instructor. Automated labels need human review before any high-stakes use.
- Sentiment ≠ meaning. Sarcasm, hedging, mixed evaluations, and discipline-specific jargon all defeat naïve sentiment models. Negation ("not at all clear") and contrast ("hard but fair") are classic failure points.
- Language and equity. Many pipelines are built and validated on English; performance can degrade for multilingual cohorts and for non-native English writers — a real concern in European higher education and one that can systematically under-represent some students'' voices. Sunar and Khalid flag translation/labelling effort precisely here.
- Construct drift and opacity. A "topic" a model finds is a statistical cluster, not necessarily a pedagogically meaningful category; and opaque models make it hard to explain why a comment was labelled a certain way — a problem for accountability.
- The model can entrench bias. If training labels carry coders'' or students'' biases (e.g. harsher language toward certain instructor groups), the automated summary can amplify rather than neutralise them. Text analytics is not a bias-removal step.
How Koji incorporates this
Koji for Education treats automated text analysis as a way to direct and substantiate human judgement, with the verbatim record always in reach.
- Automatic thematic analysis, anchored to quotes. Koji applies thematic analysis to open-text responses to surface recurring themes (assessment clarity, pace, workload, support) across an entire cohort — but keeps each theme traceable to the underlying student sentences, so a quality-assurance officer or accreditation reviewer can verify the claim rather than trust a chart. This directly answers the "themes must stay traceable" requirement above.
- Conversational depth improves the raw material. Cunningham-Nelson''s pipeline can only analyse the comments students bother to write; Koji''s AI-moderated interview elicits more and better text by probing with structured follow-ups (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) — turning a one-line gripe into a specific, codable account. Better input text is the single biggest lever on output quality, the bottleneck Sunar and Khalid identify.
- Sentiment as triage, with caveats surfaced. Where Koji highlights sentiment or theme prevalence, it is designed to frame these as signals to investigate, paired with respondent counts so a thin or skewed comment set is not mistaken for the cohort''s consensus.
- Quality scoring and bias-aware reporting. Koji applies quality scoring to responses and bias-aware framing to its summaries, mitigating the risk that automated analysis launders a biased or unrepresentative sample into an authoritative-looking result.
These mechanisms are designed to mitigate the documented limits of automated text analysis, not to claim that NLP removes the need for human reading on high-stakes decisions. For the complementary question of what open text adds over Likert scores, see our companion piece below. Koji''s core research platform at koji.so applies the same theme-extraction-with-verbatim-traceability engine to product and customer research, where over-trusting an automated sentiment score is an equally common pitfall.
A workflow for trustworthy text analytics
The research supports a disciplined workflow that captures NLP''s scale advantage while respecting its documented limits. The principle throughout is that automation prioritises human attention rather than replacing it.
- Elicit good text first. Output quality is capped by input quality. Prompt students for specifics ("what one change would most improve this course?") rather than open-ended "any comments?", which tends to yield either silence or venting. Richer source text is the highest-leverage step, upstream of any model.
- Cluster and visualise to navigate, not to conclude. Use topic modelling and sentiment to build a map of where the volume and the heat are — the contribution Cunningham-Nelson et al. demonstrate — then treat each cluster as a question, not an answer.
- Read the verbatim behind every reported theme. Before a theme enters a report, a human should read a sample of the underlying comments and confirm the label fits. This catches sarcasm, negation and discipline-specific language that defeat naïve models.
- Report representativeness alongside themes. Always pair the text analysis with the respondent count and rate, so a vocal minority is not presented as the cohort''s consensus.
- Quarantine high-stakes use. Never let a single automatically-classified comment drive a personnel decision; aggregate accuracy of ~85% still means meaningful per-comment error. Human adjudication is mandatory at the individual level.
Followed in order, this workflow lets a programme read every comment at scale without surrendering the judgement, traceability and fairness that accreditation reviewers — and the instructors being evaluated — are entitled to expect.
Related Resources
- What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
- How Many Responses Do You Need for a Reliable Course Evaluation?
- Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
- How Many Scale Points Should a Course-Evaluation Question Have?
- Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
- Turning Student Feedback into ESG / ENQA Accreditation Evidence
References
- Cunningham-Nelson, S., Baktashmotlagh, M., & Boles, W. (2019). Visualizing student opinion through text analysis. IEEE Transactions on Education, 62(4), 305–311. https://doi.org/10.1109/TE.2019.2924385
- Sunar, A. S., & Khalid, M. S. (2024). Natural language processing of student''s feedback to instructors: A systematic review. IEEE Transactions on Learning Technologies, 17, 741–753. https://doi.org/10.1109/TLT.2023.3330531
- Somers, R., Cunningham-Nelson, S., & Boles, W. (2021). Applying natural language processing to automatically assess student conceptual understanding from textual responses. Australasian Journal of Educational Technology, 37(5), 98–115. https://doi.org/10.14742/ajet.7121
Related articles
Turning Student Feedback into ESG / ENQA Accreditation Evidence
A buyer's guide mapping the ESG 2015 internal quality assurance standards to concrete, accreditation-ready evidence you can generate from student feedback — and how AI-moderated evaluation closes the loop.
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.
What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.