New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

Do Students'' Written Comments Match Their Ratings? What Concordance Tells You

Open-text comments and Likert scores usually agree — but the gaps are where the insight lives. What Alhija & Fresko (2009) and Brockx et al. (2012) found about the consistency between qualitative and quantitative course-evaluation data, and how to read it.

Koji Education Team

Product

In brief: Across the empirical record, students'' written comments are broadly consistent with the numbers they assign — courses that score low attract more critical comments, and vice versa. But comments are written by a self-selecting minority (typically 45–70%), skew general rather than specific, and lean positive. The practical value of open text is not that it confirms the scores; it is that where comment and score diverge — and where comments name causes a Likert item never asked about — you find the actionable signal.

What the research says

Two complementary European studies anchor what we know about the relationship between students'' closed-ended ratings and their free-text comments.

Alhija and Fresko (2009), in Studies in Educational Evaluation, analysed written comments from 198 classes at an Israeli institution. Their headline descriptive findings have held up remarkably well:

  • Only about 45% of students who completed the form wrote any comment at all. Open text is produced by a subset, raising an immediate representativeness question.
  • Comments were more often positive than negative, and more often general than specific ("great course") rather than diagnostic ("the second assignment had no rubric").
  • They built an inductive coding scheme of 45 subcategories grouped into eight categories spanning the course (content, assignments, general evaluation), the instructor (personal traits, teaching style, general evaluation) and the context (scheduling, student composition).
  • Comments were broadly consistent with the quantitative ratings, but the open text surfaced themes — scheduling, room, group composition — that the fixed questionnaire never asked about.

Brockx, Van Roy and Mortelmans (2012), in Procedia — Social and Behavioral Sciences, examined comment behaviour in a Belgian SET system and found a higher comment rate — around 70% of returned surveys contained a comment — and, importantly, a clear directional link: surveys with low evaluation scores were significantly more likely to carry negative comments. Students, they concluded, "take their task as commentators seriously"; the text is not random noise bolted onto the numbers but a coherent second channel.

A third strand sharpens the qualitative side. Work on the language of praise and criticism in evaluation surveys (e.g., Stupans, McGuren and Babey, 2016, Studies in Educational Evaluation) shows that even when overall tone tracks the score, the content of praise and criticism is patterned and informative — students criticise assessment and organisation in specific, recurring ways that a global satisfaction number cannot represent.

Put together, the literature supports a two-part conclusion. First, concordance is the norm: qualitative and quantitative data mostly point the same direction, which is reassuring for the basic credibility of both. Second, the residual is the point: the comments add value precisely where they are not redundant with the scores — by naming causes, surfacing un-asked topics, and flagging the cases where a middling number hides a strong opinion.

Why it matters for course evaluation in practice

If comments simply restated the scores, reading them would be a waste of committee time. They do not, and three practical consequences follow.

1. A clean numeric score is not a clean course. Because comments routinely raise issues the questionnaire never asked — timetabling clashes, a broken assessment rubric, an unhelpful room — a course can post respectable means while open text reveals a fixable, specific problem. QA processes that report only the closed-item averages systematically miss this layer.

2. Divergence between channels is a diagnostic flag, not an error. When a course''s scores are average but its comments are sharply negative (or vice versa), that gap is informative. It often signals a polarised cohort — a vocal minority with a strong, specific grievance averaging out against a quietly satisfied majority. The mean conceals exactly what the programme director needs to act on.

3. The minority who write are not the average respondent. With comment rates between 45% and 70%, and a known skew toward both general praise and motivated criticism, raw comment counts ("12 positive, 3 negative") are a poor summary. Comments should be read thematically and weighted against the response base, not tallied like votes (see text analytics for open comments).

This is why mature quality cycles treat the closed items and the open text as two instruments measuring overlapping but non-identical constructs — and read them together, looking deliberately for both agreement (which builds confidence) and disagreement (which generates action).

Limitations and honest caveats

The concordance literature carries real constraints a critical reader should hold in mind.

  • Self-selection in who comments. Across studies, fewer than three-quarters — sometimes fewer than half — of respondents write anything. Comment-derived themes describe the commenters, not the cohort. Apparent "consensus" in the text may be the consensus of the motivated. This compounds the unit non-response already present in the scores themselves (see non-response bias).
  • "Consistency" is measured coarsely. Showing that low-scoring courses attract more negative comments is a directional, aggregate claim. It does not establish that an individual student''s comment matches that same student''s rating, nor that the magnitudes line up. Genuine item-level concordance is harder to demonstrate and less often tested.
  • Coding is interpretive. Eight categories and 45 subcategories are a defensible scheme, but a different team would draw the boundaries differently. Inter-rater reliability in comment coding is achievable but not automatic (see inter-rater reliability in thematic analysis).
  • Context-bound samples. These are single-system studies (Israeli, Belgian) from before the shift to fully online, mobile, multilingual evaluation. Comment rates and tone are sensitive to mode, prompt wording and culture, so the specific percentages should be read as illustrative, not universal constants.
  • Negativity and recency dynamics. Voluntary comments are vulnerable to negativity bias and end-of-term affect, so the open channel has its own distortions, not a privileged window onto truth (see negativity bias in comments).

The defensible reading is modest: comments and scores usually agree, agreement is reassuring but not proof of validity, and the analytic payoff is in the structured, representativeness-aware reading of where they diverge.

How Koji incorporates this

Koji is designed to capture why a student gave the score they gave, in the same exchange — closing the gap between the numeric and the narrative channel that the concordance literature treats as separate instruments.

  • Probing every rating in context. Where a traditional form collects a Likert score and, separately and optionally, a free-text box, Koji''s AI-moderated conversational interview follows a rating with an adaptive "why" or "can you give a specific example?" This is designed to convert the general praise and criticism Alhija and Fresko found dominant into the specific, actionable detail that QA actually needs — and to lift the share of respondents who say something substantive beyond the self-selecting minority who volunteer comments on a static form.
  • Linking comment to score at the response level. Because the conversation and the rating come from the same respondent in one session, Koji can surface score–narrative divergence per response, not just in aggregate — flagging the polarised cases where a moderate mean hides a strong, specific grievance.
  • Automatic thematic analysis with structure. Koji applies automatic thematic analysis to open text, grouping comments into recurring themes (assessment, workload, organisation, context) rather than reducing them to a single sentiment value. This is designed to mitigate the "tally the comments" failure mode and to respect that comments raise topics the fixed questions never asked. We frame this as assistive — surfacing candidate themes for human QA review, not replacing it (see why a sentiment score is not insight).
  • Representativeness in view. Reporting keeps the comment base visible against the response base, so a vivid theme from a vocal few is not mistaken for cohort consensus.

The same conversational engine underpins Koji''s core product and customer-research platform at koji.so, where reconciling what users rate with what they say is the central analytic task.

The bottom line for practice

Treat the closed items and the open text as two instruments, not one questionnaire with a comment box tacked on. Read them for agreement, which builds confidence in both, and for divergence, which generates action — and always weigh what the vocal minority wrote against how many students actually responded. The single most valuable thing a comment can do is name a cause, or raise a topic, that no fixed question thought to ask.

Related resources

References

  • Alhija, F. N.-A., & Fresko, B. (2009). Student evaluation of instruction: What can be learned from students'' written comments? Studies in Educational Evaluation, 35(1), 37–44. https://doi.org/10.1016/j.stueduc.2009.01.002
  • Brockx, B., Van Roy, K., & Mortelmans, D. (2012). The student as a commentator: Students'' comments in student evaluations of teaching. Procedia — Social and Behavioral Sciences, 69, 1122–1133. https://doi.org/10.1016/j.sbspro.2012.12.042
  • Stupans, I., McGuren, T., & Babey, A. M. (2016). Student evaluation of teaching: A study exploring student rating instrument free-form text comments. Innovative Higher Education, 41(1), 33–42. https://doi.org/10.1007/s10755-015-9328-5
  • Hujala, M., Knutas, A., Hynninen, T., & Arminen, H. (2020). Improving the quality of teaching by utilising written student feedback: A streamlined process. Computers & Education, 157, 103965. https://doi.org/10.1016/j.compedu.2020.103965

Related articles

analysis-reporting

Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You

Natural-language processing can turn thousands of free-text course-evaluation comments into themes and sentiment at scale — but the research (Cunningham-Nelson 2019; Sunar & Khalid 2023) shows where automated analysis is reliable and where human judgement is still required.

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

analysis-reporting

How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis

When you turn thousands of free-text comments into themes and counts, how do you know the coding is trustworthy? O Connor and Joffe (2020) on intercoder reliability, Braun and Clarke on thematic analysis, and what rigorous qualitative QA looks like.

evaluation-design

How to Design Open-Ended Course Evaluation Questions That Get Useful Answers

Open-ended evaluation questions usually fail not because students have nothing to say, but because the prompt asks for too little. The survey-methodology evidence on answer-box size, verbal instructions, and specific framing — and what it means for the comment box.