How to Write Better Course Evaluation Questions: An Evidence-Based Guide
Most course-evaluation forms are quietly broken before a single student answers them — double-barrelled items, leading wording, vague referents, and scales nobody calibrated. Here is what survey methodology actually says about writing questions that produce usable evidence.
Koji Education Team
Product · June 11, 2026
The short answer: A course evaluation is only as good as its questions, and most institutional forms violate basic survey-design rules that have been settled for decades. The five most damaging mistakes are double-barrelled items (asking two things at once), leading or loaded wording, vague referents ("the course" — which part?), unlabelled or mis-calibrated response scales, and asking students to judge things they cannot validly assess. Fix those and you get evidence you can act on; ignore them and no amount of sophisticated analysis downstream can rescue the data. This guide walks through each, with the methodological sources behind it.
Why question wording is the part everyone skips
Universities pour effort into response rates, dashboards, and benchmarking — and almost none into the wording of the questions themselves. That is backwards. In survey methodology, measurement error introduced at the question-design stage is unrecoverable: if an item is ambiguous, every clever statistic computed from it inherits the ambiguity. Jon Krosnick, one of the most cited survey methodologists, frames good questionnaire design as the foundational task precisely because downstream analysis cannot undo a badly asked question (Krosnick, "Question and Questionnaire Design", Stanford). Below are the errors that do the most damage in course evaluation specifically.
Mistake 1: Double-barrelled questions
A double-barrelled item asks about two constructs but allows only one answer. "The instructor was knowledgeable and approachable" is the classic offender: a student who found the lecturer brilliant but cold has no honest way to respond. Whatever they pick, you cannot tell which half of the question it refers to (Sage Encyclopedia of Survey Research Methods). The repair is mechanical — split it into two items: "The instructor was knowledgeable" and "The instructor was approachable." Audit your existing form for the word "and" in any item; it is the single most reliable double-barrel detector.
Mistake 2: Leading and loaded wording
Leading questions nudge respondents toward an answer through phrasing, and they quietly manufacture the result the author expected (Qualtrics, double-barrelled and leading questions). "How much did this engaging course improve your skills?" presupposes the course was engaging and that skills improved. The neutral version asks the respondent to supply the judgement, not confirm yours: "How would you describe how engaging this course was?" Loaded wording also hides in agreement scales — a string of "The instructor was excellent at…" items invites acquiescence, the documented tendency to agree regardless of content, which we cover in our desk research on acquiescence bias and reverse-worded items.
Mistake 3: Vague referents and the "halo" trap
"Rate the course" is not a question; it is an invitation to rate a global feeling. Students cannot mentally disaggregate "the course" into curriculum, workload, materials, assessment, and teaching, so they collapse it into one overall impression — the halo effect — and answer every sub-item the same way. Precise referents fix this: ask about the assessment feedback, the pace of the lectures, the usefulness of the readings as distinct objects. The more concrete the referent, the more diagnostic the answer, and the less a single strong feeling contaminates everything.
Mistake 4: Response scales nobody calibrated
Three scale decisions routinely go wrong. First, number of points: too few collapses real distinctions, too many exceeds what respondents can meaningfully discriminate — the evidence favours roughly five to seven labelled points, which we unpack in the optimal number of rating-scale points. Second, labelling: every point should carry a verbal label, not just the endpoints, because unlabelled midpoints are interpreted inconsistently across students. Third, the "not applicable / no opinion" option: omitting it forces students with no basis to answer into guessing, manufacturing noise. And remember that the resulting numbers are ordinal — a point we make at length in why averaging Likert scores misleads. "4.2 out of 5" implies a precision the scale cannot support.
Mistake 5: Asking students to judge what they cannot assess
Students are expert witnesses to their own experience — whether they felt lost, whether feedback arrived in time, whether the workload was manageable. They are not in a position to judge an instructor's subject mastery, the currency of the syllabus, or whether assessment standards were rigorous. Items like "The instructor was an expert in the field" measure perception, not expertise, and perception is exactly where bias enters. The methodological discipline is to ask students only about things they directly observed, and to triangulate everything else with peer review and outcome data — see triangulation in teaching evaluation.
A short checklist before you launch
- Scan every item for "and"/"or" — split anything double-barrelled.
- Remove presuppositions; let the student supply the judgement.
- Make each item's referent concrete and singular.
- Label every scale point; include a genuine "not applicable" option.
- Cut any item that asks students to assess expertise rather than experience.
- Keep the form short enough to finish before fatigue sets in.
- Pilot it on a handful of students and ask them what each question meant to them — divergence reveals ambiguity no expert review will.
Where conversational evaluation changes the game
Even a perfectly worded static form shares one limitation: it cannot ask a follow-up. When a student rates "pace" poorly, the form cannot ask which part was too fast or why — so you are left interpreting a number. This is where AI-moderated conversational evaluation earns its place. Koji for Education combines six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) — so the calibration rules above still apply — with an AI moderator that probes beyond the rating: "You said the pace was an issue — can you give a specific example?" Every student gets the same evidence-seeking follow-up, removing the inconsistency of human facilitators, and the open-text that results is run through automatic thematic analysis tied back to verbatim comments. The same conversational engine powers the main Koji platform for customer and user research, where identical question-design rules apply.
Good questions remain the foundation: Koji does not rescue a leading or double-barrelled item, it executes well-designed ones and then goes further than a paper form ever could. Write the questions properly first — then let the instrument do more than record a score.
A worked example — and the bilingual trap
It helps to see a repair in full. Take a common institutional item: "The lecturer was well-prepared and explained difficult concepts clearly, making the subject enjoyable." This single line is triple-barrelled (preparation, clarity, enjoyment), leads toward a positive answer ("making the subject enjoyable" presupposes it was), and uses a vague global referent. A student who found the lecturer meticulous but the subject dull cannot answer honestly. The disciplined rewrite is three neutral, single-construct items, each with a concrete referent: "The lectures were well-prepared." / "Difficult concepts were explained in a way I could follow." / "How would you describe your level of interest in the subject?" Three clean signals replace one muddy one — and the third no longer assumes the answer.
There is one further trap that disproportionately affects European institutions with international cohorts: translation and reference-group effects. An item that is unambiguous in English can become double-barrelled or leading once translated, and a phrase like "the workload was reasonable" is interpreted against wildly different baselines by students from different educational cultures. If you run a form in multiple languages, each version needs its own design review, not just a literal translation — and you should expect that the same numeric scale does not mean the same thing across cohorts, a problem we examine in language bias and international students. The safest items for diverse cohorts are concrete and behavioural ("I received feedback on my work within two weeks") rather than evaluative and abstract ("the feedback was good"), because concrete referents survive translation and cross-cultural interpretation far better.
The bottom line
Most course-evaluation data is compromised at the moment of question-writing, not at analysis. The fixes are old, settled, and almost free: one idea per question, neutral wording, concrete referents, calibrated and fully-labelled scales, and humility about what students can validly judge. Get those right and even a simple form yields evidence worth acting on. Get them wrong and you have built a very precise way to measure noise.