Stop Reinventing the Questionnaire: What Validated Instruments Like SEEQ and IDEA Offer Course Evaluation
Most universities write their own evaluation form from scratch and discard decades of published validity evidence. SEEQ and IDEA already have it. Here is what a validated instrument buys you, where it still fails, and how to keep the rigour while probing beyond the number.
Koji for Education
Research & Editorial Team · July 11, 2026
Answer first: Most European universities design their course-evaluation questionnaire in-house, item by item, in a committee. Yet two instruments — Herbert Marsh''s SEEQ (Students'' Evaluation of Educational Quality) and the IDEA Student Ratings of Instruction — have decades of published evidence for their factor structure, reliability, and validity. Building your own form from scratch quietly throws that evidence away and replaces it with questions nobody has ever tested. The defensible options are to adopt a validated instrument, or to borrow its validated items as your scaffold — and then use richer methods to capture what any fixed questionnaire cannot.
The invisible default: a form nobody validated
Ask a quality-assurance office where its evaluation questions came from and the honest answer is usually "they evolved." Someone adapted a neighbouring faculty''s form a decade ago; items were added after a bad NSS year and never removed. The result is a questionnaire with no documented reliability, no tested dimensional structure, and no evidence that its items measure what their wording implies. Every downstream decision — flagging a module, comparing instructors, feeding a programme review — inherits that unmeasured uncertainty.
This is strange, because course evaluation is one of the few areas of higher education where rigorously validated, freely documented instruments already exist. We would not build our own clinical depression scale rather than use a validated one. We routinely build our own teaching-quality scale.
What "validated" actually means
SEEQ, developed by Marsh in the early 1980s, measures nine distinct dimensions of teaching: learning value, instructor enthusiasm, organisation, individual rapport, group interaction, breadth of coverage, examinations and grading, assignments and readings, and workload/difficulty. Its subscale internal-consistency reliabilities (Cronbach''s alpha) run from roughly 0.88 to 0.97, and the nine-factor structure has been recovered in separate factor analyses of nearly 5,000 classes across disciplines and levels (Marsh, 1982, British Journal of Educational Psychology; ERIC ED338629). Crucially, SEEQ ratings have been validated against external criteria — student learning on objective examinations, retrospective ratings by former students, and instructors'' own self-evaluations — not just against themselves. The structure has since replicated in Spain, China, Greece, Oman and beyond (SEEQ in Oman higher education, 2024).
IDEA, from a US non-profit centre operating since 1975, takes a different and instructive design decision: it evaluates teaching by student progress on the learning objectives the instructor selected as important for that course, rather than against a fixed template of what "good teaching" looks like. Its diagnostic form pairs 20 teaching methods with 12 learning objectives and statistically adjusts scores for course circumstances such as class size and student motivation. IDEA reports high class- and instructor-level reliability and internal consistency, and grounds validity in correlations with learning, multidimensional structure, and analysis of response processes (IDEA Student Ratings of Instruction).
Two features are worth stealing even if you adopt neither wholesale: SEEQ''s insistence that teaching is multidimensional (nine factors, not one "overall" number), and IDEA''s insistence that a course should be judged against its own objectives, with adjustment for the circumstances outside the instructor''s control.
Why borrowing beats building
A validated instrument gives you three things a home-grown form cannot:
- A defensible measurement model. You can state, with citation, that your "organisation" items form a coherent factor and cohere at a known reliability. When a head of department disputes a score, you are standing on published psychometrics, not on a committee''s intuition.
- Benchmarks that mean something. Because validated instruments are administered in standardised ways across many institutions, comparison against a discipline or sector norm is meaningful rather than an artefact of your local wording.
- Discriminant structure. Nine tested dimensions let you tell an instructor what to improve. A single home-grown "satisfaction" item cannot, because it does not separate organisation from rapport from assessment design — the jingle-jangle problem in action.
"But our context is unique" — the strongest counterargument
Sceptical readers — rightly — raise objections. Let us take the three strongest.
"These instruments are old and American; they will not fit a European, Bologna-aligned programme." Partly fair. SEEQ predates the European Standards and Guidelines and says nothing about learning-outcome alignment or graduate competences. But "old" is not the same as "invalid": the factor structure has replicated across continents for forty years, which is precisely the evidence a freshly-invented local form lacks. The right move is not to reject validated items but to supplement them with the outcome- and competence-facing questions your quality framework requires.
"Our courses are distinctive — labs, studios, clinics, placements." A genuine limitation. A general instrument built around lectures fits signature pedagogies poorly. But this argues for a validated core plus tested local modules, not for abandoning validation entirely.
"Validated instruments still measure student perception, so they inherit all the biases." Correct, and important. A validated instrument standardises the number; it does not make the number bias-free. SEEQ and IDEA scores are still susceptible to grading-leniency, halo, and the construct-irrelevant variance that plagues all student ratings. Validation buys you a trustworthy ruler; it does not tell you the thing you are measuring is the thing you care about.
What even the best instrument cannot do
This is the honest limit. A validated Likert instrument tells you that students rated organisation 3.4. It cannot tell you why, for whom, or what to change on Monday. It compresses a term of experience into a number, discards the reasoning, and leaves the open-comment box — the part faculty actually read — entirely unvalidated. Reliability of the closed items says nothing about the quality of the free text. And because every respondent answers the same fixed items in the same order, the instrument cannot follow up on the one answer that mattered. Validation solves the measurement problem. It does not solve the understanding problem.
Where Koji fits
Koji for Education is built to keep the rigour and add the understanding. You can seed a Koji evaluation with validated, multidimensional items — using the six structured question types (open-ended, scale, single-choice, multiple-choice, ranking, yes/no) so that a SEEQ-style scale item sits next to the probing it deserves. Where a static form stops at "3.4 on organisation," Koji''s AI-moderated conversational interview asks the student which part of the course felt disorganised and what would have helped — the same standardised, bias-aware follow-up for every student, with none of the inconsistency of a human interviewer. Its automatic thematic analysis then turns thousands of open responses into structured, comparable evidence, and quality scoring flags low-information answers so a committee is not misled by a confident but empty comment.
In other words: use validated instruments for the parts they do well — a trustworthy, comparable measurement backbone — and let a conversational layer capture the reasoning they were never designed to hold. The same AI interview engine underpins general user and customer research on the main Koji platform, if your institution also runs studies beyond the classroom.
Validation is not a luxury for the psychometrically fastidious. It is the difference between a score you can defend in a programme review and a number you are quietly hoping nobody questions. Start from the instruments that already earned their evidence — then go and find out what the number could never tell you.
Adopting validated items without ripping everything out
The objection that stops most offices is practical, not theoretical: "we cannot break a decade of trend data by swapping the whole instrument overnight." You do not have to. The pragmatic migration path is to run a validated core alongside your legacy items for a transition period, correlate the two, and retire the home-grown questions only once you can see what the validated versions add. Map each legacy item to the SEEQ or IDEA dimension it was gesturing at, keep the small number of local items that carry genuine institutional meaning (your own learning-outcome or graduate-attribute questions), and drop the redundant ones that were never measuring a distinct construct. Within two or three cycles you have a shorter, tested instrument, a documented crosswalk that preserves trend continuity, and — for the first time — the ability to say exactly which dimension a low score belongs to. That is not a bigger project than the annual "let us tweak the form" ritual most committees already run; it is the same effort, pointed at evidence instead of intuition.
See how Koji for Education pairs validated question design with AI conversational interviews — explore Koji for Education.