Programme Review & Revalidation: Turning Course Evaluation into Accreditation Evidence
A practical buyer's guide to using course-evaluation evidence in periodic programme review and revalidation - mapping ESG 2015 Standard 1.9 and panel expectations to concrete, audit-ready outputs.
Koji Education Team
Product
Answer first: Periodic programme review and revalidation are where course-evaluation data either earns its keep or exposes a gap. Review panels and external examiners do not want a folder of term-by-term rating averages; they want evidence that the programme monitors the student experience, interprets it, acts on it, and checks whether the action worked. This is the "you said, we did, here is the result" cycle - and most institutions can produce the first half and not the second. This guide maps the requirements of programme review and revalidation (anchored in the European Standards and Guidelines, ESG 2015, Standard 1.9 on ongoing monitoring and periodic review) to the specific outputs an evaluation system needs to generate, and shows where AI-moderated evaluation closes the loop that static surveys leave open.
It is written for QA directors, heads of teaching and learning, programme leaders, and institutional-research teams preparing for periodic review, revalidation, or a programme-level audit - regardless of which national agency (NVAO, QAA, HCERES, ANECA, ANVUR, NOKUT, AQ Austria, and others) governs the exercise.
What review panels actually look for
Periodic review and revalidation are not one-off surveys; they are a test of whether your quality cycle functions. Across the ESG and national frameworks the recurring expectations are remarkably consistent:
- Systematic monitoring - the programme routinely collects student feedback, not just when something goes wrong.
- Genuine analysis - the feedback is interpreted (themes, causes, affected cohorts), not merely tabulated.
- Action and ownership - decisions are taken in response, with a named owner and a date.
- Closing the loop - the impact of those actions is checked in a later cycle, and communicated back to students.
- Longitudinal view - trends are visible across cohorts and years, so the panel can see direction of travel, not a single snapshot.
- Triangulation - student feedback sits alongside other evidence (outcomes data, external examiner reports, employability).
ESG 2015 Standard 1.9 frames this directly: institutions should monitor and periodically review their programmes to ensure they achieve their objectives and respond to the needs of students and society, leading to continuous improvement. Standard 1.3 (student-centred learning) and the expectation that students are treated as partners reinforce that their voice must be both collected and visibly acted upon.
The requirement-to-output map
The practical question for procurement is: what must an evaluation tool actually produce to satisfy each expectation? Use this as a checklist when shortlisting.
| Review / ESG expectation | Evidence a panel expects | Concrete output your evaluation system must generate |
|---|---|---|
| Systematic monitoring (1.9) | Routine, programme-wide collection across modules and cohorts | Standardised evaluation run every term with consistent core questions and high coverage |
| Genuine analysis | Themes and causes, not just averages | Thematic analysis with representative quotes, segmented by module/cohort |
| Student-centred response (1.3) | Evidence the student voice shaped decisions | Documented actions linked to specific feedback themes, with owners |
| Closing the loop (1.9) | Proof that actions had an effect | Before/after comparison of the same theme in the next cycle + record of what was communicated to students |
| Longitudinal view | Multi-year trend, not a snapshot | Cohort-level longitudinal reporting across academic years |
| Triangulation | Feedback contextualised against other data | Exportable, structured evidence that can sit alongside outcomes and examiner reports |
| Auditability | Consistent, comparable, defensible method | Standardised, bias-aware collection method applied identically across the programme |
Where the traditional model breaks down
Most institutions run a Likert-scale Student Evaluation of Teaching (SET) each term and can hand a panel a stack of rating averages and a comment export. That satisfies "systematic monitoring" and partially satisfies "longitudinal view" - and then stalls.
- Analysis is thin. A free-text comment box collects opinions but never asks a follow-up. "Assessment was unfair" arrives with no detail; a panel cannot tell what was unfair, for whom, or what would fix it. Hand-coding hundreds of comments is slow and inconsistent between coders, which itself undermines auditability.
- Closing the loop is undocumented. Even where a programme did act on feedback, the link between this comment -> that decision -> this later improvement usually lives in a committee minute, not in a structured, retrievable record. When the panel asks for evidence of impact, the trail is cold.
- The student-as-partner expectation is unmet. If students never see that their feedback changed anything, response rates fall and the "we did" half of the cycle is invisible to the very people the cycle is meant to serve.
The result is a common failure pattern: strong on collection, weak on interpretation and impact - exactly the half that distinguishes a passing review from a commended one.
How AI-moderated evaluation generates review-ready evidence
Koji for Education is designed around the parts of the cycle that static surveys leave open. Instead of a fixed questionnaire, each student completes a short AI-moderated conversational interview that asks the standard questions and then probes - "you mentioned the workload spiked; which weeks, and what was driving it?" - using bias-aware, standardised moderation so every student is asked consistently and the data stays comparable across modules, cohorts, and years. That directly addresses the analysis and auditability gaps:
- Standardised evidence. Because the moderation method is identical for every interview, the resulting dataset is consistent and defensible - the kind of methodological consistency a panel can trust.
- Automatic thematic analysis. Koji clusters the interviews into themes with representative quotes and prevalence, segmented by module and cohort, so analysis is produced in hours, not weeks, and without inter-coder drift.
- Closing-the-loop action tracking. Themes can be linked to actions with named owners and dates, and the next cycle compares the same theme over time - turning "you said, we did, here is the result" from an aspiration into a stored, exportable record.
- Longitudinal cohort reporting. Trends are visible across academic years at programme level, giving the panel direction of travel rather than a single snapshot.
- EU/GDPR-first data handling. For European institutions, the data-protection posture is built for EU/EEA requirements, which matters when student feedback is part of a formal audit trail.
The same AI interview engine powers the main Koji platform for general customer and user research, so the conversational methodology is mature well beyond higher education.
A practical workflow for your next review
- Standardise the core instrument across all modules in the programme so the dataset is comparable - keep a small fixed core, allow module-specific additions.
- Run every term to build the longitudinal record a panel expects, rather than scrambling a one-off survey before the review.
- Convert themes to a small number of owned actions each cycle - a panel is more impressed by three closed loops than thirty open observations.
- Record the "we did" and tell students - publish a short response so the student-as-partner expectation is visibly met.
- In the next cycle, re-measure the same themes and capture the before/after as your impact evidence.
- Export a structured evidence pack that triangulates feedback themes, actions, impact, and longitudinal trends alongside your outcomes and external-examiner data.
Common pitfalls that cost marks at review
A few patterns recur in panel reports across agencies, and all are avoidable. First, treating the survey as the deliverable rather than the decisions it should drive - a wall of rating averages with no narrative of what changed reads as compliance theatre. Second, inconsistent instruments across modules, which makes cross-module and cross-cohort comparison impossible and quietly undermines the credibility of every trend you present. Third, a one-off pre-review scramble instead of a standing termly rhythm, which leaves you without the longitudinal record a panel expects. Fourth, invisible loop-closing - acting on feedback but never telling students, so response rates erode and the partnership expectation goes unmet. Building the evidence as a by-product of a routine, standardised cycle - rather than reconstructing it under deadline - is the single biggest determinant of how smoothly a review goes.
When a simpler tool is still the right call
Honesty serves this audience: you do not always need conversational evaluation. If your programme review only requires proof that feedback is collected, if you are mid-contract on a suite that already feeds your assessment reporting, or if your institution mandates a specific legacy instrument for comparability with historic data, a well-run traditional SET platform may be sufficient for the monitoring requirement. The case for AI-moderated evaluation is strongest when your reviews keep flagging weak qualitative insight, undocumented impact, or an inability to show that the student voice changed anything - the "closing the loop" findings that recur in panel reports.
Related resources
- Course evaluation evidence for NVAO accreditation
- Turning student feedback into ESG accreditation evidence
- Course evaluation evidence for UK TEF and QAA
- Course evaluation evidence for AACSB and EQUIS
- Course evaluation evidence for joint programmes (European Approach)
Bottom line
Programme review and revalidation reward institutions that can show a working quality cycle, not just a survey habit. The differentiator is the second half of the loop - interpretation, owned action, and proven impact over time. Map each review expectation to a concrete output, run a standardised instrument every term, and choose an evaluation method that produces analysis and impact evidence rather than rating averages. That is precisely the gap AI-moderated, conversational evaluation is built to close.
Related articles
Course Evaluation Evidence for the European Approach (Joint Programmes Accreditation)
How to turn student feedback and course evaluation into accreditation-ready evidence for the European Approach for Quality Assurance of Joint Programmes — with a standard-by-standard mapping to concrete outputs.
Turning Course Evaluations into NVAO Accreditation Evidence (Netherlands & Flanders)
A buyer-and-practitioner guide to NVAO accreditation: how course-evaluation evidence maps to the standards, why closing the loop is the hard part, and where an AI-moderated approach helps — with an honest note on when a traditional survey tool suffices.
Course Evaluation Evidence for the UK TEF and QAA Quality Review
A practical buyer''s guide for UK universities: how to turn course and module evaluation into defensible evidence for the OfS Teaching Excellence Framework (TEF), the OfS B-conditions, and the QAA UK Quality Code - mapping each requirement to concrete, accreditation-ready outputs.
Turning Student Feedback into ESG / ENQA Accreditation Evidence
A buyer's guide mapping the ESG 2015 internal quality assurance standards to concrete, accreditation-ready evidence you can generate from student feedback — and how AI-moderated evaluation closes the loop.