New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Illuminative Evaluation: What Parlett and Hamilton's 'Learning Milieu' Adds to Course Feedback

Parlett and Hamilton's 1972 illuminative evaluation reframes course evaluation as illuminating how a course works within its learning milieu, not just scoring objectives. Here is the evidence, its limits, and how it changes QA practice.

Koji Education Team

Product

In brief

Illuminative evaluation, proposed by Malcolm Parlett and David Hamilton in 1972, argues that the job of course evaluation is to illuminate how a course actually operates within its "learning milieu" — the whole network of cultural, social, institutional and pedagogical forces surrounding it — rather than to measure whether it hit a list of pre-specified objectives. For quality assurance, this means treating rich, contextual accounts of the student experience as primary evidence, and treating the Likert average as a starting point for enquiry rather than the conclusion. It is a qualitative, interpretive tradition that predates — and directly anticipates — today's interest in open-text and conversational feedback.

What the research says

Parlett and Hamilton first circulated Evaluation as Illumination: A New Approach to the Study of Innovatory Programs as an Edinburgh occasional paper in 1972; it was widely reprinted, including in Tawney's Curriculum Evaluation Today (1976). Their target was the then-dominant "agricultural-botany" paradigm of evaluation, which treated a course like a crop trial: define objectives, administer a standardised instrument, measure whether outcomes matched the specification, and report a mean. Parlett and Hamilton argued this paradigm systematically misses what actually happens in real classrooms because it ignores the learning milieu — the socially and institutionally negotiated environment in which teaching and learning take place.

The illuminative alternative borrows from social anthropology and naturalistic enquiry. The evaluator studies the programme "in its own terms": observing teaching, interviewing participants, reading documents, and — crucially — using progressive focusing, in which early open-ended observation gradually narrows onto the questions that matter most in this setting rather than questions imposed in advance. The aim is illumination and understanding, not verdict-by-measurement. Two organising concepts recur: the instructional system (the idealised course as designed) and the learning milieu (the course as actually experienced), with the gap between them being the most fertile territory for improvement.

This was part of a broader "qualitative turn" in educational evaluation during the 1970s. It sits alongside Robert Stake's responsive evaluation (1975), which prioritises stakeholders' issues and concerns over the programme's stated intents, and Guba and Lincoln's later Fourth Generation Evaluation (1989), which frames evaluation as a negotiated, constructivist process. All three share a conviction that experimental and psychometric methods, on their own, cannot capture the meaning of an educational experience.

Why does this matter empirically, not just philosophically? Because the quantitative core that illuminative evaluation distrusted has aged badly. Uttl, White and Gonzalez's (2017) re-analysis of decades of multisection studies found that student-evaluation ratings explain at most about 1% of the variance in actual student learning — effectively no relationship once small-study and publication bias are removed. Spooren, Brockx and Mortelmans' (2013) state-of-the-art review reached a more measured but still cautionary conclusion: validity evidence for student ratings is fragmented and context-dependent. If a single number cannot be trusted to represent learning, the illuminative instinct — go and understand what is actually happening — looks less like a relic and more like a corrective.

Why it matters for course evaluation in practice

Most institutional evaluation systems are, in Parlett and Hamilton's terms, still "agricultural-botany." They fix an item bank, compute departmental means, and benchmark instructors against norms. Illuminative evaluation suggests three practical shifts for a modern quality-assurance office:

  1. Evaluate the milieu, not just the teacher. A low score on "the course was well organised" may reflect a timetabling clash, a room change, a mismatched prerequisite, or an assessment bunching problem — none of which is the instructor's doing. Illuminative practice asks why before it ranks who, protecting against the misclassification and unfair personnel comparisons documented elsewhere in this knowledge base.

  2. Let questions emerge (progressive focusing). A fixed questionnaire can only return answers to questions someone thought to ask last year. Open-ended, iteratively focused enquiry surfaces the issue nobody anticipated — the reason a cohort is quietly disengaging, or the one lab session that transformed understanding.

  3. Treat qualitative accounts as evidence, not decoration. In many systems, open-text comments are read anecdotally, if at all, while the mean drives decisions. Illuminative evaluation inverts the hierarchy: the narrative is the evidence, and numbers are context.

Limitations and honest caveats

A PhD reader will raise several objections, and illuminative evaluation's own proponents anticipated most of them.

  • Subjectivity and evaluator bias. Naturalistic, interpretive work depends on the evaluator's judgement, which can smuggle in preconceptions. Progressive focusing can drift into confirmation-seeking. This is precisely the danger that Scriven's goal-free evaluation tries to counter, and it is why triangulation and member-checking are non-negotiable.
  • Generalisability. An illuminative study of one course, richly understood, does not license claims about the programme, the department, or the sector. It trades breadth for depth by design. Combining it with quantitative monitoring — a mixed-methods stance — is usually wiser than choosing one paradigm.
  • Resource intensity. Classical illuminative evaluation is labour-heavy: observation, interviews, and iterative analysis do not scale to every course, every term, across a faculty. Historically this confined it to special reviews rather than routine QA.
  • Reliability and auditability. Accreditors under the ESG framework expect systematic, defensible processes. Purely idiographic accounts can be hard to audit unless the analytic trail (codes, themes, saturation, inter-rater agreement) is documented — see inter-rater reliability in thematic analysis.

Naming these limits is not a retreat; it is the condition for using the approach credibly alongside, not instead of, quantitative monitoring.

How Koji incorporates this

Koji is, in effect, an attempt to make illuminative evaluation routine and scalable — to get the depth of a naturalistic enquiry at the cadence of an ordinary end-of-term cycle. Concretely:

  • AI-moderated conversational interviews replace the fixed questionnaire with a dialogue that probes beyond a Likert number. When a student says a course "felt disorganised," the interviewer follows up — what specifically, and when? — which is progressive focusing operationalised at cohort scale. This is designed to reconstruct the learning milieu, not just score the instructional system.
  • Structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let you keep the comparable numeric backbone accreditors want while the conversational layer captures the contextual "why."
  • Automatic thematic analysis of open-text responses surfaces emergent themes across a cohort, with representative quotes and quality scoring, so narrative evidence becomes systematic and auditable rather than anecdotal. This directly addresses the reliability objection above.
  • Mid-cycle and formative collection supports the iterative spirit of progressive focusing: illuminate early, act, and re-illuminate — rather than waiting for a single terminal snapshot.
  • Triangulation across cohorts and structured reporting help distinguish milieu effects (timetabling, room, assessment load) from teaching effects, mitigating the misattribution illuminative evaluators warned about.

Koji is designed to mitigate, not eliminate, the subjectivity risks Parlett and Hamilton flagged: because the same AI moderator applies a consistent protocol and the thematic layer is transparent, the interpretive trail is reviewable. The same AI-moderated interview engine powers Koji's core research platform at koji.so for product and customer research, where "understand the milieu" is simply called understanding the user.

Putting illumination to work without abandoning your metrics

Adopting illuminative principles does not require dismantling an existing survey system; it requires reordering how evidence is weighted and read. A workable mixed-methods cycle looks like this. Begin the term with a light quantitative baseline so you retain the comparability accreditors expect. At mid-cycle, open a genuinely exploratory phase — unstructured or lightly structured questions that let students describe the course in their own terms — followed by progressive focusing onto the two or three issues that recur. Feed an interim interpretation back to the teaching team while the course is still running, the formative move illuminative evaluation was built for. Close the term with a confirmatory pass that tests whether the mid-cycle changes shifted the experience, combining the numeric backbone with a final round of open accounts.

The governing discipline is that the number never travels alone. A departmental report that shows a 3.6 mean beside a short thematic account of why the assessment felt bunched, and what changed after week six, is illuminative in Parlett and Hamilton's sense even though it still contains a mean. The point was never to ban measurement; it was to stop measurement from masquerading as understanding. Institutions that treat their dashboards as the beginning of an enquiry rather than the end of one are, whether they use the label or not, practising evaluation as illumination.

Related resources

References

  • Parlett, M. & Hamilton, D. (1972/1976). Evaluation as Illumination: A New Approach to the Study of Innovatory Programs. Occasional Paper, Centre for Research in the Educational Sciences, University of Edinburgh; reprinted in D. Tawney (Ed.), Curriculum Evaluation Today: Trends and Implications. London: Macmillan. ERIC ED167634. https://eric.ed.gov/?id=ED167634
  • Stake, R. E. (1975). Evaluating the Arts in Education: A Responsive Approach. Columbus, OH: Merrill.
  • Guba, E. G. & Lincoln, Y. S. (1989). Fourth Generation Evaluation. Newbury Park, CA: Sage.
  • Uttl, B., White, C. A. & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22-42. https://doi.org/10.1016/j.stueduc.2016.08.007
  • Spooren, P., Brockx, B. & Mortelmans, D. (2013). On the validity of student evaluation of teaching: The state of the art. Review of Educational Research, 83(4), 598-642. https://doi.org/10.3102/0034654313496870

Related articles

analysis-reporting

How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis

When you turn thousands of free-text comments into themes and counts, how do you know the coding is trustworthy? O Connor and Joffe (2020) on intercoder reliability, Braun and Clarke on thematic analysis, and what rigorous qualitative QA looks like.

best-practices

Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations

Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.

best-practices

Designing Course Evaluation for Use: The Utilization-Focused Approach

The biggest failure of course evaluation is not bad data — it is data nobody acts on. Patton's Utilization-Focused Evaluation and the empirical research on evaluation use (Johnson et al. 2009) show how to design feedback for action from the start.

research-methods

What Can Open-Text Student Comments Tell You That Likert Scores Cannot?

A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.