New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

The Delphi Method for Course and Programme Evaluation: Building Expert Consensus Without the Loudest Voice Winning

How the Delphi technique produces structured, anonymous expert consensus on course-evaluation criteria and programme standards — its evidence base, its limitations, and where it fits alongside student feedback.

Koji Education Team

Product

In brief. The Delphi method is a structured, multi-round, anonymous process for turning divergent expert opinion into defensible group consensus. In course and programme evaluation it is not a substitute for student feedback — it is the tool you use before you survey anyone, to decide which criteria matter, what "good teaching" means in a discipline, and where the quality thresholds should sit. Used rigorously, it removes the dominance, conformity, and status effects that wreck ordinary committee meetings; used loosely, it manufactures a false consensus. This article summarises the evidence and shows how to run one that survives academic scrutiny.

The question this answers

Most writing about course evaluation assumes you already know what to evaluate. But someone had to decide that the questionnaire should ask about "clarity", "workload", and "assessment feedback" rather than fifteen other things. In a committee, that decision is usually made by whoever is most senior, most confident, or most persistent. The Delphi method exists to make that decision differently: through iterative, anonymous rounds in which a panel of experts rates and re-rates candidate criteria, sees the group''s aggregated response after each round, and revises — converging on consensus without ever meeting face to face.

What the research says

The Delphi technique originated at the RAND Corporation in the 1950s and 60s. The foundational empirical demonstration is Dalkey and Helmer (1963), who showed that structured, iterative questionnaires with controlled feedback produced more stable and defensible group estimates than open discussion. Its use in education is nearly as old: the technique was adopted in higher education to develop criteria for evaluating faculty, precisely the problem of deciding what counts as good teaching.

The defining features, as codified in the methodological literature, are four. Anonymity (panellists never know whose rating they are seeing) neutralises the halo of seniority and the pressure to agree with a chair. Iteration across two to four rounds lets people change their minds privately, without the social cost of a public reversal. Controlled feedback — each round returns the group''s aggregated ratings, often with the distribution and the outlying justifications — gives panellists something to react to. And statistical group response replaces "the meeting decided" with an explicit, reportable measure of agreement.

Hasson, Keeney and McKenna (2000) provide the most-cited practical guidelines: panel selection is the single most important design decision (Delphi has no probability sample, so the credibility of the result rests entirely on who the experts are), rounds should continue until stability rather than to a fixed number, and consensus must be defined in advance — commonly ≥70% or ≥80% agreement on an item — not declared after the fact. Green (2014), reviewing the technique specifically for educational research, stresses that Delphi is appropriate exactly when the question is normative or definitional ("what should a good course achieve?") and when the relevant knowledge is dispersed across people who cannot easily be convened. More recent work, such as the consensus-building case study by educational researchers in International Journal of Research & Method in Education (2024), documents how modern Delphi studies report attrition, agreement thresholds, and item stability transparently so that readers can judge whether the consensus is real.

Applied examples in higher education are abundant and verifiable: Delphi panels have been used to agree the competencies of a digitally fluent educator (a 36-expert panel converging on 14 competencies) and to define consensus-based aims, learning outcomes, and evaluation methods for specific courses. The pattern is consistent — Delphi builds the specification that a downstream evaluation instrument then measures against.

Why it matters for course evaluation in practice

Three practical problems in quality assurance are Delphi-shaped.

1. Deciding what a course-evaluation questionnaire should ask. Rather than importing a generic template or letting one committee member''s hobby-horses dominate, a Delphi panel of programme directors, external examiners, employers, and senior students can converge on the handful of dimensions that genuinely matter for this discipline. Engineering, nursing, and fine art do not share a single definition of teaching quality; a Delphi surfaces that.

2. Setting quality thresholds and programme standards. When an institution needs to decide "what score, on what indicator, should trigger a programme review?", Delphi turns an arbitrary line into a documented expert consensus — valuable evidence when an accreditation panel asks why the threshold is where it is. This connects directly to meta-evaluation of your evaluation system.

3. Reconciling stakeholders who will never sit in the same room. External examiners, industry advisors, alumni, and disability advisors hold relevant expertise but have no shared meeting. Delphi is asynchronous by design, which is also why it scales across institutions and countries.

The method complements, rather than competes with, the group techniques already in your toolkit. Where the Nominal Group Technique convenes people in one room to generate and rank ideas quickly, Delphi trades speed for anonymity, geographic reach, and a documented audit trail — and where students as partners brings learners into the design conversation, a student-inclusive Delphi panel can do the same without a confident minority speaking over everyone else.

Limitations and honest caveats

A PhD reader will raise these, so raise them first.

  • Consensus is not validity. A panel can agree strongly on something that is wrong. Delphi measures agreement among the chosen experts, nothing more. If the panel is unrepresentative, the consensus simply encodes a shared bias. This is the deepest critique of the method and it has no purely internal fix.
  • Panel selection has no sampling theory. There is no margin of error, no confidence interval on "who counts as an expert". Two defensible panels can reach different consensuses. Report your selection criteria explicitly and treat the result as conditional on them.
  • Manufactured consensus and regression to the middle. The feedback mechanism that produces convergence can also pressure dissenters toward the group mean even when their objection is correct. Well-run studies preserve and report minority justifications rather than discarding outliers.
  • Attrition across rounds. Panels shrink between rounds; if the people who drop out are systematically different (busier, more sceptical), the surviving consensus drifts. Report response rates at every round.
  • Definitional slack. "Consensus" is whatever the researcher declared it to be. A study that sets the bar at 51% and one that sets it at 80% are not comparable. Pre-register the threshold.
  • It is slow. Two to four rounds over weeks is the cost of anonymity and reflection. For a fast, in-the-room prioritisation, other methods win.

None of these makes Delphi unusable; they make transparency mandatory. A credible Delphi study reads like a credible survey study — you can see exactly how the number was produced.

How Koji incorporates this

Koji does not run the Delphi rounds for you — that is a governance process owned by your quality office. What Koji does is operationalise the output of a Delphi and close the loop that a static consensus cannot.

  • From consensus to instrument. Once a Delphi panel agrees the dimensions that matter for a programme, those become structured questions in Koji — scale, single_choice, ranking, yes_no, and crucially open_ended items — so the agreed criteria are measured consistently every cycle rather than drifting with whoever last edited the form.
  • AI-moderated conversational interviews for the panel itself. The same interview engine Koji uses with students can administer the qualitative, justification-heavy rounds of a Delphi with experts: an AI moderator probes why a panellist rated an item low, captures the reasoning that Delphi feedback depends on, and does so asynchronously across a distributed, multilingual panel.
  • Automatic thematic analysis of round-by-round free text. Delphi generates large volumes of open justification text between rounds. Koji''s thematic analysis clusters those justifications so facilitators see the emerging fault lines instead of hand-coding hundreds of comments — with a documented audit trail suited to accreditation.
  • Triangulation, not replacement. Delphi sets the standard; Koji then triangulates student feedback, cohort comparisons, and open-text evidence against it, and tracks the closing-the-loop actions. The consensus becomes a living benchmark rather than a filed report.

Koji''s core research platform at koji.so applies the same AI-moderated interview and consensus-analysis engine to product and customer research, where Delphi-style expert panels are common in forecasting and standard-setting — but for course and programme evaluation, the education-specific workflows live here.

The honest framing: Delphi is designed to reduce the dominance and conformity effects of ordinary committees, not eliminate the possibility of a wrong-but-agreed answer. Koji is designed to turn a well-run consensus into a measurable, repeatable evaluation — and to keep the messy, structured work of the rounds themselves auditable.

Related resources

References