New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

Comparing Instructors on Many Criteria at Once: Data Envelopment Analysis for Course Evaluation

Data Envelopment Analysis (DEA) rates instructors, modules, or departments against the best observed performers on several inputs and outputs at once — a fair multi-criteria alternative to ranking on one mean.

Koji Education Team

Product

In brief

Data Envelopment Analysis (DEA) is a linear-programming method from operations research that rates each unit — here an instructor, module, or department — against the best performers actually observed, using several inputs and several outputs at once, without forcing you to pre-assign weights. Applied to course evaluation, it answers a question a single mean cannot: given the class size, workload, and prior ability a teacher was handed, and the several outcomes they produced (ratings, pass rates, engagement), are they on the efficient frontier or below it? DEA is powerful for fair multi-criteria comparison, but it is deterministic, sensitive to outliers and measurement error, and can flatter a unit that is merely unusual — so it belongs alongside, not instead of, the uncertainty-aware methods you already use.

What the research says

DEA was introduced by Charnes, Cooper and Rhodes in 1978, building on Farrell to measure the relative efficiency of "decision-making units" (DMUs) that convert multiple inputs into multiple outputs (Charnes, Cooper & Rhodes, 1978). Their model — now called CCR — assumes constant returns to scale and, for each unit, solves a linear program that chooses the weights most favourable to that unit, subject to the constraint that no unit can exceed 100% efficiency under those same weights. A unit scores 1.0 (efficient) if no combination of other units produces at least as much of every output using no more of every input; otherwise it scores below 1.0, and the peers that dominate it are named explicitly.

Banker, Charnes and Cooper (1984) extended the model to variable returns to scale — the BCC model — separating pure technical efficiency from scale efficiency, which matters when small and large classes are not simply scaled copies of one another. Cook and Seiford (2009), reviewing thirty years of work, document how DEA became one of the most widely used tools in operational research, with thousands of applications and refinements (super-efficiency, cross-efficiency, network DEA).

DEA is relevant to higher education because universities are textbook multi-input, multi-output organisations, and this was recognised early. Jill Johnes (2006) applied DEA to over 100 UK economics departments and to individual graduates, arguing that its chief attraction is precisely that it handles multiple inputs and outputs without requiring prices or a pre-specified production function — a good fit for education, where outputs (learning, satisfaction, completion, research) have no common currency. She also stressed the method's fragility: results are sensitive to the choice of inputs and outputs and to sample composition.

Why it matters for course evaluation in practice

Almost every fair-comparison method in a course-evaluation toolkit reduces teaching to one number and then argues about how to adjust it. DEA takes the opposite stance: keep several outcomes, and credit a teacher for whichever mix they are genuinely strong on. An instructor with modest mean ratings but excellent pass rates and high open-text engagement, teaching a large required quantitative module, may sit on the efficient frontier even though a raw ranking of mean scores buries them.

Three practical uses stand out. First, benchmarking with peers you can name: because DEA returns the specific efficient units that dominate an inefficient one, a chair can say "these three modules achieve more on every measure with the same resources," which is far more actionable than a percentile. Second, resource-aware fairness: by entering class size, contact hours, or student prior attainment as inputs, DEA asks what was produced relative to what was available, addressing the same confounds that class-size and difficulty adjustments target. Third, target-setting: projecting an inefficient unit onto the frontier suggests concrete, feasible improvement targets rather than an abstract "raise your score."

DEA sits naturally next to the fair-comparison methods you may already use — funnel plots for sampling noise, empirical-Bayes shrinkage for small classes, and multi-criteria weighting via the Analytic Hierarchy Process. Where those adjust or combine a single signal, DEA keeps the signals separate and locates the frontier.

Limitations and honest caveats

A critical reader will raise several objections, and they are right to. DEA is deterministic: it has no error term, so every deviation from the frontier is attributed to inefficiency rather than to luck or measurement error. A single anomalous unit can define the frontier and make everyone else look inefficient — the method is notoriously outlier-sensitive. Efficiency scores depend heavily on the chosen input and output set; add a variable on which a weak unit excels and it can jump to the frontier. When inputs and outputs are many relative to the number of units, almost everyone scores 1.0 (the curse of dimensionality), so a rough rule of thumb asks for at least two or three times as many DMUs as inputs plus outputs. Because each unit is scored with its own most-favourable weights, two efficient instructors may be efficient on entirely different things, so 1.0 is not evidence of all-round excellence.

Dyson and colleagues (2001) catalogue these and other pitfalls — mixing incommensurable measures, using index numbers and ratios as inputs, including correlated variables — and propose protocols to avoid them. The prudent course is to treat DEA scores as a screening and dialogue tool, run sensitivity analyses over the variable set, bootstrap the scores to attach uncertainty (the Simar–Wilson bootstrap is standard), and never convert a raw efficiency score into a high-stakes personnel decision on its own. DEA also cannot tell you why a unit is inefficient, and it says nothing about whether the outputs are valid measures of teaching — a course-evaluation-specific worry given how weakly student ratings track learning.

How Koji incorporates this

Koji is built around the idea that teaching produces several things worth measuring, which is exactly the data structure DEA needs. Because Koji's AI-moderated conversational interviews capture structured outcomes (scale, single_choice, ranking, yes_no) alongside automatically themed open text, an evaluation record is naturally multi-output rather than a lone average. Koji's reporting layer can assemble the inputs a fair comparison requires — class size, modality, level — from study metadata, and present instructors as a portfolio of outcomes rather than a single rank.

Concretely, Koji is designed to support frontier-style benchmarking that names peer units rather than publishing a league table, to attach uncertainty to any comparative score so a below-frontier result is not read as a verdict, and to surface which outcomes drive a unit's standing so a target is specific. Koji frames these as decision-support signals for a human committee, consistent with its bias-aware reporting stance: the platform is designed to mitigate the temptation to rank on one number, not to automate a ranking. For teams doing multi-criteria efficiency work outside teaching, Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where multi-output trade-offs are just as common.

Frequently asked questions

Is DEA better than just adjusting the mean score for class size?

They answer different questions. Statistical adjustment (as in the IDEA approach) removes the influence of a confound from a single outcome; DEA keeps several outcomes and credits a unit for its best mix relative to what it was given. Use adjustment when you care about one clean score; use DEA when teaching legitimately has multiple, non-substitutable outputs.

How many instructors do I need before DEA is trustworthy?

Because DEA can label almost everyone efficient when the model is too rich, keep inputs plus outputs small relative to the number of units — a common rule of thumb is at least two to three units per variable. With four inputs and outputs you would want roughly 15 or more comparable units before scores mean much.

Does an efficiency score of 1.0 mean an instructor is excellent?

No. It means no observed combination of peers dominates them on every measure under their own most-favourable weights. Two efficient instructors can be strong on completely different outcomes, so 1.0 is "not dominated," not "best all round."

Can DEA handle the fact that student ratings are noisy?

Not on its own — standard DEA is deterministic and treats every shortfall as inefficiency. Bootstrapping the efficiency scores attaches confidence intervals, and pairing DEA with empirical-Bayes shrinkage or funnel plots keeps sampling noise from being misread as poor teaching.

Is DEA a causal method?

No. It measures relative efficiency among observed units at one time; it does not establish that a teaching practice caused an outcome. For causal claims you need designs such as interrupted time series or the other quasi-experimental methods documented in this knowledge base.

Related Resources

References