Should Every Evaluation Question Count Equally? Composite Scores and the Politics of Weighting
Most course-evaluation "overall scores" quietly average a dozen items with equal weight. Some committees then argue for weighting "clarity" more than "workload". The measurement literature has a surprising, humbling answer to who is right — and it is not the committee.
Koji Education Team
Product ·
The short answer: When a university collapses ten or twelve evaluation items into a single "overall course score", it is making a weighting decision — usually the invisible one of weighting every item equally. Departments then spend meeting hours arguing that "teaching clarity" should count more than "assessment workload", or that some items should be dropped. Six decades of measurement research deliver an uncomfortable verdict: for the correlated, same-direction items typical of course evaluation, how you weight them barely changes the result. The composite is robust to weighting — which means the weighting debate is mostly theatre, and the real decisions lie elsewhere: which items you include, whether a single composite should exist at all, and what you do with the open-text feedback a number can never carry.
The hidden decision inside every "overall score"
Almost every evaluation dashboard reports an aggregate: an overall satisfaction figure, a "teaching quality index", a department mean rolled up from item means. That aggregate is a linear composite — a weighted sum of the underlying items. When nobody specifies the weights, the default is unit (equal) weighting: every item contributes the same. This feels neutral. It is not neutral; it is a specific modelling choice that happens to be invisible because no one wrote it down.
Once people notice the choice, they want to optimise it. Surely "the instructor explained concepts clearly" should count for more than "the room was suitable"? Surely we can find the right weights — perhaps by regressing items onto overall satisfaction, or by asking a committee to assign importance? This intuition is where the methodology gets interesting, because the answer from decades of work is: for data shaped like course evaluations, the search for optimal weights is largely wasted effort.
"It don't make no nevermind": the robustness of equal weights
The canonical reference is Howard Wainer's (1976) paper in Psychological Bulletin (83[2]), memorably titled Estimating coefficients in linear models: It don't make no nevermind. Wainer showed that across a wide range of realistic conditions, replacing statistically optimal regression weights with simple equal weights costs almost nothing in predictive accuracy. Robyn Dawes made the same case famous in The robust beauty of improper linear models in decision making (American Psychologist, 1979, 34[7]), demonstrating that equal-weight ("improper") models often match or beat expert-tuned and even regression-optimal models out of sample. Einhorn and Hogarth (1975) had already laid the groundwork in Unit weighting schemes for decision making.
Why does this hold? Two reasons that both apply forcefully to course evaluation:
-
The items are positively and strongly intercorrelated. When predictors point in the same direction and correlate with each other, their weighted sums converge: shifting weight from one to another moves the total very little, because the items are partly measuring the same thing. Course-evaluation items are a textbook case — they are saturated with the halo effect, in which a single global impression bleeds across every question. When "clarity", "organisation", "fairness" and "enthusiasm" all load on one underlying "I liked this course" factor, weighting them differently is rearranging deck chairs.
-
Optimal weights are estimated with error and overfit the sample. Regression weights derived from one cohort are tuned to that cohort's noise. Applied to the next cohort, they frequently do worse than equal weights, which have no sampling error because they were never estimated. This is the same overfitting logic behind empirical-Bayes shrinkage for fair instructor comparison: unregularised, sample-specific numbers travel badly.
The practical upshot is liberating and deflating at once. If your department is arguing about whether clarity should be weighted 1.5× or 2×, the honest answer is that the overall composite will look almost identical either way. The argument is not really about measurement; it is about signalling what the department values, and it should be conducted openly as a values conversation, not disguised as a technical one.
So weighting doesn't matter? Not so fast — three things that do
The robustness result is often misread as "weighting never matters." That is too strong. Equal weighting is robust under specific conditions, and the conditions define where the real decisions live.
1. Item selection dominates weighting. What you include in the composite matters far more than how you weight what you included. Add three items about classroom facilities to a "teaching quality" index and you have changed the construct, regardless of weights. Drop the one item that captures assessment and feedback — reliably the lowest-scoring dimension — and the composite flatters the course. The composition of the item set is a substantive claim about what "quality" means. Weighting is a rounding error by comparison.
2. When items genuinely diverge, the composite lies. The robustness of equal weights depends on the items being correlated and same-signed. Where a course is bimodal — loved by some students, disliked by others — or where one dimension truly dissociates from the rest, the average conceals more than it reveals. This is the failure mode we examine in why spread and bimodality matter more than the mean. A single composite, however weighted, is the wrong summary for a divided cohort.
3. A composite may not belong in the decision at all. The deepest question is not the weights but whether a one-number index should drive anything consequential. Rolling everything into a single figure invites exactly the ranking-and-thresholding behaviour that student ratings cannot support. Our critique of Net Promoter Score for course evaluation makes the general point: compressing a multidimensional experience into one number optimises for dashboard tidiness, not for understanding or improvement.
"But we need one number for comparison and accreditation"
This is the strongest counterargument, and it is legitimate. Deans need to triage; accreditation panels ask for summary evidence; you cannot hand a governing body forty items per course. The pull toward a single composite is real.
The answer is not to refuse all aggregation — it is to be honest about what the composite can and cannot do. A single index is defensible as a screening signal: a rough, deliberately coarse flag for "look here". It is indefensible as a verdict: a precise ranking that decides tenure, closes a module, or separates a 4.2 from a 4.4. The equal-weight composite is fit for the first job precisely because it is robust and unpretentious; it fails the second job because no weighting scheme can manufacture precision the underlying data do not contain. The mistake is not building a composite. The mistake is letting a screening tool masquerade as a measurement, and then arguing about its second decimal place.
What Koji does instead of tuning weights
If differential weighting of Likert items buys almost nothing, the productive move is to stop investing in the composite and start investing in the signal underneath it. That is the design premise behind Koji.
- Themes, not just an index. Rather than compressing everything into one weighted number, Koji applies automatic thematic analysis to open-text and conversational responses, so a course is described by what students actually said — which specific things worked and which did not — instead of an aggregate that is robust to weighting precisely because it is insensitive to detail.
- Structured plus open, by design. Koji uses six structured question types (open_ended, scale, single_choice, multiple_choice, ranking and yes_no). The scale items still give you the coarse screening signal; the conversational and ranking items recover the dimension-level and priority information that an equal-weight composite averages away. Its quality scoring flags low-information responses so the composite is not quietly built on straightlined data.
- Divergence made visible. Because Koji reports distributions and themes rather than only a mean, a bimodal course does not disappear into a middling composite — the split that a single weighted number would hide is surfaced for the committee to act on.
Koji does not claim that thematic analysis is a more precise composite; it claims something different — that the right response to the weighting result is to stop pretending the number carries the information, and to read the feedback that does. The same conversational engine powers the main Koji platform for product and customer research, where teams face the identical temptation to collapse rich feedback into one satisfaction score.
The bottom line
Equal weighting of correlated evaluation items is not a lazy default to be optimised away — it is, for this kind of data, close to the best you can do, and the evidence for that is fifty years deep. The energy your committee spends debating weights is better spent on the decisions that actually move the result: which items belong in the composite, whether a divided cohort should be summarised by a single figure at all, and whether a screening number is quietly being used as a verdict. Weighting is where the argument feels rigorous. It is not where the rigor is.