New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

Which Aspects of Teaching Actually Drive the Overall Score? Relative Weights and Dominance Analysis

When clarity, workload, support, and feedback are all correlated, regression coefficients cannot tell you which one drives the overall rating. Relative weights and dominance analysis partition the explained variance fairly.

Koji Education Team

Product

In brief

A quality office often wants to know which aspect of teaching — clarity, workload, feedback, support — most drives the overall course rating. The obvious move, reading the standardised regression coefficients, is unreliable because those aspects are highly correlated: multicollinearity makes individual betas unstable, hard to interpret, and occasionally sign-flipped. Relative weights analysis (Johnson 2000) and dominance analysis (Budescu 1993) solve this by fairly partitioning the model's total explained variance among the correlated predictors, so each aspect gets a share of R-squared that sums to the whole. They answer "how much does each item contribute?" in a way raw coefficients cannot.

What the research says

The problem is old and well understood. When predictors are correlated, ordinary regression coefficients capture each variable's unique contribution holding the others constant, which systematically understates the importance of variables that share variance — and with strong multicollinearity the coefficients become unstable, so small data changes swing them wildly or reverse their sign. For course-evaluation items, which are notoriously intercorrelated, this makes the beta-ranking of "drivers" close to meaningless.

David Budescu (1993), in Psychological Bulletin, proposed dominance analysis as a principled fix. A predictor "completely dominates" another if it adds more to R-squared than its rival in every possible subset of the other predictors. By averaging each predictor's incremental R-squared contribution across all 2^(k-1) subset models, dominance analysis produces an importance measure with an intuitive meaning — average usefulness over all contexts — and the components sum to the model R-squared. Azen and Budescu (2003), in Psychological Methods, formalised three levels of dominance (complete, conditional, general) and showed the approach is more informative than comparing standardised coefficients or zero-order correlations.

Jeff Johnson (2000), in Multivariate Behavioral Research, introduced relative weights analysis (RWA), a computationally cheaper route to nearly the same answer. Johnson transforms the correlated predictors into a set of orthogonal (uncorrelated) variables that are maximally related to the originals via an eigen-decomposition of the predictor correlation matrix, regresses the outcome on those orthogonal variables (where importance is unambiguous), and then transforms the results back to the original predictors. The resulting weights are virtually identical to the all-subsets dominance averages but are feasible with many predictors, where dominance analysis becomes combinatorially expensive.

Tonidandel and LeBreton (2011), in the Journal of Business and Psychology, made RWA usable for applied researchers: they provide accessible tools, show how to rescale weights as a percentage of predicted variance, and — crucially — how to compute bootstrap confidence intervals and significance tests for the weights and for differences between weights, so an analyst can say not just that clarity ranks first but whether it is reliably above feedback.

Why it matters for course evaluation in practice

Nearly every institution runs some version of "what should we improve first?" and answers it badly.

Beyond the beta table. A committee handed a regression of overall satisfaction on eight teaching items will read the largest coefficient as the biggest lever. Because the items are correlated, that reading is often wrong, and it changes from year to year purely from multicollinearity noise. Relative weights give each item a stable share of the explained variance — "clarity accounts for 28% of the predicted variance in the overall score, feedback 9%" — that adds up to a coherent whole and survives resampling.

A fair complement to importance-performance analysis. Importance-performance maps typically use stated importance or a simple correlation as the importance axis. Relative weights supply a derived, multicollinearity-corrected importance that is a defensible replacement for that axis, sharpening the priority map without asking students to self-report what matters.

Honest prioritisation. With bootstrap intervals, a quality office can distinguish "these three drivers are statistically indistinguishable, so pick on cost" from "clarity is decisively the top lever," which is exactly the judgement a redesign budget needs.

Personnel caution. Relative weights describe how the overall score is composed, not what an instructor should be rewarded for. Used descriptively they are illuminating; used to weight a summative decision they inherit every validity caveat of the underlying ratings.

Limitations and honest caveats

A critical reader should press on several points, and the honest analyst raises them first.

First, relative importance is not causal. Relative weights and dominance analysis partition predictive variance in observational data. A high weight for "clarity" does not prove that improving clarity will raise the overall score; an omitted variable (say, prior interest) could drive both. These methods answer a variance-decomposition question, not an intervention question — for that, see the causal-inference articles in this knowledge base.

Second, the method has been challenged. Thomas, Zumbo, Kwan, and Schweitzer (2014), reanalysing Johnson's approach in Multivariate Behavioral Research, argued the relative-weights transformation does not always recover importance as cleanly as claimed and can behave unexpectedly under certain correlation structures. The practical response is to run dominance analysis as a cross-check when predictors are few, and to treat exact weight values as estimates with intervals, not precise truths.

Third, weights depend on the model you specify. Add or drop a predictor and every weight changes, because they are shares of this model's R-squared. A weight is a within-model statement, not an absolute property of an item, so the predictor set must be theoretically justified.

Fourth, shared, unallocated variance is real. These methods assign the jointly explained variance by a defensible rule, but they cannot conjure a unique answer where the data genuinely confound two predictors; near-perfectly correlated items should be combined or one dropped, not force-ranked.

Finally, stability needs sample size. With small classes or few respondents, weights and their rank order are unstable; bootstrap confidence intervals are mandatory, not optional, before acting.

How Koji incorporates this

Koji is designed to make driver analysis both possible and honest, and to reduce the reliance on a single regression that these methods exist to repair.

  • Clean, structured driver data. Koji collects the aspect-level scale and single_choice items and the overall rating in a tidy schema, so an institutional-research team can run relative weights or dominance analysis directly, with the exposure and completeness metadata needed to judge stability.
  • Evidence for mechanism, not just weight. Because relative weights are silent on why an aspect matters, Koji's AI-moderated conversational interview probes the reason behind a rating and its automatic thematic analysis surfaces the language students use — so a high weight for "feedback" comes with transcript evidence of what about feedback drove it, moving from correlation toward a testable mechanism.
  • Triangulation over a single model. Koji pairs any variance-decomposition with converging evidence — open-text themes, mid-cycle signals, cohort comparisons — so a committee is not betting a redesign on one regression whose weights a critic could reasonably contest.
  • Small-sample-aware and bias-aware reporting. Koji flags when a course has too few responses to support stable weights and reports intervals rather than point ranks, consistent with its philosophy of never presenting a fragile number as settled.

Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where relative-weights driver analysis of satisfaction is a standard technique — the education product simply points it at teaching and courses.

Frequently asked questions

Why can't I just read the regression coefficients to find the top driver?

Because course-evaluation items are highly correlated, and multicollinearity makes standardised coefficients unstable, understated for shared-variance predictors, and sometimes sign-reversed. The beta ranking can change year to year from noise alone. Relative weights and dominance analysis partition the explained variance fairly, giving a stable importance ranking coefficients cannot.

What is the difference between relative weights and dominance analysis?

They answer the same question by different routes. Dominance analysis averages each predictor's incremental R-squared across all possible subsets of the other predictors; relative weights analysis orthogonalises the predictors, computes importance where it is unambiguous, then transforms back. Their results are usually near-identical, but relative weights scale to many predictors where dominance analysis becomes computationally heavy.

Do these methods prove what will improve the overall score?

No. They decompose predictive variance in observational data, not causal effect. A high weight for an aspect does not guarantee that improving it raises the overall rating, because an omitted variable could drive both. Treat the weights as a prioritisation aid to be confirmed with causal designs or a genuine trial.

How do I know a ranking of drivers is reliable and not noise?

Compute bootstrap confidence intervals for the weights and for the differences between them, as Tonidandel and LeBreton recommend. If clarity and feedback have overlapping intervals, they are not distinguishable and you should choose between them on other grounds. Small samples make weights unstable, so intervals are essential.

How is this different from importance-performance analysis?

Importance-performance analysis usually uses stated importance or a simple correlation for its importance axis. Relative weights provide a derived, multicollinearity-corrected importance that can replace that axis, so the two are complementary: relative weights sharpen the importance dimension that an importance-performance map then plots against current performance.

Are relative weights safe to use in personnel decisions?

Use them descriptively, not as decision weights. They describe how an overall score is composed within a specific model and dataset, and they inherit every validity limitation of the underlying student ratings. High-stakes summative use requires multiple sources of evidence, not a single variance decomposition.

Related resources

References

  • Johnson, J. W. (2000). A heuristic method for estimating the relative weight of predictor variables in multiple regression. Multivariate Behavioral Research, 35(1), 1-19. https://doi.org/10.1207/S15327906MBR3501_1
  • Budescu, D. V. (1993). Dominance analysis: A new approach to the problem of relative importance of predictors in multiple regression. Psychological Bulletin, 114(3), 542-551. https://doi.org/10.1037/0033-2909.114.3.542
  • Azen, R., & Budescu, D. V. (2003). The dominance analysis approach for comparing predictors in multiple regression. Psychological Methods, 8(2), 129-148. https://doi.org/10.1037/1082-989X.8.2.129
  • Tonidandel, S., & LeBreton, J. M. (2011). Relative importance analysis: A useful supplement to regression analysis. Journal of Business and Psychology, 26(1), 1-9. https://doi.org/10.1007/s10869-010-9204-3
  • Thomas, D. R., Zumbo, B. D., Kwan, E., & Schweitzer, L. (2014). On Johnson's (2000) relative weights method for assessing variable importance: A reanalysis. Multivariate Behavioral Research, 49(4), 329-338. https://doi.org/10.1080/00273171.2014.905766

Related articles