One Number or Many? The Dimensionality Debate and How to Use Student Ratings for Personnel Decisions
Should a promotion committee use a single global teaching score or a detailed profile of many dimensions? The 1997 d'Apollonia & Abrami vs Marsh & Roche debate set the terms — and the answer depends on whether the purpose is summative judgement or formative improvement.
Koji Education Team
Product
In brief: A long-running debate, crystallised in a 1997 American Psychologist exchange, asks whether student ratings are best treated as one global dimension or many specific dimensions. d'Apollonia and Abrami (1997) argued that for summative personnel decisions, ratings effectively measure a single general instructional-skill factor (a composite of a few subskills), so a small number of global items or a validated composite is appropriate. Marsh and Roche (1997) countered that ratings are genuinely multidimensional and that the specific profile is what makes them useful for improving teaching. The resolution is purpose-driven: use a defensible global composite for high-stakes comparison, and the full dimensional profile for formative feedback.
What the research says
Every course-evaluation instrument embeds a stance on dimensionality. A one-item "overall, how effective was this instructor?" assumes teaching quality is essentially unidimensional. A 35-item SEEQ assumes it has many distinguishable facets — clarity, enthusiasm, organisation, rapport, fairness, feedback, breadth. Which is right has direct consequences for how committees should read scores.
The anchor exchange: d'Apollonia & Abrami (1997) vs Marsh & Roche (1997)
In a 1997 American Psychologist special section, Sylvia d'Apollonia and Philip Abrami advanced a deliberately practical thesis. Reviewing the factor-analytic and multisection-validity evidence, they argued that although effective instruction may be multidimensional, student ratings of instruction primarily measure a single general factor — general instructional skill — composed of three broad subskills: delivering instruction, facilitating interactions, and evaluating student learning. From their meta-analytic review of multisection validity studies, they reported that ratings predict instructor impact on learning to a moderate degree (a corrected correlation around r = .47). Their conclusion for practice: for summative decisions (tenure, promotion, comparison), institutions should rely on a small number of global rating items or an empirically weighted composite, because the fine-grained dimensions are too inter-correlated and too unreliable at the item level to support defensible distinctions between instructors.
Herbert Marsh and Lawrence Roche, in the same section, pushed back. Drawing on the SEEQ research programme, they maintained that student ratings are reliably multidimensional, that a single global score discards information, and that "making students' evaluations of teaching effectiveness effective" requires the dimensional profile — because telling an instructor which aspect (organisation? feedback? rapport?) needs work is what drives improvement. Reducing a rich profile to one number, they argued, is both psychometrically wasteful and pedagogically useless.
The synthesis that emerged
Later reviewers (e.g. Abrami, d'Apollonia & Rosenfield, 2007) framed the two positions as complementary rather than contradictory, resolved by purpose:
- For summative use — ranking, personnel decisions — a higher-order general factor dominates the variance, the specific dimensions are highly inter-correlated, and a validated global composite is the most defensible single summary. Trying to make tenure decisions on a dozen separately-reported subscales invites cherry-picking and over-interpretation of unreliable differences.
- For formative use — helping an instructor improve — the multidimensional profile is essential, because "your overall score is 3.8" gives no actionable direction, whereas "students rated your organisation high but your feedback low" does.
Why it matters for course evaluation in practice
The dimensionality question is not academic hair-splitting; it determines instrument design and reporting:
- Match the report to the decision. If a committee is making a summative comparison, give it a validated global composite with its uncertainty, not fifteen subscale means it will mine for whatever supports a predetermined view. If an instructor is seeking to improve, give them the full profile with the lowest dimensions flagged.
- Do not invent precision from inter-correlated items. Because the specific dimensions correlate highly, small differences between them are often unreliable. Reporting a rank order of an instructor's strengths and weaknesses to two decimal places overstates what the data can bear.
- Design the instrument for its primary purpose. An instrument built for personnel decisions needs a defensible overall measure; one built for development needs coverage of distinct, actionable teaching behaviours. Trying to serve both with one rigid form usually serves neither well.
- Beware the global halo. A single overall item is vulnerable to halo effects — one strong impression colouring the whole rating — which is one reason the formative profile, anchored in specific behaviours, is more diagnostic.
A practical design rule of thumb
The purpose-driven resolution translates into a simple design rule. First decide what the primary decision is. If the instrument's main job is personnel review, build it around a small number of well-validated global items, report a single composite with its uncertainty, and resist the urge to publish a dozen subscale league tables that committees will mine selectively. If its main job is teaching development, build it around distinct, behaviourally-anchored dimensions, report the full profile, and flag the lowest dimensions as the starting point for a conversation. Where one instrument must serve both — the common reality in European quality cycles — separate the reporting views rather than the data: collect rich evidence once, then render a defensible summary for the committee and an actionable profile for the instructor. The error to avoid is a single rigid form that reports fifteen noisy subscale means to everyone, satisfying neither the rigour a tenure case demands nor the specificity an improving lecturer needs.
Limitations and honest caveats
- The debate is partly about emphasis, not fact. Both camps agree ratings have some general factor and some specific structure. The disagreement is about which to privilege for which purpose — a question of use, not a clean empirical winner.
- "General factor" is sample- and instrument-dependent. How unidimensional ratings appear depends on the items, the population, and the factor-analytic method. A poorly designed instrument can manufacture a single factor by asking near-identical questions; a rich one can recover distinct dimensions.
- The validity correlations are themselves contested. The r ≈ .47 figure from the multisection validity tradition has been challenged by Uttl, White and Gonzalez (2017), who found near-zero correlations once small-sample studies are properly weighted. So the summative defensibility of even a global composite is weaker than d'Apollonia and Abrami implied.
- The exchange predates modern analytics. The 1997 debate could not anticipate automatic thematic analysis of open text, which arguably dissolves part of the trade-off — you can have a defensible summary and rich, actionable specifics without forcing students through dozens of Likert items.
- Aggregation level matters. Dimensionality at the student level and at the class-average level can differ; conclusions drawn at one level may not hold at the other.
The defensible position is conditional: use a validated global composite for high-stakes comparison and the full dimensional profile for improvement — and do not over-read small differences among highly inter-correlated dimensions either way.
How Koji incorporates this
Koji is designed to dissolve the false choice between "one number" and "many" by collecting rich, structured evidence once and reporting it differently for different purposes. These mechanisms support both uses without forcing a single rigid instrument; they mitigate, rather than eliminate, the trade-offs the 1997 debate identified.
- Both a defensible summary and an actionable profile. Koji can produce a clear overall picture for summative review and a dimension-by-dimension breakdown for formative development from the same evaluation, so committees and instructors each get the view their decision needs.
- AI-moderated conversational interviews that recover specifics without item bloat. Instead of asking students to complete dozens of Likert subscales (which inflates length and breakoff), Koji's interviewer probes the dimensions that matter conversationally, and automatic thematic analysis of the open text reconstructs the actionable profile — organisation, feedback, clarity, rapport — without the questionnaire-length penalty.
- Structured question types mapped to purpose.
scaleandsingle_choiceitems can anchor a validated composite for comparison;open_endedandrankingitems surface the specific, improvement-oriented detail. The instrument can be tuned to whether the cycle is summative or formative. - Uncertainty-aware, distribution-first reporting discourages over-interpreting small gaps between highly correlated dimensions — the precise error the dimensionality literature warns against.
- Closing-the-loop action tracking uses the dimensional profile where it is strongest: identifying what to change and following whether it improved.
Koji is designed to mitigate the one-number-versus-many tension by separating collection from reporting and tailoring the output to the decision — not by claiming a single representation is universally correct. The core research platform at koji.so applies the same dual-view philosophy to product and customer research, where teams likewise need both a headline metric and the diagnostic detail beneath it.
Related Resources
- What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and Multidimensionality
- Is One "Overall" Question Enough? Single-Item vs Multi-Item Measures
- Can You Fairly Rank Instructors by Their Course-Evaluation Scores?
- Interpreting and Reporting Student Ratings Responsibly: Linse (2017)
- Generalizability Theory and the Reliability of Student Ratings
- Peer Observation vs Student Evaluations: What Each Actually Measures
References
- d'Apollonia, S., & Abrami, P. C. (1997). Navigating student ratings of instruction. American Psychologist, 52(11), 1198–1208. https://doi.org/10.1037/0003-066X.52.11.1198
- Marsh, H. W., & Roche, L. A. (1997). Making students' evaluations of teaching effectiveness effective: The critical issues of validity, bias, and utility. American Psychologist, 52(11), 1187–1197. https://doi.org/10.1037/0003-066X.52.11.1187
- Marsh, H. W. (1987). Students' evaluations of university teaching: Research findings, methodological issues, and directions for future research. International Journal of Educational Research, 11(3), 253–388. https://doi.org/10.1016/0883-0355(87)90001-2
- Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
Related articles
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.