Multidimensional Scaling for Course Evaluation: A Perceptual Map of How Students See Your Courses
Multidimensional scaling turns a table of similarities into a two-dimensional map. Here is how MDS reveals the hidden structure in course-evaluation data that averages and factor analysis both miss.
Koji Education Team
Product
In brief
Multidimensional scaling (MDS) takes a matrix of similarities or distances — between courses, instructors, or evaluation items — and places each object as a point on a low-dimensional map so that objects students perceive as similar sit close together. Unlike averaging, which collapses every course to a single number, MDS shows the structure of perception: which courses cluster, which sit alone, and what dimensions separate them. It is a descriptive, exploratory tool, not a significance test, and the map you get depends heavily on the distance measure you feed it.
What the research says
MDS is a family of methods, not one algorithm, and it developed in three foundational steps. Warren Torgerson (1952) introduced metric MDS in Psychometrika, showing that if you have a matrix of ratio- or interval-scaled distances between objects, you can recover a set of coordinates in a small number of dimensions whose Euclidean distances reproduce the original matrix — the classical (Torgerson-Gower) solution via eigendecomposition of a double-centred distance matrix. The limitation was obvious: real perceptual data rarely arrive as clean interval distances.
Roger Shepard (1962) broke that constraint with nonmetric MDS. In "The analysis of proximities" he showed that only the rank order of the similarities needs to be preserved, not their exact magnitudes — a far weaker and more realistic assumption for human judgement data. Joseph Kruskal (1964) then made nonmetric MDS practical and principled by defining a loss function he called stress: the normalised discrepancy between the map distances and a monotonic transformation of the input proximities. Kruskal minimised stress by steepest descent, giving a reproducible numerical recipe and a fit statistic analysts still report today. Kruskal offered rough benchmarks — stress around 0.20 is poor, 0.10 fair, 0.05 good, 0.025 excellent — though he cautioned these depend on the number of points and dimensions.
The modern synthesis is Borg and Groenen (2005), Modern Multidimensional Scaling, which unifies metric and nonmetric approaches under the SMACOF majorisation algorithm and treats MDS as a general geometry-of-data method. Their central message for practitioners: MDS is a tool for seeing structure, and its value is interpretive — the recovered dimensions and clusters must be validated against theory, not read as objective facts.
Why it matters for course evaluation in practice
Most course-evaluation reporting is one-dimensional: a mean, a top-box percentage, a percentile rank. That machinery answers "how high?" but never "how are these courses different from one another?" MDS answers the second question, and three applications matter for a quality office.
Mapping courses or instructors. Build a similarity matrix — for example, correlate each course profile across the standard item battery, or ask students to rate how similar pairs of courses feel — and MDS places all courses on a single map. A programme director can then see that the two "problem" courses everyone worries about are actually far apart in perceptual space: one is seen as disorganised, the other as excessively demanding. They need different remedies, which a shared low average would have hidden.
Mapping items. Feed MDS the inter-item distance matrix (1 minus the correlation) and it shows how the evaluation instrument itself is structured — often more legibly than a scree plot. Items that students answer as interchangeable cluster tightly; a lonely item on the edge of the map is measuring something distinct, which is either a hidden dimension worth keeping or a confusing item worth cutting.
Recovering interpretable axes. Because MDS axes are arbitrary until rotated, you can rotate the solution so the horizontal axis aligns with, say, perceived workload and the vertical with perceived support — turning the map into a diagnostic quadrant that non-statistician committee members can read at a glance.
Limitations and honest caveats
MDS is seductive precisely because it always produces a picture, and a picture always looks meaningful. The methodological objections a critical reader will raise are real.
First, the map depends entirely on the distance measure. Correlation-based distances, Euclidean distances on raw means, and Gower distances on mixed data can yield qualitatively different maps from the same courses. There is no single correct choice, so the analyst must justify the metric before interpreting the geometry.
Second, low stress is necessary but not sufficient. You can always drive stress down by adding dimensions, and a two-dimensional map with stress of 0.15 may be forcing a genuinely four-dimensional structure onto a plane, creating spurious closeness. A Shepard diagram (fitted vs observed distances) and a scree plot of stress against dimensionality are mandatory diagnostics, not optional ones.
Third, the axes are not causal and often not even stable. Rotational freedom means the "dimensions" you name are interpretations imposed after the fact; a bootstrap or split-half check of point stability is the honest way to show a cluster is real rather than sampling noise. With few courses or few respondents per course, configurations can be unstable.
Fourth, MDS is exploratory, not confirmatory. It generates hypotheses about structure; it does not test them. Confirmatory factor analysis, property fitting, or an independent replication is needed before a map drives a personnel or curriculum decision. Treating an MDS picture as evidence for a ranking would be a category error — see our companion pieces on fair comparison and misclassification.
How Koji incorporates this
Koji is designed to generate the inputs MDS needs and to keep those inputs honest, rather than to hand a committee a black-box map. Concretely:
- Structured items give clean distance matrices. Koji collects
scale,single_choice,ranking, andyes_noresponses in a tidy schema, so an institutional-research team can compute course-by-course or item-by-item distances without the manual cleaning that usually corrupts a proximity matrix. - Conversational interviews add perceived-similarity evidence. Koji's AI-moderated interview can ask students directly how two courses or two aspects of teaching compare, and can probe why — producing the kind of judged-similarity data Shepard and Kruskal designed nonmetric MDS to handle, rather than forcing analysts to infer similarity only from rating correlations.
- Automatic thematic analysis labels the axes. When MDS surfaces a dimension, Koji's open-text theme extraction can tell you what students actually say along that axis (for example, "clarity of expectations" vs "assessment fairness"), so the interpretation is grounded in student language, not the analyst's guess.
- Bias-aware, small-sample-honest reporting. Koji flags when a course has too few responses for a stable position and pairs any structural view with the response-rate and representativeness context, so a map is never read as more certain than the data allow.
Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where perceptual mapping of competing products is a long-standing use of MDS — the education product simply points that engine at courses and teaching.
The right way to use MDS in a quality process is as a conversation-starter for a programme review: it shows the committee the shape of the problem before anyone argues about a number. Koji is built to make that shape trustworthy.
Frequently asked questions
What is the difference between multidimensional scaling and factor analysis?
Both reduce dimensionality, but factor analysis models the covariances among variables to find latent factors, while MDS models the distances among objects (courses, instructors, or items) to find a spatial configuration. MDS makes weaker assumptions — nonmetric MDS needs only the rank order of similarities — and produces a directly interpretable map rather than factor loadings. They often agree on item structure but answer different questions.
How many dimensions should I use?
Plot stress against the number of dimensions and look for an elbow, exactly as with a scree plot. Two dimensions are usual because they are readable, but if stress stays high in two dimensions the structure is genuinely higher-dimensional and a flat map will mislead. Always report the stress value and show a Shepard diagram so readers can judge fit for themselves.
What counts as an acceptable stress value?
Kruskal's rough guide is 0.20 poor, 0.10 fair, 0.05 good, 0.025 excellent, but these depend on the number of points — more points make low stress harder to achieve, so context matters. Treat stress as a fit diagnostic, not a pass/fail threshold, and pair it with a stability check such as a bootstrap of the configuration.
Can MDS rank instructors or courses?
No. MDS shows how objects relate, not which is better. Distance on the map is similarity, not quality, and the axes carry no inherent good/bad direction. For fair comparison and ranking questions, use empirical-Bayes shrinkage, funnel plots, or the misclassification analyses covered elsewhere in this knowledge base.
Do I need judged-similarity data, or can I use rating correlations?
Either works. Classical applications derive distances from a correlation or profile-distance matrix computed from existing ratings, which needs no extra data collection. Direct judged-similarity data (asking students how similar two things are) is richer and is what nonmetric MDS was designed for, but it costs an extra question. Koji supports both.
Is a two-dimensional map ever misleading?
Yes, whenever the true structure needs more dimensions or when a poorly chosen distance metric distorts the geometry. Points can appear close on a flat map only because a third dimension was suppressed. Guard against this with the stress-by-dimension plot, a Shepard diagram, and an independent confirmatory check before acting on the picture.
Related resources
- Beyond the Average Student: Latent Profile Analysis for Course-Evaluation Segments
- Multiple Correspondence Analysis for Categorical Course-Evaluation Data
- Network Psychometrics: Treating Course-Evaluation Items as a Network
- Importance-Performance Analysis: Turning Scores into a Priority Map
- The Dimensionality Debate and How to Use Student Ratings
- Parallel Analysis for Factor Retention
References
- Torgerson, W. S. (1952). Multidimensional scaling: I. Theory and method. Psychometrika, 17(4), 401-419. https://doi.org/10.1007/BF02288916
- Shepard, R. N. (1962). The analysis of proximities: Multidimensional scaling with an unknown distance function. I. Psychometrika, 27(2), 125-140. https://doi.org/10.1007/BF02289630
- Kruskal, J. B. (1964). Multidimensional scaling by optimizing goodness of fit to a nonmetric hypothesis. Psychometrika, 29(1), 1-27. https://doi.org/10.1007/BF02289565
- Kruskal, J. B. (1964). Nonmetric multidimensional scaling: A numerical method. Psychometrika, 29(2), 115-129. https://doi.org/10.1007/BF02289694
- Borg, I., & Groenen, P. J. F. (2005). Modern Multidimensional Scaling: Theory and Applications (2nd ed.). Springer. https://doi.org/10.1007/0-387-28981-X
Related articles
Importance-Performance Analysis: Turning Course-Evaluation Scores into a Priority Map
A course-evaluation report that lists twenty item means tells you nothing about where to act first. Importance-Performance Analysis (IPA) plots each attribute by how much it matters to students against how well you did, producing a four-quadrant map that separates urgent fixes from wasted effort.
Beyond the Average Student: Latent Profile Analysis for Course-Evaluation Segments
A course mean of 3.6 can hide two entirely different student experiences averaged into one number. Latent profile analysis (LPA) recovers those hidden subgroups from the evaluation data itself, so you can see the delighted minority and the alienated cohort that the average erased.
How Many Dimensions Does Your Evaluation Really Measure? Parallel Analysis for Factor Retention
Deciding how many factors your evaluation instrument measures with the eigenvalue-greater-than-one rule or the scree plot routinely gets the wrong answer. Parallel analysis (Horn, 1965) compares your data against random noise — and is one of the most accurate methods available.
Mapping the Patterns You Cannot Average: Multiple Correspondence Analysis for Categorical Course-Evaluation Data
Much course-evaluation data is genuinely categorical — programme, mode, agree/disagree, chosen theme. Multiple correspondence analysis (MCA) maps how those categories cluster on a two-dimensional plane, revealing response patterns that averaging destroys.