Relating a Battery of Teaching Behaviours to a Battery of Outcomes: Canonical Correlation Analysis
When you want to relate a whole set of teaching-behaviour items to a whole set of outcome items, running many pairwise correlations misleads. Canonical correlation analysis maps the joint structure.
Koji Education Team
Product
In brief
When you want to know how a whole set of teaching-behaviour items (clarity, feedback, pace, organisation, enthusiasm) relates to a whole set of outcome items (perceived learning, confidence, interest, intention to persist), running twenty separate correlations inflates error and hides the joint structure. Canonical correlation analysis (CCA), introduced by Harold Hotelling in 1936, finds the weighted combination of the teaching items and the weighted combination of the outcome items that correlate most strongly — a single honest summary of the many-to-many relationship, plus any additional independent dimensions that exist.
What the research says
Hotelling (1936), "Relations between two sets of variates" (Biometrika), generalised ordinary correlation to two sets of variables at once. CCA constructs pairs of canonical variates — one linear combination from set X (teaching behaviours) and one from set Y (outcomes) — chosen so that their correlation, the canonical correlation, is maximal. Subsequent pairs are extracted orthogonal to the first, giving up to min(p, q) independent dimensions of shared variation. It is the general parametric model that subsumes multiple regression, MANOVA, discriminant analysis and ordinary correlation as special cases, a point Knapp (1978) made in arguing that most univariate tests are special cases of CCA.
Sherry and Henson (2005), a widely cited user-friendly primer in the Journal of Personality Assessment, stress interpreting structure coefficients — the correlation of each original variable with its own canonical variate — rather than the standardized canonical function coefficients, because the latter are unstable when predictors are correlated, which is exactly the situation with overlapping teaching items. They also recommend reporting the squared canonical correlation as an effect size and cross-validating the solution. A crucial companion idea is redundancy (Stewart and Love, 1968): the canonical correlation tells you how well the two composites align, but not how much of the total variance in one battery is actually captured by the other — redundancy indices answer that, and they are often much lower than the headline canonical correlation suggests. Tabachnick and Fidell (2019) give the standard applied treatment, including the sample-size demands and diagnostic checks.
Why it matters for course evaluation in practice
Course-evaluation instruments are batteries on both sides. SEEQ-style tools measure several teaching dimensions; the outcome side includes self-reported learning, satisfaction, confidence and interest. The natural question — which bundle of teaching behaviours goes with which bundle of outcomes? — is a two-set problem, not a series of one-to-one correlations.
- It controls the error rate honestly. Running 5 x 4 = 20 separate correlations invites false positives (the exact problem multiple comparisons corrections exist to manage). CCA asks one multivariate question instead.
- It reveals independent axes. CCA can show that clarity and organisation jointly track perceived-learning and confidence on one dimension, while enthusiasm and immediacy track interest and enjoyment on a second, orthogonal dimension. A table of pairwise correlations cannot express that structure.
- It complements neighbouring tools rather than duplicating them. Mediation analysis traces one mechanism at a time; relative weights and dominance analysis rank predictors for a single outcome. CCA is the many-to-many map that sits above both, and it pairs naturally with the multidimensional view of teaching that Marsh and the SEEQ established.
Limitations and honest caveats
CCA is powerful but demanding. It is descriptive, not causal — it summarises correlational structure, so a strong canonical dimension is a pattern to explain, not proof that behaviours caused outcomes. It is sample-hungry: with rules of thumb of roughly 10 to 20 cases per variable, a battery of ten items across both sides wants a few hundred respondents, or the solution overfits and cross-validates poorly. Canonical variates beyond the first are often uninterpretable, and naming a composite is a judgement call, so structure coefficients and theory must guide interpretation rather than raw weights. The method assumes linearity and is sensitive to outliers and non-normality. Most importantly, a high canonical correlation can coexist with low redundancy — the two composites align tightly yet share little of each battery's total variance — so redundancy must be reported alongside the canonical correlation or the result flatters itself. Finally, ordinal Likert items only approximate the interval, multivariate-normal ideal CCA assumes; badly skewed or ceiling-bound items should be modelled with care.
How Koji incorporates this
Koji's structured question types produce exactly the two-battery data CCA needs — a block of teaching-behaviour items and a block of outcome items answered by the same students — without the manual stitching that usually makes this analysis rare in practice. Koji's analysis layer is designed to report joint structure rather than a wall of pairwise correlations: a committee sees "these teaching behaviours travel together with these outcomes" as coherent, labelled dimensions reported with structure coefficients and redundancy, and Koji flags when the response count is too small for a stable canonical solution instead of presenting an overfit result as fact. Because Koji favours specific, low-inference behaviour items on the predictor side, the teaching battery enters the model cleanly rather than as vague global impressions. And because Koji's AI-moderated conversational interviews probe why a behaviour maps to an outcome, each canonical dimension can be given a narrative from the open-text follow-ups, not just a coefficient — the number and the reason arrive together. These mechanisms are designed to make a multivariate, error-controlled reading of the teaching-to-outcome relationship routine; they do not remove the sample-size requirement, which constrains CCA on any platform. Teams who also run product and customer studies can apply the same two-set logic in Koji's core research platform at koji.so, relating a battery of feature perceptions to a battery of adoption outcomes.
Reading a canonical solution in practice
Suppose a faculty pools 400 responses across sections and runs CCA with five teaching-behaviour items (clarity, organisation, feedback, pace, enthusiasm) against four outcome items (perceived learning, confidence, interest, intention to continue). The analysis returns four canonical correlations, but only the first two survive a dimension-reduction test and cross-validation. On the first dimension the structure coefficients are large and positive for clarity, organisation and feedback on the teaching side, and for perceived learning and confidence on the outcome side — a readable "the course was legible and I could tell I was learning" axis. On the second, enthusiasm and pace load with interest and intention-to-continue — an "energy and momentum" axis largely independent of the first. The redundancy index then shows the teaching battery explains only about 28% of the outcome battery's variance through these dimensions, a sober reminder that much of what students report is not captured by the measured behaviours. A committee reading this learns two things a correlation table would have buried: legibility and energy are separable levers, and most outcome variance still lives outside the instrument. That combination — independent, named dimensions plus an honest ceiling on shared variance — is exactly what CCA is for, and why it belongs alongside single-outcome driver analyses rather than replacing them.
Frequently asked questions
What question does canonical correlation analysis answer?
It answers how one whole set of variables relates to another whole set at once — for example, how the battery of teaching-behaviour items relates to the battery of outcome items — instead of testing each pair separately. It finds the weighted combinations of each set that correlate most strongly.
How is CCA different from running many correlations?
Many pairwise correlations inflate the false-positive rate and cannot reveal that groups of items move together as independent dimensions. CCA asks a single multivariate question, controls the error rate, and extracts orthogonal dimensions of shared variation.
Should I interpret the weights or the structure coefficients?
Interpret the structure coefficients — each variable's correlation with its own canonical variate. Standardized canonical weights are unstable when items are correlated, which is normal for course-evaluation batteries, so weights alone can mislead.
What is redundancy and why report it?
Redundancy is the proportion of variance in one battery actually captured by the other battery's canonical variate. A high canonical correlation can still have low redundancy, so reporting redundancy prevents overstating how much the two sets really share.
How large a sample does CCA need?
It is sample-hungry: a common guideline is 10 to 20 respondents per variable across both sets. With small classes the solution overfits and does not replicate, so pool across sections or terms, or use a simpler analysis.
Does a strong canonical correlation prove teaching causes the outcome?
No. CCA is descriptive and correlational. A strong dimension identifies a pattern worth explaining, but establishing causation needs a design such as a quasi-experiment, not a canonical correlation.
References
- Hotelling, H. (1936). Relations Between Two Sets of Variates. Biometrika, 28(3/4), 321-377. https://doi.org/10.1093/biomet/28.3-4.321
- Sherry, A., & Henson, R. K. (2005). Conducting and Interpreting Canonical Correlation Analysis in Personality Research: A User-Friendly Primer. Journal of Personality Assessment, 84(1), 37-48. https://doi.org/10.1207/s15327752jpa8401_09
- Stewart, D., & Love, W. (1968). A general canonical correlation index. Psychological Bulletin, 70(3, Pt.1), 160-163. https://doi.org/10.1037/h0026143
- Knapp, T. R. (1978). Canonical correlation analysis: A general parametric significance-testing system. Psychological Bulletin, 85(2), 410-416. https://doi.org/10.1037/0033-2909.85.2.410
- Tabachnick, B. G., & Fidell, L. S. (2019). Using Multivariate Statistics (7th ed.). New York: Pearson.
Related resources
- Which Aspects of Teaching Actually Drive the Overall Score? Relative Weights and Dominance Analysis
- Why Was the Course Rated Well, Not Just Whether? Mediation Analysis
- Does the Effect Depend on Who or What? Moderation and Interaction Effects
- Multidimensional Scaling for Course Evaluation: A Perceptual Map
- Compare 60 Instructors, Expect 3 False Alarms: Multiple Comparisons and the False Discovery Rate
- What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and Multidimensional Feedback
Related articles
Compare 60 Instructors, Expect 3 False Alarms: Multiple Comparisons and the False Discovery Rate
When a QA office tests every instructor against a benchmark, chance alone produces "significant" outliers. What the multiple-comparisons literature — Bonferroni, and Benjamini & Hochberg''s false discovery rate — says about flagging course-evaluation scores fairly.
Does the Effect Depend on Who or What? Moderation and Interaction Effects in Course-Evaluation Analysis
A bias or a teaching effect on evaluations often holds only for some students, courses or conditions. Moderation analysis tests those "it depends" claims properly, and it is far harder than it looks.
Multidimensional Scaling for Course Evaluation: A Perceptual Map of How Students See Your Courses
Multidimensional scaling turns a table of similarities into a two-dimensional map. Here is how MDS reveals the hidden structure in course-evaluation data that averages and factor analysis both miss.
Which Aspects of Teaching Actually Drive the Overall Score? Relative Weights and Dominance Analysis
When clarity, workload, support, and feedback are all correlated, regression coefficients cannot tell you which one drives the overall rating. Relative weights and dominance analysis partition the explained variance fairly.