New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods12 min read

Do Student Ratings, Peer Observation and Learning Outcomes Agree? The Multitrait-Multimethod Test

If student surveys, peer observation, and learning outcomes all claim to measure teaching quality, do they converge — and do a course evaluation's sub-scores stay distinct or collapse into one halo? Campbell & Fiske's multitrait-multimethod matrix is the classic test, and it explains why triangulation, not any single method, is the honest standard.

Koji Education Team

Product

If a student survey, a peer observation, and a learning-outcome measure all claim to capture "teaching quality," a fair question is whether they actually agree. And if a single course-evaluation form reports separate scores for clarity, organisation, and fairness, another fair question is whether those are genuinely distinct or just three echoes of one overall impression. Campbell and Fiske's multitrait-multimethod (MTMM) matrix, published in 1959, is the classic apparatus for both questions — convergent validity (different methods measuring the same thing should agree) and discriminant validity (measures of different things should diverge). Applied to course evaluation, it explains why triangulating across methods is the honest standard, and why the agreement you see is often smaller — or more inflated by shared method — than it looks.

Answer box (BLUF): Different sources of evidence about teaching — student ratings, peer observation, self-reflection, learning outcomes — should converge if they measure the same underlying quality; that is convergent validity. Distinct things a form claims to measure should stay distinct; that is discriminant validity. The multitrait-multimethod matrix tests both by crossing several traits with several methods. In practice the convergence among teaching-quality measures is modest at best, and much of the apparent agreement within one survey is "method variance" (a halo of general impression), not real signal. The lesson is not to pick the one true method but to triangulate several and check whether they actually agree.

What the research says

Campbell and Fiske (1959), "Convergent and discriminant validation by the multitrait-multimethod matrix" (Psychological Bulletin), introduced convergent and discriminant validity as facets of construct validity and gave researchers a concrete procedure. You measure several traits (say, instructor clarity, course organisation, assessment fairness) using several methods (student survey, trained-observer rating, instructor self-report), then arrange all the correlations in a matrix. Four patterns matter:

  • Convergent validity appears in the monotrait–heteromethod cells: the same trait measured by different methods should correlate substantially. Student-rated clarity should track observer-rated clarity.
  • Discriminant validity requires that a trait correlate more with itself across methods than with different traits — including different traits measured by the same method.
  • Method variance shows up in the heterotrait–monomethod cells: when two different traits measured by the same method correlate highly, the shared method (not the traits) is inflating the correlation. In surveys this is the halo effect wearing a statistical coat.
  • Reliability sits on the diagonal (same trait, same method).

The SET literature has run versions of this test for decades. Feldman (1989), synthesising studies that compared teachers rated by themselves, current and former students, colleagues, administrators, and neutral observers, found the strongest agreement between current students and colleagues (or administrators), weaker agreement between students and teachers' self-ratings, and the least between self- and colleague ratings. In MTMM terms: partial convergent validity across rater "methods," never near-identity.

On the survey-versus-outcome axis, Cohen's (1981) meta-analysis of multisection validity studies found that overall instructor ratings correlated about .43 with student achievement, and overall course ratings about .47 — a moderate convergence often cited as evidence that ratings track learning. But Uttl, White, and Gonzalez (2017) re-analysed this literature and showed that once prior student ability and small-study effects are accounted for, the SET–learning correlation shrinks toward zero. That contrast is the crux: whether student ratings and learning outcomes converge is genuinely contested, which is precisely why no institution should treat any single method as the whole truth.

Why it matters for course evaluation in practice

European quality assurance increasingly asks programmes to triangulate — pairing student surveys with peer observation, learning analytics, self-evaluation, and outcome data. MTMM turns "we use multiple sources" from a slogan into a checkable claim, with two direct practical consequences.

First, convergent validity tells you whether your sources actually agree. If your student survey and your peer-observation protocol both purport to measure "active-learning use" but correlate weakly across courses, you have a problem to investigate, not a contradiction to bury: the two methods may be tapping different facets (what students feel versus what an observer sees), or one instrument may be poorly targeted. Either way, averaging them into a single "teaching score" as if they agreed would be indefensible.

Second, discriminant validity tells you whether your sub-scores are real. Many evaluation forms report five or six item averages — clarity, enthusiasm, organisation, fairness, workload. If these intercorrelate so highly that they collapse into a single general-impression factor, then reporting them as distinct is theatre: you are showing a committee five decimals of one underlying halo. MTMM (and its modern confirmatory-factor-analysis form) is how you detect that the "profile" is really a point. This directly connects to the halo-effect and common-method-bias problems documented elsewhere in this knowledge base: a single student survey, however well designed, cannot separate genuine multidimensional signal from a shared-method halo on its own.

For accreditation, the payoff is credibility. "Our judgement of this programme rests on four methods that converge on the same conclusion" is a far stronger evidence claim than four dashboards that were never checked for whether they agree.

Limitations and honest caveats

MTMM is demanding, and a rigorous reader will note where it strains:

  • It needs multiple methods on the same targets. You must measure the same traits, on the same courses or instructors, with genuinely different methods. That is expensive and rare; most institutions have one method (a survey) measured well and everything else measured sporadically.
  • The original procedure is subjective. Campbell and Fiske's "eyeball the matrix" rules were later formalised into confirmatory-factor-analysis models (Widaman, 1985), which estimate trait and method factors explicitly — but those models can be hard to identify and sometimes fail to converge, especially with few traits or methods.
  • Low convergence is not automatically failure. Different methods legitimately capture different facets: student experience, observed behaviour, and demonstrated learning are related but not interchangeable constructs. Modest correlations may reflect real, meaningful differences rather than measurement error — over-expecting convergence is its own mistake.
  • "Same trait across methods" is an assumption. Whether observer-rated "clarity" and student-rated "clarity" are the same construct is a substantive judgement the matrix cannot settle by itself.
  • Small samples and unstable estimates. With a handful of courses, MTMM correlations are noisy; treating a low cell as decisive over-reads the data.
  • Method variance can be substantive, not just nuisance. The "halo" a single method introduces sometimes carries real information about overall student experience, so partialling it out entirely can discard signal.

The honest reading is that MTMM is a diagnostic for how much to trust multi-source claims, not a verdict machine. It tells you whether your triangulation is earning its keep.

How Koji incorporates this

Koji is designed around the premise MTMM formalises: no single method is sufficient, so the platform is built to make triangulation and convergence-checking practical rather than aspirational.

  • A second method on the same trait, from the same instrument. Koji captures both structured items (scale, single_choice, ranking) and open-text, then applies automatic thematic analysis to the open responses. That gives you a text-derived reading of, say, "clarity" alongside the scale-derived reading — a within-platform convergent-validity check that a numbers-only survey cannot offer.
  • Holding survey evidence alongside other methods. Koji's reporting is designed to sit next to imported peer-observation (for example COPUS or TDOP) and learning-analytics data, so a programme can actually compare whether the student, observer, and trace-data pictures converge instead of storing them in disconnected systems.
  • Attacking method variance at the source. Common-method halo thrives when every item is a quick Likert tick coloured by overall impression. Koji's AI-moderated conversational interviews probe for specific, behaviour-anchored evidence, which is designed to loosen the halo so a "clarity" judgement is not merely a proxy for "I liked the course."
  • Flagging collapsed sub-scores. Koji's bias-aware reporting is oriented toward surfacing when reported dimensions are so intercorrelated they behave as one factor, so an institution is warned before it over-interprets a halo as a multidimensional profile.
  • Structure ready for formal analysis. Because responses are retained at the item and individual level within courses, the data supports the confirmatory-factor MTMM models (Widaman, 1985) and the generalizability-theory analyses that quantify how much variance is trait versus method.

Koji is framed throughout as designed to support convergent- and discriminant-validity checks and triangulation — it does not, and does not claim to, guarantee validity, which remains an argument the institution must make. The same multi-method engine powers Koji's core research platform at koji.so, where triangulating survey scores against interviews and behavioural data is the same discipline applied to product and customer research.

Frequently asked questions

What is the multitrait-multimethod matrix? It is a table of correlations that crosses several traits (things you want to measure) with several methods (ways of measuring them). Its layout lets you read off convergent validity (same trait, different methods should agree), discriminant validity (different traits should diverge), and method variance (inflation caused by a shared method).

What is the difference between convergent and discriminant validity? Convergent validity is the degree to which measures that should be related — the same construct measured different ways — actually correlate. Discriminant validity is the degree to which measures that should not be related stay uncorrelated. A good instrument shows both: it agrees with itself across methods and separates from things it is not measuring.

Why do student ratings and learning outcomes not always converge? Because they may measure different things and share confounders. Cohen (1981) found a moderate correlation (~.43), but Uttl et al. (2017) showed it shrinks toward zero once prior ability is controlled. Satisfaction and demonstrated learning are related but distinct constructs, so strong convergence should not be assumed.

What is method variance and why does it matter? Method variance is correlation between different traits that arises because they were measured the same way — for example, several survey items all coloured by a student's overall impression (a halo). It inflates apparent agreement within a single method, which is why within-survey correlations overstate how multidimensional your evaluation really is.

Does weak convergence mean my evaluation is broken? Not necessarily. Different methods legitimately capture different facets — what students experience, what an observer sees, what students demonstrate. Modest correlations can reflect genuine, meaningful differences rather than error. The point is to check convergence deliberately, not to expect the methods to be interchangeable.

How many methods do I need to run an MTMM check? At least two methods measuring at least two traits on the same targets, though three of each makes the patterns far clearer. In practice most institutions start by pairing their student survey with one other method — peer observation or learning analytics — and build from there.

Related resources

References

  • Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105. https://doi.org/10.1037/h0046016
  • Feldman, K. A. (1989). Instructional effectiveness of college teachers as judged by teachers themselves, current and former students, colleagues, administrators, and external (neutral) observers. Research in Higher Education, 30(2), 137–194. https://doi.org/10.1007/BF00992716
  • Cohen, P. A. (1981). Student ratings of instruction and student achievement: A meta-analysis of multisection validity studies. Review of Educational Research, 51(3), 281–309. https://doi.org/10.3102/00346543051003281
  • Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty''s teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
  • Widaman, K. F. (1985). Hierarchically nested covariance structure models for multitrait-multimethod data. Applied Psychological Measurement, 9(1), 1–26. https://doi.org/10.1177/014662168500900101

Related articles

analysis-reporting

Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough

Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.

analysis-reporting

Why a Single Student Survey Can't Stand Alone: Common-Method Bias in Course Evaluation

Common-method bias (Podsakoff et al., 2003) explains why correlations inside a single end-of-term student survey are inflated by the shared method itself - and why triangulating teaching evidence matters. A research-grounded guide for quality assurance.

research-methods

Does the Professor Students Rate Highest Teach Them Most? Value-Added Evidence vs Student Ratings

The strongest causal evidence — random-assignment studies that track follow-on course performance — shows the instructors students rate highest are often not the ones whose students learn most over time. What Carrell & West (2010) and Braga et al. (2014) found, and what it means for course evaluation.

research-methods

Validity Is About the Use, Not the Instrument: Applying Kane's Argument-Based Framework to Course Evaluation

Asking whether course evaluations are valid is the wrong question. Kane's argument-based framework asks whether a specific interpretation and use of the scores is justified. We rebuild the SET debate as an interpretation-use argument, expose where each inference breaks, and show how Koji strengthens the weak links.