New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Does a Better Researcher Make a Better Teacher? The Teaching-Research Nexus and Your Evaluations

Hattie and Marsh found the correlation between research productivity and teaching quality is essentially zero. What that means for reading course evaluations and for keeping teaching and research evidence separate.

Koji Education Team

Product

In brief: The correlation between how much research an academic produces and how well they teach is essentially zero. Hattie and Marsh's (1996) meta-analysis of 58 studies put it at about r = 0.06 — statistically indistinguishable from no relationship — and their 2002 follow-up confirmed it, while noting that research quality (as opposed to sheer output) shows a small positive link. The practical consequence for course evaluation is twofold: a strong scholarly reputation is not evidence of teaching quality and should not be allowed to halo an instructor's evaluation, and course-evaluation scores should never be read as a proxy for research standing. Teaching and research are largely independent constructs and must be evidenced separately.

What the research says

The anchor is Hattie and Marsh (1996), The Relationship Between Research and Teaching: A Meta-Analysis (Review of Educational Research, 66(4), 507–542). Synthesising 58 studies that had correlated an academic's research output with a measure of their teaching effectiveness (usually student ratings), the authors found an overall correlation of approximately 0.06 — so close to zero that, in their words, the common belief that good researchers are good teachers "is an enduring myth." The near-zero average was not hiding strong effects in both directions that cancelled out: the individual study correlations clustered tightly around nil. They laid out competing theoretical models — a "scarcity" model in which time spent on one starves the other (predicting a negative link), a "conventional wisdom" model predicting a positive link, and an independence model predicting none — and concluded the data best fit independence.

Marsh and Hattie (2002), The Relation Between Research Productivity and Teaching Effectiveness: Complementary, Antagonistic, or Independent Constructs? (The Journal of Higher Education, 73(5), 603–641), revisited the question with fresh multi-section data and a structural-equation approach. The teaching–research correlation was again close to zero. Crucially, they drew a distinction that the raw meta-analysis could not: research productivity (the quantity of output) was unrelated to teaching quality, whereas research quality showed a modest positive association. Being prolific tells you nothing about the classroom; producing genuinely good scholarship is weakly, positively linked — but the effect is small and does not license using one to infer the other.

This finding is one of the more robust in higher-education research precisely because it has been tested so many times, across disciplines, institution types and eras, with student ratings, peer ratings and learning measures — and the answer keeps coming back the same. It sits comfortably alongside the wider evidence that student ratings are multidimensional (Marsh's SEEQ work) and that global impressions can be driven by factors — charisma, fluency, warmth — that have nothing to do with either scholarship or learning.

Why it matters for course evaluation in practice

The halo runs both ways, and both are wrong. A committee that knows an instructor is a distinguished researcher may unconsciously read their evaluation scores more charitably, or discount a poor score as beneath a "serious" academic. Conversely, a brilliant teacher with a thin publication record may have their strong evaluations dismissed as mere popularity. The zero correlation says both moves are unjustified: the evaluation is evidence about teaching, and the CV is evidence about research, and neither should be used to reinterpret the other.

Do not let evaluations stand in for research assessment — or vice versa. Some institutions, under workload pressure, quietly treat strong course evaluations as a general signal of academic worth, or treat research metrics as a general signal of competence that excuses weak teaching. The nexus evidence forbids both shortcuts. A promotion or probation case needs separate teaching evidence (evaluations, peer observation, portfolios) and separate research evidence; one cannot substitute for the other because they do not track together.

Departmental staffing and course design. Because the two are independent, the best teacher of a course is not automatically the most senior researcher, and assigning teaching by research seniority optimises for the wrong variable. Evaluation data — read for what it validly captures — is a better guide to who should teach the high-stakes first-year course than the publication list is.

Limitations and honest caveats

  • Student ratings as the teaching measure. Most studies in the meta-analysis used student evaluations to operationalise "teaching effectiveness." If those ratings are themselves biased or capture charisma more than learning, the zero correlation could partly reflect noise in the teaching measure rather than true independence. That said, studies using peer ratings and learning measures find the same near-zero result, which strengthens the conclusion.
  • Productivity vs quality. The 2002 distinction matters: "research" measured as raw counts of papers behaves differently from "research" measured by quality or impact. A blanket statement that "research and teaching are unrelated" slightly overstates it; the honest version is "research output is unrelated to teaching; research quality is weakly, positively related."
  • Ecological and aggregation issues. Correlations computed across individuals within a department can differ from those computed across departments or institutions. The nexus can also be positive at the curricular level — research-rich environments can enrich what is taught — even when it is zero at the individual level. The finding is about individuals, not about whether research and teaching should be institutionally linked.
  • Range restriction. In elite research universities where almost everyone is research-active, restricted variance in research output can attenuate any correlation. The near-zero result is most secure across the full sector, less so within a narrow band.

How Koji incorporates this

Koji for Education is designed so that the evaluation measures teaching, and only teaching, as cleanly as possible — and so that reputational halos have less room to operate.

  • Multidimensional, behaviour-anchored measurement. Following the Marsh multidimensionality tradition, Koji breaks "was this good teaching?" into distinct, low-inference dimensions — clarity of explanation, organisation, feedback usefulness, intellectual challenge, inclusivity — using scale, single_choice and yes_no items. A profile of specific teaching behaviours is far harder for a "distinguished professor" halo to inflate uniformly than a single global-impression score.
  • Open-text and AI-moderated probing that stays on teaching. Koji's AI moderator asks students to describe what the instructor did — a concrete example of an explanation that helped, a piece of feedback they used — rather than to render an overall verdict. Automatic thematic analysis then surfaces teaching-specific themes, keeping the evidence anchored to classroom practice rather than reputation.
  • Separation of evidence streams. Koji is a teaching-evaluation instrument and reports it as such. It does not blend research metrics into a teaching score or present evaluations as a general "academic quality" index — reflecting the nexus finding that these are independent constructs that need independent evidence.
  • Triangulation support. Because student ratings are one imperfect window on teaching, Koji is built to sit alongside peer observation and other evidence in a portfolio, rather than to be the sole number — the responsible way to use ratings that the reliability literature demands.

Koji is careful not to overclaim: no instrument can guarantee a reviewer will not import an outside reputation into how they read a report. But by producing a specific, teaching-anchored, multidimensional profile rather than a single haloable number, Koji is designed to make the teaching evidence stand on its own. Teams that also run research and product studies use Koji's core platform at koji.so, where the same discipline applies — measure the construct you actually care about, and do not let an adjacent one leak in.

A worked example: two CVs, two evaluation profiles

Consider a promotion panel weighing two lecturers. One has forty publications and mediocre, thinly-detailed evaluations; the other has six publications and a rich, consistently strong teaching profile across several cohorts. The nexus evidence says the panel must resist two tempting but invalid inferences: that the prolific researcher is "probably fine" in the classroom despite the scores, and that the strong teacher's light publication record signals a weaker academic. Neither inference is supported — research output and teaching quality simply do not travel together at the individual level. The correct move is to evaluate each strand on its own evidence: the teaching case rests on the evaluation profile, peer observation and teaching materials; the research case rests on the quality and impact of the work, not merely its volume (recall that Marsh and Hattie found research quality, not quantity, carried what little positive signal exists). A panel that keeps the two ledgers separate will make better and fairer decisions than one that lets a strong showing in one column quietly vouch for the other. This is not pedantry; it is the direct operational consequence of a near-zero correlation.

Related resources

References

  • Hattie, J., & Marsh, H. W. (1996). The Relationship Between Research and Teaching: A Meta-Analysis. Review of Educational Research, 66(4), 507–542. https://doi.org/10.3102/00346543066004507
  • Marsh, H. W., & Hattie, J. (2002). The Relation Between Research Productivity and Teaching Effectiveness: Complementary, Antagonistic, or Independent Constructs? The Journal of Higher Education, 73(5), 603–641. https://doi.org/10.1080/00221546.2002.11777170

Related articles

best-practices

Should Student Evaluations Decide Tenure? The Ryerson Arbitration and the Limits of High-Stakes SET

A landmark 2018 Canadian arbitration ruled that student evaluations of teaching should not be used to measure teaching effectiveness for promotion and tenure. This guide explains the decision, the evidence behind it, and what it means for governing the high-stakes use of course-evaluation data.

research-methods

What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback

Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.

research-methods

The Halo Effect in Course Evaluations: When One Impression Colours Every Rating

When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.

research-methods

Stop Asking "Was the Lecturer Clear?": The Case for Low-Inference Teaching-Behaviour Items

High-inference items like "Is the instructor clear?" tell teachers nothing about what to change. Murray's low-inference behaviour research shows how concrete, observable items make evaluations diagnostic.