New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Do University Teachers Get Better With Experience? A 13-Year Growth-Model Study Says: Not Automatically

Herbert Marsh's multilevel growth model of student evaluations over 13 years found teachers' ratings neither improve nor decline with experience — individual differences dominate. What this means for tenure assumptions, development policy, and how evaluation data should track trajectories.

Koji Education Team

Product

Answer first

No — not automatically. The largest longitudinal study of the question, Herbert Marsh''s 13-year multilevel growth-model analysis of students'' evaluations of teaching, found that on average university teachers'' ratings neither improve nor deteriorate with experience: teachers who start strong stay strong, teachers who start weak stay weak, and the differences between teachers dwarf any change within a teacher over time. Teaching improves when institutions intervene — feedback, consultation, structured training — not through the passage of years. For quality assurance, this overturns a comfortable default assumption: seniority is not evidence of teaching quality, and "they''ll grow into it" is not a development plan.

What the research says

The anchor study is Marsh (2007), "Do university teachers become more effective with experience? A multilevel growth model of students'' evaluations of teaching over 13 years" (Journal of Educational Psychology, 99(4), 775–790). Marsh applied multilevel growth modelling — the appropriate technique for repeated measures nested within teachers — to over a decade of class-average student evaluations at a research university, asking whether individual teachers'' ratings trend upward as they accumulate experience.

The findings were strikingly null in the most informative way:

  • No systematic growth or decline. Average ratings were essentially flat across 13 years of a teacher''s career window studied.
  • Stable individual differences. Teachers differed substantially and consistently from one another; a teacher''s standing relative to peers was highly stable across years and across the different classes they taught.
  • Starting point did not predict trajectory. Initial status was uncorrelated with rate of change — weaker starters did not systematically catch up, which is precisely the pattern you would expect if unstructured experience alone produced learning.

This converges with the older correlational literature. Feldman (1983), "Seniority and experience of college teachers as related to evaluations they receive from students" (Research in Higher Education, 18(1), 3–124), synthesized dozens of cross-sectional studies and found seniority and experience mostly unrelated to global student evaluations — and where associations existed, age and years of experience tended to be inversely related to ratings, weakly but robustly across controls. Cross-sectional designs confound cohort and period effects, which is exactly what Marsh''s within-teacher growth model fixed — and the longitudinal answer matched: experience per se does not buy perceived teaching quality.

The contrast case makes the null result meaningful. Where researchers study structured development rather than raw experience, change appears: Gibbs and Coffey (2004) found teachers undergoing initial pedagogical training across 22 universities improved on SEEQ-rated teaching skills while untrained controls stagnated or declined. Improvement is possible; it just is not automatic.

Why it matters for course evaluation in practice

Seniority-based quality assumptions are empirically unfounded. Programme reviews sometimes discount poor evaluations for senior staff ("decades of experience") or excuse them for juniors ("still learning — scores will rise"). Marsh''s data support neither: junior teachers'' relative standing tends to persist, and senior teachers'' ratings are no better for their years. Evaluation evidence should be read at face value regardless of career stage, with uncertainty quantified, not discounted by tenure.

Waiting is not a development strategy. If initial status does not predict growth, a struggling teacher identified in year one will, absent intervention, likely still be struggling in year five — with thousands of student-experiences accumulated in between. The policy implication is early, active intervention: consultation on evaluation results, structured training, peer observation. The feedback-intervention literature shows ratings-plus-consultation moves scores where ratings alone do not.

Stability is also good news for measurement. The high between-teacher, low within-teacher variance structure means class-average ratings, aggregated sensibly, measure something reliable about a teacher''s practice rather than noise. Ironically, the same stability that indicts experience as a development mechanism validates evaluations as a measurement instrument — a teacher''s profile is a persistent signal, worth acting on.

Trajectories, not snapshots, should be the reporting unit. If the expected within-teacher trend is flat, then any genuine sustained change — improvement after a course redesign, decline after workload shocks — is diagnostic precisely because it is unusual. Evaluation systems should be built to detect and contextualise trajectory breaks, not just report this semester''s mean.

Limitations and honest caveats

Marsh''s study, though methodologically exemplary, is a single-institution analysis: one university, one instrument (SEEQ-tradition ratings), one national context, and a 13-year window that samples mid-career more densely than the first fragile semesters. Growth in the first year or two of teaching — where training studies concentrate — may be real yet largely outside the window. Student evaluations measure perceived teaching quality; a null trend in ratings is not proof that no pedagogical learning occurs, only that students do not detect it — though for a perception-based QA system, that distinction is cold comfort. Survivorship is a live concern: teachers who left (perhaps the weakest) are underrepresented in long panels, which would, if anything, bias toward finding improvement — making the null more, not less, credible. Feldman''s synthesis aggregates heterogeneous cross-sectional studies with the usual confounds. And none of this literature covers the post-2020 teaching environment; whether experience with new modalities (hybrid, AI-augmented) follows the same flat trajectory is untested.

How Koji incorporates this

Koji''s reporting and collection design assumes what this literature shows: improvement must be engineered and detected, not presumed.

  • Trajectory-first reporting. Koji tracks evaluation results for the same course and teacher across cycles, so QA staff see within-teacher trends against the flat baseline the research predicts — a sustained upward break after an intervention is visible evidence, not anecdote.
  • Career-stage-neutral interpretation. Because seniority does not predict ratings, Koji''s comparative views contextualise results by course characteristics (level, size, requiredness) rather than by the teacher''s rank — the confounds that actually move scores.
  • Diagnostic depth for intervention, not just monitoring. A flat score describes; it does not explain. Koji''s AI-moderated conversational interviews probe which specific behaviours students experience — organisation, feedback timeliness, clarity — producing the actionable, dimension-level material that consultation-based improvement requires, rather than a global number a teacher cannot act on.
  • Formative mid-cycle collection. Since improvement follows intervention, Koji supports mid-semester formative studies so a teacher can adjust within the same cohort that reported the problem — converting evaluation from an annual verdict into the feedback loop the growth literature says is the actual mechanism of change.
  • Honest uncertainty. Small classes produce noisy means; Koji''s reporting is designed to keep chance fluctuation from being read as a trajectory. A genuine trend must persist across cycles before it is treated as signal.

Institutional researchers studying their own staff development programmes can run the same longitudinal designs on Koji''s core research platform at koji.so, which applies the identical interview and analysis engine to any research population.

Why the growth model settles what cross-sections could not

It is worth pausing on why Marsh''s design is the decisive one, because committees still routinely reason from cross-sectional comparisons. Comparing this year''s senior professors with this year''s new lecturers confounds three things that a snapshot cannot separate: age effects (what experience does to a teacher), cohort effects (senior staff were hired under different criteria, into a different academic labour market, than today''s juniors), and survival effects (the seniors you observe are the ones who stayed — perhaps disproportionately the successful ones). A cross-section finding seniors rated equal to juniors is compatible with experience helping, hurting, or doing nothing, depending on how hiring and attrition changed.

A multilevel growth model dissolves the confound by following each teacher against themselves: every teacher contributes their own trajectory, and the model estimates the average within-person slope while letting both starting points and slopes vary across people. Three quantities then carry the substantive story. The mean slope answers "does the average teacher trend up?" — Marsh''s answer: no. The variance of intercepts answers "do teachers genuinely differ?" — yes, substantially and stably, which is what makes ratings informative about persons at all. The intercept–slope correlation answers "do weak starters close the gap?" — no, which is the finding with the sharpest policy edge, because it removes the comfortable assumption that time heals weak teaching.

For institutional researchers, there is a practical lesson beyond the substantive one: if your evaluation warehouse stores teacher-linked results across semesters, you can estimate the same three quantities for your own institution rather than importing Marsh''s single-site result on faith. A flat local mean slope with fat, stable between-teacher variance replicates the finding and justifies trajectory-based reporting; a positive local slope in, say, the first four semesters would be actionable evidence that your onboarding support works. Either way, the analysis converts an archive that mostly services annual reports into an empirical answer to a governance question institutions usually settle by anecdote.

Related Resources

References

  • Marsh, H. W. (2007). Do university teachers become more effective with experience? A multilevel growth model of students'' evaluations of teaching over 13 years. Journal of Educational Psychology, 99(4), 775–790. https://doi.org/10.1037/0022-0663.99.4.775
  • Feldman, K. A. (1983). Seniority and experience of college teachers as related to evaluations they receive from students. Research in Higher Education, 18(1), 3–124. https://doi.org/10.1007/BF00992080
  • Gibbs, G., & Coffey, M. (2004). The impact of training of university teachers on their teaching skills, their approach to teaching and the approach to learning of their students. Active Learning in Higher Education, 5(1), 87–100. https://doi.org/10.1177/1469787404040463

Related articles