Do University Teachers Get Better With Experience? A 13-Year Growth-Model Study Says: Not Automatically
Herbert Marsh's multilevel growth model of student evaluations over 13 years found teachers' ratings neither improve nor decline with experience — individual differences dominate. What this means for tenure assumptions, development policy, and how evaluation data should track trajectories.
Koji Education Team
Product
Answer first
No — not automatically. The largest longitudinal study of the question, Herbert Marsh''s 13-year multilevel growth-model analysis of students'' evaluations of teaching, found that on average university teachers'' ratings neither improve nor deteriorate with experience: teachers who start strong stay strong, teachers who start weak stay weak, and the differences between teachers dwarf any change within a teacher over time. Teaching improves when institutions intervene — feedback, consultation, structured training — not through the passage of years. For quality assurance, this overturns a comfortable default assumption: seniority is not evidence of teaching quality, and "they''ll grow into it" is not a development plan.
What the research says
The anchor study is Marsh (2007), "Do university teachers become more effective with experience? A multilevel growth model of students'' evaluations of teaching over 13 years" (Journal of Educational Psychology, 99(4), 775–790). Marsh applied multilevel growth modelling — the appropriate technique for repeated measures nested within teachers — to over a decade of class-average student evaluations at a research university, asking whether individual teachers'' ratings trend upward as they accumulate experience.
The findings were strikingly null in the most informative way:
- No systematic growth or decline. Average ratings were essentially flat across 13 years of a teacher''s career window studied.
- Stable individual differences. Teachers differed substantially and consistently from one another; a teacher''s standing relative to peers was highly stable across years and across the different classes they taught.
- Starting point did not predict trajectory. Initial status was uncorrelated with rate of change — weaker starters did not systematically catch up, which is precisely the pattern you would expect if unstructured experience alone produced learning.
This converges with the older correlational literature. Feldman (1983), "Seniority and experience of college teachers as related to evaluations they receive from students" (Research in Higher Education, 18(1), 3–124), synthesized dozens of cross-sectional studies and found seniority and experience mostly unrelated to global student evaluations — and where associations existed, age and years of experience tended to be inversely related to ratings, weakly but robustly across controls. Cross-sectional designs confound cohort and period effects, which is exactly what Marsh''s within-teacher growth model fixed — and the longitudinal answer matched: experience per se does not buy perceived teaching quality.
The contrast case makes the null result meaningful. Where researchers study structured development rather than raw experience, change appears: Gibbs and Coffey (2004) found teachers undergoing initial pedagogical training across 22 universities improved on SEEQ-rated teaching skills while untrained controls stagnated or declined. Improvement is possible; it just is not automatic.
Why it matters for course evaluation in practice
Seniority-based quality assumptions are empirically unfounded. Programme reviews sometimes discount poor evaluations for senior staff ("decades of experience") or excuse them for juniors ("still learning — scores will rise"). Marsh''s data support neither: junior teachers'' relative standing tends to persist, and senior teachers'' ratings are no better for their years. Evaluation evidence should be read at face value regardless of career stage, with uncertainty quantified, not discounted by tenure.
Waiting is not a development strategy. If initial status does not predict growth, a struggling teacher identified in year one will, absent intervention, likely still be struggling in year five — with thousands of student-experiences accumulated in between. The policy implication is early, active intervention: consultation on evaluation results, structured training, peer observation. The feedback-intervention literature shows ratings-plus-consultation moves scores where ratings alone do not.
Stability is also good news for measurement. The high between-teacher, low within-teacher variance structure means class-average ratings, aggregated sensibly, measure something reliable about a teacher''s practice rather than noise. Ironically, the same stability that indicts experience as a development mechanism validates evaluations as a measurement instrument — a teacher''s profile is a persistent signal, worth acting on.
Trajectories, not snapshots, should be the reporting unit. If the expected within-teacher trend is flat, then any genuine sustained change — improvement after a course redesign, decline after workload shocks — is diagnostic precisely because it is unusual. Evaluation systems should be built to detect and contextualise trajectory breaks, not just report this semester''s mean.
Limitations and honest caveats
Marsh''s study, though methodologically exemplary, is a single-institution analysis: one university, one instrument (SEEQ-tradition ratings), one national context, and a 13-year window that samples mid-career more densely than the first fragile semesters. Growth in the first year or two of teaching — where training studies concentrate — may be real yet largely outside the window. Student evaluations measure perceived teaching quality; a null trend in ratings is not proof that no pedagogical learning occurs, only that students do not detect it — though for a perception-based QA system, that distinction is cold comfort. Survivorship is a live concern: teachers who left (perhaps the weakest) are underrepresented in long panels, which would, if anything, bias toward finding improvement — making the null more, not less, credible. Feldman''s synthesis aggregates heterogeneous cross-sectional studies with the usual confounds. And none of this literature covers the post-2020 teaching environment; whether experience with new modalities (hybrid, AI-augmented) follows the same flat trajectory is untested.
How Koji incorporates this
Koji''s reporting and collection design assumes what this literature shows: improvement must be engineered and detected, not presumed.
- Trajectory-first reporting. Koji tracks evaluation results for the same course and teacher across cycles, so QA staff see within-teacher trends against the flat baseline the research predicts — a sustained upward break after an intervention is visible evidence, not anecdote.
- Career-stage-neutral interpretation. Because seniority does not predict ratings, Koji''s comparative views contextualise results by course characteristics (level, size, requiredness) rather than by the teacher''s rank — the confounds that actually move scores.
- Diagnostic depth for intervention, not just monitoring. A flat score describes; it does not explain. Koji''s AI-moderated conversational interviews probe which specific behaviours students experience — organisation, feedback timeliness, clarity — producing the actionable, dimension-level material that consultation-based improvement requires, rather than a global number a teacher cannot act on.
- Formative mid-cycle collection. Since improvement follows intervention, Koji supports mid-semester formative studies so a teacher can adjust within the same cohort that reported the problem — converting evaluation from an annual verdict into the feedback loop the growth literature says is the actual mechanism of change.
- Honest uncertainty. Small classes produce noisy means; Koji''s reporting is designed to keep chance fluctuation from being read as a trajectory. A genuine trend must persist across cycles before it is treated as signal.
Institutional researchers studying their own staff development programmes can run the same longitudinal designs on Koji''s core research platform at koji.so, which applies the identical interview and analysis engine to any research population.
Why the growth model settles what cross-sections could not
It is worth pausing on why Marsh''s design is the decisive one, because committees still routinely reason from cross-sectional comparisons. Comparing this year''s senior professors with this year''s new lecturers confounds three things that a snapshot cannot separate: age effects (what experience does to a teacher), cohort effects (senior staff were hired under different criteria, into a different academic labour market, than today''s juniors), and survival effects (the seniors you observe are the ones who stayed — perhaps disproportionately the successful ones). A cross-section finding seniors rated equal to juniors is compatible with experience helping, hurting, or doing nothing, depending on how hiring and attrition changed.
A multilevel growth model dissolves the confound by following each teacher against themselves: every teacher contributes their own trajectory, and the model estimates the average within-person slope while letting both starting points and slopes vary across people. Three quantities then carry the substantive story. The mean slope answers "does the average teacher trend up?" — Marsh''s answer: no. The variance of intercepts answers "do teachers genuinely differ?" — yes, substantially and stably, which is what makes ratings informative about persons at all. The intercept–slope correlation answers "do weak starters close the gap?" — no, which is the finding with the sharpest policy edge, because it removes the comfortable assumption that time heals weak teaching.
For institutional researchers, there is a practical lesson beyond the substantive one: if your evaluation warehouse stores teacher-linked results across semesters, you can estimate the same three quantities for your own institution rather than importing Marsh''s single-site result on faith. A flat local mean slope with fat, stable between-teacher variance replicates the finding and justifies trajectory-based reporting; a positive local slope in, say, the first four semesters would be actionable evidence that your onboarding support works. Either way, the analysis converts an archive that mostly services annual reports into an empirical answer to a governance question institutions usually settle by anecdote.
Related Resources
- Longitudinal stability of student ratings
- Regression to the mean in year-over-year evaluation changes
- Do student evaluations improve teaching? Feedback and consultation evidence
- Adjunct vs tenure-track instructor evaluations
- Student evaluations in tenure and promotion decisions
- Mid-semester feedback and consultation: the meta-analytic case
References
- Marsh, H. W. (2007). Do university teachers become more effective with experience? A multilevel growth model of students'' evaluations of teaching over 13 years. Journal of Educational Psychology, 99(4), 775–790. https://doi.org/10.1037/0022-0663.99.4.775
- Feldman, K. A. (1983). Seniority and experience of college teachers as related to evaluations they receive from students. Research in Higher Education, 18(1), 3–124. https://doi.org/10.1007/BF00992080
- Gibbs, G., & Coffey, M. (2004). The impact of training of university teachers on their teaching skills, their approach to teaching and the approach to learning of their students. Active Learning in Higher Education, 5(1), 87–100. https://doi.org/10.1177/1469787404040463
Related articles
Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations
Year-over-year movements in course-evaluation scores are mostly statistical noise regressing toward an instructor's true average. Here is why regression to the mean fools quality-assurance committees, and how to read evaluation trends responsibly.
Do Student Evaluations Actually Improve Teaching? The Feedback-Intervention Evidence
Student ratings handed back as a bare number barely change teaching (d ≈ 0.2). Paired with structured consultation, the same data produces moderate, durable improvement (d ≈ 0.6–0.7). What the intervention meta-analyses tell quality teams about closing the loop.
Should Student Evaluations Decide Tenure? The Ryerson Arbitration and the Limits of High-Stakes SET
A landmark 2018 Canadian arbitration ruled that student evaluations of teaching should not be used to measure teaching effectiveness for promotion and tenure. This guide explains the decision, the evidence behind it, and what it means for governing the high-stakes use of course-evaluation data.
Do Adjuncts Get Worse Course Evaluations Than Tenured Faculty? Employment Status as a Confound
The causal evidence says contingent faculty do not teach worse — and often teach better. Why employment status is a poor proxy for teaching quality, and how to read evaluation scores fairly across the tenure divide.