New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
best-practices10 min read

Does Pedagogical Training Improve University Teaching? What the Evidence Says — and How to Measure It

Gibbs & Coffey's eight-country study found trained university teachers improved on student ratings while untrained controls did not. What this and the wider instructional-development literature (Postareff et al. 2007; Stes et al. 2010) mean for how universities should evaluate teaching development.

Koji Education Team

Product

Answer first

Yes — the best available longitudinal evidence shows that structured pedagogical training improves university teaching as students experience it, while untrained teachers show no change or even decline. Gibbs and Coffey''s eight-country study found trained teachers improved on student-rated teaching skills and became more student-focused; Finnish follow-up work suggests the effect takes time and short courses may even depress self-ratings before improving them. The practical implication for quality assurance: teaching development programmes are among the few interventions with credible evidence of moving evaluation results — but only institutions that measure before and after, with instruments sensitive to specific behaviours, will ever see it.

What the research says

The anchor study is Gibbs and Coffey (2004), "The Impact of Training of University Teachers on Their Teaching Skills, Their Approach to Teaching and the Approach to Learning of Their Students" (Active Learning in Higher Education, 5(1), 87–100). It followed university teachers across 22 universities in eight countries through their first year of teaching, comparing those enrolled in structured initial training programmes with an untrained comparison group. Measurement was unusually rigorous for the field, triangulating three perspectives: (i) student ratings on six scales of the Student Evaluation of Educational Quality questionnaire (SEEQ) plus the ''Good Teaching'' scale of the Module Experience Questionnaire; (ii) teachers'' own approach to teaching via the Approaches to Teaching Inventory (ATI); and (iii) their students'' approaches to learning (surface vs deep) via MEQ scales.

The findings: trained teachers showed a range of positive changes — improved student-rated teaching skills and a measurable shift from teacher-focused, transmission-oriented teaching toward student-focused teaching — while the untrained control group showed no change or negative changes over the same period. Most strikingly, the students of trained teachers reported less surface learning, linking a staff-development intervention to a change in student learning behaviour, not just satisfaction.

Two corroborating lines strengthen and complicate the picture. Postareff, Lindblom-Ylänne and Nevgi (2007), "The effect of pedagogical training on teaching in higher education" (Teaching and Teacher Education, 23(5), 557–571), studied University of Helsinki staff and found that pedagogical training increased self-efficacy and student-focused approaches — but the effect was dose-dependent: meaningful change appeared only after substantial training (roughly a year of credited study), and short courses were associated with an initial dip in self-efficacy, plausibly because training makes teachers aware of what they do not yet do well. Their follow-up study (Postareff et al., Higher Education, 2008) confirmed the pattern longitudinally.

Third, Stes, Min-Leliveld, Gijbels and Van Petegem (2010), "The impact of instructional development in higher education: The state-of-the-art of the research" (Educational Research Review, 5(1), 25–49), systematically reviewed the impact literature and organised it by outcome level (teacher attitudes, knowledge, behaviour; institutional impact; student outcomes). Their sobering conclusion: most impact studies measure teacher self-report; far fewer measure behaviour; and studies measuring student-level outcomes are rare and often methodologically weak. The evidence that training works is real but concentrated at the lower rungs of the outcome ladder.

Why it matters for course evaluation in practice

European institutions invest heavily in teaching qualifications — the Dutch BKO/UTQ (Basiskwalificatie Onderwijs), UK PGCerts and Advance HE fellowships, Scandinavian mandatory university-pedagogy courses, Irish National Forum programmes. Yet few institutions can answer a basic accountability question: does our programme change what students experience? The research base offers a template and three warnings.

Measure at the student level, or you are measuring satisfaction with the workshop. Stes et al.''s outcome hierarchy maps directly onto Kirkpatrick-style logic: happy-sheet ratings of a training day say nothing about classroom change. Gibbs and Coffey''s design worked because the criterion was the students'' experience (SEEQ scales, learning approaches), collected before and after.

Expect a J-curve, not a step function. The Finnish results imply that evaluating a teacher six weeks after a short course may catch them mid-dip: newly self-critical, experimenting, not yet fluent. Summative use of evaluation scores during a development period risks punishing exactly the teachers who are improving.

Use behaviour-specific instruments. Global "overall teacher" items are noisy and halo-prone; the SEEQ scales Gibbs and Coffey used target specific dimensions (organisation, enthusiasm, rapport, feedback). A development effect concentrated in one dimension can be invisible in a global average.

Limitations and honest caveats

A critical reader should note: Gibbs and Coffey''s comparison group was not randomly assigned — teachers whose institutions required training may differ systematically from those untrained, and institutional context is confounded with training availability. Attrition across a one-year, multi-country panel was substantial, raising selection concerns. Effect sizes were modest, and SEEQ improvements, while student-reported, are still perceptions rather than direct learning measures. Postareff et al.''s work relies heavily on self-report inventories (ATI), vulnerable to social-desirability shift — training teaches you the "right" answers about student-focus. Stes et al.''s review explicitly cautions that the field''s methodological base is thin at the student-outcome level; a definitive randomized, multi-institution trial with learning outcomes as the criterion still does not exist. Finally, all headline studies predate generative AI, flipped formats, and current cohort compositions; transportability to 2026 classrooms is plausible but unproven.

How Koji incorporates this

Koji is built for exactly the measurement design this literature demands — before/after, behaviour-specific, student-level evidence of teaching development.

  • Baseline-and-follow-up by design. Koji studies can run at multiple points in a semester or across semesters, giving development units a genuine pre/post structure around a training programme rather than a single retrospective snapshot.
  • Behaviour-specific structured items. Instead of one global satisfaction number, Koji supports structured questions (scale, single_choice, multiple_choice, ranking, yes_no) mapped to specific teaching behaviours — organisation, feedback quality, clarity — mirroring the dimension-level SEEQ logic under which Gibbs and Coffey detected change.
  • Conversational probing reveals the how. A score shift after training tells you that something changed; Koji''s AI-moderated interviews ask students what changed in the teaching ("what does the instructor do when students seem lost?"), producing thematic evidence that links score movement to concrete practice — the behaviour level Stes et al. found chronically under-measured.
  • Formative framing during the dip. Because Koji supports mid-cycle, low-stakes collection, institutions can designate development-period evaluations as formative — evidence for the teacher and their consultant, not the promotion file — which is designed to mitigate the J-curve unfairness the Finnish studies imply.
  • Triangulation, not verdicts. Koji reporting is designed to sit alongside peer observation and teaching portfolios, echoing Gibbs and Coffey''s three-perspective design; no single-source score is treated as a measure of development impact.

Teaching-development units that also run staff or student research more broadly can use Koji''s core platform at koji.so — the same AI-moderated interview engine applied to any qualitative study, including programme evaluation of the development unit itself.

A minimal evaluation design for your teaching qualification programme

  1. Baseline: dimension-level student evaluation in each participant''s course the semester before training.
  2. Programme records: hours, format, credit volume (the dose matters).
  3. Follow-up: same instrument, same courses where possible, one and three semesters after completion.
  4. Comparison: matched non-participant courses, with the selection caveats above stated openly.
  5. Qualitative layer: AI-moderated student interviews probing the specific behaviours the programme targets.
  6. Report changes with uncertainty intervals per dimension, never a single global delta.

The European qualification landscape and what evidence it owes itself

Europe has quietly built the world''s densest infrastructure of university teaching qualifications: the Dutch BKO/UTQ is a de facto condition of permanent academic employment; Norway and Sweden require documented pedagogical competence for appointment and promotion; the UK routes tens of thousands of staff through PGCerts and Advance HE fellowship categories; Finland''s university pedagogy credits shaped the very studies reviewed above. This is a major, recurring institutional investment — and the research base implies it is largely evaluated at the wrong level. Completion counts and participant satisfaction dominate annual reports; student-experience evidence of the kind Gibbs and Coffey collected is rare in institutional practice, despite being the outcome the qualifications exist to change.

This creates both a risk and an opportunity for QA leaders. The risk: as funding scrutiny tightens, a development unit that cannot connect its programme to any student-level indicator is defending a cost centre with testimonials. The opportunity: the measurement apparatus already exists. Every institution running systematic course evaluation already collects the before/after instrument a Gibbs-and-Coffey-style design needs; what is missing is usually only the linkage — recording which teachers completed which programme when, and reading dimension-level evaluation trajectories around those dates with appropriate humility about selection effects. Institutions that close this loop can answer accreditation panels'' increasingly pointed questions about teaching-quality strategy with evidence rather than inventory, and can iterate their own programmes: if the workload-management module moves the "organisation" dimension but the assessment module never moves "feedback quality," that is actionable curriculum information for the development unit itself. The literature''s dose-dependence finding also gives planners a defensible answer to the perennial pressure to shrink programmes into lunchtime workshops: the evidence says brief formats change awareness, not practice.

Related Resources

References

  • Gibbs, G., & Coffey, M. (2004). The impact of training of university teachers on their teaching skills, their approach to teaching and the approach to learning of their students. Active Learning in Higher Education, 5(1), 87–100. https://doi.org/10.1177/1469787404040463
  • Postareff, L., Lindblom-Ylänne, S., & Nevgi, A. (2007). The effect of pedagogical training on teaching in higher education. Teaching and Teacher Education, 23(5), 557–571. https://doi.org/10.1016/j.tate.2006.11.013
  • Postareff, L., Lindblom-Ylänne, S., & Nevgi, A. (2008). A follow-up study of the effect of pedagogical training on teaching in higher education. Higher Education, 56(1), 29–43. https://doi.org/10.1007/s10734-007-9087-z
  • Stes, A., Min-Leliveld, M., Gijbels, D., & Van Petegem, P. (2010). The impact of instructional development in higher education: The state-of-the-art of the research. Educational Research Review, 5(1), 25–49. https://doi.org/10.1016/j.edurev.2009.07.001