New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
best-practices9 min read

Evaluating Graduate Teaching Assistants: Why GTA Feedback Needs Its Own Rules

Students judge graduate teaching assistants against a different mental template than professors, so benchmarking GTA scores on faculty norms is unfair and misleading. Here is what the research shows and how to evaluate GTA-led teaching properly.

Koji Education Team

Product

In brief. Graduate teaching assistants (GTAs) run a large share of the seminar, laboratory and problem-class teaching that undergraduates actually experience, yet most institutions evaluate them with the same instrument and the same norm-referenced benchmarks used for professors. Research shows students perceive GTAs and professors through different trait vocabularies — GTAs as approachable, relatable and interactive but also more hesitant and uncertain; professors as knowledgeable and authoritative but more distant and formal. Comparing a GTA's scores against faculty norms therefore measures the role stereotype as much as the teaching. GTA evaluation should be formative-first, instructor-type-aware, and decoupled from high-stakes faculty comparisons.

The blind spot

Walk into most European universities' quality-assurance systems and you will find a carefully governed process for evaluating the courses professors lead — and a patchwork, often an afterthought, for the doctoral researchers and teaching assistants who deliver tutorials, mark coursework, and supervise labs. GTAs frequently teach the small-group sessions where undergraduates ask their real questions, yet their teaching is either invisible to the evaluation system or folded into the parent course's score, where it cannot be seen. When GTAs are evaluated, they are usually rated on the identical form used for tenured staff and compared against the same departmental averages.

That is a measurement error waiting to happen, because students do not judge a 26-year-old doctoral tutor and a full professor on the same scale in their heads — even when the paper scale is identical.

What the research says

The most directly relevant evidence comes from Kendall and Schussler (2012), who studied how undergraduates perceive graduate teaching assistants versus professors. In Does Instructor Type Matter? Undergraduate Student Perception of Graduate Teaching Assistants and Professors (CBE—Life Sciences Education, 11(2), 187–199), they combined validated subscales with an open-ended item asking students to compare the same class taught by a professor versus a GTA. Both the quantitative and qualitative results showed that some perceptions attach to the instructor type, not the class type. Professors were described as confident, in control, organized, experienced, knowledgeable — but also distant, formal, strict, "hard," even boring, and respected. GTAs drew a strikingly different vocabulary: engaging, approachable, informal, relaxed, interactive, relatable, understanding, and able to personalize their teaching — while simultaneously being seen as more hesitant, nervous, and uncertain.

Read carefully, this is not a story about GTAs being "worse." It is a story about students applying two different mental templates. A professor and a GTA can deliver the same session and elicit systematically different descriptors — and, plausibly, different numeric ratings on items about "authority," "expertise," or "confidence" — purely because of the role the student perceives.

Kendall and Schussler's companion study, Evolving Impressions: Undergraduate Perceptions of Graduate Teaching Assistants and Faculty Members over a Semester (CBE—Life Sciences Education, 2012, 11(4), 365–376), added a temporal dimension: students' initial, stereotype-driven impressions shifted over a term as they accumulated direct experience. First impressions of a GTA are not fixed; they move with exposure. That has a sharp implication for when GTA evaluation is collected — an early snapshot may capture the stereotype more than the teaching.

The wider context comes from Park (2004), The graduate teaching assistant (GTA): lessons from North American experience (Teaching in Higher Education, 9(3), 349–361). Park reviews the GTA literature and stresses the ambiguity of the GTA role — part student, part teacher, part apprentice academic — and the associated issues of identity, self-worth, training, supervision and mentoring. The role ambiguity Park describes is precisely why a purely summative, comparative evaluation misfires: the GTA is still learning to teach, often with minimal preparation, under employment and status conditions that differ from those of faculty. Evaluating them as if they were finished practitioners misreads the developmental stage they are in.

Why it matters for course evaluation in practice

Four consequences for a quality-assurance office:

  1. Do not benchmark GTAs against faculty norms. If your reporting compares a GTA's dimension scores to the departmental average — an average built largely from professors — you are partly measuring the instructor-type stereotype Kendall and Schussler documented. Any comparison should be within-role (GTA vs GTA on comparable sessions), if it is made at all.

  2. Lead with formative purpose. Given the developmental stage Park describes, the primary function of GTA evaluation should be to help a novice teacher improve, not to rank them. Mid-cycle feedback, delivered while the GTA can still act on it, is far more useful than an end-of-term summative score that arrives after the teaching is over.

  3. Mind the timing. Because impressions of GTAs evolve over a semester, an evaluation collected too early captures the stereotype; one collected only at the end misses the chance to help. A formative mid-point plus a light end-point is better aligned with how GTA perceptions actually form.

  4. Solve the attribution problem. In a large course, a professor may deliver lectures while several GTAs run parallel labs or tutorials. Rolling everything into one course score hides the GTA contribution and can unfairly credit or blame the wrong person. Evaluation needs to attribute experience to the session and instructor the student actually had, not to the course as an undifferentiated whole — the same multi-instructor attribution challenge that arises in team-taught courses.

Limitations and honest caveats

  • Discipline and country specificity. Kendall and Schussler studied biology students in a North American setting; Park's review is explicitly North American. European GTA arrangements vary enormously — from formally employed wissenschaftliche Hilfskräfte to informal demonstrators — and the strength of the instructor-type stereotype may differ across cultures and disciplines. Treat the direction of the finding as robust and the exact magnitude as context-dependent.
  • Perception is not effectiveness. That students describe GTAs and professors differently does not establish that either teaches better. The trait vocabularies are perceptions, potentially confounded with age, gender, status and expectation. This literature documents a lens, not a ranking.
  • Selection and confounds. GTAs are typically assigned the small-group, interactive formats; professors the large lectures. Some of the "approachable vs distant" contrast may reflect the format rather than the person — though Kendall and Schussler's design specifically tried to separate instructor type from class type.
  • Small numbers, fragile statistics. A single GTA may teach two tutorial groups of fifteen. Dimension means from such small cohorts carry wide uncertainty, and the temptation to over-interpret a decimal difference is strong. Reliability guidance for small classes applies with full force.
  • Stereotypes cut both ways. The GTA "approachable" halo can inflate warmth ratings just as the professor "authority" halo can inflate competence ratings. Neither is a clean measure of teaching quality, and bias-aware interpretation is essential.

How Koji incorporates this

Koji is designed to evaluate teaching at the level students actually experience it, which makes GTA-aware evaluation a native capability rather than a workaround.

  • Session- and instructor-level attribution. Koji evaluations can be scoped to a specific tutorial, lab or seminar and the person who taught it, so a GTA's teaching is captured as its own signal instead of being averaged into the parent course. This directly addresses the attribution problem in multi-instructor courses.
  • Formative, mid-cycle collection. Because Koji supports lightweight mid-semester collection, GTAs can receive actionable feedback while there is still time to change their teaching — aligning with the developmental, improvement-first purpose the research recommends and with evidence on mid-semester consultation.
  • AI-moderated interviews that probe past the stereotype. Rather than a Likert item about "confidence" that a role stereotype can contaminate, Koji's conversational interviewer asks students what specifically helped their learning in the session and follows up for concrete examples. Probing for behaviour ("the tutor drew the diagram step by step") yields evidence about teaching rather than a rating filtered through the "GTA" or "professor" template.
  • Role-aware, bias-aware reporting. Koji's reporting is built to compare like with like and to flag when a score may reflect instructor-type or warmth stereotypes rather than teaching effectiveness, so a department does not rank a novice tutor against seasoned professors on a single unqualified number.
  • A developmental record. Structured questions (open_ended, scale, single_choice) plus thematic analysis of open text let a GTA and their supervisor track improvement across a term and across cohorts — supporting the mentoring relationship Park identifies as central.

The same conversational-interview engine powers Koji's core research platform at koji.so, where teams evaluating people at different career stages face the same trap of comparing everyone against a single benchmark.

FAQ

Should GTAs be evaluated on the same form as professors? The instrument can overlap, but the interpretation and benchmarking should not. Comparing GTA scores to faculty-dominated norms measures the instructor-type stereotype as much as the teaching. Evaluate within role and lead with formative intent.

When should GTA feedback be collected? Favour a formative mid-semester point, because student impressions of GTAs evolve with exposure and early ratings capture the stereotype. A lighter end-of-term collection can supplement it, but end-only feedback misses the chance to help a developing teacher.

Isn't a GTA's lower "confidence" score just accurate? Not necessarily. Students describe GTAs as more hesitant partly because of the role they perceive, not only because of observed behaviour. Probing for specific teaching behaviours separates genuine development needs from stereotype.

How do we stop a GTA's teaching being hidden inside the course score? Scope evaluation to the session and instructor the student actually had, rather than the course as a whole. This is the same attribution fix needed for team-taught courses and prevents the GTA contribution from being averaged away.

Should GTA evaluations feed employment decisions? Use extreme caution. Given small cohorts, wide uncertainty, developmental stage and stereotype effects, GTA ratings are poorly suited to high-stakes personnel decisions and are best used for coaching and development.

Related Resources

References

  • Kendall, K. D., & Schussler, E. E. (2012). Does instructor type matter? Undergraduate student perception of graduate teaching assistants and professors. CBE—Life Sciences Education, 11(2), 187–199. https://doi.org/10.1187/cbe.11-10-0091
  • Kendall, K. D., & Schussler, E. E. (2012). Evolving impressions: Undergraduate perceptions of graduate teaching assistants and faculty members over a semester. CBE—Life Sciences Education, 11(4), 365–376. https://doi.org/10.1187/cbe.12-07-0110
  • Park, C. (2004). The graduate teaching assistant (GTA): Lessons from North American experience. Teaching in Higher Education, 9(3), 349–361. https://doi.org/10.1080/1356251042000216660