New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Beyond the Student Survey: Structured Classroom-Observation Protocols (COPUS and TDOP) as Evaluation Evidence

What COPUS and the TDOP measure that student surveys and traditional peer visits cannot: low-inference, reliable records of what actually happens in a classroom. Their evidence base, their limits, and where they fit in a triangulated evaluation.

Koji Education Team

Product

In brief. Structured classroom-observation protocols such as COPUS (the Classroom Observation Protocol for Undergraduate STEM) and the TDOP (Teaching Dimensions Observation Protocol) record what is actually happening in a class — in two-minute time slices, using fixed behavioural codes, without asking the observer to judge quality. This makes them a fundamentally different kind of evidence from a student survey (which captures perception) or a traditional peer visit (which captures one colleague''s impression). Used alongside student feedback, they answer a question a Likert scale cannot: not "did students like it?" but "how were instructor and students spending their time?" They do not measure learning, and they are labour-intensive — but as triangulation, they are among the most defensible evidence a QA process can hold.

The gap they fill

A course-evaluation programme built only on student ratings has a structural blind spot: it measures perception of teaching, filtered through everything from instructor warmth to lecture fluency to grade expectations. Traditional peer observation adds a second view but is notoriously unreliable — one colleague''s holistic "that was a good class" is a high-inference judgement shaped by rapport, discipline norms, and what the observer values. Structured observation protocols were built to fix the reliability problem by removing judgement from the recording step entirely.

What the research says

COPUS was introduced by Smith, Jones, Gilbert and Wieman (2013) in CBE—Life Sciences Education. Its design goals are explicit and directly relevant to evaluation: after only about 1.5 hours of training, faculty observers — not education specialists — can reliably characterise how instructors and students spend class time. The protocol uses 25 codes in two categories ("What the students are doing" and "What the instructor is doing"), applied in two-minute intervals across a session. Crucially, it "does not require observers to make judgments of teaching quality" and produces clear graphical output. In the validation, seven STEM faculty observers, working in pairs after brief training, produced reliable codings of classroom practice — the low-inference design is what makes the inter-rater reliability achievable.

COPUS was itself built on the earlier and more granular TDOP (Teaching Dimensions Observation Protocol), developed by Hora and colleagues at the Wisconsin Center for Education Research. The TDOP scores the presence or absence of 47 codes across multiple dimensions of instruction at two-minute intervals, offering a richer but more demanding record. The lineage matters: COPUS deliberately traded the TDOP''s detail for trainability and reliability, on the reasoning that a protocol only useful to specialists cannot scale to routine departmental evaluation.

Later work connected these descriptive records to instructional quality. Lund et al. (2015), again in CBE—Life Sciences Education, analysed 269 class periods from 73 faculty at 28 research-intensive institutions and derived ten empirically-grounded "COPUS profiles", validating their alignment with reformed, active-learning practice using RTOP (the Reformed Teaching Observation Protocol) scores. This is important for evaluators: it turns raw COPUS counts into interpretable teaching styles, and cross-validates them against an independent quality-oriented instrument.

The through-line across this literature is the distinction between low-inference and high-inference measurement — the same principle behind the case for low-inference items in student surveys. "The instructor lectured for 38 of 50 minutes; students worked in groups for 4" is low-inference and reproducible. "The teaching was engaging" is high-inference and observer-dependent. Structured protocols are built almost entirely from the former.

Why it matters for course evaluation in practice

1. It breaks the single-method monopoly. Student ratings are a single source with well-documented biases; common-method bias means you cannot fix that by adding more survey items. COPUS/TDOP data come from a different method and a different observer, giving genuine triangulation rather than more of the same.

2. It provides formative, non-threatening feedback. Because the output describes behaviour rather than judging it, COPUS gives instructors a mirror ("you lectured 80% of the time you thought was interactive") without a verdict. Smith et al. explicitly designed it as a low-threat professional-development tool, which raises the odds that faculty act on it.

3. It supplies summative evidence student ratings cannot. For a tenure or accreditation file, "an independent, reliably-coded observation shows this instructor uses active-learning strategies in X% of class time" is evidence of practice — complementary to a teaching portfolio and far harder to dismiss than a popularity-influenced mean. It aligns naturally with active-learning frameworks like ICAP.

4. It documents change. Because the protocol is standardised, before/after observations can credibly evidence that a teaching intervention actually altered classroom practice — not just that students reported liking it more.

Limitations and honest caveats

  • Describes practice, not learning — and not quality directly. COPUS tells you what happened, not whether students learned or whether what happened was good. "Active learning" is associated with better outcomes on average, but a COPUS profile is not itself an achievement measure. Overreading counts as quality is the central misuse.
  • Reactivity (the observer effect). An instructor teaches differently when watched. A single observed session may not represent the term, and knowing the codes can prompt performative "active" behaviour.
  • Sampling. One or two 50-minute snapshots is a tiny sample of a course. Reliable characterisation needs multiple observations, which multiplies cost.
  • Labour-intensive. Even at 1.5 hours training per observer, coding every instructor every term is expensive relative to a survey that emails itself. This is the practical reason it complements rather than replaces student feedback.
  • Discipline and modality fit. COPUS was built for STEM lecture/lab settings; its codes fit a seminar, a studio, or a fully online asynchronous course less cleanly. The TDOP is more customisable but more demanding.
  • What gets counted gets valued. If a department rewards "minutes of group work", instructors optimise for the code, not the learning — a Campbell''s-Law risk that applies to any behavioural metric.

The honest position: structured observation is powerful evidence of practice and a strong reliability upgrade on casual peer visits, but it is one leg of a triangle — student perception, observed practice, and learning evidence — not a standalone verdict.

How Koji incorporates this

Koji does not send a human into the room with a stopwatch — structured observation is an in-person method and stays that way. What Koji does is make observation data usable as part of a triangulated evaluation and strengthen the perception side of the triangle so the two methods genuinely complement each other.

  • A home for triangulated evidence. Koji is designed to hold student-perception data (surveys and AI-moderated interviews) alongside imported observation results, so a programme review or accreditation file presents perception and observed practice together rather than in separate silos.
  • Lower-inference student questions. Mirroring the COPUS philosophy, Koji favours specific, behaviour-anchored open_ended and single_choice items ("How was class time usually spent?") over vague global judgements — narrowing, though never closing, the gap between what students perceive and what an observer would record.
  • Probing discrepancies. When observation shows heavy lecturing but students rate the course "highly interactive", that gap is diagnostic. Koji''s AI-moderated interviews can explore why students perceive it that way — the kind of qualitative reconciliation that turns two disagreeing methods into insight rather than a contradiction to be argued away.
  • Thematic analysis mapped to practice. Koji''s automatic thematic analysis of open text can be organised around the same behavioural dimensions a protocol tracks (instructor talk, student activity, feedback, questioning), so perception and observation speak a shared vocabulary.

Framed precisely: Koji is designed to complement structured observation and to make the student-perception leg of the triangle as behaviourally grounded as possible — not to replace the trained observer, and not to claim survey data can substitute for watching a class. Koji''s core research platform at koji.so applies the same triangulation logic to product research, pairing behavioural/usage evidence with AI-moderated interviews for the "why".

Related resources

References

  • Smith, M. K., Jones, F. H. M., Gilbert, S. L., & Wieman, C. E. (2013). The Classroom Observation Protocol for Undergraduate STEM (COPUS): A New Instrument to Characterize University STEM Classroom Practices. CBE—Life Sciences Education, 12(4), 618–627. https://doi.org/10.1187/cbe.13-08-0154
  • Lund, T. J., Pilarz, M., Velasco, J. B., Chakraverty, D., Rosploch, K., Undersander, M., & Stains, M. (2015). The Best of Both Worlds: Building on the COPUS and RTOP Observation Protocols to Easily and Reliably Measure Various Levels of Reformed Instructional Practice. CBE—Life Sciences Education, 14(2), ar18. https://doi.org/10.1187/cbe.14-10-0168
  • Hora, M. T., & Ferrare, J. J. (2013). Instructional Systems of Practice: A Multidimensional Analysis of Math and Science Undergraduate Course Planning and Classroom Teaching. Journal of the Learning Sciences, 22(2), 212–257. https://doi.org/10.1080/10508406.2012.729931
  • Sawada, D., Piburn, M. D., Judson, E., Turley, J., Falconer, K., Benford, R., & Bloom, I. (2002). Measuring Reform Practices in Science and Mathematics Classrooms: The Reformed Teaching Observation Protocol. School Science and Mathematics, 102(6), 245–253. https://doi.org/10.1111/j.1949-8594.2002.tb17883.x