New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

The Critical Incident Technique for Course Feedback

How Flanagan's Critical Incident Technique collects concrete, behaviourally-anchored student feedback that global Likert ratings cannot capture, and how it applies to course evaluation.

Koji Education Team

Product

In brief: The Critical Incident Technique (CIT) is a systematic method, formalised by John C. Flanagan in 1954, for collecting concrete, first-hand accounts of specific behaviours that made a decisive difference to an outcome — rather than abstract global ratings. Applied to course evaluation, CIT asks students to describe a specific moment when teaching helped or hindered their learning, yielding actionable, behaviourally-anchored evidence that a five-point Likert average cannot provide.

What the research says

The Critical Incident Technique was introduced to the social sciences by John C. Flanagan in his landmark 1954 monograph in Psychological Bulletin (Flanagan, 1954). Flanagan defined the technique not as a single rigid procedure but as "a set of procedures for collecting direct observations of human behavior in such a way as to facilitate their potential usefulness in solving practical problems and developing broad psychological principles." The defining move is that CIT collects incidents — concrete descriptions of behaviour observed in a specific situation — and treats those incidents as the raw data, deliberately avoiding the abstraction and generalisation that global rating scales impose on the respondent.

Origins in aviation psychology

CIT did not begin in a classroom. Its methodological roots lie in the United States Army Air Forces Aviation Psychology Program of the Second World War, which Flanagan directed. Wartime studies sought to understand why aircrew selection and combat performance so often failed to match expectations. The famous early insight was that when pilots washed out of flight training, the recorded reasons were frequently useless clichés ("lack of inherent flying ability", "poor judgement") that offered nothing an instructor could act on. Investigators therefore began collecting specific, observed instances of effective and ineffective behaviour — a bombing run, a disorientation error, a specific decision under stress. Aggregating these concrete incidents revealed the actual behavioural requirements of the job far more reliably than global evaluations did. Flanagan's 1954 paper consolidated this wartime work and a subsequent decade of studies at the American Institute for Research into a general, transferable methodology.

Flanagan's five steps

Flanagan described CIT as a flexible procedure organised around five stages, which remain the canonical framework:

  1. Determine the general aim of the activity. Before collecting anything, the researcher establishes the purpose of the activity being studied — for teaching, this is a working statement of what effective instruction is meant to achieve. This aim anchors what counts as "critical".
  2. Establish plans and specifications for collecting incidents. The researcher decides who the observers will be, what situations qualify, and the criteria that make an incident "critical" — i.e. one that made a significant positive or negative contribution to the aim.
  3. Collect the data. Incidents are gathered through interviews, questionnaires, or observation. Each incident must include the context, the specific behaviour, and the observed consequence. Flanagan stressed factual recall of a concrete event over opinion or evaluation.
  4. Analyse the data. Incidents are sorted inductively into categories to summarise and describe the content efficiently. This coding step is where the practical output — a behaviourally-anchored taxonomy — emerges.
  5. Interpret and report the requirements. The researcher reports the resulting categories and their implications, being explicit about the judgements made and the limitations of the data.

An "incident", in Flanagan's usage, is any observable human activity sufficiently complete to permit inferences about the person performing it; it becomes "critical" when it occurs in a situation where its purpose is reasonably clear and its consequences are definite enough to leave little doubt about its effects.

Corroborating and extending sources

The most cited modern appraisal of the method is Butterfield, Borgen, Amundson and Maglio's (2005) review, Fifty years of the critical incident technique: 1954–2004 and beyond, published in Qualitative Research. The authors trace CIT's evolution from Flanagan's largely quantitative, positivist origins toward a broadly accepted qualitative research method, and they codify criteria for methodological rigour and trustworthiness (credibility checks such as independent coding, exhaustiveness of categories, and participation rates). Their review is essential reading for anyone who wants to defend a CIT study against the charge that it is "just anecdotes."

In education specifically, David Tripp's Critical Incidents in Teaching: Developing Professional Judgement (1993) reframes the critical incident as a tool for reflective practice. Tripp's important clarification is that "critical" does not mean dramatic or serious — a critical incident may be an entirely ordinary classroom event that becomes critical because the analyst interprets it as significant. This interpretive stance broadened CIT from a job-analysis instrument into a lens for professional learning.

For a direct empirical application to course evaluation, Khandelwal's (2009) study in the International Journal of Teaching and Learning in Higher Education collected 237 critical incidents from 60 undergraduate students across three humanities courses and sorted them into six behavioural categories distinguishing excellent from poor teaching: rapport with students, course preparation and delivery, encouragement, fairness, spending time with students outside class, and control. Crucially, the output was a set of specific behaviours faculty could adopt or avoid — precisely the actionable specificity that global ratings fail to deliver.

Why it matters for course evaluation in practice

Conventional student evaluations of teaching lean heavily on Likert-type items ("The instructor was well organised: 1–5"). These produce tidy numbers but a thin evidential base. CIT matters for three practical reasons.

Concrete incidents beat abstract ratings. A score of 3.8 on "organisation" tells an instructor almost nothing about what to change. A student's account — "In week 6 the lecture jumped between three topics with no roadmap, and I couldn't tell which was examinable" — is diagnosable and fixable. CIT deliberately keeps the data at the level of observable events, so the evidence points at specific, changeable behaviour rather than a latent trait.

It reduces halo effects. Global ratings are notoriously contaminated by halo — a single strong impression (charisma, likeability, an easy grading reputation) bleeds across every dimension, so "organisation", "fairness" and "knowledge" scores move together regardless of the underlying reality. Anchoring feedback to a specific remembered event forces the respondent back to a concrete situation, which constrains the tendency to let one global affect colour every judgement. CIT does not eliminate bias, but by demanding an incident it makes the halo shortcut harder to take.

It surfaces actionable specifics and the "long tail". Rating averages compress and hide. A course can score respectably on average while a specific, fixable failure — an inaccessible assessment brief, a demoralising piece of feedback, a lab that consistently overruns — recurs across many students. CIT surfaces these recurring concrete problems, and the frequency of an incident category becomes a natural prioritisation signal for closing the loop.

Limitations and honest caveats

CIT is a serious method, and a serious appraisal must state its weaknesses plainly.

Recall and memory bias. CIT is retrospective. Students report incidents from memory, and memory is reconstructive: salience, recency and emotional intensity distort what is recalled and how it is described. A single distressing exam may dominate recall of an otherwise strong semester. The incidents collected are therefore a biased sample of what actually happened, weighted toward the memorable rather than the representative.

Retrospective selection and self-selection. Respondents choose which incident to report, and those who choose to respond are not a random sample. Highly satisfied and highly dissatisfied students are over-represented; the ambivalent middle stays quiet. Any frequency count of incident categories must be read as a description of what was volunteered, not a population estimate of what occurred.

Coding subjectivity and inter-rater reliability. The analysis stage is interpretive. Deciding whether an incident belongs in "fairness" or "encouragement", and whether two accounts describe the same underlying behaviour, involves judgement. Without explicit coding protocols and a reported measure of inter-rater agreement, the resulting categories can reflect the analyst as much as the data. Butterfield et al. (2005) treat independent coding and category exhaustiveness as core credibility requirements precisely because of this risk.

Generalisability. CIT yields rich, context-bound description, not statistically representative estimates. Findings from 60 students in three humanities courses (as in Khandelwal, 2009) illuminate those courses; they do not license claims about engineering cohorts or other institutions without further work. CIT answers "what specifically happened and why did it matter" far better than "how common is this across the population."

Effort and response burden. Writing a specific, contextualised incident is more cognitively demanding than ticking a box. This raises the risk of lower completion, terse non-incidents ("it was fine"), and fatigue when many incidents are requested. Instrument design has to manage this burden, or response quality degrades and the method's advantage evaporates.

How Koji incorporates this

Koji for Education is designed to operationalise CIT's core insight — elicit a specific incident, not a number — while mitigating (not eliminating) the limitations above. The mapping to real Koji mechanisms is as follows.

  • AI-moderated conversational interviews that probe for a specific incident. Rather than presenting a bare Likert grid, Koji's interviewer is designed to ask for a concrete moment — "Tell me about a specific moment when a lecture either helped or blocked your understanding" — and then follow up conversationally to recover the context, the behaviour, and the consequence, which are exactly the three elements Flanagan required of a well-formed incident. This is a design intended to counter the "it was fine" non-incident and the halo shortcut, not a guarantee against them.
  • Structured question types alongside the open probe. Koji supports open_ended, scale, single_choice, multiple_choice, ranking, and yes_no items, so a CIT-style open incident probe can be triangulated with lighter closed items in the same instrument — balancing depth against response burden.
  • Automatic thematic analysis of open text. Koji clusters incidents into recurring themes automatically, mirroring Flanagan's step 4 (inductive categorisation) at scale, and giving reviewers a first-pass taxonomy to inspect rather than code from scratch.
  • Quality scoring of responses. Koji scores response quality, which helps flag terse or non-substantive answers — a direct response to the response-burden and non-incident risks noted above.
  • Bias-aware reporting. Reporting is designed to foreground that retrospective, self-selected incident data describes what was volunteered, not a population estimate — keeping frequency counts honestly framed.
  • Triangulation across cohorts. Comparing incident themes across cohorts and time helps distinguish a recurring, structural problem from a one-off memorable event, partly offsetting recall and selection bias.
  • Closing-the-loop action tracking. Because incidents are concrete and behaviourally anchored, they convert cleanly into tracked actions, so institutions can record what changed in response to a specific recurring incident.

The same AI-moderated interview engine underpins Koji's core research platform at koji.so, where it is applied to product and customer research — the CIT principle of eliciting a specific incident rather than an abstract rating is equally powerful when a product team wants to know exactly when a user got stuck, not merely how satisfied they are on average.

None of this makes CIT bias-free. It is a structured attempt to capture better data at the source and analyse it more transparently — which is the most any honest evaluation method can claim.

References

  • Butterfield, L. D., Borgen, W. A., Amundson, N. E., & Maglio, A.-S. T. (2005). Fifty years of the critical incident technique: 1954–2004 and beyond. Qualitative Research, 5(4), 475–497. https://doi.org/10.1177/1468794105056924
  • Flanagan, J. C. (1954). The critical incident technique. Psychological Bulletin, 51(4), 327–358. https://doi.org/10.1037/h0061470
  • Khandelwal, K. A. (2009). Effective teaching behaviors in the college classroom: A critical incident technique from students' perspective. International Journal of Teaching and Learning in Higher Education, 21(3), 299–309. https://eric.ed.gov/?id=EJ909053
  • Tripp, D. (1993). Critical incidents in teaching: Developing professional judgement. London: Routledge.

Related Resources