Does Evaluating Course After Course Change How Students Answer? Panel Conditioning in Repeated Course Evaluations
Students at a European university complete dozens of course evaluations across a degree. Panel-conditioning research shows that the mere act of being surveyed repeatedly can change later answers — a threat to comparing scores across years and cohorts.
Koji Education Team
Product
In short
Repeatedly surveying the same students can change the answers they give — independently of any change in the thing being measured. This is called panel conditioning, and the strongest evidence for it comes from Halpern-Manners, Warren and Torche's (2017) analysis of the U.S. General Social Survey, which found statistically significant conditioning on a subset of attitudinal, knowledge and behavioural items. In a European degree, a student may complete 30–50 course evaluations before graduating, so their tenth evaluation is not psychologically equivalent to their first. For quality-assurance offices this is a quiet threat to the comparability of scores across years, cohorts and study stages — but it is measurable and, in part, manageable.
What the research says
Panel conditioning is the phenomenon in which prior experience of being surveyed alters a respondent's later answers. The change is not caused by anything real happening to the respondent; it is caused by the survey itself. The mechanism can be cognitive (the respondent has thought about the topic more, so their attitude has "crystallised"), motivational (they now know the survey is long and satisfice to finish quickly), or behavioural (they change the underlying behaviour because being asked about it prompted reflection).
The most rigorous single demonstration is Halpern-Manners, Warren and Torche (2017), published in Sociological Methods & Research. They exploited the rotating-panel structure of the U.S. General Social Survey (GSS), comparing two groups who were both interviewed in 2008 but differed in whether that 2008 interview was their first or their second. Because respondents were assigned to panels in a way that equated their propensity to stay in the study, the only systematic difference between the groups was prior survey experience. Across roughly 310 variables, and after correcting for multiple comparisons with a false-discovery-rate adjustment, they found conditioning effects on a meaningful subset: 8 items were significant at p < 0.05 and 19 at p < 0.10. Experienced respondents answered factual knowledge questions more accurately (for example, several science-literacy items were 4–6 percentage points higher) and gave systematically different answers on some attitudinal and household items. Individual effects were modest — typically in the single-to-low-double-digit percentage-point range — but they were real and non-random.
The GSS finding does not stand alone. Warren and Halpern-Manners (2012), in a review in the same journal, catalogue the ways longitudinal surveys are vulnerable and argue conditioning is under-diagnosed because most panels lack the design needed to detect it. Sturgis, Allum and Brunton-Smith (2009) describe the "attitude crystallisation" pathway specifically: repeated interviewing appears to increase the reliability (test–retest stability) of attitude reports, which sounds desirable but actually means the panel is drifting away from the fresh, less-rehearsed population it is supposed to represent. More recent work by Kraemer and colleagues (2025), again in Sociological Methods & Research, revisits whether attitude change over waves is real change or an artefact of repeated interviewing, and continues to find measurable conditioning on some items. The consistent conclusion across three decades is not that conditioning is huge and universal — it is that conditioning is selective, item-dependent, and easy to mistake for genuine change.
Why it matters for course evaluation in practice
A university is, structurally, one of the most intensive repeated-survey environments a young adult will ever enter. A three-year bachelor's programme with five or six modules a year generates 15–35 end-of-module evaluations per student; add mid-module surveys, the national student survey, module-choice surveys and programme reviews, and the count climbs further. By their final year, students are veteran respondents. Panel conditioning predicts three specific problems for course-evaluation data:
-
Cross-year comparisons are confounded. When you compare first-year and final-year modules, you are not only comparing different teaching — you are comparing evaluation novices with evaluation veterans. If veterans satisfice more (straightlining, central-tendency), or conversely give more crystallised and polarised answers, apparent "differences between modules" partly reflect the students' survey history, not the courses.
-
Longitudinal QA dashboards can show drift that is not pedagogical. A programme that tracks the same cohort's satisfaction across years may see a downward or upward trend that is partly conditioning. This is a cousin of regression to the mean and interacts with survey fatigue: both produce non-pedagogical movement in the numbers.
-
The "reliability increases" trap. If your scale's internal consistency improves in later years, that is not automatically good news. Crystallisation can raise reliability while reducing the representativeness of the response — the panel has learned the instrument, not necessarily reported more truthfully. This is why reliability statistics like Cronbach's alpha must be read alongside evidence about who is responding and how, not in isolation.
The practical upshot is that a course-evaluation programme should treat "student survey experience" as a variable, not a constant — much as it already treats response rate or class size as variables worth modelling.
Limitations and honest caveats
A careful reader should not over-apply this. Several caveats matter:
-
The flagship evidence is from a general social survey, not course evaluation. No large study has yet isolated panel conditioning specifically in university course-evaluation data, so we are reasoning by analogy from a well-designed adjacent literature. The direction and mechanisms transfer plausibly, but the magnitude in the course-evaluation context is genuinely unknown.
-
Effect sizes are modest and item-specific. In the GSS work, most items showed no conditioning at all. Conditioning is not a wholesale invalidation of repeated measurement; it is a selective bias on particular kinds of items (factual, sensitive, or attitudinally "loose" items). Global "overall satisfaction" items may behave differently from specific behavioural ones.
-
Conditioning and maturation are hard to separate. Students genuinely change over a degree — they become more discerning, more knowledgeable about what good teaching looks like. Distinguishing that real maturation from survey-induced conditioning requires a rotation-group or fresh-refreshment design that most institutions do not run.
-
Anonymity complicates detection. The GSS design relied on linking the same respondents across waves. Course evaluations are usually anonymous by policy, which is ethically right but makes person-level conditioning analysis difficult. You can look at cohort-level patterns, but you cannot always follow the individual.
Showing these limitations is not hedging — it is the point. Panel conditioning is a reason to interpret longitudinal course-evaluation trends cautiously, not a licence to dismiss any trend you dislike.
How Koji incorporates this
Koji for Education is designed to reduce the conditions under which conditioning does its damage, and to make it visible when it occurs. Concretely:
-
AI-moderated conversational interviews break the rote-response habit. A large part of conditioning is respondents learning the shape of a fixed questionnaire and answering on autopilot. Koji's evaluations are conducted as adaptive, AI-moderated conversations that probe beyond a Likert number rather than presenting the identical grid every term. Because the follow-up probes are generated from what the student actually says, a veteran respondent cannot pattern-match their way through — which is designed to mitigate (not eliminate) the satisficing pathway of conditioning.
-
Structured question types with rotation. Koji supports open_ended, scale, single_choice, multiple_choice, ranking and yes_no items, and can vary item order and framing across administrations. Varying presentation is one of the few practical defences against the "learned instrument" form of conditioning.
-
Bias-aware, cohort-segmented reporting. Koji's analytics let a QA office segment results by study stage and cohort, so that a final-year dip can be inspected for whether it tracks with survey-experience rather than teaching. Making the confound visible is the first step to not being fooled by it.
-
Formative, mid-cycle collection reduces total survey load. Over-surveying is what turns students into conditioned veterans in the first place. Koji's mid-cycle and conversational formats are designed to gather richer data from fewer, better-timed touchpoints, easing the survey fatigue that feeds conditioning.
-
Triangulation, not single-source trust. Koji encourages combining conversational evaluation with other evidence rather than treating a repeated Likert item as ground truth — consistent with the research message that no single repeated instrument should stand alone.
For teams who also run general product or customer research, Koji's core research platform at koji.so applies the same AI-moderated interview engine to panels of customers, where longitudinal conditioning is an equally live concern.
Related resources
- Are Student Ratings Just a Mood? The Longitudinal Stability Evidence
- Survey Fatigue: Why Over-Surveying Students Quietly Wrecks Your Response Rates
- Why a Low-Scoring Course Usually "Improves" Next Year: Regression to the Mean
- Did the Teaching Change Actually Work? Interrupted Time Series for Course-Evaluation Trends
- Why Students Click Straight Down the Middle: Satisficing in Course Evaluations
References
- Halpern-Manners, A., Warren, J. R., & Torche, F. (2017). Panel Conditioning in the General Social Survey. Sociological Methods & Research, 46(1), 103–124. https://doi.org/10.1177/0049124114532445
- Warren, J. R., & Halpern-Manners, A. (2012). Panel Conditioning in Longitudinal Social Science Surveys. Sociological Methods & Research, 41(4), 491–534. https://doi.org/10.1177/0049124112460374
- Sturgis, P., Allum, N., & Brunton-Smith, I. (2009). Attitudes Over Time: The Psychology of Panel Conditioning. In P. Lynn (Ed.), Methodology of Longitudinal Surveys (pp. 113–126). Wiley. https://doi.org/10.1002/9780470743874.ch7
- Kraemer, F., Lugtig, P., Struminskaya, B., Silber, H., Weiß, B., Bosnjak, M. (2025). Monitoring Attitudes Over Time: Real Change or the Result of Repeated Interviewing? Sociological Methods & Research. https://doi.org/10.1177/00491241251372503
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations
Year-over-year movements in course-evaluation scores are mostly statistical noise regressing toward an instructor's true average. Here is why regression to the mean fools quality-assurance committees, and how to read evaluation trends responsibly.
Did the Teaching Change Actually Work? Interrupted Time Series for Course-Evaluation Trends
Comparing this year''s evaluation mean to last year''s cannot tell you whether a curriculum redesign worked. Interrupted time series with segmented regression can — here is how to apply it, and where it breaks.
Survey Fatigue: Why Over-Surveying Students Quietly Wrecks Your Response Rates
Porter, Whitcomb & Weitzer (2004) showed that administering multiple surveys in one year suppresses later response rates. A research-grounded guide to survey fatigue in course evaluation — what causes it, what the evidence shows, and how to design around it.