New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends8 min read

Your Workload and Difficulty Ratings Are Now Measuring ChatGPT

When most students use generative AI on coursework, the "workload" and "difficulty" items on your course evaluation stop measuring the course and start measuring how much AI a student used. This is a live construct-validity threat — here is why it matters and how to evaluate around it.

Koji Education Team

Product · July 15, 2026

BLUF: Two of the most consequential items on a standard course evaluation — perceived workload and perceived difficulty — quietly assume that the effort a student spends is a property of the course. Since 2023 that assumption has broken. When a large majority of students now use generative AI to summarise readings, draft assignments, and explain concepts, the hours a student reports spending are partly a function of how much AI they used, not how demanding the course is. Two students in identical seats can now honestly report very different workloads. That is textbook construct-irrelevant variance, and it silently corrupts the workload-calibration, difficulty-benchmarking, and ECTS-audit uses that institutions build on those items. This piece sets out the scale of the shift, why it is a measurement problem rather than a moral panic, the strongest counterargument, and how to evaluate courses in a world where the student and the AI are a team.

The shift, in numbers

Student use of generative AI has moved from fringe to near-universal in roughly two academic years. The Higher Education Policy Institute's 2025 survey of 1,041 UK undergraduates found that the share using AI for assessments jumped from 53% to 88% in a single year, and that the share using any AI tool rose from 66% to 92% (HEPI/Kortext Student Generative AI Survey 2025). The most common uses — explaining concepts, summarising articles, suggesting ideas — are precisely the activities that used to constitute course workload. When AI compresses the reading, the summarising, and the first draft, the hours a diligent student logs fall, even though the course has not changed.

This matters because "hours spent" and "how difficult did you find this" are not idle questions. They feed real decisions.

What those items are actually used for

The workload and difficulty items are load-bearing in ways that make their corruption expensive:

  • ECTS calibration. In the European Credit Transfer and Accumulation System, credits are defined by notional student workload. Institutions routinely use evaluation-reported hours to check whether a module's real workload matches its credit allocation — the exact use we examined in auditing ECTS workload with course evaluation. If reported hours now embed an unknown, uneven dose of AI assistance, the audit is measuring the wrong thing.
  • Difficulty benchmarking. Difficulty ratings are compared across cohorts and years to flag modules that are too easy or too punishing. A year-on-year drop in perceived difficulty may now reflect AI adoption, not a change in the course — a confound that will fool any naive trend line.
  • Interpreting satisfaction. Difficulty and workload are standard covariates when people try to make sense of satisfaction scores. If the covariate is contaminated, so is the adjustment.

Why this is construct-irrelevant variance, not a moral panic

The instinct is to frame student AI use as cheating and reach for detection. That is a different (and largely losing) battle. The measurement point is narrower and more robust: a course evaluation is supposed to measure the course. When the workload item's variance is driven substantially by how much AI each student chose to use, the score reflects something other than the intended construct — the definition of construct-irrelevant variance in Messick's validity framework. Crucially, this is true even for entirely legitimate, permitted AI use. A student who uses an AI tutor to understand a proof faster is not cheating; they are simply making the "hours spent" item mean something different from what it meant in 2019.

The result is a loss of measurement invariance: the same questionnaire no longer measures the same latent quantity across students or across years. Comparisons that assume it does — the cross-cohort, cross-module, cross-time comparisons universities run constantly — become unsafe.

The counterargument: "workload was always self-reported and noisy"

The fair objection is that self-reported workload was never precise. Students misremember hours; recall is biased; the item always had wide error bars. So why single out AI?

Because AI does not add noise — it adds bias with structure. Random measurement error averages out across a cohort; a systematic, adoption-correlated reduction in reported effort does not. Worse, AI use is not evenly distributed. It correlates with confidence, discipline, language background, and access, so the contamination falls unevenly across exactly the subgroups equity-minded evaluation is supposed to protect. A noisy-but-unbiased item can still support valid group comparisons; a biased item whose bias tracks student characteristics cannot. The pre-2023 workload item was noisy. The post-2023 item is noisy and differentially biased — a categorically harder problem.

A second objection: "then just ask students to report effort excluding AI." Useful, but limited — students genuinely cannot partition their cognition that cleanly, and the instruction invites social-desirability distortion. Better to redesign what we ask.

How to evaluate when the student and the AI are a team

The productive response is not detection but redesign:

  • Ask about AI use directly and non-judgementally. Make "how, and how much, did you use AI tools in this course?" a first-class evaluation item, framed as legitimate. You cannot control for a variable you refuse to measure. This turns the confound into data.
  • Shift from hours to learning behaviours and outcomes. "How many hours did you spend" is now ambiguous; "which activities deepened your understanding, and where did you get stuck" is not. Anchor items to specific learning processes and to whether stated learning outcomes were met, which AI assistance does not automatically inflate.
  • Probe the quality of struggle, not just its quantity. Desirable difficulty — productive struggle — is the thing worth measuring. A conversational follow-up can distinguish "AI removed busywork so I could think harder about the hard part" from "AI let me skip the thinking entirely." A single difficulty number cannot.
  • Re-baseline trends. Treat 2023 as a structural break. Difficulty and workload series that straddle it should not be compared as if the instrument were stable.

A concrete example

Consider a quantitative-methods module rated "very heavy" for workload for years. In the latest cohort its workload score drops sharply and its satisfaction rises. The tempting reading — the teaching team finally fixed the pacing — may be wrong. If AI now handles the routine computation and code-debugging that used to consume students' evenings, the felt workload fell without any change to the module itself. Acting on the number — cutting content because "students found it lighter and enjoyed it more" — could quietly hollow out the very rigour that the drop concealed. The lesson is not to distrust every apparent improvement, but to ask, before acting on any workload or difficulty trend, a new prior question: did the course change, or did the instrument?

Where Koji fits

A static Likert form cannot tell whether a low workload number means an easy course or a heavy AI user. Koji for Education is designed for exactly this ambiguity. Its AI-moderated conversational interviews probe behind the number: when a student reports light workload, the moderator can ask what they actually did, where they used AI, and what they still found hard — turning an uninterpretable rating into structured, thematically analysed evidence. AI use becomes a measured, reportable dimension rather than an invisible confound. The six question types (open-ended, scale, single-choice, multiple-choice, ranking, yes/no) let programmes ask directly about tool use while still capturing the nuance in students' own words, and automatic thematic analysis surfaces patterns — "students used AI to compress the reading but struggled with the applied problem set" — that no fixed scale would reveal. Koji frames this honestly: it surfaces and helps interpret the AI confound; it does not pretend to strip AI back out of a self-reported hour.

The same interpretive problem — self-reported effort and behaviour distorted by tools — shows up across user and customer research, which is why teams use the main Koji platform and its shared AI interview engine to understand not just what people report but how they actually did it.

The takeaway

Do not quietly keep trending your workload and difficulty items as though nothing happened. In a cohort where 88% of students use AI on assessments, those items now measure a blend of course demand and tool adoption. Name the confound, measure AI use directly, move from hours to learning, and treat pre-2023 baselines with suspicion. The course did not necessarily get easier — your instrument started measuring something new.

Want course evaluation that can tell an easy course from a heavy AI user? Explore Koji for Education.