Learning Analytics vs Student Ratings: Can Behavioural Data Replace the Course Evaluation?
Every click, login and video-scrub is now logged. It is tempting to think this behavioural exhaust is a more objective substitute for the course-evaluation survey — no self-report, no bias, no non-response. That intuition is wrong in an instructive way. Learning analytics answers a different question than a course evaluation, predicts learning only weakly and unstably, and carries a heavier ethical load. It belongs beside the survey, not in its place.
Koji Education Team
Product ·
The short answer: Learning analytics — the data trail students leave in the virtual learning environment (VLE/LMS) — is a genuinely useful complement to course evaluation, but it cannot replace it. Behavioural logs measure engagement proxies (clicks, logins, time-on-page), which predict academic performance only weakly and unstably, tell you nothing about why a course worked or failed, and carry a materially heavier data-protection and ethics burden than a consented survey. The mature position is not "dashboards over surveys" but "behaviour tells you where to look; asking students tells you what it means." Use analytics to target your evaluation, not to retire it.
The seductive pitch — and what is wrong with it
The argument for replacing surveys with behavioural data is easy to state and superficially compelling. Course evaluations suffer from non-response bias, social desirability, recall error and every rating-scale artefact in the methodology canon. Log data, by contrast, is a census — every student, every session, automatically captured, no questionnaire fatigue. Why ask people what they did when you can just see what they did?
Because seeing what they did is not the same as knowing what they learned, and it is a long way from knowing whether the teaching was any good. Three problems make behavioural data a poor substitute for — though a fine partner to — the course evaluation.
Problem one: engagement metrics predict learning weakly and unstably
The foundational hope of learning analytics is that behaviour predicts outcomes: more engagement, more learning. The evidence is far more equivocal than the dashboards suggest. Reviews of VLE log data find that the variance in final grades explained by log features ranges widely — from roughly 8% to 37% across studies — and that simple counts like total clicks and total logins are among the weakest predictors. What predicts performance better is not volume of activity but quality of it: timely submission, consistency, self-regulation. Crucially, the models are highly context-sensitive and port badly from one course or platform to another, so a predictor that works in one module can fail in the next (Journal of Computers in Education, 2025).
This matters for evaluation because it breaks the core assumption. If you wanted to use engagement as a proxy for teaching quality, you would need engagement to reliably track learning, and learning to reliably track teaching. Both links are weak and unstable. Worse, high engagement can signal the opposite of good teaching: a student who logs 40 hours in the VLE may be struggling with badly-designed material, not thriving. A confused cohort clicks a lot. Behaviour is ambiguous without the meaning students attach to it.
Problem two: analytics cannot answer the evaluation question
A course evaluation, at its best, answers a causal-explanatory question: what about this course helped or hindered learning, and what should change? Behavioural logs are structurally incapable of answering it. They can tell you that video views collapsed in week 7. They cannot tell you why — whether the topic was harder, the recording was inaudible, the assessment deadline pulled attention elsewhere, or the content was so clear students did not need to rewatch. Every one of those explanations implies a different action, and the log is silent on which is true.
This is the difference between a symptom and a diagnosis. Analytics is excellent at flagging symptoms at scale — an early-warning system for disengagement. Diagnosis requires asking the people involved. This is also why learning analytics and course evaluation measure different constructs and should not be scored against each other: one is behavioural trace, the other is elicited experience and judgement. Treating the former as a replacement for the latter is a category error — closer to replacing a patient interview with a pedometer.
Problem three: the ethics and data-protection load is heavier, not lighter
There is a comfortable assumption that behavioural data is ethically cleaner than a survey because it is "just system logs." Under European data-protection norms the reverse is closer to the truth. Continuous behavioural monitoring of identifiable students engages the GDPR in ways a consented, purpose-limited evaluation survey does not: questions of lawful basis, purpose limitation, transparency, profiling and the risk of function creep, where data gathered to "support learning" migrates into surveillance or into judgements about staff.
Europe has thought hard about this. The cross-European SHEILA project — built on interviews with 78 senior managers across 51 higher-education institutions in 16 countries — produced a policy framework precisely because privacy, ethics, transparency and student agency are unresolved governance challenges, not solved ones (SHEILA policy framework). You cannot bolt learning analytics onto evaluation as a frictionless upgrade; you inherit a governance problem that a well-designed, consented, GDPR-compliant survey largely avoids.
But doesn't behavioural data at least fix non-response and honesty bias?
This is the strongest argument for the analytics-first position, and it deserves a fair hearing. Yes — log data is a census, so it sidesteps non-response bias, and because it is not self-reported it is immune to social-desirability and recall effects. Those are real advantages, and any honest comparison must grant them.
But notice what the trade buys and what it costs. You escape self-report bias by giving up access to the one thing self-report provides: the student's interpretation of their own experience. You trade a known, well-studied set of survey biases — which the field has decades of tools to mitigate — for a newer, less understood set of analytics pitfalls: proxy invalidity (clicks are not learning), model instability across contexts, algorithmic bias in early-warning systems that can disadvantage exactly the students who behave differently for legitimate reasons (carers, part-time workers, students with disabilities, those using their own offline notes). "Objective" behavioural data is not bias-free; it relocates the bias from the response to the model and the metric, where it is harder to see.
A second fair point: analytics is timely in a way an end-of-term survey is not — it flags problems while the course is running. Granted. But that is an argument for mid-cycle evaluation, not for abandoning evaluation. The timeliness advantage is best captured by moving your asking earlier, not by replacing asking with watching. We make that case in why experience sampling beats retrospective recall.
The mature model: analytics targets, evaluation explains
The two data sources are complements with a natural division of labour:
- Learning analytics is the smoke detector. It scans the whole cohort cheaply and continuously and says "look here" — this module, this week, this at-risk group.
- Course evaluation is the diagnosis. Once analytics has narrowed the field, you ask the students in that module, that week, what actually happened and what to change.
Used this way, analytics makes evaluation sharper — you can target a conversational follow-up at the exact point where engagement dropped, rather than surveying everyone about everything at the end and drowning the signal. This is triangulation done properly: two different kinds of evidence, each doing what it is good at.
How Koji fits the complement model
Koji for Education is designed to be the diagnosis layer that behavioural data cannot provide. Where analytics flags that something happened, Koji's AI-moderated conversational interviews find out why — probing beyond a number to capture the student's own account of what confused them, what helped, and what should change. Because these interviews can run mid-cycle, they slot neatly behind an analytics early-warning signal: the dashboard says week 7 disengagement, Koji asks that cohort about week 7 while they can still remember and while you can still act.
Automatic thematic analysis turns those open-text explanations into structured, programme-level insight at scale — the interpretive layer logs lack — and closing-the-loop action tracking records what you changed in response. The AI moderation is standardised and bias-aware, and the whole pipeline is GDPR/AVG-compliant with consented, purpose-limited data collection — which is precisely the governance profile that continuous behavioural surveillance struggles to match. Koji does not pretend to read minds from clicks; it asks, consistently and at scale, and makes the answers actionable.
Legacy analytics-heavy strategies chase the illusion that enough behavioural exhaust adds up to understanding. It does not: it adds up to symptoms. Understanding still comes from asking well — which is exactly what Koji's conversational engine, shared with the main Koji research platform used for wider user and staff research, is built to do.
If your institution is investing in dashboards, invest equally in the layer that makes them meaningful. See how Koji for Education turns behavioural signals into explanations you can act on.