New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Do the Clicks Confirm the Comments? Triangulating Course Evaluations with LMS Learning-Analytics Data

Learning-analytics trace data from your VLE looks like an objective check on what students say in course evaluations. The research says it is a weaker and more course-specific signal than most institutions assume — here is how to triangulate it honestly.

Koji Education Team

Product

In short: LMS trace data (logins, resource views, forum posts, video watch time) is a genuine second source of evidence about a course, but it is not an objective validation of student ratings. Across 17 blended courses with 4,989 students, LMS variables explained between 8% and 37% of the variance in student performance depending on the course (Conijn et al., 2017) — meaning a model that works in one module can be near-useless in the next. Use trace data to generate hypotheses that your course evaluation then explains, not to confirm or overrule what students told you.

Why this question keeps coming up

Every quality-assurance office eventually notices that it holds two datasets about the same course. One is the course evaluation: what students said about clarity, workload, assessment, and support. The other is the virtual learning environment (VLE/LMS) log: what students did — when they logged in, which readings they opened, how far into the lecture capture they watched, whether they posted in the forum.

The temptation is obvious. Course evaluations are self-report, they arrive from a self-selected subset of students, and they are vulnerable to every bias documented elsewhere in this knowledge base. Trace data is a census, it is unobtrusive, and nobody has to remember anything to produce it. So: why not use the logs to check the survey?

Because the logs are measuring something different, less stable, and more dependent on how the course was designed than the enthusiasm for this idea usually admits.

What the research says

Learning analytics are about learning, not about clicks

Gašević, Dawson and Siemens (2015) wrote what remains the field's most-cited corrective. Their argument is that learning analytics drifted toward whatever was easiest to count — logins, clicks, time-on-page — and away from theoretically grounded constructs about how people actually learn. A count of resource views is a proxy for engagement only under assumptions about the course that are rarely stated and frequently false. In a course where all readings are distributed as a PDF pack in week one, low LMS resource-view counts mean nothing about engagement at all.

Their prescription is that trace measures must be interpreted against the instructional conditions of the specific course, not pooled across an institution as if a click meant the same thing everywhere.

The same model does not transfer between courses

The empirical version of that argument came from Gašević, Dawson, Rogers and Gašević (2016), who modelled academic success across nine undergraduate blended courses (n = 4,134). Course-specific models and generalised models produced meaningfully different results: which predictors mattered, and how much they mattered, varied by course. A generalised institution-wide model obscured effects that were real inside particular modules and manufactured apparent effects that were not.

Conijn, Snijders, Kleingeld and Matzat (2017) put a number on the instability. Across 17 blended courses at a single institution (4,989 students, Moodle), the variance in performance explained by LMS predictors ranged from 8% to 37%. Same institution, same platform, same extraction pipeline — a fourfold difference in explanatory power depending on which course you looked at. Their conclusion is the operationally important one: LMS-based prediction is not a portable instrument you can switch on across a portfolio.

And student ratings are not the ground truth either

It is worth being symmetrical about this. Uttl, White and Gonzalez (2017), re-analysing the multisection literature, found that once small-sample and publication-bias artefacts are corrected, student evaluation ratings and student learning are essentially unrelated. So when trace data and evaluation scores disagree, you do not have one valid instrument and one invalid one. You have two partial indicators, each measuring something real and neither measuring "course quality" directly.

That is precisely the case for triangulation rather than arbitration.

Why it matters for course evaluation in practice

Four practical consequences follow.

1. Never use trace data to overrule a student's account. If evaluation comments say the lecture recordings were hard to follow, and the logs show high watch time, the resolution is not "the students are wrong." High watch time is equally consistent with students re-watching because the first pass was incomprehensible. Behaviour is ambiguous; the comment is the disambiguating evidence, not the other way round.

2. Benchmark within a course over time, not across courses. Given the 8%–37% spread, cross-module comparison of engagement metrics is close to meaningless. Comparing this year's cohort against last year's in the same module, with the same design, is defensible.

3. Treat divergence as a research question. The valuable pattern is not agreement — it is disagreement. A module where ratings are strong but forum participation collapsed mid-term has something happening that neither dataset explains alone. That is the module where a follow-up conversation with students earns its cost.

4. Model the instructional condition explicitly. Before interpreting any trace metric, record what the course design implies it should look like. A flipped course and a lecture-plus-reading course should not be expected to produce comparable log signatures, and a dip in one may be by design.

Limitations and honest caveats

A critical reader should press on several points, and we would agree with most of them.

  • Prediction is not explanation. Every study cited above predicts performance, not teaching quality. Even a well-fitting model tells you which students are struggling, not whether the teaching was good. Conflating the two is the central error in this area.
  • Grades are a contaminated criterion. The outcome variable in this literature is course grade, which is itself affected by assessment design, grading standards, and the same leniency dynamics discussed in our work on grading leniency. A model predicting grades inherits every problem the grade has.
  • The findings are platform- and era-specific. Conijn et al. studied Moodle at one Dutch institution in a particular blended configuration. Generalising to a different VLE, a different national context, or a post-2020 teaching mix requires argument, not assumption.
  • Trace data has its own non-response problem. Students who work from downloaded materials, share notes, or use non-institutional tools are invisible to the logs. This is not a census of learning; it is a census of platform-mediated learning.
  • Ethics and data protection are not incidental. Linking identifiable behavioural data to evaluation responses can break the anonymity guarantee the evaluation was collected under. See re-identification and k-anonymity before any such join is contemplated. In most European institutions the defensible design is aggregate-to-aggregate comparison, never record linkage.

How Koji incorporates this

Koji does not ingest VLE logs and does not claim to. What it does is make the other side of the triangle strong enough to be worth triangulating with — because the failure mode in most institutions is not weak analytics, it is an evaluation instrument too thin to explain anything the analytics surfaced.

  • AI-moderated conversational interviews follow up on a rating rather than stopping at it. Where a Likert item records that lecture materials scored 3.4, the interview probes what the student did with them — whether they used the VLE at all, what they substituted, where they got stuck. That is the interpretive layer trace data cannot supply about itself.
  • Structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let you ask directly about study behaviour — which resources were used, in what order, alongside what external material — so behavioural claims are collected as data rather than inferred from clicks.
  • Automatic thematic analysis of open text lets a QA officer test a hypothesis generated from the logs against what students actually wrote, at portfolio scale, without hand-coding thousands of comments.
  • Mid-cycle collection means the divergence between what the logs show and what students report can be surfaced while the module is still running, rather than in a report written after the cohort has left.
  • Bias-aware reporting is designed to keep small-n and confounded comparisons from being read as differences — relevant here because engagement metrics invite exactly the cross-module ranking the research says is unsafe.

These mechanisms are designed to mitigate the interpretive gap in trace data. They do not eliminate it, and no evaluation instrument can substitute for a properly specified analytics model where one is genuinely needed.

Institutions running parallel product or service research will find the same AI-moderated interview engine at koji.so, where the equivalent problem — reconciling behavioural telemetry with what users say — is the everyday case.

Related resources

References

  • Conijn, R., Snijders, C., Kleingeld, A., & Matzat, U. (2017). Predicting student performance from LMS data: A comparison of 17 blended courses using Moodle LMS. IEEE Transactions on Learning Technologies, 10(1), 17–29. https://doi.org/10.1109/TLT.2016.2616312
  • Gašević, D., Dawson, S., & Siemens, G. (2015). Let's not forget: Learning analytics are about learning. TechTrends, 59(1), 64–71. https://doi.org/10.1007/s11528-014-0822-x
  • Gašević, D., Dawson, S., Rogers, T., & Gašević, D. (2016). Learning analytics should not promote one size fits all: The effects of instructional conditions in predicting academic success. The Internet and Higher Education, 28, 68–84. https://doi.org/10.1016/j.iheduc.2015.10.002
  • Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007