New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Graduate outcomes10 min read

Direct vs Indirect Measures: Why Your Course Evaluation Is Not Evidence of Learning (and What Accreditors Want Beside It)

A course evaluation tells you what students think they learned and how they felt about the course. It does not tell you what they can actually do. That distinction — indirect versus direct evidence of learning — is central to how accreditors read a programme file, and it is where most institutions quietly overreach. The point of a good evaluation is not to prove learning happened. It is to explain why the direct evidence looks the way it does.

Koji Education Team

Product ·

The short answer: In assessment terms, a course evaluation is an indirect measure of learning — it captures students' perceptions of what they learned and their satisfaction with how it was taught. Direct measures capture what students can actually do: exams, portfolios, performances, rubric-scored work. Accreditors and quality frameworks increasingly insist that programmes lead with direct evidence and use indirect evidence only in support, because self-reported learning is a weak proxy for real learning. The mistake is not using course evaluations — it is treating them as proof of learning. Their real job is explanatory: when the direct evidence shows a gap, a good evaluation tells you why.

The distinction accreditors actually use

Assessment professionals draw a bright line between two kinds of evidence, and it is worth stating precisely because so much course-evaluation practice blurs it.

  • Direct measures are demonstrations of learning — student work that shows, rather than reports, achievement of an outcome. Exam answers, coursework marked against a rubric, capstone projects, clinical or lab performances, portfolios, juried exhibitions. The assessor looks at the actual product of learning.
  • Indirect measures are perceptions or proxies for learning — students' opinions about what they learned, satisfaction ratings, self-reported confidence, participation counts, graduate-survey responses. They are evidence that students believe they learned, which is related to, but not the same as, learning (Northern Illinois University CITL).

A standard course evaluation — "How much do you feel you learned?", "Rate the instructor's effectiveness" — sits squarely in the indirect column. It is a perception instrument. And the assessment literature is consistent that indirect measures should not stand alone: most outcomes should be evidenced by a direct measure, with indirect measures used as supporting context, not primary proof (UNLV Office of Academic Assessment). The professional accreditation world is explicit about this: business-school accreditors, for example, treat direct assessment of learning as the backbone of "assurance of learning", with indirect measures such as surveys playing a supporting role (AACSB, 2024).

Why perception is a weak proxy for learning

This is not merely a taxonomic nicety — there is a substantive reason accreditors demand direct evidence. Students are unreliable judges of their own learning.

The multisection meta-analytic literature is the sharpest illustration. When Uttl, White and Gonzalez re-analysed the studies that pair student ratings against a common final exam, they found student ratings are essentially unrelated to how much students actually learned (Uttl, White & Gonzalez, 2017). If the satisfaction signal barely tracks the learning signal, then reporting high evaluation scores as though they demonstrate a programme is achieving its outcomes is, strictly, a non-sequitur.

There is a deeper cognitive reason too, well documented in the feeling-of-learning literature: the conditions that produce the sensation of learning (fluent lectures, ease, entertainment) are not the conditions that produce durable learning (desirable difficulty, retrieval, struggle). Students sometimes rate the courses they learned most from lower, because effortful learning feels worse in the moment. A course evaluation faithfully records the feeling — and the feeling can point the wrong way. Self-reported learning also runs into the self-assessment accuracy problem: the students least equipped in a subject are often least able to judge how much they have mastered.

But surely satisfied students and good scores mean something?

Yes — and this is the counterargument that keeps course evaluations rightly central, so it deserves a straight answer rather than a dismissal. Indirect measures carry real, distinctive value that direct measures cannot supply:

  1. They measure things direct assessment cannot. Belonging, confidence, perceived inclusivity, workload experience, whether students would recommend the programme — these are legitimate outcomes in their own right, and the only valid way to measure a perception is to ask. A programme that produces competent but miserable, disengaged graduates has a real problem that no exam will reveal.
  2. They explain the direct evidence. This is the crucial move. When a rubric-scored capstone shows students consistently fail to synthesise across modules, the exam tells you that there is a gap; only the evaluation — asking students about their experience — tells you why (say, the modules never referenced one another, or the sequencing buried the integrative assignment in an overloaded week). Direct measures locate the problem; indirect measures diagnose its cause.
  3. They are early and cheap. You can gather perception mid-course, before the direct evidence even exists, and act while it still matters.

So the honest position is not "indirect evidence is inferior" but "indirect evidence is answering a different question." The error is only when an institution submits satisfaction scores as if they were evidence of achieved learning outcomes — which is exactly what a thin evaluation report does when it leads an accreditation file with "92% of students agreed the course met its objectives" and calls that outcome assurance.

A second fair concession: direct measures have their own validity problems. A poorly-designed rubric, an unreliable marker, or an exam that tests recall rather than the stated outcome can make "direct" evidence no better than the survey. Direct does not automatically mean valid. The strongest programme files triangulate — direct and indirect, each checking the other — which is precisely why the two must be kept distinct rather than conflated.

What this means for how you run evaluations

If the course evaluation's real job is explanatory — to make sense of the direct evidence — then a five-point Likert grid is badly suited to it. Averaging "I feel I learned a lot" across a cohort produces a number that is neither valid proof of learning nor a usable explanation of anything. To play its proper supporting role, an evaluation needs to surface reasons: which parts of the course built which capabilities, where the integration failed, why a cohort that scored well on knowledge fell short on application.

That reframes the design brief. You are not trying to squeeze a learning-proof out of a satisfaction survey. You are trying to collect rich, structured, outcome-linked explanation that sits alongside direct evidence in a programme-level quality file — the kind of evidence that maps onto learning outcomes and stands up under ESG-aligned quality assurance and professional accreditation.

How Koji strengthens the indirect layer

Koji for Education is built to make indirect evidence do its real job — explanation — instead of masquerading as proof. Its AI-moderated conversational interviews probe beyond a rating to capture why students experienced a course as they did, mapped to specific learning outcomes rather than a generic satisfaction score. When your direct assessment reveals a gap, that is exactly the interpretive evidence an accreditor wants beside it. Six structured question types (open-ended, scale, single- and multiple-choice, ranking, yes/no) let you align items directly to programme outcomes, and automatic thematic analysis turns open-text at scale into structured, outcome-linked themes — the difference between "students were 4.1/5 satisfied" and "students report the research-methods thread never connected to the dissertation, which explains the synthesis gap in the capstone."

Quality scoring flags thin or low-effort responses so your indirect evidence is trustworthy; programme- and institution-level reporting assembles it into the accreditation-ready, outcome-mapped picture that ESG and professional bodies expect; and closing-the-loop action tracking documents the response, which is itself evidence of a functioning quality cycle. Koji is careful about what it claims: it does not assert that a conversation proves learning — that is what your direct measures are for. It makes the indirect layer rigorous, interpretable and outcome-aligned, so it supports the direct evidence instead of pretending to be it.

Legacy SET tools were built to average a Likert score — the one thing the direct/indirect distinction says an evaluation should not be reduced to. Koji treats course evaluation as the explanatory partner to direct assessment, which is what the assessment profession has always said it is. The same conversational engine underpins the main Koji research platform for teams running wider stakeholder and outcomes research.

If your accreditation file leans on satisfaction scores to stand in for evidence of learning, you have an exposure. See how Koji for Education produces outcome-linked evidence that strengthens — rather than substitutes for — your direct measures.