New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods12 min read

Validity Is About the Use, Not the Instrument: Applying Kane's Argument-Based Framework to Course Evaluation

Asking whether course evaluations are valid is the wrong question. Kane's argument-based framework asks whether a specific interpretation and use of the scores is justified. We rebuild the SET debate as an interpretation-use argument, expose where each inference breaks, and show how Koji strengthens the weak links.

Koji Education Team

Product

In brief: "Are course evaluations valid?" is a malformed question. In modern measurement theory, validity is not a property of an instrument but of a specific interpretation and use of its scores. Michael Kane's argument-based framework (Kane, 2013) reframes validation as building and then critically evaluating an interpretation/use argument (IUA): an explicit chain of inferences from a student's clicks to the decision a committee makes. The same evaluation data can be highly valid for one use (formative feedback to an instructor) and indefensible for another (ranking instructors for promotion). Applying this lens to the Student Evaluation of Teaching (SET) literature shows the chain usually breaks at the extrapolation inference — the leap from "students rated this highly" to "this instructor produced more learning" — which the meta-analytic evidence does not support. The disciplined response is not to abolish evaluation but to match each use to the inferences the evidence can actually bear.

What the research says

Michael Kane's Validating the interpretations and uses of test scores (Kane, 2013, Journal of Educational Measurement, 50(1), 1–73) is the most influential modern statement of validity theory. Building on Messick and the Standards for Educational and Psychological Testing (AERA, APA & NCME, 2014), Kane makes a deceptively simple move: validity is not something a test has; it is something a particular interpretation and use of scores either earns or fails to earn. Validation is therefore a two-step enterprise. First, lay out the interpretation/use argument (IUA): the explicit chain of inferences and assumptions that takes you from observed responses to the claim and decision you want to make. Second, build the validity argument: an evaluation of how plausible each link in that chain is, given the available evidence.

Kane decomposes the chain into a sequence of inferences. For a course-evaluation programme they read roughly as:

  1. Scoring — from a student's individual responses to a score (does summing/averaging these items yield a meaningful number?).
  2. Generalisation — from this sample of students and items to the universe of students and items (is the score reliable across raters, items, and occasions?).
  3. Extrapolation — from the test score to the real-world construct of interest (does a high rating reflect actual teaching effectiveness / student learning?).
  4. Implication / decision — from the construct claim to the action taken (is it justified to deny tenure, rank a department, or trigger a teaching review on this basis?).

A score-use is valid only if every link holds. The power of the framework is that it tells you where to look and forces stakeholders to state, in advance, what claim they are actually making.

When you map the SET evidence onto this chain, the breaks are not evenly distributed. The scoring and generalisation inferences are often defensible: well-designed SET instruments can be internally consistent and, with enough raters across enough occasions, reasonably reliable (the generalizability-theory literature quantifies exactly how many). It is the extrapolation inference that collapses. Uttl, White and Gonzalez (2017), Meta-analysis of faculty's teaching effectiveness (Studies in Educational Evaluation, 54, 22–42), re-analysed the multisection validity literature and found that, once sample size is properly accounted for, SET ratings are not related to student learning. Spooren, Brockx and Mortelmans (2013), On the validity of student evaluation of teaching: The state of the art (Review of Educational Research, 83(4), 598–642), conclude that SET is a complex, multidimensional construct whose validity for high-stakes inferences is far from established and is contaminated by numerous biases. In Kane's terms: the data can survive scoring and generalisation but cannot support the extrapolation to "teaching effectiveness," and therefore cannot license the high-stakes implication of ranking or promotion.

Why it matters for course evaluation in practice

The argument-based framework is the single most useful reframing available to a quality-assurance office, because it dissolves a sterile debate ("evaluations are biased rubbish" versus "evaluations are essential") into a set of answerable questions.

  • It separates uses that share an instrument. The same evaluation can validly support a formative use (an instructor reading their own comments to improve next term) while being indefensible for a summative one (a dean ranking instructors). Formative use rests mainly on scoring and generalisation; it does not require the fragile extrapolation to "effectiveness." This is why "abolish SET" and "SET is fine" are both wrong — they fail to name the use.

  • It tells you what evidence to collect. If your IUA depends on the generalisation inference, you need reliability evidence (how many responses, across how many items and occasions). If it depends on extrapolation, you need external evidence linking scores to learning — exactly the evidence the meta-analyses say is missing. The framework converts vague anxiety into a checklist.

  • It is audit- and accreditation-ready. ESG/ENQA and national agencies increasingly ask not "do you collect feedback?" but "is your use of it justified?" An explicit IUA — stating the claim, the inferences, and the supporting evidence for each — is precisely the documentation a critical reviewer wants, and it pre-empts the legal and equity challenges that follow from using biased scores for personnel decisions.

Limitations and honest caveats

The framework is powerful but not a panacea, and a PhD reader will press on several points.

  • It is a structure, not evidence. Kane tells you how to organise a validity argument; it does not supply the missing data. A beautifully specified IUA with no evidence for its weakest inference is still invalid. The framework can give a false sense of rigour if the validity argument is thin.

  • Defining the construct is contested. The extrapolation inference presupposes you can define "teaching effectiveness." The multidimensionality debate (Marsh; d'Apollonia and Abrami) shows there is no single agreed construct, and different definitions move the goalposts for what counts as supporting evidence.

  • Consequential and value implications are hard. Kane (following Messick) folds the consequences of use into validity — including washback effects such as grade inflation and reduced rigour when SET becomes high-stakes (the incentive literature, e.g., Stroebe). These consequences are real but difficult to quantify, and reasonable people disagree about how much weight they should carry.

  • Practicality versus completeness. A fully specified IUA for every score-use is labour-intensive. Institutions trade completeness for feasibility, and that trade-off is itself a judgement that the framework does not make for you.

  • Generalisation can fail too. Although scoring and generalisation are often defensible, small classes, low response rates, and non-response bias can break the generalisation inference long before extrapolation is even reached — so "the data are fine for formative use" is not automatic either.

The honest conclusion: Kane's framework is the right way to reason about course-evaluation validity, but it raises the bar for evidence rather than lowering it.

How Koji incorporates this

Koji for Education is designed around the same premise Kane formalises: the legitimacy of an evaluation depends on the claim you make from it. Several mechanisms are built to strengthen the specific inferences in the chain.

  • Use-aware evaluation design. Koji distinguishes formative cycles (mid-semester, instructor-facing) from summative, programme-level reporting, so that data collected for improvement is not silently repurposed for high-stakes ranking — the exact misuse Kane's framework exposes. Question design and reporting are matched to the intended use.

  • Strengthening the generalisation inference. Because generalisation depends on reliability across raters and occasions, Koji supports mid-cycle and repeated collection, surfaces response-rate and non-response signals, and is designed to report results with uncertainty rather than treating one low-N cohort as definitive.

  • Targeting the weak extrapolation link with richer evidence. A single Likert number cannot support the leap to "effective teaching." Koji's AI-moderated conversational interviews probe behavioural specifics (what the instructor did, what the student could or could not do as a result), and its automatic thematic analysis structures that qualitative evidence — designed to provide the kind of construct-relevant detail that a bare rating lacks, while never claiming to prove learning.

  • Triangulation for the implication inference. Because high-stakes decisions require more than one fragile inference, Koji is built to triangulate student perceptions with other evidence sources over time, supporting the "do not decide on a single source" conclusion that the validity literature demands.

  • Documentation for the validity argument itself. Koji's structured outputs (scores, themes, action tracking, cohort comparisons) are designed to produce exactly the audit trail an accreditation reviewer needs to see which inferences a claim rests on and what supports them — the practical embodiment of an interpretation/use argument.

These are mechanisms designed to strengthen specific inferences in the chain; they cannot manufacture an extrapolation to learning that the underlying data do not support, and Koji's reporting is explicit about that boundary. Koji's core research platform at koji.so applies the same interpretation-aware approach to product and customer research, where the discipline of matching a claim to the evidence behind it is equally decisive.

Related Resources

References

  1. Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
  2. Spooren, P., Brockx, B., & Mortelmans, D. (2013). On the validity of student evaluation of teaching: The state of the art. Review of Educational Research, 83(4), 598–642. https://doi.org/10.3102/0034654313496870
  3. Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
  4. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. Washington, DC: AERA.

Related articles

analysis-reporting

Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved

Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.

analysis-reporting

Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough

Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

research-methods

Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited

A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.