Is Your Course Accessible to Every Learner? Universal Design for Learning as an Inclusive-Evaluation Lens
Universal Design for Learning asks whether a course is usable by the full range of learners by default, not by retrofitted accommodation. Here is how to evaluate inclusivity with UDL — and why the evidence demands you do it honestly.
Koji Education Team
Product
The short answer
An aggregate course-evaluation score can be perfectly healthy while a subgroup of students quietly struggles because the course was designed for an imagined "average" learner who does not exist. Universal Design for Learning (UDL) — a framework built on providing multiple means of engagement, representation, and action/expression — gives evaluation an inclusivity question the mean hides: was this course usable, from the start, by the full range of learners in it? Using UDL as an evaluation lens means disaggregating experience by learner variability and asking whether flexibility was built in by default or bolted on as an afterthought. The evidence base, however, comes with an unusually important caveat, so this lens must be used honestly.
What the research says
UDL grew out of the work of CAST and is articulated most fully in Meyer, Rose and Gordon's Universal Design for Learning: Theory and Practice (2014). Its founding premise, borrowed from architecture's universal design movement, is that variability among learners is the norm, not the exception, and that curricula should therefore be designed from the outset to be flexible along three axes:
- Multiple means of engagement — the "why" of learning: varied ways to recruit interest, sustain effort, and self-regulate.
- Multiple means of representation — the "what" of learning: information presented in more than one modality and format.
- Multiple means of action and expression — the "how" of learning: more than one way for students to demonstrate what they know.
The claim is that designing for the "edges" — students with disabilities, non-native speakers, students balancing work and caring responsibilities — produces a course that is better for everyone, just as kerb cuts designed for wheelchair users benefit parents with prams and travellers with luggage.
Here is where an honest evaluation lens has to be careful about the evidence. Capp's 2017 meta-analysis in the International Journal of Inclusive Education ("The effectiveness of universal design for learning: a meta-analysis of literature between 2013 and 2016," N = 18 studies) concluded that UDL appears to be "an effective teaching methodology for improving the learning process" — but explicitly noted that its impact on educational outcomes has not been demonstrated. In other words, the strongest available synthesis supports UDL's effect on the process and experience of learning far more than on measured attainment. Later syntheses echo the pattern: promising engagement and accessibility effects, thin causal evidence on grades.
And there is a pointed critique worth taking seriously. Boysen's 2021 paper in Scholarship of Teaching and Learning in Psychology ("Lessons (not) learned: the troubling similarities between learning styles and universal design for learning") argues that UDL shares uncomfortable features with the discredited learning-styles myth: both overemphasise diversity in how students learn over universal principles of learning, both lean on oversimplified neuroscience, and — most importantly — both "lack evidence showing that their implementation increases student learning." Whether or not one accepts Boysen's conclusion, the critique sets the correct evidential bar: evaluate UDL for what it demonstrably delivers (access, engagement, reduced need for individual accommodation) and do not overclaim attainment gains it has not earned.
Why it matters for course evaluation in practice
Most evaluation reporting leads with the mean and maybe a distribution. Neither reveals whether a course systematically disadvantaged a subgroup. A course delivered in a single modality (dense lectures, one high-stakes written exam) can post a solid average while non-native speakers, students with disabilities, and neurodivergent students report a markedly worse experience — a gap that vanishes in the aggregate. UDL gives evaluators a principled reason to disaggregate, and a vocabulary for what to ask about: Were materials available in more than one format? Was there more than one legitimate way to demonstrate learning? Were there flexible, low-stakes entry points, or a single narrow on-ramp?
This turns "inclusivity" from a vague aspiration into specific, checkable design features and specific, comparable experience data. Critically, it distinguishes universal design (flexibility built in for everyone) from accommodation (individual retrofits granted on request). A course that scores well only because a disability office bolted on accommodations after the fact is, by UDL's standard, still a poorly designed course — and an evaluation should be able to see that difference. Under European quality frameworks (ESG) and equality legislation, evidence that a course works for the full learner population — not just its median — is increasingly what accreditation panels and equality reviews expect.
Limitations and honest caveats
The attainment evidence is genuinely thin — do not overclaim. As Capp's meta-analysis states outright, UDL's effect on measured outcomes is not established. An evaluation built on UDL can credibly report on access, engagement, and experience equity; it cannot credibly claim that adopting UDL raised grades. Presenting it otherwise repeats exactly the error Boysen warns about.
The learning-styles critique should not be dismissed. UDL is not the learning-styles myth — it argues for offering options to all learners, not for matching a teaching mode to a fixed learner "type," which is the specific, falsified claim of learning styles. But the critique correctly disciplines how you evaluate: measure the provision of options and their uptake, not a fictitious match between modality and learner type.
Subgroup analysis needs care with small numbers and anonymity. Disaggregating by disability, language, or other protected characteristics in a small class can re-identify individuals and can itself feel othering. Sample sizes for subgroups are often tiny, making quantitative comparison unstable. Qualitative, consent-based inquiry is frequently more appropriate and more ethical than slicing a small quantitative dataset.
UDL describes design, not worth. A course can be beautifully flexible and still teach little of value. UDL evaluates the accessibility of the learning environment, not the quality or rigour of what is being learned; it complements outcome- and alignment-focused evaluation rather than replacing it.
How Koji incorporates this
Koji is built to collect equity-of-experience evidence without forcing evaluators into unsafe subgroup slicing, and to keep claims honest about what the data supports.
- Experience questions mapped to the three UDL principles. A Koji evaluation can carry
scaleandyes_noitems on whether materials were available in multiple formats (representation), whether students had a genuine choice in how to demonstrate learning (action/expression), and whether there were varied ways to engage (engagement) — turning "inclusive?" into concrete, answerable design questions. - AI-moderated conversational inclusion probes. For sensitive equity topics, Koji's conversational moderator can invite students to describe, in their own words and at their own comfort level, any point where the course's format made participation harder — surfacing barriers a fixed subgroup checkbox would miss, and doing so without demanding a student label themselves.
- Anonymity-preserving analysis. Because subgroup slicing in small classes risks re-identification, Koji's thematic analysis clusters barrier experiences ("only one way to submit", "no captions on the recordings") across the cohort rather than requiring identity-tagged breakdowns, and the platform's small-sample and re-identification safeguards apply.
- Honest, non-overclaiming reporting. Koji reports UDL evidence as access-and-experience evidence — where the research is strong — and is designed not to fabricate an attainment claim the literature does not support.
- Formative timing. Because a format barrier is best fixed while the course runs, Koji's mid-cycle collection can surface a UDL problem in week four rather than in a post-mortem.
- Closing-the-loop tracking records which barrier a team removed and lets the next cohort confirm the fix.
Koji is designed to mitigate exclusion by revealing where a course's default design failed part of its cohort — framed as "designed to surface barriers," never as a guarantee of equal outcomes. The same AI-moderated interview engine powers Koji's core research platform at koji.so for inclusive product and customer research, where designing for the full range of users follows the identical logic.
Related Resources
- Should Course Evaluations Ask Whether Teaching Matched a Student's Learning Style?
- Evaluating Online and Blended Teaching: The Community of Inquiry Framework
- Students as Partners, Not Just Respondents: Co-Designing Course Evaluation
- Can a Student Be Re-Identified From Their Course Feedback? Small-Class Anonymity, k-Anonymity and GDPR
- Evaluating Graduate Teaching Assistants: Why GTA Feedback Needs Its Own Rules
- Abusive Open-Text Comments in Course Evaluations: The Evidence and the Duty of Care
References
- Capp, M. J. (2017). The effectiveness of universal design for learning: a meta-analysis of literature between 2013 and 2016. International Journal of Inclusive Education, 21(8), 791–807. https://doi.org/10.1080/13603116.2017.1325074
- Meyer, A., Rose, D. H., & Gordon, D. (2014). Universal Design for Learning: Theory and Practice. Wakefield, MA: CAST Professional Publishing.
- Boysen, G. A. (2021). Lessons (not) learned: The troubling similarities between learning styles and universal design for learning. Scholarship of Teaching and Learning in Psychology. https://doi.org/10.1037/stl0000280
Related articles
Abusive Open-Text Comments in Course Evaluations: The Evidence and the Duty of Care
Lakeman et al. (2022) found that 91% of surveyed Australian academics received non-constructive anonymous comments — insults, threats, remarks on appearance. A research-grounded look at abusive student feedback, its effect on staff wellbeing, and how to moderate open text responsibly.
Can a Student Be Re-Identified From Their Course Feedback? Small-Class Anonymity, k-Anonymity and GDPR
"Anonymous" course evaluations in small classes often are not. What re-identification research and GDPR actually require — and how to report small-cohort feedback without breaching either confidentiality or law.
Students as Partners, Not Just Respondents: Co-Designing Course Evaluation
Healey, Flint and Harrington's students-as-partners framework reframes evaluation from something done *to* students into something done *with* them. What partnership adds to course feedback, its limits, and how to build it in.
Evaluating Graduate Teaching Assistants: Why GTA Feedback Needs Its Own Rules
Students judge graduate teaching assistants against a different mental template than professors, so benchmarking GTA scores on faculty norms is unfair and misleading. Here is what the research shows and how to evaluate GTA-led teaching properly.