Which Combinations of Course Features Produce High Ratings? Qualitative Comparative Analysis for Programme Evaluation
Qualitative Comparative Analysis (QCA) finds which combinations of course conditions are consistently sufficient for a good outcome, capturing equifinality and conjunctural causation that net-effect regression misses — but it is sensitive to calibration choices, limited case diversity and contradictions, and is descriptive of your cases, not a causal proof.
Koji Education Team
Product
In brief
Qualitative Comparative Analysis (QCA) asks a question regression cannot: which combinations of conditions reliably produce the outcome? Instead of estimating the average independent effect of each variable, QCA treats each course (or section, or programme) as a configuration of conditions — small class and experienced instructor and active learning — and uses set theory and Boolean minimisation to find the configurations that are consistently sufficient for a high rating, and any condition that is necessary across all good cases. Its strengths are exactly the things net-effect models suppress: equifinality (several different recipes can lead to the same good outcome) and conjunctural causation (a condition matters only in combination with others). Its weaknesses are equally specific: results are sensitive to how you calibrate conditions into set membership, to limited case diversity, and to contradictory configurations — and QCA describes patterns across your cases, it does not by itself prove causation.
What the research says
Ragin (1987, The Comparative Method) introduced QCA to bridge case-oriented and variable-oriented research for the medium-N problem — too many cases to hold in the head, too few for inferential statistics. The core apparatus: express each condition and the outcome as set membership, build a truth table with one row per logically possible combination of conditions, record which rows are consistently associated with the outcome, and apply Boolean minimisation to reduce them to the simplest sufficient combinations. Ragin (2008, Redesigning Social Inquiry: Fuzzy Sets and Beyond) generalised crisp (in/out) sets to fuzzy sets, where a case can be, say, 0.8 in the set of "high-workload courses", and formalised the two key fit statistics: consistency (how reliably a configuration is sufficient for the outcome) and coverage (how much of the outcome the configuration accounts for).
Two ideas from this literature matter most for evaluators. First, necessity and sufficiency are asymmetric and separate: a condition can be necessary (present in every good course) without being sufficient (not enough on its own), and QCA tests each explicitly — unlike a regression coefficient, which conflates them. Second, causation is understood as INUS-style: a condition is typically an Insufficient but Necessary part of a combination that is itself Unnecessary but Sufficient. Schneider and Wagemann (2012, Set-Theoretic Methods for the Social Sciences) is the standard rigorous treatment, formalising calibration, the treatment of limited diversity and logical remainders (combinations with no observed cases), and the analytic moves that keep a QCA defensible. Greckhamer, Furnari, Fiss and Aguilera (2018, Strategic Organization 16(4):482–495) distilled best practices — on model building, sampling, calibration, consistency thresholds and reporting — and, importantly, catalogued the ways QCA is done badly, which is the checklist a critical reader will apply to your analysis.
Why it matters for course evaluation in practice
Programme-level evaluation constantly runs into the net-effects trap. A regression says "class size has a small negative average coefficient", and a committee concludes size does not matter much — when the truth may be that small size is decisive only for discussion-based courses and irrelevant for lectures. QCA is built for that reality. Feed it a set of courses coded on a handful of theoretically chosen conditions (size, modality, instructor experience, assessment type, active-learning intensity) and their outcome (high overall rating), and it returns statements of the form: courses achieve high ratings either when they are small and discussion-based, or when they are large but have strong feedback systems — two distinct recipes, each with a consistency and coverage figure. That is far more actionable for a dean than a table of average coefficients, because it tells you which bundles of features to protect or build, and it accepts that there is more than one route to a good course.
QCA is distinct from the platform's other pattern methods, and the distinction is not cosmetic. Latent profile analysis finds empirical clusters of respondents from their response patterns; QCA works at the level of cases (courses/programmes) and tests explicit set-theoretic sufficiency against an outcome. Realist evaluation asks "what works, for whom, in what circumstances" through context-mechanism-outcome theory; QCA is a natural quantitative partner to that logic, turning "for whom, in what circumstances" into testable configurations. And unlike contribution analysis, which builds a narrative causal argument for a single programme, QCA compares many cases to find the combinations that travel.
Limitations and honest caveats
QCA invites specific, well-known objections, and pre-empting them is what makes an analysis credible. Calibration drives the result. Deciding the thresholds that make a course "small" or "high-rated" is a judgement, and modest changes can alter the solution; you must justify calibration anchors theoretically and run robustness checks. Limited diversity is pervasive. With five conditions there are 32 possible configurations but you may observe only a dozen; the empty rows (logical remainders) force assumptions, and the "conservative", "intermediate" and "parsimonious" solutions handle them differently — report which you used and why. Contradictions happen — the same configuration sometimes produces the outcome and sometimes not, signalling a missing condition or miscalibration, not a result to paper over. QCA is not causal proof. Consistency and coverage describe set relations among your cases; generalisation and causal claims need theory, case knowledge and ideally a design, echoing the specification-curve lesson that analytic choices move conclusions. And it is a medium-N method — with thousands of student-level rows it is the wrong tool; QCA lives at the level of dozens of courses or programmes, not individual respondents.
How Koji incorporates this
Koji is positioned to feed a QCA rather than to replace the analyst's judgement. Because its structured questions produce clean, comparable dimension scores per course, and its automatic thematic analysis codes open-text into consistent conditions (for example, "students report strong feedback" as present/absent), Koji can assemble the case-by-condition matrix QCA needs across a programme's courses or across cohorts — the assembly step that is usually the most tedious part of a configurational study. Koji is designed to support transparent calibration, storing the thresholds that turn a continuous rating or theme frequency into set membership so the choice is auditable and can be varied for robustness, and to report consistency and coverage for each configuration rather than a single verdict. Crucially, Koji frames QCA output as "designed to surface candidate recipes for review", never as proof that a configuration causes good ratings — the platform pairs the configurations with the students' own words so a committee can sanity-check each recipe against lived accounts, which is exactly the case-knowledge QCA demands. The same configurational-analysis capability is available in Koji's core research platform at koji.so, where "which combination of product conditions yields loyal customers" is the identical question in a different domain.
Frequently asked questions
How is QCA different from regression?
Regression estimates the average, independent effect of each variable holding others constant. QCA finds combinations of conditions that are jointly sufficient for an outcome, allowing several different recipes (equifinality) and conditions that only matter in combination (conjunctural causation). They answer different questions and can be used together.
What are consistency and coverage in QCA?
Consistency measures how reliably a configuration is sufficient for the outcome — how often cases with that configuration actually show the outcome. Coverage measures how much of the outcome the configuration accounts for. High consistency with modest coverage means a reliable but partial recipe.
What is the difference between crisp-set and fuzzy-set QCA?
Crisp-set QCA codes each condition as fully in or out of a set (0 or 1). Fuzzy-set QCA (Ragin 2008) allows partial membership between 0 and 1, so a course can be 0.7 in the set of high-workload courses, preserving more information from continuous measures.
How many cases do I need for QCA?
QCA is a medium-N method, typically applied to somewhere between roughly a dozen and a few hundred cases such as courses or programmes. It is not designed for thousands of individual student responses, and with too few cases relative to conditions limited diversity becomes severe.
Does a high-consistency configuration prove it causes good ratings?
No. Consistency describes a set relation among your observed cases. Causal claims need theory, case knowledge and ideally a design. Report QCA as evidence of a pattern worth investigating, not as proof of cause.
What is limited diversity and why does it matter?
Limited diversity means many logically possible condition combinations have no observed cases. QCA must then make assumptions about those empty rows, and the conservative, intermediate and parsimonious solutions treat them differently, which can change the result — so you must report which solution you used.
Related Resources
- Realist Evaluation: What Works, for Whom, in What Circumstances — the theory QCA operationalises
- Latent Profile Analysis for Course-Evaluation Segments
- Contribution Analysis: Credible Causal Claims from Course Evaluation
- Multiple Correspondence Analysis for Categorical Course-Evaluation Data
- Stufflebeam's CIPP Model for Programme-Level Evaluation
References
- Ragin, C. C. (1987). The Comparative Method: Moving Beyond Qualitative and Quantitative Strategies. University of California Press. ISBN 978-0520058347.
- Ragin, C. C. (2008). Redesigning Social Inquiry: Fuzzy Sets and Beyond. University of Chicago Press. https://doi.org/10.7208/9780226702797
- Schneider, C. Q., & Wagemann, C. (2012). Set-Theoretic Methods for the Social Sciences: A Guide to Qualitative Comparative Analysis. Cambridge University Press. https://doi.org/10.1017/CBO9781139004244
- Greckhamer, T., Furnari, S., Fiss, P. C., & Aguilera, R. V. (2018). Studying configurations with qualitative comparative analysis: Best practices in strategy and organization research. Strategic Organization, 16(4), 482–495. https://doi.org/10.1177/1476127018786487
Related articles
Beyond the Average Student: Latent Profile Analysis for Course-Evaluation Segments
A course mean of 3.6 can hide two entirely different student experiences averaged into one number. Latent profile analysis (LPA) recovers those hidden subgroups from the evaluation data itself, so you can see the delighted minority and the alienated cohort that the average erased.
One Instructor, Two Hundred Defensible Scores: Specification-Curve Analysis for Course Evaluation
Every course-evaluation "score" is the product of dozens of defensible analytic choices — which items count, how to weight, mean or median, whether to adjust for class size. Specification-curve and multiverse analysis compute the result across all of them, so a personnel decision rests on the pattern, not on one analyst's arbitrary path.
Mapping the Patterns You Cannot Average: Multiple Correspondence Analysis for Categorical Course-Evaluation Data
Much course-evaluation data is genuinely categorical — programme, mode, agree/disagree, chosen theme. Multiple correspondence analysis (MCA) maps how those categories cluster on a two-dimensional plane, revealing response patterns that averaging destroys.
Realist Evaluation for Course Feedback: 'What Works, for Whom, in What Circumstances'
Pawson and Tilley's realist evaluation replaces 'did it work?' with 'what works, for whom, in what circumstances?' using context-mechanism-outcome configurations. Here is the theory, the evidence, the caveats, and how to apply it to course evaluation.