Numbers First, Then the Why: Explanatory Sequential Mixed Methods for Course Evaluation
A Likert average tells you a course scored 3.4; it never tells you why. The explanatory sequential mixed methods design (Creswell & Plano Clark) fixes this by using a quantitative survey to decide exactly which students to follow up with qualitatively — turning an unexplained number into an evidenced account.
Koji Education Team
Product
In brief
An explanatory sequential mixed methods design collects quantitative course-evaluation data first, then uses those results to purposively select whom to interview and what to ask, so a second qualitative phase can explain the numbers. It is the most defensible way to answer the question a Likert mean cannot: why did this course score the way it did? For a quality-assurance office, the design converts an ambiguous 3.4/5 into an evidenced account of the mechanism behind it — and, crucially, it disciplines the follow-up so you interview the right students about the right result rather than gathering anecdotes at random.
What the research says
The explanatory sequential design (originally the "sequential explanatory" design) was formalised by John Creswell, Vicki Plano Clark and colleagues and set out step-by-step by Ivankova, Creswell & Stick (2006) in Field Methods. Their model has two connected phases. Phase one is quantitative: a survey is administered and analysed. Phase two is qualitative: the researcher follows up on specific quantitative results — outliers, unexpected findings, or highly significant effects — with interviews, focus groups or open-ended probes. The two phases are joined at an explicit connection point: the quantitative results are used to build the qualitative sampling frame and the interview protocol. Ivankova and colleagues stress three design decisions that make or break the study: priority (which strand carries more weight — usually the quantitative), sequence (quantitative strictly before qualitative), and the point of interface where results are connected and later merged.
The logic is that the strands do different jobs. Numbers establish what and how much — how many students were dissatisfied, on which dimension, by how much relative to a benchmark. Words establish why and how — the reasoning, context and lived experience behind the number. Creswell & Plano Clark (2018), in the standard reference text Designing and Conducting Mixed Methods Research, describe this design as especially useful when a quantitative result is surprising or needs elaboration, and when the quantitative phase can guide purposeful sampling for the qualitative phase. Applied higher-education studies use it exactly this way: a survey identifies a subgroup or an anomalous result, and follow-up interviews with members of that subgroup explain it. The design has become one of the most widely taught mixed-methods templates precisely because the hand-off from numbers to interviews is concrete and auditable.
Two features distinguish a genuine explanatory sequential study from "we ran a survey and also did some interviews." First, the qualitative sample is derived from the quantitative results (for example, deliberately interviewing the students who scored a dimension lowest, not a convenience sample). Second, the interview questions are written to interrogate the quantitative finding, not a fresh topic. When both hold, the qualitative phase is genuinely explanatory rather than merely additional.
Why it matters for course evaluation in practice
Most institutional course evaluation stops at phase one. A programme director receives a dashboard: "Organisation 4.1, Assessment 3.2, Overall 3.7." The low assessment score is a flag, but the number is mute about mechanism. Was assessment rated low because the workload was crushing, because feedback arrived too late to be useful, because the rubric was opaque, or because one high-stakes exam dominated the grade? Each diagnosis implies a different fix, and the mean cannot distinguish them. Acting on the number alone is guessing.
An explanatory sequential approach makes the second step routine and rigorous. The quantitative evaluation is not the end of the process but the sampling instrument for a focused qualitative follow-up: the office reads the distribution, identifies the specific result that needs explaining, selects the students best placed to explain it, and probes only that. This is the difference between "assessment scored 3.2" and "assessment scored 3.2 because summative feedback consistently arrived after the next assignment was due, so students could not act on it" — a finding a QA panel or accreditor can actually act upon. It also protects scarce qualitative effort: instead of reading 400 free-text comments hoping a theme emerges, you interview the 12 students whose pattern of answers makes them most informative about the flagged result.
Limitations and honest caveats
The design is demanding, and a PhD reader will raise several objections. Time and sequence. Because phase two cannot begin until phase one is analysed, the design is slow; by the time interviews are scheduled, the cohort may have dispersed and memories decayed. Selecting the follow-up sample is a judgement call that can bias the account — interview only the angriest students and you will over-explain the negative tail. Ivankova and colleagues note that deciding which results to pursue, and whom to recruit, requires explicit, defensible criteria, not convenience. Reduced anonymity. A student who agrees to be interviewed about their own low ratings is no longer anonymous, which can suppress candour or deter participation, especially in small cohorts where re-identification is easy. Integration is where studies fail. Many nominally mixed-methods evaluations never truly merge the strands; they report a survey and, separately, some quotes. Without a genuine connection point, you have two mono-method studies stapled together, not an explanatory design. Finally, generalisability of phase two is limited by design — the interviews explain the quantitative pattern for the sampled students; they are not a representative census of opinion, and should not be reported as one. Treat the qualitative phase as mechanism-finding, not prevalence-estimating.
How Koji incorporates this
Koji is built to run the explanatory sequential logic inside a single instrument rather than across two disconnected surveys, which removes the main practical cost — the time gap and the re-recruitment problem — while keeping the methodological discipline.
- Quantitative first strand. Koji collects structured responses using scale, single_choice, multiple_choice, yes_no and ranking questions, producing the numeric distribution that identifies which dimension or subgroup needs explaining.
- Adaptive connection point. Rather than waiting weeks to select an interview sample, Koji's AI-moderated conversational interview branches in the moment: when a respondent rates a dimension low (or unusually high), the interview follows up with an open_ended probe targeted at that specific result — the connection point is executed live, per respondent, on the result that actually needs elaboration.
- Qualitative second strand at scale. The conversational probes do the work of the interview phase for every relevant respondent, not just a hand-picked twelve, so the "why" is captured while the experience is fresh and while the student is still anonymous.
- Merging the strands. Koji's automatic thematic analysis codes the open-text explanations and links them back to the numeric ratings that triggered them, so a report can show not just that assessment scored 3.2 but the ranked themes explaining it — the merged inference the design is meant to produce.
- Bias-aware reporting. Because Koji records which quantitative result each qualitative theme is explaining, it can flag when an account is drawn disproportionately from the dissatisfied tail, guarding against the "interview only the angriest" distortion above.
This is designed to approximate and streamline the explanatory sequential design, not to replace a full multi-phase research study; deep, longitudinal explanatory work still benefits from dedicated interviews. Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where the "survey then explain the outliers" pattern is equally standard.
Frequently asked questions
Is this the same as just adding open-text boxes to a survey? No. Open-text boxes gather unstructured comments from everyone regardless of their ratings. An explanatory sequential design uses the ratings to decide whom to probe and about what, so the qualitative data is targeted at the specific result that needs explaining.
Which phase should carry more weight? In most course-evaluation uses the quantitative strand is primary (it establishes prevalence and flags the issue) and the qualitative strand is explanatory. Ivankova, Creswell & Stick (2006) treat priority as an explicit design decision you should state, not leave implicit.
Does the qualitative phase let me claim how common a view is? No. The follow-up explains the mechanism behind a quantitative result for the students you sampled; it does not estimate prevalence. Use the quantitative strand for "how many" and the qualitative strand for "why."
How does this differ from a concurrent (convergent) mixed-methods design? In a convergent design both strands are collected at once and compared. In an explanatory sequential design the quantitative results drive the qualitative sampling and questions, so the sequence is essential — the numbers must come first.
Can a conversational AI interview really substitute for a human interview phase? It substitutes for the structured probing function — following up on a specific result with targeted open questions at scale. It does not replace an in-depth researcher-led interview for complex, exploratory topics, and should be framed as a streamlined implementation of the design.
Related resources
- Why a Single Student Survey Can't Stand Alone: Common-Method Bias in Course Evaluation
- What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
- Do Interactive Follow-Up Probes Improve Open-Text Course Feedback Quality?
- Engagement Surveys vs Course Evaluations: What NSSE Measures
- The Most Significant Change Technique: Story-Based Course Evaluation
- The Framework Method for Open-Text Course Feedback
References
- Ivankova, N. V., Creswell, J. W., & Stick, S. L. (2006). Using Mixed-Methods Sequential Explanatory Design: From Theory to Practice. Field Methods, 18(1), 3-20. https://doi.org/10.1177/1525822X05282260
- Creswell, J. W., & Plano Clark, V. L. (2018). Designing and Conducting Mixed Methods Research (3rd ed.). SAGE Publications.
- Creswell, J. W., Plano Clark, V. L., Gutmann, M. L., & Hanson, W. E. (2003). Advanced mixed methods research designs. In A. Tashakkori & C. Teddlie (Eds.), Handbook of Mixed Methods in Social and Behavioral Research (pp. 209-240). SAGE Publications.
Related articles
Why a Single Student Survey Can't Stand Alone: Common-Method Bias in Course Evaluation
Common-method bias (Podsakoff et al., 2003) explains why correlations inside a single end-of-term student survey are inflated by the shared method itself - and why triangulating teaching evidence matters. A research-grounded guide for quality assurance.
The Framework Method for Open-Text Course Feedback: A Structured Alternative to Thematic Coding
When a quality committee — not a lone qualitative researcher — has to make sense of thousands of student comments, the Framework Method offers a transparent, auditable matrix-based approach. What Gale et al. (2013) actually proposed, and how to use it.
What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.
Engagement Surveys vs Course Evaluations: What NSSE Measures and Why Porter Says Be Careful
How student-engagement surveys such as NSSE differ from course evaluations, what Kuh (2009) argues they capture, why Porter (2011) questions their validity, and how to triangulate engagement and evaluation evidence without over-trusting either.