New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Learn What Students Actually Weigh: The Factorial Survey (Vignette) Experiment for Course Evaluation

A factorial survey embeds a randomised experiment inside a survey: students rate realistic course vignettes whose features are varied independently, so the weight of each feature on their judgement can be recovered cleanly. This article explains the design, its distinction from conjoint and anchoring vignettes, and how to avoid its pitfalls.

Koji Education Team

Product

In brief

If you want to know what students actually weigh when they judge a course — rather than what they say they value when asked directly — present them with short, realistic course scenarios ("vignettes") whose features you have experimentally varied, and have them rate each one. Because the features are randomised independently, a regression of ratings on features recovers the causal weight each factor carries in students' judgements, cleanly separated from all the others. This is the factorial survey experiment (Rossi & Anderson, 1982; Auspurg & Hinz, 2015). It measures judgement principles that direct questions distort, and it is distinct from choice-based conjoint and from cross-cultural anchoring vignettes.

What the research says

The factorial survey — also called a vignette experiment — was formalised by Peter Rossi and colleagues and introduced programmatically by Rossi and Anderson (1982) in Measuring Social Judgments. The core idea is to embed a randomised experiment inside a survey. Each respondent reads one or more vignettes: compact descriptions of a hypothetical case built from several dimensions (attributes), each of which takes one of several levels. The universe of all dimension-by-level combinations is the "vignette population"; each respondent judges a sample from it. Because the levels are assigned independently of one another (an orthogonal or D-efficient design), the dimensions are uncorrelated by construction — which means the effect of each dimension on the rating can be estimated without the multicollinearity that plagues observational surveys.

Auspurg and Hinz (2015), in the standard SAGE monograph, lay out the design decisions: how to choose dimensions and levels, how to build an efficient fractional design when the full factorial is too large to present, how many vignettes to give each respondent before fatigue sets in, and how to analyse the resulting data with multilevel models (vignettes nested within respondents). Wallander (2009), reviewing 25 years and 106 sociology studies, documents both the method's spread and its recurring design weaknesses — implausible attribute combinations, over-long vignette decks, and inconsistent analysis choices. Atzmüller and Steiner (2010) give a compact primer framing the factorial survey explicitly as an experimental design that combines the internal validity of randomisation with the external reach of a survey sample.

The method's signature strength is measuring trade-offs and normative judgements that respondents cannot or will not report directly. Asked "how important is workload to your rating?", students give a socially-shaped, introspective answer. Asked to rate a series of realistic courses that happen to differ in workload (among other things), they reveal the weight through their ratings — and that revealed weight is far less contaminated by self-presentation, because the respondent is not aware of which dimension is under study.

Why it matters for course evaluation in practice

Course-evaluation design constantly faces the question: which aspects of a course actually drive students' overall judgement? Answering it from live evaluation data is treacherous, because the drivers are entangled — hard courses tend to be the quantitative ones, small classes tend to be the advanced ones — so an observational regression cannot cleanly separate "difficulty" from "discipline" from "level." A factorial survey breaks the entanglement by design: you can present a small, hard, quantitative course and a small, easy, quantitative course to different students at random, isolating the pure effect of difficulty.

This makes the method valuable for several QA tasks. It can weight an evaluation instrument by revealing which teaching attributes genuinely move overall satisfaction, informing which items deserve prominence. It can probe fairness questions experimentally — by varying an instructor's name or described attributes across otherwise-identical vignettes, an institution can test whether students apply different standards, complementing the natural-experiment evidence on bias. And it can pre-test policy — for instance, gauging how students would judge a course that adopts a new assessment model before rolling it out. Because it elicits judgement principles rather than reactions to one real course, it generalises in a way a single cohort's ratings cannot.

It is worth being precise about how the factorial survey differs from neighbours already in the toolkit. A discrete choice experiment asks respondents to choose between whole alternatives, recovering preferences from choices; a factorial survey asks them to rate full-profile vignettes, recovering the weight of each attribute from ratings. Anchoring vignettes serve a different purpose entirely — calibrating self-reports across groups who use scales differently — rather than measuring judgement weights. And unlike MaxDiff / best-worst scaling, the factorial survey preserves the realistic, multi-attribute context of a whole course description.

Limitations and honest caveats

The factorial survey trades some realism for control, and the honest analyst acknowledges the costs.

  1. Hypothetical is not real. Respondents rate described courses, not lived ones. Stated judgements of vignettes may diverge from behaviour in an actual course; the method measures judgement principles, and their correspondence to real ratings is an empirical question, not a guarantee.
  2. Implausible combinations threaten validity. A fully orthogonal design can generate nonsensical vignettes (a "large, highly personalised, one-to-one seminar"). Wallander (2009) flags this as a common flaw; constrained or D-efficient designs that suppress impossible cells help but complicate estimation.
  3. Respondent burden and fatigue. Vignette decks that are too long produce satisficing and order effects; each respondent can only judge a handful of cases well, so the design must sample the vignette universe efficiently.
  4. Number-of-dimensions limits. Students can integrate only a few attributes at once; loading a vignette with ten dimensions invites simplification heuristics rather than genuine trade-off, biasing the recovered weights.
  5. Social desirability is reduced, not eliminated. Sensitive dimensions (an instructor's gender or ethnicity) can still trigger guarded responses if respondents infer the study's purpose; careful masking within a broader set of dimensions is essential.
  6. Analysis is non-trivial. Vignettes are nested within respondents, so the correct analysis is multilevel; treating each vignette rating as independent understates standard errors.

How Koji incorporates this

The factorial survey is a design method, and its output is only as good as the instrument that delivers the vignettes and the analysis that follows. Koji's platform maps naturally onto both.

  • Delivering the vignettes and capturing structured ratings. Koji's structured question types (scale, single_choice, ranking, open_ended) can present vignette descriptions and record the ratings a factorial design needs, while its survey logic supports assigning different respondents different vignette samples - the randomised delivery that makes the design work.
  • Adding the "why" a pure factorial survey misses. A classic factorial survey recovers how much each attribute weighs but not why. Koji's AI-moderated conversational follow-up can probe a student's reasoning immediately after a vignette rating, and its automatic thematic analysis codes those open responses - turning a weight estimate into an explanation, and helping catch when respondents found a vignette implausible.
  • Efficient, fatigue-aware collection. Because Koji's conversational format is designed to feel lighter than a long static grid, it helps keep vignette decks within the burden limits Wallander (2009) warns about, protecting against the satisficing that corrupts recovered weights.
  • A single engine across research modes. The same conversational research engine at koji.so runs factorial-survey-style trade-off studies for product and customer research, where separating what users say they value from what they actually weigh is the identical problem.

Koji is designed to operationalise factorial-survey thinking - randomised vignette delivery plus conversational depth - rather than to replace careful experimental design; the validity of the recovered judgement weights still rests on the dimensions, levels, and design the researcher specifies.

Frequently asked questions

What is a factorial survey (vignette experiment)?

It is a survey with a randomised experiment inside it: respondents rate short, realistic scenarios ("vignettes") whose features have been varied independently. Because the features are uncorrelated by design, a regression of the ratings on the features recovers the weight each feature carries in respondents' judgements.

How is it different from a discrete choice experiment?

A discrete choice experiment asks respondents to choose between whole alternatives and infers preferences from those choices. A factorial survey asks respondents to rate full-profile vignettes and infers the weight of each attribute from the ratings. They answer related questions with different response tasks.

Why not just ask students directly how much workload or clarity matters?

Direct importance questions are distorted by introspection and self-presentation - students report what they think they should value. A factorial survey has students rate whole courses that happen to differ in workload and clarity, revealing the true weight without their awareness of which attribute is under study.

How is this different from anchoring vignettes?

Anchoring vignettes calibrate how different groups use a rating scale so their self-reports become comparable. A factorial survey uses vignettes to measure the causal weight of course attributes on judgement. Same word, different purpose.

What is the biggest design mistake to avoid?

Two: creating implausible attribute combinations (a "large one-to-one seminar") and overloading respondents with too many dimensions or too many vignettes. Both corrupt the recovered weights - the first by breaking realism, the second by triggering satisficing.

How should factorial-survey data be analysed?

With multilevel (mixed) models, because multiple vignette ratings are nested within each respondent. Treating each rating as an independent observation understates the standard errors and overstates significance.

References

  • Rossi, P. H., & Anderson, A. B. (1982). The factorial survey approach: An introduction. In P. H. Rossi & S. L. Nock (Eds.), Measuring social judgments: The factorial survey approach (pp. 15-67). Sage Publications.
  • Auspurg, K., & Hinz, T. (2015). Factorial survey experiments (Quantitative Applications in the Social Sciences, No. 175). SAGE Publications. https://doi.org/10.4135/9781483398075
  • Wallander, L. (2009). 25 years of factorial surveys in sociology: A review. Social Science Research, 38(3), 505-520. https://doi.org/10.1016/j.ssresearch.2009.03.004
  • Atzmüller, C., & Steiner, P. M. (2010). Experimental vignette studies in survey research. Methodology, 6(3), 128-138. https://doi.org/10.1027/1614-2241/a000014

Related resources