New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting11 min read

One Instructor, Two Hundred Defensible Scores: Specification-Curve Analysis for Course Evaluation

Every course-evaluation "score" is the product of dozens of defensible analytic choices — which items count, how to weight, mean or median, whether to adjust for class size. Specification-curve and multiverse analysis compute the result across all of them, so a personnel decision rests on the pattern, not on one analyst's arbitrary path.

Koji Education Team

Product

There is no such thing as the score for an instructor. The number on a personnel report is one path through a forest of defensible choices — which items to average, how to handle non-response, mean or median, whether to adjust for class size and discipline, how to treat the neutral midpoint, whether to shrink small classes toward the mean. Each choice is arguable; together they can produce hundreds of different "scores" from the same raw responses. Specification-curve analysis (and its sibling, multiverse analysis) computes the result across that whole set of defensible specifications and displays the distribution, so a high-stakes decision can rest on how robust the conclusion is across choices — not on which fork one analyst happened to take.

Answer box (BLUF): When a committee reads "Dr. Alvarez scored 3.9, below the 4.0 benchmark," that verdict depends on analytic choices that were never surfaced: different but equally defensible decisions about weighting, non-response, averaging method, and adjustment can move the same instructor above or below the line. Specification-curve analysis makes this visible by (1) enumerating the set of defensible, non-redundant specifications, (2) plotting the resulting estimates from most to least favourable, and (3) drawing inference from the whole curve. If the "below benchmark" conclusion holds across most defensible specifications, trust it; if it flips depending on arbitrary choices, the decision is arbitrary, and you should say so.

What the research says

The method was formalised by Simonsohn, Simmons, and Nelson (2020) in Nature Human Behaviour, "Specification curve analysis." They describe three steps: identify the set of theoretically justified, statistically valid, and non-redundant specifications; display the results graphically so readers can see which analytic decisions are consequential; and conduct joint inference across all specifications rather than cherry-picking one. The central insight is that empirical results "hinge on analytical decisions that are defensible, arbitrary and motivated" — and those decisions introduce variability that a single standard error never captures.

Steegen, Tuerlinckx, Gelman, and Vanpaemel (2016), "Increasing transparency through a multiverse analysis" (Perspectives on Psychological Science), make the same argument from the data-processing side. Before any model is fit, choices about excluding, transforming, and coding responses create a "multiverse" of datasets. Running the analysis across all reasonable versions shows how much the conclusion depends on arbitrary construction choices, and which choices matter most.

The most vivid empirical demonstration is Silberzahn et al. (2018), "Many analysts, one data set" (Advances in Methods and Practices in Psychological Science). Twenty-nine independent teams were given the same dataset and the same question — are soccer referees more likely to red-card dark-skinned players? Estimated effect sizes ranged from an odds ratio of 0.89 to 2.93; twenty teams (69%) found a significant positive effect and nine (31%) did not. Analyst expertise and prior beliefs did not explain the spread. If experts diverge that widely on one dataset, a single instructor "score" computed by one office one way should be read as one draw from a wide distribution, not as ground truth.

Underpinning all of this is Gelman and Loken's (2014) "garden of forking paths" argument (American Scientist): even an honest analyst who runs only one analysis has implicitly chosen it from many they would have run had the data looked different — so the reported precision overstates the certainty.

Why it matters for course evaluation in practice

Course-evaluation numbers are unusually rich in forking paths, and unusually consequential when they cross a threshold. Consider a partial list of defensible-but-arbitrary choices baked into a single instructor average:

  • Which items? Overall-satisfaction only, or a composite of organisation, clarity, and workload items — and if a composite, equally or differentially weighted?
  • Central tendency. Mean, median, or top-box percentage ("% agree or strongly agree")?
  • Non-response. Analyse respondents as-is, weight to the enrolled population, or restrict to classes above a response-rate floor?
  • Small classes. Report raw averages, or shrink small classes toward the departmental mean (empirical-Bayes)?
  • Adjustment. Leave scores raw, or adjust for class size, level, discipline, and required-vs-elective status — each a documented confound?
  • Scale treatment. Treat the Likert response as an interval number, or model it ordinally?
  • Outliers and straight-lining. Keep every response, or filter careless responders?

Any single report picks one cell from this grid silently. Specification-curve analysis picks all the defensible cells and asks: does the conclusion survive? For personnel and quality decisions, this reframes the question from "what is the score?" to "how fragile is the verdict?" A rating that sits below benchmark under 95% of defensible specifications is a signal worth acting on. A rating that crosses the benchmark in half of them is a coin flip dressed up as a measurement, and treating it as decisive is indefensible — and, where evaluations feed promotion, potentially discriminatory if the fragile cases cluster by gender or discipline.

Limitations and honest caveats

Specification-curve analysis is a discipline, not a magic wand, and a rigorous reader will note several boundaries:

  • "Defensible" is itself a judgment. The curve is only as honest as the set of specifications you admit. Padding it with weak specifications to dilute an uncomfortable result, or excluding defensible ones that would overturn a comfortable one, reintroduces exactly the bias the method is meant to expose. The specification set must be pre-registered or at least transparently justified.
  • It does not fix bad data or confounding. If every specification shares the same flaw — a biased instrument, a non-representative sample, an unmeasured confounder — the whole curve is biased in the same direction. Robustness across specifications is not validity.
  • Joint inference rests on assumptions. The permutation-based tests Simonsohn et al. propose assume the specifications are exchangeable under the null in a particular way; with heavily correlated specifications the test can be over- or under-powered.
  • Not all forks are equal. Some choices (mean vs. median) are cosmetic; others (adjust for discipline or not) are substantive. A curve treats them as points on the same axis, which can obscure that one decision, not the spread, is what matters. Reading which specifications move the result is often more informative than the overall curve.
  • Communication burden. A dean does not want a 200-point plot. The method's value in a QA setting is often internal — to decide whether a verdict is safe to report — rather than something to hand to every committee.

The honest framing is modest: specification-curve analysis tells you whether a conclusion is robust to arbitrary choices. It cannot tell you the conclusion is right.

How Koji incorporates this

Koji cannot make an evaluation decision robust on your behalf, but it is built so that robustness is checkable rather than hidden.

  • Item-level data is preserved, not pre-collapsed. Because Koji retains structured, item-level responses (scale, single_choice, multiple_choice, open_ended) rather than only a stored average, every defensible specification can be recomputed from the same source. You cannot run a multiverse on a spreadsheet that only kept the mean.
  • Configurable, side-by-side reporting. Koji's reporting is designed to show a result under more than one defensible lens — for example, raw mean alongside a small-class-shrunk estimate, or interval alongside ordinal treatment — so a fragile verdict reveals itself instead of hiding behind a single figure. This is the practical, decision-scale version of a specification curve.
  • Bias-aware framing. Rather than presenting one number as authoritative, Koji's reporting philosophy is to flag when a comparison sits close to a threshold or depends heavily on how small classes are handled — the cases where the fork matters most.
  • Reducing reliance on any single fragile number. Koji's AI-moderated conversational interviews add qualitative, mechanism-level evidence that does not live or die by one averaging choice, so a borderline quantitative verdict need not be the sole basis for a decision.
  • A reproducible audit trail. The specification you did use — items, filters, adjustments — is recorded, so a later reviewer (or an accreditation panel) can re-run alternatives and see whether the story holds.

Koji is designed to surface analytic fragility, not to certify robustness it cannot guarantee. The same structured-data engine underlies Koji's core research platform at koji.so, where product teams face identical forking-path risk whenever a single metric crosses a launch threshold.

Frequently asked questions

What is a specification curve in one sentence? It is a plot of the result you would have obtained under every defensible, non-redundant way of analysing the same data, ordered from most to least favourable, so you can see whether your conclusion depends on the analytic path.

How is this different from just running a few robustness checks? Ad hoc robustness checks let the analyst pick which alternatives to show, which invites bias. A specification curve commits in advance to the full set of defensible specifications and reports all of them jointly, so a cherry-picked "robust" result cannot survive undetected.

Does specification-curve analysis mean every result is meaningless? No — the opposite. Many results are robust: they hold across most or all defensible specifications, and the method certifies that. It only undermines conclusions that were fragile all along and happened to be reported from a favourable fork.

Can the method be gamed? Yes. Padding the curve with weak specifications to dilute a result, or omitting defensible ones that would overturn it, reintroduces bias. That is why the specification set should be justified transparently or pre-registered, not chosen after seeing the answer.

How many specifications do I need for a course-evaluation decision? There is no fixed number; it is the count of genuinely defensible, non-redundant choices in your pipeline. Even four or five substantive forks — averaging method, small-class shrinkage, adjustment, scale treatment — are enough to reveal whether a near-threshold verdict is stable.

Does robustness across specifications prove the score is valid? No. If every specification shares the same underlying flaw — a biased instrument or an unmeasured confounder — the whole curve is biased together. Robustness guards against arbitrary analytic choices, not against invalid measurement.

Related resources

References

  • Simonsohn, U., Simmons, J. P., & Nelson, L. D. (2020). Specification curve analysis. Nature Human Behaviour, 4(11), 1208–1214. https://doi.org/10.1038/s41562-020-0912-z
  • Steegen, S., Tuerlinckx, F., Gelman, A., & Vanpaemel, W. (2016). Increasing transparency through a multiverse analysis. Perspectives on Psychological Science, 11(5), 702–712. https://doi.org/10.1177/1745691616658637
  • Silberzahn, R., Uhlmann, E. L., Martin, D. P., et al. (2018). Many analysts, one data set: Making transparent how variations in analytic choices affect results. Advances in Methods and Practices in Psychological Science, 1(3), 337–356. https://doi.org/10.1177/2515245917747646
  • Gelman, A., & Loken, E. (2014). The statistical crisis in science. American Scientist, 102(6), 460–465. https://doi.org/10.1511/2014.111.460

Related articles

analysis-reporting

Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores

Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.

analysis-reporting

Which Instructor Scores Are Genuinely Unusual? Funnel Plots for Fair Course-Evaluation Comparison

Ranking instructors by mean evaluation score turns sampling noise into a league table. Funnel plots — Spiegelhalter's method for institutional comparison — show which scores are genuinely unusual and which are just small-sample wobble.

analysis-reporting

Reading Evaluations to Confirm What You Already Believe: Confirmation Bias in Interpreting Course Feedback

Whoever reads a course evaluation already has a hypothesis about the instructor. Confirmation bias shapes which comments they weight, how ambiguity is resolved, and what the "data" is taken to show. Here is the evidence and the guardrails.

analysis-reporting

Is That Score Gap Real or Just Luck of the Draw? Permutation Tests for Course-Evaluation Comparisons

Comparing two instructors' evaluation averages with a t-test quietly assumes normal, equal-variance data you rarely have with small, skewed Likert samples. Permutation tests, formalised by Fisher and reviewed by Ernst, answer the comparison question by shuffling the data itself — with almost no distributional assumptions. Here is when and how to use them.