Sliders, Visual-Analogue, or Radio Buttons? The Evidence on Course-Evaluation Response Formats
Slider widgets look modern, but the survey-methodology evidence says they cost you data. Funke (2016), Couper et al. (2006) and Bosch et al. (2019) on choosing a response widget for online course evaluations.
Koji Education Team
Product
In short: The widget students click — a slider, a continuous visual-analogue scale (VAS), or plain radio buttons — measurably affects your data, not just the look of the survey. The strongest experimental evidence (Funke 2016) finds slider scales perform worst: they lower response rates, distort the sample, increase missing data and slow respondents down, especially on mobile. Plain radio buttons and well-built VAS are the safe choices. For online course evaluations completed largely on phones, the practical recommendation is unambiguous: default to radio buttons; avoid slider widgets.
What the research says
The anchor is Frederik Funke's "A Web Experiment Showing Negative Effects of Slider Scales Compared to Visual Analogue Scales and Radio Button Scales" (Social Science Computer Review, 2016, 34(2), 244–254). Funke ran a controlled web experiment comparing three widgets that look similar but behave differently: HTML radio buttons, a visual-analogue scale (a continuous line on which respondents mark a point), and a slider (a draggable handle that often starts at a default position). Each was tested with three, five, or seven response options where relevant.
The results were one-sided against sliders. Compared with radio buttons, slider scales lowered the response rate (worst on mobile devices), altered the sample composition, distorted the distribution of values, and increased response times. They also produced more incomplete data — a recurring failure mode is that respondents simply do not move the handle, so the survey cannot distinguish "I chose the default" from "I didn't answer." Funke's recommendation is explicit: avoid slider scales; VAS and radio buttons can be used without these side effects, even on touchscreens.
Two corroborating sources triangulate the picture:
- Couper, Tourangeau, Conrad and Singer (2006), "Evaluating the Effectiveness of Visual Analog Scales: A Web Experiment" (Social Science Computer Review, 24(2), 227–245), compared a VAS against radio buttons and numeric text entry. The VAS response distributions did not differ statistically from the other formats, but the VAS had higher rates of missing data and longer completion times — an early signal that fancier continuous widgets carry a burden cost even when they don't bias the central tendency.
- Bosch, Revilla, DeCastellarnau and Weber (2019), "Measurement Reliability, Validity, and Quality of Slider Versus Radio Button Scales in an Online Probability-Based Panel in Norway" (Social Science Computer Review, 37(3), 401–420), assessed psychometric quality directly and found sliders did not outperform radio buttons on reliability and validity, while remaining more demanding for some respondents — reinforcing that the modern widget buys you no measurement advantage.
The convergent conclusion across a decade of experiments: continuous and draggable widgets add respondent burden and missing-data risk without improving measurement. Radio buttons remain the workhorse; VAS is acceptable when a continuous scale is genuinely needed; sliders are the format to avoid.
Why it matters for course evaluation in practice
Course evaluations are now overwhelmingly online and completed on phones (see device effects), which is exactly where Funke found slider penalties to be largest. The consequences map directly onto the things QA teams care about:
- Response rate and representativeness. Anything that adds friction — a fiddly handle that is hard to place precisely on a small touchscreen — costs completions, and lost completions are rarely random. Lower, more selective response undermines the non-response story you must tell accreditors.
- Missing data masquerading as a default. A slider that starts at the midpoint and is never moved silently manufactures a pile of "3 out of 5" ratings that no student actually chose. That is a validity hazard, not a cosmetic one.
- False precision. A VAS yields a 0–100 number that looks more precise than a five-point scale, but the underlying judgment is no finer-grained — and over-precise reporting invites the over-interpretation of trivial differences that responsible reporting warns against.
- Accessibility. Draggable handles are harder for students using assistive technology, keyboard navigation, or motor-impaired input, raising both an inclusion and a compliance concern.
The practical rule for an online SET instrument: radio buttons for ordinal scale items; reserve VAS only where a genuinely continuous response is meaningful; do not use sliders.
Limitations & honest caveats
A careful reader should weigh several caveats:
- Implementation matters as much as format. "Slider" covers a range of designs; a slider with no default handle position, a visible numeric readout, and large touch targets behaves better than a naïve one. Some of the penalty Funke found is attributable to specific implementations, not the abstract concept.
- Technology has moved. The 2006 and 2016 experiments predate the latest mobile browsers and touch refinements; modern responsive widgets may narrow the gap, though the 2019 Norwegian panel still found no slider advantage.
- Context dependence. These experiments used general-population web panels, not students rating courses; the direction of the effects should transfer, but exact magnitudes for a university cohort are an extrapolation.
- Central tendency is often unaffected. Couper et al. found VAS distributions did not differ from radio buttons — so the harm is concentrated in missing data, burden, and dropout, not necessarily in biased averages. If your only concern is the mean of completers, the format effect is smaller than the response-rate effect.
- Not the biggest lever. Widget choice is a hygiene factor. It will not offset the dominant threats — timing, response rate, grading and identity biases — documented elsewhere in this knowledge base.
How Koji incorporates this
Koji's design follows the evidence and is built to mitigate format-induced data loss rather than chase visual novelty:
- Robust, accessible defaults. Koji's
scale,single_choice,multiple_choice,yes_noandrankingquestion types render as clear, tappable controls suited to mobile-first completion — the radio-button-style reliability Funke and Bosch et al. endorse — rather than fragile draggable handles. - No silent defaults. Where Koji uses scale inputs, an unanswered item stays genuinely unanswered, so missing data is recorded as missing rather than being manufactured into a midpoint — directly addressing the slider failure mode of "never moved the handle".
- Less dependence on the widget at all. Koji's AI-moderated conversational interview elicits the substance in natural language and probes follow-ups, so the most important judgments do not hinge on how precisely a student can place a marker on a line. The widget stops being the bottleneck.
- Honest precision. Koji's reporting and automatic thematic analysis present findings at a granularity the data can support, helping QA teams avoid the false precision a 0–100 VAS invites, and supporting bias-aware interpretation of small differences.
These are framed as designed to reduce — not eliminate — format and burden effects. The same engine runs product and customer research on Koji's core platform at koji.so, where slider-heavy "rate this 0–100" surveys create the identical dropout-and-missing-data problem.
Related Resources
- Do students on phones give worse course evaluations? Device effects
- Online vs paper course evaluations and response rates
- How many scale points should a question have?
- Should every point on a course-evaluation scale be labelled?
- How long should a course evaluation be?
- Why students click straight down the middle: satisficing
A response-format checklist for online evaluation teams
When you configure or procure an online course-evaluation tool, the widget decisions below turn the experimental evidence into practice:
-
Default to radio-button-style scale controls. For ordinal rating items they are the best-evidenced choice — robust across Funke (2016) and Bosch et al. (2019), fast on mobile, and accessible to keyboard and assistive-technology users.
-
Treat sliders as a red flag. If a vendor demo leads with draggable sliders, ask how they handle the "never moved the handle" case. A slider that starts mid-scale and records that default as a real answer is silently fabricating data. Prefer a control where no answer is no answer.
-
Reserve VAS for genuinely continuous judgments. A visual-analogue scale is defensible when the underlying quantity really is continuous and you have a reason to want fine gradation. For a five-point "how clear was the lecturer" item it adds burden and missing-data risk without measurement gain.
-
Test on a real phone, not just a laptop. The slider penalty Funke found was largest on mobile, which is where most students now complete evaluations. A widget that feels fine on a 27-inch monitor can be unusable with a thumb on a moving train.
-
Make "no response" recoverable, not default. Ensure unanswered scale items are stored as missing rather than imputed to a midpoint. This protects both your non-response analysis and the honesty of your means.
-
Match reporting precision to the input. If you do use a 0–100 VAS, resist reporting differences of one or two points as meaningful. The continuous readout looks precise, but Couper et al. (2006) showed its distribution is no finer-grained than a radio-button scale — so report at a granularity the data can actually support.
Run through this list once at instrument-build time and the format question is settled for every cohort that follows, with the data-quality risks closed off before a single student clicks submit.
References
- Funke, F. (2016). A web experiment showing negative effects of slider scales compared to visual analogue scales and radio button scales. Social Science Computer Review, 34(2), 244–254. https://doi.org/10.1177/0894439315575477
- Couper, M. P., Tourangeau, R., Conrad, F. G., & Singer, E. (2006). Evaluating the effectiveness of visual analog scales: A web experiment. Social Science Computer Review, 24(2), 227–245. https://doi.org/10.1177/0894439305281503
- Bosch, O. J., Revilla, M., DeCastellarnau, A., & Weber, W. (2019). Measurement reliability, validity, and quality of slider versus radio button scales in an online probability-based panel in Norway. Social Science Computer Review, 37(3), 401–420. https://doi.org/10.1177/0894439317750089
Related articles
How Many Scale Points Should a Course-Evaluation Question Have?
What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.
Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
Online course evaluations consistently draw lower response rates than in-class paper forms, but the research shows the resulting scores are largely equivalent. Here is what Dommeyer, Nulty, and Stowell actually found, and what an adequate response rate really requires.
Why Students Click Straight Down the Middle: Satisficing in Course Evaluations
A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.
How Long Should a Course Evaluation Be? Questionnaire Length, Breakoff, and Answer Quality
What the survey-methodology evidence says about questionnaire length: longer instruments depress participation and degrade answers to later questions, but ruthless shortening is not automatically the answer. A research-grounded guide for designing course evaluations.