Which Evaluation Questions Are Redundant? Mutual Information and mRMR for Trimming Your Questionnaire
Two evaluation items are redundant to the degree that one predicts the other — mutual information. The mRMR rule keeps items that are maximally relevant to your decision and minimally redundant with each other, yielding a shorter, sharper questionnaire.
Koji Education Team
Product
Course-evaluation questionnaires accrete. Every review adds an item someone felt strongly about; almost none removes one. The result is a form where "the lecturer explained clearly", "the lecturer communicated the material well", and "the teaching was easy to follow" all ask, from the student's point of view, the same thing — and each extra near-duplicate lengthens the survey, depresses response rates, and inflates the apparent internal consistency of the instrument without adding information. The question "which of these items are redundant?" has a precise, model-free answer from information theory: mutual information, and the item-selection rule built on it, minimum-redundancy–maximum-relevance (mRMR).
The short answer (BLUF)
Two evaluation items are redundant to the extent that knowing the answer to one tells you the answer to the other — a quantity called mutual information. An item earns its place on the questionnaire only if it carries information the other items do not. The mRMR criterion (Peng, Long, and Ding, 2005) formalises the trade-off: keep items that are highly relevant (high mutual information with the outcome you care about, such as the overall rating or a decision) while being minimally redundant (low mutual information with items already selected). Unlike correlation, mutual information captures non-linear and non-monotone dependence; unlike item-response-theory information, it needs no fitted measurement model. The payoff is a shorter, sharper questionnaire that measures the same constructs with fewer questions — which typically raises response rates and data quality rather than trading them away.
What the research says
The foundational quantity is Shannon's (1948) entropy, the amount of uncertainty in a variable, and its extension mutual information, the reduction in uncertainty about one variable from knowing another. Mutual information is zero if and only if two variables are statistically independent, and — unlike the Pearson correlation — it detects any form of dependence, not just linear association. This makes it well suited to Likert data, where relationships are often monotone-but-curved or concentrated at the scale ends.
Peng, Long, and Ding (2005) turned mutual information into a practical feature-selection method. Their insight was that choosing features (here, items) purely for individual relevance produces redundant sets — the top items are often near-copies of one another. mRMR instead maximises relevance to the target while penalising redundancy among the chosen features, and they showed it yields more compact, better-performing feature sets than relevance-only selection across diverse datasets. The method has since become a standard tool wherever many correlated measurements must be pruned to an informative few.
Two adjacent literatures reinforce the point for questionnaires. The psychometric work on short-form construction shows that well-chosen subsets of items can retain most of a scale's information — the redundancy in long instruments is real and removable. And the survey-methodology evidence on questionnaire length shows that longer surveys reduce response rates and increase satisficing and break-off, so every redundant item carries a data-quality cost, not just a time cost. Information-theoretic pruning targets exactly the items that cost response quality while adding least.
Why it matters for course evaluation in practice
Redundant items masquerade as reliability. Cronbach's alpha rises mechanically as you add items that correlate with each other, so three paraphrases of "clear teaching" inflate alpha while adding almost no information — a classic case where a reliability statistic rewards redundancy. Mutual information exposes the paraphrases as near-duplicates that alpha was quietly double-counting.
Shorter forms get more, and more honest, responses. Cutting a 30-item instrument to the dozen items that carry distinct information reduces the burden that drives students to straightline or abandon the survey. The information you lose is, by construction, the information the retained items already carried.
Relevance is defined by your decision. mRMR asks "relevant to what?" — you choose the target. If the goal is predicting the overall rating, keep items that inform it; if the goal is diagnosing why a course underperforms, keep items that carry distinct diagnostic information. The same corpus can yield different optimal short forms for different QA purposes, made explicit rather than left to committee preference.
It is a principled complement to dimensionality analysis. Parallel analysis tells you how many dimensions exist; mutual information and mRMR tell you the fewest items that still span them without repetition — the operational next step after you know the structure.
Limitations and honest caveats
Estimating mutual information from modest samples is genuinely hard: with a handful of respondents and several scale points, empirical mutual information is biased upward and unstable, so item-pruning decisions from small classes are unreliable and should pool across many classes and terms. Mutual information is also purely statistical — an item can be low-information yet essential for content-validity, accreditation, or legal reasons (an accessibility item you are required to ask), and no information criterion should override that; redundancy analysis informs the human decision, it does not make it. Because mutual information ignores measurement structure, it will not, by itself, tell you whether the retained items still form a coherent scale — it should be run alongside a factor or IRT check, not instead of one. mRMR's greedy selection is a heuristic, not a guaranteed optimum, and its results depend on how you discretise or bin the responses. And "relevance to the overall rating" is only a good target if the overall rating is itself a good criterion — optimising items to predict a biased summary measure will faithfully reproduce the bias. Information theory optimises efficiency of measurement, which is necessary but not sufficient for validity of measurement.
How Koji incorporates this
Koji is designed to keep questionnaires short and non-redundant rather than letting them accrete. Because Koji stores every item response in structured form across many courses and cohorts, it has the pooled data needed to estimate mutual information reliably — the scale at which redundancy analysis actually works — and to flag items that duplicate the information already captured by others. When an institution's instrument has grown three near-identical "clarity" questions, Koji's analysis can surface that they carry overlapping information, so a programme can retain the one that best informs its chosen decision and retire the rest, guided by mRMR-style relevance-versus-redundancy reasoning rather than by which item a committee happens to favour. This directly serves data quality: shorter, information-dense forms reduce the survey fatigue that drives straightlining and break-off. Crucially, Koji does not treat efficiency as the only criterion — items required for accreditation, accessibility, or content-validity are preserved regardless of their statistical information, because redundancy analysis is designed to inform the human editing decision, not automate it. Koji's AI-moderated conversational interview changes the calculus too: rather than adding another fixed item to chase a nuance, the interview probes it dynamically only when relevant, so the standing questionnaire can stay lean while coverage stays deep. The same efficiency-of-measurement thinking runs through Koji's core research platform at koji.so, where product and customer surveys face the identical pressure to measure more with fewer questions. As with every method here, Koji frames this as mitigating redundancy and burden — not as a guarantee that a shorter form is automatically a more valid one.
Frequently asked questions
What is mutual information, in plain terms?
Mutual information measures how much knowing the answer to one question reduces your uncertainty about the answer to another. If two evaluation items are near-duplicates, knowing one tells you the other, so their mutual information is high — the signal that one of them is redundant. If two items are independent, their mutual information is zero.
How is mutual information different from correlation?
Correlation only detects linear (straight-line) association and can miss curved or end-loaded relationships common in Likert data. Mutual information detects any statistical dependence, linear or not, and is zero only when two items are truly independent. It is a more general redundancy detector.
What does mRMR actually do?
mRMR selects a subset of items that are individually informative about a target (maximum relevance) while being minimally redundant with each other (minimum redundancy). It prevents the common failure where the "top" items are all near-copies, giving you a compact set that spans the information rather than repeating it.
Won't cutting items lower my Cronbach's alpha and hurt reliability?
Alpha often falls when you remove redundant items, but that fall is largely cosmetic: alpha rises mechanically with redundancy, so a high alpha built on paraphrases overstates reliability. Removing duplicates costs little real information. Check coefficient omega and the factor structure rather than trusting alpha alone.
Can I use this to shorten the survey for small classes?
Not from small classes directly — mutual information estimates are unstable with few respondents. Pool responses across many classes and terms to decide which items are redundant, then apply the shortened form everywhere, including small classes.
Should information content ever override keeping an item?
No. An item can be statistically low-information yet mandatory for accreditation, accessibility, or content-validity reasons. Redundancy analysis informs the human decision about what to keep; it should never automatically delete an item that exists for a non-statistical reason.
Related resources
- Computerized Adaptive Testing for Shorter Course Evaluations
- Parallel Analysis for Factor Retention
- Planned Missingness and Split-Questionnaire Designs
- Random Forests and Variable Importance for Course-Evaluation Drivers
- Relative Weights and Dominance Analysis
- The Total Survey Error Framework
References
- Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27(3), 379–423. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x
- Peng, H., Long, F., & Ding, C. (2005). Feature selection based on mutual information: Criteria of max-dependency, max-relevance, and min-redundancy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 27(8), 1226–1238. https://doi.org/10.1109/TPAMI.2005.159
- Cover, T. M., & Thomas, J. A. (2006). Elements of Information Theory (2nd ed.). Wiley. https://doi.org/10.1002/047174882X
- Rolstad, S., Adler, J., & Rydén, A. (2011). Response burden and questionnaire length: Is shorter better? A review and meta-analysis. Value in Health, 14(8), 1101–1108. https://doi.org/10.1016/j.jval.2011.06.003
Related articles
How Many Dimensions Does Your Evaluation Really Measure? Parallel Analysis for Factor Retention
Deciding how many factors your evaluation instrument measures with the eigenvalue-greater-than-one rule or the scree plot routinely gets the wrong answer. Parallel analysis (Horn, 1965) compares your data against random noise — and is one of the most accurate methods available.
Which Aspects of Teaching Actually Drive the Overall Score? Relative Weights and Dominance Analysis
When clarity, workload, support, and feedback are all correlated, regression coefficients cannot tell you which one drives the overall rating. Relative weights and dominance analysis partition the explained variance fairly.
Which Teaching Behaviours Drive the Overall Score — Nonlinearly? Random Forests and Variable Importance for Course Evaluation
Random forests and their variable-importance measures reveal which teaching items predict the overall rating when the relationships are nonlinear and interacting — but Gini importance is biased toward correlated and high-cardinality predictors, so use conditional permutation importance and read the output as association, not cause.
Shorter Surveys Without Losing Coverage: Planned Missingness for Course Evaluations
Planned missing data designs let you cover more questions while each student answers fewer. Here is how the three-form design works, what the evidence says, and where it fits in course evaluation.