Do Your Evaluation Items Actually Cover Teaching? Content Validity, the CVR and the CVI
An evaluation form can be reliable and still measure the wrong things. Content validity, quantified with Lawshe's CVR and the Content Validity Index, tests whether your items cover the teaching domain.
Koji Education Team
Product
In brief: Content validity is the degree to which the items on your evaluation form actually represent the domain of teaching you intend to measure — no important facet left out, nothing irrelevant let in. It is established before data collection, by expert judgment, and can be quantified with Lawshe's Content Validity Ratio (CVR) and the Content Validity Index (CVI). A form can have excellent reliability (consistent scores) and still have poor content validity (consistently measuring the wrong things), so content validation is a separate, prior obligation that most course-evaluation instruments quietly skip.
Reliability gets almost all the psychometric attention in course evaluation — Cronbach's alpha, omega, generalizability coefficients. But reliability only tells you a scale measures something consistently. Content validity asks the prior question: does the scale measure the right something, completely? This article explains content validity, shows how to quantify it with the CVR and CVI, and warns about the ways those indices can mislead.
What the research says
Content validity concerns the match between an instrument's items and the construct domain it is meant to cover. The authoritative conceptual treatment is Haynes, Richard and Kubany (1995), "Content validity in psychological assessment: A functional approach to concepts and methods" (Psychological Assessment, 7(3), 238–247). They define content validity as "the degree to which elements of an assessment instrument are relevant to and representative of the targeted construct," and stress two failure modes drawn from Messick's validity theory: construct underrepresentation (the form omits important facets — say, it asks about lecture clarity but never about assessment or feedback) and construct-irrelevant variance (the form includes items that tap something else — the instructor's warmth, the room, the timetable). Crucially, they argue content validity is not a property an instrument has forever; it is conditional on the specific construct definition and the intended use, and must be argued, not assumed.
The most widely used quantitative index is Lawshe's (1975) Content Validity Ratio, from "A quantitative approach to content validity" (Personnel Psychology, 28(4), 563–575). A panel of subject-matter experts rates each item as "essential," "useful but not essential," or "not necessary." For each item, the CVR is computed as CVR = (n_e − N/2) / (N/2), where n_e is the number of experts calling the item essential and N is the panel size. The index runs from −1 (no expert says essential) to +1 (all do); Lawshe supplied a table of minimum CVR values that exceed chance agreement for a given panel size. Items below the threshold are candidates for revision or deletion.
The complementary tool is the Content Validity Index, dissected in Polit and Beck (2006), "The content validity index: Are you sure you know what's being reported? Critique and recommendations" (Research in Nursing & Health, 29(5), 489–497). Experts rate each item's relevance on a scale (typically 1–4). The item-level CVI (I-CVI) is the proportion of experts rating an item 3 or 4. The scale-level CVI can be computed two ways that Polit and Beck show are routinely confused: S-CVI/UA (universal agreement — the proportion of items every expert endorsed) and the more lenient S-CVI/Ave (the average of the I-CVIs). Their central methodological warning is that authors rarely say which they used, producing incomparable numbers, and that raw agreement indices ignore chance; they recommend reporting a modified kappa that adjusts I-CVI for the probability of chance agreement. Their commonly cited benchmarks — I-CVI ≥ 0.78 for panels of six or more, S-CVI/Ave ≥ 0.90 — are useful but, they caution, not magic thresholds.
These tools sit alongside, not inside, the broader validity argument. Content validity supports the item-selection step in Kane's argument-based framework (see our argument-based validity article) and is distinct from convergent and discriminant evidence (the multitrait-multimethod test) and from the dimensional structure of ratings (the dimensionality debate).
Why it matters for course evaluation in practice
Most course-evaluation forms were never content-validated. They accreted over decades, borrowed items from other institutions, or codified whatever a committee found intuitive. Three concrete consequences:
-
Construct underrepresentation. A form heavy on "the lecturer explained clearly" and "the lecturer was enthusiastic" but silent on assessment design, feedback quality, inclusivity, or intellectual challenge measures a narrow slice of teaching and then reports it as an overall verdict. Decisions built on it systematically undervalue the facets the form forgot — which is one reason "assessment and feedback" so often scores oddly and why low-inference behavioural items and multidimensional models like the SEEQ matter.
-
Construct-irrelevant variance. Items that tap likeability, physical setting, or difficulty import confounds into what is reported as a teaching score. Content validation by experts is the front-line screen for such items before they ever reach a student.
-
A defensible audit trail. When an instructor challenges a low score in a personnel or promotion case, "our items were rated essential by a panel of teaching experts, with documented CVR and CVI" is a far stronger position than "these are the questions we have always asked." Content validity evidence is exactly the kind of documentation that quality-assurance and accreditation reviewers increasingly expect.
The practical workflow is modest: define the teaching-quality domain explicitly (ideally from a teaching framework), assemble a small expert panel, have them rate each candidate item for essentiality and relevance, compute CVR and I-CVI, and revise or cut items that fall short. Done once when a form is designed or overhauled, it improves every future response.
Limitations and honest caveats
- Necessary, not sufficient. Content validity establishes that items cover the intended domain; it says nothing about whether students interpret them correctly, whether the scale is reliable, or whether scores predict learning. It is one strand of a validity argument, not the whole rope.
- Only as good as the domain definition. If the teaching-quality construct is defined narrowly or wrongly, a panel can certify high content validity for a form that still measures the wrong thing. Garbage domain, garbage validation.
- Expert-dependent and potentially biased. CVR and CVI reflect the particular panel. A homogeneous panel can share blind spots, and disciplinary norms differ on what "essential" teaching looks like. Panel composition is a substantive choice, not a formality.
- The indices ignore chance and are sensitive to panel size. As Polit and Beck (2006) show, high raw agreement can arise by chance with small panels; a modified kappa is needed to correct for it, and the S-CVI benchmark you meet depends on which of the two definitions you use. Reporting a bare "CVI = 0.9" without specifying the method is uninformative.
- Static snapshot. Content validity is conditional on a construct and context that evolve — new modalities, new expectations. A form validated a decade ago may now underrepresent the domain, so validation is periodic, not permanent.
Content validity is a powerful, cheap discipline for instrument design, provided you treat its numbers as structured expert judgment rather than objective proof.
How Koji incorporates this
The Koji for Education platform is built to make domain coverage explicit and to catch the gaps content validation targets.
- Domain-first study design. Koji encourages defining what a study is meant to measure up front and building typed questions (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) against that domain, rather than pasting in a legacy item list. This operationalises the first and most important step of content validation: an explicit construct definition.
- Coverage checking through open text and thematic analysis. Koji's AI-moderated conversational interviews collect open-ended responses that its automatic thematic analysis then clusters. When students repeatedly raise a theme the structured items never asked about, that is direct, data-driven evidence of construct underrepresentation — a gap in content coverage surfacing itself, complementing the a-priori expert panel.
- Probing to reduce construct-irrelevant variance. Because the conversational moderator can ask why a rating was given, Koji helps distinguish teaching-relevant signal from irrelevant noise (mood, room, timetable), supporting the same goal the expert relevance rating serves.
- A documented instrument trail. Koji's structured configuration provides the audit trail — which items, measuring which facets — that content-validity documentation and accreditation review require.
- Iteration across cycles. Mid-cycle and formative collection let a revised item set be piloted and its coverage re-checked, matching the guidance that content validity is periodic, not permanent.
Koji is careful not to overclaim: the platform is designed to support content validation and to surface coverage gaps empirically, but it does not replace a deliberate expert panel or the substantive judgment of what the teaching-quality domain should include. For teams running product or customer research, Koji's core platform at koji.so applies the same domain-first design and thematic coverage analysis to those instruments.
Frequently asked questions
What is the difference between content validity and reliability? Reliability is about consistency — whether the scale gives stable, internally coherent scores. Content validity is about coverage — whether the items represent the intended domain without omissions or irrelevant additions. A form can be highly reliable yet have poor content validity, because it consistently measures the wrong things.
How do I calculate Lawshe's CVR? Have expert panellists rate each item as essential, useful-but-not-essential, or not necessary. For each item, CVR = (n_e − N/2) / (N/2), where n_e is the number saying essential and N is the panel size. Compare each item's CVR to Lawshe's (1975) critical value for your panel size and cut or revise items that fall below it.
What is the difference between I-CVI and S-CVI? The item-level CVI (I-CVI) is the proportion of experts rating a single item as relevant. The scale-level CVI (S-CVI) summarises the whole instrument, either as universal agreement across items (S-CVI/UA) or the average of the I-CVIs (S-CVI/Ave). Polit and Beck (2006) stress you must report which you used, and ideally a chance-corrected modified kappa.
How many experts do I need for a content-validity panel? There is no fixed number, but small panels make chance agreement more likely and thresholds stricter. Lawshe's critical CVR values rise as panels shrink, and Polit and Beck recommend benchmarks (I-CVI ≥ 0.78) calibrated for panels of six or more. Diversity of expertise matters as much as size.
Does high content validity mean the evaluation predicts learning? No. Content validity only certifies that items cover the intended domain. Whether scores relate to actual learning is a separate, criterion-related question addressed by multisection and value-added studies, not by the CVR or CVI.
Can content validity be established after data collection? The core judgment — do these items represent the domain? — is made by experts before or independent of the data. Response data can reveal coverage gaps after the fact (students raising unasked themes), but that supplements rather than replaces the a-priori expert review.
Related resources
- Validity Is About the Use, Not the Instrument: Kane's Argument-Based Framework
- What Do Student Evaluations Actually Measure? Marsh and the SEEQ
- The Case for Low-Inference Teaching-Behaviour Items
- The Dimensionality Debate: One Number or Many?
- Measurement Invariance and Differential Item Functioning
- The Multitrait-Multimethod Test of Convergent and Discriminant Validity
References
- Lawshe, C. H. (1975). A quantitative approach to content validity. Personnel Psychology, 28(4), 563–575. https://doi.org/10.1111/j.1744-6570.1975.tb01393.x
- Polit, D. F., & Beck, C. T. (2006). The content validity index: Are you sure you know what's being reported? Critique and recommendations. Research in Nursing & Health, 29(5), 489–497. https://doi.org/10.1002/nur.20147
- Haynes, S. N., Richard, D. C. S., & Kubany, E. S. (1995). Content validity in psychological assessment: A functional approach to concepts and methods. Psychological Assessment, 7(3), 238–247. https://doi.org/10.1037/1040-3590.7.3.238
- Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
Related articles
Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?
Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.
Stop Asking "Was the Lecturer Clear?": The Case for Low-Inference Teaching-Behaviour Items
High-inference items like "Is the instructor clear?" tell teachers nothing about what to change. Murray's low-inference behaviour research shows how concrete, observable items make evaluations diagnostic.
One Number or Many? The Dimensionality Debate and How to Use Student Ratings for Personnel Decisions
Should a promotion committee use a single global teaching score or a detailed profile of many dimensions? The 1997 d'Apollonia & Abrami vs Marsh & Roche debate set the terms — and the answer depends on whether the purpose is summative judgement or formative improvement.