New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Bifactor and ESEM Models: When a Multidimensional Course Evaluation Still Justifies One Overall Score

Bifactor and exploratory structural equation models let you estimate a general teaching-quality factor and specific factors at once, then test with omega-hierarchical and ECV whether reporting a single overall course-evaluation score is defensible.

Koji Education Team

Product

In brief

A bifactor model estimates one general factor (call it overall teaching quality) and several specific factors (clarity, workload, feedback, assessment) from the same set of course-evaluation items at the same time. Two indices then answer the question every quality office actually cares about: omega-hierarchical (ω_h) — how much of the reliable variance in the total score comes from the general factor — and Explained Common Variance (ECV) — the share of common variance the general factor carries. When ω_h is roughly above 0.75 and ECV above 0.70, a single overall number is a defensible summary despite the instrument being multidimensional; below that, averaging everything into one score throws away the very information a programme director needs. Bifactor and its cousin ESEM (exploratory structural equation modelling) give you an empirical test for the oldest argument in student evaluation: one number, or a profile?

What the research says

The dimensionality of student evaluations of teaching (SET) has been contested for forty years. Marsh built the Students Evaluations of Educational Quality (SEEQ) as an explicitly nine-factor instrument, while others argued a single global impression dominates. The dimensionality debate has usually been fought with ordinary confirmatory factor analysis (CFA) — fit a one-factor model, fit a correlated-factors model, compare. The bifactor model reframes the question, because it lets both structures coexist and be measured against each other.

Reise (2012), in "The Rediscovery of Bifactor Measurement Models" (Multivariate Behavioral Research, 47(5), 667-696, doi:10.1080/00273171.2012.715555), set out the modern case: many psychological scales are multidimensional but still support a meaningful total score, and the bifactor model is the natural tool for deciding whether that is true. Each item loads on the general factor and on exactly one specific (group) factor; the factors are orthogonal. The general loadings tell you how strongly the total score is saturated by one common dimension; the group loadings tell you what reliable, specific variance survives after the general factor is removed.

Rodriguez, Reise and Haviland (2016) operationalised this into reportable numbers in two companion papers — "Evaluating bifactor models: Calculating and interpreting statistical indices" (Psychological Methods, 21(2), 137-150, doi:10.1037/met0000045) and "Applying bifactor statistical indices in the evaluation of psychological measures" (Journal of Personality Assessment, 98(3), 223-237, doi:10.1080/00223891.2015.1089249). The key indices:

  • ω_h (omega-hierarchical): proportion of total-score variance attributable to the single general factor. High ω_h means an overall average is dominated by one thing and is interpretable as such.
  • ω_hs (omega-hierarchical subscale): the reliable variance a subscale adds after removing the general factor. Low ω_hs means a subscale score is mostly redundant with the total.
  • ECV: proportion of common variance explained by the general factor; a rough unidimensionality gauge.
  • PUC (percentage of uncontaminated correlations) and the H index (construct replicability): how trustworthy the factor structure is likely to be in a new sample.

Reise, Moore and Haviland (2010), "Bifactor models and rotations" (Journal of Personality Assessment, 92(6), 544-559, doi:10.1080/00223891.2010.496477), showed how bifactor structure can be recovered exploratorily, which matters because a hand-specified CFA often mis-assigns items. That exploratory spirit connects to ESEM (Marsh, Morin, Parker and Kaur, 2014, "Exploratory structural equation modeling," Annual Review of Clinical Psychology, 10, 85-110, doi:10.1146/annurev-clinpsy-032813-153700), which relaxes CFA's requirement that every cross-loading be exactly zero. Real course-evaluation items are rarely so clean — a "the lecturer explained clearly" item loads a little on the feedback factor too — and forcing those cross-loadings to zero inflates the correlations between factors, which is exactly what makes a scale look more unidimensional than it is. ESEM and bifactor-ESEM let the cross-loadings be small-but-nonzero and give a less distorted picture.

Why it matters for course evaluation in practice

The bifactor question is not academic. Whenever an institution reduces an evaluation to a single "overall" mean for a promotion file, a league table, or a threshold flag, it is asserting that one number faithfully represents the instrument. Bifactor indices let you check that assertion instead of assuming it.

  • If ω_h and ECV are high: the total score is empirically defensible. Report it, but you now have a warrant, not a habit.
  • If ω_h is high but a specific ω_hs is also substantial: report the overall score and flag that subscale as carrying unique, decision-relevant information (often "assessment and feedback," which reliably scores lowest).
  • If ω_h is low: a single average is misleading. The instrument measures several distinct things and should be reported as a profile, and personnel decisions built on the mean are on weak ground — echoing the misclassification risk that fair-ranking work has documented.

This complements rather than replaces reliability reporting. Coefficient omega tells you the total-score reliability; ω_h tells you how much of that reliability is the general factor rather than a blend of distinct dimensions. And where parallel analysis answers "how many factors are there," bifactor answers the follow-up "given that there are several, does one of them dominate enough to summarise?"

Limitations and honest caveats

A PhD reader will immediately raise the strongest objection: bifactor models tend to fit well even when they are the wrong model. Bonifay and Cai, and later Greene and colleagues, showed that the bifactor specification has high fitting propensity — it can absorb misfit and produce good fit statistics for data that were not generated by a bifactor process. Good fit is therefore weak evidence for bifactor structure. The defensible use is comparative and index-based (ω_h, ECV, H), not "the bifactor CFA fit, therefore one score is fine."

Other caveats:

  • Anomalous results are common: vanishing or negative group-factor loadings, or a general factor that is really a method artefact rather than substantive teaching quality.
  • Sample size and item count: stable bifactor estimation needs many respondents and several items per specific factor — a hard ask for a 12-item form in a class of 30. Aggregate across sections or terms, and mind that one class is rarely enough for stable estimates.
  • Orthogonality is an assumption: forcing specific factors to be uncorrelated with the general factor is a modelling choice, not a fact about teaching.
  • It does not fix bias: bifactor structure says nothing about whether the general factor is contaminated by grading leniency or charisma. Establish measurement invariance before comparing groups, whatever the dimensional structure.

Used honestly, bifactor analysis is a decision aid for reporting, not a certificate of validity.

How Koji incorporates this

Koji is built to produce the item-level, dimension-tagged data these models require, and to report at the altitude the evidence supports.

  • Structured, dimension-mapped items: Koji studies use typed questions (scale, single_choice, multiple_choice, ranking, open_ended, yes_no) that can be tagged to intended dimensions — clarity, workload, assessment, inclusion — so the exported matrix is ready for a bifactor or ESEM CFA rather than an undifferentiated blob.
  • Full item-level export: raw per-respondent, per-item data exports let an institutional-research team run ω_h, ω_hs and ECV in their own tooling; Koji does not force a pre-averaged total that would pre-empt the very question bifactor answers.
  • Report at the right altitude: when the general factor dominates, Koji dashboards can surface a defensible overall indicator; when specific factors carry unique signal, its thematic and dimension-level views keep those distinct instead of drowning them in a mean.
  • Conversational probing for the specific factors: the AI-moderated interview follows a low overall number into the specific dimension that caused it, generating the qualitative evidence that a bare subscale score only points at. Koji's core research platform at koji.so applies the same engine to product and customer research, where the general-vs-specific question recurs as "one satisfaction score or a driver profile?"

Koji is designed to mitigate the habit of over-summarising, not to certify any particular factor structure — the model still has to be fit and defended on your own data.

Frequently asked questions

What is the difference between a bifactor model and a higher-order factor model?

A higher-order model explains the correlations among first-order factors with a second-order factor, so the general factor only influences items indirectly. A bifactor model lets the general factor load on every item directly, alongside the specific factors. The bifactor is more flexible and yields ω_h and ECV directly; the higher-order model is a constrained special case of it.

What thresholds of omega-hierarchical and ECV justify a single total score?

Common rules of thumb from Rodriguez, Reise and Haviland are ω_h above about 0.75 and ECV above about 0.70, often alongside PUC, to treat a scale as "essentially unidimensional" for practical scoring. Treat these as guides, not hard cutoffs — report the values and let readers judge.

Is a bifactor model the same as saying the scale is unidimensional?

No. It explicitly models multidimensionality; it just quantifies how much one general dimension dominates. A high ω_h means a total score is interpretable, not that the specific factors are absent.

Why not just run a normal confirmatory factor analysis?

Ordinary CFA forces all cross-loadings to zero, which inflates inter-factor correlations and can make a scale look more unidimensional than it is. Bifactor and ESEM relax that and separate general from specific variance, so they answer the reporting question more directly.

Do I have enough data to fit one on a single class?

Usually not. Bifactor estimation needs many respondents and several items per specific factor; a small class with a short form will not support stable estimates. Pool across sections or terms, and treat single-class results with caution.

Does a good bifactor fit prove the instrument is valid?

No. Bifactor models fit well even for data not generated by a bifactor process, so good fit is weak evidence. Validity still depends on invariance, bias checks, and how the scores are used.

Related resources

References

  • Reise, S. P. (2012). The rediscovery of bifactor measurement models. Multivariate Behavioral Research, 47(5), 667-696. doi:10.1080/00273171.2012.715555
  • Reise, S. P., Moore, T. M., & Haviland, M. G. (2010). Bifactor models and rotations: Exploring the extent to which multidimensional data yield univocal scale scores. Journal of Personality Assessment, 92(6), 544-559. doi:10.1080/00223891.2010.496477
  • Rodriguez, A., Reise, S. P., & Haviland, M. G. (2016). Evaluating bifactor models: Calculating and interpreting statistical indices. Psychological Methods, 21(2), 137-150. doi:10.1037/met0000045
  • Rodriguez, A., Reise, S. P., & Haviland, M. G. (2016). Applying bifactor statistical indices in the evaluation of psychological measures. Journal of Personality Assessment, 98(3), 223-237. doi:10.1080/00223891.2015.1089249
  • Marsh, H. W., Morin, A. J. S., Parker, P. D., & Kaur, G. (2014). Exploratory structural equation modeling: An integration of the best features of exploratory and confirmatory factor analysis. Annual Review of Clinical Psychology, 10, 85-110. doi:10.1146/annurev-clinpsy-032813-153700
  • Bonifay, W., & Cai, L. (2017). On the complexity of item response theory models. Multivariate Behavioral Research, 52(4), 465-484. doi:10.1080/00273171.2017.1309262

Related articles

analysis-reporting

Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough

Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.

analysis-reporting

Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?

Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.

analysis-reporting

Coefficient Omega vs Cronbach's Alpha: Reporting the Reliability of a Course-Evaluation Scale

Cronbach's alpha assumes every item measures the construct equally well — an assumption course-evaluation subscales rarely meet. Here is why McDonald's omega is the more defensible reliability coefficient, and how to report it.

analysis-reporting

How Many Dimensions Does Your Evaluation Really Measure? Parallel Analysis for Factor Retention

Deciding how many factors your evaluation instrument measures with the eigenvalue-greater-than-one rule or the scree plot routinely gets the wrong answer. Parallel analysis (Horn, 1965) compares your data against random noise — and is one of the most accurate methods available.