Skip to main content
v2026.11,610 entries · CC-BY 4.0

Structural Equation Modeling (SEM): Fit Indices, Cut-Offs and Model Evaluation

A judgment-layer guide to evaluating structural equation models: which fit indices to report, why the widely cited Hu and Bentler cut-offs differ from what that paper actually recommended, modification-index discipline, and how to justify sample size.

Ask about Structural Equation Modeling (SEM): Fit Indices, Cut-Offs and Model Evaluation

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

Structural equation modeling (SEM) combines a measurement model — latent constructs and their indicators — with a structural model of directed relationships among those constructs, and evaluates both simultaneously against an observed covariance matrix. This guide is not an introduction to specifying one. It is about the decision that actually consumes reviewer attention: deciding whether your model fits, defending that decision, and knowing which parts of it the literature does not actually support.

The short version, which the rest of this page documents: the fit-index cut-offs that nearly every methods section cites as Hu & Bentler (1999) are not what Hu and Bentler recommended, the rule they did recommend is a different kind of rule from the one in common use, and the authors said in their own conclusion that a fixed cut-off per index is not defensible. None of that means fit indices are useless. It means reporting them as a pass/fail gate is a claim the source literature does not support.

Disambiguation. In this guide SEM means structural equation modeling. In measurement and psychometrics the same acronym means standard error of measurement, an entirely different quantity used to compute change thresholds — see our guide to minimal detectable change (MDC), which is built on that other SEM. The two share nothing but three letters.

What a fit index is actually testing

Every global fit index in common use is a function of the discrepancy between the observed covariance matrix and the covariance matrix implied by your fitted model. That single fact carries two consequences most tutorials skip.

First, fit is a statement about the whole model, not about any path in it. A structural equation model contains a measurement component and a structural component. A global index collapses misfit from both into one number. A model can therefore return an excellent CFI while the specific directed path your hypothesis is about is misspecified, non-significant, or in the wrong direction. Good global fit is a necessary condition for interpreting the structural estimates; it is not evidence for any of them.

Second, different indices are sensitive to different kinds of misspecification — and this is the part that matters most for SEM specifically, as opposed to a pure confirmatory factor analysis. Hu and Bentler reported that the ML-based SRMR is the index most sensitive to models with misspecified factor covariances or latent structure, while TLI, CFI, RNI and RMSEA are the indices most sensitive to misspecified factor loadings. In SEM terms: CFI and RMSEA are largely reporting on your measurement model; SRMR is the index carrying most of the signal about the relations among your constructs.

The practical implication is uncomfortable for a very common habit. Reporting CFI together with RMSEA — the pairing seen constantly in published methods sections — pairs two indices from the same sensitivity cluster. Hu and Bentler’s own factor analysis of index behaviour grouped TLI, BL89, RNI, CFI, Mc and RMSEA together as highly intercorrelated, with SRMR behaving least like the others. A CFI-plus-RMSEA pair is close to reporting the same information twice, and it is the pair least able to detect an omitted relationship between constructs.

The cut-off controversy

What is usually cited

The near-universal convention, attributed to Hu, L., & Bentler, P. M. (1999), Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives, Structural Equation Modeling, 6(1), 1–55, is stated as a three-part gate:

  • CFI (and TLI) ≥ .95
  • RMSEA ≤ .06
  • SRMR ≤ .08

Models clearing all three are declared to fit; models failing any one are declared not to. This is the version that appears in reviewer comments, in supervisor guidance, and in the majority of ranking tutorial pages.

What the 1999 paper actually recommends

Three specific discrepancies are worth knowing before you cite the paper.

1. Those three values are the single-index cut-offs, and the paper explicitly argues against using single-index cut-offs. The values close to .95 (CFI/TLI), .06 (RMSEA) and .08 (SRMR) appear in the paper as the best-performing thresholds when only one index is used. The paper’s actual recommendation is the two-index presentation strategy, in which SRMR is always reported and supplemented by one index from the other cluster.

2. The recommended combinational cut-offs are different numbers. In their conclusion, Hu and Bentler recommend “a cutoff value close to .95 for TLI (BL89, RNI, CFI, or Gamma Hat) in combination with a cutoff value close to .09 for SRMR”, and report that a CFI/TLI value of .96 combined with SRMR above .09 or .10 produced the least sum of Type I and Type II error rates. For the RMSEA pairing, the best-performing rule was RMSEA > .06 combined with SRMR > .09 (or .10). The widely cited SRMR ≤ .08 is not the value the paper’s own preferred rules use.

3. The combinational rule is a conjunction for rejection, not a set of hurdles to clear. The rules are evaluated in the paper in the form “reject if SRMR > .08 and RMSEA > .05″, “reject if SRMR > .08 and CFI < .95″, and so on. A model is flagged when both indices indicate poor fit. The common practice of requiring a model to simultaneously satisfy CFI ≥ .95 and RMSEA ≤ .06 and SRMR ≤ .08 is a materially stricter and structurally different decision rule from the one that was actually validated.

The authors’ own caveat

The conclusion section states the limitation directly:

“Although it is difficult to designate a specific cutoff value for each fit index because it does not work equally well with various conditions, a cutoff value close to .95 for the ML-based TLI, BL89, CFI, RNI, and Gamma Hat … seem to result in lower Type II error rates (with acceptable costs of Type I error rates).” — Hu & Bentler (1999), Conclusion and Recommendation

They further note that when N < 250 there is a trade-off between Type I and Type II error rates for every recommended combinational rule, that the appropriate rule therefore depends on which error is less acceptable in a given research area, and that under non-normality the Satorra–Bentler scaled test statistic should be used alongside the rules. “Close to” and “depends on your conditions” is not the language of a threshold table.

How often the conventional cut-off actually catches a misspecified model

The paper’s simulation results are the strongest argument against pass/fail use, and they are rarely quoted. Hu and Bentler distinguished two misspecification types: in the “simple” condition a true non-zero factor covariance was fixed to zero — the analogue of an omitted relationship between constructs — and in the “complex” condition a true cross-loading was omitted, a measurement-model error.

Against the omitted-factor-covariance condition, using CFI at .95:

  • At N < 500, only about 29.8% to 71.5% of misspecified models were rejected.
  • At N ≥ 1,000, only 4.6% to 57.6% were rejected.

Read that second line carefully: with a large sample, a CFI ≥ .95 gate can miss more than 95% of models in which a real relationship between constructs has been omitted — and it gets worse as N grows, not better. SRMR at .08, by contrast, rejected 99.9% to 100% of those same models, while catching only 0.9% to 66.5% of the omitted-cross-loading condition. Neither index is good at both jobs. That asymmetry is the entire reason the two-index strategy exists, and it is invisible if the cut-offs are treated as an interchangeable checklist.

The critiques that followed

Three lines of subsequent work are worth citing by name, because a reviewer who knows the literature will expect you to know them.

  • Marsh, H. W., Hau, K.-T., & Wen, Z. (2004). In search of golden rules: Comment on hypothesis-testing approaches to setting cutoff values for fit indexes and dangers in overgeneralizing Hu and Bentler’s (1999) findings. Structural Equation Modeling, 11(3), 320–341. The core argument is that the cut-offs were derived under a specific simulation design and generalized far past it, and that a hypothesis-testing rationale suited to significance testing was imported into approximate-fit assessment where it does not belong.
  • Nye, C. D., & Drasgow, F. (2011). Assessing goodness of fit: Simple rules of thumb simply do not work. Organizational Research Methods, 14(3), 548–570. Demonstrates that index values depend on properties of the model and data — number of indicators, communality, sample size — that have nothing to do with whether the model is correct.
  • McNeish, D., & Wolf, M. G. (2023). Dynamic fit index cutoffs for confirmatory factor analysis models. Psychological Methods, 28(1), 61–88 (doi:10.1037/met0000425). Rather than argue about which fixed number is right, this work simulates cut-offs for your specific model — its size, loadings and sample size — and derives the value that separates correctly specified from misspecified versions of it at controlled error rates. Implemented in the dynamic R package (Wolf & McNeish, 2023, Multivariate Behavioral Research, 58(1), 189–194) and at dynamicfit.app.

Dynamic fit indices are the most consequential development here, and the honest position on them is that they are well founded, increasingly expected in psychometrics journals, and not yet universal — the published work to date targets factor-analytic and related covariance structure models rather than every arbitrary structural model. If your model is within their scope, running them and reporting both the dynamic and conventional cut-offs is a stronger methods section than either alone.

What to report — and how to justify it

A defensible reporting set is not a threshold table. It is a set of quantities plus a stated decision rule you commit to in advance.

Report Why How to state the standard
χ² with df and p The only exact test of the model. Report it even though it will reject at large N — omitting it looks like concealment. Report the value; do not use it alone as the decision.
SRMR The index most sensitive to misspecified relations among constructs — the part of an SEM your hypotheses live in. State the cut-off you used and cite the rule it comes from.
CFI (and TLI) Incremental fit; sensitive primarily to measurement-model misspecification. Pair with SRMR, not with RMSEA alone.
RMSEA with its 90% CI The interval is the informative part; a point estimate hides the precision. Browne and Cudeck’s ranges (<.05 close, .05–.08 fair, >.10 poor) and MacCallum, Browne and Sugawara’s .08–.10 “mediocre” band are the conventional anchors. Report the interval, not just the point value.
Estimator and how non-normality was handled Cut-off behaviour is estimator-dependent; the 1999 values are ML-based. Name MLR / Satorra–Bentler / WLSMV explicitly.
Dynamic cut-offs, if applicable Model-specific rather than borrowed. Report alongside conventional values.

The framing that survives review is “we adopted rule X, specified before estimation, for reason Y” — not “CFI = .953, therefore the model fits.” Fit indices support an argument; they do not replace one.

Global fit does not license your structural paths

This is the failure mode most specific to SEM and the reason a CFA-level treatment of fit is not sufficient here.

Because global indices are dominated by measurement-model information, a model with a strong, clean measurement structure — high loadings, well-behaved indicators — can absorb a badly specified structural component and still return attractive fit statistics. The two-step approach associated with Anderson and Gerbing (1988, Psychological Bulletin) exists precisely to separate these: establish and evaluate the measurement model first, then impose the structural model, and treat the change in fit between them as the evidence about the structural specification. A single fit statistic from the combined model cannot make that distinction for you.

Three checks that do bear on the structural model, none of which are global fit indices:

  • The measurement-to-structural fit comparison. Fit that degrades materially when the structural constraints are imposed is direct evidence about those constraints.
  • Local fit. Standardized residuals and the pattern of large modification indices tell you where misfit sits. A model can meet every global cut-off while carrying a systematically misfit block of residuals.
  • Equivalent models. For most structural models there exist alternative specifications — including ones that reverse causal direction — that reproduce the same covariance matrix and therefore produce identical fit statistics. Fit cannot adjudicate among them. Only design, temporal ordering and theory can, and a reviewer is entitled to ask which equivalent models you considered.

Modification indices: the discipline

A modification index estimates how much χ² would fall if a currently fixed parameter were freed. Freeing parameters because the software suggested them is the fastest route to a model that fits this sample and replicates in no other. Our CFA guide covers the mechanics; the SEM-specific rules are:

  1. Theory first, index second. A freed parameter needs a substantive reason that would have justified specifying it in advance. “The MI was large” is not a reason.
  2. Never free a structural path on an MI. Correlated errors are a measurement decision. Adding a directed path between constructs because an index suggested it converts a confirmatory test into an exploratory search and should be reported as such.
  3. Report every modification. Report the original model’s fit as well as the modified model’s, state what was changed and why, and label the result exploratory.
  4. Cross-validate. A respecified model’s fit in the sample that generated the respecification is not evidence. A holdout sample, or preregistration of the original specification, is.

This is a research-integrity issue as much as a statistical one — the same category as p-hacking. Preregistering the model specification (see preregistration of a study protocol) is the cleanest defence available.

Sample size

Sample-size guidance for SEM is the second area where conventions are repeated with more confidence than the evidence supports. Being explicit about the status of each rule:

  • Fixed minimums (“N = 200”, “N = 100”). Conventions with no general justification. Required N depends on model size, loading strength, indicator count and distribution, and varies enormously across models. Treat these as folklore.
  • The N:q ratio — a ratio of observations to estimated parameters, commonly quoted at 10:1 or 20:1 and associated with Bentler and Chou (1987) and with Jackson (2003). This at least scales with model complexity, which fixed minimums do not, but empirical support for any particular ratio is mixed. Useful as a sanity check, not as a justification.
  • RMSEA-based power analysis (MacCallum, Browne & Sugawara, 1996, Psychological Methods, 1(2), 130–149) — computes power for tests of close fit and not-close fit given df and target RMSEA values. A real, model-aware calculation, and the most commonly accepted formal justification.
  • Monte Carlo power analysis for your specific model — simulate data from your hypothesised parameter values and determine the N at which the structural parameters you care about are recovered with adequate power and acceptable bias. This is the strongest justification, and the only one that targets the estimates your conclusions rest on rather than global fit.

Note the mismatch the first three share: they justify a sample size for detecting model misfit, while the claims in the paper usually concern the size and significance of specific paths. Adequate power for a close-fit test is not adequate power for a particular structural coefficient. See our guides to power analysis and sample size calculation and statistical power analysis for the general framework.

Small samples also degrade the fit indices themselves. Hu and Bentler’s simulations spanned N of 150, 250, 500, 1,000, 2,500 and 5,000, and they flagged that below N = 250 every recommended combinational rule involves a Type I / Type II trade-off, with rules based on RMSEA or TLI less preferable at those sizes. A cut-off borrowed for an N = 120 model is being used outside the conditions that produced it.

What reviewers actually challenge

  1. Cut-offs cited without their source’s conditions. Citing Hu and Bentler for a WLSMV-estimated model with ordinal indicators, when the values were derived for ML with continuous ones.
  2. A missing or unexplained χ². Reporting approximate indices only reads as selective.
  3. RMSEA without a confidence interval.
  4. Undisclosed respecification. A model that fits suspiciously well with a correlated-error structure that has no theoretical rationale.
  5. No equivalent-model discussion in a cross-sectional model making directional claims.
  6. Sample size justified by a rule of thumb rather than by a power analysis.
  7. Fit reported, structural estimates under-reported. Standardized coefficients, standard errors and interval estimates for the paths are the actual result; fit is the precondition.

A pre-submission checklist

  1. Specification, including every constraint, fixed before seeing the data — ideally preregistered.
  2. Sample size justified by a power analysis targeting the structural parameters.
  3. Measurement model established and evaluated before the structural model is imposed.
  4. Estimator named, with the non-normality or ordinality treatment stated.
  5. χ², df, p reported.
  6. SRMR reported, paired with CFI/TLI; RMSEA reported with its 90% CI.
  7. Decision rule stated in advance, with the citation it comes from and an acknowledgement that the cut-offs are conventions.
  8. Dynamic fit index cut-offs computed where the model is within their scope.
  9. Local fit examined: standardized residuals and modification-index pattern.
  10. Every respecification disclosed and labelled exploratory.
  11. Equivalent models acknowledged.
  12. Full structural estimates reported with uncertainty, not just fit statistics.

Frequently asked questions

What are acceptable fit indices for SEM?

The conventional answer is CFI/TLI ≥ .95, RMSEA ≤ .06 and SRMR ≤ .08, attributed to Hu and Bentler (1999). The accurate answer is that those are the paper’s single-index values, that its recommended combinational rules use SRMR close to .09 with CFI/TLI close to .95 or .96, and that the paper’s own conclusion states a specific cut-off per index cannot be designated because performance varies by condition. Report the indices, state the rule you adopted and why, and do not present the values as a pass/fail test.

Is CFI = .93 a failing model?

Not automatically, and treating it as such is the error the critique literature is about. CFI depends on the baseline model, indicator count and loading magnitude, so the same value means different things across models. A CFI of .93 with clean local fit and a theoretically specified model is more defensible than a CFI of .96 reached by freeing four unjustified correlated errors. Explain the value; do not launder it.

Why does my chi-square reject when every other index looks fine?

χ² tests exact fit — the hypothesis that your model reproduces the population covariance matrix perfectly — and its power grows with N, so at large samples it detects trivially small discrepancies. This is expected, not a defect. Report it, explain it, and rely on approximate indices and local fit for the substantive judgment.

Should I report CFI and RMSEA together?

It is the most common pairing and the weakest one. Hu and Bentler found CFI, TLI, RNI and RMSEA to be highly intercorrelated and primarily sensitive to misspecified loadings, while SRMR is the index most sensitive to misspecified factor covariances and latent structure. Their two-index strategy pairs SRMR with one index from the other cluster. If you report only two, make SRMR one of them.

What are dynamic fit indices and should I use them?

Dynamic fit index cut-offs (McNeish & Wolf, 2023) are simulated for your specific model and sample rather than borrowed from a 1999 simulation of different models. Where your model is within the scope of the published methods and available tooling, computing them and reporting them alongside the conventional values is a clear improvement. They have not replaced the conventions in general practice, so report both.

How is SEM different from confirmatory factor analysis?

A CFA is the measurement half of an SEM: latent factors, their indicators, and covariances among factors, with no directed paths. A full SEM adds directed structural relationships among the latent variables. Every fit consideration in our CFA guide applies, plus the structural-model concerns on this page — most importantly that global indices carry limited information about the directed paths.

Can fit indices tell me my causal model is correct?

No. For most structural models, equivalent models exist that fit identically, including specifications that reverse the direction of a path. Good fit establishes that your model is consistent with the covariance structure; it does not establish that it is the only model that is, or that the directions are right. See correlation vs causation.

Related CASRAI guides

Verification note: the Hu and Bentler (1999) figures, quotations and recommended combinational rules on this page were read directly from the published article rather than from secondary summaries, because the secondary summaries are where the discrepancy described above originates. Where guidance on this page is convention rather than a finding from a source we verified — notably the N:q sample-size ratios — it is labelled as such in the text.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →