What is the difference between exploratory and confirmatory factor analysis?

Factor analysis is central to questionnaire-based research. It helps you check whether your items measure the constructs you intend them to measure. But there are two main types, exploratory factor analysis (EFA) and confirmatory factor analysis (CFA), and students often wonder which one they should use, whether they need both, and how they differ in practice.

Choosing the wrong one, or using them in the wrong way, is a common source of examiner comments. This post explains the key differences between EFA and CFA, when to use each, how they can be combined, and how to report them correctly.

The big picture

Both EFA and CFA are based on the idea that correlations among observed items can be explained by a smaller number of latent (unobserved) factors. The difference lies in how much you specify in advance.

  • EFA is data-driven. You do not specify which items load on which factor. The analysis explores the data and suggests a structure.
  • CFA is theory-driven. You specify in advance how many factors there are, which items load on each factor, and whether factors are correlated. The analysis tests whether your data fit that model.

A simple way to remember: EFA asks, “What structure is in my data?” CFA asks, “Does my data fit the structure I expect?”

Exploratory factor analysis (EFA)

Purpose

EFA is used to discover the underlying dimensions of a set of items when the structure is unknown or uncertain. It is common in the early stages of scale development and when using a scale in a very new context.

How it works

  • Every item is allowed to load on every factor.
  • You decide how many factors to retain using criteria such as parallel analysis, scree plots and interpretability.
  • Rotation (such as Promax or Oblimin) is used to make the factor structure easier to interpret.
  • You examine factor loadings, cross-loadings and communalities, and may remove problematic items.

Typical software

SPSS (Analyze → Dimension Reduction → Factor), R (psych package), Jamovi, JASP and others.

Output

A pattern of loadings showing which items relate most strongly to which factors, eigenvalues, variance explained and factor correlations.

Confirmatory factor analysis (CFA)

Purpose

CFA tests whether a hypothesised measurement model fits the data. It is used when you have a clear theoretical or empirical basis for the factor structure, such as when using an established scale or confirming a structure identified by EFA in a previous sample.

How it works

  • You specify which items load on which factor. Items normally load only on their intended factor; other loadings are fixed to zero.
  • You specify whether factors are allowed to correlate.
  • The software estimates the model and produces fit indices that indicate how well the model reproduces the observed data.
  • You can compare alternative models, such as a one-factor model versus a three-factor model.

Typical software

AMOS, Mplus, the lavaan package in R, LISREL, Stata and JASP.

Output

Standardised factor loadings, model fit indices, factor correlations, and information for assessing reliability and validity.

Key differences at a glance

  • Starting point: EFA starts with no fixed structure; CFA starts with a specified model.
  • Role of theory: EFA is guided by theory in interpretation; CFA is guided by theory in model specification.
  • Item loadings: In EFA, all items can load on all factors; in CFA, items load only on specified factors.
  • Model fit: EFA does not usually focus on overall fit (though some methods provide it); CFA is evaluated primarily through fit indices.
  • Typical use: EFA for exploring and developing; CFA for testing and confirming.
  • Software: EFA is available in general statistical packages; CFA usually requires structural equation modelling software.

Assessing CFA model fit

CFA uses several fit indices, each capturing different aspects of fit. Commonly reported indices include:

  • Chi-square (χ²): tests exact fit. A non-significant result suggests good fit, but it is very sensitive to sample size and is almost always significant in large samples. Report it, but do not rely on it alone.
  • Comparative Fit Index (CFI) and Tucker–Lewis Index (TLI): compare your model with a baseline model. Values close to .95 or above are often treated as good fit, following Hu and Bentler (1999), while values above .90 are sometimes described as acceptable.
  • Root Mean Square Error of Approximation (RMSEA): values of about .06 or below are often considered good, and up to .08 acceptable. Report its 90% confidence interval.
  • Standardised Root Mean Square Residual (SRMR): values of about .08 or below are often considered good.

These cut-offs are widely used but also widely debated; they were derived from specific simulation conditions and should not be applied as rigid rules. Report several indices, cite your source for cut-offs, and interpret fit holistically.

Reliability and validity in CFA

CFA also provides information to assess construct reliability and validity, especially in management and business research:

  • Standardised loadings: ideally substantial, commonly .50 or above, with .70 or above often considered strong.
  • Composite reliability (CR): often expected to be .70 or above.
  • Average variance extracted (AVE): often expected to be .50 or above as evidence of convergent validity (Fornell and Larcker, 1981).
  • Discriminant validity: assessed using approaches such as the Fornell–Larcker criterion or the heterotrait–monotrait ratio (HTMT), which has been proposed as a more sensitive test.

Using EFA and CFA together

EFA and CFA are often complementary. A common sequence in scale development is:

  1. Develop items based on theory and expert input.
  2. Collect data from Sample 1 and run EFA to explore the structure.
  3. Refine the item set based on EFA results.
  4. Collect data from Sample 2 (or split a large sample randomly into two halves) and run CFA to test the structure identified in EFA.

An important warning

Running EFA and CFA on the same dataset and presenting the CFA as confirmation is problematic. The CFA will tend to fit well because the model was derived from those same data. This is sometimes called capitalising on chance. If you cannot collect a second sample, split your sample randomly (if it is large enough), or clearly acknowledge the limitation.

A worked example: from EFA to CFA

Consider a researcher developing a new 18-item scale of “academic belonging” for international postgraduate students. The process might look like this (all figures are illustrative):

  1. Sample 1 (n = 280), EFA: KMO = .89 and Bartlett's test is significant. Parallel analysis supports three factors. Using principal axis factoring with Promax rotation, the researcher identifies three interpretable factors: connection with peers, connection with academic staff, and sense of fit with the institution. Three items with weak loadings or strong cross-loadings are removed, leaving 15 items.
  2. Sample 2 (n = 320), CFA: The three-factor, 15-item model is specified in lavaan. Fit is acceptable (for example, CFI = .95, TLI = .94, RMSEA = .055 with 90% CI [.045, .066], SRMR = .046). All standardised loadings exceed .60. Composite reliability ranges from .82 to .88, and AVE values exceed .50 for each factor.
  3. Alternative models: A one-factor model and a two-factor model (combining peers and staff) are also tested and fit clearly worse, supporting the three-factor structure.
  4. Discriminant validity: HTMT ratios between factors are below commonly suggested thresholds, suggesting the three factors are distinct.

This sequence shows the complementary roles of EFA and CFA: the first sample generates the structure, and the second independently tests it. Testing competing models also strengthens the argument, because it shows the chosen structure is better than plausible alternatives, not just acceptable on its own.

Which should you use in your study?

Use EFA when:

  • You are developing a new scale.
  • The factor structure of an existing scale is unclear or inconsistent across studies.
  • You have adapted a scale substantially and are unsure whether the original structure holds.
  • You want to explore dimensionality before theory is well developed.

Use CFA when:

  • You use established scales with a clear, previously validated structure.
  • You want to test a specific measurement model before structural equation modelling.
  • You want to compare competing models, such as one-factor versus multi-factor structures.
  • You want to test measurement invariance across groups, such as men and women, or different countries.

Use both when:

  • You are developing and validating a new instrument and have enough data for independent samples.

What if CFA shows poor fit?

Poor fit means the specified model does not adequately reproduce the observed data. Possible steps include:

  • Checking for data problems, such as coding errors or outliers.
  • Inspecting standardised loadings for weak items.
  • Examining modification indices, which suggest changes that might improve fit, such as allowing correlated errors.

Be careful with modification indices. Only make changes that are theoretically justified, for example, correlated errors between items with very similar wording. Making many data-driven changes turns CFA into an exploratory exercise, and results may not replicate. Report all modifications and the reasons for them.

Reporting in your thesis

For EFA: sample suitability (KMO, Bartlett's test), extraction and rotation methods, criteria for the number of factors, loading thresholds, item removal decisions, final loadings table, variance explained and factor correlations.

For CFA: the specified model (often with a diagram), estimation method, sample size, fit indices with cut-off sources, standardised loadings, factor correlations, reliability and validity statistics, and any model modifications with justification.

Common mistakes to avoid

  • Using EFA on well-established scales without explanation when CFA would be more appropriate.
  • Running EFA and CFA on the same data and calling the CFA confirmation.
  • Reporting only chi-square or only one fit index.
  • Treating fit cut-offs as absolute rules.
  • Making many modification-index-based changes without theory.
  • Calling principal component analysis “exploratory factor analysis”.

Final thoughts

EFA and CFA are two sides of the same coin. EFA helps you discover the structure of your measures when it is uncertain; CFA tests whether an expected structure fits your data. Choose based on how much you already know about your measures, use independent samples when combining them, and report your decisions and results transparently. With the right approach, factor analysis gives strong support for the validity of your measurement and, in turn, the credibility of your whole study.

Not sure whether your study needs EFA, CFA or both? Describe your scales in the comments and we can help you decide.