Analysis Types

Factor Analysis – Definition, Types, Steps and Examples

Table of Contents

Factor analysis is a family of multivariate statistical methods used to explain correlations among observed variables through a smaller number of unobserved factors. Researchers use it to identify latent constructs, develop or validate measurement scales, reduce redundant variables, and test whether questionnaire items represent theoretically expected dimensions.

Factor Analysis

Factor analysis is widely used in psychology, education, healthcare, business, sociology, marketing and other fields in which important concepts cannot be measured directly. Constructs such as anxiety, job satisfaction, academic engagement and perceived service quality are usually represented by multiple questions or indicators. Factor analysis helps researchers examine whether these indicators reflect one or more underlying dimensions.

This guide explains what factor analysis is, how exploratory and confirmatory approaches differ, what assumptions must be considered, how an analysis is conducted, and how the results should be interpreted and reported.

Key Takeaways

  • Factor analysis examines whether correlations among measured variables can be explained by a smaller number of latent factors.
  • Exploratory factor analysis discovers a plausible structure, while confirmatory factor analysis evaluates a prespecified measurement model.
  • Factor analysis and principal component analysis are related but answer different questions.
  • Extraction, factor retention, rotation and correlation type should be chosen according to the research question and data.
  • Statistical criteria must be combined with theory, interpretability and evidence of solution stability.
  • An exploratory solution should ideally be evaluated in new data before it is described as confirmed.

What Is Factor Analysis?

Factor analysis is a statistical method for identifying latent variables that may account for patterns of correlation among observed variables.

An observed variable is directly measured, such as a questionnaire response, test score or clinical rating. A latent factor is not measured directly. Instead, it is inferred from relationships among multiple observed variables.

Suppose a university survey contains questions about:

  • Participating in class discussions
  • Completing optional exercises
  • Concentrating during lectures
  • Feeling connected to the course
  • Enjoying academic challenges
  • Planning study time

Responses to these questions may reflect a smaller set of constructs, such as behavioural engagement, emotional engagement and study regulation. Factor analysis helps determine whether such a structure is supported by the correlation pattern.

A factor is therefore not simply a group created by the software. It is a statistical dimension that must also be interpreted using item content, theory and previous research.

How Does Factor Analysis Work?

Factor analysis starts with a correlation or covariance matrix. If several variables are correlated because they reflect the same underlying construct, the method attempts to represent their shared variance through one or more factors.

A simplified common-factor model can be written as:

x = Λf + ε

Where:

  • x represents the observed variables.
  • Λ is the matrix of factor loadings.
  • f represents the latent factors.
  • ε represents unique variance and measurement error.

A factor loading describes the relationship between an observed variable and a factor. Larger absolute loadings indicate a stronger relationship, but interpretation should not depend on one universal cutoff.

The model separates two broad sources of variance:

  1. Common variance: Variance shared with other variables and potentially explained by common factors.
  2. Unique variance: Variance specific to an indicator, including specific influences and measurement error.

For a standardized variable, its communality estimates the proportion of variance explained by the retained common factors. Its uniqueness is the remaining variance not explained by those factors.

Important Factor-Analysis Terms

TermMeaning
Observed variableA directly measured item, score or indicator
Latent factorAn unobserved construct inferred from several indicators
Factor loadingThe strength and direction of the relationship between an indicator and a factor
CommunalityThe proportion of an indicator’s variance explained by the retained common factors
UniquenessVariance not explained by the common factors
EigenvalueA quantity related to how much variance a dimension represents
ExtractionThe method used to estimate an initial factor solution
RotationA transformation used to make the loading pattern easier to interpret
Cross-loadingA substantial relationship between one indicator and more than one factor
Factor correlationThe estimated relationship between two obliquely rotated factors
Factor scoreAn estimated score representing a participant’s position on a factor
ResidualThe difference between an observed relationship and the relationship reproduced by the model

Why Is Factor Analysis Used?

Factor analysis has several related purposes.

Identifying latent constructs

Researchers can use correlated indicators to investigate concepts that cannot be observed directly. For example, several questions about worry, tension and difficulty relaxing may reflect a latent anxiety factor.

Developing measurement instruments

During scale development, exploratory factor analysis can show whether proposed items form coherent dimensions and whether some items load weakly or ambiguously.

Evaluating construct validity

Factor analysis contributes evidence about internal structure. It can help researchers evaluate whether an instrument’s dimensional structure is consistent with its theoretical interpretation.

Factor analysis alone does not establish every form of validity. Content validity, relationships with external variables, response processes and consequences of score use may also require investigation.

Reducing redundancy

A large set of correlated variables may contain overlapping information. A smaller factor representation can make the structure easier to understand, although PCA may be more appropriate when the sole objective is efficient data compression.

Testing a measurement model

Confirmatory factor analysis evaluates whether the observed covariance structure is consistent with a model specified before examining the results.

When Should Factor Analysis Be Used?

Factor analysis is suitable when:

  • The research concerns one or more latent constructs.
  • Several observed variables are expected to reflect those constructs.
  • Variables show meaningful correlations.
  • The sample and indicators provide sufficient information for stable estimation.
  • The proposed model is substantively interpretable.
  • The measurement level and estimator are compatible.

Factor analysis may not be appropriate when:

  • Variables have little meaningful correlation.
  • There are too few indicators to define the proposed factors.
  • Items are formative causes rather than reflective indicators of a construct.
  • The sample is too small or unrepresentative for the intended model.
  • The only goal is to create uncorrelated mathematical summaries of total variance.
  • Variables consist of unrelated demographic or categorical categories without an appropriate latent-variable model.

Types of Factor Analysis

Types of Factor Analysis

Exploratory Factor Analysis

Exploratory factor analysis, or EFA, is used to investigate the number and nature of latent dimensions when the structure is uncertain or only partially specified.

In EFA, each observed variable may initially load on each factor. The researcher decides:

  • Which correlation matrix to analyse
  • Which extraction method to use
  • How many factors to retain
  • Which rotation to apply
  • How to interpret and label the factors
  • Whether the solution is stable and theoretically meaningful

EFA is not theory-free. Item construction, factor interpretation and model evaluation should still be informed by substantive knowledge.

Confirmatory Factor Analysis

Confirmatory factor analysis, or CFA, tests whether a prespecified factor model is compatible with the observed data.

Before estimation, the researcher specifies matters such as:

  • The number of factors
  • Which indicators load on each factor
  • Whether factors are correlated
  • Whether residual relationships are permitted
  • How the model is identified

CFA estimates the model and evaluates parameter estimates, residuals, global fit, local fit and theoretical plausibility. A poorly fitting model is evidence that the proposed structure, estimator, data assumptions or measurement design may need reconsideration.

EFA vs. CFA

FeatureEFACFA
Main purposeDiscover or explore a plausible structureEvaluate a prespecified structure
Starting pointUncertain or partially understood dimensionalityStrong theory or previous empirical evidence
Cross-loadingsGenerally estimated before rotationCommonly fixed to zero unless specified
Main outputRotated loading pattern and factor relationshipsModel parameters, residuals and fit evidence
Typical useEarly scale developmentValidation, replication and theory testing
Best follow-upReplication or CFA using new dataCross-validation, invariance testing or model refinement

Running CFA immediately after EFA on the same observations may describe the same sample well without demonstrating genuine confirmation. Where possible, researchers should use an independent validation sample, collect new data, or apply a defensible sample-splitting strategy.

Factor Analysis vs. Principal Component Analysis

Factor analysis and principal component analysis are often confused because both can produce loadings and reduce dimensionality.

The central difference is their objective.

Factor analysisPrincipal component analysis
Models common variance attributed to latent factorsCreates weighted combinations that summarize total observed variance
Includes unique or error variance in the modelDoes not separate common variance from unique variance in the same way
Used to investigate latent constructsUsed mainly for mathematical data reduction
Produces factorsProduces components
Requires a latent-variable interpretationDoes not require latent causes

Use factor analysis when the research question concerns underlying constructs that are believed to produce correlations among indicators. Use PCA when the principal objective is to compress a set of variables into fewer composite dimensions while retaining as much total variance as possible.

PCA may appear as an extraction option in software menus labelled “factor analysis,” but this does not make PCA and common-factor analysis conceptually identical.

Assumptions and Data Requirements

No single checklist applies equally to every extraction method or estimator. Researchers should evaluate the requirements of the specific model they intend to estimate.

Meaningful correlations

The variables should contain enough shared information to justify a latent-factor representation. Researchers should inspect the correlation matrix rather than relying only on significance tests.

Appropriate correlation type

Pearson correlations may be appropriate for reasonably continuous variables when their relationships are adequately represented linearly.

For ordered categorical items, particularly items with few response categories or strongly skewed distributions, polychoric correlations and ordinal estimators may be more appropriate. The decision should consider category counts, distributions, sample size and software capabilities.

Independence of observations

Standard factor models generally assume that observations are independent. Students nested within classes, patients nested within hospitals or repeated measurements from the same individual may require multilevel, longitudinal or otherwise dependent-data models.

Linearity

Traditional continuous-variable factor analysis represents relationships through a linear latent-variable model. Serious nonlinear relationships may not be represented adequately by an ordinary Pearson correlation matrix.

Absence of extreme singularity

Very high correlations or duplicate indicators can create unstable estimates. Variables that are almost exact combinations of other variables should be investigated.

Distributional requirements

Multivariate normality is especially relevant to conventional maximum-likelihood standard errors, tests and fit statistics. Some alternative estimators are less dependent on normality. The appropriate response to nonnormality is not automatically to abandon factor analysis; researchers can consider robust estimators, ordinal models or sensitivity analyses.

Adequately defined factors

A factor is more defensible when it is represented by multiple strong, conceptually coherent indicators. Factors defined by only one or two weak indicators are often unstable or unidentified unless additional constraints and strong justification are available.

KMO and Bartlett’s Test

The Kaiser–Meyer–Olkin measure evaluates whether the pattern of correlations is sufficiently compact for factor analysis. Higher values generally indicate that common factors may be estimated more effectively.

The Bartlett test of sphericity tests the null hypothesis that the population correlation matrix is an identity matrix. A statistically significant result suggests that the variables are not all uncorrelated.

Neither statistic should make the decision alone:

  • Bartlett’s test can become significant in large samples even when correlations are not substantively useful.
  • An acceptable overall KMO can hide problematic individual variables.
  • Factorability should also be assessed through correlations, anti-image information, communalities, residuals and the quality of the resulting solution.

How Large Should the Sample Be?

There is no universally correct minimum sample size for factor analysis. Sample requirements depend on factor loadings, communalities, number of indicators, number of factors, factor correlations, data distributions, missingness and the chosen estimator.

Simple participant-to-item ratios can be used for rough planning, but they should not be treated as laws. A solution with strong loadings, high communalities and several indicators per factor may be recoverable with fewer cases than a model containing weak loadings, low communalities and poorly defined factors.

Researchers should preferably:

  1. Base planning on the expected model.
  2. Review simulation evidence for similar conditions.
  3. Conduct a Monte Carlo power or parameter-recovery study for consequential CFA projects.
  4. Allow for missing data and exclusions.
  5. Use a larger sample when factors are weak, data are ordinal, or the model is complex.
  6. Reserve data for validation when EFA and CFA are both planned.

A sample of approximately 200 is often used as a practical planning reference in applied research, but it is neither automatically sufficient nor always necessary. The characteristics of the measurement model matter more than a single universal number.

How to Conduct Exploratory Factor Analysis

Step 1: Define the construct and purpose

State whether the objective is scale development, dimensionality assessment, construct exploration or another purpose. Explain why a latent-variable model is appropriate.

Step 2: Examine the indicators

Review item wording, response ranges, reverse-coded items, missing values, floor and ceiling effects, and descriptive distributions.

Statistical analysis cannot repair a poorly designed item pool. The indicators should adequately represent the theoretical content of the construct.

Step 3: Choose the correlation matrix

Possible choices include:

  • Pearson correlations for suitable continuous variables
  • Polychoric correlations for ordered categorical variables
  • Tetrachoric correlations for suitable binary indicators
  • Mixed correlations when variable types differ

The same measurement assumptions should be used consistently in factor retention and extraction where possible.

Step 4: Evaluate factorability

Inspect:

  • The correlation matrix
  • KMO measures
  • Bartlett’s test
  • Anti-image information
  • Initial communalities
  • Possible multicollinearity or singularity
  • Variables with almost no relationship to the rest of the item pool

Do not continue only because one significance test passes.

Step 5: Choose an extraction method

Common choices include:

Maximum likelihood: Useful when its distributional assumptions are sufficiently met and when inferential tests or likelihood-based comparisons are required.

Principal-axis factoring: Commonly used to estimate factors from shared variance and less dependent than conventional maximum likelihood on multivariate normality.

Minimum residual or unweighted least squares: Useful alternatives available in several modern programs.

Ordinal least-squares methods: Often considered for ordered categorical indicators.

PCA should be selected only when the research objective is component-based data reduction rather than common-factor modelling.

Step 6: Determine how many factors to retain

Use several forms of evidence:

  1. Theoretical expectations
  2. Parallel analysis
  3. Scree-plot shape
  4. Model fit where available
  5. Residual correlations
  6. Factor interpretability
  7. Number and strength of indicators per factor
  8. Stability across plausible analytical choices

The eigenvalue-greater-than-one rule should not be the only criterion. It can retain too many or too few dimensions under some conditions.

Step 7: Select a rotation

Rotation improves interpretability but does not change the underlying fit of a solution with a fixed number of factors.

Orthogonal rotation, such as Varimax, constrains factors to be uncorrelated.

Oblique rotation, such as Oblimin or Promax, allows factors to correlate.

Because constructs in social, behavioural and educational research are often related, oblique rotation is frequently more realistic. An orthogonal solution should be justified by theory or the intended use rather than selected automatically.

Step 8: Interpret the factor pattern

Interpret indicators using:

  • Loading magnitude
  • Loading direction
  • Cross-loadings
  • Communalities
  • Factor correlations
  • Item wording
  • Conceptual coherence
  • Residual relationships

After oblique rotation, distinguish the pattern matrix from the structure matrix. Pattern coefficients show an indicator’s unique relationship with a factor while accounting for other factors. Structure coefficients are correlations between indicators and factors and can be influenced by correlations among factors.

Step 9: Refine carefully

An item should not be deleted merely because it misses an arbitrary loading cutoff. Consider:

  • The item’s theoretical importance
  • Content coverage
  • Whether it is ambiguously worded
  • Its primary and secondary loadings
  • Its communality
  • Consequences for reliability and validity
  • Whether deleting it would narrow the construct
  • Whether the decision is stable in another sample

When items are removed, rerun the analysis and document each decision. Extensive data-driven deletion can overfit the instrument to one sample.

Step 10: Evaluate stability and validate

Where resources permit:

  • Repeat the analysis under reasonable extraction and rotation choices.
  • Bootstrap loadings or evaluate sampling variability.
  • Examine whether the solution appears in subgroups or resamples.
  • Conduct CFA in an independent sample.
  • Test measurement invariance when scores will be compared across groups or time.

Hypothetical Example of Factor Analysis

Imagine that researchers develop a nine-item student-engagement questionnaire. They expect three dimensions but initially conduct EFA because the items are new.

The following simplified loadings are hypothetical:

Questionnaire itemBehavioural engagementEmotional engagementStudy regulation
I participate in class activities.76.12.09
I complete optional learning tasks.71.14.21
I contribute to group discussions.68.18.06
I enjoy learning difficult material.11.78.16
I feel interested during lessons.20.73.13
I feel connected to my course.17.65.12
I plan when I will study.08.10.81
I monitor whether I understand.19.16.72
I change my strategy when necessary.22.12.67

This pattern suggests three interpretable factors:

  1. Behavioural engagement
  2. Emotional engagement
  3. Study regulation

The labels are based on the meaning of the items, not only the numerical output.

Before accepting the structure, the researchers should also examine communalities, secondary loadings, factor correlations, residuals and solution stability. They should then test the three-factor model using new data rather than claiming that the same exploratory sample has confirmed it.

How to Interpret Factor Loadings

A loading indicates how strongly an indicator is related to a factor under the estimated model.

Loadings near zero show little relationship. Larger positive or negative loadings show stronger relationships. The sign indicates direction, although factor signs can be reversed without changing the substantive solution if all associated signs are reversed consistently.

Rules such as .30, .40, .50 or .70 are sometimes used as descriptive reference points. They should not be treated as universal pass–fail standards. Interpretation depends on:

  • Sample size
  • Measurement quality
  • Research stage
  • Number of indicators
  • Communalities
  • Cross-loadings
  • Estimation uncertainty
  • The theoretical importance of the item

A loading of .45 may be useful in early exploratory work but insufficient for a high-stakes instrument. Conversely, deleting every item below .70 can remove important content and produce an artificially narrow construct.

How to Interpret Cross-Loadings

A cross-loading occurs when an indicator has notable relationships with multiple factors.

Cross-loadings may indicate:

  • Ambiguous wording
  • Overlapping constructs
  • A broad item that genuinely reflects several dimensions
  • A missing general or method factor
  • Too many or too few retained factors
  • An inappropriate rotation
  • Poor construct differentiation

Researchers should not automatically assign the item to whichever factor has the slightly larger loading. The content and purpose of the item must be considered.

Confirmatory Factor Analysis in Practice

CFA usually involves the following process:

  1. Specify the model before examining modification suggestions.
  2. Select an estimator that matches the data.
  3. Identify the model appropriately.
  4. Inspect convergence and inadmissible estimates.
  5. Review standardized loadings and residual variances.
  6. Evaluate global and local fit.
  7. Compare theoretically meaningful alternatives where justified.
  8. Report any post-hoc modifications transparently.
  9. Cross-validate the final model.

Common CFA fit statistics include:

  • Model chi-square
  • Comparative Fit Index
  • Tucker–Lewis Index
  • Root Mean Square Error of Approximation
  • Standardized Root Mean Square Residual

Values such as CFI or TLI near .95, RMSEA near .06 and SRMR near .08 are often quoted from simulation-based recommendations. These values are not universal laws. Their behaviour varies with model size, loadings, estimator, degrees of freedom, distribution and type of misspecification.

Fit indices should therefore be interpreted together with parameter estimates, residuals, theory, data quality and the consequences of model error.

Modification indices can identify possible sources of misfit, but blindly following them can overfit a model. A residual covariance or cross-loading should be added only when it has a defensible substantive or methodological explanation and should ideally be tested in new data.

Factor Scores

Factor scores estimate each participant’s position on a latent factor. They may be used in later analyses, but they are estimates rather than perfectly observed values.

Common score-estimation methods include regression and Bartlett approaches. Scores can differ according to the extraction method, rotation, scoring algorithm and sample.

Before replacing original indicators with factor scores, researchers should consider:

  • Whether the score is sufficiently determinate
  • Whether a validated scale mean or sum would be easier to interpret
  • Whether uncertainty in estimated scores affects later analyses
  • Whether the scoring method will be applied consistently in new samples

Advantages of Factor Analysis

Factor analysis can:

  • Reveal patterns that are difficult to see in individual correlations.
  • Represent complex constructs more parsimoniously.
  • Support scale development and validation.
  • Identify redundant or poorly functioning indicators.
  • Separate shared variance from unique variance.
  • Test theoretically specified measurement structures.
  • Provide evidence about the internal structure of an instrument.

Limitations of Factor Analysis

Factor analysis also has important limitations:

  • Results depend on the variables included.
  • Different defensible analytical decisions can produce different solutions.
  • A factor is not automatically a real-world causal entity.
  • Weak samples or indicators produce unstable results.
  • Factor naming involves substantive judgement.
  • EFA can be overfit through repeated deletion and reanalysis.
  • CFA can be overfit through unjustified modification indices.
  • Good model fit does not prove that a model is uniquely correct.
  • Factor scores contain estimation error.
  • Cross-sectional factor analysis does not establish causality.

Common Factor-Analysis Mistakes

Treating PCA as EFA without explanation

Software defaults do not determine the correct method. The extraction approach should match the research objective.

Using only the eigenvalue-above-one rule

Factor retention should combine parallel analysis, theory, scree evidence, fit, interpretability and stability.

Selecting Varimax automatically

Varimax assumes uncorrelated factors. Related psychological, educational and social constructs often require an oblique rotation.

Using Pearson correlations for every Likert item

Ordered categories may require polychoric correlations or ordinal estimators, particularly when the number of categories is small or responses are skewed.

Applying one loading cutoff mechanically

A cutoff cannot replace evaluation of uncertainty, cross-loadings, communalities and content validity.

Deleting items only to improve statistics

Removing items can reduce conceptual coverage even when model fit improves.

Conducting EFA and CFA on the same data without disclosure

This does not provide strong independent confirmation and may exaggerate replicability.

Ignoring factor correlations

When oblique rotation is used, factor correlations are part of the substantive result.

Reporting only the rotated matrix

A reproducible report should state the data, correlation type, estimator, retention criteria, rotation, missing-data method, software and version, and all important item decisions.

Factor Analysis in Modern Research

Modern factor analysis extends beyond a simple SPSS procedure.

Ordinal factor models

Researchers increasingly use polychoric correlations and weighted least-squares estimators for ordered categorical questionnaire items.

Exploratory structural equation modelling

ESEM combines exploratory loading structures with features of structural equation modelling. It can be useful when strict CFA assumptions that all secondary loadings equal zero are unrealistic.

Hierarchical and bifactor models

These models examine whether indicators reflect a broad general factor, narrower specific factors, or both. They require careful identification and interpretation; a well-fitting bifactor model does not automatically justify one total score.

Multilevel factor analysis

When respondents are nested within schools, organizations or regions, the factor structure may differ within and between groups. Multilevel models can separate these sources of covariance.

Measurement invariance

Invariance testing evaluates whether a measurement model functions comparably across groups, languages, cultures or time points. This is important before interpreting score differences as substantive differences in the construct.

Bayesian factor analysis

Bayesian approaches allow researchers to incorporate prior information, estimate complex models and express uncertainty through posterior distributions. Their validity depends on appropriate prior specification, model checking and transparent reporting.

Replicability and sensitivity analysis

Researchers increasingly examine whether conclusions change across plausible estimators, correlation types, retention criteria, rotations and samples. A factor solution that appears only under one narrow set of choices should be interpreted cautiously.

Factor-Analysis Software

SoftwareTypical uses
SPSSMenu-based EFA, extraction, rotation and factor scores
R psychEFA, polychoric correlations, parallel analysis and psychometric functions
R lavaanCFA, SEM, ordinal estimators, multigroup models and ESEM capabilities
JASPGraphical EFA and CFA with accessible output
jamoviGraphical EFA and CFA with reproducible analysis options
MplusAdvanced CFA, ESEM, categorical, multilevel and mixture models
StataFactor analysis, SEM and related post-estimation tools
PythonProbabilistic factor-analysis and machine-learning workflows, although additional tools may be needed for a complete psychometric analysis

The software name is not enough for reproducibility. Researchers should report the version, package, estimator, correlation type, rotation, retention procedure, missing-data method and any non-default settings.

How Artificial Intelligence Can Assist

Generative AI can support factor-analysis work by:

  • Explaining software output in simpler language
  • Drafting R, Python, SPSS or Stata syntax
  • Translating a model specification between programs
  • Creating a reporting checklist
  • Identifying missing documentation
  • Commenting code
  • Suggesting sensitivity analyses

AI should not independently decide:

  • Which construct is being measured
  • How many factors are theoretically meaningful
  • Which items should be deleted
  • Whether a post-hoc modification is defensible
  • Whether a model has established validity
  • Which references exist

AI-generated code must be tested against official documentation and validated using known data or independently reviewed output. References and DOIs should be checked individually because generative systems can produce nonexistent citations.

Researchers should not upload confidential participant data to an AI service unless institutional approval, informed-consent requirements, data-protection rules and the service’s privacy conditions permit it. De-identified or simulated data are safer for code development.

How to Report Exploratory Factor Analysis

A complete report should identify:

  • Purpose of the EFA
  • Sample and indicators
  • Missing-data treatment
  • Correlation matrix
  • Factorability evidence
  • Extraction method
  • Retention criteria
  • Rotation
  • Number of factors
  • Factor correlations
  • Loading and cross-loading criteria
  • Communalities
  • Item-removal decisions
  • Variance information where relevant
  • Software and version
  • Validation or sensitivity analyses

Adaptable reporting example

“An exploratory factor analysis was conducted on [number] items using data from [sample size] participants. Because the indicators were [continuous/ordinal], the analysis used a [Pearson/polychoric] correlation matrix and [extraction method] estimation. Factorability was evaluated using the correlation matrix, KMO measures and Bartlett’s test of sphericity. The number of factors was determined through parallel analysis, scree-plot inspection, model adequacy and theoretical interpretability. A [rotation] rotation was applied because the factors were expected to be [correlated/uncorrelated]. The retained [number]-factor solution was interpreted using pattern coefficients, communalities, cross-loadings, factor correlations and item content. All item exclusions and repeated analyses are reported in [table or supplement].”

Numerical findings should then be added, including confidence intervals or other uncertainty information where the chosen method provides them.

FAQs

1. What is factor analysis in simple terms?

Factor analysis examines whether several correlated measurements can be explained by a smaller number of underlying dimensions. For example, ten questionnaire questions may represent two latent constructs rather than ten entirely separate characteristics.

2. What are the two main types of factor analysis?

The two main types are exploratory factor analysis, which investigates a possible factor structure, and confirmatory factor analysis, which tests a structure specified before analysing the results.

3. What is considered a good factor loading?

There is no universal cutoff. Values such as .30, .40, .50 or .70 are sometimes used as reference points, but adequacy depends on sample size, uncertainty, communality, cross-loadings, research stage and theoretical importance.

4. Is 100 participants enough for factor analysis?

It may be enough under favourable conditions, but it may be inadequate when loadings are weak, communalities are low, factors have few indicators, data are highly ordinal, or the model is complex. Sample planning should be model-specific.

5. Can factor analysis be used with Likert-scale data?

Yes. However, individual Likert-type items are ordered categorical variables. Polychoric correlations and ordinal estimators may be preferable when there are few categories, skewed distributions or other violations of continuous-variable assumptions.

6. What do KMO and Bartlett’s test indicate?

KMO evaluates sampling adequacy based on the correlation pattern. Bartlett’s test evaluates whether the correlation matrix differs from an identity matrix. They are useful diagnostics but do not prove that the resulting factor solution is valid. IBM’s documentation similarly defines Bartlett’s test in relation to the identity matrix.

7. What is the difference between Varimax and Oblimin rotation?

Varimax is an orthogonal rotation that constrains factor correlations to zero. Oblimin is an oblique rotation that permits factors to correlate. Oblique rotation is often more realistic when the measured constructs are theoretically related.

8. How many factors should be retained?

Use parallel analysis, theoretical expectations, scree evidence, model adequacy, factor interpretability and solution stability. Do not rely exclusively on the eigenvalue-greater-than-one rule.

9. Can EFA and CFA be conducted on the same sample?

They can be computed on the same data, but doing so provides weak evidence of confirmation because the CFA is evaluating a structure derived from those observations. Independent data or a defensible validation split is preferable.

10. Does factor analysis prove that a latent factor is real?

No. It shows that a statistical latent-variable model may reproduce important relationships among the indicators. The interpretation of a factor must also be supported by theory, measurement design, external evidence and replication.

Conclusion

Factor analysis helps researchers investigate whether correlations among measured variables can be represented by a smaller set of latent dimensions. EFA is primarily used to explore structure, whereas CFA evaluates a structure specified in advance.

A credible analysis requires more than clicking a software option. Researchers must justify the indicators, correlation matrix, extraction or estimator, factor-retention method, rotation, interpretation and validation strategy. Statistical output should be evaluated alongside theory, measurement quality, uncertainty and evidence that the factor solution can be reproduced.


References

  • Bartlett, M. S. (1950). Tests of significance in factor analysis. British Journal of Statistical Psychology, 3(2), 77–85.
  • Fabrigar, L. R., Wegener, D. T., MacCallum, R. C., & Strahan, E. J. (1999). Evaluating the use of exploratory factor analysis in psychological research. Psychological Methods, 4(3), 272–299.
  • Flora, D. B., & Curran, P. J. (2004). An empirical evaluation of alternative methods of estimation for confirmatory factor analysis with ordinal data. Psychological Methods, 9(4), 466–491.
  • Hayton, J. C., Allen, D. G., & Scarpello, V. (2004). Factor retention decisions in exploratory factor analysis: A tutorial on parallel analysis. Organizational Research Methods, 7(2), 191–205.
  • Hu, L.-T., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling: A Multidisciplinary Journal, 6(1), 1–55.
  • MacCallum, R. C., Widaman, K. F., Zhang, S., & Hong, S. (1999). Sample size in factor analysis. Psychological Methods, 4(1), 84–99.
  • Rosseel, Y. (2012). lavaan: An R package for structural equation modeling. Journal of Statistical Software, 48(2), 1–36.
  • Schmitt, T. A. (2011). Current methodological considerations in exploratory and confirmatory factor analysis. Journal of Psychoeducational Assessment, 29(4), 304–321.
  • Watkins, M. W. (2018). Exploratory factor analysis: A guide to best practice. Journal of Black Psychology, 44(3), 219–246.

Software and Methodological Resources

  • Columbia University Mailman School of Public Health. (n.d.). Exploratory factor analysis.
  • IBM. (n.d.). Factor analysis: IBM SPSS Statistics documentation.
  • JASP Team. (n.d.). JASP features and factor-analysis documentation.
  • Revelle, W. (n.d.). psych: Procedures for psychological, psychometric, and personality research.
  • Rosseel, Y. (n.d.). The lavaan tutorial.

About the author

Muhammad Hassan

Muhammad Hassan writes about research design, academic methods and data-analysis concepts for ResearchMethod.net. His work focuses on presenting methodological topics in clear language for students and early-career researchers. Articles are developed from recognized methodological literature and official software documentation.