Discriminant analysis is a multivariate statistical method used to distinguish predefined groups and classify new observations from quantitative predictors. It creates one or more weighted functions that maximize separation between groups. Linear discriminant analysis assumes a common within-group covariance matrix; quadratic discriminant analysis permits different covariance matrices and curved decision boundaries.

Introduction
Researchers frequently encounter questions in which the outcome is not a numerical score but membership in a known category. A university may want to classify applicants into academic-support groups, a biologist may need to distinguish related species, or a medical researcher may examine whether measured characteristics separate diagnostic categories.
Discriminant analysis provides a structured way to address such questions. It can determine which combinations of predictors separate known groups, summarize the dimensions along which the groups differ, and classify new observations.
This article explains the meaning, purposes, formulas, assumptions, types, procedures, interpretation, limitations, software, and reporting of discriminant analysis. It also shows how the method differs from logistic regression, MANOVA, principal component analysis, and cluster analysis.
Key Takeaways
- Discriminant analysis predicts or explains membership in two or more predefined groups.
- The outcome is categorical, while classical discriminant-analysis predictors are usually quantitative.
- LDA assumes a common within-group covariance matrix; QDA estimates a separate covariance matrix for each group.
- The maximum number of discriminant functions is the smaller of the number of predictors and the number of groups minus one.
- Training accuracy should not be treated as evidence of performance on new data; use cross-validation or independent validation.
- Coefficients, structure correlations, centroids, posterior probabilities, and classification metrics answer different interpretive questions.
What Is Discriminant Analysis?
Discriminant analysis is a supervised multivariate technique that finds combinations of predictor variables that separate known groups and can use those combinations to classify observations.
The groups must be defined before the model is fitted. For example, a dataset may contain patients already classified as having mild, moderate, or severe disease. The analysis then examines whether measurements such as age, blood pressure, and laboratory values distinguish those categories.
The term is used in two closely related ways:
- Discrimination: Finding dimensions that explain how known groups differ.
- Classification: Assigning observations to groups from their predictor values.
These purposes should not be confused. A statistically detectable group difference does not automatically imply that individual cases can be classified accurately. Groups can have different average profiles while still overlapping considerably.
The Two Main Purposes of Discriminant Analysis
1. Explaining group separation
Canonical discriminant analysis finds linear combinations of predictors that maximize separation between group means relative to variation within groups.
This purpose is interpretive. The researcher examines questions such as:
- Which variables contribute most to group separation?
- How many meaningful dimensions distinguish the groups?
- Which groups are separated by each function?
- Are group differences primarily related to one predictor or a broader profile?
2. Classifying observations
Classification discriminant analysis estimates rules for assigning a case to one of the known groups. The assignment may be based on discriminant scores, distances, prior probabilities, posterior probabilities, or explicit misclassification costs.
This purpose is predictive. The central question is not merely whether the groups differ, but whether a model can classify new cases accurately enough for its intended use.
Important Terms
| Term | Meaning |
|---|---|
| Grouping variable | The categorical outcome indicating known group membership |
| Predictor | A measured characteristic used to distinguish or classify the groups |
| Discriminant function | A weighted combination of predictors |
| Discriminant coefficient | A weight assigned to a predictor in a function |
| Discriminant score | A case’s calculated value on a discriminant function |
| Group centroid | The mean discriminant score for a group |
| Prior probability | The probability of group membership before considering a case’s predictors |
| Posterior probability | The updated probability of group membership after considering the predictors |
| Classification function | A group-specific scoring equation used to assign cases |
| Misclassification | Assigning an observation to the wrong group |
| Structure coefficient | Correlation between a predictor and a discriminant function |
| Confusion matrix | Table comparing actual and predicted group membership |
How Does Discriminant Analysis Work?
Discriminant analysis looks for a direction or set of directions in the predictor space that produces substantial differences between groups while keeping observations within each group relatively close together.
Imagine two groups plotted using two predictors. The groups may overlap when viewed along either original axis. A new axis created from a weighted combination of both predictors may separate them more clearly.
For two groups, Fisher’s idea can be expressed as maximizing the ratio:
[
J(\mathbf{w})=
\frac{\mathbf{w}^{T}S_B\mathbf{w}}
{\mathbf{w}^{T}S_W\mathbf{w}}
]
where:
- (\mathbf{w}) is the vector of weights,
- (S_B) represents between-group variation, and
- (S_W) represents within-group variation.
A useful direction has large between-group variation and comparatively small within-group variation.
The Discriminant Function Formula
A canonical linear discriminant function can be written as:
[
Z_m=a_m+b_{m1}X_1+b_{m2}X_2+\cdots+b_{mp}X_p
]
where:
- (Z_m) is the score on function (m),
- (a_m) is a constant,
- (b_{mj}) is the coefficient for predictor (j),
- (X_j) is a predictor value, and
- (p) is the number of predictors.
Each observation receives a score on every retained function. Group centroids show where the average member of each group lies on those dimensions.
LDA classification score
Under a multivariate Gaussian model with a common covariance matrix, the score for class (k) may be written as:
[
\delta_k(\mathbf{x})=
\mathbf{x}^{T}\Sigma^{-1}\boldsymbol{\mu}_k
-\frac{1}{2}\boldsymbol{\mu}_k^{T}\Sigma^{-1}\boldsymbol{\mu}_k
+\log(\pi_k)
]
where:
- (\mathbf{x}) is the predictor vector for the observation,
- (\boldsymbol{\mu}_k) is the mean vector for class (k),
- (\Sigma) is the common covariance matrix, and
- (\pi_k) is the prior probability for class (k).
The observation is assigned to the class with the largest score.
QDA classification score
Quadratic discriminant analysis allows each class to have its own covariance matrix:
[
\delta_k(\mathbf{x})=
-\frac{1}{2}\log|\Sigma_k|
-\frac{1}{2}(\mathbf{x}-\boldsymbol{\mu}_k)^T
\Sigma_k^{-1}
(\mathbf{x}-\boldsymbol{\mu}_k)
+\log(\pi_k)
]
Because (\Sigma_k) differs between groups, quadratic terms do not cancel. The resulting classification boundaries can therefore be curved.
How Many Discriminant Functions Are Produced?
The maximum number of discriminant functions is:
[
\min(p,g-1)
]
where:
- (p) is the number of predictors, and
- (g) is the number of groups.
Examples:
- Two groups and five predictors produce no more than one function.
- Three groups and five predictors produce no more than two functions.
- Four groups and two predictors produce no more than two functions.
Functions are ordered by the amount of group separation they explain. The first provides the strongest separation, the second explains the strongest remaining separation while being distinct from the first, and so forth.
Types of Discriminant Analysis
Linear Discriminant Analysis
Linear discriminant analysis assumes that the predictor distributions within the groups have different means but share a common covariance matrix.
LDA creates linear decision boundaries. It is relatively economical because only one pooled covariance matrix must be estimated.
LDA is appropriate when:
- Groups are known in advance.
- Predictors are quantitative.
- Within-group distributions are reasonably compatible with a multivariate Gaussian model.
- Covariance structures are similar enough for the research purpose.
- The sample is sufficient to estimate the pooled covariance matrix.
LDA can also be used for supervised dimensionality reduction by projecting data onto the discriminant directions.
Quadratic Discriminant Analysis
Quadratic discriminant analysis allows each group to have a different covariance matrix and can produce curved decision boundaries.
QDA is more flexible than LDA, but that flexibility has a cost. A separate covariance matrix must be estimated for every group, so QDA generally requires more data and is more vulnerable to unstable estimates.
QDA may be considered when:
- Covariance patterns differ substantially between groups.
- Curved boundaries are plausible.
- Each group contains enough observations.
- Cross-validation supports improved out-of-sample performance.
A significant Box’s M test alone should not automatically determine the choice. Model assumptions, sample size, scientific plausibility, and validated prediction should all be considered.
Two-Group Discriminant Analysis
Two-group analysis has one categorical outcome with exactly two categories. Only one canonical discriminant function can be produced.
Examples include:
- Completer versus non-completer.
- Diseased versus non-diseased.
- Genuine versus counterfeit.
- Customer retained versus customer lost.
Multiple or Canonical Discriminant Analysis
Multiple discriminant analysis is used when the outcome contains three or more groups. It can produce several functions, up to (\min(p,g-1)).
The functions may distinguish different contrasts. In a three-group problem, the first function may separate one group from the other two, while the second distinguishes the two remaining groups.
Regularized Discriminant Analysis
Regularized discriminant analysis modifies covariance estimates to improve stability. It creates a continuum between highly pooled and highly group-specific covariance structures and may also shrink covariance estimates toward simpler forms.
Regularization is valuable when:
- There are many predictors.
- Some predictors are strongly correlated.
- Group samples are limited.
- Ordinary covariance matrices are unstable or singular.
- Classical LDA or QDA overfits.
The amount of regularization should be selected using resampling within the training data rather than chosen after examining the test results.
Stepwise Discriminant Analysis
Stepwise analysis automatically adds or removes predictors according to a statistical criterion, often based on Wilks’ lambda or an F statistic.
Although convenient, stepwise selection can:
- Produce unstable variable sets.
- Overstate statistical significance.
- Exploit chance patterns.
- Make coefficients difficult to reproduce.
- Create optimistic classification estimates when selection and evaluation use the same data.
Pre-specified predictors, penalized methods, or feature selection performed entirely inside cross-validation are usually more defensible.
Partial Least Squares Discriminant Analysis
PLS-DA adapts partial least squares methods to categorical outcomes and is common in chemometrics, spectroscopy, and other high-dimensional fields.
It can be useful when predictors are numerous and highly correlated, but it is susceptible to overfitting. The number of components and all preprocessing choices must be validated using an appropriate resampling design.
Assumptions of Discriminant Analysis
1. Groups are predefined
The grouping variable must represent categories that are known before the model is fitted. Discriminant analysis does not discover unknown groups.
Use cluster analysis or latent-class methods when the objective is to identify groups from the data.
2. Categories are mutually exclusive
Each observation should belong to only one outcome category. If people can belong to several categories simultaneously, ordinary single-label discriminant analysis is not appropriate.
3. Categories are collectively meaningful
The candidate groups should represent the relevant classification possibilities. A model forced to choose between species A and B will still assign an observation from species C to one of them unless rejection or novelty-detection procedures are added.
4. Observations are independent
One person’s measurements should not determine another person’s values. Repeated observations, matched pairs, clustered students, family data, or measurements from several locations within the same participant require methods that account for dependence.
5. Predictors are normally quantitative
Classical LDA and QDA are formulated for quantitative predictors. Dummy variables may sometimes be processed by software, but binary predictors cannot follow a multivariate normal distribution in the classical sense.
When predictors are mainly categorical, binary, counts, or strongly non-Gaussian, logistic or other generalized classification models may be more suitable.
6. Within-group multivariate normality
The relevant assumption is:
[
\mathbf{X}\mid Y=k \sim N(\boldsymbol{\mu}_k,\Sigma_k)
]
That is, the joint predictor distribution is multivariate normal within each group.
Checking every variable separately is not sufficient. Researchers should inspect:
- Within-group histograms and Q–Q plots.
- Scatterplots.
- Multivariate outliers.
- Skewness and heavy tails.
- Transformations supported by the measurement scale.
- Sensitivity of conclusions to alternative models.
For prediction, the practical consequence of non-normality should be assessed through out-of-sample performance. For inferential claims, violations may affect test statistics and interpretation more directly.
7. Equality of covariance matrices for LDA
LDA assumes:
[
\Sigma_1=\Sigma_2=\cdots=\Sigma_g
]
QDA relaxes this assumption by estimating (\Sigma_k) separately for each group.
Box’s M test is often reported, but it can be sensitive to sample size and departures from normality. Researchers should not use it as the only model-selection criterion.
8. No severe multicollinearity or singularity
Highly correlated or redundant predictors make covariance inversion unstable. Problems can also occur when a predictor is an exact linear combination of others.
Possible remedies include:
- Removing redundant variables.
- Combining conceptually overlapping measures.
- Using a carefully validated dimension-reduction procedure.
- Applying covariance shrinkage or regularization.
- Collecting more observations.
9. Adequate observations relative to predictors
There is no universally valid sample-size ratio. Adequacy depends on the number of predictors, number of groups, covariance complexity, class balance, signal strength, and desired precision.
For ordinary QDA, each group covariance matrix must be estimable. When a group has too few observations relative to the predictors, its covariance matrix becomes singular or highly unstable.
For ordinary LDA, the pooled covariance estimate also requires sufficient effective rank. Even when the software completes the calculation, small samples may produce unstable functions and overly optimistic accuracy.
10. No highly influential outliers
Means and covariance matrices are sensitive to extreme observations. Researchers should inspect whether unusual cases represent:
- Data-entry errors.
- Measurement failure.
- Valid members of a heavy-tailed population.
- A different population not represented by the specified groups.
Valid observations should not be deleted merely to improve the model. Robust or alternative classification methods may be more appropriate.
Assumptions summary
| Assumption | Why it matters | Practical response |
|---|---|---|
| Known groups | The method is supervised | Use clustering when groups are unknown |
| Independent observations | Standard covariance and tests assume independence | Use grouped or repeated-measures methods when necessary |
| Quantitative predictors | Classical Gaussian model concerns continuous measurements | Consider logistic models for categorical predictors |
| Within-group multivariate normality | Supports Gaussian discriminant rules and inference | Inspect distributions and compare alternatives |
| Equal covariance matrices for LDA | Produces linear boundaries and pooled estimates | Compare LDA with QDA or regularized methods |
| Sufficient sample size | Covariance estimation otherwise becomes unstable | Reduce predictors, regularize, or collect more data |
| Limited multicollinearity | Required for stable covariance inversion | Remove redundancy or apply shrinkage |
| No dominant outliers | Outliers distort means and covariance matrices | Verify, investigate, and perform sensitivity analysis |
When Should Discriminant Analysis Be Used?
Use discriminant analysis when:
- The outcome contains two or more known categories.
- The main predictors are quantitative.
- The study aims to explain group separation or classify cases.
- The model assumptions are plausible enough for the intended inference.
- Group labels are reliable.
- Adequate training data are available.
- Prediction can be evaluated on observations not used to fit the model.
When Should It Not Be Used?
Discriminant analysis may be unsuitable when:
- The outcome is continuous.
- Groups have not yet been identified.
- Observations can belong to several classes at once.
- Most predictors are categorical.
- Dependence between observations is ignored.
- The number of predictors is large relative to the sample and no regularization is used.
- Group labels contain serious measurement error.
- A new case may belong to none of the candidate groups.
- Only in-sample classification accuracy will be reported.
How to Conduct Discriminant Analysis
Step 1: Define the research question
Specify whether the main purpose is:
- Testing or describing group separation.
- Interpreting the variables that distinguish groups.
- Predicting membership for new observations.
- Developing a practical screening or classification tool.
The analysis and reporting should reflect the primary objective.
Step 2: Define the groups carefully
Explain how group membership was established and whether the categories are:
- Scientifically meaningful.
- Reliably measured.
- Mutually exclusive.
- Available before predictor modelling.
- Representative of the intended future population.
A classifier cannot be more valid than the labels used to train it.
Step 3: Select predictors before examining final performance
Predictors should have theoretical, clinical, educational, or operational relevance. Avoid adding large numbers of variables merely because they are available.
Predictor selection based on the complete dataset can leak information into evaluation. When feature selection is data-driven, it must occur separately within every training fold.
Step 4: Clean and partition the data
Before fitting the model:
- Check coding and units.
- Identify missing values.
- Investigate impossible values.
- Document exclusions.
- Address repeated observations.
- Reserve a test set when the sample permits.
- Keep preprocessing inside the training workflow.
Scaling is not required for the mathematical validity of ordinary LDA in the same way it is for distance methods using unadjusted Euclidean distance. However, standardization can improve numerical conditioning and make some coefficients easier to compare.
Step 5: Explore the groups
Examine:
- Group sizes.
- Predictor means and standard deviations.
- Within-group distributions.
- Correlations.
- Scatterplots.
- Potential nonlinear boundaries.
- Covariance patterns.
- Class imbalance.
- Outliers.
Exploration should guide model choice without using test-set outcomes.
Step 6: Assess model assumptions
Check independence from the study design rather than from a software test alone.
Evaluate within-group distributions, covariance structures, collinearity, effective sample size, and influential observations. Document whether the model is being used primarily for inference, description, or prediction.
Step 7: Fit candidate models
Depending on the research problem, compare:
- LDA.
- QDA.
- Regularized LDA or RDA.
- Binary or multinomial logistic regression.
- Other pre-specified classifiers.
Priors should reflect the deployment population or the decision objective. Equal priors should not be used automatically when real-world prevalence differs substantially.
Step 8: Interpret the functions and predictions
For interpretation, examine:
- Eigenvalues.
- Canonical correlations.
- Wilks’ lambda.
- Standardized coefficients.
- Structure coefficients.
- Group centroids.
- Plots of canonical scores.
For prediction, examine:
- Posterior probabilities.
- Confusion matrices.
- Sensitivity and specificity.
- Balanced accuracy.
- Class-specific errors.
- Calibration.
- Performance on unseen data.
Step 9: Validate the entire workflow
Use one or more of the following:
- Leave-one-out cross-validation.
- Repeated k-fold cross-validation.
- A held-out test set.
- Temporal validation.
- External validation in another institution or population.
All variable selection, transformation, imputation, scaling, and tuning must be repeated within the training portion of each resample.
Step 10: Report decisions transparently
Report:
- How groups were defined.
- Predictor selection.
- Missing-data handling.
- Assumption assessment.
- Priors and misclassification costs.
- Model type.
- Validation design.
- Performance metrics.
- Uncertainty or confidence intervals where available.
- Limitations affecting transportability.
Worked Hypothetical Example
Suppose a university wants to distinguish students who are likely to complete a foundation programme from those at risk of non-completion.
The outcome has two groups:
- Group 0: non-completion.
- Group 1: completion.
The predictors are:
- Entry assessment score.
- Prior GPA.
- Number of absences during the first month.
A hypothetical fitted function is:
[
D=-7.40+0.045(\text{assessment})
+1.10(\text{GPA})
-0.080(\text{absences})
]
Consider a student with:
- Assessment score = 70.
- GPA = 3.0.
- Absences = 2.
The score is:
[
D=-7.40+0.045(70)+1.10(3.0)-0.080(2)
]
[
D=-7.40+3.15+3.30-0.16=-1.11
]
Assume the estimated group centroids are:
- Non-completion centroid: (-1.35).
- Completion centroid: (0.95).
The score of (-1.11) is closer to the non-completion centroid, so a simple nearest-centroid rule would assign the student to the non-completion group.
This calculation is illustrative. In a real analysis, classification should also account for priors, covariance, posterior probabilities, decision costs, and validated performance. A classification should not automatically be treated as a definitive judgement about an individual.
How to Interpret Discriminant-Analysis Output
Software packages organize results differently, but the following sequence is common.
Group statistics
Group statistics show the number of observations, means, and standard deviations for each predictor within each group.
Use them to understand:
- Which variables show visible mean differences.
- Whether group sizes are imbalanced.
- Whether standard deviations differ markedly.
- Whether values appear implausible.
These statistics are descriptive and do not establish that the complete multivariate model classifies accurately.
Tests of equality of group means
SPSS may provide a univariate Wilks’ lambda and F test for each predictor.
A significant result suggests that the predictor mean differs across groups when considered separately. However:
- A non-significant variable may still contribute jointly.
- A significant variable may add little once correlated predictors are considered.
- Multiple testing can inflate false-positive findings.
- Statistical significance is not the same as predictive value.
Pooled within-group correlation matrix
This matrix shows correlations among predictors after accounting for group membership.
Very high correlations may indicate:
- Redundant predictors.
- Unstable coefficients.
- Difficulty assigning unique importance to variables.
- Potential covariance singularity.
Box’s M test
Box’s M evaluates equality of covariance matrices.
A small p-value indicates evidence against exact equality. It does not by itself prove that LDA is unusable or that QDA will predict better.
Interpret the result alongside:
- Group sample sizes.
- Within-group normality.
- Covariance estimates and plots.
- Model stability.
- Cross-validated comparisons.
Eigenvalues
Each canonical function has an eigenvalue representing its relative discriminatory strength. Larger eigenvalues indicate that a function separates groups more strongly relative to within-group variation.
The percentage associated with each eigenvalue indicates its share of the total discriminatory information represented by the retained functions.
Canonical correlation
The canonical correlation measures the association between a discriminant function and group membership.
It can be derived from the eigenvalue (\lambda):
[
R_c=\sqrt{\frac{\lambda}{1+\lambda}}
]
Larger values indicate stronger separation. The squared canonical correlation should be interpreted carefully as a measure of association for the discriminant dimension, not as ordinary regression (R^2) for a continuous outcome.
Wilks’ lambda
Wilks’ lambda measures the proportion of multivariate variation not explained by group differences on the tested functions. Lower values generally indicate stronger group separation.
Software normally converts lambda into an approximate chi-square statistic to test whether the remaining functions contribute significant discrimination.
With several functions, the tests are sequential. The first row may test all functions together, while later rows test the functions remaining after earlier ones are removed.
A significant Wilks’ lambda test indicates evidence of group separation. It does not show that classification accuracy is practically useful.
Standardized canonical coefficients
These coefficients resemble standardized regression weights. They show each predictor’s contribution while controlling for the others.
Large absolute coefficients may indicate stronger unique contributions, but interpretation can become unstable when predictors are correlated.
The sign is meaningful only relative to:
- The coding of variables.
- The signs of the other coefficients.
- Group centroids.
- The arbitrary orientation of the function.
Multiplying every coefficient by (-1) represents the same discriminant dimension in the opposite direction.
Structure matrix
The structure matrix contains correlations between predictors and discriminant functions. These values are often called structure coefficients or discriminant loadings.
They show which variables are most strongly associated with each function. When predictors are correlated, structure coefficients may be more stable for interpretation than standardized coefficients.
A sound interpretation considers both:
- Standardized coefficients for unique contribution.
- Structure coefficients for overall association.
Group centroids
A centroid is the mean discriminant score for a group.
Centroids help identify which groups are separated by a function. For example:
- A large positive centroid for group A and negative centroids for B and C suggests that the function mainly distinguishes A from the other groups.
- If B and C are separated on the second function, that function explains a different contrast.
Canonical discriminant plot
For three or more groups, plotting the first two functions can show:
- Group separation.
- Overlap.
- Outliers.
- Which groups differ on each dimension.
- Whether a linear representation appears reasonable.
A visually appealing plot does not replace validation.
Classification-function coefficients
Classification coefficients create a separate equation for every group. A case is commonly assigned to the group with the largest classification-function score.
Do not confuse classification-function coefficients with canonical coefficients. They are designed for assigning groups, not for explaining canonical dimensions.
Prior probabilities
Prior probabilities specify the expected group probabilities before considering predictors.
Common choices are:
- Equal priors.
- Priors based on training-group proportions.
- Priors based on known population prevalence.
- Decision-adjusted priors reflecting policy or costs.
The choice should be reported because it can change classifications.
Posterior probabilities
Posterior probabilities estimate class membership after combining:
- Predictor values.
- Estimated group distributions.
- Prior probabilities.
They express model-based uncertainty, but they are not automatically well calibrated. Calibration should be assessed when probabilities will inform decisions.
Classification table
A classification or confusion matrix compares actual and predicted groups.
For two groups it contains:
- True positives.
- False positives.
- True negatives.
- False negatives.
The overall percentage correctly classified can be misleading when one group dominates. Class-specific metrics should also be reported.
Cross-validated classification
Cross-validated results estimate performance when each prediction is generated without using that observation in the model fitting for that prediction.
Leave-one-out cross-validation is common in traditional software, but repeated k-fold cross-validation may provide more informative stability estimates in many applications.
Feature selection and tuning must occur inside each resample. Otherwise, cross-validated accuracy remains optimistic.
Evaluating Classification Performance
Accuracy
[
\text{Accuracy}=
\frac{\text{Correct predictions}}
{\text{All predictions}}
]
Accuracy is easy to understand but can be misleading with imbalanced classes.
Sensitivity
[
\text{Sensitivity}=
\frac{TP}{TP+FN}
]
Sensitivity is the proportion of actual positive cases correctly identified.
Specificity
[
\text{Specificity}=
\frac{TN}{TN+FP}
]
Specificity is the proportion of actual negative cases correctly identified.
Precision
[
\text{Precision}=
\frac{TP}{TP+FP}
]
Precision is the proportion of predicted positive cases that are actually positive.
F1 score
[
F1=
2\times
\frac{\text{Precision}\times\text{Sensitivity}}
{\text{Precision}+\text{Sensitivity}}
]
The F1 score balances precision and sensitivity but does not account directly for true negatives.
Balanced accuracy
Balanced accuracy averages sensitivity across classes. It is useful when group frequencies differ.
ROC AUC
For binary classification, the receiver operating characteristic curve evaluates sensitivity and false-positive rate across probability thresholds.
AUC measures ranking discrimination, not probability calibration or usefulness at a particular decision threshold.
Calibration
Calibration examines whether predicted probabilities correspond to observed event rates. A model can rank cases well but produce poorly calibrated probabilities.
Misclassification cost
Not all errors have equal consequences. Failing to identify a serious condition may be more costly than sending an unaffected person for additional assessment.
Costs should be specified from the research or decision context rather than inferred solely from statistical convenience.
LDA vs. QDA
| Feature | LDA | QDA |
|---|---|---|
| Covariance assumption | Common covariance matrix | Separate covariance matrix for every group |
| Decision boundary | Linear | Usually quadratic or curved |
| Number of covariance parameters | Lower | Higher |
| Data requirement | Generally lower | Generally higher |
| Flexibility | Lower | Higher |
| Variance of estimates | Usually lower | Usually higher |
| Risk with small samples | Moderate | Greater |
| Best use | Similar covariance structures or limited data | Substantial covariance differences with adequate data |
| Model selection | Validate on unseen data | Validate on unseen data |
QDA is not automatically superior because it is more flexible. When data are limited, its separate covariance estimates may introduce more estimation error than the flexibility removes.
Discriminant Analysis vs. Logistic Regression
| Question | Discriminant analysis | Logistic regression |
|---|---|---|
| Outcome | Categorical | Categorical |
| Predictor distribution | Models class-conditional predictor distributions | Does not require predictors themselves to be normally distributed |
| Boundary | Linear for LDA; curved for QDA | Linear in the specified logit unless nonlinear terms are added |
| Multiclass use | Inherently multiclass | Requires multinomial or related extension |
| Categorical predictors | Awkward under classical assumptions | Easily included through indicator coding |
| Covariance modelling | Explicit | Not required |
| Interpretive emphasis | Group separation and classification | Conditional outcome probabilities and odds |
| Sensitivity to distributional model | Potentially substantial | Different assumptions, including correct logit specification |
| Preferred choice | Plausible Gaussian structure and useful covariance modelling | Mixed predictor types or weaker interest in Gaussian assumptions |
Neither method is universally better. The choice should reflect:
- Research purpose.
- Predictor types.
- Distributional plausibility.
- Sample size.
- Need for interpretability.
- Probability calibration.
- Validated predictive performance.
Discriminant Analysis vs. MANOVA
MANOVA and discriminant analysis use related multivariate mathematics but reverse the practical emphasis.
- MANOVA: Tests whether groups differ across several continuous outcomes.
- Discriminant analysis: Uses quantitative measurements to explain or predict categorical group membership.
Canonical discriminant functions are often used after a significant MANOVA to describe the dimensions underlying group differences. However, a significant MANOVA does not guarantee accurate individual classification.
Discriminant Analysis vs. PCA
| Feature | Discriminant analysis | PCA |
|---|---|---|
| Uses group labels | Yes | No |
| Learning type | Supervised | Unsupervised |
| Main objective | Maximize group separation | Maximize total variance |
| Maximum dimensions | At most (g-1) and (p) | Up to (p) |
| Use | Classification and supervised projection | Compression, visualization, and structure discovery |
A high-variance PCA direction may be unrelated to group separation. Conversely, an LDA direction may explain relatively little total variance while distinguishing the classes effectively.
Discriminant Analysis vs. Cluster Analysis
- Discriminant analysis begins with known labels.
- Cluster analysis attempts to discover groups from similarities among observations.
A common workflow is to develop clusters in one sample and then build a classifier for assigning later cases. However, treating clusters as unquestionably true labels can exaggerate certainty because the labels were themselves estimated.
Advantages of Discriminant Analysis
- Directly addresses categorical group membership.
- Combines multiple predictors into interpretable functions.
- Supports two-group and multiclass classification.
- Provides group centroids and visual discriminant dimensions.
- Includes prior probabilities and decision costs.
- Has closed-form solutions for classical LDA and QDA.
- Can perform supervised dimensionality reduction.
- Connects classical multivariate statistics with modern classification.
- Can work well with limited tuning when its structure is plausible.
- Provides interpretable relationships between predictors and group separation.
Limitations of Discriminant Analysis
- Classical forms depend on restrictive distributional assumptions.
- Covariance estimates can be unstable in small or high-dimensional samples.
- Outliers can substantially alter means and covariance matrices.
- Categorical predictors do not fit the classical Gaussian framework naturally.
- Training accuracy is often optimistic.
- Stepwise procedures can be unstable and difficult to reproduce.
- The method forces a case into one of the available groups unless rejection rules are added.
- Group labels may be uncertain or incorrectly measured.
- Coefficients can be difficult to interpret under multicollinearity.
- Good average performance can hide poor results for minority or important subgroups.
Common Mistakes
Confusing LDA with latent Dirichlet allocation
In machine learning and text analysis, “LDA” can also mean latent Dirichlet allocation. Always write the full method name at first use.
Using discriminant analysis to discover groups
Discriminant analysis requires known groups. Use cluster analysis for exploratory group discovery.
Testing normality on the pooled sample
The model concerns predictor distributions within each group, not only the overall sample.
Treating Box’s M as a pass-or-fail gate
Exact covariance equality is rarely a realistic literal condition. Evaluate practical consequences through diagnostics and model comparison.
Reporting resubstitution accuracy only
Accuracy on the training sample is not a trustworthy estimate of performance on new cases.
Selecting variables before cross-validation
Feature selection using all observations leaks information into the validation process.
Interpreting only standardized coefficients
Use coefficients, structure correlations, group centroids, theory, and stability together.
Ignoring prior probabilities
Equal priors can be inappropriate when groups have very different population prevalence.
Relying only on overall accuracy
Report class-specific errors, balanced metrics, uncertainty, and calibration when relevant.
Assuming a significant function guarantees practical value
Statistical evidence of separation may coexist with extensive overlap and weak classification.
Software for Discriminant Analysis
SPSS
SPSS provides classical canonical linear discriminant analysis through:
Analyze > Classify > Discriminant
Researchers can request:
- Group statistics.
- Box’s M.
- Within-group correlations.
- Standardized coefficients.
- Structure matrices.
- Centroids.
- Classification functions.
- Posterior probabilities.
- Leave-one-out classification.
SPSS’s standard discriminant procedure primarily reflects classical linear discriminant analysis. Researchers needing flexible regularization or broader resampling workflows may require other software.
R
The MASS package provides lda() and qda().
library(MASS)
model_lda <- lda(group ~ x1 + x2 + x3, data = training_data)
pred_lda <- predict(model_lda, newdata = test_data)
model_qda <- qda(group ~ x1 + x2 + x3, data = training_data)
pred_qda <- predict(model_qda, newdata = test_data)
A complete workflow should separately preprocess the data, tune any regularization, and evaluate predictions on resampled or held-out observations.
Python
scikit-learn provides LinearDiscriminantAnalysis and QuadraticDiscriminantAnalysis.
from sklearn.discriminant_analysis import (
LinearDiscriminantAnalysis,
QuadraticDiscriminantAnalysis,
)
from sklearn.model_selection import cross_validate
lda = LinearDiscriminantAnalysis()
qda = QuadraticDiscriminantAnalysis(reg_param=0.1)
scoring = ["accuracy", "balanced_accuracy", "f1_macro"]
lda_results = cross_validate(
lda,
X,
y,
cv=5,
scoring=scoring,
return_train_score=False,
)
Preprocessing, feature selection, and parameter tuning should be placed in a pipeline so that they are learned separately in each training fold.
SAS
SAS provides:
PROC DISCRIMfor classification using linear, quadratic, and related discriminant rules.PROC CANDISCfor canonical discriminant analysis.
The POOL= option can control whether covariance matrices are pooled or estimated separately in normal-theory classification.
Minitab
Minitab offers menu-based discriminant analysis, cross-validation, linear or quadratic functions, group predictions, and classification summaries.
XLSTAT
XLSTAT provides Excel-based LDA and QDA, priors, model-selection options, classification tables, ROC analysis, cross-validation, and canonical visualizations.
OriginPro
OriginPro includes assumption checking, linear and quadratic methods, canonical score plots, and classification results.
How Discriminant Analysis Is Used in Modern Research
Discriminant analysis remains relevant because it is interpretable, computationally efficient, multiclass, and closely connected to probabilistic classification.
Contemporary applications include:
- Biological species classification.
- Medical and laboratory classification.
- Spectroscopy and chemometrics.
- Remote sensing.
- Credit and risk modelling.
- Educational placement.
- Pattern recognition.
- Customer segmentation.
- Quality control.
- Supervised feature extraction.
Modern research has extended classical discriminant analysis through:
- Shrinkage covariance estimation.
- Sparse discriminant analysis.
- Robust estimation.
- Kernel discriminant analysis.
- Regularized discriminant analysis.
- High-dimensional discriminant methods.
- Ensemble and deep-learning extensions.
These extensions should not be treated as automatic improvements. Their tuning, feature selection, and performance must be evaluated under a leakage-free validation design.
Artificial Intelligence and Discriminant Analysis
Generative AI tools can assist with:
- Drafting R, Python, SPSS, or SAS syntax.
- Explaining software output in simpler language.
- Creating simulated teaching examples.
- Suggesting diagnostic plots.
- Converting an analysis plan into a reporting checklist.
- Checking whether a results section omits essential information.
However, AI-generated statistical guidance can contain incorrect formulas, invented references, unsuitable assumptions, or code that evaluates the training data rather than unseen data.
Researchers should therefore:
- Verify code against official documentation.
- Check formulas against a statistical text or primary source.
- Never upload confidential data without appropriate authorization.
- Preserve an auditable script and software environment.
- Distinguish AI-generated explanations from verified statistical conclusions.
- Seek expert review for high-stakes medical, legal, financial, or policy decisions.
AI can support the workflow, but it cannot determine whether the research design, labels, sampling process, or causal interpretation is valid.
How to Report Discriminant Analysis
A complete report should include:
- Purpose of the analysis.
- Definition and size of each group.
- Predictors entered.
- Missing-data treatment.
- Assessment of independence, distributions, outliers, covariance, and collinearity.
- Model type and software.
- Prior probabilities.
- Variable-selection procedure.
- Number and significance of functions.
- Eigenvalues and canonical correlations.
- Main structure coefficients or standardized coefficients.
- Group centroids.
- Validation procedure.
- Confusion matrix and class-specific metrics.
- Limitations and intended population.
Adaptable APA-style reporting template
“A [linear/quadratic/regularized] discriminant analysis was conducted to determine whether [predictors] distinguished among [groups] and to evaluate classification performance. Group membership was defined using [criterion]. Assumptions were assessed using [procedures], and prior probabilities were set to [choice and reason].
The analysis produced [number] discriminant function(s). The first function had an eigenvalue of [value], a canonical correlation of [value], and significantly differentiated the groups, Wilks’ (\Lambda) = [value], (\chi^2)([df]) = [value], (p) = [value]. Variables most strongly associated with the function were [variables and structure coefficients].
Using [cross-validation/held-out/external validation], the model correctly classified [percentage]% of observations. Class-specific [sensitivity/specificity/balanced accuracy] values were [values]. These results indicate that [careful interpretation], although performance may be limited by [sample size, overlap, imbalance, assumption violation, or validation limitation].”
Do not report only the percentage classified correctly. The validation design and class-specific errors are essential for judging credibility.
Conclusion
Discriminant analysis is a supervised multivariate method for explaining differences among predefined groups and classifying observations from their predictor values. LDA provides a parsimonious linear model with a shared covariance matrix, while QDA offers more flexible boundaries by estimating group-specific covariance matrices.
A defensible analysis requires more than running a software procedure. Researchers must define reliable groups, assess the data structure, choose priors deliberately, prevent information leakage, validate performance on unseen data, interpret functions carefully, and report both statistical separation and practical classification accuracy.
