Skip to content
intermediate

Linear Discriminant Analysis: Derive the Gaussian Classifier

You fit LinearDiscriminantAnalysis, plot the boundary, and get a straight line. Then someone asks why it is straight, and the honest answer is "because the…

Published 2026-10-02Updated 2026-10-0410 min read
Close-up of a businessman analyzing colorful statistical data in an office setting.
Close-up of a businessman analyzing colorful statistical data in an office setting. Photo by ANTONI SHKRABA production on Pexels.

You fit LinearDiscriminantAnalysis, plot the boundary, and get a straight line. Then someone asks why it is straight, and the honest answer is "because the library says so." That is not understanding. That is a magic trick you happen to be holding.

Here is the stronger model: LDA is a generative Gaussian classifier. It assumes each class is a cloud of points drawn from a Gaussian distribution, and that all clouds share one covariance matrix. The straight boundary is not a design choice. It is a consequence of that single shared-covariance assumption. Everything below is the algebra that turns the assumption into the line.

Why a Gaussian Model Produces a Straight Line

LDA takes a generative stance. Instead of modeling the decision boundary directly, it models the class-conditional distribution P(x∣y=k)P(x \mid y=k) and uses Bayes' rule to turn that into a decision. If you have already worked through the generative-versus-discriminative distinction, this is the same fork in the road: model the data, then derive the decision.

The setup is simple. For each class kk, assume the data comes from a multivariate Gaussian with mean μk\mu_k and covariance Σk\Sigma_k. Bayes' rule gives the posterior:

P(y=k∣x)=P(x∣y=k) P(y=k)P(x)P(y=k \mid x) = \frac{P(x \mid y=k)\,P(y=k)}{P(x)}

You pick the class with the largest posterior. That is the whole classifier.

Now the central claim: when every class shares one covariance matrix Σ\Sigma, the quadratic terms in the posterior cancel, and the comparison collapses to a linear function of xx. The boundary between any two classes becomes a hyperplane.

QDA is the sibling that keeps per-class covariances Σk\Sigma_k. It keeps the quadratic terms, and its boundary curves. LDA is the special case where those terms vanish.

Notation and Assumptions Before the Algebra

Define the symbols once so the derivation reads cleanly.

SymbolMeaning
x∈Rdx \in \mathbb{R}^dfeature vector for one sample
kkclass label, k=1,…,Kk = 1, \dots, K
μk∈Rd\mu_k \in \mathbb{R}^dmean of class kk
Σ\Sigmashared covariance matrix, d×dd \times d
P(y=k)P(y=k)class prior
P(y=k∣x)P(y=k \mid x)posterior probability

Four assumptions are doing real work here:

  1. Class-conditional Gaussianity. Each class's features follow a Gaussian. This gives the exponential form that later becomes a quadratic.
  2. Shared covariance (homoscedasticity). Σk=Σ\Sigma_k = \Sigma for all kk. This is the assumption that kills the quadratic term.
  3. Full-rank Σ\Sigma. The inverse Σ−1\Sigma^{-1} must exist. If features are collinear or you have more features than samples, it does not.
  4. Priors from training frequencies. P(y=k)P(y=k) is estimated as the fraction of training samples in class kk.

These are claims about the data-generating process, not implementation details. When you call the estimator, it estimates μk\mu_k, Σ\Sigma, and the priors from your training set, then plugs them into the formula we are about to derive.

From Bayes' Rule to the Log Posterior

Start with the posterior and drop the evidence term. Since P(x)P(x) does not depend on kk, it is identical for every class and cannot change the argmax:

P(y=k∣x)∝P(x∣y=k) P(y=k)P(y=k \mid x) \propto P(x \mid y=k)\,P(y=k)

Now substitute the multivariate Gaussian density for P(x∣y=k)P(x \mid y=k):

P(x∣y=k)=1(2π)d/2 ∣Σ∣1/2exp⁡ ⁣(−12(x−μk)TΣ−1(x−μk))P(x \mid y=k) = \frac{1}{(2\pi)^{d/2}\,|\Sigma|^{1/2}} \exp\!\left(-\tfrac{1}{2}(x-\mu_k)^T \Sigma^{-1}(x-\mu_k)\right)

The exponent contains a Mahalanobis distance: how far xx is from μk\mu_k, measured in units of the covariance. That distance is the heart of the classifier.

Take the log. The exponential disappears, products become sums, and the normalizing constant becomes an additive term:

log⁡P(y=k∣x)=−12(x−μk)TΣ−1(x−μk)+log⁡P(y=k)+C\log P(y=k \mid x) = -\tfrac{1}{2}(x-\mu_k)^T \Sigma^{-1}(x-\mu_k) + \log P(y=k) + C

Here CC collects the terms that do not depend on kk: the −d2log⁡(2π)-\tfrac{d}{2}\log(2\pi) and the −12log⁡∣Σ∣-\tfrac{1}{2}\log|\Sigma|. Because Σ\Sigma is shared, ∣Σ∣|\Sigma| is the same for every class, so it folds into CC and disappears from the comparison. Keep CC visible for one more step. It is about to do its job.

Knowledge check

Check your understanding

Answer this question before you continue.

When comparing classes for a fixed feature vector x, why can the denominator P(x) in Bayes' rule be dropped?
Single Choice

Focus: Explain why the evidence term can be omitted when selecting the most probable class.

Where the Quadratic Term Cancels

Two side-by-side Gaussian contour plots compare shared and unequal covariance. Identically shaped class contours on the left are separated by a straight line; differently shaped contours on the right have a curved boundary.
With shared covariance, the common quadratic term cancels and leaves a straight boundary; class-specific covariances retain curvature.

This is the single algebraic move that makes LDA linear. Expand the Mahalanobis term:

(x−μk)TΣ−1(x−μk)=xTΣ−1x−2μkTΣ−1x+μkTΣ−1μk(x-\mu_k)^T \Sigma^{-1}(x-\mu_k) = x^T\Sigma^{-1}x - 2\mu_k^T\Sigma^{-1}x + \mu_k^T\Sigma^{-1}\mu_k

Look at each piece:

  • xTΣ−1xx^T\Sigma^{-1}x contains no kk. It is the same for every class.
  • −2μkTΣ−1x-2\mu_k^T\Sigma^{-1}x is linear in xx, and it depends on kk through μk\mu_k.
  • μkTΣ−1μk\mu_k^T\Sigma^{-1}\mu_k is a constant per class.

Substitute this back into the log posterior. The xTΣ−1xx^T\Sigma^{-1}x term is identical for every class, so it cannot change which class wins the argmax. Drop it, along with the other class-independent constants, and collect what remains:

log⁡P(y=k∣x)=μkTΣ−1x−12μkTΣ−1μk+log⁡P(y=k)+C′\log P(y=k \mid x) = \mu_k^T\Sigma^{-1}x - \tfrac{1}{2}\mu_k^T\Sigma^{-1}\mu_k + \log P(y=k) + C'

What survives is linear in xx. Define the weight vector and offset:

wk=Σ−1μk,bk=−12μkTΣ−1μk+log⁡P(y=k)w_k = \Sigma^{-1}\mu_k, \qquad b_k = -\tfrac{1}{2}\mu_k^T\Sigma^{-1}\mu_k + \log P(y=k)

The score is now a clean linear function:

δk(x)=wkTx+bk\delta_k(x) = w_k^T x + b_k

Assign xx to the class with the largest δk(x)\delta_k(x). Because each score is linear, the set of points where two scores are equal is a hyperplane. That is the LDA decision boundary, and it is straight for exactly one reason: the shared covariance made the quadratic term cancel.

Knowledge check

Check your understanding

Answer this question before you continue.

In the expanded Mahalanobis term, which part can be removed from the class comparison because it is identical for every class?
Misconception Check

Focus: Identify the class-independent quadratic term whose cancellation yields a linear LDA score.

Reading the Score: Means, Covariance, and Priors

The formula is short enough to hold in your head. Each piece has a job.

The mean term. wk=Σ−1μkw_k = \Sigma^{-1}\mu_k points the score toward class kk's center, but stretched by the inverse covariance. In the sphered space where Σ=I\Sigma = I, this reduces to plain Euclidean distance to the mean. LDA is, at bottom, a nearest-mean classifier under a learned metric.

The covariance. Σ−1\Sigma^{-1} reweights feature directions. Directions with small variance get large weight, because a small deviation there is informative. Directions with large variance get down-weighted. This is the model's built-in scale handling: a feature measured in the thousands is not automatically more influential than one measured in fractions, because the inverse covariance rescales each direction by its spread. What still bites in practice is estimation. When features are nearly collinear, or when the sample size is small relative to the number of features, the estimated Σ−1\Sigma^{-1} becomes unstable and the weights swing wildly. That is a data problem, not a units problem, and shrinkage or regularization of the covariance is the fix.

The priors. log⁡P(y=k)\log P(y=k) shifts the offset. A more likely class gets a higher baseline, so the boundary slides toward the less likely class. With equal priors, the boundary sits at the midpoint in Mahalanobis space. Change a prior, and the line translates without rotating.

There is a clean geometric reading of all this: sphere the data so the covariance becomes the identity, then classify by nearest mean in Euclidean distance, adjusted by the log prior. The sphering step is what turns a tilted, stretched Gaussian cloud into a symmetric ball.

Knowledge check

Check your understanding

Answer this question before you continue.

Under the article's LDA interpretation, what happens to a feature direction with large variance, all else equal?
Scenario Interpretation

Focus: Interpret how inverse covariance changes the influence of feature directions in the LDA score.

A Worked Two-Class Example

Take two classes in two dimensions. Suppose:

μ1=[00],μ2=[22],Σ=[1001]\mu_1 = \begin{bmatrix} 0 \\ 0 \end{bmatrix}, \quad \mu_2 = \begin{bmatrix} 2 \\ 2 \end{bmatrix}, \quad \Sigma = \begin{bmatrix} 1 & 0 \\ 0 & 1 \end{bmatrix}

With Σ=I\Sigma = I, the inverse is the identity, so wk=μkw_k = \mu_k. The scores are:

δ1(x)=0,δ2(x)=2x1+2x2−4+log⁡P(y=2)P(y=1)\delta_1(x) = 0, \qquad \delta_2(x) = 2x_1 + 2x_2 - 4 + \log\frac{P(y=2)}{P(y=1)}

Set the priors equal. The log⁡\log term is zero, and the boundary is where δ1=δ2\delta_1 = \delta_2:

2x1+2x2−4=0⇒x1+x2=22x_1 + 2x_2 - 4 = 0 \quad\Rightarrow\quad x_1 + x_2 = 2

That is the perpendicular bisector between the two means. Now make class 2 rarer: P(y=2)=0.2P(y=2) = 0.2, P(y=1)=0.8P(y=1) = 0.8. The log ratio is log⁡(0.25)≈−1.386\log(0.25) \approx -1.386, so the boundary becomes:

2x1+2x2−4−1.386=0⇒x1+x2≈2.692x_1 + 2x_2 - 4 - 1.386 = 0 \quad\Rightarrow\quad x_1 + x_2 \approx 2.69

The line slid toward class 2. Same orientation, new intercept. That is the prior doing its work: it makes the classifier demand more evidence before committing to the rare class.

Now change the covariance to Σ=[4001]\Sigma = \begin{bmatrix} 4 & 0 \\ 0 & 1 \end{bmatrix}. The inverse is [0.25001]\begin{bmatrix} 0.25 & 0 \\ 0 & 1 \end{bmatrix}, so the first feature is down-weighted. The weight vector rotates, and the boundary tilts. The covariance controls orientation and scale; the priors control position.

You can verify the hand result against the library in a few lines:

import numpy as np
from sklearn.discriminant_analysis import LinearDiscriminantAnalysis

X = np.array([[0, 0], [2, 2], [0.5, -0.5], [2.5, 1.5]])
y = np.array([0, 1, 0, 1])

lda = LinearDiscriminantAnalysis(priors=[0.8, 0.2]).fit(X, y)
print(lda.coef_)        # weight vector, up to scale
print(lda.intercept_)   # offset, includes the log-prior shift

The coefficients will not match your hand calculation exactly, because the library estimates Σ\Sigma from four points rather than using the identity. But the structure is the same: a weight vector and an offset. Change the priors argument and watch the intercept move while the coefficients barely budge.

Knowledge check

Check your understanding

Answer this question before you continue.

In the worked example, changing the priors from equal to P(y=2)=0.2 and P(y=1)=0.8 changes the boundary from x₁+x₂=2 to approximately x₁+x₂=2.69. What does this shift indicate?
Output Prediction

Focus: Predict the direction of the boundary shift when class 2 has a lower prior in the worked two-class example.

When the Assumptions Break

The derivation is a diagnostic checklist. Each assumption, when violated, produces a specific failure.

Unequal covariances. The quadratic terms no longer cancel. Under the Gaussian model, the true boundary curves, and LDA fits a straight line through a curved problem. It underfits. QDA keeps the per-class covariances and recovers the curve, at the cost of estimating more parameters.

Non-Gaussian or multimodal classes. A single mean and covariance cannot represent a class with two clusters. The score will be confidently wrong in the tails, where the Gaussian assigns probability mass the data never had. The boundary may still be linear, but it is the wrong boundary for the data.

Near-singular Σ\Sigma. When features are nearly collinear or dd approaches the sample size, Σ−1\Sigma^{-1} becomes unstable. Small changes in the data produce large swings in the weights. Shrinkage or regularization of the covariance is the standard fix.

Class imbalance. The priors dominate the offset. The boundary drifts toward the minority class. Sometimes that is exactly what you want; sometimes it is a silent failure that inflates accuracy while destroying recall on the class you care about.

Common mistake: Treating LDA as "just a linear classifier like logistic regression." LDA makes a strong generative assumption about the shape of each class. Logistic regression does not. That difference is why LDA can be more stable in small samples and more fragile when the Gaussian assumption is wrong.

When to Reach for LDA

Use LDA when the classes are roughly Gaussian, roughly equally spread, and your sample size is small relative to the number of features. The shared covariance is a strong, low-variance estimator, and strong assumptions buy stability when data is scarce. It also works as an interpretable baseline and as a dimensionality-reduction step before another classifier.

Do not reach for it as a default on complex nonlinear problems. If the linearity assumption is clearly wrong, the derivation above tells you exactly why, and that is your signal to move on.

Compared with logistic regression, the tradeoff is assumption versus flexibility. LDA assumes a generative model and can outperform logistic regression in small samples when that model holds. Logistic regression makes fewer distributional assumptions and is the safer default when you cannot justify the Gaussian shape.

The next time you fit LDA and get a straight line, you will know it is straight because one shared covariance canceled the quadratic term, and you will know which knob to turn to move it. Perturb the priors in the worked example and watch the intercept slide. Perturb the covariance and watch the boundary rotate. When the shared-covariance assumption finally breaks, the natural next steps are QDA and regularized discriminant analysis, both of which relax the assumption this derivation depends on.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A dataset has approximately Gaussian classes with clearly different covariance matrices, and the class boundary appears curved. Which interpretation best follows from the article?
Question 1 of 2Comparison Reasoning

Focus: Relate unequal class covariances to the limitation of LDA and the role of QDA.

Which situation best matches the article's guidance for trying LDA?
Question 2 of 2Scenario Interpretation

Focus: Choose when LDA's shared-covariance assumptions and lower-variance estimation make it a reasonable model to try.

References

  1. 1.2. Linear and Quadratic Discriminant Analysis — scikit-learn 1.9.1 documentationscikit-learn.org
  2. Characterization of a Family of Algorithms for Generalized ...jmlr.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.