Skip to content
intermediate

Derive PCA: Variance-Maximizing Directions and Covariance Eigenvectors

You can call PCA(), read explained_variance_ratio_, and get a working result. Then someone asks why the principal directions are eigenvectors of the…

Published 2026-10-02Updated 2026-10-0411 min read
Minimalist image of a crescent moon in a clear blue sky over Samsun, Türkiye.
Minimalist image of a crescent moon in a clear blue sky over Samsun, Türkiye. Photo by Nrlh Semiz on Pexels.

You can call PCA(), read explained_variance_ratio_, and get a working result. Then someone asks why the principal directions are eigenvectors of the covariance matrix, and the room goes quiet. This article closes that gap. We will turn "find the direction of most variance" into a constrained optimization problem, solve it, and watch the covariance matrix fall out of the algebra.

Why the Eigenvector Answer Feels Like a Trick

Most learners meet PCA in the reverse order. First they see the code, then the explained variance output, then a sentence that says the components are eigenvectors of the covariance matrix. The sentence gets memorized. It never gets derived.

That is a fragile way to know something. If you cannot reconstruct why the answer is an eigenvector, you cannot reason about what happens when you standardize your features, when you keep three components instead of two, or when a high-variance direction turns out to be useless for prediction.

So here is the plan. We define notation, write the variance objective, add the constraint that makes it solvable, and let the eigenvector equation appear as a consequence rather than an assumption. Then we extend the argument to the second component and beyond, and finish by explaining how centering and scaling change the answer — and why preserved variance is not the same as predictive relevance.

Notation and Assumptions Before Any Algebra

Let XX be your data matrix with nn rows (samples) and pp columns (features). Each row is one observation. Each column is one variable.

Define the centered data matrix:

Xc=X−1xˉ⊤X_c = X - \mathbf{1}\bar{x}^\top

where xˉ\bar{x} is the vector of column means and 1\mathbf{1} is a column of ones. Centering subtracts the mean of each feature. This is a precondition, not a stylistic choice — we will see exactly why in a moment.

The covariance matrix is:

S=1nXc⊤XcS = \frac{1}{n} X_c^\top X_c

Some texts use 1n−1\frac{1}{n-1} instead. The choice only rescales every eigenvalue by the same constant. It does not change the eigenvectors, so it does not change the directions. Pick a convention and stay consistent.

Now project a centered sample xx onto a unit vector ww. The projection score is w⊤xw^\top x. The variance of those scores across all samples is:

w⊤Sww^\top S w

That single expression is the whole game. Everything that follows is about maximizing it.

Three assumptions are baked in. First, the structure we care about is linear — components are linear combinations of features. Second, variance is the quantity worth maximizing. Third, directions are only meaningful relative to the scaling you chose. That last one will matter a lot later.

Note: If you have not yet built intuition for what a principal component is, start with the concept article on PCA and dimensionality reduction. Here we assume you already know components are directions of variance and that projection is how you move data onto them.

The Variance Objective and Why the Constraint Is Mandatory

Write the goal plainly:

max⁡w  w⊤Sw\max_{w} \; w^\top S w

Try to solve it as stated and you hit a wall immediately. Scale ww by any constant cc. The objective becomes c2 w⊤Swc^2 \, w^\top S w. Make cc large and the variance grows without bound. The unconstrained problem has no maximum — only a direction that runs off to infinity.

The fix is to fix the length:

max⁡w  w⊤Swsubject to∥w∥2=w⊤w=1\max_{w} \; w^\top S w \quad \text{subject to} \quad \|w\|^2 = w^\top w = 1

This is not a technicality. It is what makes the question well-posed. The constraint says: I do not care how long the vector is, only which way it points. Direction is the free variable. Length is pinned.

Now use a Lagrange multiplier. Define:

L(w,λ)=w⊤Sw−λ(w⊤w−1)L(w, \lambda) = w^\top S w - \lambda(w^\top w - 1)

Take the gradient with respect to ww and set it to zero:

∂L∂w=2Sw−2λw=0\frac{\partial L}{\partial w} = 2Sw - 2\lambda w = 0

Divide by two:

Sw=λwSw = \lambda w

There it is. The eigenvector equation did not get imported from a linear algebra textbook. It fell out of the optimization. The principal directions are eigenvectors because that is what a constrained variance maximizer has to satisfy.

Knowledge check

Check your understanding

Answer this question before you continue.

Why must the variance-maximization problem constrain the length of the direction vector?
Misconception Check

Focus: Explain why PCA constrains a candidate direction to unit length when maximizing projected variance.

From Eigenvalue Equation to Largest Eigenvalue

A tilted elliptical data cloud with perpendicular principal axes. The long axis is labeled w₁ with variance λ₁ ≈ 5.56; the short axis is labeled w₂ with variance λ₂ ≈ 1.44.
The leading eigenvector points along the cloud’s greatest spread; its eigenvalue is the variance measured in that direction.

The equation Sw=λwSw = \lambda w tells us ww is an eigenvector, but it does not yet tell us which one. Substitute it back into the objective:

w⊤Sw=w⊤(λw)=λ w⊤w=λw^\top S w = w^\top (\lambda w) = \lambda \, w^\top w = \lambda

So the variance captured by a candidate direction equals its eigenvalue. To maximize variance, pick the eigenvector with the largest eigenvalue. That is the first principal component.

Because SS is symmetric, its eigenvectors can be chosen orthonormal — mutually perpendicular and unit length. This is not a lucky accident. It is the property that makes the next section clean.

A worked 2×2 example

Take a covariance matrix:

S=[4223]S = \begin{bmatrix} 4 & 2 \\ 2 & 3 \end{bmatrix}

The eigenvalues solve det⁡(S−λI)=0\det(S - \lambda I) = 0:

(4−λ)(3−λ)−4=λ2−7λ+8=0(4-\lambda)(3-\lambda) - 4 = \lambda^2 - 7\lambda + 8 = 0

λ=7±49−322=7±172≈5.56 and 1.44\lambda = \frac{7 \pm \sqrt{49 - 32}}{2} = \frac{7 \pm \sqrt{17}}{2} \approx 5.56 \text{ and } 1.44

For λ1≈5.56\lambda_1 \approx 5.56, solve (S−λ1I)w=0(S - \lambda_1 I)w = 0:

[−1.5622−2.56]w=0  ⟹  w≈[0.790.62]\begin{bmatrix} -1.56 & 2 \\ 2 & -2.56 \end{bmatrix} w = 0 \implies w \approx \begin{bmatrix} 0.79 \\ 0.62 \end{bmatrix}

Project your centered data onto this direction and the sample variance will be approximately 5.56. Project onto the second eigenvector and you get approximately 1.44. The first direction wins because it is the eigenvector of the larger eigenvalue — exactly as the derivation predicted.

Geometrically, the leading eigenvector points along the long axis of the data cloud's ellipsoid, and the eigenvalue is that axis's squared length. The second eigenvector points along the short axis. PCA is choosing the axes of the ellipsoid that best describe the spread.

Knowledge check

Check your understanding

Answer this question before you continue.

For the article's 2×2 covariance example, which conclusion follows when choosing the first principal direction?
Scenario Interpretation

Focus: Use the eigenvalue-variance relationship to identify the leading direction in the worked covariance example.

Extending to the Second Component and Beyond

The second component faces a harder problem. It must maximize variance, stay unit length, and be orthogonal to the first component. Otherwise it would just rediscover the same direction.

Write the objective with two constraints:

max⁡w  w⊤Swsubject tow⊤w=1,w⊤w1=0\max_{w} \; w^\top S w \quad \text{subject to} \quad w^\top w = 1, \quad w^\top w_1 = 0

The Lagrange setup now carries two multipliers:

L(w,λ,μ)=w⊤Sw−λ(w⊤w−1)−μ(w⊤w1)L(w, \lambda, \mu) = w^\top S w - \lambda(w^\top w - 1) - \mu(w^\top w_1)

Setting the gradient to zero gives the stationarity condition:

2Sw−2λw−μw1=02Sw - 2\lambda w - \mu w_1 = 0

Now multiply both sides by w1⊤w_1^\top from the left:

2w1⊤Sw−2λw1⊤w−μw1⊤w1=02w_1^\top S w - 2\lambda w_1^\top w - \mu w_1^\top w_1 = 0

Three terms, three facts. Since SS is symmetric and w1w_1 is an eigenvector, w1⊤S=λ1w1⊤w_1^\top S = \lambda_1 w_1^\top, so w1⊤Sw=λ1w1⊤ww_1^\top S w = \lambda_1 w_1^\top w. The orthogonality constraint makes w1⊤w=0w_1^\top w = 0. And w1⊤w1=1w_1^\top w_1 = 1 because w1w_1 is unit length. Substitute all three:

2λ1⋅0−2λ⋅0−μ⋅1=0  ⟹  μ=02\lambda_1 \cdot 0 - 2\lambda \cdot 0 - \mu \cdot 1 = 0 \implies \mu = 0

The orthogonality multiplier vanishes. Drop it from the stationarity condition and you are left with Sw=λwSw = \lambda w — the same eigenvector equation, now restricted to the subspace orthogonal to w1w_1. Maximizing within that subspace selects the eigenvector with the second-largest eigenvalue.

The pattern generalizes. The kk-th principal component is the eigenvector of the kk-th largest eigenvalue, and all components are mutually orthogonal. This is why components come out ordered by variance and why truncation keeps the top kk — you are keeping the directions that capture the most spread, in descending order.

The bookkeeping is clean too. The eigenvalues sum to the total variance in the data, so each explained variance ratio is just that eigenvalue's share of the sum. If λ1=5.56\lambda_1 = 5.56 and λ2=1.44\lambda_2 = 1.44, the first component explains 5.56/7.00≈79%5.56 / 7.00 \approx 79\% of the variance.

Knowledge check

Check your understanding

Answer this question before you continue.

Once the first component is fixed, what optimization characterizes the second component in the article's derivation?
Single Choice

Focus: Describe how the second principal component is selected after the first component is fixed.

Centering, Scaling, and How They Change the Eigenvectors

Here is where the derivation earns its keep. Two preprocessing choices change the eigenvectors themselves, not just the numbers.

Centering. The objective w⊤Sww^\top S w measures spread around the mean. If you skip centering, you are no longer eigendecomposing the covariance matrix. You are eigendecomposing the uncentered second-moment matrix, 1nX⊤X\frac{1}{n}X^\top X, which mixes two things: how the data varies and where the data sits relative to the origin. A feature with a large mean contributes a large diagonal entry that reflects its offset, not its spread. The leading direction can end up pointing toward the mean offset instead of the structure you care about. Centering removes that distraction, which is why ordinary PCA pipelines center first.

Scaling. Standardizing each feature to unit variance is equivalent to eigendecomposing the correlation matrix instead of the covariance matrix. The two matrices have different eigenvectors whenever features have different scales.

The consequence is concrete. Suppose one feature is measured in millimeters and another in meters. The millimeter feature has values roughly a thousand times larger, so its variance dominates the covariance matrix. The first principal component will point almost entirely at that feature — not because it carries more information, but because its units are bigger. That is a scale artifact masquerading as a finding.

Common mistake: Running PCA on unscaled features with different units and then interpreting the first component as "the most important direction." It is the most spread-out direction, which often just means the largest-unit feature.

My decision rule:

SituationPreprocessing
Features have different units or arbitrary scalesStandardize to unit variance
Features share comparable units and relative magnitude is meaningfulKeep raw covariance
Sparse data where centering would densify the matrixSkip centering only if you accept the uncentered objective, or use a method built for sparse input

That last row is a real constraint, not a loophole. Centering a sparse matrix turns zeros into nonzero means, which can blow up memory. Some pipelines avoid centering for exactly this reason, accepting a different objective in exchange for tractability. That is a deliberate trade, not a defective covariance calculation.

Knowledge check

Check your understanding

Answer this question before you continue.

Two features have very different units. What change does standardizing them make to the PCA analysis described in the article?
Comparison Reasoning

Focus: Explain how standardizing features affects the matrix and directions used by PCA.

What the Derivation Does Not Promise

The derivation optimizes variance. It does not optimize class separation, prediction accuracy, or causal structure. Those are different objectives, and confusing them is the most common way PCA gets misused.

A high-variance direction can be pure noise or an irrelevant nuisance variable. A low-variance direction can carry the signal you actually need. Explained variance is a property of the input representation — it says nothing about your target. A component that explains 40% of the variance in your features might explain 0% of the variation in the label you care about.

Warning: Fitting scaling or PCA on data that includes validation or test rows leaks information and inflates apparent quality. Fit on training data only, then transform the rest.

Use PCA for compression, denoising, or visualization. Validate any downstream claim with a held-out evaluation. The math guarantees preserved variance. It does not guarantee preserved usefulness.

Verify the Derivation in a Few Lines

You can confirm the theory against computation without building a full pipeline. The idea is to check that the eigenvector of the covariance matrix really does maximize projected variance.

import numpy as np

rng = np.random.default_rng(0)
X = rng.normal(size=(200, 2)) @ np.array([[2.0, 1.0], [0.0, 1.0]])

Xc = X - X.mean(axis=0)
S = Xc.T @ Xc / len(Xc)

eigvals, eigvecs = np.linalg.eigh(S)
w1 = eigvecs[:, -1]  # largest eigenvalue

projected_variance = (Xc @ w1) ** 2
print(eigvals[-1])                 # eigenvalue
print(projected_variance.mean())   # projected variance

The two printed numbers match. That is the derivation made observable: the eigenvalue is the variance along its eigenvector. Rerun the same block with standardized features — divide each column by its standard deviation — and the eigenvectors will shift, tying the code directly back to the scaling discussion.

Keep this short. The code verifies the derivation. It does not replace it.

Where This Leaves You

The eigenvectors are not a library convention. They are the solution to a constrained variance problem, and that is exactly why preprocessing choices and the variance-versus-relevance distinction matter. Once you own the derivation, scaling stops being an arbitrary rule and becomes a decision about which matrix you are eigendecomposing.

Your next step: take a real dataset, run the same verification on it, and watch how standardizing changes the leading direction. Then move to the practical PCA experiment to see these effects on real data — scaling, component count, and the gap between retained variation and predictive value.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A component captures a large share of the input-feature variance. What can you conclude about its usefulness for predicting a target, based on the article?
Question 1 of 2Misconception Check

Focus: Distinguish variance preserved by PCA from relevance to a prediction target.

A learner skips centering before computing a matrix product for PCA. Which interpretation best matches the article's warning about this change?
Question 2 of 2Scenario Interpretation

Focus: Identify how skipping centering changes the matrix PCA analyzes and what the resulting directions may reflect.

References

  1. [1404.1100] A Tutorial on Principal Component Analysis - ar5ivar5iv.labs.arxiv.org
  2. PCA, SVD, and Centering of Dataarxiv.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.