Every Matrix Tells the Same Story
The Singular Value Decomposition says that every matrix can be written as:
where is orthogonal, is orthogonal, and is diagonal with non-negative entries .
In words: every linear transformation is a rotation, followed by a scaling along coordinate axes, followed by another rotation. No matter how tangled the matrix looks, its action decomposes into these three clean steps.
1. The Geometry: Ellipsoids From Spheres
Apply to the unit sphere in . The image is an ellipsoid in .
- The directions of the ellipsoid’s axes are the left singular vectors (columns of ).
- The lengths of the semi-axes are the singular values .
- The orientations on the input sphere that map to these axes are the right singular vectors (columns of ).
The rank of is the number of nonzero singular values—the dimension of the ellipsoid. The condition number measures its eccentricity: a large condition number means the ellipsoid is extremely elongated, and the linear system is sensitive to perturbations along the thin directions.
2. Eckart-Young: Optimal Compression
Theorem: The best rank- approximation to in both Frobenius and operator norm is:
“Best” means minimizing over all matrices with . The SVD solves this by keeping the largest singular values and discarding the rest.
Why this matters: The fraction of “energy” captured by rank is:
If the singular values decay rapidly, a few components capture most of the structure. This is not a heuristic—it is the provably optimal low-rank approximation.
Connection to projection: just as conditional expectation projects onto the subspace generated by (keeping the “explained” component and discarding orthogonal noise), the truncated SVD projects the data matrix onto its -dimensional principal subspace.
3. PCA Is SVD in Disguise
Given a data matrix (rows = samples, columns = features), centered so each column has mean zero. The sample covariance is .
The eigendecomposition of gives the principal components. But , and if , then .
So the right singular vectors of are the principal components, and the singular values of determine the explained variance: .
In practice, you never form the covariance matrix. You compute the SVD of directly—which is numerically more stable and works even when .
4. The Four Fundamental Subspaces
The SVD reveals the complete anatomy of a linear map .
Let . The singular vectors split the domain and codomain into orthogonal complements:
is an isomorphism from Row to Col—it maps to . On Null, it annihilates. On Null, nothing lands. The SVD makes this perfectly explicit: provides the basis for domain, for codomain, and governs the stretching between them.
5. Pseudoinverse and Least Squares
When has no solution (overdetermined system), the least squares solution minimizes . The SVD gives it directly:
where inverts the nonzero singular values and zeros out the rest. This is the Moore-Penrose pseudoinverse.
The geometry: project onto Col (via ), then invert the map on the row space. If has components in Null, those are silently discarded. If the system is underdetermined, the pseudoinverse chooses the minimum-norm solution.
The pseudoinverse is the SVD’s answer to the question: “What is the best you can do with an imperfect linear system?”
Why SVD Is Universal
The SVD is not one algorithm—it is the structural theorem behind:
| Application | What SVD reveals |
|---|---|
| PCA | Principal directions of variance |
| Least Squares | Minimum-norm projection |
| Low-rank approximation | Optimal compression (Eckart-Young) |
| Condition number | Sensitivity to perturbation |
| Pseudoinverse | Best “undo” for a non-invertible map |
| Matrix norms | , |
Any time you have a matrix and need to understand its geometry—what it stretches, what it kills, and what subspace captures the action—the SVD is the answer.