The Decomposition Everyone Gets Wrong
The bias-variance tradeoff is usually presented as a vague heuristic: simple models underfit, complex models overfit, find the sweet spot. This misses the actual mathematics—which is both cleaner and more revealing than the cartoon.
1. Setup: What Are We Decomposing?
Fix an input . The target is where and is noise with , .
We train a model on a random training set . The expected prediction error at is:
The expectation is over both the noise in and the randomness in .
2. The Decomposition
Expand the squared loss:
Three terms:
-
Irreducible error : The noise in that no model can predict. This is —the exact quantity from the Law of Total Variance.
-
Bias²: How far the average prediction (over all possible training sets) deviates from the truth. This is systematic error—the model class cannot express .
-
Variance: How much the prediction fluctuates across training sets. This is sensitivity to the particular sample drawn.
3. The Geometric View
Recall from Conditional Expectation as Projection: in space, is the orthogonal projection of onto the subspace of functions of .
Now introduce a model class (linear functions, decision trees, neural networks). This is a further restriction—a subspace within the subspace of all measurable functions of .
The bias measures the distance from to the model subspace . A richer (higher-dimensional subspace) reduces this distance—at the cost of variance, because a higher-dimensional subspace is harder to estimate from finite data.
The tradeoff is geometric: the subspace that best approximates is not the one most reliably estimable from samples.
4. Concrete Example: Polynomial Regression
Fit degree- polynomials to data from with noise.
- (linear): High bias (a line cannot approximate a sine wave), low variance (only 2 parameters to estimate).
- : Low bias (the polynomial can fit the sine wave closely), high variance (16 parameters estimated from noisy data → wildly different fits for different samples).
- Optimal : Somewhere in between. The sweet spot depends on (more data → can afford higher ) and (more noise → need lower ).
The singular values of the design matrix (cf. SVD) reveal this directly: the -th singular value determines how well the -th polynomial component can be estimated. When is small relative to the noise level, that component’s estimate is dominated by variance.
5. Why This Tradeoff Is Fundamental
The bias-variance decomposition is not specific to any algorithm. It is a property of the estimation problem itself.
For any estimator :
- If is deterministic (e.g., ): variance is zero, bias can be large.
- If interpolates the training data exactly: bias is zero on training points, variance can be enormous.
- The Bayes-optimal predictor has zero bias, zero variance—but it requires knowing the true conditional distribution, which is exactly what we don’t have.
The irreducible error is the floor. No model, no matter how flexible, no matter how much data, can go below . This connects directly to the total variance decomposition: is the maximum achievable explained variance. The rest——is noise forever.
The Modern Wrinkle: Double Descent
Classical wisdom: error is U-shaped as model complexity increases (underfit → sweet spot → overfit). But modern neural networks often exhibit double descent: error decreases, rises, then decreases again as model size grows far past the interpolation threshold.
This does not violate the bias-variance decomposition—both terms still sum correctly. What changes is the implicit regularization: overparameterized models trained with gradient descent converge to the minimum-norm interpolant, which has surprising statistical properties. The variance term, which should explode at the interpolation threshold, gets tamed by the geometry of gradient descent.
The decomposition remains the right lens. The surprise is in how each term behaves, not whether they exist.