The double descent phenomenon refers to a non-monotonic behavior of generalization error as model capacity increases—contradicting the classical bias–variance tradeoff from statistical learning theory.
1. Classical Regime: Bias–Variance Tradeoff
In traditional statistical learning theory, expected test error decomposes as:

- High capacity models → low bias, high variance
- Low capacity models → high bias, low variance
This predicts a U-shaped curve:
- Error decreases (better fit)
- Then increases (overfitting)
2. Interpolation Threshold
- When model capacity ≈ dataset size
- The model just fits the training data perfectly (zero training error)
This is called the interpolation threshold.
At this point:
- The solution is highly unstable
- Small perturbations → large parameter changes
- Variance spikes sharply
This causes the first peak in test error.
3. Overparameterized Regime (Second Descent)
As capacity increases beyond interpolation, something surprising happens:
- The model can fit the data in many different ways
- Optimization (e.g., SGD) implicitly selects low-complexity solutions
This leads to:
- Reduced variance
- Improved generalization
Test error decreases again → second descent.
4. Statistical Learning Theory Perspective
- Function Class Expansion
Let:

Before interpolation
- Approximation error ↓
- Estimation error ↑
- Classical regime holds.
At interpolation
- Estimation error blows up
- Model is maximally sensitive
Beyond interpolation
- Model selects minimum norm / margin-maximizing solutions
- Effective complexity ≠ parameter count
5. Final Shape of Risk Curve
Instead of U-shape, we get:



