← Bias-Variance All techniques Next: TS Decomposition →

21 — Ridge vs Lasso Regularization

Both shrink coefficients toward zero — but Lasso (L1) can push them exactly to zero (built-in feature selection). Ridge (L2) shrinks them smoothly, never killing them.

Regularization α0.50
||β||₁ Ridge
||β||₁ Lasso
Lasso β=0 count
RIDGE (L2 penalty)
LASSO (L1 penalty)
0.50
How to choose: use Lasso when you want a sparse model — only a handful of features survive. Use Ridge when most features are weakly useful and you'd rather keep all of them but tamed. Elastic Net mixes both.