rewriting slides from week 35
This commit is contained in:
@@ -200,24 +200,6 @@ Automatically generated HTML file from DocOnce source
|
||||
2,
|
||||
None,
|
||||
'meet-the-covariance-matrix'),
|
||||
('Ridge and LASSO Regression',
|
||||
2,
|
||||
None,
|
||||
'ridge-and-lasso-regression'),
|
||||
('More on Ridge Regression', 2, None, 'more-on-ridge-regression'),
|
||||
('Interpreting the Ridge results',
|
||||
2,
|
||||
None,
|
||||
'interpreting-the-ridge-results'),
|
||||
('More interpretations', 2, None, 'more-interpretations'),
|
||||
('A better understanding of regularization',
|
||||
2,
|
||||
None,
|
||||
'a-better-understanding-of-regularization'),
|
||||
('Decomposing the OLS and Ridge expressions',
|
||||
2,
|
||||
None,
|
||||
'decomposing-the-ols-and-ridge-expressions'),
|
||||
('Introducing the Covariance and Correlation functions',
|
||||
2,
|
||||
None,
|
||||
@@ -243,6 +225,24 @@ Automatically generated HTML file from DocOnce source
|
||||
2,
|
||||
None,
|
||||
'rewriting-the-covariance-and-or-correlation-matrix'),
|
||||
('Ridge and LASSO Regression',
|
||||
2,
|
||||
None,
|
||||
'ridge-and-lasso-regression'),
|
||||
('More on Ridge Regression', 2, None, 'more-on-ridge-regression'),
|
||||
('Interpreting the Ridge results',
|
||||
2,
|
||||
None,
|
||||
'interpreting-the-ridge-results'),
|
||||
('More interpretations', 2, None, 'more-interpretations'),
|
||||
('A better understanding of regularization',
|
||||
2,
|
||||
None,
|
||||
'a-better-understanding-of-regularization'),
|
||||
('Decomposing the OLS and Ridge expressions',
|
||||
2,
|
||||
None,
|
||||
'decomposing-the-ols-and-ridge-expressions'),
|
||||
('Mathematical Properties', 2, None, 'mathematical-properties'),
|
||||
('Exercises for week 36, September 6-10',
|
||||
2,
|
||||
@@ -345,19 +345,19 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs050.html#setting-up-the-matrix-to-be-inverted" style="font-size: 80%;"><b>Setting up the Matrix to be inverted</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs051.html#further-properties-important-for-our-analyses-later" style="font-size: 80%;"><b>Further properties (important for our analyses later)</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs052.html#meet-the-covariance-matrix" style="font-size: 80%;"><b>Meet the Covariance Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs053.html#ridge-and-lasso-regression" style="font-size: 80%;"><b>Ridge and LASSO Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs054.html#more-on-ridge-regression" style="font-size: 80%;"><b>More on Ridge Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs055.html#interpreting-the-ridge-results" style="font-size: 80%;"><b>Interpreting the Ridge results</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs056.html#more-interpretations" style="font-size: 80%;"><b>More interpretations</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs057.html#a-better-understanding-of-regularization" style="font-size: 80%;"><b>A better understanding of regularization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs058.html#decomposing-the-ols-and-ridge-expressions" style="font-size: 80%;"><b>Decomposing the OLS and Ridge expressions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs059.html#introducing-the-covariance-and-correlation-functions" style="font-size: 80%;"><b>Introducing the Covariance and Correlation functions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs060.html#correlation-function-and-design-feature-matrix" style="font-size: 80%;"><b>Correlation Function and Design/Feature Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs061.html#covariance-matrix-examples" style="font-size: 80%;"><b>Covariance Matrix Examples</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs062.html#correlation-matrix" style="font-size: 80%;"><b>Correlation Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs063.html#correlation-matrix-with-pandas" style="font-size: 80%;"><b>Correlation Matrix with Pandas</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs064.html#correlation-matrix-with-pandas-and-the-franke-function" style="font-size: 80%;"><b>Correlation Matrix with Pandas and the Franke function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs065.html#rewriting-the-covariance-and-or-correlation-matrix" style="font-size: 80%;"><b>Rewriting the Covariance and/or Correlation Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs053.html#introducing-the-covariance-and-correlation-functions" style="font-size: 80%;"><b>Introducing the Covariance and Correlation functions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs054.html#correlation-function-and-design-feature-matrix" style="font-size: 80%;"><b>Correlation Function and Design/Feature Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs055.html#covariance-matrix-examples" style="font-size: 80%;"><b>Covariance Matrix Examples</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs056.html#correlation-matrix" style="font-size: 80%;"><b>Correlation Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs057.html#correlation-matrix-with-pandas" style="font-size: 80%;"><b>Correlation Matrix with Pandas</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs058.html#correlation-matrix-with-pandas-and-the-franke-function" style="font-size: 80%;"><b>Correlation Matrix with Pandas and the Franke function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs059.html#rewriting-the-covariance-and-or-correlation-matrix" style="font-size: 80%;"><b>Rewriting the Covariance and/or Correlation Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs060.html#ridge-and-lasso-regression" style="font-size: 80%;"><b>Ridge and LASSO Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs061.html#more-on-ridge-regression" style="font-size: 80%;"><b>More on Ridge Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs062.html#interpreting-the-ridge-results" style="font-size: 80%;"><b>Interpreting the Ridge results</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs063.html#more-interpretations" style="font-size: 80%;"><b>More interpretations</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs064.html#a-better-understanding-of-regularization" style="font-size: 80%;"><b>A better understanding of regularization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs065.html#decomposing-the-ols-and-ridge-expressions" style="font-size: 80%;"><b>Decomposing the OLS and Ridge expressions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs066.html#mathematical-properties" style="font-size: 80%;"><b>Mathematical Properties</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs067.html#exercises-for-week-36-september-6-10" style="font-size: 80%;"><b>Exercises for week 36, September 6-10</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs067.html#exercise-1-adding-ridge-and-lasso-regression" style="font-size: 80%;"><b>Exercise 1: Adding Ridge and Lasso Regression</b></a></li>
|
||||
|
||||
@@ -2283,265 +2283,12 @@ terms of the singular values.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="ridge-and-lasso-regression">Ridge and LASSO Regression </h2>
|
||||
|
||||
<p>
|
||||
Let us remind ourselves about the expression for the standard Mean Squared Error (MSE) which we used to define our cost function and the equations for the ordinary least squares (OLS) method, that is
|
||||
our optimization problem is
|
||||
<p> <br>
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in {\mathbb{R}}^{p}}}\frac{1}{n}\left\{\left(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\right)^T\left(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\right)\right\}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
or we can state it as
|
||||
<p> <br>
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\sum_{i=0}^{n-1}\left(y_i-\tilde{y}_i\right)^2=\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
where we have used the definition of a norm-2 vector, that is
|
||||
<p> <br>
|
||||
$$
|
||||
\vert\vert \boldsymbol{x}\vert\vert_2 = \sqrt{\sum_i x_i^2}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
By minimizing the above equation with respect to the parameters
|
||||
\( \boldsymbol{\beta} \) we could then obtain an analytical expression for the
|
||||
parameters \( \boldsymbol{\beta} \). We can add a regularization parameter \( \lambda \) by
|
||||
defining a new cost function to be optimized, that is
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_2^2
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
which leads to the Ridge regression minimization problem where we
|
||||
require that \( \vert\vert \boldsymbol{\beta}\vert\vert_2^2\le t \), where \( t \) is
|
||||
a finite number larger than zero. By defining
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_1,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
we have a new optimization equation
|
||||
<p> <br>
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_1
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
which leads to Lasso regression. Lasso stands for least absolute shrinkage and selection operator.
|
||||
|
||||
<p>
|
||||
Here we have defined the norm-1 as
|
||||
<p> <br>
|
||||
$$
|
||||
\vert\vert \boldsymbol{x}\vert\vert_1 = \sum_i \vert x_i\vert.
|
||||
$$
|
||||
<p> <br>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="more-on-ridge-regression">More on Ridge Regression </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
by taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of \( \lambda \) does not suffer from singularity problems. We obtain
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{Ridge}} = \left(\boldsymbol{X}^T\boldsymbol{X}+\lambda\boldsymbol{I}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
with \( \boldsymbol{I} \) being a \( p\times p \) identity matrix with the constraint that
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\sum_{i=0}^{p-1} \beta_i^2 \leq t,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
with \( t \) a finite positive number.
|
||||
|
||||
<p>
|
||||
We see that Ridge regression is nothing but the standard
|
||||
OLS with a modified diagonal term added to \( \boldsymbol{X}^T\boldsymbol{X} \). The
|
||||
consequences, in particular for our discussion of the bias-variance tradeoff
|
||||
are rather interesting.
|
||||
|
||||
<p>
|
||||
Furthermore, if we use the result above in terms of the SVD decomposition (our analysis was done for the OLS method), we had
|
||||
<p> <br>
|
||||
$$
|
||||
(\boldsymbol{X}\boldsymbol{X}^T)\boldsymbol{U} = \boldsymbol{U}\boldsymbol{D}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
We can analyse the OLS solutions in terms of the eigenvectors (the columns) of the right singular value matrix \( \boldsymbol{U} \) as
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{X}\boldsymbol{\beta} = \boldsymbol{X}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T \right)^{-1}\boldsymbol{X}^T\boldsymbol{y}=\boldsymbol{U\Sigma V^T}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T \right)^{-1}(\boldsymbol{U\Sigma V^T})^T\boldsymbol{y}=\boldsymbol{U}\boldsymbol{U}^T\boldsymbol{y}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
For Ridge regression this becomes
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{X}\boldsymbol{\beta}^{\mathrm{Ridge}} = \boldsymbol{U\Sigma V^T}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T+\lambda\boldsymbol{I} \right)^{-1}(\boldsymbol{U\Sigma V^T})^T\boldsymbol{y}=\sum_{j=0}^{p-1}\boldsymbol{u}_j\boldsymbol{u}_j^T\frac{\sigma_j^2}{\sigma_j^2+\lambda}\boldsymbol{y},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
with the vectors \( \boldsymbol{u}_j \) being the columns of \( \boldsymbol{U} \).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="interpreting-the-ridge-results">Interpreting the Ridge results </h2>
|
||||
|
||||
<p>
|
||||
Since \( \lambda \geq 0 \), it means that compared to OLS, we have
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\frac{\sigma_j^2}{\sigma_j^2+\lambda} \leq 1.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
Ridge regression finds the coordinates of \( \boldsymbol{y} \) with respect to the
|
||||
orthonormal basis \( \boldsymbol{U} \), it then shrinks the coordinates by
|
||||
\( \frac{\sigma_j^2}{\sigma_j^2+\lambda} \). Recall that the SVD has
|
||||
eigenvalues ordered in a descending way, that is \( \sigma_i \geq
|
||||
\sigma_{i+1} \).
|
||||
|
||||
<p>
|
||||
For small eigenvalues \( \sigma_i \) it means that their contributions become less important, a fact which can be used to reduce the number of degrees of freedom.
|
||||
Actually, calculating the variance of \( \boldsymbol{X}\boldsymbol{v}_j \) shows that this quantity is equal to \( \sigma_j^2/n \).
|
||||
With a parameter \( \lambda \) we can thus shrink the role of specific parameters.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="more-interpretations">More interpretations </h2>
|
||||
|
||||
<p>
|
||||
For the sake of simplicity, let us assume that the design matrix is orthonormal, that is
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}=(\boldsymbol{X}^T\boldsymbol{X})^{-1} =\boldsymbol{I}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
In this case the standard OLS results in
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{OLS}} = \boldsymbol{X}^T\boldsymbol{y}=\sum_{i=0}^{p-1}\boldsymbol{u}_j\boldsymbol{u}_j^T\boldsymbol{y},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
and
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{Ridge}} = \left(\boldsymbol{I}+\lambda\boldsymbol{I}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y}=\left(1+\lambda\right)^{-1}\boldsymbol{\beta}^{\mathrm{OLS}},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
that is the Ridge estimator scales the OLS estimator by the inverse of a factor \( 1+\lambda \), and
|
||||
the Ridge estimator converges to zero when the hyperparameter goes to
|
||||
infinity.
|
||||
|
||||
<p>
|
||||
We will come back to more interpreations after we have gone through some of the statistical analysis part.
|
||||
|
||||
<p>
|
||||
For more discussions of Ridge and Lasso regression, <a href="https://arxiv.org/abs/1509.09169" target="_blank">Wessel van Wieringen's</a> article is highly recommended.
|
||||
Similarly, <a href="https://arxiv.org/abs/1803.08823" target="_blank">Mehta et al's article</a> is also recommended.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="a-better-understanding-of-regularization">A better understanding of regularization </h2>
|
||||
|
||||
<p>
|
||||
The parameter \( \lambda \) that we have introduced in the Ridge (and
|
||||
Lasso as well) regression is often called a regularization parameter
|
||||
or shrinkage parameter. It is common to call it a hyperparameter. What does it mean mathemtically?
|
||||
|
||||
<p>
|
||||
Here we will first look at how to analyze the difference between the
|
||||
standard OLS equations and the Ridge expressions in terms of a linear
|
||||
algebra analysis using the SVD algorithm. Thereafter, we will link
|
||||
(see the material on the bias-variance tradeoff below) these
|
||||
observation to the statisical analysis of the results. In particular
|
||||
we consider how the variance of the parameters \( \boldsymbol{\beta} \) is
|
||||
affected by changing the parameter \( \lambda \).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="decomposing-the-ols-and-ridge-expressions">Decomposing the OLS and Ridge expressions </h2>
|
||||
|
||||
<p>
|
||||
We have our design matrix
|
||||
\( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \). With the SVD we decompose it as
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{X} = \boldsymbol{U\Sigma V^T},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
with \( \boldsymbol{U}\in {\mathbb{R}}^{n\times n} \), \( \boldsymbol{\Sigma}\in {\mathbb{R}}^{n\times p} \)
|
||||
and \( \boldsymbol{V}\in {\mathbb{R}}^{p\times p} \).
|
||||
|
||||
<p>
|
||||
The matrices \( \boldsymbol{U} \) and \( \boldsymbol{V} \) are unitary/orthonormal matrices, that is in case the matrices are real we have \( \boldsymbol{U}^T\boldsymbol{U}=\boldsymbol{U}\boldsymbol{U}^T=\boldsymbol{I} \) and \( \boldsymbol{V}^T\boldsymbol{V}=\boldsymbol{V}\boldsymbol{V}^T=\boldsymbol{I} \).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="introducing-the-covariance-and-correlation-functions">Introducing the Covariance and Correlation functions </h2>
|
||||
|
||||
<p>
|
||||
Before we discuss the link between for example Ridge regression and the singular value decomposition, we need to remind ourselves about
|
||||
the definition of the covariance and the correlation function. These are quantities
|
||||
the definition of the covariance and the correlation function. These are quantities that play a central role in machine learning methods.
|
||||
|
||||
<p>
|
||||
Suppose we have defined two vectors
|
||||
@@ -2914,6 +2661,259 @@ It is easy to generalize this to a matrix \( \boldsymbol{X}\in {\mathbb{R}}^{n\t
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="ridge-and-lasso-regression">Ridge and LASSO Regression </h2>
|
||||
|
||||
<p>
|
||||
Let us remind ourselves about the expression for the standard Mean Squared Error (MSE) which we used to define our cost function and the equations for the ordinary least squares (OLS) method, that is
|
||||
our optimization problem is
|
||||
<p> <br>
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in {\mathbb{R}}^{p}}}\frac{1}{n}\left\{\left(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\right)^T\left(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\right)\right\}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
or we can state it as
|
||||
<p> <br>
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\sum_{i=0}^{n-1}\left(y_i-\tilde{y}_i\right)^2=\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
where we have used the definition of a norm-2 vector, that is
|
||||
<p> <br>
|
||||
$$
|
||||
\vert\vert \boldsymbol{x}\vert\vert_2 = \sqrt{\sum_i x_i^2}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
By minimizing the above equation with respect to the parameters
|
||||
\( \boldsymbol{\beta} \) we could then obtain an analytical expression for the
|
||||
parameters \( \boldsymbol{\beta} \). We can add a regularization parameter \( \lambda \) by
|
||||
defining a new cost function to be optimized, that is
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_2^2
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
which leads to the Ridge regression minimization problem where we
|
||||
require that \( \vert\vert \boldsymbol{\beta}\vert\vert_2^2\le t \), where \( t \) is
|
||||
a finite number larger than zero. By defining
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_1,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
we have a new optimization equation
|
||||
<p> <br>
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_1
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
which leads to Lasso regression. Lasso stands for least absolute shrinkage and selection operator.
|
||||
|
||||
<p>
|
||||
Here we have defined the norm-1 as
|
||||
<p> <br>
|
||||
$$
|
||||
\vert\vert \boldsymbol{x}\vert\vert_1 = \sum_i \vert x_i\vert.
|
||||
$$
|
||||
<p> <br>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="more-on-ridge-regression">More on Ridge Regression </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
by taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of \( \lambda \) does not suffer from singularity problems. We obtain
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{Ridge}} = \left(\boldsymbol{X}^T\boldsymbol{X}+\lambda\boldsymbol{I}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
with \( \boldsymbol{I} \) being a \( p\times p \) identity matrix with the constraint that
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\sum_{i=0}^{p-1} \beta_i^2 \leq t,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
with \( t \) a finite positive number.
|
||||
|
||||
<p>
|
||||
We see that Ridge regression is nothing but the standard
|
||||
OLS with a modified diagonal term added to \( \boldsymbol{X}^T\boldsymbol{X} \). The
|
||||
consequences, in particular for our discussion of the bias-variance tradeoff
|
||||
are rather interesting.
|
||||
|
||||
<p>
|
||||
Furthermore, if we use the result above in terms of the SVD decomposition (our analysis was done for the OLS method), we had
|
||||
<p> <br>
|
||||
$$
|
||||
(\boldsymbol{X}\boldsymbol{X}^T)\boldsymbol{U} = \boldsymbol{U}\boldsymbol{D}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
We can analyse the OLS solutions in terms of the eigenvectors (the columns) of the right singular value matrix \( \boldsymbol{U} \) as
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{X}\boldsymbol{\beta} = \boldsymbol{X}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T \right)^{-1}\boldsymbol{X}^T\boldsymbol{y}=\boldsymbol{U\Sigma V^T}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T \right)^{-1}(\boldsymbol{U\Sigma V^T})^T\boldsymbol{y}=\boldsymbol{U}\boldsymbol{U}^T\boldsymbol{y}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
For Ridge regression this becomes
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{X}\boldsymbol{\beta}^{\mathrm{Ridge}} = \boldsymbol{U\Sigma V^T}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T+\lambda\boldsymbol{I} \right)^{-1}(\boldsymbol{U\Sigma V^T})^T\boldsymbol{y}=\sum_{j=0}^{p-1}\boldsymbol{u}_j\boldsymbol{u}_j^T\frac{\sigma_j^2}{\sigma_j^2+\lambda}\boldsymbol{y},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
with the vectors \( \boldsymbol{u}_j \) being the columns of \( \boldsymbol{U} \).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="interpreting-the-ridge-results">Interpreting the Ridge results </h2>
|
||||
|
||||
<p>
|
||||
Since \( \lambda \geq 0 \), it means that compared to OLS, we have
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\frac{\sigma_j^2}{\sigma_j^2+\lambda} \leq 1.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
Ridge regression finds the coordinates of \( \boldsymbol{y} \) with respect to the
|
||||
orthonormal basis \( \boldsymbol{U} \), it then shrinks the coordinates by
|
||||
\( \frac{\sigma_j^2}{\sigma_j^2+\lambda} \). Recall that the SVD has
|
||||
eigenvalues ordered in a descending way, that is \( \sigma_i \geq
|
||||
\sigma_{i+1} \).
|
||||
|
||||
<p>
|
||||
For small eigenvalues \( \sigma_i \) it means that their contributions become less important, a fact which can be used to reduce the number of degrees of freedom.
|
||||
Actually, calculating the variance of \( \boldsymbol{X}\boldsymbol{v}_j \) shows that this quantity is equal to \( \sigma_j^2/n \).
|
||||
With a parameter \( \lambda \) we can thus shrink the role of specific parameters.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="more-interpretations">More interpretations </h2>
|
||||
|
||||
<p>
|
||||
For the sake of simplicity, let us assume that the design matrix is orthonormal, that is
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}=(\boldsymbol{X}^T\boldsymbol{X})^{-1} =\boldsymbol{I}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
In this case the standard OLS results in
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{OLS}} = \boldsymbol{X}^T\boldsymbol{y}=\sum_{i=0}^{p-1}\boldsymbol{u}_j\boldsymbol{u}_j^T\boldsymbol{y},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
and
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{Ridge}} = \left(\boldsymbol{I}+\lambda\boldsymbol{I}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y}=\left(1+\lambda\right)^{-1}\boldsymbol{\beta}^{\mathrm{OLS}},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
that is the Ridge estimator scales the OLS estimator by the inverse of a factor \( 1+\lambda \), and
|
||||
the Ridge estimator converges to zero when the hyperparameter goes to
|
||||
infinity.
|
||||
|
||||
<p>
|
||||
We will come back to more interpreations after we have gone through some of the statistical analysis part.
|
||||
|
||||
<p>
|
||||
For more discussions of Ridge and Lasso regression, <a href="https://arxiv.org/abs/1509.09169" target="_blank">Wessel van Wieringen's</a> article is highly recommended.
|
||||
Similarly, <a href="https://arxiv.org/abs/1803.08823" target="_blank">Mehta et al's article</a> is also recommended.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="a-better-understanding-of-regularization">A better understanding of regularization </h2>
|
||||
|
||||
<p>
|
||||
The parameter \( \lambda \) that we have introduced in the Ridge (and
|
||||
Lasso as well) regression is often called a regularization parameter
|
||||
or shrinkage parameter. It is common to call it a hyperparameter. What does it mean mathemtically?
|
||||
|
||||
<p>
|
||||
Here we will first look at how to analyze the difference between the
|
||||
standard OLS equations and the Ridge expressions in terms of a linear
|
||||
algebra analysis using the SVD algorithm. Thereafter, we will link
|
||||
(see the material on the bias-variance tradeoff below) these
|
||||
observation to the statisical analysis of the results. In particular
|
||||
we consider how the variance of the parameters \( \boldsymbol{\beta} \) is
|
||||
affected by changing the parameter \( \lambda \).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="decomposing-the-ols-and-ridge-expressions">Decomposing the OLS and Ridge expressions </h2>
|
||||
|
||||
<p>
|
||||
We have our design matrix
|
||||
\( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \). With the SVD we decompose it as
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{X} = \boldsymbol{U\Sigma V^T},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
with \( \boldsymbol{U}\in {\mathbb{R}}^{n\times n} \), \( \boldsymbol{\Sigma}\in {\mathbb{R}}^{n\times p} \)
|
||||
and \( \boldsymbol{V}\in {\mathbb{R}}^{p\times p} \).
|
||||
|
||||
<p>
|
||||
The matrices \( \boldsymbol{U} \) and \( \boldsymbol{V} \) are unitary/orthonormal matrices, that is in case the matrices are real we have \( \boldsymbol{U}^T\boldsymbol{U}=\boldsymbol{U}\boldsymbol{U}^T=\boldsymbol{I} \) and \( \boldsymbol{V}^T\boldsymbol{V}=\boldsymbol{V}\boldsymbol{V}^T=\boldsymbol{I} \).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="mathematical-properties">Mathematical Properties </h2>
|
||||
|
||||
|
||||
@@ -220,24 +220,6 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
2,
|
||||
None,
|
||||
'meet-the-covariance-matrix'),
|
||||
('Ridge and LASSO Regression',
|
||||
2,
|
||||
None,
|
||||
'ridge-and-lasso-regression'),
|
||||
('More on Ridge Regression', 2, None, 'more-on-ridge-regression'),
|
||||
('Interpreting the Ridge results',
|
||||
2,
|
||||
None,
|
||||
'interpreting-the-ridge-results'),
|
||||
('More interpretations', 2, None, 'more-interpretations'),
|
||||
('A better understanding of regularization',
|
||||
2,
|
||||
None,
|
||||
'a-better-understanding-of-regularization'),
|
||||
('Decomposing the OLS and Ridge expressions',
|
||||
2,
|
||||
None,
|
||||
'decomposing-the-ols-and-ridge-expressions'),
|
||||
('Introducing the Covariance and Correlation functions',
|
||||
2,
|
||||
None,
|
||||
@@ -263,6 +245,24 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
2,
|
||||
None,
|
||||
'rewriting-the-covariance-and-or-correlation-matrix'),
|
||||
('Ridge and LASSO Regression',
|
||||
2,
|
||||
None,
|
||||
'ridge-and-lasso-regression'),
|
||||
('More on Ridge Regression', 2, None, 'more-on-ridge-regression'),
|
||||
('Interpreting the Ridge results',
|
||||
2,
|
||||
None,
|
||||
'interpreting-the-ridge-results'),
|
||||
('More interpretations', 2, None, 'more-interpretations'),
|
||||
('A better understanding of regularization',
|
||||
2,
|
||||
None,
|
||||
'a-better-understanding-of-regularization'),
|
||||
('Decomposing the OLS and Ridge expressions',
|
||||
2,
|
||||
None,
|
||||
'decomposing-the-ols-and-ridge-expressions'),
|
||||
('Mathematical Properties', 2, None, 'mathematical-properties'),
|
||||
('Exercises for week 36, September 6-10',
|
||||
2,
|
||||
@@ -2315,228 +2315,11 @@ terms of the singular values.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="ridge-and-lasso-regression">Ridge and LASSO Regression </h2>
|
||||
|
||||
<p>
|
||||
Let us remind ourselves about the expression for the standard Mean Squared Error (MSE) which we used to define our cost function and the equations for the ordinary least squares (OLS) method, that is
|
||||
our optimization problem is
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in {\mathbb{R}}^{p}}}\frac{1}{n}\left\{\left(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\right)^T\left(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\right)\right\}.
|
||||
$$
|
||||
|
||||
or we can state it as
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\sum_{i=0}^{n-1}\left(y_i-\tilde{y}_i\right)^2=\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2,
|
||||
$$
|
||||
|
||||
where we have used the definition of a norm-2 vector, that is
|
||||
$$
|
||||
\vert\vert \boldsymbol{x}\vert\vert_2 = \sqrt{\sum_i x_i^2}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
By minimizing the above equation with respect to the parameters
|
||||
\( \boldsymbol{\beta} \) we could then obtain an analytical expression for the
|
||||
parameters \( \boldsymbol{\beta} \). We can add a regularization parameter \( \lambda \) by
|
||||
defining a new cost function to be optimized, that is
|
||||
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_2^2
|
||||
$$
|
||||
|
||||
<p>
|
||||
which leads to the Ridge regression minimization problem where we
|
||||
require that \( \vert\vert \boldsymbol{\beta}\vert\vert_2^2\le t \), where \( t \) is
|
||||
a finite number larger than zero. By defining
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_1,
|
||||
$$
|
||||
|
||||
<p>
|
||||
we have a new optimization equation
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_1
|
||||
$$
|
||||
|
||||
which leads to Lasso regression. Lasso stands for least absolute shrinkage and selection operator.
|
||||
|
||||
<p>
|
||||
Here we have defined the norm-1 as
|
||||
$$
|
||||
\vert\vert \boldsymbol{x}\vert\vert_1 = \sum_i \vert x_i\vert.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="more-on-ridge-regression">More on Ridge Regression </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
$$
|
||||
|
||||
<p>
|
||||
by taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of \( \lambda \) does not suffer from singularity problems. We obtain
|
||||
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{Ridge}} = \left(\boldsymbol{X}^T\boldsymbol{X}+\lambda\boldsymbol{I}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y},
|
||||
$$
|
||||
|
||||
<p>
|
||||
with \( \boldsymbol{I} \) being a \( p\times p \) identity matrix with the constraint that
|
||||
|
||||
$$
|
||||
\sum_{i=0}^{p-1} \beta_i^2 \leq t,
|
||||
$$
|
||||
|
||||
<p>
|
||||
with \( t \) a finite positive number.
|
||||
|
||||
<p>
|
||||
We see that Ridge regression is nothing but the standard
|
||||
OLS with a modified diagonal term added to \( \boldsymbol{X}^T\boldsymbol{X} \). The
|
||||
consequences, in particular for our discussion of the bias-variance tradeoff
|
||||
are rather interesting.
|
||||
|
||||
<p>
|
||||
Furthermore, if we use the result above in terms of the SVD decomposition (our analysis was done for the OLS method), we had
|
||||
$$
|
||||
(\boldsymbol{X}\boldsymbol{X}^T)\boldsymbol{U} = \boldsymbol{U}\boldsymbol{D}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
We can analyse the OLS solutions in terms of the eigenvectors (the columns) of the right singular value matrix \( \boldsymbol{U} \) as
|
||||
$$
|
||||
\boldsymbol{X}\boldsymbol{\beta} = \boldsymbol{X}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T \right)^{-1}\boldsymbol{X}^T\boldsymbol{y}=\boldsymbol{U\Sigma V^T}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T \right)^{-1}(\boldsymbol{U\Sigma V^T})^T\boldsymbol{y}=\boldsymbol{U}\boldsymbol{U}^T\boldsymbol{y}
|
||||
$$
|
||||
|
||||
<p>
|
||||
For Ridge regression this becomes
|
||||
|
||||
$$
|
||||
\boldsymbol{X}\boldsymbol{\beta}^{\mathrm{Ridge}} = \boldsymbol{U\Sigma V^T}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T+\lambda\boldsymbol{I} \right)^{-1}(\boldsymbol{U\Sigma V^T})^T\boldsymbol{y}=\sum_{j=0}^{p-1}\boldsymbol{u}_j\boldsymbol{u}_j^T\frac{\sigma_j^2}{\sigma_j^2+\lambda}\boldsymbol{y},
|
||||
$$
|
||||
|
||||
<p>
|
||||
with the vectors \( \boldsymbol{u}_j \) being the columns of \( \boldsymbol{U} \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="interpreting-the-ridge-results">Interpreting the Ridge results </h2>
|
||||
|
||||
<p>
|
||||
Since \( \lambda \geq 0 \), it means that compared to OLS, we have
|
||||
|
||||
$$
|
||||
\frac{\sigma_j^2}{\sigma_j^2+\lambda} \leq 1.
|
||||
$$
|
||||
|
||||
<p>
|
||||
Ridge regression finds the coordinates of \( \boldsymbol{y} \) with respect to the
|
||||
orthonormal basis \( \boldsymbol{U} \), it then shrinks the coordinates by
|
||||
\( \frac{\sigma_j^2}{\sigma_j^2+\lambda} \). Recall that the SVD has
|
||||
eigenvalues ordered in a descending way, that is \( \sigma_i \geq
|
||||
\sigma_{i+1} \).
|
||||
|
||||
<p>
|
||||
For small eigenvalues \( \sigma_i \) it means that their contributions become less important, a fact which can be used to reduce the number of degrees of freedom.
|
||||
Actually, calculating the variance of \( \boldsymbol{X}\boldsymbol{v}_j \) shows that this quantity is equal to \( \sigma_j^2/n \).
|
||||
With a parameter \( \lambda \) we can thus shrink the role of specific parameters.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="more-interpretations">More interpretations </h2>
|
||||
|
||||
<p>
|
||||
For the sake of simplicity, let us assume that the design matrix is orthonormal, that is
|
||||
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}=(\boldsymbol{X}^T\boldsymbol{X})^{-1} =\boldsymbol{I}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
In this case the standard OLS results in
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{OLS}} = \boldsymbol{X}^T\boldsymbol{y}=\sum_{i=0}^{p-1}\boldsymbol{u}_j\boldsymbol{u}_j^T\boldsymbol{y},
|
||||
$$
|
||||
|
||||
<p>
|
||||
and
|
||||
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{Ridge}} = \left(\boldsymbol{I}+\lambda\boldsymbol{I}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y}=\left(1+\lambda\right)^{-1}\boldsymbol{\beta}^{\mathrm{OLS}},
|
||||
$$
|
||||
|
||||
<p>
|
||||
that is the Ridge estimator scales the OLS estimator by the inverse of a factor \( 1+\lambda \), and
|
||||
the Ridge estimator converges to zero when the hyperparameter goes to
|
||||
infinity.
|
||||
|
||||
<p>
|
||||
We will come back to more interpreations after we have gone through some of the statistical analysis part.
|
||||
|
||||
<p>
|
||||
For more discussions of Ridge and Lasso regression, <a href="https://arxiv.org/abs/1509.09169" target="_blank">Wessel van Wieringen's</a> article is highly recommended.
|
||||
Similarly, <a href="https://arxiv.org/abs/1803.08823" target="_blank">Mehta et al's article</a> is also recommended.
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="a-better-understanding-of-regularization">A better understanding of regularization </h2>
|
||||
|
||||
<p>
|
||||
The parameter \( \lambda \) that we have introduced in the Ridge (and
|
||||
Lasso as well) regression is often called a regularization parameter
|
||||
or shrinkage parameter. It is common to call it a hyperparameter. What does it mean mathemtically?
|
||||
|
||||
<p>
|
||||
Here we will first look at how to analyze the difference between the
|
||||
standard OLS equations and the Ridge expressions in terms of a linear
|
||||
algebra analysis using the SVD algorithm. Thereafter, we will link
|
||||
(see the material on the bias-variance tradeoff below) these
|
||||
observation to the statisical analysis of the results. In particular
|
||||
we consider how the variance of the parameters \( \boldsymbol{\beta} \) is
|
||||
affected by changing the parameter \( \lambda \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="decomposing-the-ols-and-ridge-expressions">Decomposing the OLS and Ridge expressions </h2>
|
||||
|
||||
<p>
|
||||
We have our design matrix
|
||||
\( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \). With the SVD we decompose it as
|
||||
|
||||
$$
|
||||
\boldsymbol{X} = \boldsymbol{U\Sigma V^T},
|
||||
$$
|
||||
|
||||
<p>
|
||||
with \( \boldsymbol{U}\in {\mathbb{R}}^{n\times n} \), \( \boldsymbol{\Sigma}\in {\mathbb{R}}^{n\times p} \)
|
||||
and \( \boldsymbol{V}\in {\mathbb{R}}^{p\times p} \).
|
||||
|
||||
<p>
|
||||
The matrices \( \boldsymbol{U} \) and \( \boldsymbol{V} \) are unitary/orthonormal matrices, that is in case the matrices are real we have \( \boldsymbol{U}^T\boldsymbol{U}=\boldsymbol{U}\boldsymbol{U}^T=\boldsymbol{I} \) and \( \boldsymbol{V}^T\boldsymbol{V}=\boldsymbol{V}\boldsymbol{V}^T=\boldsymbol{I} \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="introducing-the-covariance-and-correlation-functions">Introducing the Covariance and Correlation functions </h2>
|
||||
|
||||
<p>
|
||||
Before we discuss the link between for example Ridge regression and the singular value decomposition, we need to remind ourselves about
|
||||
the definition of the covariance and the correlation function. These are quantities
|
||||
the definition of the covariance and the correlation function. These are quantities that play a central role in machine learning methods.
|
||||
|
||||
<p>
|
||||
Suppose we have defined two vectors
|
||||
@@ -2875,6 +2658,223 @@ It is easy to generalize this to a matrix \( \boldsymbol{X}\in {\mathbb{R}}^{n\t
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="ridge-and-lasso-regression">Ridge and LASSO Regression </h2>
|
||||
|
||||
<p>
|
||||
Let us remind ourselves about the expression for the standard Mean Squared Error (MSE) which we used to define our cost function and the equations for the ordinary least squares (OLS) method, that is
|
||||
our optimization problem is
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in {\mathbb{R}}^{p}}}\frac{1}{n}\left\{\left(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\right)^T\left(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\right)\right\}.
|
||||
$$
|
||||
|
||||
or we can state it as
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\sum_{i=0}^{n-1}\left(y_i-\tilde{y}_i\right)^2=\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2,
|
||||
$$
|
||||
|
||||
where we have used the definition of a norm-2 vector, that is
|
||||
$$
|
||||
\vert\vert \boldsymbol{x}\vert\vert_2 = \sqrt{\sum_i x_i^2}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
By minimizing the above equation with respect to the parameters
|
||||
\( \boldsymbol{\beta} \) we could then obtain an analytical expression for the
|
||||
parameters \( \boldsymbol{\beta} \). We can add a regularization parameter \( \lambda \) by
|
||||
defining a new cost function to be optimized, that is
|
||||
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_2^2
|
||||
$$
|
||||
|
||||
<p>
|
||||
which leads to the Ridge regression minimization problem where we
|
||||
require that \( \vert\vert \boldsymbol{\beta}\vert\vert_2^2\le t \), where \( t \) is
|
||||
a finite number larger than zero. By defining
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_1,
|
||||
$$
|
||||
|
||||
<p>
|
||||
we have a new optimization equation
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_1
|
||||
$$
|
||||
|
||||
which leads to Lasso regression. Lasso stands for least absolute shrinkage and selection operator.
|
||||
|
||||
<p>
|
||||
Here we have defined the norm-1 as
|
||||
$$
|
||||
\vert\vert \boldsymbol{x}\vert\vert_1 = \sum_i \vert x_i\vert.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="more-on-ridge-regression">More on Ridge Regression </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
$$
|
||||
|
||||
<p>
|
||||
by taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of \( \lambda \) does not suffer from singularity problems. We obtain
|
||||
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{Ridge}} = \left(\boldsymbol{X}^T\boldsymbol{X}+\lambda\boldsymbol{I}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y},
|
||||
$$
|
||||
|
||||
<p>
|
||||
with \( \boldsymbol{I} \) being a \( p\times p \) identity matrix with the constraint that
|
||||
|
||||
$$
|
||||
\sum_{i=0}^{p-1} \beta_i^2 \leq t,
|
||||
$$
|
||||
|
||||
<p>
|
||||
with \( t \) a finite positive number.
|
||||
|
||||
<p>
|
||||
We see that Ridge regression is nothing but the standard
|
||||
OLS with a modified diagonal term added to \( \boldsymbol{X}^T\boldsymbol{X} \). The
|
||||
consequences, in particular for our discussion of the bias-variance tradeoff
|
||||
are rather interesting.
|
||||
|
||||
<p>
|
||||
Furthermore, if we use the result above in terms of the SVD decomposition (our analysis was done for the OLS method), we had
|
||||
$$
|
||||
(\boldsymbol{X}\boldsymbol{X}^T)\boldsymbol{U} = \boldsymbol{U}\boldsymbol{D}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
We can analyse the OLS solutions in terms of the eigenvectors (the columns) of the right singular value matrix \( \boldsymbol{U} \) as
|
||||
$$
|
||||
\boldsymbol{X}\boldsymbol{\beta} = \boldsymbol{X}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T \right)^{-1}\boldsymbol{X}^T\boldsymbol{y}=\boldsymbol{U\Sigma V^T}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T \right)^{-1}(\boldsymbol{U\Sigma V^T})^T\boldsymbol{y}=\boldsymbol{U}\boldsymbol{U}^T\boldsymbol{y}
|
||||
$$
|
||||
|
||||
<p>
|
||||
For Ridge regression this becomes
|
||||
|
||||
$$
|
||||
\boldsymbol{X}\boldsymbol{\beta}^{\mathrm{Ridge}} = \boldsymbol{U\Sigma V^T}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T+\lambda\boldsymbol{I} \right)^{-1}(\boldsymbol{U\Sigma V^T})^T\boldsymbol{y}=\sum_{j=0}^{p-1}\boldsymbol{u}_j\boldsymbol{u}_j^T\frac{\sigma_j^2}{\sigma_j^2+\lambda}\boldsymbol{y},
|
||||
$$
|
||||
|
||||
<p>
|
||||
with the vectors \( \boldsymbol{u}_j \) being the columns of \( \boldsymbol{U} \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="interpreting-the-ridge-results">Interpreting the Ridge results </h2>
|
||||
|
||||
<p>
|
||||
Since \( \lambda \geq 0 \), it means that compared to OLS, we have
|
||||
|
||||
$$
|
||||
\frac{\sigma_j^2}{\sigma_j^2+\lambda} \leq 1.
|
||||
$$
|
||||
|
||||
<p>
|
||||
Ridge regression finds the coordinates of \( \boldsymbol{y} \) with respect to the
|
||||
orthonormal basis \( \boldsymbol{U} \), it then shrinks the coordinates by
|
||||
\( \frac{\sigma_j^2}{\sigma_j^2+\lambda} \). Recall that the SVD has
|
||||
eigenvalues ordered in a descending way, that is \( \sigma_i \geq
|
||||
\sigma_{i+1} \).
|
||||
|
||||
<p>
|
||||
For small eigenvalues \( \sigma_i \) it means that their contributions become less important, a fact which can be used to reduce the number of degrees of freedom.
|
||||
Actually, calculating the variance of \( \boldsymbol{X}\boldsymbol{v}_j \) shows that this quantity is equal to \( \sigma_j^2/n \).
|
||||
With a parameter \( \lambda \) we can thus shrink the role of specific parameters.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="more-interpretations">More interpretations </h2>
|
||||
|
||||
<p>
|
||||
For the sake of simplicity, let us assume that the design matrix is orthonormal, that is
|
||||
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}=(\boldsymbol{X}^T\boldsymbol{X})^{-1} =\boldsymbol{I}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
In this case the standard OLS results in
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{OLS}} = \boldsymbol{X}^T\boldsymbol{y}=\sum_{i=0}^{p-1}\boldsymbol{u}_j\boldsymbol{u}_j^T\boldsymbol{y},
|
||||
$$
|
||||
|
||||
<p>
|
||||
and
|
||||
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{Ridge}} = \left(\boldsymbol{I}+\lambda\boldsymbol{I}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y}=\left(1+\lambda\right)^{-1}\boldsymbol{\beta}^{\mathrm{OLS}},
|
||||
$$
|
||||
|
||||
<p>
|
||||
that is the Ridge estimator scales the OLS estimator by the inverse of a factor \( 1+\lambda \), and
|
||||
the Ridge estimator converges to zero when the hyperparameter goes to
|
||||
infinity.
|
||||
|
||||
<p>
|
||||
We will come back to more interpreations after we have gone through some of the statistical analysis part.
|
||||
|
||||
<p>
|
||||
For more discussions of Ridge and Lasso regression, <a href="https://arxiv.org/abs/1509.09169" target="_blank">Wessel van Wieringen's</a> article is highly recommended.
|
||||
Similarly, <a href="https://arxiv.org/abs/1803.08823" target="_blank">Mehta et al's article</a> is also recommended.
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="a-better-understanding-of-regularization">A better understanding of regularization </h2>
|
||||
|
||||
<p>
|
||||
The parameter \( \lambda \) that we have introduced in the Ridge (and
|
||||
Lasso as well) regression is often called a regularization parameter
|
||||
or shrinkage parameter. It is common to call it a hyperparameter. What does it mean mathemtically?
|
||||
|
||||
<p>
|
||||
Here we will first look at how to analyze the difference between the
|
||||
standard OLS equations and the Ridge expressions in terms of a linear
|
||||
algebra analysis using the SVD algorithm. Thereafter, we will link
|
||||
(see the material on the bias-variance tradeoff below) these
|
||||
observation to the statisical analysis of the results. In particular
|
||||
we consider how the variance of the parameters \( \boldsymbol{\beta} \) is
|
||||
affected by changing the parameter \( \lambda \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="decomposing-the-ols-and-ridge-expressions">Decomposing the OLS and Ridge expressions </h2>
|
||||
|
||||
<p>
|
||||
We have our design matrix
|
||||
\( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \). With the SVD we decompose it as
|
||||
|
||||
$$
|
||||
\boldsymbol{X} = \boldsymbol{U\Sigma V^T},
|
||||
$$
|
||||
|
||||
<p>
|
||||
with \( \boldsymbol{U}\in {\mathbb{R}}^{n\times n} \), \( \boldsymbol{\Sigma}\in {\mathbb{R}}^{n\times p} \)
|
||||
and \( \boldsymbol{V}\in {\mathbb{R}}^{p\times p} \).
|
||||
|
||||
<p>
|
||||
The matrices \( \boldsymbol{U} \) and \( \boldsymbol{V} \) are unitary/orthonormal matrices, that is in case the matrices are real we have \( \boldsymbol{U}^T\boldsymbol{U}=\boldsymbol{U}\boldsymbol{U}^T=\boldsymbol{I} \) and \( \boldsymbol{V}^T\boldsymbol{V}=\boldsymbol{V}\boldsymbol{V}^T=\boldsymbol{I} \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="mathematical-properties">Mathematical Properties </h2>
|
||||
|
||||
<p>
|
||||
|
||||
+236
-236
@@ -225,24 +225,6 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
2,
|
||||
None,
|
||||
'meet-the-covariance-matrix'),
|
||||
('Ridge and LASSO Regression',
|
||||
2,
|
||||
None,
|
||||
'ridge-and-lasso-regression'),
|
||||
('More on Ridge Regression', 2, None, 'more-on-ridge-regression'),
|
||||
('Interpreting the Ridge results',
|
||||
2,
|
||||
None,
|
||||
'interpreting-the-ridge-results'),
|
||||
('More interpretations', 2, None, 'more-interpretations'),
|
||||
('A better understanding of regularization',
|
||||
2,
|
||||
None,
|
||||
'a-better-understanding-of-regularization'),
|
||||
('Decomposing the OLS and Ridge expressions',
|
||||
2,
|
||||
None,
|
||||
'decomposing-the-ols-and-ridge-expressions'),
|
||||
('Introducing the Covariance and Correlation functions',
|
||||
2,
|
||||
None,
|
||||
@@ -268,6 +250,24 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
2,
|
||||
None,
|
||||
'rewriting-the-covariance-and-or-correlation-matrix'),
|
||||
('Ridge and LASSO Regression',
|
||||
2,
|
||||
None,
|
||||
'ridge-and-lasso-regression'),
|
||||
('More on Ridge Regression', 2, None, 'more-on-ridge-regression'),
|
||||
('Interpreting the Ridge results',
|
||||
2,
|
||||
None,
|
||||
'interpreting-the-ridge-results'),
|
||||
('More interpretations', 2, None, 'more-interpretations'),
|
||||
('A better understanding of regularization',
|
||||
2,
|
||||
None,
|
||||
'a-better-understanding-of-regularization'),
|
||||
('Decomposing the OLS and Ridge expressions',
|
||||
2,
|
||||
None,
|
||||
'decomposing-the-ols-and-ridge-expressions'),
|
||||
('Mathematical Properties', 2, None, 'mathematical-properties'),
|
||||
('Exercises for week 36, September 6-10',
|
||||
2,
|
||||
@@ -2320,228 +2320,11 @@ terms of the singular values.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="ridge-and-lasso-regression">Ridge and LASSO Regression </h2>
|
||||
|
||||
<p>
|
||||
Let us remind ourselves about the expression for the standard Mean Squared Error (MSE) which we used to define our cost function and the equations for the ordinary least squares (OLS) method, that is
|
||||
our optimization problem is
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in {\mathbb{R}}^{p}}}\frac{1}{n}\left\{\left(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\right)^T\left(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\right)\right\}.
|
||||
$$
|
||||
|
||||
or we can state it as
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\sum_{i=0}^{n-1}\left(y_i-\tilde{y}_i\right)^2=\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2,
|
||||
$$
|
||||
|
||||
where we have used the definition of a norm-2 vector, that is
|
||||
$$
|
||||
\vert\vert \boldsymbol{x}\vert\vert_2 = \sqrt{\sum_i x_i^2}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
By minimizing the above equation with respect to the parameters
|
||||
\( \boldsymbol{\beta} \) we could then obtain an analytical expression for the
|
||||
parameters \( \boldsymbol{\beta} \). We can add a regularization parameter \( \lambda \) by
|
||||
defining a new cost function to be optimized, that is
|
||||
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_2^2
|
||||
$$
|
||||
|
||||
<p>
|
||||
which leads to the Ridge regression minimization problem where we
|
||||
require that \( \vert\vert \boldsymbol{\beta}\vert\vert_2^2\le t \), where \( t \) is
|
||||
a finite number larger than zero. By defining
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_1,
|
||||
$$
|
||||
|
||||
<p>
|
||||
we have a new optimization equation
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_1
|
||||
$$
|
||||
|
||||
which leads to Lasso regression. Lasso stands for least absolute shrinkage and selection operator.
|
||||
|
||||
<p>
|
||||
Here we have defined the norm-1 as
|
||||
$$
|
||||
\vert\vert \boldsymbol{x}\vert\vert_1 = \sum_i \vert x_i\vert.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="more-on-ridge-regression">More on Ridge Regression </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
$$
|
||||
|
||||
<p>
|
||||
by taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of \( \lambda \) does not suffer from singularity problems. We obtain
|
||||
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{Ridge}} = \left(\boldsymbol{X}^T\boldsymbol{X}+\lambda\boldsymbol{I}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y},
|
||||
$$
|
||||
|
||||
<p>
|
||||
with \( \boldsymbol{I} \) being a \( p\times p \) identity matrix with the constraint that
|
||||
|
||||
$$
|
||||
\sum_{i=0}^{p-1} \beta_i^2 \leq t,
|
||||
$$
|
||||
|
||||
<p>
|
||||
with \( t \) a finite positive number.
|
||||
|
||||
<p>
|
||||
We see that Ridge regression is nothing but the standard
|
||||
OLS with a modified diagonal term added to \( \boldsymbol{X}^T\boldsymbol{X} \). The
|
||||
consequences, in particular for our discussion of the bias-variance tradeoff
|
||||
are rather interesting.
|
||||
|
||||
<p>
|
||||
Furthermore, if we use the result above in terms of the SVD decomposition (our analysis was done for the OLS method), we had
|
||||
$$
|
||||
(\boldsymbol{X}\boldsymbol{X}^T)\boldsymbol{U} = \boldsymbol{U}\boldsymbol{D}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
We can analyse the OLS solutions in terms of the eigenvectors (the columns) of the right singular value matrix \( \boldsymbol{U} \) as
|
||||
$$
|
||||
\boldsymbol{X}\boldsymbol{\beta} = \boldsymbol{X}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T \right)^{-1}\boldsymbol{X}^T\boldsymbol{y}=\boldsymbol{U\Sigma V^T}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T \right)^{-1}(\boldsymbol{U\Sigma V^T})^T\boldsymbol{y}=\boldsymbol{U}\boldsymbol{U}^T\boldsymbol{y}
|
||||
$$
|
||||
|
||||
<p>
|
||||
For Ridge regression this becomes
|
||||
|
||||
$$
|
||||
\boldsymbol{X}\boldsymbol{\beta}^{\mathrm{Ridge}} = \boldsymbol{U\Sigma V^T}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T+\lambda\boldsymbol{I} \right)^{-1}(\boldsymbol{U\Sigma V^T})^T\boldsymbol{y}=\sum_{j=0}^{p-1}\boldsymbol{u}_j\boldsymbol{u}_j^T\frac{\sigma_j^2}{\sigma_j^2+\lambda}\boldsymbol{y},
|
||||
$$
|
||||
|
||||
<p>
|
||||
with the vectors \( \boldsymbol{u}_j \) being the columns of \( \boldsymbol{U} \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="interpreting-the-ridge-results">Interpreting the Ridge results </h2>
|
||||
|
||||
<p>
|
||||
Since \( \lambda \geq 0 \), it means that compared to OLS, we have
|
||||
|
||||
$$
|
||||
\frac{\sigma_j^2}{\sigma_j^2+\lambda} \leq 1.
|
||||
$$
|
||||
|
||||
<p>
|
||||
Ridge regression finds the coordinates of \( \boldsymbol{y} \) with respect to the
|
||||
orthonormal basis \( \boldsymbol{U} \), it then shrinks the coordinates by
|
||||
\( \frac{\sigma_j^2}{\sigma_j^2+\lambda} \). Recall that the SVD has
|
||||
eigenvalues ordered in a descending way, that is \( \sigma_i \geq
|
||||
\sigma_{i+1} \).
|
||||
|
||||
<p>
|
||||
For small eigenvalues \( \sigma_i \) it means that their contributions become less important, a fact which can be used to reduce the number of degrees of freedom.
|
||||
Actually, calculating the variance of \( \boldsymbol{X}\boldsymbol{v}_j \) shows that this quantity is equal to \( \sigma_j^2/n \).
|
||||
With a parameter \( \lambda \) we can thus shrink the role of specific parameters.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="more-interpretations">More interpretations </h2>
|
||||
|
||||
<p>
|
||||
For the sake of simplicity, let us assume that the design matrix is orthonormal, that is
|
||||
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}=(\boldsymbol{X}^T\boldsymbol{X})^{-1} =\boldsymbol{I}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
In this case the standard OLS results in
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{OLS}} = \boldsymbol{X}^T\boldsymbol{y}=\sum_{i=0}^{p-1}\boldsymbol{u}_j\boldsymbol{u}_j^T\boldsymbol{y},
|
||||
$$
|
||||
|
||||
<p>
|
||||
and
|
||||
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{Ridge}} = \left(\boldsymbol{I}+\lambda\boldsymbol{I}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y}=\left(1+\lambda\right)^{-1}\boldsymbol{\beta}^{\mathrm{OLS}},
|
||||
$$
|
||||
|
||||
<p>
|
||||
that is the Ridge estimator scales the OLS estimator by the inverse of a factor \( 1+\lambda \), and
|
||||
the Ridge estimator converges to zero when the hyperparameter goes to
|
||||
infinity.
|
||||
|
||||
<p>
|
||||
We will come back to more interpreations after we have gone through some of the statistical analysis part.
|
||||
|
||||
<p>
|
||||
For more discussions of Ridge and Lasso regression, <a href="https://arxiv.org/abs/1509.09169" target="_blank">Wessel van Wieringen's</a> article is highly recommended.
|
||||
Similarly, <a href="https://arxiv.org/abs/1803.08823" target="_blank">Mehta et al's article</a> is also recommended.
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="a-better-understanding-of-regularization">A better understanding of regularization </h2>
|
||||
|
||||
<p>
|
||||
The parameter \( \lambda \) that we have introduced in the Ridge (and
|
||||
Lasso as well) regression is often called a regularization parameter
|
||||
or shrinkage parameter. It is common to call it a hyperparameter. What does it mean mathemtically?
|
||||
|
||||
<p>
|
||||
Here we will first look at how to analyze the difference between the
|
||||
standard OLS equations and the Ridge expressions in terms of a linear
|
||||
algebra analysis using the SVD algorithm. Thereafter, we will link
|
||||
(see the material on the bias-variance tradeoff below) these
|
||||
observation to the statisical analysis of the results. In particular
|
||||
we consider how the variance of the parameters \( \boldsymbol{\beta} \) is
|
||||
affected by changing the parameter \( \lambda \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="decomposing-the-ols-and-ridge-expressions">Decomposing the OLS and Ridge expressions </h2>
|
||||
|
||||
<p>
|
||||
We have our design matrix
|
||||
\( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \). With the SVD we decompose it as
|
||||
|
||||
$$
|
||||
\boldsymbol{X} = \boldsymbol{U\Sigma V^T},
|
||||
$$
|
||||
|
||||
<p>
|
||||
with \( \boldsymbol{U}\in {\mathbb{R}}^{n\times n} \), \( \boldsymbol{\Sigma}\in {\mathbb{R}}^{n\times p} \)
|
||||
and \( \boldsymbol{V}\in {\mathbb{R}}^{p\times p} \).
|
||||
|
||||
<p>
|
||||
The matrices \( \boldsymbol{U} \) and \( \boldsymbol{V} \) are unitary/orthonormal matrices, that is in case the matrices are real we have \( \boldsymbol{U}^T\boldsymbol{U}=\boldsymbol{U}\boldsymbol{U}^T=\boldsymbol{I} \) and \( \boldsymbol{V}^T\boldsymbol{V}=\boldsymbol{V}\boldsymbol{V}^T=\boldsymbol{I} \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="introducing-the-covariance-and-correlation-functions">Introducing the Covariance and Correlation functions </h2>
|
||||
|
||||
<p>
|
||||
Before we discuss the link between for example Ridge regression and the singular value decomposition, we need to remind ourselves about
|
||||
the definition of the covariance and the correlation function. These are quantities
|
||||
the definition of the covariance and the correlation function. These are quantities that play a central role in machine learning methods.
|
||||
|
||||
<p>
|
||||
Suppose we have defined two vectors
|
||||
@@ -2880,6 +2663,223 @@ It is easy to generalize this to a matrix \( \boldsymbol{X}\in {\mathbb{R}}^{n\t
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="ridge-and-lasso-regression">Ridge and LASSO Regression </h2>
|
||||
|
||||
<p>
|
||||
Let us remind ourselves about the expression for the standard Mean Squared Error (MSE) which we used to define our cost function and the equations for the ordinary least squares (OLS) method, that is
|
||||
our optimization problem is
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in {\mathbb{R}}^{p}}}\frac{1}{n}\left\{\left(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\right)^T\left(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\right)\right\}.
|
||||
$$
|
||||
|
||||
or we can state it as
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\sum_{i=0}^{n-1}\left(y_i-\tilde{y}_i\right)^2=\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2,
|
||||
$$
|
||||
|
||||
where we have used the definition of a norm-2 vector, that is
|
||||
$$
|
||||
\vert\vert \boldsymbol{x}\vert\vert_2 = \sqrt{\sum_i x_i^2}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
By minimizing the above equation with respect to the parameters
|
||||
\( \boldsymbol{\beta} \) we could then obtain an analytical expression for the
|
||||
parameters \( \boldsymbol{\beta} \). We can add a regularization parameter \( \lambda \) by
|
||||
defining a new cost function to be optimized, that is
|
||||
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_2^2
|
||||
$$
|
||||
|
||||
<p>
|
||||
which leads to the Ridge regression minimization problem where we
|
||||
require that \( \vert\vert \boldsymbol{\beta}\vert\vert_2^2\le t \), where \( t \) is
|
||||
a finite number larger than zero. By defining
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_1,
|
||||
$$
|
||||
|
||||
<p>
|
||||
we have a new optimization equation
|
||||
$$
|
||||
{\displaystyle \min_{\boldsymbol{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta}\vert\vert_2^2+\lambda\vert\vert \boldsymbol{\beta}\vert\vert_1
|
||||
$$
|
||||
|
||||
which leads to Lasso regression. Lasso stands for least absolute shrinkage and selection operator.
|
||||
|
||||
<p>
|
||||
Here we have defined the norm-1 as
|
||||
$$
|
||||
\vert\vert \boldsymbol{x}\vert\vert_1 = \sum_i \vert x_i\vert.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="more-on-ridge-regression">More on Ridge Regression </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
$$
|
||||
|
||||
<p>
|
||||
by taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of \( \lambda \) does not suffer from singularity problems. We obtain
|
||||
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{Ridge}} = \left(\boldsymbol{X}^T\boldsymbol{X}+\lambda\boldsymbol{I}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y},
|
||||
$$
|
||||
|
||||
<p>
|
||||
with \( \boldsymbol{I} \) being a \( p\times p \) identity matrix with the constraint that
|
||||
|
||||
$$
|
||||
\sum_{i=0}^{p-1} \beta_i^2 \leq t,
|
||||
$$
|
||||
|
||||
<p>
|
||||
with \( t \) a finite positive number.
|
||||
|
||||
<p>
|
||||
We see that Ridge regression is nothing but the standard
|
||||
OLS with a modified diagonal term added to \( \boldsymbol{X}^T\boldsymbol{X} \). The
|
||||
consequences, in particular for our discussion of the bias-variance tradeoff
|
||||
are rather interesting.
|
||||
|
||||
<p>
|
||||
Furthermore, if we use the result above in terms of the SVD decomposition (our analysis was done for the OLS method), we had
|
||||
$$
|
||||
(\boldsymbol{X}\boldsymbol{X}^T)\boldsymbol{U} = \boldsymbol{U}\boldsymbol{D}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
We can analyse the OLS solutions in terms of the eigenvectors (the columns) of the right singular value matrix \( \boldsymbol{U} \) as
|
||||
$$
|
||||
\boldsymbol{X}\boldsymbol{\beta} = \boldsymbol{X}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T \right)^{-1}\boldsymbol{X}^T\boldsymbol{y}=\boldsymbol{U\Sigma V^T}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T \right)^{-1}(\boldsymbol{U\Sigma V^T})^T\boldsymbol{y}=\boldsymbol{U}\boldsymbol{U}^T\boldsymbol{y}
|
||||
$$
|
||||
|
||||
<p>
|
||||
For Ridge regression this becomes
|
||||
|
||||
$$
|
||||
\boldsymbol{X}\boldsymbol{\beta}^{\mathrm{Ridge}} = \boldsymbol{U\Sigma V^T}\left(\boldsymbol{V}\boldsymbol{D}\boldsymbol{V}^T+\lambda\boldsymbol{I} \right)^{-1}(\boldsymbol{U\Sigma V^T})^T\boldsymbol{y}=\sum_{j=0}^{p-1}\boldsymbol{u}_j\boldsymbol{u}_j^T\frac{\sigma_j^2}{\sigma_j^2+\lambda}\boldsymbol{y},
|
||||
$$
|
||||
|
||||
<p>
|
||||
with the vectors \( \boldsymbol{u}_j \) being the columns of \( \boldsymbol{U} \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="interpreting-the-ridge-results">Interpreting the Ridge results </h2>
|
||||
|
||||
<p>
|
||||
Since \( \lambda \geq 0 \), it means that compared to OLS, we have
|
||||
|
||||
$$
|
||||
\frac{\sigma_j^2}{\sigma_j^2+\lambda} \leq 1.
|
||||
$$
|
||||
|
||||
<p>
|
||||
Ridge regression finds the coordinates of \( \boldsymbol{y} \) with respect to the
|
||||
orthonormal basis \( \boldsymbol{U} \), it then shrinks the coordinates by
|
||||
\( \frac{\sigma_j^2}{\sigma_j^2+\lambda} \). Recall that the SVD has
|
||||
eigenvalues ordered in a descending way, that is \( \sigma_i \geq
|
||||
\sigma_{i+1} \).
|
||||
|
||||
<p>
|
||||
For small eigenvalues \( \sigma_i \) it means that their contributions become less important, a fact which can be used to reduce the number of degrees of freedom.
|
||||
Actually, calculating the variance of \( \boldsymbol{X}\boldsymbol{v}_j \) shows that this quantity is equal to \( \sigma_j^2/n \).
|
||||
With a parameter \( \lambda \) we can thus shrink the role of specific parameters.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="more-interpretations">More interpretations </h2>
|
||||
|
||||
<p>
|
||||
For the sake of simplicity, let us assume that the design matrix is orthonormal, that is
|
||||
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}=(\boldsymbol{X}^T\boldsymbol{X})^{-1} =\boldsymbol{I}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
In this case the standard OLS results in
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{OLS}} = \boldsymbol{X}^T\boldsymbol{y}=\sum_{i=0}^{p-1}\boldsymbol{u}_j\boldsymbol{u}_j^T\boldsymbol{y},
|
||||
$$
|
||||
|
||||
<p>
|
||||
and
|
||||
|
||||
$$
|
||||
\boldsymbol{\beta}^{\mathrm{Ridge}} = \left(\boldsymbol{I}+\lambda\boldsymbol{I}\right)^{-1}\boldsymbol{X}^T\boldsymbol{y}=\left(1+\lambda\right)^{-1}\boldsymbol{\beta}^{\mathrm{OLS}},
|
||||
$$
|
||||
|
||||
<p>
|
||||
that is the Ridge estimator scales the OLS estimator by the inverse of a factor \( 1+\lambda \), and
|
||||
the Ridge estimator converges to zero when the hyperparameter goes to
|
||||
infinity.
|
||||
|
||||
<p>
|
||||
We will come back to more interpreations after we have gone through some of the statistical analysis part.
|
||||
|
||||
<p>
|
||||
For more discussions of Ridge and Lasso regression, <a href="https://arxiv.org/abs/1509.09169" target="_blank">Wessel van Wieringen's</a> article is highly recommended.
|
||||
Similarly, <a href="https://arxiv.org/abs/1803.08823" target="_blank">Mehta et al's article</a> is also recommended.
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="a-better-understanding-of-regularization">A better understanding of regularization </h2>
|
||||
|
||||
<p>
|
||||
The parameter \( \lambda \) that we have introduced in the Ridge (and
|
||||
Lasso as well) regression is often called a regularization parameter
|
||||
or shrinkage parameter. It is common to call it a hyperparameter. What does it mean mathemtically?
|
||||
|
||||
<p>
|
||||
Here we will first look at how to analyze the difference between the
|
||||
standard OLS equations and the Ridge expressions in terms of a linear
|
||||
algebra analysis using the SVD algorithm. Thereafter, we will link
|
||||
(see the material on the bias-variance tradeoff below) these
|
||||
observation to the statisical analysis of the results. In particular
|
||||
we consider how the variance of the parameters \( \boldsymbol{\beta} \) is
|
||||
affected by changing the parameter \( \lambda \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="decomposing-the-ols-and-ridge-expressions">Decomposing the OLS and Ridge expressions </h2>
|
||||
|
||||
<p>
|
||||
We have our design matrix
|
||||
\( \boldsymbol{X}\in {\mathbb{R}}^{n\times p} \). With the SVD we decompose it as
|
||||
|
||||
$$
|
||||
\boldsymbol{X} = \boldsymbol{U\Sigma V^T},
|
||||
$$
|
||||
|
||||
<p>
|
||||
with \( \boldsymbol{U}\in {\mathbb{R}}^{n\times n} \), \( \boldsymbol{\Sigma}\in {\mathbb{R}}^{n\times p} \)
|
||||
and \( \boldsymbol{V}\in {\mathbb{R}}^{p\times p} \).
|
||||
|
||||
<p>
|
||||
The matrices \( \boldsymbol{U} \) and \( \boldsymbol{V} \) are unitary/orthonormal matrices, that is in case the matrices are real we have \( \boldsymbol{U}^T\boldsymbol{U}=\boldsymbol{U}\boldsymbol{U}^T=\boldsymbol{I} \) and \( \boldsymbol{V}^T\boldsymbol{V}=\boldsymbol{V}\boldsymbol{V}^T=\boldsymbol{I} \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="mathematical-properties">Mathematical Properties </h2>
|
||||
|
||||
<p>
|
||||
|
||||
Binary file not shown.
+369
-369
@@ -2903,378 +2903,10 @@
|
||||
"terms of the singular values.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Ridge and LASSO Regression\n",
|
||||
"\n",
|
||||
"Let us remind ourselves about the expression for the standard Mean Squared Error (MSE) which we used to define our cost function and the equations for the ordinary least squares (OLS) method, that is \n",
|
||||
"our optimization problem is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"{\\displaystyle \\min_{\\boldsymbol{\\beta}\\in {\\mathbb{R}}^{p}}}\\frac{1}{n}\\left\\{\\left(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta}\\right)^T\\left(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta}\\right)\\right\\}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"or we can state it as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"{\\displaystyle \\min_{\\boldsymbol{\\beta}\\in\n",
|
||||
"{\\mathbb{R}}^{p}}}\\frac{1}{n}\\sum_{i=0}^{n-1}\\left(y_i-\\tilde{y}_i\\right)^2=\\frac{1}{n}\\vert\\vert \\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta}\\vert\\vert_2^2,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"where we have used the definition of a norm-2 vector, that is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\vert\\vert \\boldsymbol{x}\\vert\\vert_2 = \\sqrt{\\sum_i x_i^2}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"By minimizing the above equation with respect to the parameters\n",
|
||||
"$\\boldsymbol{\\beta}$ we could then obtain an analytical expression for the\n",
|
||||
"parameters $\\boldsymbol{\\beta}$. We can add a regularization parameter $\\lambda$ by\n",
|
||||
"defining a new cost function to be optimized, that is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"{\\displaystyle \\min_{\\boldsymbol{\\beta}\\in\n",
|
||||
"{\\mathbb{R}}^{p}}}\\frac{1}{n}\\vert\\vert \\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta}\\vert\\vert_2^2+\\lambda\\vert\\vert \\boldsymbol{\\beta}\\vert\\vert_2^2\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"which leads to the Ridge regression minimization problem where we\n",
|
||||
"require that $\\vert\\vert \\boldsymbol{\\beta}\\vert\\vert_2^2\\le t$, where $t$ is\n",
|
||||
"a finite number larger than zero. By defining"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"C(\\boldsymbol{X},\\boldsymbol{\\beta})=\\frac{1}{n}\\vert\\vert \\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta}\\vert\\vert_2^2+\\lambda\\vert\\vert \\boldsymbol{\\beta}\\vert\\vert_1,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"we have a new optimization equation"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"{\\displaystyle \\min_{\\boldsymbol{\\beta}\\in\n",
|
||||
"{\\mathbb{R}}^{p}}}\\frac{1}{n}\\vert\\vert \\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta}\\vert\\vert_2^2+\\lambda\\vert\\vert \\boldsymbol{\\beta}\\vert\\vert_1\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"which leads to Lasso regression. Lasso stands for least absolute shrinkage and selection operator. \n",
|
||||
"\n",
|
||||
"Here we have defined the norm-1 as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\vert\\vert \\boldsymbol{x}\\vert\\vert_1 = \\sum_i \\vert x_i\\vert.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## More on Ridge Regression\n",
|
||||
"\n",
|
||||
"Using the matrix-vector expression for Ridge regression,"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"C(\\boldsymbol{X},\\boldsymbol{\\beta})=\\frac{1}{n}\\left\\{(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta})^T(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta})\\right\\}+\\lambda\\boldsymbol{\\beta}^T\\boldsymbol{\\beta},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"by taking the derivatives with respect to $\\boldsymbol{\\beta}$ we obtain then\n",
|
||||
"a slightly modified matrix inversion problem which for finite values\n",
|
||||
"of $\\lambda$ does not suffer from singularity problems. We obtain"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{\\beta}^{\\mathrm{Ridge}} = \\left(\\boldsymbol{X}^T\\boldsymbol{X}+\\lambda\\boldsymbol{I}\\right)^{-1}\\boldsymbol{X}^T\\boldsymbol{y},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"with $\\boldsymbol{I}$ being a $p\\times p$ identity matrix with the constraint that"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\sum_{i=0}^{p-1} \\beta_i^2 \\leq t,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"with $t$ a finite positive number. \n",
|
||||
"\n",
|
||||
"We see that Ridge regression is nothing but the standard\n",
|
||||
"OLS with a modified diagonal term added to $\\boldsymbol{X}^T\\boldsymbol{X}$. The\n",
|
||||
"consequences, in particular for our discussion of the bias-variance tradeoff \n",
|
||||
"are rather interesting.\n",
|
||||
"\n",
|
||||
"Furthermore, if we use the result above in terms of the SVD decomposition (our analysis was done for the OLS method), we had"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"(\\boldsymbol{X}\\boldsymbol{X}^T)\\boldsymbol{U} = \\boldsymbol{U}\\boldsymbol{D}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"We can analyse the OLS solutions in terms of the eigenvectors (the columns) of the right singular value matrix $\\boldsymbol{U}$ as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{X}\\boldsymbol{\\beta} = \\boldsymbol{X}\\left(\\boldsymbol{V}\\boldsymbol{D}\\boldsymbol{V}^T \\right)^{-1}\\boldsymbol{X}^T\\boldsymbol{y}=\\boldsymbol{U\\Sigma V^T}\\left(\\boldsymbol{V}\\boldsymbol{D}\\boldsymbol{V}^T \\right)^{-1}(\\boldsymbol{U\\Sigma V^T})^T\\boldsymbol{y}=\\boldsymbol{U}\\boldsymbol{U}^T\\boldsymbol{y}\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"For Ridge regression this becomes"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{X}\\boldsymbol{\\beta}^{\\mathrm{Ridge}} = \\boldsymbol{U\\Sigma V^T}\\left(\\boldsymbol{V}\\boldsymbol{D}\\boldsymbol{V}^T+\\lambda\\boldsymbol{I} \\right)^{-1}(\\boldsymbol{U\\Sigma V^T})^T\\boldsymbol{y}=\\sum_{j=0}^{p-1}\\boldsymbol{u}_j\\boldsymbol{u}_j^T\\frac{\\sigma_j^2}{\\sigma_j^2+\\lambda}\\boldsymbol{y},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"with the vectors $\\boldsymbol{u}_j$ being the columns of $\\boldsymbol{U}$. \n",
|
||||
"\n",
|
||||
"## Interpreting the Ridge results\n",
|
||||
"\n",
|
||||
"Since $\\lambda \\geq 0$, it means that compared to OLS, we have"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\frac{\\sigma_j^2}{\\sigma_j^2+\\lambda} \\leq 1.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Ridge regression finds the coordinates of $\\boldsymbol{y}$ with respect to the\n",
|
||||
"orthonormal basis $\\boldsymbol{U}$, it then shrinks the coordinates by\n",
|
||||
"$\\frac{\\sigma_j^2}{\\sigma_j^2+\\lambda}$. Recall that the SVD has\n",
|
||||
"eigenvalues ordered in a descending way, that is $\\sigma_i \\geq\n",
|
||||
"\\sigma_{i+1}$.\n",
|
||||
"\n",
|
||||
"For small eigenvalues $\\sigma_i$ it means that their contributions become less important, a fact which can be used to reduce the number of degrees of freedom.\n",
|
||||
"Actually, calculating the variance of $\\boldsymbol{X}\\boldsymbol{v}_j$ shows that this quantity is equal to $\\sigma_j^2/n$.\n",
|
||||
"With a parameter $\\lambda$ we can thus shrink the role of specific parameters. \n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## More interpretations\n",
|
||||
"\n",
|
||||
"For the sake of simplicity, let us assume that the design matrix is orthonormal, that is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{X}^T\\boldsymbol{X}=(\\boldsymbol{X}^T\\boldsymbol{X})^{-1} =\\boldsymbol{I}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"In this case the standard OLS results in"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{\\beta}^{\\mathrm{OLS}} = \\boldsymbol{X}^T\\boldsymbol{y}=\\sum_{i=0}^{p-1}\\boldsymbol{u}_j\\boldsymbol{u}_j^T\\boldsymbol{y},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{\\beta}^{\\mathrm{Ridge}} = \\left(\\boldsymbol{I}+\\lambda\\boldsymbol{I}\\right)^{-1}\\boldsymbol{X}^T\\boldsymbol{y}=\\left(1+\\lambda\\right)^{-1}\\boldsymbol{\\beta}^{\\mathrm{OLS}},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"that is the Ridge estimator scales the OLS estimator by the inverse of a factor $1+\\lambda$, and\n",
|
||||
"the Ridge estimator converges to zero when the hyperparameter goes to\n",
|
||||
"infinity.\n",
|
||||
"\n",
|
||||
"We will come back to more interpreations after we have gone through some of the statistical analysis part. \n",
|
||||
"\n",
|
||||
"For more discussions of Ridge and Lasso regression, [Wessel van Wieringen's](https://arxiv.org/abs/1509.09169) article is highly recommended.\n",
|
||||
"Similarly, [Mehta et al's article](https://arxiv.org/abs/1803.08823) is also recommended.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"<!-- !split -->\n",
|
||||
"## A better understanding of regularization\n",
|
||||
"\n",
|
||||
"The parameter $\\lambda$ that we have introduced in the Ridge (and\n",
|
||||
"Lasso as well) regression is often called a regularization parameter\n",
|
||||
"or shrinkage parameter. It is common to call it a hyperparameter. What does it mean mathemtically?\n",
|
||||
"\n",
|
||||
"Here we will first look at how to analyze the difference between the\n",
|
||||
"standard OLS equations and the Ridge expressions in terms of a linear\n",
|
||||
"algebra analysis using the SVD algorithm. Thereafter, we will link\n",
|
||||
"(see the material on the bias-variance tradeoff below) these\n",
|
||||
"observation to the statisical analysis of the results. In particular\n",
|
||||
"we consider how the variance of the parameters $\\boldsymbol{\\beta}$ is\n",
|
||||
"affected by changing the parameter $\\lambda$.\n",
|
||||
"\n",
|
||||
"## Decomposing the OLS and Ridge expressions\n",
|
||||
"\n",
|
||||
"We have our design matrix\n",
|
||||
" $\\boldsymbol{X}\\in {\\mathbb{R}}^{n\\times p}$. With the SVD we decompose it as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{X} = \\boldsymbol{U\\Sigma V^T},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"with $\\boldsymbol{U}\\in {\\mathbb{R}}^{n\\times n}$, $\\boldsymbol{\\Sigma}\\in {\\mathbb{R}}^{n\\times p}$\n",
|
||||
"and $\\boldsymbol{V}\\in {\\mathbb{R}}^{p\\times p}$.\n",
|
||||
"\n",
|
||||
"The matrices $\\boldsymbol{U}$ and $\\boldsymbol{V}$ are unitary/orthonormal matrices, that is in case the matrices are real we have $\\boldsymbol{U}^T\\boldsymbol{U}=\\boldsymbol{U}\\boldsymbol{U}^T=\\boldsymbol{I}$ and $\\boldsymbol{V}^T\\boldsymbol{V}=\\boldsymbol{V}\\boldsymbol{V}^T=\\boldsymbol{I}$.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Introducing the Covariance and Correlation functions\n",
|
||||
"\n",
|
||||
"Before we discuss the link between for example Ridge regression and the singular value decomposition, we need to remind ourselves about\n",
|
||||
"the definition of the covariance and the correlation function. These are quantities \n",
|
||||
"the definition of the covariance and the correlation function. These are quantities that play a central role in machine learning methods.\n",
|
||||
"\n",
|
||||
"Suppose we have defined two vectors\n",
|
||||
"$\\hat{x}$ and $\\hat{y}$ with $n$ elements each. The covariance matrix $\\boldsymbol{C}$ is defined as"
|
||||
@@ -3796,6 +3428,374 @@
|
||||
"It is easy to generalize this to a matrix $\\boldsymbol{X}\\in {\\mathbb{R}}^{n\\times p}$.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Ridge and LASSO Regression\n",
|
||||
"\n",
|
||||
"Let us remind ourselves about the expression for the standard Mean Squared Error (MSE) which we used to define our cost function and the equations for the ordinary least squares (OLS) method, that is \n",
|
||||
"our optimization problem is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"{\\displaystyle \\min_{\\boldsymbol{\\beta}\\in {\\mathbb{R}}^{p}}}\\frac{1}{n}\\left\\{\\left(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta}\\right)^T\\left(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta}\\right)\\right\\}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"or we can state it as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"{\\displaystyle \\min_{\\boldsymbol{\\beta}\\in\n",
|
||||
"{\\mathbb{R}}^{p}}}\\frac{1}{n}\\sum_{i=0}^{n-1}\\left(y_i-\\tilde{y}_i\\right)^2=\\frac{1}{n}\\vert\\vert \\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta}\\vert\\vert_2^2,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"where we have used the definition of a norm-2 vector, that is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\vert\\vert \\boldsymbol{x}\\vert\\vert_2 = \\sqrt{\\sum_i x_i^2}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"By minimizing the above equation with respect to the parameters\n",
|
||||
"$\\boldsymbol{\\beta}$ we could then obtain an analytical expression for the\n",
|
||||
"parameters $\\boldsymbol{\\beta}$. We can add a regularization parameter $\\lambda$ by\n",
|
||||
"defining a new cost function to be optimized, that is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"{\\displaystyle \\min_{\\boldsymbol{\\beta}\\in\n",
|
||||
"{\\mathbb{R}}^{p}}}\\frac{1}{n}\\vert\\vert \\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta}\\vert\\vert_2^2+\\lambda\\vert\\vert \\boldsymbol{\\beta}\\vert\\vert_2^2\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"which leads to the Ridge regression minimization problem where we\n",
|
||||
"require that $\\vert\\vert \\boldsymbol{\\beta}\\vert\\vert_2^2\\le t$, where $t$ is\n",
|
||||
"a finite number larger than zero. By defining"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"C(\\boldsymbol{X},\\boldsymbol{\\beta})=\\frac{1}{n}\\vert\\vert \\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta}\\vert\\vert_2^2+\\lambda\\vert\\vert \\boldsymbol{\\beta}\\vert\\vert_1,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"we have a new optimization equation"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"{\\displaystyle \\min_{\\boldsymbol{\\beta}\\in\n",
|
||||
"{\\mathbb{R}}^{p}}}\\frac{1}{n}\\vert\\vert \\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta}\\vert\\vert_2^2+\\lambda\\vert\\vert \\boldsymbol{\\beta}\\vert\\vert_1\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"which leads to Lasso regression. Lasso stands for least absolute shrinkage and selection operator. \n",
|
||||
"\n",
|
||||
"Here we have defined the norm-1 as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\vert\\vert \\boldsymbol{x}\\vert\\vert_1 = \\sum_i \\vert x_i\\vert.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## More on Ridge Regression\n",
|
||||
"\n",
|
||||
"Using the matrix-vector expression for Ridge regression,"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"C(\\boldsymbol{X},\\boldsymbol{\\beta})=\\frac{1}{n}\\left\\{(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta})^T(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta})\\right\\}+\\lambda\\boldsymbol{\\beta}^T\\boldsymbol{\\beta},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"by taking the derivatives with respect to $\\boldsymbol{\\beta}$ we obtain then\n",
|
||||
"a slightly modified matrix inversion problem which for finite values\n",
|
||||
"of $\\lambda$ does not suffer from singularity problems. We obtain"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{\\beta}^{\\mathrm{Ridge}} = \\left(\\boldsymbol{X}^T\\boldsymbol{X}+\\lambda\\boldsymbol{I}\\right)^{-1}\\boldsymbol{X}^T\\boldsymbol{y},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"with $\\boldsymbol{I}$ being a $p\\times p$ identity matrix with the constraint that"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\sum_{i=0}^{p-1} \\beta_i^2 \\leq t,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"with $t$ a finite positive number. \n",
|
||||
"\n",
|
||||
"We see that Ridge regression is nothing but the standard\n",
|
||||
"OLS with a modified diagonal term added to $\\boldsymbol{X}^T\\boldsymbol{X}$. The\n",
|
||||
"consequences, in particular for our discussion of the bias-variance tradeoff \n",
|
||||
"are rather interesting.\n",
|
||||
"\n",
|
||||
"Furthermore, if we use the result above in terms of the SVD decomposition (our analysis was done for the OLS method), we had"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"(\\boldsymbol{X}\\boldsymbol{X}^T)\\boldsymbol{U} = \\boldsymbol{U}\\boldsymbol{D}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"We can analyse the OLS solutions in terms of the eigenvectors (the columns) of the right singular value matrix $\\boldsymbol{U}$ as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{X}\\boldsymbol{\\beta} = \\boldsymbol{X}\\left(\\boldsymbol{V}\\boldsymbol{D}\\boldsymbol{V}^T \\right)^{-1}\\boldsymbol{X}^T\\boldsymbol{y}=\\boldsymbol{U\\Sigma V^T}\\left(\\boldsymbol{V}\\boldsymbol{D}\\boldsymbol{V}^T \\right)^{-1}(\\boldsymbol{U\\Sigma V^T})^T\\boldsymbol{y}=\\boldsymbol{U}\\boldsymbol{U}^T\\boldsymbol{y}\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"For Ridge regression this becomes"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{X}\\boldsymbol{\\beta}^{\\mathrm{Ridge}} = \\boldsymbol{U\\Sigma V^T}\\left(\\boldsymbol{V}\\boldsymbol{D}\\boldsymbol{V}^T+\\lambda\\boldsymbol{I} \\right)^{-1}(\\boldsymbol{U\\Sigma V^T})^T\\boldsymbol{y}=\\sum_{j=0}^{p-1}\\boldsymbol{u}_j\\boldsymbol{u}_j^T\\frac{\\sigma_j^2}{\\sigma_j^2+\\lambda}\\boldsymbol{y},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"with the vectors $\\boldsymbol{u}_j$ being the columns of $\\boldsymbol{U}$. \n",
|
||||
"\n",
|
||||
"## Interpreting the Ridge results\n",
|
||||
"\n",
|
||||
"Since $\\lambda \\geq 0$, it means that compared to OLS, we have"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\frac{\\sigma_j^2}{\\sigma_j^2+\\lambda} \\leq 1.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Ridge regression finds the coordinates of $\\boldsymbol{y}$ with respect to the\n",
|
||||
"orthonormal basis $\\boldsymbol{U}$, it then shrinks the coordinates by\n",
|
||||
"$\\frac{\\sigma_j^2}{\\sigma_j^2+\\lambda}$. Recall that the SVD has\n",
|
||||
"eigenvalues ordered in a descending way, that is $\\sigma_i \\geq\n",
|
||||
"\\sigma_{i+1}$.\n",
|
||||
"\n",
|
||||
"For small eigenvalues $\\sigma_i$ it means that their contributions become less important, a fact which can be used to reduce the number of degrees of freedom.\n",
|
||||
"Actually, calculating the variance of $\\boldsymbol{X}\\boldsymbol{v}_j$ shows that this quantity is equal to $\\sigma_j^2/n$.\n",
|
||||
"With a parameter $\\lambda$ we can thus shrink the role of specific parameters. \n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## More interpretations\n",
|
||||
"\n",
|
||||
"For the sake of simplicity, let us assume that the design matrix is orthonormal, that is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{X}^T\\boldsymbol{X}=(\\boldsymbol{X}^T\\boldsymbol{X})^{-1} =\\boldsymbol{I}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"In this case the standard OLS results in"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{\\beta}^{\\mathrm{OLS}} = \\boldsymbol{X}^T\\boldsymbol{y}=\\sum_{i=0}^{p-1}\\boldsymbol{u}_j\\boldsymbol{u}_j^T\\boldsymbol{y},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{\\beta}^{\\mathrm{Ridge}} = \\left(\\boldsymbol{I}+\\lambda\\boldsymbol{I}\\right)^{-1}\\boldsymbol{X}^T\\boldsymbol{y}=\\left(1+\\lambda\\right)^{-1}\\boldsymbol{\\beta}^{\\mathrm{OLS}},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"that is the Ridge estimator scales the OLS estimator by the inverse of a factor $1+\\lambda$, and\n",
|
||||
"the Ridge estimator converges to zero when the hyperparameter goes to\n",
|
||||
"infinity.\n",
|
||||
"\n",
|
||||
"We will come back to more interpreations after we have gone through some of the statistical analysis part. \n",
|
||||
"\n",
|
||||
"For more discussions of Ridge and Lasso regression, [Wessel van Wieringen's](https://arxiv.org/abs/1509.09169) article is highly recommended.\n",
|
||||
"Similarly, [Mehta et al's article](https://arxiv.org/abs/1803.08823) is also recommended.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"<!-- !split -->\n",
|
||||
"## A better understanding of regularization\n",
|
||||
"\n",
|
||||
"The parameter $\\lambda$ that we have introduced in the Ridge (and\n",
|
||||
"Lasso as well) regression is often called a regularization parameter\n",
|
||||
"or shrinkage parameter. It is common to call it a hyperparameter. What does it mean mathemtically?\n",
|
||||
"\n",
|
||||
"Here we will first look at how to analyze the difference between the\n",
|
||||
"standard OLS equations and the Ridge expressions in terms of a linear\n",
|
||||
"algebra analysis using the SVD algorithm. Thereafter, we will link\n",
|
||||
"(see the material on the bias-variance tradeoff below) these\n",
|
||||
"observation to the statisical analysis of the results. In particular\n",
|
||||
"we consider how the variance of the parameters $\\boldsymbol{\\beta}$ is\n",
|
||||
"affected by changing the parameter $\\lambda$.\n",
|
||||
"\n",
|
||||
"## Decomposing the OLS and Ridge expressions\n",
|
||||
"\n",
|
||||
"We have our design matrix\n",
|
||||
" $\\boldsymbol{X}\\in {\\mathbb{R}}^{n\\times p}$. With the SVD we decompose it as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{X} = \\boldsymbol{U\\Sigma V^T},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"with $\\boldsymbol{U}\\in {\\mathbb{R}}^{n\\times n}$, $\\boldsymbol{\\Sigma}\\in {\\mathbb{R}}^{n\\times p}$\n",
|
||||
"and $\\boldsymbol{V}\\in {\\mathbb{R}}^{p\\times p}$.\n",
|
||||
"\n",
|
||||
"The matrices $\\boldsymbol{U}$ and $\\boldsymbol{V}$ are unitary/orthonormal matrices, that is in case the matrices are real we have $\\boldsymbol{U}^T\\boldsymbol{U}=\\boldsymbol{U}\\boldsymbol{U}^T=\\boldsymbol{I}$ and $\\boldsymbol{V}^T\\boldsymbol{V}=\\boldsymbol{V}\\boldsymbol{V}^T=\\boldsymbol{I}$.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Mathematical Properties\n",
|
||||
"\n",
|
||||
"There are several interesting mathematical properties which will be\n",
|
||||
|
||||
+221
-221
@@ -1821,231 +1821,11 @@ the eigenvalues of the covariance matrix and the Hessian matrix in
|
||||
terms of the singular values.
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== Ridge and LASSO Regression =====
|
||||
|
||||
Let us remind ourselves about the expression for the standard Mean Squared Error (MSE) which we used to define our cost function and the equations for the ordinary least squares (OLS) method, that is
|
||||
our optimization problem is
|
||||
!bt
|
||||
\[
|
||||
{\displaystyle \min_{\bm{\beta}\in {\mathbb{R}}^{p}}}\frac{1}{n}\left\{\left(\bm{y}-\bm{X}\bm{\beta}\right)^T\left(\bm{y}-\bm{X}\bm{\beta}\right)\right\}.
|
||||
\]
|
||||
!et
|
||||
or we can state it as
|
||||
!bt
|
||||
\[
|
||||
{\displaystyle \min_{\bm{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\sum_{i=0}^{n-1}\left(y_i-\tilde{y}_i\right)^2=\frac{1}{n}\vert\vert \bm{y}-\bm{X}\bm{\beta}\vert\vert_2^2,
|
||||
\]
|
||||
!et
|
||||
where we have used the definition of a norm-2 vector, that is
|
||||
!bt
|
||||
\[
|
||||
\vert\vert \bm{x}\vert\vert_2 = \sqrt{\sum_i x_i^2}.
|
||||
\]
|
||||
!et
|
||||
|
||||
By minimizing the above equation with respect to the parameters
|
||||
$\bm{\beta}$ we could then obtain an analytical expression for the
|
||||
parameters $\bm{\beta}$. We can add a regularization parameter $\lambda$ by
|
||||
defining a new cost function to be optimized, that is
|
||||
|
||||
!bt
|
||||
\[
|
||||
{\displaystyle \min_{\bm{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \bm{y}-\bm{X}\bm{\beta}\vert\vert_2^2+\lambda\vert\vert \bm{\beta}\vert\vert_2^2
|
||||
\]
|
||||
!et
|
||||
|
||||
which leads to the Ridge regression minimization problem where we
|
||||
require that $\vert\vert \bm{\beta}\vert\vert_2^2\le t$, where $t$ is
|
||||
a finite number larger than zero. By defining
|
||||
|
||||
!bt
|
||||
\[
|
||||
C(\bm{X},\bm{\beta})=\frac{1}{n}\vert\vert \bm{y}-\bm{X}\bm{\beta}\vert\vert_2^2+\lambda\vert\vert \bm{\beta}\vert\vert_1,
|
||||
\]
|
||||
!et
|
||||
|
||||
we have a new optimization equation
|
||||
!bt
|
||||
\[
|
||||
{\displaystyle \min_{\bm{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \bm{y}-\bm{X}\bm{\beta}\vert\vert_2^2+\lambda\vert\vert \bm{\beta}\vert\vert_1
|
||||
\]
|
||||
!et
|
||||
which leads to Lasso regression. Lasso stands for least absolute shrinkage and selection operator.
|
||||
|
||||
Here we have defined the norm-1 as
|
||||
!bt
|
||||
\[
|
||||
\vert\vert \bm{x}\vert\vert_1 = \sum_i \vert x_i\vert.
|
||||
\]
|
||||
!et
|
||||
|
||||
|
||||
!split
|
||||
===== More on Ridge Regression =====
|
||||
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
|
||||
!bt
|
||||
\[
|
||||
C(\bm{X},\bm{\beta})=\frac{1}{n}\left\{(\bm{y}-\bm{X}\bm{\beta})^T(\bm{y}-\bm{X}\bm{\beta})\right\}+\lambda\bm{\beta}^T\bm{\beta},
|
||||
\]
|
||||
!et
|
||||
|
||||
by taking the derivatives with respect to $\bm{\beta}$ we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of $\lambda$ does not suffer from singularity problems. We obtain
|
||||
|
||||
!bt
|
||||
\[
|
||||
\bm{\beta}^{\mathrm{Ridge}} = \left(\bm{X}^T\bm{X}+\lambda\bm{I}\right)^{-1}\bm{X}^T\bm{y},
|
||||
\]
|
||||
!et
|
||||
|
||||
with $\bm{I}$ being a $p\times p$ identity matrix with the constraint that
|
||||
|
||||
!bt
|
||||
\[
|
||||
\sum_{i=0}^{p-1} \beta_i^2 \leq t,
|
||||
\]
|
||||
!et
|
||||
|
||||
with $t$ a finite positive number.
|
||||
|
||||
We see that Ridge regression is nothing but the standard
|
||||
OLS with a modified diagonal term added to $\bm{X}^T\bm{X}$. The
|
||||
consequences, in particular for our discussion of the bias-variance tradeoff
|
||||
are rather interesting.
|
||||
|
||||
Furthermore, if we use the result above in terms of the SVD decomposition (our analysis was done for the OLS method), we had
|
||||
!bt
|
||||
\[
|
||||
(\bm{X}\bm{X}^T)\bm{U} = \bm{U}\bm{D}.
|
||||
\]
|
||||
!et
|
||||
|
||||
We can analyse the OLS solutions in terms of the eigenvectors (the columns) of the right singular value matrix $\bm{U}$ as
|
||||
!bt
|
||||
\[
|
||||
\bm{X}\bm{\beta} = \bm{X}\left(\bm{V}\bm{D}\bm{V}^T \right)^{-1}\bm{X}^T\bm{y}=\bm{U\Sigma V^T}\left(\bm{V}\bm{D}\bm{V}^T \right)^{-1}(\bm{U\Sigma V^T})^T\bm{y}=\bm{U}\bm{U}^T\bm{y}
|
||||
\]
|
||||
!et
|
||||
|
||||
|
||||
For Ridge regression this becomes
|
||||
|
||||
!bt
|
||||
\[
|
||||
\bm{X}\bm{\beta}^{\mathrm{Ridge}} = \bm{U\Sigma V^T}\left(\bm{V}\bm{D}\bm{V}^T+\lambda\bm{I} \right)^{-1}(\bm{U\Sigma V^T})^T\bm{y}=\sum_{j=0}^{p-1}\bm{u}_j\bm{u}_j^T\frac{\sigma_j^2}{\sigma_j^2+\lambda}\bm{y},
|
||||
\]
|
||||
!et
|
||||
|
||||
with the vectors $\bm{u}_j$ being the columns of $\bm{U}$.
|
||||
|
||||
!split
|
||||
===== Interpreting the Ridge results =====
|
||||
|
||||
Since $\lambda \geq 0$, it means that compared to OLS, we have
|
||||
|
||||
!bt
|
||||
\[
|
||||
\frac{\sigma_j^2}{\sigma_j^2+\lambda} \leq 1.
|
||||
\]
|
||||
!et
|
||||
|
||||
Ridge regression finds the coordinates of $\bm{y}$ with respect to the
|
||||
orthonormal basis $\bm{U}$, it then shrinks the coordinates by
|
||||
$\frac{\sigma_j^2}{\sigma_j^2+\lambda}$. Recall that the SVD has
|
||||
eigenvalues ordered in a descending way, that is $\sigma_i \geq
|
||||
\sigma_{i+1}$.
|
||||
|
||||
For small eigenvalues $\sigma_i$ it means that their contributions become less important, a fact which can be used to reduce the number of degrees of freedom.
|
||||
Actually, calculating the variance of $\bm{X}\bm{v}_j$ shows that this quantity is equal to $\sigma_j^2/n$.
|
||||
With a parameter $\lambda$ we can thus shrink the role of specific parameters.
|
||||
|
||||
|
||||
!split
|
||||
===== More interpretations =====
|
||||
|
||||
For the sake of simplicity, let us assume that the design matrix is orthonormal, that is
|
||||
|
||||
!bt
|
||||
\[
|
||||
\bm{X}^T\bm{X}=(\bm{X}^T\bm{X})^{-1} =\bm{I}.
|
||||
\]
|
||||
!et
|
||||
|
||||
In this case the standard OLS results in
|
||||
!bt
|
||||
\[
|
||||
\bm{\beta}^{\mathrm{OLS}} = \bm{X}^T\bm{y}=\sum_{i=0}^{p-1}\bm{u}_j\bm{u}_j^T\bm{y},
|
||||
\]
|
||||
!et
|
||||
|
||||
and
|
||||
|
||||
!bt
|
||||
\[
|
||||
\bm{\beta}^{\mathrm{Ridge}} = \left(\bm{I}+\lambda\bm{I}\right)^{-1}\bm{X}^T\bm{y}=\left(1+\lambda\right)^{-1}\bm{\beta}^{\mathrm{OLS}},
|
||||
\]
|
||||
!et
|
||||
|
||||
that is the Ridge estimator scales the OLS estimator by the inverse of a factor $1+\lambda$, and
|
||||
the Ridge estimator converges to zero when the hyperparameter goes to
|
||||
infinity.
|
||||
|
||||
We will come back to more interpreations after we have gone through some of the statistical analysis part.
|
||||
|
||||
For more discussions of Ridge and Lasso regression, "Wessel van Wieringen's":"https://arxiv.org/abs/1509.09169" article is highly recommended.
|
||||
Similarly, "Mehta et al's article":"https://arxiv.org/abs/1803.08823" is also recommended.
|
||||
|
||||
|
||||
!split
|
||||
===== A better understanding of regularization =====
|
||||
|
||||
The parameter $\lambda$ that we have introduced in the Ridge (and
|
||||
Lasso as well) regression is often called a regularization parameter
|
||||
or shrinkage parameter. It is common to call it a hyperparameter. What does it mean mathemtically?
|
||||
|
||||
Here we will first look at how to analyze the difference between the
|
||||
standard OLS equations and the Ridge expressions in terms of a linear
|
||||
algebra analysis using the SVD algorithm. Thereafter, we will link
|
||||
(see the material on the bias-variance tradeoff below) these
|
||||
observation to the statisical analysis of the results. In particular
|
||||
we consider how the variance of the parameters $\bm{\beta}$ is
|
||||
affected by changing the parameter $\lambda$.
|
||||
|
||||
!split
|
||||
===== Decomposing the OLS and Ridge expressions =====
|
||||
|
||||
We have our design matrix
|
||||
$\bm{X}\in {\mathbb{R}}^{n\times p}$. With the SVD we decompose it as
|
||||
|
||||
!bt
|
||||
\[
|
||||
\bm{X} = \bm{U\Sigma V^T},
|
||||
\]
|
||||
!et
|
||||
|
||||
with $\bm{U}\in {\mathbb{R}}^{n\times n}$, $\bm{\Sigma}\in {\mathbb{R}}^{n\times p}$
|
||||
and $\bm{V}\in {\mathbb{R}}^{p\times p}$.
|
||||
|
||||
The matrices $\bm{U}$ and $\bm{V}$ are unitary/orthonormal matrices, that is in case the matrices are real we have $\bm{U}^T\bm{U}=\bm{U}\bm{U}^T=\bm{I}$ and $\bm{V}^T\bm{V}=\bm{V}\bm{V}^T=\bm{I}$.
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== Introducing the Covariance and Correlation functions =====
|
||||
|
||||
Before we discuss the link between for example Ridge regression and the singular value decomposition, we need to remind ourselves about
|
||||
the definition of the covariance and the correlation function. These are quantities
|
||||
the definition of the covariance and the correlation function. These are quantities that play a central role in machine learning methods.
|
||||
|
||||
Suppose we have defined two vectors
|
||||
$\hat{x}$ and $\hat{y}$ with $n$ elements each. The covariance matrix $\bm{C}$ is defined as
|
||||
@@ -2376,6 +2156,226 @@ where we wrote $$\bm{C}[\bm{x}_0,\bm{x}_1] = \bm{C}[\bm{x}]$$ to indicate that t
|
||||
It is easy to generalize this to a matrix $\bm{X}\in {\mathbb{R}}^{n\times p}$.
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== Ridge and LASSO Regression =====
|
||||
|
||||
Let us remind ourselves about the expression for the standard Mean Squared Error (MSE) which we used to define our cost function and the equations for the ordinary least squares (OLS) method, that is
|
||||
our optimization problem is
|
||||
!bt
|
||||
\[
|
||||
{\displaystyle \min_{\bm{\beta}\in {\mathbb{R}}^{p}}}\frac{1}{n}\left\{\left(\bm{y}-\bm{X}\bm{\beta}\right)^T\left(\bm{y}-\bm{X}\bm{\beta}\right)\right\}.
|
||||
\]
|
||||
!et
|
||||
or we can state it as
|
||||
!bt
|
||||
\[
|
||||
{\displaystyle \min_{\bm{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\sum_{i=0}^{n-1}\left(y_i-\tilde{y}_i\right)^2=\frac{1}{n}\vert\vert \bm{y}-\bm{X}\bm{\beta}\vert\vert_2^2,
|
||||
\]
|
||||
!et
|
||||
where we have used the definition of a norm-2 vector, that is
|
||||
!bt
|
||||
\[
|
||||
\vert\vert \bm{x}\vert\vert_2 = \sqrt{\sum_i x_i^2}.
|
||||
\]
|
||||
!et
|
||||
|
||||
By minimizing the above equation with respect to the parameters
|
||||
$\bm{\beta}$ we could then obtain an analytical expression for the
|
||||
parameters $\bm{\beta}$. We can add a regularization parameter $\lambda$ by
|
||||
defining a new cost function to be optimized, that is
|
||||
|
||||
!bt
|
||||
\[
|
||||
{\displaystyle \min_{\bm{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \bm{y}-\bm{X}\bm{\beta}\vert\vert_2^2+\lambda\vert\vert \bm{\beta}\vert\vert_2^2
|
||||
\]
|
||||
!et
|
||||
|
||||
which leads to the Ridge regression minimization problem where we
|
||||
require that $\vert\vert \bm{\beta}\vert\vert_2^2\le t$, where $t$ is
|
||||
a finite number larger than zero. By defining
|
||||
|
||||
!bt
|
||||
\[
|
||||
C(\bm{X},\bm{\beta})=\frac{1}{n}\vert\vert \bm{y}-\bm{X}\bm{\beta}\vert\vert_2^2+\lambda\vert\vert \bm{\beta}\vert\vert_1,
|
||||
\]
|
||||
!et
|
||||
|
||||
we have a new optimization equation
|
||||
!bt
|
||||
\[
|
||||
{\displaystyle \min_{\bm{\beta}\in
|
||||
{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \bm{y}-\bm{X}\bm{\beta}\vert\vert_2^2+\lambda\vert\vert \bm{\beta}\vert\vert_1
|
||||
\]
|
||||
!et
|
||||
which leads to Lasso regression. Lasso stands for least absolute shrinkage and selection operator.
|
||||
|
||||
Here we have defined the norm-1 as
|
||||
!bt
|
||||
\[
|
||||
\vert\vert \bm{x}\vert\vert_1 = \sum_i \vert x_i\vert.
|
||||
\]
|
||||
!et
|
||||
|
||||
|
||||
!split
|
||||
===== More on Ridge Regression =====
|
||||
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
|
||||
!bt
|
||||
\[
|
||||
C(\bm{X},\bm{\beta})=\frac{1}{n}\left\{(\bm{y}-\bm{X}\bm{\beta})^T(\bm{y}-\bm{X}\bm{\beta})\right\}+\lambda\bm{\beta}^T\bm{\beta},
|
||||
\]
|
||||
!et
|
||||
|
||||
by taking the derivatives with respect to $\bm{\beta}$ we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of $\lambda$ does not suffer from singularity problems. We obtain
|
||||
|
||||
!bt
|
||||
\[
|
||||
\bm{\beta}^{\mathrm{Ridge}} = \left(\bm{X}^T\bm{X}+\lambda\bm{I}\right)^{-1}\bm{X}^T\bm{y},
|
||||
\]
|
||||
!et
|
||||
|
||||
with $\bm{I}$ being a $p\times p$ identity matrix with the constraint that
|
||||
|
||||
!bt
|
||||
\[
|
||||
\sum_{i=0}^{p-1} \beta_i^2 \leq t,
|
||||
\]
|
||||
!et
|
||||
|
||||
with $t$ a finite positive number.
|
||||
|
||||
We see that Ridge regression is nothing but the standard
|
||||
OLS with a modified diagonal term added to $\bm{X}^T\bm{X}$. The
|
||||
consequences, in particular for our discussion of the bias-variance tradeoff
|
||||
are rather interesting.
|
||||
|
||||
Furthermore, if we use the result above in terms of the SVD decomposition (our analysis was done for the OLS method), we had
|
||||
!bt
|
||||
\[
|
||||
(\bm{X}\bm{X}^T)\bm{U} = \bm{U}\bm{D}.
|
||||
\]
|
||||
!et
|
||||
|
||||
We can analyse the OLS solutions in terms of the eigenvectors (the columns) of the right singular value matrix $\bm{U}$ as
|
||||
!bt
|
||||
\[
|
||||
\bm{X}\bm{\beta} = \bm{X}\left(\bm{V}\bm{D}\bm{V}^T \right)^{-1}\bm{X}^T\bm{y}=\bm{U\Sigma V^T}\left(\bm{V}\bm{D}\bm{V}^T \right)^{-1}(\bm{U\Sigma V^T})^T\bm{y}=\bm{U}\bm{U}^T\bm{y}
|
||||
\]
|
||||
!et
|
||||
|
||||
|
||||
For Ridge regression this becomes
|
||||
|
||||
!bt
|
||||
\[
|
||||
\bm{X}\bm{\beta}^{\mathrm{Ridge}} = \bm{U\Sigma V^T}\left(\bm{V}\bm{D}\bm{V}^T+\lambda\bm{I} \right)^{-1}(\bm{U\Sigma V^T})^T\bm{y}=\sum_{j=0}^{p-1}\bm{u}_j\bm{u}_j^T\frac{\sigma_j^2}{\sigma_j^2+\lambda}\bm{y},
|
||||
\]
|
||||
!et
|
||||
|
||||
with the vectors $\bm{u}_j$ being the columns of $\bm{U}$.
|
||||
|
||||
!split
|
||||
===== Interpreting the Ridge results =====
|
||||
|
||||
Since $\lambda \geq 0$, it means that compared to OLS, we have
|
||||
|
||||
!bt
|
||||
\[
|
||||
\frac{\sigma_j^2}{\sigma_j^2+\lambda} \leq 1.
|
||||
\]
|
||||
!et
|
||||
|
||||
Ridge regression finds the coordinates of $\bm{y}$ with respect to the
|
||||
orthonormal basis $\bm{U}$, it then shrinks the coordinates by
|
||||
$\frac{\sigma_j^2}{\sigma_j^2+\lambda}$. Recall that the SVD has
|
||||
eigenvalues ordered in a descending way, that is $\sigma_i \geq
|
||||
\sigma_{i+1}$.
|
||||
|
||||
For small eigenvalues $\sigma_i$ it means that their contributions become less important, a fact which can be used to reduce the number of degrees of freedom.
|
||||
Actually, calculating the variance of $\bm{X}\bm{v}_j$ shows that this quantity is equal to $\sigma_j^2/n$.
|
||||
With a parameter $\lambda$ we can thus shrink the role of specific parameters.
|
||||
|
||||
|
||||
!split
|
||||
===== More interpretations =====
|
||||
|
||||
For the sake of simplicity, let us assume that the design matrix is orthonormal, that is
|
||||
|
||||
!bt
|
||||
\[
|
||||
\bm{X}^T\bm{X}=(\bm{X}^T\bm{X})^{-1} =\bm{I}.
|
||||
\]
|
||||
!et
|
||||
|
||||
In this case the standard OLS results in
|
||||
!bt
|
||||
\[
|
||||
\bm{\beta}^{\mathrm{OLS}} = \bm{X}^T\bm{y}=\sum_{i=0}^{p-1}\bm{u}_j\bm{u}_j^T\bm{y},
|
||||
\]
|
||||
!et
|
||||
|
||||
and
|
||||
|
||||
!bt
|
||||
\[
|
||||
\bm{\beta}^{\mathrm{Ridge}} = \left(\bm{I}+\lambda\bm{I}\right)^{-1}\bm{X}^T\bm{y}=\left(1+\lambda\right)^{-1}\bm{\beta}^{\mathrm{OLS}},
|
||||
\]
|
||||
!et
|
||||
|
||||
that is the Ridge estimator scales the OLS estimator by the inverse of a factor $1+\lambda$, and
|
||||
the Ridge estimator converges to zero when the hyperparameter goes to
|
||||
infinity.
|
||||
|
||||
We will come back to more interpreations after we have gone through some of the statistical analysis part.
|
||||
|
||||
For more discussions of Ridge and Lasso regression, "Wessel van Wieringen's":"https://arxiv.org/abs/1509.09169" article is highly recommended.
|
||||
Similarly, "Mehta et al's article":"https://arxiv.org/abs/1803.08823" is also recommended.
|
||||
|
||||
|
||||
!split
|
||||
===== A better understanding of regularization =====
|
||||
|
||||
The parameter $\lambda$ that we have introduced in the Ridge (and
|
||||
Lasso as well) regression is often called a regularization parameter
|
||||
or shrinkage parameter. It is common to call it a hyperparameter. What does it mean mathemtically?
|
||||
|
||||
Here we will first look at how to analyze the difference between the
|
||||
standard OLS equations and the Ridge expressions in terms of a linear
|
||||
algebra analysis using the SVD algorithm. Thereafter, we will link
|
||||
(see the material on the bias-variance tradeoff below) these
|
||||
observation to the statisical analysis of the results. In particular
|
||||
we consider how the variance of the parameters $\bm{\beta}$ is
|
||||
affected by changing the parameter $\lambda$.
|
||||
|
||||
!split
|
||||
===== Decomposing the OLS and Ridge expressions =====
|
||||
|
||||
We have our design matrix
|
||||
$\bm{X}\in {\mathbb{R}}^{n\times p}$. With the SVD we decompose it as
|
||||
|
||||
!bt
|
||||
\[
|
||||
\bm{X} = \bm{U\Sigma V^T},
|
||||
\]
|
||||
!et
|
||||
|
||||
with $\bm{U}\in {\mathbb{R}}^{n\times n}$, $\bm{\Sigma}\in {\mathbb{R}}^{n\times p}$
|
||||
and $\bm{V}\in {\mathbb{R}}^{p\times p}$.
|
||||
|
||||
The matrices $\bm{U}$ and $\bm{V}$ are unitary/orthonormal matrices, that is in case the matrices are real we have $\bm{U}^T\bm{U}=\bm{U}\bm{U}^T=\bm{I}$ and $\bm{V}^T\bm{V}=\bm{V}\bm{V}^T=\bm{I}$.
|
||||
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== Mathematical Properties =====
|
||||
|
||||
|
||||
Reference in New Issue
Block a user