final update of week 35
This commit is contained in:
@@ -248,6 +248,10 @@ Automatically generated HTML file from DocOnce source
|
||||
None,
|
||||
'interpreting-the-ridge-results'),
|
||||
('More interpretations', 2, None, 'more-interpretations'),
|
||||
('Deriving the Lasso Regression Equations',
|
||||
2,
|
||||
None,
|
||||
'deriving-the-lasso-regression-equations'),
|
||||
('Exercises for week 36, September 6-10',
|
||||
2,
|
||||
None,
|
||||
@@ -364,9 +368,10 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs065.html#deriving-the-ridge-regression-equations" style="font-size: 80%;"><b>Deriving the Ridge Regression Equations</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs066.html#interpreting-the-ridge-results" style="font-size: 80%;"><b>Interpreting the Ridge results</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs067.html#more-interpretations" style="font-size: 80%;"><b>More interpretations</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs068.html#exercises-for-week-36-september-6-10" style="font-size: 80%;"><b>Exercises for week 36, September 6-10</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs068.html#exercise-1-adding-ridge-and-lasso-regression" style="font-size: 80%;"><b>Exercise 1: Adding Ridge and Lasso Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs068.html#exercise-linear-regression-for-a-two-dimensional-function" style="font-size: 80%;"> Exercise: Linear Regression for a two-dimensional function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs068.html#deriving-the-lasso-regression-equations" style="font-size: 80%;"><b>Deriving the Lasso Regression Equations</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs069.html#exercises-for-week-36-september-6-10" style="font-size: 80%;"><b>Exercises for week 36, September 6-10</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs069.html#exercise-1-adding-ridge-and-lasso-regression" style="font-size: 80%;"><b>Exercise 1: Adding Ridge and Lasso Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs069.html#exercise-linear-regression-for-a-two-dimensional-function" style="font-size: 80%;"> Exercise: Linear Regression for a two-dimensional function</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -425,7 +430,7 @@ MathJax.Hub.Config({
|
||||
<li><a href="._week35-bs008.html">9</a></li>
|
||||
<li><a href="._week35-bs009.html">10</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._week35-bs068.html">69</a></li>
|
||||
<li><a href="._week35-bs069.html">70</a></li>
|
||||
<li><a href="._week35-bs001.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -2899,16 +2899,16 @@ $$
|
||||
<h2 id="deriving-the-ridge-regression-equations">Deriving the Ridge Regression Equations </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
Using the matrix-vector expression for Ridge regression and dropping the parameter \( 1/n \) in front of the standard means squared error equation, we have
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
by taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
and
|
||||
taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of \( \lambda \) does not suffer from singularity problems. We obtain
|
||||
the optimal parameters
|
||||
@@ -3039,6 +3039,45 @@ Similarly, <a href="https://arxiv.org/abs/1803.08823" target="_blank">Mehta et a
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="deriving-the-lasso-regression-equations">Deriving the Lasso Regression Equations </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Lasso regression and dropping the parameter \( 1/n \) in front of the standard means squared error equation, we have the following <b>cost</b> function
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\vert\vert\boldsymbol{\beta}\vert\vert_1,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
Taking the derivative with respect to \( \boldsymbol{\beta} \) and recalling that the derivative of the absolute value is (we drop the boldfaced vector symbol for simplicty)
|
||||
<p> <br>
|
||||
$$
|
||||
\frac{d \vert \beta\vert}{d \boldsymbol{\beta}}=\mathrm{sgn}(\boldsymbol{\beta})=\left\{\begin{array}{cc} 1 & \beta > 0 \\ 0 & \beta =0\\-1 & \beta < 0, \end{array}\right.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
we have that the derivative of the cost function is
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\frac{\partial C(\boldsymbol{X},\boldsymbol{\beta})}{\partial \boldsymbol{\beta}}=-2\boldsymbol{X}^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})+\lambda sgn(\boldsymbol{\beta})=0,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
and reordering we have
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}\boldsymbol{\beta})+\lambda sgn(\boldsymbol{\beta})=2\boldsymbol{X}^T(\boldsymbol{y}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
This equation does not lead to a nice analytical equation as in either Ridge regression or ordinary least squares. This equation can however be solved by using standard convex optimization algorithms using for example the Python package <a href="https://cvxopt.org/" target="_blank">CVXOPT</a>. We will discuss this later.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="exercises-for-week-36-september-6-10">Exercises for week 36, September 6-10 </h2>
|
||||
|
||||
|
||||
@@ -268,6 +268,10 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
None,
|
||||
'interpreting-the-ridge-results'),
|
||||
('More interpretations', 2, None, 'more-interpretations'),
|
||||
('Deriving the Lasso Regression Equations',
|
||||
2,
|
||||
None,
|
||||
'deriving-the-lasso-regression-equations'),
|
||||
('Exercises for week 36, September 6-10',
|
||||
2,
|
||||
None,
|
||||
@@ -2863,14 +2867,14 @@ $$
|
||||
<h2 id="deriving-the-ridge-regression-equations">Deriving the Ridge Regression Equations </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
Using the matrix-vector expression for Ridge regression and dropping the parameter \( 1/n \) in front of the standard means squared error equation, we have
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
$$
|
||||
|
||||
<p>
|
||||
by taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
and
|
||||
taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of \( \lambda \) does not suffer from singularity problems. We obtain
|
||||
the optimal parameters
|
||||
@@ -2984,6 +2988,37 @@ Similarly, <a href="https://arxiv.org/abs/1803.08823" target="_blank">Mehta et a
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="deriving-the-lasso-regression-equations">Deriving the Lasso Regression Equations </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Lasso regression and dropping the parameter \( 1/n \) in front of the standard means squared error equation, we have the following <b>cost</b> function
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\vert\vert\boldsymbol{\beta}\vert\vert_1,
|
||||
$$
|
||||
|
||||
<p>
|
||||
Taking the derivative with respect to \( \boldsymbol{\beta} \) and recalling that the derivative of the absolute value is (we drop the boldfaced vector symbol for simplicty)
|
||||
$$
|
||||
\frac{d \vert \beta\vert}{d \boldsymbol{\beta}}=\mathrm{sgn}(\boldsymbol{\beta})=\left\{\begin{array}{cc} 1 & \beta > 0 \\ 0 & \beta =0\\-1 & \beta < 0, \end{array}\right.
|
||||
$$
|
||||
|
||||
we have that the derivative of the cost function is
|
||||
|
||||
$$
|
||||
\frac{\partial C(\boldsymbol{X},\boldsymbol{\beta})}{\partial \boldsymbol{\beta}}=-2\boldsymbol{X}^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})+\lambda sgn(\boldsymbol{\beta})=0,
|
||||
$$
|
||||
|
||||
and reordering we have
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}\boldsymbol{\beta})+\lambda sgn(\boldsymbol{\beta})=2\boldsymbol{X}^T(\boldsymbol{y}.
|
||||
$$
|
||||
|
||||
This equation does not lead to a nice analytical equation as in either Ridge regression or ordinary least squares. This equation can however be solved by using standard convex optimization algorithms using for example the Python package <a href="https://cvxopt.org/" target="_blank">CVXOPT</a>. We will discuss this later.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="exercises-for-week-36-september-6-10">Exercises for week 36, September 6-10 </h2>
|
||||
|
||||
<p>
|
||||
|
||||
@@ -273,6 +273,10 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
None,
|
||||
'interpreting-the-ridge-results'),
|
||||
('More interpretations', 2, None, 'more-interpretations'),
|
||||
('Deriving the Lasso Regression Equations',
|
||||
2,
|
||||
None,
|
||||
'deriving-the-lasso-regression-equations'),
|
||||
('Exercises for week 36, September 6-10',
|
||||
2,
|
||||
None,
|
||||
@@ -2868,14 +2872,14 @@ $$
|
||||
<h2 id="deriving-the-ridge-regression-equations">Deriving the Ridge Regression Equations </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
Using the matrix-vector expression for Ridge regression and dropping the parameter \( 1/n \) in front of the standard means squared error equation, we have
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
$$
|
||||
|
||||
<p>
|
||||
by taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
and
|
||||
taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of \( \lambda \) does not suffer from singularity problems. We obtain
|
||||
the optimal parameters
|
||||
@@ -2989,6 +2993,37 @@ Similarly, <a href="https://arxiv.org/abs/1803.08823" target="_blank">Mehta et a
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="deriving-the-lasso-regression-equations">Deriving the Lasso Regression Equations </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Lasso regression and dropping the parameter \( 1/n \) in front of the standard means squared error equation, we have the following <b>cost</b> function
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\vert\vert\boldsymbol{\beta}\vert\vert_1,
|
||||
$$
|
||||
|
||||
<p>
|
||||
Taking the derivative with respect to \( \boldsymbol{\beta} \) and recalling that the derivative of the absolute value is (we drop the boldfaced vector symbol for simplicty)
|
||||
$$
|
||||
\frac{d \vert \beta\vert}{d \boldsymbol{\beta}}=\mathrm{sgn}(\boldsymbol{\beta})=\left\{\begin{array}{cc} 1 & \beta > 0 \\ 0 & \beta =0\\-1 & \beta < 0, \end{array}\right.
|
||||
$$
|
||||
|
||||
we have that the derivative of the cost function is
|
||||
|
||||
$$
|
||||
\frac{\partial C(\boldsymbol{X},\boldsymbol{\beta})}{\partial \boldsymbol{\beta}}=-2\boldsymbol{X}^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})+\lambda sgn(\boldsymbol{\beta})=0,
|
||||
$$
|
||||
|
||||
and reordering we have
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}\boldsymbol{\beta})+\lambda sgn(\boldsymbol{\beta})=2\boldsymbol{X}^T(\boldsymbol{y}.
|
||||
$$
|
||||
|
||||
This equation does not lead to a nice analytical equation as in either Ridge regression or ordinary least squares. This equation can however be solved by using standard convex optimization algorithms using for example the Python package <a href="https://cvxopt.org/" target="_blank">CVXOPT</a>. We will discuss this later.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="exercises-for-week-36-september-6-10">Exercises for week 36, September 6-10 </h2>
|
||||
|
||||
<p>
|
||||
|
||||
Reference in New Issue
Block a user