final update of week 35
This commit is contained in:
@@ -248,6 +248,10 @@ Automatically generated HTML file from DocOnce source
|
||||
None,
|
||||
'interpreting-the-ridge-results'),
|
||||
('More interpretations', 2, None, 'more-interpretations'),
|
||||
('Deriving the Lasso Regression Equations',
|
||||
2,
|
||||
None,
|
||||
'deriving-the-lasso-regression-equations'),
|
||||
('Exercises for week 36, September 6-10',
|
||||
2,
|
||||
None,
|
||||
@@ -364,9 +368,10 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs065.html#deriving-the-ridge-regression-equations" style="font-size: 80%;"><b>Deriving the Ridge Regression Equations</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs066.html#interpreting-the-ridge-results" style="font-size: 80%;"><b>Interpreting the Ridge results</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs067.html#more-interpretations" style="font-size: 80%;"><b>More interpretations</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs068.html#exercises-for-week-36-september-6-10" style="font-size: 80%;"><b>Exercises for week 36, September 6-10</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs068.html#exercise-1-adding-ridge-and-lasso-regression" style="font-size: 80%;"><b>Exercise 1: Adding Ridge and Lasso Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs068.html#exercise-linear-regression-for-a-two-dimensional-function" style="font-size: 80%;"> Exercise: Linear Regression for a two-dimensional function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs068.html#deriving-the-lasso-regression-equations" style="font-size: 80%;"><b>Deriving the Lasso Regression Equations</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs069.html#exercises-for-week-36-september-6-10" style="font-size: 80%;"><b>Exercises for week 36, September 6-10</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs069.html#exercise-1-adding-ridge-and-lasso-regression" style="font-size: 80%;"><b>Exercise 1: Adding Ridge and Lasso Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs069.html#exercise-linear-regression-for-a-two-dimensional-function" style="font-size: 80%;"> Exercise: Linear Regression for a two-dimensional function</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -425,7 +430,7 @@ MathJax.Hub.Config({
|
||||
<li><a href="._week35-bs008.html">9</a></li>
|
||||
<li><a href="._week35-bs009.html">10</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._week35-bs068.html">69</a></li>
|
||||
<li><a href="._week35-bs069.html">70</a></li>
|
||||
<li><a href="._week35-bs001.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -2899,16 +2899,16 @@ $$
|
||||
<h2 id="deriving-the-ridge-regression-equations">Deriving the Ridge Regression Equations </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
Using the matrix-vector expression for Ridge regression and dropping the parameter \( 1/n \) in front of the standard means squared error equation, we have
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
by taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
and
|
||||
taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of \( \lambda \) does not suffer from singularity problems. We obtain
|
||||
the optimal parameters
|
||||
@@ -3039,6 +3039,45 @@ Similarly, <a href="https://arxiv.org/abs/1803.08823" target="_blank">Mehta et a
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="deriving-the-lasso-regression-equations">Deriving the Lasso Regression Equations </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Lasso regression and dropping the parameter \( 1/n \) in front of the standard means squared error equation, we have the following <b>cost</b> function
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\vert\vert\boldsymbol{\beta}\vert\vert_1,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
Taking the derivative with respect to \( \boldsymbol{\beta} \) and recalling that the derivative of the absolute value is (we drop the boldfaced vector symbol for simplicty)
|
||||
<p> <br>
|
||||
$$
|
||||
\frac{d \vert \beta\vert}{d \boldsymbol{\beta}}=\mathrm{sgn}(\boldsymbol{\beta})=\left\{\begin{array}{cc} 1 & \beta > 0 \\ 0 & \beta =0\\-1 & \beta < 0, \end{array}\right.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
we have that the derivative of the cost function is
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\frac{\partial C(\boldsymbol{X},\boldsymbol{\beta})}{\partial \boldsymbol{\beta}}=-2\boldsymbol{X}^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})+\lambda sgn(\boldsymbol{\beta})=0,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
and reordering we have
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}\boldsymbol{\beta})+\lambda sgn(\boldsymbol{\beta})=2\boldsymbol{X}^T(\boldsymbol{y}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
This equation does not lead to a nice analytical equation as in either Ridge regression or ordinary least squares. This equation can however be solved by using standard convex optimization algorithms using for example the Python package <a href="https://cvxopt.org/" target="_blank">CVXOPT</a>. We will discuss this later.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="exercises-for-week-36-september-6-10">Exercises for week 36, September 6-10 </h2>
|
||||
|
||||
|
||||
@@ -268,6 +268,10 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
None,
|
||||
'interpreting-the-ridge-results'),
|
||||
('More interpretations', 2, None, 'more-interpretations'),
|
||||
('Deriving the Lasso Regression Equations',
|
||||
2,
|
||||
None,
|
||||
'deriving-the-lasso-regression-equations'),
|
||||
('Exercises for week 36, September 6-10',
|
||||
2,
|
||||
None,
|
||||
@@ -2863,14 +2867,14 @@ $$
|
||||
<h2 id="deriving-the-ridge-regression-equations">Deriving the Ridge Regression Equations </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
Using the matrix-vector expression for Ridge regression and dropping the parameter \( 1/n \) in front of the standard means squared error equation, we have
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
$$
|
||||
|
||||
<p>
|
||||
by taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
and
|
||||
taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of \( \lambda \) does not suffer from singularity problems. We obtain
|
||||
the optimal parameters
|
||||
@@ -2984,6 +2988,37 @@ Similarly, <a href="https://arxiv.org/abs/1803.08823" target="_blank">Mehta et a
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="deriving-the-lasso-regression-equations">Deriving the Lasso Regression Equations </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Lasso regression and dropping the parameter \( 1/n \) in front of the standard means squared error equation, we have the following <b>cost</b> function
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\vert\vert\boldsymbol{\beta}\vert\vert_1,
|
||||
$$
|
||||
|
||||
<p>
|
||||
Taking the derivative with respect to \( \boldsymbol{\beta} \) and recalling that the derivative of the absolute value is (we drop the boldfaced vector symbol for simplicty)
|
||||
$$
|
||||
\frac{d \vert \beta\vert}{d \boldsymbol{\beta}}=\mathrm{sgn}(\boldsymbol{\beta})=\left\{\begin{array}{cc} 1 & \beta > 0 \\ 0 & \beta =0\\-1 & \beta < 0, \end{array}\right.
|
||||
$$
|
||||
|
||||
we have that the derivative of the cost function is
|
||||
|
||||
$$
|
||||
\frac{\partial C(\boldsymbol{X},\boldsymbol{\beta})}{\partial \boldsymbol{\beta}}=-2\boldsymbol{X}^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})+\lambda sgn(\boldsymbol{\beta})=0,
|
||||
$$
|
||||
|
||||
and reordering we have
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}\boldsymbol{\beta})+\lambda sgn(\boldsymbol{\beta})=2\boldsymbol{X}^T(\boldsymbol{y}.
|
||||
$$
|
||||
|
||||
This equation does not lead to a nice analytical equation as in either Ridge regression or ordinary least squares. This equation can however be solved by using standard convex optimization algorithms using for example the Python package <a href="https://cvxopt.org/" target="_blank">CVXOPT</a>. We will discuss this later.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="exercises-for-week-36-september-6-10">Exercises for week 36, September 6-10 </h2>
|
||||
|
||||
<p>
|
||||
|
||||
@@ -273,6 +273,10 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
None,
|
||||
'interpreting-the-ridge-results'),
|
||||
('More interpretations', 2, None, 'more-interpretations'),
|
||||
('Deriving the Lasso Regression Equations',
|
||||
2,
|
||||
None,
|
||||
'deriving-the-lasso-regression-equations'),
|
||||
('Exercises for week 36, September 6-10',
|
||||
2,
|
||||
None,
|
||||
@@ -2868,14 +2872,14 @@ $$
|
||||
<h2 id="deriving-the-ridge-regression-equations">Deriving the Ridge Regression Equations </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Ridge regression,
|
||||
Using the matrix-vector expression for Ridge regression and dropping the parameter \( 1/n \) in front of the standard means squared error equation, we have
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\frac{1}{n}\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\boldsymbol{\beta}^T\boldsymbol{\beta},
|
||||
$$
|
||||
|
||||
<p>
|
||||
by taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
and
|
||||
taking the derivatives with respect to \( \boldsymbol{\beta} \) we obtain then
|
||||
a slightly modified matrix inversion problem which for finite values
|
||||
of \( \lambda \) does not suffer from singularity problems. We obtain
|
||||
the optimal parameters
|
||||
@@ -2989,6 +2993,37 @@ Similarly, <a href="https://arxiv.org/abs/1803.08823" target="_blank">Mehta et a
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="deriving-the-lasso-regression-equations">Deriving the Lasso Regression Equations </h2>
|
||||
|
||||
<p>
|
||||
Using the matrix-vector expression for Lasso regression and dropping the parameter \( 1/n \) in front of the standard means squared error equation, we have the following <b>cost</b> function
|
||||
|
||||
$$
|
||||
C(\boldsymbol{X},\boldsymbol{\beta})=\left\{(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})\right\}+\lambda\vert\vert\boldsymbol{\beta}\vert\vert_1,
|
||||
$$
|
||||
|
||||
<p>
|
||||
Taking the derivative with respect to \( \boldsymbol{\beta} \) and recalling that the derivative of the absolute value is (we drop the boldfaced vector symbol for simplicty)
|
||||
$$
|
||||
\frac{d \vert \beta\vert}{d \boldsymbol{\beta}}=\mathrm{sgn}(\boldsymbol{\beta})=\left\{\begin{array}{cc} 1 & \beta > 0 \\ 0 & \beta =0\\-1 & \beta < 0, \end{array}\right.
|
||||
$$
|
||||
|
||||
we have that the derivative of the cost function is
|
||||
|
||||
$$
|
||||
\frac{\partial C(\boldsymbol{X},\boldsymbol{\beta})}{\partial \boldsymbol{\beta}}=-2\boldsymbol{X}^T(\boldsymbol{y}-\boldsymbol{X}\boldsymbol{\beta})+\lambda sgn(\boldsymbol{\beta})=0,
|
||||
$$
|
||||
|
||||
and reordering we have
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}\boldsymbol{\beta})+\lambda sgn(\boldsymbol{\beta})=2\boldsymbol{X}^T(\boldsymbol{y}.
|
||||
$$
|
||||
|
||||
This equation does not lead to a nice analytical equation as in either Ridge regression or ordinary least squares. This equation can however be solved by using standard convex optimization algorithms using for example the Python package <a href="https://cvxopt.org/" target="_blank">CVXOPT</a>. We will discuss this later.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="exercises-for-week-36-september-6-10">Exercises for week 36, September 6-10 </h2>
|
||||
|
||||
<p>
|
||||
|
||||
Binary file not shown.
@@ -3790,7 +3790,7 @@
|
||||
"source": [
|
||||
"## Deriving the Ridge Regression Equations\n",
|
||||
"\n",
|
||||
"Using the matrix-vector expression for Ridge regression,"
|
||||
"Using the matrix-vector expression for Ridge regression and dropping the parameter $1/n$ in front of the standard means squared error equation, we have"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -3798,7 +3798,7 @@
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"C(\\boldsymbol{X},\\boldsymbol{\\beta})=\\frac{1}{n}\\left\\{(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta})^T(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta})\\right\\}+\\lambda\\boldsymbol{\\beta}^T\\boldsymbol{\\beta},\n",
|
||||
"C(\\boldsymbol{X},\\boldsymbol{\\beta})=\\left\\{(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta})^T(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta})\\right\\}+\\lambda\\boldsymbol{\\beta}^T\\boldsymbol{\\beta},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
@@ -3806,7 +3806,8 @@
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"by taking the derivatives with respect to $\\boldsymbol{\\beta}$ we obtain then\n",
|
||||
"and \n",
|
||||
"taking the derivatives with respect to $\\boldsymbol{\\beta}$ we obtain then\n",
|
||||
"a slightly modified matrix inversion problem which for finite values\n",
|
||||
"of $\\lambda$ does not suffer from singularity problems. We obtain\n",
|
||||
"the optimal parameters"
|
||||
@@ -3991,8 +3992,73 @@
|
||||
"For more discussions of Ridge and Lasso regression, [Wessel van Wieringen's](https://arxiv.org/abs/1509.09169) article is highly recommended.\n",
|
||||
"Similarly, [Mehta et al's article](https://arxiv.org/abs/1803.08823) is also recommended.\n",
|
||||
"\n",
|
||||
"## Deriving the Lasso Regression Equations\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Using the matrix-vector expression for Lasso regression and dropping the parameter $1/n$ in front of the standard means squared error equation, we have the following **cost** function"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"C(\\boldsymbol{X},\\boldsymbol{\\beta})=\\left\\{(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta})^T(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta})\\right\\}+\\lambda\\vert\\vert\\boldsymbol{\\beta}\\vert\\vert_1,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Taking the derivative with respect to $\\boldsymbol{\\beta}$ and recalling that the derivative of the absolute value is (we drop the boldfaced vector symbol for simplicty)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\frac{d \\vert \\beta\\vert}{d \\boldsymbol{\\beta}}=\\mathrm{sgn}(\\boldsymbol{\\beta})=\\left\\{\\begin{array}{cc} 1 & \\beta > 0 \\\\ 0 & \\beta =0\\\\-1 & \\beta < 0, \\end{array}\\right.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"we have that the derivative of the cost function is"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\frac{\\partial C(\\boldsymbol{X},\\boldsymbol{\\beta})}{\\partial \\boldsymbol{\\beta}}=-2\\boldsymbol{X}^T(\\boldsymbol{y}-\\boldsymbol{X}\\boldsymbol{\\beta})+\\lambda sgn(\\boldsymbol{\\beta})=0,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and reordering we have"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{X}^T\\boldsymbol{X}\\boldsymbol{\\beta})+\\lambda sgn(\\boldsymbol{\\beta})=2\\boldsymbol{X}^T(\\boldsymbol{y}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"This equation does not lead to a nice analytical equation as in either Ridge regression or ordinary least squares. This equation can however be solved by using standard convex optimization algorithms using for example the Python package [CVXOPT](https://cvxopt.org/). We will discuss this later. \n",
|
||||
"\n",
|
||||
"## Exercises for week 36, September 6-10\n",
|
||||
"\n",
|
||||
|
||||
Reference in New Issue
Block a user