typos in gradient methods
This commit is contained in:
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -193,7 +193,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Sep 20, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Sep 21, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -185,11 +185,13 @@ direction of the negative gradient \( -\nabla F(\mathbf{x}) \).
|
||||
<p>
|
||||
It can be shown that if
|
||||
$$
|
||||
\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), \ \ \gamma_k > 0
|
||||
\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k),
|
||||
$$
|
||||
|
||||
with \( \gamma_k > 0 \).
|
||||
|
||||
<p>
|
||||
for \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq
|
||||
For \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq
|
||||
F(\mathbf{x}_k) \). This means that for a sufficiently small \( \gamma_k \)
|
||||
we are always moving towards smaller function values, i.e a minimum.
|
||||
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -177,7 +177,7 @@ MathJax.Hub.Config({
|
||||
<h2 id="___sec3" class="anchor">The ideal </h2>
|
||||
|
||||
<p>
|
||||
Ideally the sequence \( \{ \mathbf{x}_k \}_{k=0} \) converges to a global
|
||||
Ideally the sequence \( \{\mathbf{x}_k \}_{k=0} \) converges to a global
|
||||
minimum of the function \( F \). In general we do not know if we are in a
|
||||
global or local minimum. In the special case when \( F \) is a convex
|
||||
function, all local minima are also global minima, so in this case
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -177,7 +177,8 @@ MathJax.Hub.Config({
|
||||
<h2 id="___sec4" class="anchor">The sensitiveness of the gradient descent </h2>
|
||||
|
||||
<p>
|
||||
GD is sensitive to the choice of learning rate \( \gamma_k \). This is due
|
||||
The gradient descent method
|
||||
is sensitive to the choice of learning rate \( \gamma_k \). This is due
|
||||
to the fact that we are only guaranteed that \( F(\mathbf{x}_{k+1}) \leq
|
||||
F(\mathbf{x}_k) \) for sufficiently small \( \gamma_k \). The problem is to
|
||||
determine an optimal learning rate. If the learning rate is chosen too
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -172,54 +172,25 @@ MathJax.Hub.Config({
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0006"></a>
|
||||
<!-- !split -->
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec5" class="anchor">Gradient Descent Example </h2>
|
||||
<h2 id="___sec5" class="anchor">Convex functions </h2>
|
||||
|
||||
<p>
|
||||
We revisit now our simple linear regression example with a linear polynomial.
|
||||
Ideally we want our cost/loss function to be convex(concave).
|
||||
|
||||
<p>
|
||||
First we give the definition of a convex set: A set \( C \) in
|
||||
\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and
|
||||
all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to
|
||||
C. Geometrically this means that every point on the line segment
|
||||
connecting \( x \) and \( y \) is in \( C \) as discussed below.
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Importing various packages</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">random</span> <span style="color: #008000; font-weight: bold">import</span> random, seed
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">mpl_toolkits.mplot3d</span> <span style="color: #008000; font-weight: bold">import</span> Axes3D
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">matplotlib</span> <span style="color: #008000; font-weight: bold">import</span> cm
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">matplotlib.ticker</span> <span style="color: #008000; font-weight: bold">import</span> LinearLocator, FormatStrFormatter
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">sys</span>
|
||||
<p>
|
||||
The convex subsets of \( \mathbb{R} \) are the intervals of
|
||||
\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the
|
||||
regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
|
||||
x <span style="color: #666666">=</span> <span style="color: #666666">2*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">4+3*</span>x<span style="color: #666666">+</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
|
||||
xb <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)), x]
|
||||
beta_linreg <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb))<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>T)<span style="color: #666666">.</span>dot(y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_linreg)
|
||||
beta <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">2</span>,<span style="color: #666666">1</span>)
|
||||
|
||||
eta <span style="color: #666666">=</span> <span style="color: #666666">0.1</span>
|
||||
Niterations <span style="color: #666666">=</span> <span style="color: #666666">1000</span>
|
||||
m <span style="color: #666666">=</span> <span style="color: #666666">100</span>
|
||||
|
||||
<span style="color: #008000; font-weight: bold">for</span> <span style="color: #008000">iter</span> <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(Niterations):
|
||||
gradients <span style="color: #666666">=</span> <span style="color: #666666">2.0/</span>m<span style="color: #666666">*</span>xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>dot(beta)<span style="color: #666666">-</span>y)
|
||||
beta <span style="color: #666666">-=</span> eta<span style="color: #666666">*</span>gradients
|
||||
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta)
|
||||
xnew <span style="color: #666666">=</span> np<span style="color: #666666">.</span>array([[<span style="color: #666666">0</span>],[<span style="color: #666666">2</span>]])
|
||||
xbnew <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">2</span>,<span style="color: #666666">1</span>)), xnew]
|
||||
ypredict <span style="color: #666666">=</span> xbnew<span style="color: #666666">.</span>dot(beta)
|
||||
ypredict2 <span style="color: #666666">=</span> xbnew<span style="color: #666666">.</span>dot(beta_linreg)
|
||||
plt<span style="color: #666666">.</span>plot(xnew, ypredict, <span style="color: #BA2121">"r-"</span>)
|
||||
plt<span style="color: #666666">.</span>plot(xnew, ypredict2, <span style="color: #BA2121">"b-"</span>)
|
||||
plt<span style="color: #666666">.</span>plot(x, y ,<span style="color: #BA2121">'ro'</span>)
|
||||
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">0</span>,<span style="color: #666666">2.0</span>,<span style="color: #666666">0</span>, <span style="color: #666666">15.0</span>])
|
||||
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r'$x$'</span>)
|
||||
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r'$y$'</span>)
|
||||
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">r'Gradient descent example'</span>)
|
||||
plt<span style="color: #666666">.</span>show()
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -174,27 +174,11 @@ MathJax.Hub.Config({
|
||||
<a name="part0007"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec6" class="anchor">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
<h2 id="___sec6" class="anchor">Convex function </h2>
|
||||
|
||||
<p>
|
||||
<b>Convex function</b>: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below.
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Importing various packages</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">random</span> <span style="color: #008000; font-weight: bold">import</span> random, seed
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.linear_model</span> <span style="color: #008000; font-weight: bold">import</span> SGDRegressor
|
||||
|
||||
x <span style="color: #666666">=</span> <span style="color: #666666">2*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">4+3*</span>x<span style="color: #666666">+</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
|
||||
xb <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)), x]
|
||||
beta_linreg <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb))<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>T)<span style="color: #666666">.</span>dot(y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_linreg)
|
||||
sgdreg <span style="color: #666666">=</span> SGDRegressor(n_iter <span style="color: #666666">=</span> <span style="color: #666666">50</span>, penalty<span style="color: #666666">=</span><span style="color: #008000">None</span>, eta0<span style="color: #666666">=0.1</span>)
|
||||
sgdreg<span style="color: #666666">.</span>fit(x,y<span style="color: #666666">.</span>ravel())
|
||||
<span style="color: #008000; font-weight: bold">print</span>(sgdreg<span style="color: #666666">.</span>intercept_, sgdreg<span style="color: #666666">.</span>coef_)
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -172,24 +172,48 @@ MathJax.Hub.Config({
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0008"></a>
|
||||
<!-- !split -->
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec7" class="anchor">Convex functions </h2>
|
||||
<h2 id="___sec7" class="anchor">Conditions on convex functions </h2>
|
||||
|
||||
<p>
|
||||
Ideally we want our cost/loss function to be convex(concave).
|
||||
In the following we state first and second-order conditions which
|
||||
ensures convexity of a function \( f \). We write \( D_f \) to denote the
|
||||
domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more
|
||||
details and proofs we refer to: <a href="http://stanford.edu/boyd/cvxbook/, 2004" target="_self">S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press</a>.
|
||||
|
||||
<p>
|
||||
First we give the definition of a convex set: A set \( C \) in
|
||||
\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and
|
||||
all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to
|
||||
C. Geometrically this means that every point on the line segment
|
||||
connecting \( x \) and \( y \) is in \( C \) as discussed below.
|
||||
<div class="panel panel-default">
|
||||
<div class="panel-body">
|
||||
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
|
||||
Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for
|
||||
all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \)
|
||||
is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
|
||||
for all \( x,y \in D_f \). This condition means that for a convex function
|
||||
the first order Taylor expansion (right hand side above) at any point
|
||||
a global under estimator of the function. To convince yourself you can
|
||||
make a drawing of \( f(x) = x^2+1 \) and draw the tangent line to \( f(x) \) and
|
||||
note that it is always below the graph.
|
||||
</div>
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
The convex subsets of \( \mathbb{R} \) are the intervals of
|
||||
\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the
|
||||
regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
<div class="panel panel-default">
|
||||
<div class="panel-body">
|
||||
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
|
||||
Assume that \( f \) is twice
|
||||
differentiable, i.e the Hessian matrix exists at each point in
|
||||
\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its
|
||||
Hessian is positive semi-definite for all \( x\in D_f \). For a
|
||||
single-variable function this reduces to \( f''(x) \geq 0 \). Geometrically this means that \( f \) has nonnegative curvature
|
||||
everywhere.
|
||||
</div>
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -174,10 +174,33 @@ MathJax.Hub.Config({
|
||||
<a name="part0009"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec8" class="anchor">Convex function </h2>
|
||||
<h2 id="___sec8" class="anchor">More on convex functions </h2>
|
||||
|
||||
<p>
|
||||
<b>Convex function</b>: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below.
|
||||
The next result is of great importance to us and the reason why we are
|
||||
going on about convex functions. In machine learning we frequently
|
||||
have to minimize a loss/cost function in order to find the best
|
||||
parameters for the model we are considering.
|
||||
|
||||
<p>
|
||||
Ideally we want the
|
||||
global minimum (for high-dimensional models it is hard to know
|
||||
if we have local or global minimum). However, if the cost/loss function
|
||||
is convex the following result provides invaluable information:
|
||||
|
||||
<p>
|
||||
<div class="panel panel-default">
|
||||
<div class="panel-body">
|
||||
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
|
||||
Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \)
|
||||
is minimal, where \( f \) is convex and differentiable. Then, any point
|
||||
\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum.
|
||||
</div>
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -174,47 +174,29 @@ MathJax.Hub.Config({
|
||||
<a name="part0010"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec9" class="anchor">Conditions on convex functions </h2>
|
||||
<h2 id="___sec9" class="anchor">Some simple problems </h2>
|
||||
|
||||
<p>
|
||||
In the following we state first and second-order conditions which
|
||||
ensures convexity of a function \( f \). We write \( D_f \) to denote the
|
||||
domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more
|
||||
details and proofs we refer to: <a href="http://stanford.edu/boyd/cvxbook/, 2004" target="_self">S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press</a>.
|
||||
<ol>
|
||||
<li> Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.</li>
|
||||
<li> Using the second order condition show that the following functions are convex on the specified domain.</li>
|
||||
|
||||
<p>
|
||||
<div class="panel panel-default">
|
||||
<div class="panel-body">
|
||||
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
|
||||
Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for
|
||||
all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \)
|
||||
is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
|
||||
for all \( x,y \in D_f \). This condition means that for a convex function
|
||||
the first order Taylor expansion (right hand side above) at any point
|
||||
a global under estimator of the function. To convince yourself you can
|
||||
make a drawing of f(x) = x^2+1 and draw the tangent line to \( f(x) \) and
|
||||
note that it is always below the graph.
|
||||
</div>
|
||||
</div>
|
||||
<ul>
|
||||
<li> \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).</li>
|
||||
<li> \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).</li>
|
||||
</ul>
|
||||
|
||||
<li> Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.</li>
|
||||
<li> A norm is any function that satisfy the following properties</li>
|
||||
|
||||
<p>
|
||||
<div class="panel panel-default">
|
||||
<div class="panel-body">
|
||||
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
|
||||
Assume that \( f \) is twice
|
||||
differentiable, i.e the Hessian matrix exists at each point in
|
||||
\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its
|
||||
Hessian is positive semi-definite for all \( x\in D_f \). For a
|
||||
single-variable function this reduces to \( f''(x) \geq
|
||||
0 \). Geometrically this means that \( f \) has nonnegative curvature
|
||||
everywhere.
|
||||
</div>
|
||||
</div>
|
||||
<ul>
|
||||
<li> \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).</li>
|
||||
<li> \( f(x+y) \leq f(x) + f(y) \)</li>
|
||||
<li> \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)</li>
|
||||
</ul>
|
||||
|
||||
</ol>
|
||||
|
||||
<p>
|
||||
This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
|
||||
Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
|
||||
|
||||
<p>
|
||||
<p>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -172,35 +172,37 @@ MathJax.Hub.Config({
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0011"></a>
|
||||
<!-- !split -->
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec10" class="anchor">More on convex functions </h2>
|
||||
<h2 id="___sec10" class="anchor">Revisiting our first homework </h2>
|
||||
|
||||
<p>
|
||||
The next result is of great importance to us and the reason why we are
|
||||
going on about convex functions. In machine learning we frequently
|
||||
have to minimize a loss/cost function in order to find the best
|
||||
parameters for the model we are considering.
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
<p>
|
||||
Ideally we want the
|
||||
global minimum (for high-dimensional models it is hard to know
|
||||
if we have local or global minimum). However, if the cost/loss function
|
||||
is convex the following result provides invaluable information:
|
||||
<ol>
|
||||
<li> An analytical solution (recall homework set 1).</li>
|
||||
<li> The gradient can be computed analytically.</li>
|
||||
<li> The cost function is convex which guarantees that gradient descent converges for small enough learning rates</li>
|
||||
</ol>
|
||||
|
||||
<p>
|
||||
<div class="panel panel-default">
|
||||
<div class="panel-body">
|
||||
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
|
||||
Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \)
|
||||
is minimal, where \( f \) is convex and differentiable. Then, any point
|
||||
\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum.
|
||||
</div>
|
||||
</div>
|
||||
We revisit the example from homework set 1 where we had
|
||||
$$
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
$$
|
||||
|
||||
with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
|
||||
The linear regression model is given by
|
||||
$$
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
$$
|
||||
|
||||
<p>
|
||||
This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
|
||||
such that
|
||||
$$
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<p>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -172,31 +172,29 @@ MathJax.Hub.Config({
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0012"></a>
|
||||
<!-- !split -->
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec11" class="anchor">Some simple problems </h2>
|
||||
<h2 id="___sec11" class="anchor">Gradient descent example </h2>
|
||||
|
||||
<ol>
|
||||
<li> Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x) + (1-\lambda)f(y) - f(\lambda x + (1-\lambda) y ) \geq 0. $</li>
|
||||
<li> Using the second order condition show that the following functions are convex on the specified domain.</li>
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
|
||||
<ul>
|
||||
<li> \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).</li>
|
||||
<li> \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).</li>
|
||||
</ul>
|
||||
<p>
|
||||
It is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by
|
||||
$$
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
$$
|
||||
|
||||
<li> Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.</li>
|
||||
<li> A norm is any function that satisfy the following properties</li>
|
||||
The loss function is given by
|
||||
$$
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
$$
|
||||
|
||||
<ul>
|
||||
<li> \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).</li>
|
||||
<li> \( f(x+y) \leq f(x) + f(y) \)</li>
|
||||
<li> \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)</li>
|
||||
</ul>
|
||||
|
||||
</ol>
|
||||
|
||||
Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
|
||||
and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -172,37 +172,19 @@ MathJax.Hub.Config({
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0013"></a>
|
||||
<!-- !split -->
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec12" class="anchor">Revisiting our first homework </h2>
|
||||
<h2 id="___sec12" class="anchor">The derivative of the cost/loss function </h2>
|
||||
|
||||
<p>
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
<ol>
|
||||
<li> An analytical solution (recall homework set 1).</li>
|
||||
<li> The gradient can be computed analytically.</li>
|
||||
<li> The cost function is convex which guarantees that gradient descent converges for small enough learning rates</li>
|
||||
</ol>
|
||||
|
||||
We revisit the example from homework set 1 where we had
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
$$
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
$$
|
||||
|
||||
with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
|
||||
The linear regression model is given by
|
||||
$$
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
$$
|
||||
|
||||
such that
|
||||
$$
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
$$
|
||||
where \( X \) is the design matrix defined above.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -172,32 +172,18 @@ MathJax.Hub.Config({
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0014"></a>
|
||||
<!-- !split -->
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec13" class="anchor">Gradient descent example </h2>
|
||||
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
|
||||
<p>
|
||||
t is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by
|
||||
<h2 id="___sec13" class="anchor">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
$$
|
||||
\begin{equation}
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
\tag{1}
|
||||
\end{equation}
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
$$
|
||||
|
||||
The loss function is given by
|
||||
$$
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
$$
|
||||
|
||||
and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -174,18 +174,44 @@ MathJax.Hub.Config({
|
||||
<a name="part0015"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec14" class="anchor">The derivative of the cost/loss function </h2>
|
||||
<h2 id="___sec14" class="anchor">Simple program </h2>
|
||||
|
||||
<p>
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
$$
|
||||
\nabla_\beta C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
$$
|
||||
|
||||
where \( X \) is the design matrix defined above.
|
||||
<p>
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
|
||||
when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \).
|
||||
|
||||
<p>
|
||||
And finally we can compare our solution for \( \beta \) with the analytic result given by
|
||||
\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
<span style="color: #BA2121; font-style: italic">The following setup is just a suggestion, feel free to write it the way you like.</span>
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Setup problem described in the exercise</span>
|
||||
N <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#Nr of datapoints</span>
|
||||
M <span style="color: #666666">=</span> <span style="color: #666666">2</span> <span style="color: #408080; font-style: italic">#Nr of features</span>
|
||||
x <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(N) <span style="color: #408080; font-style: italic">#Uniformly generated x-values in [0,1]</span>
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">5*</span>x<span style="color: #666666">**2</span> <span style="color: #666666">+</span> <span style="color: #666666">0.1*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(N)
|
||||
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones(N),x] <span style="color: #408080; font-style: italic">#Construct design matrix</span>
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Compute beta according to normal equations to compare with GD solution</span>
|
||||
Xt_X_inv <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,X))
|
||||
Xt_y <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>transpose(),y)
|
||||
beta_NE <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(Xt_X_inv,Xt_y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_NE)
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -174,17 +174,52 @@ MathJax.Hub.Config({
|
||||
<a name="part0016"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec15" class="anchor">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
$$
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
$$
|
||||
<h2 id="___sec15" class="anchor">Gradient Descent Example </h2>
|
||||
|
||||
This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
|
||||
<p>
|
||||
Another simple example is here
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Importing various packages</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">random</span> <span style="color: #008000; font-weight: bold">import</span> random, seed
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">mpl_toolkits.mplot3d</span> <span style="color: #008000; font-weight: bold">import</span> Axes3D
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">matplotlib</span> <span style="color: #008000; font-weight: bold">import</span> cm
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">matplotlib.ticker</span> <span style="color: #008000; font-weight: bold">import</span> LinearLocator, FormatStrFormatter
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">sys</span>
|
||||
|
||||
x <span style="color: #666666">=</span> <span style="color: #666666">2*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">4+3*</span>x<span style="color: #666666">+</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
|
||||
xb <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)), x]
|
||||
beta_linreg <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb))<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>T)<span style="color: #666666">.</span>dot(y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_linreg)
|
||||
beta <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">2</span>,<span style="color: #666666">1</span>)
|
||||
|
||||
eta <span style="color: #666666">=</span> <span style="color: #666666">0.1</span>
|
||||
Niterations <span style="color: #666666">=</span> <span style="color: #666666">1000</span>
|
||||
m <span style="color: #666666">=</span> <span style="color: #666666">100</span>
|
||||
|
||||
<span style="color: #008000; font-weight: bold">for</span> <span style="color: #008000">iter</span> <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(Niterations):
|
||||
gradients <span style="color: #666666">=</span> <span style="color: #666666">2.0/</span>m<span style="color: #666666">*</span>xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>dot(beta)<span style="color: #666666">-</span>y)
|
||||
beta <span style="color: #666666">-=</span> eta<span style="color: #666666">*</span>gradients
|
||||
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta)
|
||||
xnew <span style="color: #666666">=</span> np<span style="color: #666666">.</span>array([[<span style="color: #666666">0</span>],[<span style="color: #666666">2</span>]])
|
||||
xbnew <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">2</span>,<span style="color: #666666">1</span>)), xnew]
|
||||
ypredict <span style="color: #666666">=</span> xbnew<span style="color: #666666">.</span>dot(beta)
|
||||
ypredict2 <span style="color: #666666">=</span> xbnew<span style="color: #666666">.</span>dot(beta_linreg)
|
||||
plt<span style="color: #666666">.</span>plot(xnew, ypredict, <span style="color: #BA2121">"r-"</span>)
|
||||
plt<span style="color: #666666">.</span>plot(xnew, ypredict2, <span style="color: #BA2121">"b-"</span>)
|
||||
plt<span style="color: #666666">.</span>plot(x, y ,<span style="color: #BA2121">'ro'</span>)
|
||||
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">0</span>,<span style="color: #666666">2.0</span>,<span style="color: #666666">0</span>, <span style="color: #666666">15.0</span>])
|
||||
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r'$x$'</span>)
|
||||
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r'$y$'</span>)
|
||||
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">r'Gradient descent example'</span>)
|
||||
plt<span style="color: #666666">.</span>show()
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -174,43 +174,26 @@ MathJax.Hub.Config({
|
||||
<a name="part0017"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec16" class="anchor">Simple program </h2>
|
||||
<h2 id="___sec16" class="anchor">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
|
||||
<p>
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
$$
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
$$
|
||||
|
||||
<p>
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
|
||||
when \( ||\nabla_\beta C(\beta_k) || < \epsilon = 10^{-8} \).
|
||||
|
||||
<p>
|
||||
And finally we can compare our solution for \( \beta \) with the analytic result given by
|
||||
\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Importing various packages</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">random</span> <span style="color: #008000; font-weight: bold">import</span> random, seed
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.linear_model</span> <span style="color: #008000; font-weight: bold">import</span> SGDRegressor
|
||||
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
<span style="color: #BA2121; font-style: italic">The following setup is just a suggestion, feel free to write it the way you like.</span>
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
x <span style="color: #666666">=</span> <span style="color: #666666">2*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">4+3*</span>x<span style="color: #666666">+</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Setup problem described in the exercise</span>
|
||||
N <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#Nr of datapoints</span>
|
||||
M <span style="color: #666666">=</span> <span style="color: #666666">2</span> <span style="color: #408080; font-style: italic">#Nr of features</span>
|
||||
x <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(N) <span style="color: #408080; font-style: italic">#Uniformly generated x-values in [0,1]</span>
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">5*</span>x<span style="color: #666666">**2</span> <span style="color: #666666">+</span> <span style="color: #666666">0.1*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(N)
|
||||
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones(N),x] <span style="color: #408080; font-style: italic">#Construct design matrix</span>
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Compute beta according to normal equations to compare with GD solution</span>
|
||||
Xt_X_inv <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,X))
|
||||
Xt_y <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>transpose(),y)
|
||||
beta_NE <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(Xt_X_inv,Xt_y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_NE)
|
||||
xb <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)), x]
|
||||
beta_linreg <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb))<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>T)<span style="color: #666666">.</span>dot(y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_linreg)
|
||||
sgdreg <span style="color: #666666">=</span> SGDRegressor(n_iter <span style="color: #666666">=</span> <span style="color: #666666">50</span>, penalty<span style="color: #666666">=</span><span style="color: #008000">None</span>, eta0<span style="color: #666666">=0.1</span>)
|
||||
sgdreg<span style="color: #666666">.</span>fit(x,y<span style="color: #666666">.</span>ravel())
|
||||
<span style="color: #008000; font-weight: bold">print</span>(sgdreg<span style="color: #666666">.</span>intercept_, sgdreg<span style="color: #666666">.</span>coef_)
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -183,7 +183,7 @@ the shortcomings of the Gradient descent method discussed above.
|
||||
<p>
|
||||
The underlying idea of SGD comes from the observation that the cost
|
||||
function, which we want to minimize, can almost always be written as a
|
||||
sum over \( n \) datapoints \( \{\mathbf{x}_i\}_{i=1}^n \),
|
||||
sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \),
|
||||
$$
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}).
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -187,7 +187,7 @@ $$
|
||||
<p>
|
||||
Stochasticity/randomness is introduced by only taking the
|
||||
gradient on a subset of the data called minibatches. If there are \( n \)
|
||||
datapoints and the size of each minibatch is \( M \), there will be \( n/M \)
|
||||
data points and the size of each minibatch is \( M \), there will be \( n/M \)
|
||||
minibatches. We denote these minibatches by \( B_k \) where
|
||||
\( k=1,\cdots,n/M \).
|
||||
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -175,21 +175,21 @@ MathJax.Hub.Config({
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec20" class="anchor">SGD example </h2>
|
||||
As an example, suppose we have \( 10 \) datapoints \( ( \mathbf{x}_1,
|
||||
\cdots, \mathbf{x}_{10} ) \) and we choose to have \( M=5 \) minibathces,
|
||||
then each minibatch contains two datapoints. In particular we have
|
||||
As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \)
|
||||
and we choose to have \( M=5 \) minibathces,
|
||||
then each minibatch contains two data points. In particular we have
|
||||
\( B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 =
|
||||
(\mathbf{x}_9,\mathbf{x}_{10}) \). Note that if you choose \( M=1 \) you
|
||||
have only a single batch with all datapoints and on the other extreme,
|
||||
have only a single batch with all data points and on the other extreme,
|
||||
you may choose \( M=n \) resulting in a minibatch for each datapoint, i.e
|
||||
\( B_k = \mathbf{x}_k \).
|
||||
|
||||
<p>
|
||||
The idea is now to approximate the gradient by replacing the sum over
|
||||
all datapoints with a sum over the datapoints in one the minibatches
|
||||
all data points with a sum over the data points in one the minibatches
|
||||
picked at random in each gradient descent step
|
||||
$$
|
||||
\nabla_\beta
|
||||
\nabla_{\beta}
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta
|
||||
c_i(\mathbf{x}_i, \mathbf{\beta}).
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -199,8 +199,8 @@ Taking the gradient only on a subset of the data has two important
|
||||
benefits. First, it introduces randomness which decreases the chance
|
||||
that our opmization scheme gets stuck in a local minima. Second, if
|
||||
the size of the minibatches are small relative to the number of
|
||||
datapoints (\( M < n \)), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the k-th minibatch and not
|
||||
datapoints (\( M < n \)), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the \( k-th \) minibatch and not
|
||||
all \( n \) datapoints.
|
||||
|
||||
<p>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -182,7 +182,7 @@ number of epochs in such a way that it becomes very small after a
|
||||
reasonable time such that we do not move at all.
|
||||
|
||||
<p>
|
||||
As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in <em>time</em> \( t \).
|
||||
As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in <em>time</em> \( t \).
|
||||
|
||||
<p>
|
||||
In this way we can fix the number of epochs, compute \( \beta \) and
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
|
||||
@@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -128,18 +128,18 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs003.html#___sec2" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs004.html#___sec3" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs005.html#___sec4" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs006.html#___sec5" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs007.html#___sec6" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs008.html#___sec7" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs009.html#___sec8" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs010.html#___sec9" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs011.html#___sec10" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs012.html#___sec11" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs013.html#___sec12" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs014.html#___sec13" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs015.html#___sec14" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -193,7 +193,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Sep 20, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Sep 21, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
|
||||
|
||||
@@ -148,7 +148,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p> <br>
|
||||
<center><h4>Sep 20, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Sep 21, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
|
||||
@@ -186,12 +186,14 @@ direction of the negative gradient \( -\nabla F(\mathbf{x}) \).
|
||||
It can be shown that if
|
||||
<p> <br>
|
||||
$$
|
||||
\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), \ \ \gamma_k > 0
|
||||
\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k),
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
with \( \gamma_k > 0 \).
|
||||
|
||||
<p>
|
||||
for \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq
|
||||
For \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq
|
||||
F(\mathbf{x}_k) \). This means that for a sufficiently small \( \gamma_k \)
|
||||
we are always moving towards smaller function values, i.e a minimum.
|
||||
</section>
|
||||
@@ -222,7 +224,7 @@ the learning rate within the context of Machine Learning.
|
||||
<h2 id="___sec3">The ideal </h2>
|
||||
|
||||
<p>
|
||||
Ideally the sequence \( \{ \mathbf{x}_k \}_{k=0} \) converges to a global
|
||||
Ideally the sequence \( \{\mathbf{x}_k \}_{k=0} \) converges to a global
|
||||
minimum of the function \( F \). In general we do not know if we are in a
|
||||
global or local minimum. In the special case when \( F \) is a convex
|
||||
function, all local minima are also global minima, so in this case
|
||||
@@ -248,7 +250,8 @@ Note that the gradient is a function of \( \mathbf{x} =
|
||||
<h2 id="___sec4">The sensitiveness of the gradient descent </h2>
|
||||
|
||||
<p>
|
||||
GD is sensitive to the choice of learning rate \( \gamma_k \). This is due
|
||||
The gradient descent method
|
||||
is sensitive to the choice of learning rate \( \gamma_k \). This is due
|
||||
to the fact that we are only guaranteed that \( F(\mathbf{x}_{k+1}) \leq
|
||||
F(\mathbf{x}_k) \) for sufficiently small \( \gamma_k \). The problem is to
|
||||
determine an optimal learning rate. If the learning rate is chosen too
|
||||
@@ -263,10 +266,284 @@ randomness. One such method is that of Stochastic Gradient Descent
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec5">Gradient Descent Example </h2>
|
||||
<h2 id="___sec5">Convex functions </h2>
|
||||
|
||||
<p>
|
||||
We revisit now our simple linear regression example with a linear polynomial.
|
||||
Ideally we want our cost/loss function to be convex(concave).
|
||||
|
||||
<p>
|
||||
First we give the definition of a convex set: A set \( C \) in
|
||||
\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and
|
||||
all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to
|
||||
C. Geometrically this means that every point on the line segment
|
||||
connecting \( x \) and \( y \) is in \( C \) as discussed below.
|
||||
|
||||
<p>
|
||||
The convex subsets of \( \mathbb{R} \) are the intervals of
|
||||
\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the
|
||||
regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec6">Convex function </h2>
|
||||
|
||||
<p>
|
||||
<b>Convex function</b>: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if <p> <br>
|
||||
$$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$
|
||||
<p> <br> for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec7">Conditions on convex functions </h2>
|
||||
|
||||
<p>
|
||||
In the following we state first and second-order conditions which
|
||||
ensures convexity of a function \( f \). We write \( D_f \) to denote the
|
||||
domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more
|
||||
details and proofs we refer to: <a href="http://stanford.edu/boyd/cvxbook/, 2004" target="_blank">S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press</a>.
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>First order condition.</b>
|
||||
<p>
|
||||
Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for
|
||||
all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \)
|
||||
is a convex set and <p> <br>
|
||||
$$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$
|
||||
<p> <br> holds
|
||||
for all \( x,y \in D_f \). This condition means that for a convex function
|
||||
the first order Taylor expansion (right hand side above) at any point
|
||||
a global under estimator of the function. To convince yourself you can
|
||||
make a drawing of \( f(x) = x^2+1 \) and draw the tangent line to \( f(x) \) and
|
||||
note that it is always below the graph.
|
||||
</div>
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>Second order condition.</b>
|
||||
<p>
|
||||
Assume that \( f \) is twice
|
||||
differentiable, i.e the Hessian matrix exists at each point in
|
||||
\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its
|
||||
Hessian is positive semi-definite for all \( x\in D_f \). For a
|
||||
single-variable function this reduces to \( f''(x) \geq 0 \). Geometrically this means that \( f \) has nonnegative curvature
|
||||
everywhere.
|
||||
</div>
|
||||
|
||||
<p>
|
||||
This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec8">More on convex functions </h2>
|
||||
|
||||
<p>
|
||||
The next result is of great importance to us and the reason why we are
|
||||
going on about convex functions. In machine learning we frequently
|
||||
have to minimize a loss/cost function in order to find the best
|
||||
parameters for the model we are considering.
|
||||
|
||||
<p>
|
||||
Ideally we want the
|
||||
global minimum (for high-dimensional models it is hard to know
|
||||
if we have local or global minimum). However, if the cost/loss function
|
||||
is convex the following result provides invaluable information:
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>Any minimum is global for convex functions.</b>
|
||||
<p>
|
||||
Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \)
|
||||
is minimal, where \( f \) is convex and differentiable. Then, any point
|
||||
\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum.
|
||||
</div>
|
||||
|
||||
<p>
|
||||
This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec9">Some simple problems </h2>
|
||||
|
||||
<ol>
|
||||
<p><li> Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.</li>
|
||||
<p><li> Using the second order condition show that the following functions are convex on the specified domain.</li>
|
||||
|
||||
<ul>
|
||||
<p><li> \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).</li>
|
||||
<p><li> \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).</li>
|
||||
</ul>
|
||||
<p><li> Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.</li>
|
||||
<p><li> A norm is any function that satisfy the following properties</li>
|
||||
|
||||
<ul>
|
||||
<p><li> \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).</li>
|
||||
<p><li> \( f(x+y) \leq f(x) + f(y) \)</li>
|
||||
<p><li> \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)</li>
|
||||
</ul>
|
||||
<p>
|
||||
</ol>
|
||||
<p>
|
||||
|
||||
Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec10">Revisiting our first homework </h2>
|
||||
|
||||
<p>
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
<ol>
|
||||
<p><li> An analytical solution (recall homework set 1).</li>
|
||||
<p><li> The gradient can be computed analytically.</li>
|
||||
<p><li> The cost function is convex which guarantees that gradient descent converges for small enough learning rates</li>
|
||||
</ol>
|
||||
<p>
|
||||
|
||||
We revisit the example from homework set 1 where we had
|
||||
<p> <br>
|
||||
$$
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
|
||||
The linear regression model is given by
|
||||
<p> <br>
|
||||
$$
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
such that
|
||||
<p> <br>
|
||||
$$
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
$$
|
||||
<p> <br>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec11">Gradient descent example </h2>
|
||||
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
|
||||
<p>
|
||||
It is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by
|
||||
<p> <br>
|
||||
$$
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
The loss function is given by
|
||||
<p> <br>
|
||||
$$
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec12">The derivative of the cost/loss function </h2>
|
||||
|
||||
<p>
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
<p> <br>
|
||||
$$
|
||||
\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
where \( X \) is the design matrix defined above.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec13">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
<p> <br>
|
||||
$$
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec14">Simple program </h2>
|
||||
|
||||
<p>
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
<p> <br>
|
||||
$$
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
|
||||
when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \).
|
||||
|
||||
<p>
|
||||
And finally we can compare our solution for \( \beta \) with the analytic result given by
|
||||
\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
|
||||
<div class="highlight" style="background: #eeeedd"><pre style="font-size: 80%; line-height: 125%"><span></span><span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">numpy</span> <span style="color: #8B008B; font-weight: bold">as</span> <span style="color: #008b45; text-decoration: underline">np</span>
|
||||
|
||||
<span style="color: #CD5555">"""</span>
|
||||
<span style="color: #CD5555">The following setup is just a suggestion, feel free to write it the way you like.</span>
|
||||
<span style="color: #CD5555">"""</span>
|
||||
|
||||
<span style="color: #228B22">#Setup problem described in the exercise</span>
|
||||
N = <span style="color: #B452CD">100</span> <span style="color: #228B22">#Nr of datapoints</span>
|
||||
M = <span style="color: #B452CD">2</span> <span style="color: #228B22">#Nr of features</span>
|
||||
x = np.random.rand(N) <span style="color: #228B22">#Uniformly generated x-values in [0,1]</span>
|
||||
y = <span style="color: #B452CD">5</span>*x**<span style="color: #B452CD">2</span> + <span style="color: #B452CD">0.1</span>*np.random.randn(N)
|
||||
X = np.c_[np.ones(N),x] <span style="color: #228B22">#Construct design matrix</span>
|
||||
|
||||
<span style="color: #228B22">#Compute beta according to normal equations to compare with GD solution</span>
|
||||
Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
|
||||
Xt_y = np.dot(X.transpose(),y)
|
||||
beta_NE = np.dot(Xt_X_inv,Xt_y)
|
||||
<span style="color: #8B008B; font-weight: bold">print</span>(beta_NE)
|
||||
</pre></div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec15">Gradient Descent Example </h2>
|
||||
|
||||
<p>
|
||||
Another simple example is here
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
|
||||
@@ -313,7 +590,7 @@ plt.show()
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec6">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
<h2 id="___sec16">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
|
||||
<p>
|
||||
|
||||
@@ -337,284 +614,6 @@ sgdreg.fit(x,y.ravel())
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec7">Convex functions </h2>
|
||||
|
||||
<p>
|
||||
Ideally we want our cost/loss function to be convex(concave).
|
||||
|
||||
<p>
|
||||
First we give the definition of a convex set: A set \( C \) in
|
||||
\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and
|
||||
all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to
|
||||
C. Geometrically this means that every point on the line segment
|
||||
connecting \( x \) and \( y \) is in \( C \) as discussed below.
|
||||
|
||||
<p>
|
||||
The convex subsets of \( \mathbb{R} \) are the intervals of
|
||||
\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the
|
||||
regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec8">Convex function </h2>
|
||||
|
||||
<p>
|
||||
<b>Convex function</b>: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if <p> <br>
|
||||
$$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$
|
||||
<p> <br> for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec9">Conditions on convex functions </h2>
|
||||
|
||||
<p>
|
||||
In the following we state first and second-order conditions which
|
||||
ensures convexity of a function \( f \). We write \( D_f \) to denote the
|
||||
domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more
|
||||
details and proofs we refer to: <a href="http://stanford.edu/boyd/cvxbook/, 2004" target="_blank">S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press</a>.
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>First order condition.</b>
|
||||
<p>
|
||||
Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for
|
||||
all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \)
|
||||
is a convex set and <p> <br>
|
||||
$$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$
|
||||
<p> <br> holds
|
||||
for all \( x,y \in D_f \). This condition means that for a convex function
|
||||
the first order Taylor expansion (right hand side above) at any point
|
||||
a global under estimator of the function. To convince yourself you can
|
||||
make a drawing of f(x) = x^2+1 and draw the tangent line to \( f(x) \) and
|
||||
note that it is always below the graph.
|
||||
</div>
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>Second order condition.</b>
|
||||
<p>
|
||||
Assume that \( f \) is twice
|
||||
differentiable, i.e the Hessian matrix exists at each point in
|
||||
\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its
|
||||
Hessian is positive semi-definite for all \( x\in D_f \). For a
|
||||
single-variable function this reduces to \( f''(x) \geq
|
||||
0 \). Geometrically this means that \( f \) has nonnegative curvature
|
||||
everywhere.
|
||||
</div>
|
||||
|
||||
<p>
|
||||
This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec10">More on convex functions </h2>
|
||||
|
||||
<p>
|
||||
The next result is of great importance to us and the reason why we are
|
||||
going on about convex functions. In machine learning we frequently
|
||||
have to minimize a loss/cost function in order to find the best
|
||||
parameters for the model we are considering.
|
||||
|
||||
<p>
|
||||
Ideally we want the
|
||||
global minimum (for high-dimensional models it is hard to know
|
||||
if we have local or global minimum). However, if the cost/loss function
|
||||
is convex the following result provides invaluable information:
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>Any minimum is global for convex functions.</b>
|
||||
<p>
|
||||
Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \)
|
||||
is minimal, where \( f \) is convex and differentiable. Then, any point
|
||||
\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum.
|
||||
</div>
|
||||
|
||||
<p>
|
||||
This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec11">Some simple problems </h2>
|
||||
|
||||
<ol>
|
||||
<p><li> Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x) + (1-\lambda)f(y) - f(\lambda x + (1-\lambda) y ) \geq 0. $</li>
|
||||
<p><li> Using the second order condition show that the following functions are convex on the specified domain.</li>
|
||||
|
||||
<ul>
|
||||
<p><li> \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).</li>
|
||||
<p><li> \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).</li>
|
||||
</ul>
|
||||
<p><li> Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.</li>
|
||||
<p><li> A norm is any function that satisfy the following properties</li>
|
||||
|
||||
<ul>
|
||||
<p><li> \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).</li>
|
||||
<p><li> \( f(x+y) \leq f(x) + f(y) \)</li>
|
||||
<p><li> \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)</li>
|
||||
</ul>
|
||||
<p>
|
||||
</ol>
|
||||
<p>
|
||||
|
||||
Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec12">Revisiting our first homework </h2>
|
||||
|
||||
<p>
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
<ol>
|
||||
<p><li> An analytical solution (recall homework set 1).</li>
|
||||
<p><li> The gradient can be computed analytically.</li>
|
||||
<p><li> The cost function is convex which guarantees that gradient descent converges for small enough learning rates</li>
|
||||
</ol>
|
||||
<p>
|
||||
|
||||
We revisit the example from homework set 1 where we had
|
||||
<p> <br>
|
||||
$$
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
|
||||
The linear regression model is given by
|
||||
<p> <br>
|
||||
$$
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
such that
|
||||
<p> <br>
|
||||
$$
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
$$
|
||||
<p> <br>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec13">Gradient descent example </h2>
|
||||
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
|
||||
<p>
|
||||
t is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation}
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
\tag{1}
|
||||
\end{equation}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
The loss function is given by
|
||||
<p> <br>
|
||||
$$
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec14">The derivative of the cost/loss function </h2>
|
||||
|
||||
<p>
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
<p> <br>
|
||||
$$
|
||||
\nabla_\beta C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
where \( X \) is the design matrix defined above.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec15">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
<p> <br>
|
||||
$$
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec16">Simple program </h2>
|
||||
|
||||
<p>
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
<p> <br>
|
||||
$$
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
|
||||
when \( ||\nabla_\beta C(\beta_k) || < \epsilon = 10^{-8} \).
|
||||
|
||||
<p>
|
||||
And finally we can compare our solution for \( \beta \) with the analytic result given by
|
||||
\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
|
||||
<div class="highlight" style="background: #eeeedd"><pre style="font-size: 80%; line-height: 125%"><span></span><span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">numpy</span> <span style="color: #8B008B; font-weight: bold">as</span> <span style="color: #008b45; text-decoration: underline">np</span>
|
||||
|
||||
<span style="color: #CD5555">"""</span>
|
||||
<span style="color: #CD5555">The following setup is just a suggestion, feel free to write it the way you like.</span>
|
||||
<span style="color: #CD5555">"""</span>
|
||||
|
||||
<span style="color: #228B22">#Setup problem described in the exercise</span>
|
||||
N = <span style="color: #B452CD">100</span> <span style="color: #228B22">#Nr of datapoints</span>
|
||||
M = <span style="color: #B452CD">2</span> <span style="color: #228B22">#Nr of features</span>
|
||||
x = np.random.rand(N) <span style="color: #228B22">#Uniformly generated x-values in [0,1]</span>
|
||||
y = <span style="color: #B452CD">5</span>*x**<span style="color: #B452CD">2</span> + <span style="color: #B452CD">0.1</span>*np.random.randn(N)
|
||||
X = np.c_[np.ones(N),x] <span style="color: #228B22">#Construct design matrix</span>
|
||||
|
||||
<span style="color: #228B22">#Compute beta according to normal equations to compare with GD solution</span>
|
||||
Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
|
||||
Xt_y = np.dot(X.transpose(),y)
|
||||
beta_NE = np.dot(Xt_X_inv,Xt_y)
|
||||
<span style="color: #8B008B; font-weight: bold">print</span>(beta_NE)
|
||||
</pre></div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec17">Gradient descent and Ridge </h2>
|
||||
|
||||
@@ -689,7 +688,7 @@ the shortcomings of the Gradient descent method discussed above.
|
||||
<p>
|
||||
The underlying idea of SGD comes from the observation that the cost
|
||||
function, which we want to minimize, can almost always be written as a
|
||||
sum over \( n \) datapoints \( \{\mathbf{x}_i\}_{i=1}^n \),
|
||||
sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \),
|
||||
<p> <br>
|
||||
$$
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i,
|
||||
@@ -715,7 +714,7 @@ $$
|
||||
<p>
|
||||
Stochasticity/randomness is introduced by only taking the
|
||||
gradient on a subset of the data called minibatches. If there are \( n \)
|
||||
datapoints and the size of each minibatch is \( M \), there will be \( n/M \)
|
||||
data points and the size of each minibatch is \( M \), there will be \( n/M \)
|
||||
minibatches. We denote these minibatches by \( B_k \) where
|
||||
\( k=1,\cdots,n/M \).
|
||||
</section>
|
||||
@@ -723,22 +722,22 @@ minibatches. We denote these minibatches by \( B_k \) where
|
||||
|
||||
<section>
|
||||
<h2 id="___sec20">SGD example </h2>
|
||||
As an example, suppose we have \( 10 \) datapoints \( ( \mathbf{x}_1,
|
||||
\cdots, \mathbf{x}_{10} ) \) and we choose to have \( M=5 \) minibathces,
|
||||
then each minibatch contains two datapoints. In particular we have
|
||||
As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \)
|
||||
and we choose to have \( M=5 \) minibathces,
|
||||
then each minibatch contains two data points. In particular we have
|
||||
\( B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 =
|
||||
(\mathbf{x}_9,\mathbf{x}_{10}) \). Note that if you choose \( M=1 \) you
|
||||
have only a single batch with all datapoints and on the other extreme,
|
||||
have only a single batch with all data points and on the other extreme,
|
||||
you may choose \( M=n \) resulting in a minibatch for each datapoint, i.e
|
||||
\( B_k = \mathbf{x}_k \).
|
||||
|
||||
<p>
|
||||
The idea is now to approximate the gradient by replacing the sum over
|
||||
all datapoints with a sum over the datapoints in one the minibatches
|
||||
all data points with a sum over the data points in one the minibatches
|
||||
picked at random in each gradient descent step
|
||||
<p> <br>
|
||||
$$
|
||||
\nabla_\beta
|
||||
\nabla_{\beta}
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta
|
||||
c_i(\mathbf{x}_i, \mathbf{\beta}).
|
||||
@@ -794,8 +793,8 @@ Taking the gradient only on a subset of the data has two important
|
||||
benefits. First, it introduces randomness which decreases the chance
|
||||
that our opmization scheme gets stuck in a local minima. Second, if
|
||||
the size of the minibatches are small relative to the number of
|
||||
datapoints (\( M < n \)), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the k-th minibatch and not
|
||||
datapoints (\( M < n \)), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the \( k-th \) minibatch and not
|
||||
all \( n \) datapoints.
|
||||
</section>
|
||||
|
||||
@@ -826,7 +825,7 @@ number of epochs in such a way that it becomes very small after a
|
||||
reasonable time such that we do not move at all.
|
||||
|
||||
<p>
|
||||
As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function <p> <br>
|
||||
As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function <p> <br>
|
||||
$$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$
|
||||
<p> <br> goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in <em>time</em> \( t \).
|
||||
|
||||
|
||||
@@ -69,21 +69,21 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -147,7 +147,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Sep 20, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Sep 21, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -178,11 +178,13 @@ direction of the negative gradient \( -\nabla F(\mathbf{x}) \).
|
||||
<p>
|
||||
It can be shown that if
|
||||
$$
|
||||
\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), \ \ \gamma_k > 0
|
||||
\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k),
|
||||
$$
|
||||
|
||||
with \( \gamma_k > 0 \).
|
||||
|
||||
<p>
|
||||
for \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq
|
||||
For \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq
|
||||
F(\mathbf{x}_k) \). This means that for a sufficiently small \( \gamma_k \)
|
||||
we are always moving towards smaller function values, i.e a minimum.
|
||||
|
||||
@@ -211,7 +213,7 @@ the learning rate within the context of Machine Learning.
|
||||
<h2 id="___sec3">The ideal </h2>
|
||||
|
||||
<p>
|
||||
Ideally the sequence \( \{ \mathbf{x}_k \}_{k=0} \) converges to a global
|
||||
Ideally the sequence \( \{\mathbf{x}_k \}_{k=0} \) converges to a global
|
||||
minimum of the function \( F \). In general we do not know if we are in a
|
||||
global or local minimum. In the special case when \( F \) is a convex
|
||||
function, all local minima are also global minima, so in this case
|
||||
@@ -237,7 +239,8 @@ Note that the gradient is a function of \( \mathbf{x} =
|
||||
<h2 id="___sec4">The sensitiveness of the gradient descent </h2>
|
||||
|
||||
<p>
|
||||
GD is sensitive to the choice of learning rate \( \gamma_k \). This is due
|
||||
The gradient descent method
|
||||
is sensitive to the choice of learning rate \( \gamma_k \). This is due
|
||||
to the fact that we are only guaranteed that \( F(\mathbf{x}_{k+1}) \leq
|
||||
F(\mathbf{x}_k) \) for sufficiently small \( \gamma_k \). The problem is to
|
||||
determine an optimal learning rate. If the learning rate is chosen too
|
||||
@@ -250,12 +253,267 @@ randomness. One such method is that of Stochastic Gradient Descent
|
||||
(SGD), see below.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec5">Gradient Descent Example </h2>
|
||||
<h2 id="___sec5">Convex functions </h2>
|
||||
|
||||
<p>
|
||||
We revisit now our simple linear regression example with a linear polynomial.
|
||||
Ideally we want our cost/loss function to be convex(concave).
|
||||
|
||||
<p>
|
||||
First we give the definition of a convex set: A set \( C \) in
|
||||
\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and
|
||||
all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to
|
||||
C. Geometrically this means that every point on the line segment
|
||||
connecting \( x \) and \( y \) is in \( C \) as discussed below.
|
||||
|
||||
<p>
|
||||
The convex subsets of \( \mathbb{R} \) are the intervals of
|
||||
\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the
|
||||
regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec6">Convex function </h2>
|
||||
|
||||
<p>
|
||||
<b>Convex function</b>: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec7">Conditions on convex functions </h2>
|
||||
|
||||
<p>
|
||||
In the following we state first and second-order conditions which
|
||||
ensures convexity of a function \( f \). We write \( D_f \) to denote the
|
||||
domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more
|
||||
details and proofs we refer to: <a href="http://stanford.edu/boyd/cvxbook/, 2004" target="_blank">S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press</a>.
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>First order condition.</b>
|
||||
<p>
|
||||
Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for
|
||||
all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \)
|
||||
is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
|
||||
for all \( x,y \in D_f \). This condition means that for a convex function
|
||||
the first order Taylor expansion (right hand side above) at any point
|
||||
a global under estimator of the function. To convince yourself you can
|
||||
make a drawing of \( f(x) = x^2+1 \) and draw the tangent line to \( f(x) \) and
|
||||
note that it is always below the graph.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>Second order condition.</b>
|
||||
<p>
|
||||
Assume that \( f \) is twice
|
||||
differentiable, i.e the Hessian matrix exists at each point in
|
||||
\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its
|
||||
Hessian is positive semi-definite for all \( x\in D_f \). For a
|
||||
single-variable function this reduces to \( f''(x) \geq 0 \). Geometrically this means that \( f \) has nonnegative curvature
|
||||
everywhere.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec8">More on convex functions </h2>
|
||||
|
||||
<p>
|
||||
The next result is of great importance to us and the reason why we are
|
||||
going on about convex functions. In machine learning we frequently
|
||||
have to minimize a loss/cost function in order to find the best
|
||||
parameters for the model we are considering.
|
||||
|
||||
<p>
|
||||
Ideally we want the
|
||||
global minimum (for high-dimensional models it is hard to know
|
||||
if we have local or global minimum). However, if the cost/loss function
|
||||
is convex the following result provides invaluable information:
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>Any minimum is global for convex functions.</b>
|
||||
<p>
|
||||
Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \)
|
||||
is minimal, where \( f \) is convex and differentiable. Then, any point
|
||||
\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec9">Some simple problems </h2>
|
||||
|
||||
<ol>
|
||||
<li> Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.</li>
|
||||
<li> Using the second order condition show that the following functions are convex on the specified domain.</li>
|
||||
|
||||
<ul>
|
||||
<li> \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).</li>
|
||||
<li> \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).</li>
|
||||
</ul>
|
||||
|
||||
<li> Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.</li>
|
||||
<li> A norm is any function that satisfy the following properties</li>
|
||||
|
||||
<ul>
|
||||
<li> \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).</li>
|
||||
<li> \( f(x+y) \leq f(x) + f(y) \)</li>
|
||||
<li> \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)</li>
|
||||
</ul>
|
||||
|
||||
</ol>
|
||||
|
||||
Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec10">Revisiting our first homework </h2>
|
||||
|
||||
<p>
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
<ol>
|
||||
<li> An analytical solution (recall homework set 1).</li>
|
||||
<li> The gradient can be computed analytically.</li>
|
||||
<li> The cost function is convex which guarantees that gradient descent converges for small enough learning rates</li>
|
||||
</ol>
|
||||
|
||||
We revisit the example from homework set 1 where we had
|
||||
$$
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
$$
|
||||
|
||||
with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
|
||||
The linear regression model is given by
|
||||
$$
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
$$
|
||||
|
||||
such that
|
||||
$$
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec11">Gradient descent example </h2>
|
||||
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
|
||||
<p>
|
||||
It is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by
|
||||
$$
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
$$
|
||||
|
||||
The loss function is given by
|
||||
$$
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
$$
|
||||
|
||||
and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec12">The derivative of the cost/loss function </h2>
|
||||
|
||||
<p>
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
$$
|
||||
\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
$$
|
||||
|
||||
where \( X \) is the design matrix defined above.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec13">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
$$
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
$$
|
||||
|
||||
This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec14">Simple program </h2>
|
||||
|
||||
<p>
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
$$
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
$$
|
||||
|
||||
<p>
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
|
||||
when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \).
|
||||
|
||||
<p>
|
||||
And finally we can compare our solution for \( \beta \) with the analytic result given by
|
||||
\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
|
||||
<div class="highlight" style="background: #eeeedd"><pre style="line-height: 125%"><span></span><span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">numpy</span> <span style="color: #8B008B; font-weight: bold">as</span> <span style="color: #008b45; text-decoration: underline">np</span>
|
||||
|
||||
<span style="color: #CD5555">"""</span>
|
||||
<span style="color: #CD5555">The following setup is just a suggestion, feel free to write it the way you like.</span>
|
||||
<span style="color: #CD5555">"""</span>
|
||||
|
||||
<span style="color: #228B22">#Setup problem described in the exercise</span>
|
||||
N = <span style="color: #B452CD">100</span> <span style="color: #228B22">#Nr of datapoints</span>
|
||||
M = <span style="color: #B452CD">2</span> <span style="color: #228B22">#Nr of features</span>
|
||||
x = np.random.rand(N) <span style="color: #228B22">#Uniformly generated x-values in [0,1]</span>
|
||||
y = <span style="color: #B452CD">5</span>*x**<span style="color: #B452CD">2</span> + <span style="color: #B452CD">0.1</span>*np.random.randn(N)
|
||||
X = np.c_[np.ones(N),x] <span style="color: #228B22">#Construct design matrix</span>
|
||||
|
||||
<span style="color: #228B22">#Compute beta according to normal equations to compare with GD solution</span>
|
||||
Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
|
||||
Xt_y = np.dot(X.transpose(),y)
|
||||
beta_NE = np.dot(Xt_X_inv,Xt_y)
|
||||
<span style="color: #8B008B; font-weight: bold">print</span>(beta_NE)
|
||||
</pre></div>
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec15">Gradient Descent Example </h2>
|
||||
|
||||
<p>
|
||||
Another simple example is here
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
|
||||
@@ -301,7 +559,7 @@ plt.show()
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec6">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
<h2 id="___sec16">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
|
||||
<p>
|
||||
|
||||
@@ -325,265 +583,6 @@ sgdreg.fit(x,y.ravel())
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec7">Convex functions </h2>
|
||||
|
||||
<p>
|
||||
Ideally we want our cost/loss function to be convex(concave).
|
||||
|
||||
<p>
|
||||
First we give the definition of a convex set: A set \( C \) in
|
||||
\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and
|
||||
all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to
|
||||
C. Geometrically this means that every point on the line segment
|
||||
connecting \( x \) and \( y \) is in \( C \) as discussed below.
|
||||
|
||||
<p>
|
||||
The convex subsets of \( \mathbb{R} \) are the intervals of
|
||||
\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the
|
||||
regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec8">Convex function </h2>
|
||||
|
||||
<p>
|
||||
<b>Convex function</b>: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec9">Conditions on convex functions </h2>
|
||||
|
||||
<p>
|
||||
In the following we state first and second-order conditions which
|
||||
ensures convexity of a function \( f \). We write \( D_f \) to denote the
|
||||
domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more
|
||||
details and proofs we refer to: <a href="http://stanford.edu/boyd/cvxbook/, 2004" target="_blank">S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press</a>.
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>First order condition.</b>
|
||||
<p>
|
||||
Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for
|
||||
all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \)
|
||||
is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
|
||||
for all \( x,y \in D_f \). This condition means that for a convex function
|
||||
the first order Taylor expansion (right hand side above) at any point
|
||||
a global under estimator of the function. To convince yourself you can
|
||||
make a drawing of f(x) = x^2+1 and draw the tangent line to \( f(x) \) and
|
||||
note that it is always below the graph.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>Second order condition.</b>
|
||||
<p>
|
||||
Assume that \( f \) is twice
|
||||
differentiable, i.e the Hessian matrix exists at each point in
|
||||
\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its
|
||||
Hessian is positive semi-definite for all \( x\in D_f \). For a
|
||||
single-variable function this reduces to \( f''(x) \geq
|
||||
0 \). Geometrically this means that \( f \) has nonnegative curvature
|
||||
everywhere.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec10">More on convex functions </h2>
|
||||
|
||||
<p>
|
||||
The next result is of great importance to us and the reason why we are
|
||||
going on about convex functions. In machine learning we frequently
|
||||
have to minimize a loss/cost function in order to find the best
|
||||
parameters for the model we are considering.
|
||||
|
||||
<p>
|
||||
Ideally we want the
|
||||
global minimum (for high-dimensional models it is hard to know
|
||||
if we have local or global minimum). However, if the cost/loss function
|
||||
is convex the following result provides invaluable information:
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>Any minimum is global for convex functions.</b>
|
||||
<p>
|
||||
Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \)
|
||||
is minimal, where \( f \) is convex and differentiable. Then, any point
|
||||
\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec11">Some simple problems </h2>
|
||||
|
||||
<ol>
|
||||
<li> Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x) + (1-\lambda)f(y) - f(\lambda x + (1-\lambda) y ) \geq 0. $</li>
|
||||
<li> Using the second order condition show that the following functions are convex on the specified domain.</li>
|
||||
|
||||
<ul>
|
||||
<li> \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).</li>
|
||||
<li> \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).</li>
|
||||
</ul>
|
||||
|
||||
<li> Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.</li>
|
||||
<li> A norm is any function that satisfy the following properties</li>
|
||||
|
||||
<ul>
|
||||
<li> \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).</li>
|
||||
<li> \( f(x+y) \leq f(x) + f(y) \)</li>
|
||||
<li> \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)</li>
|
||||
</ul>
|
||||
|
||||
</ol>
|
||||
|
||||
Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec12">Revisiting our first homework </h2>
|
||||
|
||||
<p>
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
<ol>
|
||||
<li> An analytical solution (recall homework set 1).</li>
|
||||
<li> The gradient can be computed analytically.</li>
|
||||
<li> The cost function is convex which guarantees that gradient descent converges for small enough learning rates</li>
|
||||
</ol>
|
||||
|
||||
We revisit the example from homework set 1 where we had
|
||||
$$
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
$$
|
||||
|
||||
with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
|
||||
The linear regression model is given by
|
||||
$$
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
$$
|
||||
|
||||
such that
|
||||
$$
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec13">Gradient descent example </h2>
|
||||
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
|
||||
<p>
|
||||
t is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by
|
||||
$$
|
||||
\begin{equation}
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
\label{_auto1}
|
||||
\end{equation}
|
||||
$$
|
||||
|
||||
The loss function is given by
|
||||
$$
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
$$
|
||||
|
||||
and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec14">The derivative of the cost/loss function </h2>
|
||||
|
||||
<p>
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
$$
|
||||
\nabla_\beta C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
$$
|
||||
|
||||
where \( X \) is the design matrix defined above.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec15">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
$$
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
$$
|
||||
|
||||
This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec16">Simple program </h2>
|
||||
|
||||
<p>
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
$$
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
$$
|
||||
|
||||
<p>
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
|
||||
when \( ||\nabla_\beta C(\beta_k) || < \epsilon = 10^{-8} \).
|
||||
|
||||
<p>
|
||||
And finally we can compare our solution for \( \beta \) with the analytic result given by
|
||||
\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "perldoc" -->
|
||||
<div class="highlight" style="background: #eeeedd"><pre style="line-height: 125%"><span></span><span style="color: #8B008B; font-weight: bold">import</span> <span style="color: #008b45; text-decoration: underline">numpy</span> <span style="color: #8B008B; font-weight: bold">as</span> <span style="color: #008b45; text-decoration: underline">np</span>
|
||||
|
||||
<span style="color: #CD5555">"""</span>
|
||||
<span style="color: #CD5555">The following setup is just a suggestion, feel free to write it the way you like.</span>
|
||||
<span style="color: #CD5555">"""</span>
|
||||
|
||||
<span style="color: #228B22">#Setup problem described in the exercise</span>
|
||||
N = <span style="color: #B452CD">100</span> <span style="color: #228B22">#Nr of datapoints</span>
|
||||
M = <span style="color: #B452CD">2</span> <span style="color: #228B22">#Nr of features</span>
|
||||
x = np.random.rand(N) <span style="color: #228B22">#Uniformly generated x-values in [0,1]</span>
|
||||
y = <span style="color: #B452CD">5</span>*x**<span style="color: #B452CD">2</span> + <span style="color: #B452CD">0.1</span>*np.random.randn(N)
|
||||
X = np.c_[np.ones(N),x] <span style="color: #228B22">#Construct design matrix</span>
|
||||
|
||||
<span style="color: #228B22">#Compute beta according to normal equations to compare with GD solution</span>
|
||||
Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
|
||||
Xt_y = np.dot(X.transpose(),y)
|
||||
beta_NE = np.dot(Xt_X_inv,Xt_y)
|
||||
<span style="color: #8B008B; font-weight: bold">print</span>(beta_NE)
|
||||
</pre></div>
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec17">Gradient descent and Ridge </h2>
|
||||
|
||||
<p>
|
||||
@@ -650,7 +649,7 @@ the shortcomings of the Gradient descent method discussed above.
|
||||
<p>
|
||||
The underlying idea of SGD comes from the observation that the cost
|
||||
function, which we want to minimize, can almost always be written as a
|
||||
sum over \( n \) datapoints \( \{\mathbf{x}_i\}_{i=1}^n \),
|
||||
sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \),
|
||||
$$
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}).
|
||||
@@ -672,7 +671,7 @@ $$
|
||||
<p>
|
||||
Stochasticity/randomness is introduced by only taking the
|
||||
gradient on a subset of the data called minibatches. If there are \( n \)
|
||||
datapoints and the size of each minibatch is \( M \), there will be \( n/M \)
|
||||
data points and the size of each minibatch is \( M \), there will be \( n/M \)
|
||||
minibatches. We denote these minibatches by \( B_k \) where
|
||||
\( k=1,\cdots,n/M \).
|
||||
|
||||
@@ -680,21 +679,21 @@ minibatches. We denote these minibatches by \( B_k \) where
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec20">SGD example </h2>
|
||||
As an example, suppose we have \( 10 \) datapoints \( ( \mathbf{x}_1,
|
||||
\cdots, \mathbf{x}_{10} ) \) and we choose to have \( M=5 \) minibathces,
|
||||
then each minibatch contains two datapoints. In particular we have
|
||||
As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \)
|
||||
and we choose to have \( M=5 \) minibathces,
|
||||
then each minibatch contains two data points. In particular we have
|
||||
\( B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 =
|
||||
(\mathbf{x}_9,\mathbf{x}_{10}) \). Note that if you choose \( M=1 \) you
|
||||
have only a single batch with all datapoints and on the other extreme,
|
||||
have only a single batch with all data points and on the other extreme,
|
||||
you may choose \( M=n \) resulting in a minibatch for each datapoint, i.e
|
||||
\( B_k = \mathbf{x}_k \).
|
||||
|
||||
<p>
|
||||
The idea is now to approximate the gradient by replacing the sum over
|
||||
all datapoints with a sum over the datapoints in one the minibatches
|
||||
all data points with a sum over the data points in one the minibatches
|
||||
picked at random in each gradient descent step
|
||||
$$
|
||||
\nabla_\beta
|
||||
\nabla_{\beta}
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta
|
||||
c_i(\mathbf{x}_i, \mathbf{\beta}).
|
||||
@@ -747,8 +746,8 @@ Taking the gradient only on a subset of the data has two important
|
||||
benefits. First, it introduces randomness which decreases the chance
|
||||
that our opmization scheme gets stuck in a local minima. Second, if
|
||||
the size of the minibatches are small relative to the number of
|
||||
datapoints (\( M < n \)), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the k-th minibatch and not
|
||||
datapoints (\( M < n \)), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the \( k-th \) minibatch and not
|
||||
all \( n \) datapoints.
|
||||
|
||||
<p>
|
||||
@@ -779,7 +778,7 @@ number of epochs in such a way that it becomes very small after a
|
||||
reasonable time such that we do not move at all.
|
||||
|
||||
<p>
|
||||
As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in <em>time</em> \( t \).
|
||||
As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in <em>time</em> \( t \).
|
||||
|
||||
<p>
|
||||
In this way we can fix the number of epochs, compute \( \beta \) and
|
||||
|
||||
+290
-291
@@ -74,21 +74,21 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
('More on Steepest descent', 2, None, '___sec2'),
|
||||
('The ideal', 2, None, '___sec3'),
|
||||
('The sensitiveness of the gradient descent', 2, None, '___sec4'),
|
||||
('Gradient Descent Example', 2, None, '___sec5'),
|
||||
('Convex functions', 2, None, '___sec5'),
|
||||
('Convex function', 2, None, '___sec6'),
|
||||
('Conditions on convex functions', 2, None, '___sec7'),
|
||||
('More on convex functions', 2, None, '___sec8'),
|
||||
('Some simple problems', 2, None, '___sec9'),
|
||||
('Revisiting our first homework', 2, None, '___sec10'),
|
||||
('Gradient descent example', 2, None, '___sec11'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec12'),
|
||||
('The Hessian matrix', 2, None, '___sec13'),
|
||||
('Simple program', 2, None, '___sec14'),
|
||||
('Gradient Descent Example', 2, None, '___sec15'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('Convex functions', 2, None, '___sec7'),
|
||||
('Convex function', 2, None, '___sec8'),
|
||||
('Conditions on convex functions', 2, None, '___sec9'),
|
||||
('More on convex functions', 2, None, '___sec10'),
|
||||
('Some simple problems', 2, None, '___sec11'),
|
||||
('Revisiting our first homework', 2, None, '___sec12'),
|
||||
('Gradient descent example', 2, None, '___sec13'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec14'),
|
||||
('The Hessian matrix', 2, None, '___sec15'),
|
||||
('Simple program', 2, None, '___sec16'),
|
||||
'___sec16'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec17'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec18'),
|
||||
('Computation of gradients', 2, None, '___sec19'),
|
||||
@@ -152,7 +152,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Sep 20, 2018</h4></center> <!-- date -->
|
||||
<center><h4>Sep 21, 2018</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -183,11 +183,13 @@ direction of the negative gradient \( -\nabla F(\mathbf{x}) \).
|
||||
<p>
|
||||
It can be shown that if
|
||||
$$
|
||||
\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), \ \ \gamma_k > 0
|
||||
\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k),
|
||||
$$
|
||||
|
||||
with \( \gamma_k > 0 \).
|
||||
|
||||
<p>
|
||||
for \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq
|
||||
For \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq
|
||||
F(\mathbf{x}_k) \). This means that for a sufficiently small \( \gamma_k \)
|
||||
we are always moving towards smaller function values, i.e a minimum.
|
||||
|
||||
@@ -216,7 +218,7 @@ the learning rate within the context of Machine Learning.
|
||||
<h2 id="___sec3">The ideal </h2>
|
||||
|
||||
<p>
|
||||
Ideally the sequence \( \{ \mathbf{x}_k \}_{k=0} \) converges to a global
|
||||
Ideally the sequence \( \{\mathbf{x}_k \}_{k=0} \) converges to a global
|
||||
minimum of the function \( F \). In general we do not know if we are in a
|
||||
global or local minimum. In the special case when \( F \) is a convex
|
||||
function, all local minima are also global minima, so in this case
|
||||
@@ -242,7 +244,8 @@ Note that the gradient is a function of \( \mathbf{x} =
|
||||
<h2 id="___sec4">The sensitiveness of the gradient descent </h2>
|
||||
|
||||
<p>
|
||||
GD is sensitive to the choice of learning rate \( \gamma_k \). This is due
|
||||
The gradient descent method
|
||||
is sensitive to the choice of learning rate \( \gamma_k \). This is due
|
||||
to the fact that we are only guaranteed that \( F(\mathbf{x}_{k+1}) \leq
|
||||
F(\mathbf{x}_k) \) for sufficiently small \( \gamma_k \). The problem is to
|
||||
determine an optimal learning rate. If the learning rate is chosen too
|
||||
@@ -255,12 +258,267 @@ randomness. One such method is that of Stochastic Gradient Descent
|
||||
(SGD), see below.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec5">Gradient Descent Example </h2>
|
||||
<h2 id="___sec5">Convex functions </h2>
|
||||
|
||||
<p>
|
||||
We revisit now our simple linear regression example with a linear polynomial.
|
||||
Ideally we want our cost/loss function to be convex(concave).
|
||||
|
||||
<p>
|
||||
First we give the definition of a convex set: A set \( C \) in
|
||||
\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and
|
||||
all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to
|
||||
C. Geometrically this means that every point on the line segment
|
||||
connecting \( x \) and \( y \) is in \( C \) as discussed below.
|
||||
|
||||
<p>
|
||||
The convex subsets of \( \mathbb{R} \) are the intervals of
|
||||
\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the
|
||||
regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec6">Convex function </h2>
|
||||
|
||||
<p>
|
||||
<b>Convex function</b>: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec7">Conditions on convex functions </h2>
|
||||
|
||||
<p>
|
||||
In the following we state first and second-order conditions which
|
||||
ensures convexity of a function \( f \). We write \( D_f \) to denote the
|
||||
domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more
|
||||
details and proofs we refer to: <a href="http://stanford.edu/boyd/cvxbook/, 2004" target="_blank">S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press</a>.
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>First order condition.</b>
|
||||
<p>
|
||||
Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for
|
||||
all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \)
|
||||
is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
|
||||
for all \( x,y \in D_f \). This condition means that for a convex function
|
||||
the first order Taylor expansion (right hand side above) at any point
|
||||
a global under estimator of the function. To convince yourself you can
|
||||
make a drawing of \( f(x) = x^2+1 \) and draw the tangent line to \( f(x) \) and
|
||||
note that it is always below the graph.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>Second order condition.</b>
|
||||
<p>
|
||||
Assume that \( f \) is twice
|
||||
differentiable, i.e the Hessian matrix exists at each point in
|
||||
\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its
|
||||
Hessian is positive semi-definite for all \( x\in D_f \). For a
|
||||
single-variable function this reduces to \( f''(x) \geq 0 \). Geometrically this means that \( f \) has nonnegative curvature
|
||||
everywhere.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec8">More on convex functions </h2>
|
||||
|
||||
<p>
|
||||
The next result is of great importance to us and the reason why we are
|
||||
going on about convex functions. In machine learning we frequently
|
||||
have to minimize a loss/cost function in order to find the best
|
||||
parameters for the model we are considering.
|
||||
|
||||
<p>
|
||||
Ideally we want the
|
||||
global minimum (for high-dimensional models it is hard to know
|
||||
if we have local or global minimum). However, if the cost/loss function
|
||||
is convex the following result provides invaluable information:
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>Any minimum is global for convex functions.</b>
|
||||
<p>
|
||||
Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \)
|
||||
is minimal, where \( f \) is convex and differentiable. Then, any point
|
||||
\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec9">Some simple problems </h2>
|
||||
|
||||
<ol>
|
||||
<li> Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.</li>
|
||||
<li> Using the second order condition show that the following functions are convex on the specified domain.</li>
|
||||
|
||||
<ul>
|
||||
<li> \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).</li>
|
||||
<li> \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).</li>
|
||||
</ul>
|
||||
|
||||
<li> Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.</li>
|
||||
<li> A norm is any function that satisfy the following properties</li>
|
||||
|
||||
<ul>
|
||||
<li> \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).</li>
|
||||
<li> \( f(x+y) \leq f(x) + f(y) \)</li>
|
||||
<li> \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)</li>
|
||||
</ul>
|
||||
|
||||
</ol>
|
||||
|
||||
Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec10">Revisiting our first homework </h2>
|
||||
|
||||
<p>
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
<ol>
|
||||
<li> An analytical solution (recall homework set 1).</li>
|
||||
<li> The gradient can be computed analytically.</li>
|
||||
<li> The cost function is convex which guarantees that gradient descent converges for small enough learning rates</li>
|
||||
</ol>
|
||||
|
||||
We revisit the example from homework set 1 where we had
|
||||
$$
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
$$
|
||||
|
||||
with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
|
||||
The linear regression model is given by
|
||||
$$
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
$$
|
||||
|
||||
such that
|
||||
$$
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec11">Gradient descent example </h2>
|
||||
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
|
||||
<p>
|
||||
It is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by
|
||||
$$
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
$$
|
||||
|
||||
The loss function is given by
|
||||
$$
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
$$
|
||||
|
||||
and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec12">The derivative of the cost/loss function </h2>
|
||||
|
||||
<p>
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
$$
|
||||
\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
$$
|
||||
|
||||
where \( X \) is the design matrix defined above.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec13">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
$$
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
$$
|
||||
|
||||
This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec14">Simple program </h2>
|
||||
|
||||
<p>
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
$$
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
$$
|
||||
|
||||
<p>
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
|
||||
when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \).
|
||||
|
||||
<p>
|
||||
And finally we can compare our solution for \( \beta \) with the analytic result given by
|
||||
\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
<span style="color: #BA2121; font-style: italic">The following setup is just a suggestion, feel free to write it the way you like.</span>
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Setup problem described in the exercise</span>
|
||||
N <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#Nr of datapoints</span>
|
||||
M <span style="color: #666666">=</span> <span style="color: #666666">2</span> <span style="color: #408080; font-style: italic">#Nr of features</span>
|
||||
x <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(N) <span style="color: #408080; font-style: italic">#Uniformly generated x-values in [0,1]</span>
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">5*</span>x<span style="color: #666666">**2</span> <span style="color: #666666">+</span> <span style="color: #666666">0.1*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(N)
|
||||
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones(N),x] <span style="color: #408080; font-style: italic">#Construct design matrix</span>
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Compute beta according to normal equations to compare with GD solution</span>
|
||||
Xt_X_inv <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,X))
|
||||
Xt_y <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>transpose(),y)
|
||||
beta_NE <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(Xt_X_inv,Xt_y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_NE)
|
||||
</pre></div>
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec15">Gradient Descent Example </h2>
|
||||
|
||||
<p>
|
||||
Another simple example is here
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
@@ -306,7 +564,7 @@ plt<span style="color: #666666">.</span>show()
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec6">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
<h2 id="___sec16">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
|
||||
<p>
|
||||
|
||||
@@ -330,265 +588,6 @@ sgdreg<span style="color: #666666">.</span>fit(x,y<span style="color: #666666">.
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec7">Convex functions </h2>
|
||||
|
||||
<p>
|
||||
Ideally we want our cost/loss function to be convex(concave).
|
||||
|
||||
<p>
|
||||
First we give the definition of a convex set: A set \( C \) in
|
||||
\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and
|
||||
all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to
|
||||
C. Geometrically this means that every point on the line segment
|
||||
connecting \( x \) and \( y \) is in \( C \) as discussed below.
|
||||
|
||||
<p>
|
||||
The convex subsets of \( \mathbb{R} \) are the intervals of
|
||||
\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the
|
||||
regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec8">Convex function </h2>
|
||||
|
||||
<p>
|
||||
<b>Convex function</b>: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec9">Conditions on convex functions </h2>
|
||||
|
||||
<p>
|
||||
In the following we state first and second-order conditions which
|
||||
ensures convexity of a function \( f \). We write \( D_f \) to denote the
|
||||
domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more
|
||||
details and proofs we refer to: <a href="http://stanford.edu/boyd/cvxbook/, 2004" target="_blank">S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press</a>.
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>First order condition.</b>
|
||||
<p>
|
||||
Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for
|
||||
all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \)
|
||||
is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
|
||||
for all \( x,y \in D_f \). This condition means that for a convex function
|
||||
the first order Taylor expansion (right hand side above) at any point
|
||||
a global under estimator of the function. To convince yourself you can
|
||||
make a drawing of f(x) = x^2+1 and draw the tangent line to \( f(x) \) and
|
||||
note that it is always below the graph.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>Second order condition.</b>
|
||||
<p>
|
||||
Assume that \( f \) is twice
|
||||
differentiable, i.e the Hessian matrix exists at each point in
|
||||
\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its
|
||||
Hessian is positive semi-definite for all \( x\in D_f \). For a
|
||||
single-variable function this reduces to \( f''(x) \geq
|
||||
0 \). Geometrically this means that \( f \) has nonnegative curvature
|
||||
everywhere.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec10">More on convex functions </h2>
|
||||
|
||||
<p>
|
||||
The next result is of great importance to us and the reason why we are
|
||||
going on about convex functions. In machine learning we frequently
|
||||
have to minimize a loss/cost function in order to find the best
|
||||
parameters for the model we are considering.
|
||||
|
||||
<p>
|
||||
Ideally we want the
|
||||
global minimum (for high-dimensional models it is hard to know
|
||||
if we have local or global minimum). However, if the cost/loss function
|
||||
is convex the following result provides invaluable information:
|
||||
|
||||
<p>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b>Any minimum is global for convex functions.</b>
|
||||
<p>
|
||||
Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \)
|
||||
is minimal, where \( f \) is convex and differentiable. Then, any point
|
||||
\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec11">Some simple problems </h2>
|
||||
|
||||
<ol>
|
||||
<li> Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x) + (1-\lambda)f(y) - f(\lambda x + (1-\lambda) y ) \geq 0. $</li>
|
||||
<li> Using the second order condition show that the following functions are convex on the specified domain.</li>
|
||||
|
||||
<ul>
|
||||
<li> \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).</li>
|
||||
<li> \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).</li>
|
||||
</ul>
|
||||
|
||||
<li> Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.</li>
|
||||
<li> A norm is any function that satisfy the following properties</li>
|
||||
|
||||
<ul>
|
||||
<li> \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).</li>
|
||||
<li> \( f(x+y) \leq f(x) + f(y) \)</li>
|
||||
<li> \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)</li>
|
||||
</ul>
|
||||
|
||||
</ol>
|
||||
|
||||
Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec12">Revisiting our first homework </h2>
|
||||
|
||||
<p>
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
<ol>
|
||||
<li> An analytical solution (recall homework set 1).</li>
|
||||
<li> The gradient can be computed analytically.</li>
|
||||
<li> The cost function is convex which guarantees that gradient descent converges for small enough learning rates</li>
|
||||
</ol>
|
||||
|
||||
We revisit the example from homework set 1 where we had
|
||||
$$
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
$$
|
||||
|
||||
with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
|
||||
The linear regression model is given by
|
||||
$$
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
$$
|
||||
|
||||
such that
|
||||
$$
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec13">Gradient descent example </h2>
|
||||
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
|
||||
<p>
|
||||
t is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by
|
||||
$$
|
||||
\begin{equation}
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
\label{_auto1}
|
||||
\end{equation}
|
||||
$$
|
||||
|
||||
The loss function is given by
|
||||
$$
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
$$
|
||||
|
||||
and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec14">The derivative of the cost/loss function </h2>
|
||||
|
||||
<p>
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
$$
|
||||
\nabla_\beta C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
$$
|
||||
|
||||
where \( X \) is the design matrix defined above.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec15">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
$$
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
$$
|
||||
|
||||
This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec16">Simple program </h2>
|
||||
|
||||
<p>
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
$$
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
$$
|
||||
|
||||
<p>
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
|
||||
when \( ||\nabla_\beta C(\beta_k) || < \epsilon = 10^{-8} \).
|
||||
|
||||
<p>
|
||||
And finally we can compare our solution for \( \beta \) with the analytic result given by
|
||||
\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
<span style="color: #BA2121; font-style: italic">The following setup is just a suggestion, feel free to write it the way you like.</span>
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Setup problem described in the exercise</span>
|
||||
N <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#Nr of datapoints</span>
|
||||
M <span style="color: #666666">=</span> <span style="color: #666666">2</span> <span style="color: #408080; font-style: italic">#Nr of features</span>
|
||||
x <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(N) <span style="color: #408080; font-style: italic">#Uniformly generated x-values in [0,1]</span>
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">5*</span>x<span style="color: #666666">**2</span> <span style="color: #666666">+</span> <span style="color: #666666">0.1*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(N)
|
||||
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones(N),x] <span style="color: #408080; font-style: italic">#Construct design matrix</span>
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Compute beta according to normal equations to compare with GD solution</span>
|
||||
Xt_X_inv <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,X))
|
||||
Xt_y <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>transpose(),y)
|
||||
beta_NE <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(Xt_X_inv,Xt_y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_NE)
|
||||
</pre></div>
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec17">Gradient descent and Ridge </h2>
|
||||
|
||||
<p>
|
||||
@@ -655,7 +654,7 @@ the shortcomings of the Gradient descent method discussed above.
|
||||
<p>
|
||||
The underlying idea of SGD comes from the observation that the cost
|
||||
function, which we want to minimize, can almost always be written as a
|
||||
sum over \( n \) datapoints \( \{\mathbf{x}_i\}_{i=1}^n \),
|
||||
sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \),
|
||||
$$
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}).
|
||||
@@ -677,7 +676,7 @@ $$
|
||||
<p>
|
||||
Stochasticity/randomness is introduced by only taking the
|
||||
gradient on a subset of the data called minibatches. If there are \( n \)
|
||||
datapoints and the size of each minibatch is \( M \), there will be \( n/M \)
|
||||
data points and the size of each minibatch is \( M \), there will be \( n/M \)
|
||||
minibatches. We denote these minibatches by \( B_k \) where
|
||||
\( k=1,\cdots,n/M \).
|
||||
|
||||
@@ -685,21 +684,21 @@ minibatches. We denote these minibatches by \( B_k \) where
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec20">SGD example </h2>
|
||||
As an example, suppose we have \( 10 \) datapoints \( ( \mathbf{x}_1,
|
||||
\cdots, \mathbf{x}_{10} ) \) and we choose to have \( M=5 \) minibathces,
|
||||
then each minibatch contains two datapoints. In particular we have
|
||||
As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \)
|
||||
and we choose to have \( M=5 \) minibathces,
|
||||
then each minibatch contains two data points. In particular we have
|
||||
\( B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 =
|
||||
(\mathbf{x}_9,\mathbf{x}_{10}) \). Note that if you choose \( M=1 \) you
|
||||
have only a single batch with all datapoints and on the other extreme,
|
||||
have only a single batch with all data points and on the other extreme,
|
||||
you may choose \( M=n \) resulting in a minibatch for each datapoint, i.e
|
||||
\( B_k = \mathbf{x}_k \).
|
||||
|
||||
<p>
|
||||
The idea is now to approximate the gradient by replacing the sum over
|
||||
all datapoints with a sum over the datapoints in one the minibatches
|
||||
all data points with a sum over the data points in one the minibatches
|
||||
picked at random in each gradient descent step
|
||||
$$
|
||||
\nabla_\beta
|
||||
\nabla_{\beta}
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta
|
||||
c_i(\mathbf{x}_i, \mathbf{\beta}).
|
||||
@@ -752,8 +751,8 @@ Taking the gradient only on a subset of the data has two important
|
||||
benefits. First, it introduces randomness which decreases the chance
|
||||
that our opmization scheme gets stuck in a local minima. Second, if
|
||||
the size of the minibatches are small relative to the number of
|
||||
datapoints (\( M < n \)), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the k-th minibatch and not
|
||||
datapoints (\( M < n \)), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the \( k-th \) minibatch and not
|
||||
all \( n \) datapoints.
|
||||
|
||||
<p>
|
||||
@@ -784,7 +783,7 @@ number of epochs in such a way that it becomes very small after a
|
||||
reasonable time such that we do not move at all.
|
||||
|
||||
<p>
|
||||
As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in <em>time</em> \( t \).
|
||||
As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in <em>time</em> \( t \).
|
||||
|
||||
<p>
|
||||
In this way we can fix the number of epochs, compute \( \beta \) and
|
||||
|
||||
+118
-121
@@ -10,7 +10,7 @@
|
||||
"<!-- Author: --> \n",
|
||||
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
|
||||
"\n",
|
||||
"Date: **Sep 20, 2018**\n",
|
||||
"Date: **Sep 21, 2018**\n",
|
||||
"\n",
|
||||
"Copyright 1999-2018, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
|
||||
"\n",
|
||||
@@ -43,7 +43,7 @@
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\mathbf{x}_{k+1} = \\mathbf{x}_k - \\gamma_k \\nabla F(\\mathbf{x}_k), \\ \\ \\gamma_k > 0\n",
|
||||
"\\mathbf{x}_{k+1} = \\mathbf{x}_k - \\gamma_k \\nabla F(\\mathbf{x}_k),\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
@@ -51,7 +51,9 @@
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"for $\\gamma_k$ small enough, then $F(\\mathbf{x}_{k+1}) \\leq\n",
|
||||
"with $\\gamma_k > 0$.\n",
|
||||
"\n",
|
||||
"For $\\gamma_k$ small enough, then $F(\\mathbf{x}_{k+1}) \\leq\n",
|
||||
"F(\\mathbf{x}_k)$. This means that for a sufficiently small $\\gamma_k$\n",
|
||||
"we are always moving towards smaller function values, i.e a minimum.\n",
|
||||
"\n",
|
||||
@@ -83,7 +85,7 @@
|
||||
"<!-- !split -->\n",
|
||||
"## The ideal\n",
|
||||
"\n",
|
||||
"Ideally the sequence $\\{ \\mathbf{x}_k \\}_{k=0}$ converges to a global\n",
|
||||
"Ideally the sequence $\\{\\mathbf{x}_k \\}_{k=0}$ converges to a global\n",
|
||||
"minimum of the function $F$. In general we do not know if we are in a\n",
|
||||
"global or local minimum. In the special case when $F$ is a convex\n",
|
||||
"function, all local minima are also global minima, so in this case\n",
|
||||
@@ -105,7 +107,8 @@
|
||||
"<!-- !split -->\n",
|
||||
"## The sensitiveness of the gradient descent\n",
|
||||
"\n",
|
||||
"GD is sensitive to the choice of learning rate $\\gamma_k$. This is due\n",
|
||||
"The gradient descent method \n",
|
||||
"is sensitive to the choice of learning rate $\\gamma_k$. This is due\n",
|
||||
"to the fact that we are only guaranteed that $F(\\mathbf{x}_{k+1}) \\leq\n",
|
||||
"F(\\mathbf{x}_k)$ for sufficiently small $\\gamma_k$. The problem is to\n",
|
||||
"determine an optimal learning rate. If the learning rate is chosen too\n",
|
||||
@@ -116,98 +119,7 @@
|
||||
"randomness. One such method is that of Stochastic Gradient Descent\n",
|
||||
"(SGD), see below.\n",
|
||||
"\n",
|
||||
"## Gradient Descent Example\n",
|
||||
"\n",
|
||||
"We revisit now our simple linear regression example with a linear polynomial."
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"%matplotlib inline\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"# Importing various packages\n",
|
||||
"from random import random, seed\n",
|
||||
"import numpy as np\n",
|
||||
"import matplotlib.pyplot as plt\n",
|
||||
"from mpl_toolkits.mplot3d import Axes3D\n",
|
||||
"from matplotlib import cm\n",
|
||||
"from matplotlib.ticker import LinearLocator, FormatStrFormatter\n",
|
||||
"import sys\n",
|
||||
"\n",
|
||||
"x = 2*np.random.rand(100,1)\n",
|
||||
"y = 4+3*x+np.random.randn(100,1)\n",
|
||||
"\n",
|
||||
"xb = np.c_[np.ones((100,1)), x]\n",
|
||||
"beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n",
|
||||
"print(beta_linreg)\n",
|
||||
"beta = np.random.randn(2,1)\n",
|
||||
"\n",
|
||||
"eta = 0.1\n",
|
||||
"Niterations = 1000\n",
|
||||
"m = 100\n",
|
||||
"\n",
|
||||
"for iter in range(Niterations):\n",
|
||||
" gradients = 2.0/m*xb.T.dot(xb.dot(beta)-y)\n",
|
||||
" beta -= eta*gradients\n",
|
||||
"\n",
|
||||
"print(beta)\n",
|
||||
"xnew = np.array([[0],[2]])\n",
|
||||
"xbnew = np.c_[np.ones((2,1)), xnew]\n",
|
||||
"ypredict = xbnew.dot(beta)\n",
|
||||
"ypredict2 = xbnew.dot(beta_linreg)\n",
|
||||
"plt.plot(xnew, ypredict, \"r-\")\n",
|
||||
"plt.plot(xnew, ypredict2, \"b-\")\n",
|
||||
"plt.plot(x, y ,'ro')\n",
|
||||
"plt.axis([0,2.0,0, 15.0])\n",
|
||||
"plt.xlabel(r'$x$')\n",
|
||||
"plt.ylabel(r'$y$')\n",
|
||||
"plt.title(r'Gradient descent example')\n",
|
||||
"plt.show()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## And a corresponding example using **scikit-learn**"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Importing various packages\n",
|
||||
"from random import random, seed\n",
|
||||
"import numpy as np\n",
|
||||
"import matplotlib.pyplot as plt\n",
|
||||
"from sklearn.linear_model import SGDRegressor\n",
|
||||
"\n",
|
||||
"x = 2*np.random.rand(100,1)\n",
|
||||
"y = 4+3*x+np.random.randn(100,1)\n",
|
||||
"\n",
|
||||
"xb = np.c_[np.ones((100,1)), x]\n",
|
||||
"beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n",
|
||||
"print(beta_linreg)\n",
|
||||
"sgdreg = SGDRegressor(n_iter = 50, penalty=None, eta0=0.1)\n",
|
||||
"sgdreg.fit(x,y.ravel())\n",
|
||||
"print(sgdreg.intercept_, sgdreg.coef_)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"<!-- !split -->\n",
|
||||
"## Convex functions\n",
|
||||
"\n",
|
||||
@@ -242,7 +154,7 @@
|
||||
"for all $x,y \\in D_f$. This condition means that for a convex function\n",
|
||||
"the first order Taylor expansion (right hand side above) at any point\n",
|
||||
"a global under estimator of the function. To convince yourself you can\n",
|
||||
"make a drawing of f(x) = x^2+1 and draw the tangent line to $f(x)$ and\n",
|
||||
"make a drawing of $f(x) = x^2+1$ and draw the tangent line to $f(x)$ and\n",
|
||||
"note that it is always below the graph.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
@@ -253,8 +165,7 @@
|
||||
"differentiable, i.e the Hessian matrix exists at each point in\n",
|
||||
"$D_f$. Then $f$ is convex if and only if $D_f$ is a convex set and its\n",
|
||||
"Hessian is positive semi-definite for all $x\\in D_f$. For a\n",
|
||||
"single-variable function this reduces to $f''(x) \\geq\n",
|
||||
"0$. Geometrically this means that $f$ has nonnegative curvature\n",
|
||||
"single-variable function this reduces to $f''(x) \\geq 0$. Geometrically this means that $f$ has nonnegative curvature\n",
|
||||
"everywhere.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
@@ -285,7 +196,7 @@
|
||||
"\n",
|
||||
"## Some simple problems\n",
|
||||
"\n",
|
||||
"1. Show that $f(x)=x^2$ is convex for $x \\in \\mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \\in D_f$ and any $\\lambda \\in [0,1] $ $\\lambda f(x) + (1-\\lambda)f(y) - f(\\lambda x + (1-\\lambda) y ) \\geq 0. $\n",
|
||||
"1. Show that $f(x)=x^2$ is convex for $x \\in \\mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \\in D_f$ and any $\\lambda \\in [0,1]$ $\\lambda f(x)+(1-\\lambda)f(y)-f(\\lambda x + (1-\\lambda) y ) \\geq 0$.\n",
|
||||
"\n",
|
||||
"2. Using the second order condition show that the following functions are convex on the specified domain.\n",
|
||||
"\n",
|
||||
@@ -375,25 +286,19 @@
|
||||
"\n",
|
||||
"Let $\\mathbf{y} = (y_1,\\cdots,y_n)^T$, $\\mathbf{\\hat{y}} = (\\hat{y}_1,\\cdots,\\hat{y}_n)^T$ and $\\beta = (\\beta_0, \\beta_1)^T$\n",
|
||||
"\n",
|
||||
"t is convenient to write $\\mathbf{\\hat{y}} = X\\beta$ where $X \\in \\mathbb{R}^{100 \\times 2} $ is the design matrix given by"
|
||||
"It is convenient to write $\\mathbf{\\hat{y}} = X\\beta$ where $X \\in \\mathbb{R}^{100 \\times 2} $ is the design matrix given by"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"<!-- Equation labels as ordinary links -->\n",
|
||||
"<div id=\"_auto1\"></div>\n",
|
||||
"\n",
|
||||
"$$\n",
|
||||
"\\begin{equation}\n",
|
||||
"X \\equiv \\begin{bmatrix}\n",
|
||||
"1 & x_1 \\\\\n",
|
||||
"\\vdots & \\vdots \\\\\n",
|
||||
"1 & x_{100} & \\\\\n",
|
||||
"\\end{bmatrix}.\n",
|
||||
"\\label{_auto1} \\tag{1}\n",
|
||||
"\\end{equation}\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
@@ -429,7 +334,7 @@
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\nabla_\\beta C(\\beta) = (\\partial C(\\beta) / \\partial \\beta_0, \\partial C(\\beta) / \\partial \\beta_1)^T = 2\\begin{bmatrix} \\sum_{i=1}^{100} \\left(\\beta_0+\\beta_1x_i-y_i\\right) \\\\\n",
|
||||
"\\nabla_{\\beta} C(\\beta) = (\\partial C(\\beta) / \\partial \\beta_0, \\partial C(\\beta) / \\partial \\beta_1)^T = 2\\begin{bmatrix} \\sum_{i=1}^{100} \\left(\\beta_0+\\beta_1x_i-y_i\\right) \\\\\n",
|
||||
"\\sum_{i=1}^{100}\\left( x_i (\\beta_0+\\beta_1x_i)-y_ix_i\\right) \\\\\n",
|
||||
"\\end{bmatrix} = 2X^T(X\\beta - \\mathbf{y}),\n",
|
||||
"$$"
|
||||
@@ -483,7 +388,7 @@
|
||||
"source": [
|
||||
"We can use the expression we computed for the gradient and let use a\n",
|
||||
"$\\beta_0$ be chosen randomly and let $\\gamma = 0.001$. Stop iterating\n",
|
||||
"when $||\\nabla_\\beta C(\\beta_k) || < \\epsilon = 10^{-8}$. \n",
|
||||
"when $||\\nabla_\\beta C(\\beta_k) || \\leq \\epsilon = 10^{-8}$. \n",
|
||||
"\n",
|
||||
"And finally we can compare our solution for $\\beta$ with the analytic result given by \n",
|
||||
"$\\beta= (X^TX)^{-1} X^T \\mathbf{y}$."
|
||||
@@ -491,7 +396,7 @@
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"execution_count": 1,
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
@@ -517,6 +422,98 @@
|
||||
"print(beta_NE)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Gradient Descent Example\n",
|
||||
"\n",
|
||||
"Another simple example is here"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"%matplotlib inline\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"# Importing various packages\n",
|
||||
"from random import random, seed\n",
|
||||
"import numpy as np\n",
|
||||
"import matplotlib.pyplot as plt\n",
|
||||
"from mpl_toolkits.mplot3d import Axes3D\n",
|
||||
"from matplotlib import cm\n",
|
||||
"from matplotlib.ticker import LinearLocator, FormatStrFormatter\n",
|
||||
"import sys\n",
|
||||
"\n",
|
||||
"x = 2*np.random.rand(100,1)\n",
|
||||
"y = 4+3*x+np.random.randn(100,1)\n",
|
||||
"\n",
|
||||
"xb = np.c_[np.ones((100,1)), x]\n",
|
||||
"beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n",
|
||||
"print(beta_linreg)\n",
|
||||
"beta = np.random.randn(2,1)\n",
|
||||
"\n",
|
||||
"eta = 0.1\n",
|
||||
"Niterations = 1000\n",
|
||||
"m = 100\n",
|
||||
"\n",
|
||||
"for iter in range(Niterations):\n",
|
||||
" gradients = 2.0/m*xb.T.dot(xb.dot(beta)-y)\n",
|
||||
" beta -= eta*gradients\n",
|
||||
"\n",
|
||||
"print(beta)\n",
|
||||
"xnew = np.array([[0],[2]])\n",
|
||||
"xbnew = np.c_[np.ones((2,1)), xnew]\n",
|
||||
"ypredict = xbnew.dot(beta)\n",
|
||||
"ypredict2 = xbnew.dot(beta_linreg)\n",
|
||||
"plt.plot(xnew, ypredict, \"r-\")\n",
|
||||
"plt.plot(xnew, ypredict2, \"b-\")\n",
|
||||
"plt.plot(x, y ,'ro')\n",
|
||||
"plt.axis([0,2.0,0, 15.0])\n",
|
||||
"plt.xlabel(r'$x$')\n",
|
||||
"plt.ylabel(r'$y$')\n",
|
||||
"plt.title(r'Gradient descent example')\n",
|
||||
"plt.show()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## And a corresponding example using **scikit-learn**"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {
|
||||
"collapsed": false
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Importing various packages\n",
|
||||
"from random import random, seed\n",
|
||||
"import numpy as np\n",
|
||||
"import matplotlib.pyplot as plt\n",
|
||||
"from sklearn.linear_model import SGDRegressor\n",
|
||||
"\n",
|
||||
"x = 2*np.random.rand(100,1)\n",
|
||||
"y = 4+3*x+np.random.randn(100,1)\n",
|
||||
"\n",
|
||||
"xb = np.c_[np.ones((100,1)), x]\n",
|
||||
"beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n",
|
||||
"print(beta_linreg)\n",
|
||||
"sgdreg = SGDRegressor(n_iter = 50, penalty=None, eta0=0.1)\n",
|
||||
"sgdreg.fit(x,y.ravel())\n",
|
||||
"print(sgdreg.intercept_, sgdreg.coef_)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
@@ -624,7 +621,7 @@
|
||||
"\n",
|
||||
"The underlying idea of SGD comes from the observation that the cost\n",
|
||||
"function, which we want to minimize, can almost always be written as a\n",
|
||||
"sum over $n$ datapoints $\\{\\mathbf{x}_i\\}_{i=1}^n$,"
|
||||
"sum over $n$ data points $\\{\\mathbf{x}_i\\}_{i=1}^n$,"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -663,22 +660,22 @@
|
||||
"source": [
|
||||
"Stochasticity/randomness is introduced by only taking the\n",
|
||||
"gradient on a subset of the data called minibatches. If there are $n$\n",
|
||||
"datapoints and the size of each minibatch is $M$, there will be $n/M$\n",
|
||||
"data points and the size of each minibatch is $M$, there will be $n/M$\n",
|
||||
"minibatches. We denote these minibatches by $B_k$ where\n",
|
||||
"$k=1,\\cdots,n/M$.\n",
|
||||
"\n",
|
||||
"## SGD example\n",
|
||||
"As an example, suppose we have $10$ datapoints $( \\mathbf{x}_1,\n",
|
||||
"\\cdots, \\mathbf{x}_{10} )$ and we choose to have $M=5$ minibathces,\n",
|
||||
"then each minibatch contains two datapoints. In particular we have\n",
|
||||
"As an example, suppose we have $10$ data points $(\\mathbf{x}_1,\\cdots, \\mathbf{x}_{10})$ \n",
|
||||
"and we choose to have $M=5$ minibathces,\n",
|
||||
"then each minibatch contains two data points. In particular we have\n",
|
||||
"$B_1 = (\\mathbf{x}_1,\\mathbf{x}_2), \\cdots, B_5 =\n",
|
||||
"(\\mathbf{x}_9,\\mathbf{x}_{10})$. Note that if you choose $M=1$ you\n",
|
||||
"have only a single batch with all datapoints and on the other extreme,\n",
|
||||
"have only a single batch with all data points and on the other extreme,\n",
|
||||
"you may choose $M=n$ resulting in a minibatch for each datapoint, i.e\n",
|
||||
"$B_k = \\mathbf{x}_k$.\n",
|
||||
"\n",
|
||||
"The idea is now to approximate the gradient by replacing the sum over\n",
|
||||
"all datapoints with a sum over the datapoints in one the minibatches\n",
|
||||
"all data points with a sum over the data points in one the minibatches\n",
|
||||
"picked at random in each gradient descent step"
|
||||
]
|
||||
},
|
||||
@@ -687,7 +684,7 @@
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\nabla_\\beta\n",
|
||||
"\\nabla_{\\beta}\n",
|
||||
"C(\\mathbf{\\beta}) = \\sum_{i=1}^n \\nabla_\\beta c_i(\\mathbf{x}_i,\n",
|
||||
"\\mathbf{\\beta}) \\rightarrow \\sum_{i \\in B_k}^n \\nabla_\\beta\n",
|
||||
"c_i(\\mathbf{x}_i, \\mathbf{\\beta}).\n",
|
||||
@@ -758,8 +755,8 @@
|
||||
"benefits. First, it introduces randomness which decreases the chance\n",
|
||||
"that our opmization scheme gets stuck in a local minima. Second, if\n",
|
||||
"the size of the minibatches are small relative to the number of\n",
|
||||
"datapoints ($M < n$), the computation of the gradient is much\n",
|
||||
"cheaper since we sum over the datapoints in the k-th minibatch and not\n",
|
||||
"datapoints ($M < n$), the computation of the gradient is much\n",
|
||||
"cheaper since we sum over the datapoints in the $k-th$ minibatch and not\n",
|
||||
"all $n$ datapoints.\n",
|
||||
"\n",
|
||||
"## When do we stop?\n",
|
||||
@@ -781,7 +778,7 @@
|
||||
"number of epochs in such a way that it becomes very small after a\n",
|
||||
"reasonable time such that we do not move at all.\n",
|
||||
"\n",
|
||||
"As an example, let $e = 0,1,2,3,\\cdots$ denote the current epoch and let $t_0, t_1 > 0$ be two fixed numbers. Furthermore, let $t = e \\cdot m + i$ where $m$ is the number of minibatches and $i=0,\\cdots,m-1$. Then the function $$\\gamma_j(t; t_0, t_1) = \\frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length $\\gamma_j (0; t_0, t_1) = t_0/t_1$ which decays in *time* $t$.\n",
|
||||
"As an example, let $e = 0,1,2,3,\\cdots$ denote the current epoch and let $t_0, t_1 > 0$ be two fixed numbers. Furthermore, let $t = e \\cdot m + i$ where $m$ is the number of minibatches and $i=0,\\cdots,m-1$. Then the function $$\\gamma_j(t; t_0, t_1) = \\frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length $\\gamma_j (0; t_0, t_1) = t_0/t_1$ which decays in *time* $t$.\n",
|
||||
"\n",
|
||||
"In this way we can fix the number of epochs, compute $\\beta$ and\n",
|
||||
"evaluate the cost function at the end. Repeating the computation will\n",
|
||||
|
||||
Binary file not shown.
Binary file not shown.
+225
-223
@@ -26,11 +26,12 @@ direction of the negative gradient $-\nabla F(\mathbf{x})$.
|
||||
It can be shown that if
|
||||
!bt
|
||||
\[
|
||||
\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), \ \ \gamma_k > 0
|
||||
\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k),
|
||||
\]
|
||||
!et
|
||||
with $\gamma_k > 0$.
|
||||
|
||||
for $\gamma_k$ small enough, then $F(\mathbf{x}_{k+1}) \leq
|
||||
For $\gamma_k$ small enough, then $F(\mathbf{x}_{k+1}) \leq
|
||||
F(\mathbf{x}_k)$. This means that for a sufficiently small $\gamma_k$
|
||||
we are always moving towards smaller function values, i.e a minimum.
|
||||
|
||||
@@ -54,7 +55,7 @@ the learning rate within the context of Machine Learning.
|
||||
!split
|
||||
===== The ideal =====
|
||||
|
||||
Ideally the sequence $\{ \mathbf{x}_k \}_{k=0}$ converges to a global
|
||||
Ideally the sequence $\{\mathbf{x}_k \}_{k=0}$ converges to a global
|
||||
minimum of the function $F$. In general we do not know if we are in a
|
||||
global or local minimum. In the special case when $F$ is a convex
|
||||
function, all local minima are also global minima, so in this case
|
||||
@@ -76,7 +77,8 @@ Note that the gradient is a function of $\mathbf{x} =
|
||||
!split
|
||||
===== The sensitiveness of the gradient descent =====
|
||||
|
||||
GD is sensitive to the choice of learning rate $\gamma_k$. This is due
|
||||
The gradient descent method
|
||||
is sensitive to the choice of learning rate $\gamma_k$. This is due
|
||||
to the fact that we are only guaranteed that $F(\mathbf{x}_{k+1}) \leq
|
||||
F(\mathbf{x}_k)$ for sufficiently small $\gamma_k$. The problem is to
|
||||
determine an optimal learning rate. If the learning rate is chosen too
|
||||
@@ -87,10 +89,217 @@ Many of these shortcomings can be alleviated by introducing
|
||||
randomness. One such method is that of Stochastic Gradient Descent
|
||||
(SGD), see below.
|
||||
|
||||
|
||||
!split
|
||||
===== Convex functions =====
|
||||
|
||||
Ideally we want our cost/loss function to be convex(concave).
|
||||
|
||||
First we give the definition of a convex set: A set $C$ in
|
||||
$\mathbb{R}^n$ is said to be convex if, for all $x$ and $y$ in $C$ and
|
||||
all $t \in (0,1)$ , the point $(1 − t)x + ty$ also belongs to
|
||||
C. Geometrically this means that every point on the line segment
|
||||
connecting $x$ and $y$ is in $C$ as discussed below.
|
||||
|
||||
The convex subsets of $\mathbb{R}$ are the intervals of
|
||||
$\mathbb{R}$. Examples of convex sets of $\mathbb{R}^2$ are the
|
||||
regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
|
||||
!split
|
||||
===== Convex function =====
|
||||
|
||||
_Convex function_: Let $X \subset \mathbb{R}^n$ be a convex set. Assume that the function $f: X \rightarrow \mathbb{R}$ is continuous, then $f$ is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all $x_1, x_2 \in X$ and for all $t \in [0,1]$. If $\leq$ is replaced with a strict inequaltiy in the definition, we demand $x_1 \neq x_2$ and $t\in(0,1)$ then $f$ is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting $f(x_1)$ and $f(x_2)$, the value of the function on the interval $[x_1,x_2]$ is always below the line as illustrated below.
|
||||
|
||||
!split
|
||||
===== Conditions on convex functions =====
|
||||
|
||||
In the following we state first and second-order conditions which
|
||||
ensures convexity of a function $f$. We write $D_f$ to denote the
|
||||
domain of $f$, i.e the subset of $R^n$ where $f$ is defined. For more
|
||||
details and proofs we refer to: "S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press":"http://stanford.edu/boyd/cvxbook/, 2004".
|
||||
|
||||
!bblock First order condition
|
||||
Suppose $f$ is differentiable (i.e $\nabla f(x)$ is well defined for
|
||||
all $x$ in the domain of $f$). Then $f$ is convex if and only if $D_f$
|
||||
is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
|
||||
for all $x,y \in D_f$. This condition means that for a convex function
|
||||
the first order Taylor expansion (right hand side above) at any point
|
||||
a global under estimator of the function. To convince yourself you can
|
||||
make a drawing of $f(x) = x^2+1$ and draw the tangent line to $f(x)$ and
|
||||
note that it is always below the graph.
|
||||
!eblock
|
||||
|
||||
!bblock Second order condition
|
||||
Assume that $f$ is twice
|
||||
differentiable, i.e the Hessian matrix exists at each point in
|
||||
$D_f$. Then $f$ is convex if and only if $D_f$ is a convex set and its
|
||||
Hessian is positive semi-definite for all $x\in D_f$. For a
|
||||
single-variable function this reduces to $f''(x) \geq 0$. Geometrically this means that $f$ has nonnegative curvature
|
||||
everywhere.
|
||||
!eblock
|
||||
|
||||
This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
|
||||
|
||||
!split
|
||||
===== More on convex functions =====
|
||||
|
||||
The next result is of great importance to us and the reason why we are
|
||||
going on about convex functions. In machine learning we frequently
|
||||
have to minimize a loss/cost function in order to find the best
|
||||
parameters for the model we are considering.
|
||||
|
||||
Ideally we want the
|
||||
global minimum (for high-dimensional models it is hard to know
|
||||
if we have local or global minimum). However, if the cost/loss function
|
||||
is convex the following result provides invaluable information:
|
||||
|
||||
!bblock Any minimum is global for convex functions
|
||||
Consider the problem of finding $x \in \mathbb{R}^n$ such that $f(x)$
|
||||
is minimal, where $f$ is convex and differentiable. Then, any point
|
||||
$x^*$ that satisfies $\nabla f(x^*) = 0$ is a global minimum.
|
||||
!eblock
|
||||
|
||||
This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
|
||||
|
||||
!split
|
||||
===== Some simple problems =====
|
||||
|
||||
o Show that $f(x)=x^2$ is convex for $x \in \mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \in D_f$ and any $\lambda \in [0,1]$ $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.
|
||||
|
||||
o Using the second order condition show that the following functions are convex on the specified domain.
|
||||
* $f(x) = e^x$ is convex for $x \in \mathbb{R}$.
|
||||
* $g(x) = -\ln(x)$ is convex for $x \in (0,\infty)$.
|
||||
o Let $f(x) = x^2$ and $g(x) = e^x$. Show that $f(g(x))$ and $g(f(x))$ is convex for $x \in \mathbb{R}$. Also show that if $f(x)$ is any convex function than $h(x) = e^{f(x)}$ is convex.
|
||||
|
||||
o A norm is any function that satisfy the following properties
|
||||
* $f(\alpha x) = |\alpha| f(x)$ for all $\alpha \in \mathbb{R}$.
|
||||
* $f(x+y) \leq f(x) + f(y)$
|
||||
* $f(x) \leq 0$ for all $x \in \mathbb{R}^n$ with equality if and only if $x = 0$
|
||||
|
||||
Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
|
||||
|
||||
!split
|
||||
===== Revisiting our first homework =====
|
||||
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
o An analytical solution (recall homework set 1).
|
||||
o The gradient can be computed analytically.
|
||||
o The cost function is convex which guarantees that gradient descent converges for small enough learning rates
|
||||
|
||||
We revisit the example from homework set 1 where we had
|
||||
!bt
|
||||
\[
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
\]
|
||||
!et
|
||||
with $x_i \in [0,1] $ chosen randomly with a uniform distribution. Additionally $\xi_i$ represents stochastic noise chosen according to a normal distribution $\cal {N}(0,1)$.
|
||||
The linear regression model is given by
|
||||
!bt
|
||||
\[
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
\]
|
||||
!et
|
||||
such that
|
||||
!bt
|
||||
\[
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
\]
|
||||
!et
|
||||
|
||||
!split
|
||||
===== Gradient descent example =====
|
||||
|
||||
Let $\mathbf{y} = (y_1,\cdots,y_n)^T$, $\mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T$ and $\beta = (\beta_0, \beta_1)^T$
|
||||
|
||||
It is convenient to write $\mathbf{\hat{y}} = X\beta$ where $X \in \mathbb{R}^{100 \times 2} $ is the design matrix given by
|
||||
!bt
|
||||
\[
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
\]
|
||||
!et
|
||||
The loss function is given by
|
||||
!bt
|
||||
\[
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
\]
|
||||
!et
|
||||
and we want to find $\beta$ such that $C(\beta)$ is minimized.
|
||||
|
||||
!split
|
||||
===== The derivative of the cost/loss function =====
|
||||
|
||||
Computing $\partial C(\beta) / \partial \beta_0$ and $\partial C(\beta) / \partial \beta_1$ we can show that the gradient can be written as
|
||||
!bt
|
||||
\[
|
||||
\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
\]
|
||||
!et
|
||||
where $X$ is the design matrix defined above.
|
||||
|
||||
!split
|
||||
===== The Hessian matrix =====
|
||||
The Hessian matrix of $C(\beta)$ is given by
|
||||
!bt
|
||||
\[
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
\]
|
||||
!et
|
||||
This result implies that $C(\beta)$ is a convex function since the matrix $X^T X$ always is positive semi-definite.
|
||||
|
||||
!split
|
||||
===== Simple program =====
|
||||
|
||||
We can now write a program that minimizes $C(\beta)$ using the gradient descent method with a constant learning rate $\gamma$ according to
|
||||
!bt
|
||||
\[
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
\]
|
||||
!et
|
||||
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
$\beta_0$ be chosen randomly and let $\gamma = 0.001$. Stop iterating
|
||||
when $||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8}$.
|
||||
|
||||
And finally we can compare our solution for $\beta$ with the analytic result given by
|
||||
$\beta= (X^TX)^{-1} X^T \mathbf{y}$.
|
||||
!bc pycod
|
||||
import numpy as np
|
||||
|
||||
"""
|
||||
The following setup is just a suggestion, feel free to write it the way you like.
|
||||
"""
|
||||
|
||||
#Setup problem described in the exercise
|
||||
N = 100 #Nr of datapoints
|
||||
M = 2 #Nr of features
|
||||
x = np.random.rand(N) #Uniformly generated x-values in [0,1]
|
||||
y = 5*x**2 + 0.1*np.random.randn(N)
|
||||
X = np.c_[np.ones(N),x] #Construct design matrix
|
||||
|
||||
#Compute beta according to normal equations to compare with GD solution
|
||||
Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
|
||||
Xt_y = np.dot(X.transpose(),y)
|
||||
beta_NE = np.dot(Xt_X_inv,Xt_y)
|
||||
print(beta_NE)
|
||||
!ec
|
||||
|
||||
!split
|
||||
===== Gradient Descent Example =====
|
||||
|
||||
We revisit now our simple linear regression example with a linear polynomial.
|
||||
Another simple example is here
|
||||
!bc pycod
|
||||
|
||||
# Importing various packages
|
||||
@@ -156,214 +365,7 @@ print(sgdreg.intercept_, sgdreg.coef_)
|
||||
|
||||
!ec
|
||||
|
||||
!split
|
||||
===== Convex functions =====
|
||||
|
||||
Ideally we want our cost/loss function to be convex(concave).
|
||||
|
||||
First we give the definition of a convex set: A set $C$ in
|
||||
$\mathbb{R}^n$ is said to be convex if, for all $x$ and $y$ in $C$ and
|
||||
all $t \in (0,1)$ , the point $(1 − t)x + ty$ also belongs to
|
||||
C. Geometrically this means that every point on the line segment
|
||||
connecting $x$ and $y$ is in $C$ as discussed below.
|
||||
|
||||
The convex subsets of $\mathbb{R}$ are the intervals of
|
||||
$\mathbb{R}$. Examples of convex sets of $\mathbb{R}^2$ are the
|
||||
regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
|
||||
!split
|
||||
===== Convex function =====
|
||||
|
||||
_Convex function_: Let $X \subset \mathbb{R}^n$ be a convex set. Assume that the function $f: X \rightarrow \mathbb{R}$ is continuous, then $f$ is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all $x_1, x_2 \in X$ and for all $t \in [0,1]$. If $\leq$ is replaced with a strict inequaltiy in the definition, we demand $x_1 \neq x_2$ and $t\in(0,1)$ then $f$ is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting $f(x_1)$ and $f(x_2)$, the value of the function on the interval $[x_1,x_2]$ is always below the line as illustrated below.
|
||||
|
||||
!split
|
||||
===== Conditions on convex functions =====
|
||||
|
||||
In the following we state first and second-order conditions which
|
||||
ensures convexity of a function $f$. We write $D_f$ to denote the
|
||||
domain of $f$, i.e the subset of $R^n$ where $f$ is defined. For more
|
||||
details and proofs we refer to: "S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press":"http://stanford.edu/boyd/cvxbook/, 2004".
|
||||
|
||||
!bblock First order condition
|
||||
Suppose $f$ is differentiable (i.e $\nabla f(x)$ is well defined for
|
||||
all $x$ in the domain of $f$). Then $f$ is convex if and only if $D_f$
|
||||
is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
|
||||
for all $x,y \in D_f$. This condition means that for a convex function
|
||||
the first order Taylor expansion (right hand side above) at any point
|
||||
a global under estimator of the function. To convince yourself you can
|
||||
make a drawing of f(x) = x^2+1 and draw the tangent line to $f(x)$ and
|
||||
note that it is always below the graph.
|
||||
!eblock
|
||||
|
||||
!bblock Second order condition
|
||||
Assume that $f$ is twice
|
||||
differentiable, i.e the Hessian matrix exists at each point in
|
||||
$D_f$. Then $f$ is convex if and only if $D_f$ is a convex set and its
|
||||
Hessian is positive semi-definite for all $x\in D_f$. For a
|
||||
single-variable function this reduces to $f''(x) \geq
|
||||
0$. Geometrically this means that $f$ has nonnegative curvature
|
||||
everywhere.
|
||||
!eblock
|
||||
|
||||
This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
|
||||
|
||||
!split
|
||||
===== More on convex functions =====
|
||||
|
||||
The next result is of great importance to us and the reason why we are
|
||||
going on about convex functions. In machine learning we frequently
|
||||
have to minimize a loss/cost function in order to find the best
|
||||
parameters for the model we are considering.
|
||||
|
||||
Ideally we want the
|
||||
global minimum (for high-dimensional models it is hard to know
|
||||
if we have local or global minimum). However, if the cost/loss function
|
||||
is convex the following result provides invaluable information:
|
||||
|
||||
!bblock Any minimum is global for convex functions
|
||||
Consider the problem of finding $x \in \mathbb{R}^n$ such that $f(x)$
|
||||
is minimal, where $f$ is convex and differentiable. Then, any point
|
||||
$x^*$ that satisfies $\nabla f(x^*) = 0$ is a global minimum.
|
||||
!eblock
|
||||
|
||||
This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
|
||||
|
||||
!split
|
||||
===== Some simple problems =====
|
||||
|
||||
o Show that $f(x)=x^2$ is convex for $x \in \mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \in D_f$ and any $\lambda \in [0,1] $ $\lambda f(x) + (1-\lambda)f(y) - f(\lambda x + (1-\lambda) y ) \geq 0. $
|
||||
|
||||
o Using the second order condition show that the following functions are convex on the specified domain.
|
||||
* $f(x) = e^x$ is convex for $x \in \mathbb{R}$.
|
||||
* $g(x) = -\ln(x)$ is convex for $x \in (0,\infty)$.
|
||||
o Let $f(x) = x^2$ and $g(x) = e^x$. Show that $f(g(x))$ and $g(f(x))$ is convex for $x \in \mathbb{R}$. Also show that if $f(x)$ is any convex function than $h(x) = e^{f(x)}$ is convex.
|
||||
|
||||
o A norm is any function that satisfy the following properties
|
||||
* $f(\alpha x) = |\alpha| f(x)$ for all $\alpha \in \mathbb{R}$.
|
||||
* $f(x+y) \leq f(x) + f(y)$
|
||||
* $f(x) \leq 0$ for all $x \in \mathbb{R}^n$ with equality if and only if $x = 0$
|
||||
|
||||
Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
|
||||
|
||||
!split
|
||||
===== Revisiting our first homework =====
|
||||
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
o An analytical solution (recall homework set 1).
|
||||
o The gradient can be computed analytically.
|
||||
o The cost function is convex which guarantees that gradient descent converges for small enough learning rates
|
||||
|
||||
We revisit the example from homework set 1 where we had
|
||||
!bt
|
||||
\[
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
\]
|
||||
!et
|
||||
with $x_i \in [0,1] $ chosen randomly with a uniform distribution. Additionally $\xi_i$ represents stochastic noise chosen according to a normal distribution $\cal {N}(0,1)$.
|
||||
The linear regression model is given by
|
||||
!bt
|
||||
\[
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
\]
|
||||
!et
|
||||
such that
|
||||
!bt
|
||||
\[
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
\]
|
||||
!et
|
||||
|
||||
!split
|
||||
===== Gradient descent example =====
|
||||
|
||||
Let $\mathbf{y} = (y_1,\cdots,y_n)^T$, $\mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T$ and $\beta = (\beta_0, \beta_1)^T$
|
||||
|
||||
t is convenient to write $\mathbf{\hat{y}} = X\beta$ where $X \in \mathbb{R}^{100 \times 2} $ is the design matrix given by
|
||||
!bt
|
||||
\[
|
||||
\begin{equation}
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
\end{equation}
|
||||
\]
|
||||
!et
|
||||
The loss function is given by
|
||||
!bt
|
||||
\[
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
\]
|
||||
!et
|
||||
and we want to find $\beta$ such that $C(\beta)$ is minimized.
|
||||
|
||||
!split
|
||||
===== The derivative of the cost/loss function =====
|
||||
|
||||
Computing $\partial C(\beta) / \partial \beta_0$ and $\partial C(\beta) / \partial \beta_1$ we can show that the gradient can be written as
|
||||
!bt
|
||||
\[
|
||||
\nabla_\beta C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
\]
|
||||
!et
|
||||
where $X$ is the design matrix defined above.
|
||||
|
||||
!split
|
||||
===== The Hessian matrix =====
|
||||
The Hessian matrix of $C(\beta)$ is given by
|
||||
!bt
|
||||
\[
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
\]
|
||||
!et
|
||||
This result implies that $C(\beta)$ is a convex function since the matrix $X^T X$ always is positive semi-definite.
|
||||
|
||||
!split
|
||||
===== Simple program =====
|
||||
|
||||
We can now write a program that minimizes $C(\beta)$ using the gradient descent method with a constant learning rate $\gamma$ according to
|
||||
!bt
|
||||
\[
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
\]
|
||||
!et
|
||||
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
$\beta_0$ be chosen randomly and let $\gamma = 0.001$. Stop iterating
|
||||
when $||\nabla_\beta C(\beta_k) || < \epsilon = 10^{-8}$.
|
||||
|
||||
And finally we can compare our solution for $\beta$ with the analytic result given by
|
||||
$\beta= (X^TX)^{-1} X^T \mathbf{y}$.
|
||||
!bc pycod
|
||||
import numpy as np
|
||||
|
||||
"""
|
||||
The following setup is just a suggestion, feel free to write it the way you like.
|
||||
"""
|
||||
|
||||
#Setup problem described in the exercise
|
||||
N = 100 #Nr of datapoints
|
||||
M = 2 #Nr of features
|
||||
x = np.random.rand(N) #Uniformly generated x-values in [0,1]
|
||||
y = 5*x**2 + 0.1*np.random.randn(N)
|
||||
X = np.c_[np.ones(N),x] #Construct design matrix
|
||||
|
||||
#Compute beta according to normal equations to compare with GD solution
|
||||
Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
|
||||
Xt_y = np.dot(X.transpose(),y)
|
||||
beta_NE = np.dot(Xt_X_inv,Xt_y)
|
||||
print(beta_NE)
|
||||
!ec
|
||||
|
||||
!split
|
||||
===== Gradient descent and Ridge =====
|
||||
@@ -432,7 +434,7 @@ the shortcomings of the Gradient descent method discussed above.
|
||||
|
||||
The underlying idea of SGD comes from the observation that the cost
|
||||
function, which we want to minimize, can almost always be written as a
|
||||
sum over $n$ datapoints $\{\mathbf{x}_i\}_{i=1}^n$,
|
||||
sum over $n$ data points $\{\mathbf{x}_i\}_{i=1}^n$,
|
||||
!bt
|
||||
\[
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i,
|
||||
@@ -454,27 +456,27 @@ computed as a sum over $i$-gradients
|
||||
|
||||
Stochasticity/randomness is introduced by only taking the
|
||||
gradient on a subset of the data called minibatches. If there are $n$
|
||||
datapoints and the size of each minibatch is $M$, there will be $n/M$
|
||||
data points and the size of each minibatch is $M$, there will be $n/M$
|
||||
minibatches. We denote these minibatches by $B_k$ where
|
||||
$k=1,\cdots,n/M$.
|
||||
|
||||
!split
|
||||
===== SGD example =====
|
||||
As an example, suppose we have $10$ datapoints $( \mathbf{x}_1,
|
||||
\cdots, \mathbf{x}_{10} )$ and we choose to have $M=5$ minibathces,
|
||||
then each minibatch contains two datapoints. In particular we have
|
||||
As an example, suppose we have $10$ data points $(\mathbf{x}_1,\cdots, \mathbf{x}_{10})$
|
||||
and we choose to have $M=5$ minibathces,
|
||||
then each minibatch contains two data points. In particular we have
|
||||
$B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 =
|
||||
(\mathbf{x}_9,\mathbf{x}_{10})$. Note that if you choose $M=1$ you
|
||||
have only a single batch with all datapoints and on the other extreme,
|
||||
have only a single batch with all data points and on the other extreme,
|
||||
you may choose $M=n$ resulting in a minibatch for each datapoint, i.e
|
||||
$B_k = \mathbf{x}_k$.
|
||||
|
||||
The idea is now to approximate the gradient by replacing the sum over
|
||||
all datapoints with a sum over the datapoints in one the minibatches
|
||||
all data points with a sum over the data points in one the minibatches
|
||||
picked at random in each gradient descent step
|
||||
!bt
|
||||
\[
|
||||
\nabla_\beta
|
||||
\nabla_{\beta}
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta
|
||||
c_i(\mathbf{x}_i, \mathbf{\beta}).
|
||||
@@ -522,8 +524,8 @@ Taking the gradient only on a subset of the data has two important
|
||||
benefits. First, it introduces randomness which decreases the chance
|
||||
that our opmization scheme gets stuck in a local minima. Second, if
|
||||
the size of the minibatches are small relative to the number of
|
||||
datapoints ($M < n$), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the k-th minibatch and not
|
||||
datapoints ($M < n$), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the $k-th$ minibatch and not
|
||||
all $n$ datapoints.
|
||||
|
||||
!split
|
||||
@@ -547,7 +549,7 @@ Another approach is to let the step length $\gamma_j$ depend on the
|
||||
number of epochs in such a way that it becomes very small after a
|
||||
reasonable time such that we do not move at all.
|
||||
|
||||
As an example, let $e = 0,1,2,3,\cdots$ denote the current epoch and let $t_0, t_1 > 0$ be two fixed numbers. Furthermore, let $t = e \cdot m + i$ where $m$ is the number of minibatches and $i=0,\cdots,m-1$. Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length $\gamma_j (0; t_0, t_1) = t_0/t_1$ which decays in *time* $t$.
|
||||
As an example, let $e = 0,1,2,3,\cdots$ denote the current epoch and let $t_0, t_1 > 0$ be two fixed numbers. Furthermore, let $t = e \cdot m + i$ where $m$ is the number of minibatches and $i=0,\cdots,m-1$. Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length $\gamma_j (0; t_0, t_1) = t_0/t_1$ which decays in *time* $t$.
|
||||
|
||||
In this way we can fix the number of epochs, compute $\beta$ and
|
||||
evaluate the cost function at the end. Repeating the computation will
|
||||
|
||||
Reference in New Issue
Block a user