more on steepest descent

This commit is contained in:
mhjensen
2018-09-27 04:47:50 +02:00
parent 20e8b8fbec
commit a55882e5a4
45 changed files with 2914 additions and 4593 deletions
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -249,7 +234,7 @@ MathJax.Hub.Config({
<li><a href="._Splines-bs008.html">9</a></li>
<li><a href="._Splines-bs009.html">10</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs001.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -235,7 +220,7 @@ some approximative/numerical method to compute the minimum.
<li><a href="._Splines-bs009.html">10</a></li>
<li><a href="._Splines-bs010.html">11</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs002.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -243,7 +228,7 @@ where \( \hat{\beta} \) are the weights we wish to extract from data, in our cas
<li><a href="._Splines-bs010.html">11</a></li>
<li><a href="._Splines-bs011.html">12</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs003.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -248,7 +233,7 @@ This defines what we call the Hessian.
<li><a href="._Splines-bs011.html">12</a></li>
<li><a href="._Splines-bs012.html">13</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs004.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -249,7 +234,7 @@ If we can compute these matrices, in particular the Hessian, the above is often
<li><a href="._Splines-bs012.html">13</a></li>
<li><a href="._Splines-bs013.html">14</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs005.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -243,7 +228,7 @@ discourage the use of this method.
<li><a href="._Splines-bs013.html">14</a></li>
<li><a href="._Splines-bs014.html">15</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs006.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -261,7 +246,7 @@ $$
<li><a href="._Splines-bs014.html">15</a></li>
<li><a href="._Splines-bs015.html">16</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs007.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -244,7 +229,7 @@ vanishes, then Newton-Raphson may fail totally
<li><a href="._Splines-bs015.html">16</a></li>
<li><a href="._Splines-bs016.html">17</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs008.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -282,7 +267,7 @@ more than two non-linear equations. In our case, the Jacobian matrix is given by
<li><a href="._Splines-bs016.html">17</a></li>
<li><a href="._Splines-bs017.html">18</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs009.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -252,7 +237,7 @@ we are always moving towards smaller function values, i.e a minimum.
<li><a href="._Splines-bs017.html">18</a></li>
<li><a href="._Splines-bs018.html">19</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs010.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -248,7 +233,7 @@ the learning rate within the context of Machine Learning.
<li><a href="._Splines-bs018.html">19</a></li>
<li><a href="._Splines-bs019.html">20</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs011.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -255,7 +240,7 @@ Note that the gradient is a function of \( \mathbf{x} =
<li><a href="._Splines-bs019.html">20</a></li>
<li><a href="._Splines-bs020.html">21</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs012.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -248,7 +233,7 @@ randomness. One such method is that of Stochastic Gradient Descent
<li><a href="._Splines-bs020.html">21</a></li>
<li><a href="._Splines-bs021.html">22</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs013.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -249,7 +234,7 @@ regular polygons (triangles, rectangles, pentagons, etc...).
<li><a href="._Splines-bs021.html">22</a></li>
<li><a href="._Splines-bs022.html">23</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs014.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -237,7 +222,7 @@ MathJax.Hub.Config({
<li><a href="._Splines-bs022.html">23</a></li>
<li><a href="._Splines-bs023.html">24</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs015.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -273,7 +258,7 @@ This condition is particularly useful since it gives us an procedure for determi
<li><a href="._Splines-bs023.html">24</a></li>
<li><a href="._Splines-bs024.html">25</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs016.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -260,7 +245,7 @@ This result means that if we know that the cost/loss function is convex and we a
<li><a href="._Splines-bs024.html">25</a></li>
<li><a href="._Splines-bs025.html">26</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs017.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -256,7 +241,7 @@ Using the definition of convexity, try to show that a function satisfying the pr
<li><a href="._Splines-bs025.html">26</a></li>
<li><a href="._Splines-bs026.html">27</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs018.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -263,7 +248,7 @@ When we have found the exact solution, \( \hat{r}=0 \).
<li><a href="._Splines-bs026.html">27</a></li>
<li><a href="._Splines-bs027.html">28</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs019.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+48 -66
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -206,7 +191,7 @@ MathJax.Hub.Config({
<a name="part0019"></a>
<!-- !split -->
<h2 id="___sec18" class="anchor">Conjugate gradient method </h2>
<h2 id="___sec18" class="anchor">Gradient method </h2>
<p>
The residual is zero when we reach the minimum of the quadratic equation
@@ -223,9 +208,6 @@ variance, then the matrix \( \hat{A} \), which is called the Hessian, is
given by the second-derivative of the function we want to minimize.
This quantity is always positive definite.
<p>
More details will be added here soon.
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
@@ -252,7 +234,7 @@ More details will be added here soon.
<li><a href="._Splines-bs027.html">28</a></li>
<li><a href="._Splines-bs028.html">29</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs020.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+63 -100
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -206,47 +191,25 @@ MathJax.Hub.Config({
<a name="part0020"></a>
<!-- !split -->
<h2 id="___sec19" class="anchor">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
<div class="panel panel-default">
<div class="panel-body">
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
<h2 id="___sec19" class="anchor">Steepest descent method </h2>
<p>
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
We can assume without loss of generality that
$$
\begin{equation*}
\hat{x}_0=0,
\end{equation*}
$$
<!-- code=c++ (!bc cppcod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic">&lt;cmath&gt;</span><span style="color: #BC7A00"></span>
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic">&lt;iostream&gt;</span><span style="color: #BC7A00"></span>
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic">&lt;fstream&gt;</span><span style="color: #BC7A00"></span>
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic">&lt;iomanip&gt;</span><span style="color: #BC7A00"></span>
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic">&quot;vectormatrixclass.h&quot;</span><span style="color: #BC7A00"></span>
<span style="color: #008000; font-weight: bold">using</span> <span style="color: #008000; font-weight: bold">namespace</span> std;
<span style="color: #408080; font-style: italic">// Main function begins here</span>
<span style="color: #B00040">int</span> <span style="color: #0000FF">main</span>(<span style="color: #B00040">int</span> argc, <span style="color: #B00040">char</span> <span style="color: #666666">*</span> argv[]){
<span style="color: #B00040">int</span> dim <span style="color: #666666">=</span> <span style="color: #666666">2</span>;
Vector x(dim),xsd(dim), b(dim),x0(dim);
Matrix A(dim,dim);
<span style="color: #408080; font-style: italic">// Set our initial guess</span>
x0(<span style="color: #666666">0</span>) <span style="color: #666666">=</span> x0(<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">0</span>;
<span style="color: #408080; font-style: italic">// Set the matrix</span>
A(<span style="color: #666666">0</span>,<span style="color: #666666">0</span>) <span style="color: #666666">=</span> <span style="color: #666666">3</span>; A(<span style="color: #666666">1</span>,<span style="color: #666666">0</span>) <span style="color: #666666">=</span> <span style="color: #666666">2</span>; A(<span style="color: #666666">0</span>,<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">2</span>; A(<span style="color: #666666">1</span>,<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">6</span>;
b(<span style="color: #666666">0</span>) <span style="color: #666666">=</span> <span style="color: #666666">2</span>; b(<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">-8</span>;
cout <span style="color: #666666">&lt;&lt;</span> <span style="color: #BA2121">&quot;The Matrix A that we are using: &quot;</span> <span style="color: #666666">&lt;&lt;</span> endl;
A.Print();
cout <span style="color: #666666">&lt;&lt;</span> endl;
x <span style="color: #666666">=</span> ConjugateGradient(A,b,x0);
xsd <span style="color: #666666">=</span> SteepestDescent(A,b,x0);
cout <span style="color: #666666">&lt;&lt;</span> <span style="color: #BA2121">&quot;The approximate solution using Conjugate Gradient is: &quot;</span> <span style="color: #666666">&lt;&lt;</span> endl;
x.Print();
cout <span style="color: #666666">&lt;&lt;</span> endl;
cout <span style="color: #666666">&lt;&lt;</span> <span style="color: #BA2121">&quot;The approximate solution using Steepest Descent is: &quot;</span> <span style="color: #666666">&lt;&lt;</span> endl;
xsd.Print();
cout <span style="color: #666666">&lt;&lt;</span> endl;
}
</pre></div>
<p>
</div>
</div>
or consider the system
$$
\begin{equation*}
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
\end{equation*}
$$
instead.
<p>
<p>
@@ -274,7 +237,7 @@ MathJax.Hub.Config({
<li><a href="._Splines-bs028.html">29</a></li>
<li><a href="._Splines-bs029.html">30</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs021.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+66 -87
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -206,35 +191,29 @@ MathJax.Hub.Config({
<a name="part0021"></a>
<!-- !split -->
<h2 id="___sec20" class="anchor">The routine for the steepest descent method </h2>
<h2 id="___sec20" class="anchor">Steepest descent method </h2>
<div class="panel panel-default">
<div class="panel-body">
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
<p>
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
$$
\begin{equation*}
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
\end{equation*}
$$
This suggests taking the first basis vector \( \hat{p}_1 \)
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
which equals
$$
\begin{equation*}
\hat{A}\hat{x}_0-\hat{b},
\end{equation*}
$$
and
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
<!-- code=c++ (!bc cppcod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span>Vector <span style="color: #0000FF">SteepestDescent</span>(Matrix A, Vector b, Vector x0){
<span style="color: #B00040">int</span> IterMax, i;
<span style="color: #B00040">int</span> dim <span style="color: #666666">=</span> x0.Dimension();
<span style="color: #008000; font-weight: bold">const</span> <span style="color: #B00040">double</span> tolerance <span style="color: #666666">=</span> <span style="color: #666666">1.0e-14</span>;
Vector x(dim),f(dim),z(dim);
<span style="color: #B00040">double</span> c,alpha,d;
IterMax <span style="color: #666666">=</span> <span style="color: #666666">30</span>;
x <span style="color: #666666">=</span> x0;
f <span style="color: #666666">=</span> A<span style="color: #666666">*</span>x<span style="color: #666666">-</span>b;
i <span style="color: #666666">=</span> <span style="color: #666666">0</span>;
<span style="color: #008000; font-weight: bold">while</span> (i <span style="color: #666666">&lt;=</span> IterMax){
z <span style="color: #666666">=</span> A<span style="color: #666666">*</span>f;
c <span style="color: #666666">=</span> dot(f,f);
alpha <span style="color: #666666">=</span> c<span style="color: #666666">/</span>dot(f,z);
x <span style="color: #666666">=</span> x <span style="color: #666666">-</span> alpha<span style="color: #666666">*</span>f;
f <span style="color: #666666">=</span> A<span style="color: #666666">*</span>x<span style="color: #666666">-</span>b;
<span style="color: #008000; font-weight: bold">if</span>(sqrt(dot(f,f)) <span style="color: #666666">&lt;</span> tolerance) <span style="color: #008000; font-weight: bold">break</span>;
i<span style="color: #666666">++</span>;
}
<span style="color: #008000; font-weight: bold">return</span> x;
}
</pre></div>
<p>
</div>
</div>
@@ -266,7 +245,7 @@ MathJax.Hub.Config({
<li><a href="._Splines-bs029.html">30</a></li>
<li><a href="._Splines-bs030.html">31</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs022.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+65 -86
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -204,37 +189,31 @@ MathJax.Hub.Config({
<p>&nbsp;</p><p>&nbsp;</p><p>&nbsp;</p> <!-- add vertical space -->
<a name="part0022"></a>
<!-- !split -->
<!-- !split -->
<h2 id="___sec21" class="anchor">Revisiting our first homework </h2>
<p>
We will use linear regression as a case study for the gradient descent
methods. Linear regression is a great test case for the gradient
descent methods discussed in the lectures since it has several
desirable properties such as:
<ol>
<li> An analytical solution (recall homework set 1).</li>
<li> The gradient can be computed analytically.</li>
<li> The cost function is convex which guarantees that gradient descent converges for small enough learning rates</li>
</ol>
We revisit the example from homework set 1 where we had
<h2 id="___sec21" class="anchor">Gradient descent method </h2>
<div class="panel panel-default">
<div class="panel-body">
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
$$
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
\begin{equation*}
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
\end{equation*}
$$
with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
The linear regression model is given by
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
\( \hat{x}=\hat{x}_k \),
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
This gives the following expression
$$
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
\begin{equation*}
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
\end{equation*}
$$
</div>
</div>
such that
$$
\hat{y}_i = \beta_0 + \beta_1 x_i.
$$
<p>
<p>
@@ -262,7 +241,7 @@ $$
<li><a href="._Splines-bs030.html">31</a></li>
<li><a href="._Splines-bs031.html">32</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs023.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+77 -78
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -204,29 +189,43 @@ MathJax.Hub.Config({
<p>&nbsp;</p><p>&nbsp;</p><p>&nbsp;</p> <!-- add vertical space -->
<a name="part0023"></a>
<!-- !split -->
<!-- !split -->
<h2 id="___sec22" class="anchor">Gradient descent example </h2>
<p>
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
<p>
It is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by
<h2 id="___sec22" class="anchor">Final expressions </h2>
<div class="panel panel-default">
<div class="panel-body">
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
We can also compute the residual iteratively as
$$
X \equiv \begin{bmatrix}
1 &amp; x_1 \\
\vdots &amp; \vdots \\
1 &amp; x_{100} &amp; \\
\end{bmatrix}.
\begin{equation*}
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
\end{equation*}
$$
The loss function is given by
which equals
$$
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
\begin{equation*}
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
\end{equation*}
$$
and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
or
$$
\begin{equation*}
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
\end{equation*}
$$
which gives
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
\end{equation*}
$$
</div>
</div>
<p>
<p>
@@ -254,7 +253,7 @@ and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
<li><a href="._Splines-bs031.html">32</a></li>
<li><a href="._Splines-bs032.html">33</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs024.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+48 -73
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -206,17 +191,7 @@ MathJax.Hub.Config({
<a name="part0024"></a>
<!-- !split -->
<h2 id="___sec23" class="anchor">The derivative of the cost/loss function </h2>
<p>
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
$$
\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
$$
where \( X \) is the design matrix defined above.
<h2 id="___sec23" class="anchor">The Steepest descent algorithm </h2>
<p>
<p>
@@ -244,7 +219,7 @@ where \( X \) is the design matrix defined above.
<li><a href="._Splines-bs032.html">33</a></li>
<li><a href="._Splines-bs033.html">34</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs025.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+83 -71
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -206,16 +191,43 @@ MathJax.Hub.Config({
<a name="part0025"></a>
<!-- !split -->
<h2 id="___sec24" class="anchor">The Hessian matrix </h2>
The Hessian matrix of \( C(\beta) \) is given by
$$
\hat{H} \equiv \begin{bmatrix}
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} &amp; \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} &amp; \frac{\partial^2 C(\beta)}{\partial \beta_1^2} &amp; \\
\end{bmatrix} = 2X^T X.
$$
<h2 id="___sec24" class="anchor">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
<div class="panel panel-default">
<div class="panel-body">
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
<p>
<!-- code=c++ (!bc cppcod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic">&lt;cmath&gt;</span><span style="color: #BC7A00"></span>
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic">&lt;iostream&gt;</span><span style="color: #BC7A00"></span>
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic">&lt;fstream&gt;</span><span style="color: #BC7A00"></span>
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic">&lt;iomanip&gt;</span><span style="color: #BC7A00"></span>
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic">&quot;vectormatrixclass.h&quot;</span><span style="color: #BC7A00"></span>
<span style="color: #008000; font-weight: bold">using</span> <span style="color: #008000; font-weight: bold">namespace</span> std;
<span style="color: #408080; font-style: italic">// Main function begins here</span>
<span style="color: #B00040">int</span> <span style="color: #0000FF">main</span>(<span style="color: #B00040">int</span> argc, <span style="color: #B00040">char</span> <span style="color: #666666">*</span> argv[]){
<span style="color: #B00040">int</span> dim <span style="color: #666666">=</span> <span style="color: #666666">2</span>;
Vector x(dim),xsd(dim), b(dim),x0(dim);
Matrix A(dim,dim);
<span style="color: #408080; font-style: italic">// Set our initial guess</span>
x0(<span style="color: #666666">0</span>) <span style="color: #666666">=</span> x0(<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">0</span>;
<span style="color: #408080; font-style: italic">// Set the matrix</span>
A(<span style="color: #666666">0</span>,<span style="color: #666666">0</span>) <span style="color: #666666">=</span> <span style="color: #666666">3</span>; A(<span style="color: #666666">1</span>,<span style="color: #666666">0</span>) <span style="color: #666666">=</span> <span style="color: #666666">2</span>; A(<span style="color: #666666">0</span>,<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">2</span>; A(<span style="color: #666666">1</span>,<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">6</span>;
b(<span style="color: #666666">0</span>) <span style="color: #666666">=</span> <span style="color: #666666">2</span>; b(<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">-8</span>;
cout <span style="color: #666666">&lt;&lt;</span> <span style="color: #BA2121">&quot;The Matrix A that we are using: &quot;</span> <span style="color: #666666">&lt;&lt;</span> endl;
A.Print();
cout <span style="color: #666666">&lt;&lt;</span> endl;
xsd <span style="color: #666666">=</span> SteepestDescent(A,b,x0);
cout <span style="color: #666666">&lt;&lt;</span> <span style="color: #BA2121">&quot;The approximate solution using Steepest Descent is: &quot;</span> <span style="color: #666666">&lt;&lt;</span> endl;
xsd.Print();
cout <span style="color: #666666">&lt;&lt;</span> endl;
}
</pre></div>
<p>
</div>
</div>
This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
<p>
<p>
@@ -243,7 +255,7 @@ This result implies that \( C(\beta) \) is a convex function since the matrix \(
<li><a href="._Splines-bs033.html">34</a></li>
<li><a href="._Splines-bs034.html">35</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs026.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+78 -97
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -206,44 +191,40 @@ MathJax.Hub.Config({
<a name="part0026"></a>
<!-- !split -->
<h2 id="___sec25" class="anchor">Simple program </h2>
<p>
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
$$
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
$$
<p>
We can use the expression we computed for the gradient and let use a
\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \).
<p>
And finally we can compare our solution for \( \beta \) with the analytic result given by
\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
<h2 id="___sec25" class="anchor">The routine for the steepest descent method </h2>
<div class="panel panel-default">
<div class="panel-body">
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
<p>
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
<span style="color: #BA2121; font-style: italic">&quot;&quot;&quot;</span>
<span style="color: #BA2121; font-style: italic">The following setup is just a suggestion, feel free to write it the way you like.</span>
<span style="color: #BA2121; font-style: italic">&quot;&quot;&quot;</span>
<span style="color: #408080; font-style: italic">#Setup problem described in the exercise</span>
N <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#Nr of datapoints</span>
M <span style="color: #666666">=</span> <span style="color: #666666">2</span> <span style="color: #408080; font-style: italic">#Nr of features</span>
x <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(N) <span style="color: #408080; font-style: italic">#Uniformly generated x-values in [0,1]</span>
y <span style="color: #666666">=</span> <span style="color: #666666">5*</span>x<span style="color: #666666">**2</span> <span style="color: #666666">+</span> <span style="color: #666666">0.1*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(N)
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones(N),x] <span style="color: #408080; font-style: italic">#Construct design matrix</span>
<span style="color: #408080; font-style: italic">#Compute beta according to normal equations to compare with GD solution</span>
Xt_X_inv <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,X))
Xt_y <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>transpose(),y)
beta_NE <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(Xt_X_inv,Xt_y)
<span style="color: #008000; font-weight: bold">print</span>(beta_NE)
<!-- code=c++ (!bc cppcod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span>Vector <span style="color: #0000FF">SteepestDescent</span>(Matrix A, Vector b, Vector x0){
<span style="color: #B00040">int</span> IterMax, i;
<span style="color: #B00040">int</span> dim <span style="color: #666666">=</span> x0.Dimension();
<span style="color: #008000; font-weight: bold">const</span> <span style="color: #B00040">double</span> tolerance <span style="color: #666666">=</span> <span style="color: #666666">1.0e-14</span>;
Vector x(dim),f(dim),z(dim);
<span style="color: #B00040">double</span> c,alpha,d;
IterMax <span style="color: #666666">=</span> <span style="color: #666666">30</span>;
x <span style="color: #666666">=</span> x0;
f <span style="color: #666666">=</span> A<span style="color: #666666">*</span>x<span style="color: #666666">-</span>b;
i <span style="color: #666666">=</span> <span style="color: #666666">0</span>;
<span style="color: #008000; font-weight: bold">while</span> (i <span style="color: #666666">&lt;=</span> IterMax){
z <span style="color: #666666">=</span> A<span style="color: #666666">*</span>f;
c <span style="color: #666666">=</span> dot(f,f);
alpha <span style="color: #666666">=</span> c<span style="color: #666666">/</span>dot(f,z);
x <span style="color: #666666">=</span> x <span style="color: #666666">-</span> alpha<span style="color: #666666">*</span>f;
f <span style="color: #666666">=</span> A<span style="color: #666666">*</span>x<span style="color: #666666">-</span>b;
<span style="color: #008000; font-weight: bold">if</span>(sqrt(dot(f,f)) <span style="color: #666666">&lt;</span> tolerance) <span style="color: #008000; font-weight: bold">break</span>;
i<span style="color: #666666">++</span>;
}
<span style="color: #008000; font-weight: bold">return</span> x;
}
</pre></div>
<p>
</div>
</div>
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
@@ -270,7 +251,7 @@ beta_NE <span style="color: #666666">=</span> np<span style="color: #666666">.</
<li><a href="._Splines-bs034.html">35</a></li>
<li><a href="._Splines-bs035.html">36</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs027.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+71 -102
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -204,54 +189,38 @@ MathJax.Hub.Config({
<p>&nbsp;</p><p>&nbsp;</p><p>&nbsp;</p> <!-- add vertical space -->
<a name="part0027"></a>
<!-- !split -->
<!-- !split -->
<h2 id="___sec26" class="anchor">Gradient Descent Example </h2>
<h2 id="___sec26" class="anchor">Revisiting our first homework </h2>
<p>
Another simple example is here
<p>
We will use linear regression as a case study for the gradient descent
methods. Linear regression is a great test case for the gradient
descent methods discussed in the lectures since it has several
desirable properties such as:
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Importing various packages</span>
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">random</span> <span style="color: #008000; font-weight: bold">import</span> random, seed
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">mpl_toolkits.mplot3d</span> <span style="color: #008000; font-weight: bold">import</span> Axes3D
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">matplotlib</span> <span style="color: #008000; font-weight: bold">import</span> cm
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">matplotlib.ticker</span> <span style="color: #008000; font-weight: bold">import</span> LinearLocator, FormatStrFormatter
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">sys</span>
<ol>
<li> An analytical solution (recall homework set 1).</li>
<li> The gradient can be computed analytically.</li>
<li> The cost function is convex which guarantees that gradient descent converges for small enough learning rates</li>
</ol>
x <span style="color: #666666">=</span> <span style="color: #666666">2*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
y <span style="color: #666666">=</span> <span style="color: #666666">4+3*</span>x<span style="color: #666666">+</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
We revisit the example from homework set 1 where we had
$$
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
$$
xb <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)), x]
beta_linreg <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb))<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>T)<span style="color: #666666">.</span>dot(y)
<span style="color: #008000; font-weight: bold">print</span>(beta_linreg)
beta <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">2</span>,<span style="color: #666666">1</span>)
with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
The linear regression model is given by
$$
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
$$
eta <span style="color: #666666">=</span> <span style="color: #666666">0.1</span>
Niterations <span style="color: #666666">=</span> <span style="color: #666666">1000</span>
m <span style="color: #666666">=</span> <span style="color: #666666">100</span>
such that
$$
\hat{y}_i = \beta_0 + \beta_1 x_i.
$$
<span style="color: #008000; font-weight: bold">for</span> <span style="color: #008000">iter</span> <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(Niterations):
gradients <span style="color: #666666">=</span> <span style="color: #666666">2.0/</span>m<span style="color: #666666">*</span>xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>dot(beta)<span style="color: #666666">-</span>y)
beta <span style="color: #666666">-=</span> eta<span style="color: #666666">*</span>gradients
<span style="color: #008000; font-weight: bold">print</span>(beta)
xnew <span style="color: #666666">=</span> np<span style="color: #666666">.</span>array([[<span style="color: #666666">0</span>],[<span style="color: #666666">2</span>]])
xbnew <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">2</span>,<span style="color: #666666">1</span>)), xnew]
ypredict <span style="color: #666666">=</span> xbnew<span style="color: #666666">.</span>dot(beta)
ypredict2 <span style="color: #666666">=</span> xbnew<span style="color: #666666">.</span>dot(beta_linreg)
plt<span style="color: #666666">.</span>plot(xnew, ypredict, <span style="color: #BA2121">&quot;r-&quot;</span>)
plt<span style="color: #666666">.</span>plot(xnew, ypredict2, <span style="color: #BA2121">&quot;b-&quot;</span>)
plt<span style="color: #666666">.</span>plot(x, y ,<span style="color: #BA2121">&#39;ro&#39;</span>)
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">0</span>,<span style="color: #666666">2.0</span>,<span style="color: #666666">0</span>, <span style="color: #666666">15.0</span>])
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r&#39;$x$&#39;</span>)
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r&#39;$y$&#39;</span>)
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">r&#39;Gradient descent example&#39;</span>)
plt<span style="color: #666666">.</span>show()
</pre></div>
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
@@ -278,7 +247,7 @@ plt<span style="color: #666666">.</span>show()
<li><a href="._Splines-bs035.html">36</a></li>
<li><a href="._Splines-bs036.html">37</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs028.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+65 -79
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -204,29 +189,30 @@ MathJax.Hub.Config({
<p>&nbsp;</p><p>&nbsp;</p><p>&nbsp;</p> <!-- add vertical space -->
<a name="part0028"></a>
<!-- !split -->
<!-- !split -->
<h2 id="___sec27" class="anchor">And a corresponding example using <b>scikit-learn</b> </h2>
<h2 id="___sec27" class="anchor">Gradient descent example </h2>
<p>
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Importing various packages</span>
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">random</span> <span style="color: #008000; font-weight: bold">import</span> random, seed
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.linear_model</span> <span style="color: #008000; font-weight: bold">import</span> SGDRegressor
<p>
It is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by
$$
X \equiv \begin{bmatrix}
1 &amp; x_1 \\
\vdots &amp; \vdots \\
1 &amp; x_{100} &amp; \\
\end{bmatrix}.
$$
x <span style="color: #666666">=</span> <span style="color: #666666">2*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
y <span style="color: #666666">=</span> <span style="color: #666666">4+3*</span>x<span style="color: #666666">+</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
The loss function is given by
$$
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
$$
and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
xb <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)), x]
beta_linreg <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb))<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>T)<span style="color: #666666">.</span>dot(y)
<span style="color: #008000; font-weight: bold">print</span>(beta_linreg)
sgdreg <span style="color: #666666">=</span> SGDRegressor(n_iter <span style="color: #666666">=</span> <span style="color: #666666">50</span>, penalty<span style="color: #666666">=</span><span style="color: #008000">None</span>, eta0<span style="color: #666666">=0.1</span>)
sgdreg<span style="color: #666666">.</span>fit(x,y<span style="color: #666666">.</span>ravel())
<span style="color: #008000; font-weight: bold">print</span>(sgdreg<span style="color: #666666">.</span>intercept_, sgdreg<span style="color: #666666">.</span>coef_)
</pre></div>
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
@@ -253,7 +239,7 @@ sgdreg<span style="color: #666666">.</span>fit(x,y<span style="color: #666666">.
<li><a href="._Splines-bs036.html">37</a></li>
<li><a href="._Splines-bs037.html">38</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs029.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+53 -110
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -204,62 +189,20 @@ MathJax.Hub.Config({
<p>&nbsp;</p><p>&nbsp;</p><p>&nbsp;</p> <!-- add vertical space -->
<a name="part0029"></a>
<!-- !split -->
<!-- !split -->
<h2 id="___sec28" class="anchor">Gradient descent and Ridge </h2>
<h2 id="___sec28" class="anchor">The derivative of the cost/loss function </h2>
<p>
We have also discussed Ridge regression where the loss function contains a regularized given by the \( L_2 \) norm of \( \beta \),
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
$$
C_{\text{ridge}}(\beta) = ||X\beta -\mathbf{y}||^2 + \lambda ||\beta||^2, \ \lambda \geq 0.
$$
<p>
In order to minimize \( C_{\text{ridge}}(\beta) \) using GD we only have adjust the gradient as follows
$$
\nabla_\beta C_{\text{ridge}}(\beta) = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
\end{bmatrix} + 2\lambda\begin{bmatrix} \beta_0 \\ \beta_1\end{bmatrix} = 2 (X^T(X\beta - \mathbf{y})+\lambda \beta).
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
$$
<p>
We can now extend our program to minimize \( C_{\text{ridge}}(\beta) \) using gradient descent and compare with the analytical solution given by
$$
\beta_{\text{ridge}} = \left(X^T X + \lambda I_{2 \times 2} \right)^{-1} X^T \mathbf{y},
$$
where \( X \) is the design matrix defined above.
for \( \lambda = {0,1,10,50,100} \) (\( \lambda = 0 \) corresponds to ordinary least squares).
We can then compute \( ||\beta_{\text{ridge}}|| \) for each \( \lambda \).
<p>
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
<span style="color: #BA2121; font-style: italic">&quot;&quot;&quot;</span>
<span style="color: #BA2121; font-style: italic">The following setup is just a suggestion, feel free to write it the way you like.</span>
<span style="color: #BA2121; font-style: italic">&quot;&quot;&quot;</span>
<span style="color: #408080; font-style: italic">#Setup problem described in the exercise</span>
N <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#Nr of datapoints</span>
M <span style="color: #666666">=</span> <span style="color: #666666">2</span> <span style="color: #408080; font-style: italic">#Nr of features</span>
x <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(N)
y <span style="color: #666666">=</span> <span style="color: #666666">5*</span>x<span style="color: #666666">**2</span> <span style="color: #666666">+</span> <span style="color: #666666">0.1*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(N)
<span style="color: #408080; font-style: italic">#Compute analytic beta for Ridge regression </span>
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones(N),x]
XT_X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,X)
l <span style="color: #666666">=</span> <span style="color: #666666">0.1</span> <span style="color: #408080; font-style: italic">#Ridge parameter lambda</span>
Id <span style="color: #666666">=</span> np<span style="color: #666666">.</span>eye(XT_X<span style="color: #666666">.</span>shape[<span style="color: #666666">0</span>])
Z <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(XT_X<span style="color: #666666">+</span>l<span style="color: #666666">*</span>Id)
beta_ridge <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(Z,np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,y))
<span style="color: #008000; font-weight: bold">print</span>(beta_ridge)
<span style="color: #008000; font-weight: bold">print</span>(np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>norm(beta_ridge)) <span style="color: #408080; font-style: italic">#||beta||</span>
</pre></div>
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
@@ -286,7 +229,7 @@ beta_ridge <span style="color: #666666">=</span> np<span style="color: #666666">
<li><a href="._Splines-bs037.html">38</a></li>
<li><a href="._Splines-bs038.html">39</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs030.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+55 -74
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -206,21 +191,17 @@ MathJax.Hub.Config({
<a name="part0030"></a>
<!-- !split -->
<h2 id="___sec29" class="anchor">Stochastic Gradient Descent </h2>
<p>
Stochastic gradient descent (SGD) and variants thereof address some of
the shortcomings of the Gradient descent method discussed above.
<p>
The underlying idea of SGD comes from the observation that the cost
function, which we want to minimize, can almost always be written as a
sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \),
<h2 id="___sec29" class="anchor">The Hessian matrix </h2>
The Hessian matrix of \( C(\beta) \) is given by
$$
C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i,
\mathbf{\beta}).
\hat{H} \equiv \begin{bmatrix}
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} &amp; \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} &amp; \frac{\partial^2 C(\beta)}{\partial \beta_1^2} &amp; \\
\end{bmatrix} = 2X^T X.
$$
This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
@@ -247,7 +228,7 @@ $$
<li><a href="._Splines-bs038.html">39</a></li>
<li><a href="._Splines-bs039.html">40</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs031.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+78 -72
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -206,23 +191,44 @@ MathJax.Hub.Config({
<a name="part0031"></a>
<!-- !split -->
<h2 id="___sec30" class="anchor">Computation of gradients </h2>
<h2 id="___sec30" class="anchor">Simple program </h2>
<p>
This in turn means that the gradient can be
computed as a sum over \( i \)-gradients
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
$$
\nabla_\beta C(\mathbf{\beta}) = \sum_i^n \nabla_\beta c_i(\mathbf{x}_i,
\mathbf{\beta}).
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
$$
<p>
Stochasticity/randomness is introduced by only taking the
gradient on a subset of the data called minibatches. If there are \( n \)
data points and the size of each minibatch is \( M \), there will be \( n/M \)
minibatches. We denote these minibatches by \( B_k \) where
\( k=1,\cdots,n/M \).
We can use the expression we computed for the gradient and let use a
\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \).
<p>
And finally we can compare our solution for \( \beta \) with the analytic result given by
\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
<p>
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
<span style="color: #BA2121; font-style: italic">&quot;&quot;&quot;</span>
<span style="color: #BA2121; font-style: italic">The following setup is just a suggestion, feel free to write it the way you like.</span>
<span style="color: #BA2121; font-style: italic">&quot;&quot;&quot;</span>
<span style="color: #408080; font-style: italic">#Setup problem described in the exercise</span>
N <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#Nr of datapoints</span>
M <span style="color: #666666">=</span> <span style="color: #666666">2</span> <span style="color: #408080; font-style: italic">#Nr of features</span>
x <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(N) <span style="color: #408080; font-style: italic">#Uniformly generated x-values in [0,1]</span>
y <span style="color: #666666">=</span> <span style="color: #666666">5*</span>x<span style="color: #666666">**2</span> <span style="color: #666666">+</span> <span style="color: #666666">0.1*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(N)
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones(N),x] <span style="color: #408080; font-style: italic">#Construct design matrix</span>
<span style="color: #408080; font-style: italic">#Compute beta according to normal equations to compare with GD solution</span>
Xt_X_inv <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,X))
Xt_y <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>transpose(),y)
beta_NE <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(Xt_X_inv,Xt_y)
<span style="color: #008000; font-weight: bold">print</span>(beta_NE)
</pre></div>
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
@@ -249,7 +255,7 @@ minibatches. We denote these minibatches by \( B_k \) where
<li><a href="._Splines-bs039.html">40</a></li>
<li><a href="._Splines-bs040.html">41</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs032.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+89 -81
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -206,27 +191,52 @@ MathJax.Hub.Config({
<a name="part0032"></a>
<!-- !split -->
<h2 id="___sec31" class="anchor">SGD example </h2>
As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \)
and we choose to have \( M=5 \) minibathces,
then each minibatch contains two data points. In particular we have
\( B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 =
(\mathbf{x}_9,\mathbf{x}_{10}) \). Note that if you choose \( M=1 \) you
have only a single batch with all data points and on the other extreme,
you may choose \( M=n \) resulting in a minibatch for each datapoint, i.e
\( B_k = \mathbf{x}_k \).
<h2 id="___sec31" class="anchor">Gradient Descent Example </h2>
<p>
The idea is now to approximate the gradient by replacing the sum over
all data points with a sum over the data points in one the minibatches
picked at random in each gradient descent step
$$
\nabla_{\beta}
C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i,
\mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta
c_i(\mathbf{x}_i, \mathbf{\beta}).
$$
Another simple example is here
<p>
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Importing various packages</span>
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">random</span> <span style="color: #008000; font-weight: bold">import</span> random, seed
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">mpl_toolkits.mplot3d</span> <span style="color: #008000; font-weight: bold">import</span> Axes3D
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">matplotlib</span> <span style="color: #008000; font-weight: bold">import</span> cm
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">matplotlib.ticker</span> <span style="color: #008000; font-weight: bold">import</span> LinearLocator, FormatStrFormatter
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">sys</span>
x <span style="color: #666666">=</span> <span style="color: #666666">2*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
y <span style="color: #666666">=</span> <span style="color: #666666">4+3*</span>x<span style="color: #666666">+</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
xb <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)), x]
beta_linreg <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb))<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>T)<span style="color: #666666">.</span>dot(y)
<span style="color: #008000; font-weight: bold">print</span>(beta_linreg)
beta <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">2</span>,<span style="color: #666666">1</span>)
eta <span style="color: #666666">=</span> <span style="color: #666666">0.1</span>
Niterations <span style="color: #666666">=</span> <span style="color: #666666">1000</span>
m <span style="color: #666666">=</span> <span style="color: #666666">100</span>
<span style="color: #008000; font-weight: bold">for</span> <span style="color: #008000">iter</span> <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(Niterations):
gradients <span style="color: #666666">=</span> <span style="color: #666666">2.0/</span>m<span style="color: #666666">*</span>xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>dot(beta)<span style="color: #666666">-</span>y)
beta <span style="color: #666666">-=</span> eta<span style="color: #666666">*</span>gradients
<span style="color: #008000; font-weight: bold">print</span>(beta)
xnew <span style="color: #666666">=</span> np<span style="color: #666666">.</span>array([[<span style="color: #666666">0</span>],[<span style="color: #666666">2</span>]])
xbnew <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">2</span>,<span style="color: #666666">1</span>)), xnew]
ypredict <span style="color: #666666">=</span> xbnew<span style="color: #666666">.</span>dot(beta)
ypredict2 <span style="color: #666666">=</span> xbnew<span style="color: #666666">.</span>dot(beta_linreg)
plt<span style="color: #666666">.</span>plot(xnew, ypredict, <span style="color: #BA2121">&quot;r-&quot;</span>)
plt<span style="color: #666666">.</span>plot(xnew, ypredict2, <span style="color: #BA2121">&quot;b-&quot;</span>)
plt<span style="color: #666666">.</span>plot(x, y ,<span style="color: #BA2121">&#39;ro&#39;</span>)
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">0</span>,<span style="color: #666666">2.0</span>,<span style="color: #666666">0</span>, <span style="color: #666666">15.0</span>])
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r&#39;$x$&#39;</span>)
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r&#39;$y$&#39;</span>)
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">r&#39;Gradient descent example&#39;</span>)
plt<span style="color: #666666">.</span>show()
</pre></div>
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
@@ -252,8 +262,6 @@ $$
<li><a href="._Splines-bs039.html">40</a></li>
<li><a href="._Splines-bs040.html">41</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs033.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+63 -76
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -206,22 +191,27 @@ MathJax.Hub.Config({
<a name="part0033"></a>
<!-- !split -->
<h2 id="___sec32" class="anchor">The gradient step </h2>
<h2 id="___sec32" class="anchor">And a corresponding example using <b>scikit-learn</b> </h2>
<p>
Thus a gradient descent step now looks like
$$
\beta_{j+1} = \beta_j - \gamma_j \sum_{i \in B_k}^n \nabla_\beta c_i(\mathbf{x}_i,
\mathbf{\beta})
$$
<p>
where \( k \) is picked at random with equal
probability from \( [1,n/M] \). An iteration over the number of
minibathces (n/M) is commonly referred to as an epoch. Thus it is
typical to choose a number of epochs and for each epoch iterate over
the number of minibatches, as exemplified in the code below.
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Importing various packages</span>
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">random</span> <span style="color: #008000; font-weight: bold">import</span> random, seed
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.linear_model</span> <span style="color: #008000; font-weight: bold">import</span> SGDRegressor
x <span style="color: #666666">=</span> <span style="color: #666666">2*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
y <span style="color: #666666">=</span> <span style="color: #666666">4+3*</span>x<span style="color: #666666">+</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
xb <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)), x]
beta_linreg <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb))<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>T)<span style="color: #666666">.</span>dot(y)
<span style="color: #008000; font-weight: bold">print</span>(beta_linreg)
sgdreg <span style="color: #666666">=</span> SGDRegressor(n_iter <span style="color: #666666">=</span> <span style="color: #666666">50</span>, penalty<span style="color: #666666">=</span><span style="color: #008000">None</span>, eta0<span style="color: #666666">=0.1</span>)
sgdreg<span style="color: #666666">.</span>fit(x,y<span style="color: #666666">.</span>ravel())
<span style="color: #008000; font-weight: bold">print</span>(sgdreg<span style="color: #666666">.</span>intercept_, sgdreg<span style="color: #666666">.</span>coef_)
</pre></div>
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
@@ -246,9 +236,6 @@ the number of minibatches, as exemplified in the code below.
<li><a href="._Splines-bs039.html">40</a></li>
<li><a href="._Splines-bs040.html">41</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs042.html">43</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs034.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+94 -88
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -204,37 +189,62 @@ MathJax.Hub.Config({
<p>&nbsp;</p><p>&nbsp;</p><p>&nbsp;</p> <!-- add vertical space -->
<a name="part0034"></a>
<!-- !split -->
<!-- !split -->
<h2 id="___sec33" class="anchor">Simple example code </h2>
<h2 id="___sec33" class="anchor">Gradient descent and Ridge </h2>
<p>
We have also discussed Ridge regression where the loss function contains a regularized given by the \( L_2 \) norm of \( \beta \),
$$
C_{\text{ridge}}(\beta) = ||X\beta -\mathbf{y}||^2 + \lambda ||\beta||^2, \ \lambda \geq 0.
$$
<p>
In order to minimize \( C_{\text{ridge}}(\beta) \) using GD we only have adjust the gradient as follows
$$
\nabla_\beta C_{\text{ridge}}(\beta) = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
\end{bmatrix} + 2\lambda\begin{bmatrix} \beta_0 \\ \beta_1\end{bmatrix} = 2 (X^T(X\beta - \mathbf{y})+\lambda \beta).
$$
<p>
We can now extend our program to minimize \( C_{\text{ridge}}(\beta) \) using gradient descent and compare with the analytical solution given by
$$
\beta_{\text{ridge}} = \left(X^T X + \lambda I_{2 \times 2} \right)^{-1} X^T \mathbf{y},
$$
for \( \lambda = {0,1,10,50,100} \) (\( \lambda = 0 \) corresponds to ordinary least squares).
We can then compute \( ||\beta_{\text{ridge}}|| \) for each \( \lambda \).
<p>
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
n <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#100 datapoints </span>
M <span style="color: #666666">=</span> <span style="color: #666666">5</span> <span style="color: #408080; font-style: italic">#size of each minibatch</span>
m <span style="color: #666666">=</span> <span style="color: #008000">int</span>(n<span style="color: #666666">/</span>M) <span style="color: #408080; font-style: italic">#number of minibatches</span>
n_epochs <span style="color: #666666">=</span> <span style="color: #666666">10</span> <span style="color: #408080; font-style: italic">#number of epochs</span>
<span style="color: #BA2121; font-style: italic">&quot;&quot;&quot;</span>
<span style="color: #BA2121; font-style: italic">The following setup is just a suggestion, feel free to write it the way you like.</span>
<span style="color: #BA2121; font-style: italic">&quot;&quot;&quot;</span>
j <span style="color: #666666">=</span> <span style="color: #666666">0</span>
<span style="color: #008000; font-weight: bold">for</span> epoch <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(<span style="color: #666666">1</span>,n_epochs<span style="color: #666666">+1</span>):
<span style="color: #008000; font-weight: bold">for</span> i <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(m):
k <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randint(m) <span style="color: #408080; font-style: italic">#Pick the k-th minibatch at random</span>
<span style="color: #408080; font-style: italic">#Compute the gradient using the data in minibatch Bk</span>
<span style="color: #408080; font-style: italic">#Compute new suggestion for </span>
j <span style="color: #666666">+=</span> <span style="color: #666666">1</span>
<span style="color: #408080; font-style: italic">#Setup problem described in the exercise</span>
N <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#Nr of datapoints</span>
M <span style="color: #666666">=</span> <span style="color: #666666">2</span> <span style="color: #408080; font-style: italic">#Nr of features</span>
x <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(N)
y <span style="color: #666666">=</span> <span style="color: #666666">5*</span>x<span style="color: #666666">**2</span> <span style="color: #666666">+</span> <span style="color: #666666">0.1*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(N)
<span style="color: #408080; font-style: italic">#Compute analytic beta for Ridge regression </span>
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones(N),x]
XT_X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,X)
l <span style="color: #666666">=</span> <span style="color: #666666">0.1</span> <span style="color: #408080; font-style: italic">#Ridge parameter lambda</span>
Id <span style="color: #666666">=</span> np<span style="color: #666666">.</span>eye(XT_X<span style="color: #666666">.</span>shape[<span style="color: #666666">0</span>])
Z <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(XT_X<span style="color: #666666">+</span>l<span style="color: #666666">*</span>Id)
beta_ridge <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(Z,np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,y))
<span style="color: #008000; font-weight: bold">print</span>(beta_ridge)
<span style="color: #008000; font-weight: bold">print</span>(np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>norm(beta_ridge)) <span style="color: #408080; font-style: italic">#||beta||</span>
</pre></div>
<p>
Taking the gradient only on a subset of the data has two important
benefits. First, it introduces randomness which decreases the chance
that our opmization scheme gets stuck in a local minima. Second, if
the size of the minibatches are small relative to the number of
datapoints (\( M < n \)), the computation of the gradient is much
cheaper since we sum over the datapoints in the \( k-th \) minibatch and not
all \( n \) datapoints.
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
@@ -258,10 +268,6 @@ all \( n \) datapoints.
<li><a href="._Splines-bs039.html">40</a></li>
<li><a href="._Splines-bs040.html">41</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs042.html">43</a></li>
<li><a href="._Splines-bs043.html">44</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs035.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+58 -77
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -206,19 +191,20 @@ MathJax.Hub.Config({
<a name="part0035"></a>
<!-- !split -->
<h2 id="___sec34" class="anchor">When do we stop? </h2>
<h2 id="___sec34" class="anchor">Stochastic Gradient Descent </h2>
<p>
A natural question is when do we stop the search for a new minimum?
One possibility is to compute the full gradient after a given number
of epochs and check if the norm of the gradient is smaller than some
threshold and stop if true. However, the condition that the gradient
is zero is valid also for local minima, so this would only tell us
that we are close to a local/global minimum. However, we could also
evaluate the cost function at this point, store the result and
continue the search. If the test kicks in at a later stage we can
compare the values of the cost function and keep the \( \beta \) that
gave the lowest value.
Stochastic gradient descent (SGD) and variants thereof address some of
the shortcomings of the Gradient descent method discussed above.
<p>
The underlying idea of SGD comes from the observation that the cost
function, which we want to minimize, can almost always be written as a
sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \),
$$
C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i,
\mathbf{\beta}).
$$
<p>
<p>
@@ -242,11 +228,6 @@ gave the lowest value.
<li><a href="._Splines-bs039.html">40</a></li>
<li><a href="._Splines-bs040.html">41</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs042.html">43</a></li>
<li><a href="._Splines-bs043.html">44</a></li>
<li><a href="._Splines-bs044.html">45</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs036.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+58 -107
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -206,51 +191,23 @@ MathJax.Hub.Config({
<a name="part0036"></a>
<!-- !split -->
<h2 id="___sec35" class="anchor">Slightly different approach </h2>
<h2 id="___sec35" class="anchor">Computation of gradients </h2>
<p>
Another approach is to let the step length \( \gamma_j \) depend on the
number of epochs in such a way that it becomes very small after a
reasonable time such that we do not move at all.
This in turn means that the gradient can be
computed as a sum over \( i \)-gradients
$$
\nabla_\beta C(\mathbf{\beta}) = \sum_i^n \nabla_\beta c_i(\mathbf{x}_i,
\mathbf{\beta}).
$$
<p>
As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in <em>time</em> \( t \).
Stochasticity/randomness is introduced by only taking the
gradient on a subset of the data called minibatches. If there are \( n \)
data points and the size of each minibatch is \( M \), there will be \( n/M \)
minibatches. We denote these minibatches by \( B_k \) where
\( k=1,\cdots,n/M \).
<p>
In this way we can fix the number of epochs, compute \( \beta \) and
evaluate the cost function at the end. Repeating the computation will
give a different result since the scheme is random by design. Then we
pick the final \( \beta \) that gives the lowest value of the cost
function.
<p>
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">step_length</span>(t,t0,t1):
<span style="color: #008000; font-weight: bold">return</span> t0<span style="color: #666666">/</span>(t<span style="color: #666666">+</span>t1)
n <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#100 datapoints </span>
M <span style="color: #666666">=</span> <span style="color: #666666">5</span> <span style="color: #408080; font-style: italic">#size of each minibatch</span>
m <span style="color: #666666">=</span> <span style="color: #008000">int</span>(n<span style="color: #666666">/</span>M) <span style="color: #408080; font-style: italic">#number of minibatches</span>
n_epochs <span style="color: #666666">=</span> <span style="color: #666666">500</span> <span style="color: #408080; font-style: italic">#number of epochs</span>
t0 <span style="color: #666666">=</span> <span style="color: #666666">1.0</span>
t1 <span style="color: #666666">=</span> <span style="color: #666666">10</span>
gamma_j <span style="color: #666666">=</span> t0<span style="color: #666666">/</span>t1
j <span style="color: #666666">=</span> <span style="color: #666666">0</span>
<span style="color: #008000; font-weight: bold">for</span> epoch <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(<span style="color: #666666">1</span>,n_epochs<span style="color: #666666">+1</span>):
<span style="color: #008000; font-weight: bold">for</span> i <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(m):
k <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randint(m) <span style="color: #408080; font-style: italic">#Pick the k-th minibatch at random</span>
<span style="color: #408080; font-style: italic">#Compute the gradient using the data in minibatch Bk</span>
<span style="color: #408080; font-style: italic">#Compute new suggestion for beta</span>
t <span style="color: #666666">=</span> epoch<span style="color: #666666">*</span>m<span style="color: #666666">+</span>i
gamma_j <span style="color: #666666">=</span> step_length(t,t0,t1)
j <span style="color: #666666">+=</span> <span style="color: #666666">1</span>
<span style="color: #008000; font-weight: bold">print</span>(<span style="color: #BA2121">&quot;gamma_j after </span><span style="color: #BB6688; font-weight: bold">%d</span><span style="color: #BA2121"> epochs: </span><span style="color: #BB6688; font-weight: bold">%g</span><span style="color: #BA2121">&quot;</span> <span style="color: #666666">%</span> (n_epochs,gamma_j))
</pre></div>
<p>
<p>
<!-- navigation buttons at the bottom of the page -->
@@ -272,12 +229,6 @@ j <span style="color: #666666">=</span> <span style="color: #666666">0</span>
<li><a href="._Splines-bs039.html">40</a></li>
<li><a href="._Splines-bs040.html">41</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs042.html">43</a></li>
<li><a href="._Splines-bs043.html">44</a></li>
<li><a href="._Splines-bs044.html">45</a></li>
<li><a href="._Splines-bs045.html">46</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs037.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+47 -62
View File
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
</ul>
</li>
@@ -249,7 +234,7 @@ MathJax.Hub.Config({
<li><a href="._Splines-bs008.html">9</a></li>
<li><a href="._Splines-bs009.html">10</a></li>
<li><a href="">...</a></li>
<li><a href="._Splines-bs047.html">48</a></li>
<li><a href="._Splines-bs041.html">42</a></li>
<li><a href="._Splines-bs001.html">&raquo;</a></li>
</ul>
<!-- ------------------- end of main content --------------- -->
+157 -374
View File
@@ -662,7 +662,7 @@ When we have found the exact solution, \( \hat{r}=0 \).
<section>
<h2 id="___sec18">Conjugate gradient method </h2>
<h2 id="___sec18">Gradient method </h2>
<p>
The residual is zero when we reach the minimum of the quadratic equation
@@ -680,14 +680,150 @@ symmetric. If we search for a minimum of the quantum mechanical
variance, then the matrix \( \hat{A} \), which is called the Hessian, is
given by the second-derivative of the function we want to minimize.
This quantity is always positive definite.
<p>
More details will be added here soon.
</section>
<section>
<h2 id="___sec19">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
<h2 id="___sec19">Steepest descent method </h2>
<p>
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
We can assume without loss of generality that
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{x}_0=0,
\end{equation*}
$$
<p>&nbsp;<br>
or consider the system
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
\end{equation*}
$$
<p>&nbsp;<br>
instead.
</section>
<section>
<h2 id="___sec20">Steepest descent method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
<p>&nbsp;<br>
$$
\begin{equation*}
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
\end{equation*}
$$
<p>&nbsp;<br>
This suggests taking the first basis vector \( \hat{p}_1 \)
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
which equals
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{A}\hat{x}_0-\hat{b},
\end{equation*}
$$
<p>&nbsp;<br>
and
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
</div>
</section>
<section>
<h2 id="___sec21">Gradient descent method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
\end{equation*}
$$
<p>&nbsp;<br>
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
\( \hat{x}=\hat{x}_k \),
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
This gives the following expression
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
\end{equation*}
$$
<p>&nbsp;<br>
</div>
</section>
<section>
<h2 id="___sec22">Final expressions </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
We can also compute the residual iteratively as
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
\end{equation*}
$$
<p>&nbsp;<br>
which equals
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
\end{equation*}
$$
<p>&nbsp;<br>
or
<p>&nbsp;<br>
$$
\begin{equation*}
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
\end{equation*}
$$
<p>&nbsp;<br>
which gives
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
\end{equation*}
$$
<p>&nbsp;<br>
</div>
</section>
<section>
<h2 id="___sec23">The Steepest descent algorithm </h2>
</section>
<section>
<h2 id="___sec24">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
@@ -713,11 +849,7 @@ More details will be added here soon.
cout &lt;&lt; <span style="color: #CD5555">&quot;The Matrix A that we are using: &quot;</span> &lt;&lt; endl;
A.Print();
cout &lt;&lt; endl;
x = ConjugateGradient(A,b,x0);
xsd = SteepestDescent(A,b,x0);
cout &lt;&lt; <span style="color: #CD5555">&quot;The approximate solution using Conjugate Gradient is: &quot;</span> &lt;&lt; endl;
x.Print();
cout &lt;&lt; endl;
cout &lt;&lt; <span style="color: #CD5555">&quot;The approximate solution using Steepest Descent is: &quot;</span> &lt;&lt; endl;
xsd.Print();
cout &lt;&lt; endl;
@@ -729,7 +861,7 @@ More details will be added here soon.
<section>
<h2 id="___sec20">The routine for the steepest descent method </h2>
<h2 id="___sec25">The routine for the steepest descent method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
@@ -763,7 +895,7 @@ More details will be added here soon.
<section>
<h2 id="___sec21">Revisiting our first homework </h2>
<h2 id="___sec26">Revisiting our first homework </h2>
<p>
We will use linear regression as a case study for the gradient descent
@@ -803,7 +935,7 @@ $$
<section>
<h2 id="___sec22">Gradient descent example </h2>
<h2 id="___sec27">Gradient descent example </h2>
<p>
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
@@ -832,7 +964,7 @@ and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
<section>
<h2 id="___sec23">The derivative of the cost/loss function </h2>
<h2 id="___sec28">The derivative of the cost/loss function </h2>
<p>
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
@@ -849,7 +981,7 @@ where \( X \) is the design matrix defined above.
<section>
<h2 id="___sec24">The Hessian matrix </h2>
<h2 id="___sec29">The Hessian matrix </h2>
The Hessian matrix of \( C(\beta) \) is given by
<p>&nbsp;<br>
$$
@@ -865,7 +997,7 @@ This result implies that \( C(\beta) \) is a convex function since the matrix \(
<section>
<h2 id="___sec25">Simple program </h2>
<h2 id="___sec30">Simple program </h2>
<p>
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
@@ -909,7 +1041,7 @@ beta_NE = np.dot(Xt_X_inv,Xt_y)
<section>
<h2 id="___sec26">Gradient Descent Example </h2>
<h2 id="___sec31">Gradient Descent Example </h2>
<p>
Another simple example is here
@@ -959,7 +1091,7 @@ plt.show()
<section>
<h2 id="___sec27">And a corresponding example using <b>scikit-learn</b> </h2>
<h2 id="___sec32">And a corresponding example using <b>scikit-learn</b> </h2>
<p>
@@ -984,7 +1116,7 @@ sgdreg.fit(x,y.ravel())
<section>
<h2 id="___sec28">Gradient descent and Ridge </h2>
<h2 id="___sec33">Gradient descent and Ridge </h2>
<p>
We have also discussed Ridge regression where the loss function contains a regularized given by the \( L_2 \) norm of \( \beta \),
@@ -1048,7 +1180,7 @@ beta_ridge = np.dot(Z,np.dot(X.T,y))
<section>
<h2 id="___sec29">Stochastic Gradient Descent </h2>
<h2 id="___sec34">Stochastic Gradient Descent </h2>
<p>
Stochastic gradient descent (SGD) and variants thereof address some of
@@ -1068,7 +1200,7 @@ $$
<section>
<h2 id="___sec30">Computation of gradients </h2>
<h2 id="___sec35">Computation of gradients </h2>
<p>
This in turn means that the gradient can be
@@ -1090,7 +1222,7 @@ minibatches. We denote these minibatches by \( B_k \) where
<section>
<h2 id="___sec31">SGD example </h2>
<h2 id="___sec36">SGD example </h2>
As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \)
and we choose to have \( M=5 \) minibathces,
then each minibatch contains two data points. In particular we have
@@ -1116,7 +1248,7 @@ $$
<section>
<h2 id="___sec32">The gradient step </h2>
<h2 id="___sec37">The gradient step </h2>
<p>
Thus a gradient descent step now looks like
@@ -1137,7 +1269,7 @@ the number of minibatches, as exemplified in the code below.
<section>
<h2 id="___sec33">Simple example code </h2>
<h2 id="___sec38">Simple example code </h2>
<p>
@@ -1169,7 +1301,7 @@ all \( n \) datapoints.
<section>
<h2 id="___sec34">When do we stop? </h2>
<h2 id="___sec39">When do we stop? </h2>
<p>
A natural question is when do we stop the search for a new minimum?
@@ -1186,7 +1318,7 @@ gave the lowest value.
<section>
<h2 id="___sec35">Slightly different approach </h2>
<h2 id="___sec40">Slightly different approach </h2>
<p>
Another approach is to let the step length \( \gamma_j \) depend on the
@@ -1236,355 +1368,6 @@ j = <span style="color: #B452CD">0</span>
</section>
<section>
<h2 id="___sec36">Conjugate gradient (CG) method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
The success of the CG method for finding solutions of non-linear problems is based
on the theory of conjugate gradients for linear systems of equations. It belongs
to the class of iterative methods for solving problems from linear algebra of the type
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{A}\hat{x} = \hat{b}.
\end{equation*}
$$
<p>&nbsp;<br>
In the iterative process we end up with a problem like
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{r}= \hat{b}-\hat{A}\hat{x},
\end{equation*}
$$
<p>&nbsp;<br>
where \( \hat{r} \) is the so-called residual or error in the iterative process.
<p>
When we have found the exact solution, \( \hat{r}=0 \).
</div>
</section>
<section>
<h2 id="___sec37">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
The residual is zero when we reach the minimum of the quadratic equation
<p>&nbsp;<br>
$$
\begin{equation*}
P(\hat{x})=\frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T\hat{b},
\end{equation*}
$$
<p>&nbsp;<br>
with the constraint that the matrix \( \hat{A} \) is positive definite and symmetric.
If we search for a minimum of the quantum mechanical variance, then the matrix
\( \hat{A} \), which is called the Hessian, is given by the second-derivative of the function we want to minimize. This quantity is always positive definite. In our case this corresponds normally to the second derivative of the energy.
</div>
</section>
<section>
<h2 id="___sec38">Conjugate gradient method, Newton's method first </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
We seek the minimum of the energy or the variance as function of various variational parameters.
In our case we have thus a function \( f \) whose minimum we are seeking.
In Newton's method we set \( \nabla f = 0 \) and we can thus compute the next iteration point
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{x}-\hat{x}_i=\hat{A}^{-1}\nabla f(\hat{x}_i).
\end{equation*}
$$
<p>&nbsp;<br>
Subtracting this equation from that of \( \hat{x}_{i+1} \) we have
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{x}_{i+1}-\hat{x}_i=\hat{A}^{-1}(\nabla f(\hat{x}_{i+1})-\nabla f(\hat{x}_i)).
\end{equation*}
$$
<p>&nbsp;<br>
</div>
</section>
<section>
<h2 id="___sec39">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
In the CG method we define so-called conjugate directions and two vectors
\( \hat{s} \) and \( \hat{t} \)
are said to be
conjugate if
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{s}^T\hat{A}\hat{t}= 0.
\end{equation*}
$$
<p>&nbsp;<br>
The philosophy of the CG method is to perform searches in various conjugate directions
of our vectors \( \hat{x}_i \) obeying the above criterion, namely
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{x}_i^T\hat{A}\hat{x}_j= 0.
\end{equation*}
$$
<p>&nbsp;<br>
Two vectors are conjugate if they are orthogonal with respect to
this inner product. Being conjugate is a symmetric relation: if \( \hat{s} \) is conjugate to \( \hat{t} \), then \( \hat{t} \) is conjugate to \( \hat{s} \).
</div>
</section>
<section>
<h2 id="___sec40">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
An example is given by the eigenvectors of the matrix
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{v}_i^T\hat{A}\hat{v}_j= \lambda\hat{v}_i^T\hat{v}_j,
\end{equation*}
$$
<p>&nbsp;<br>
which is zero unless \( i=j \).
</div>
</section>
<section>
<h2 id="___sec41">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
Assume now that we have a symmetric positive-definite matrix \( \hat{A} \) of size
\( n\times n \). At each iteration \( i+1 \) we obtain the conjugate direction of a vector
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{x}_{i+1}=\hat{x}_{i}+\alpha_i\hat{p}_{i}.
\end{equation*}
$$
<p>&nbsp;<br>
We assume that \( \hat{p}_{i} \) is a sequence of \( n \) mutually conjugate directions.
Then the \( \hat{p}_{i} \) form a basis of \( R^n \) and we can expand the solution
$ \hat{A}\hat{x} = \hat{b}$ in this basis, namely
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{x} = \sum^{n}_{i=1} \alpha_i \hat{p}_i.
\end{equation*}
$$
<p>&nbsp;<br>
</div>
</section>
<section>
<h2 id="___sec42">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
The coefficients are given by
<p>&nbsp;<br>
$$
\begin{equation*}
\mathbf{A}\mathbf{x} = \sum^{n}_{i=1} \alpha_i \mathbf{A} \mathbf{p}_i = \mathbf{b}.
\end{equation*}
$$
<p>&nbsp;<br>
Multiplying with \( \hat{p}_k^T \) from the left gives
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{p}_k^T \hat{A}\hat{x} = \sum^{n}_{i=1} \alpha_i\hat{p}_k^T \hat{A}\hat{p}_i= \hat{p}_k^T \hat{b},
\end{equation*}
$$
<p>&nbsp;<br>
and we can define the coefficients \( \alpha_k \) as
<p>&nbsp;<br>
$$
\begin{equation*}
\alpha_k = \frac{\hat{p}_k^T \hat{b}}{\hat{p}_k^T \hat{A} \hat{p}_k}
\end{equation*}
$$
<p>&nbsp;<br>
</div>
</section>
<section>
<h2 id="___sec43">Conjugate gradient method and iterations </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
If we choose the conjugate vectors \( \hat{p}_k \) carefully,
then we may not need all of them to obtain a good approximation to the solution
\( \hat{x} \).
We want to regard the conjugate gradient method as an iterative method.
This will us to solve systems where \( n \) is so large that the direct
method would take too much time.
<p>
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
We can assume without loss of generality that
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{x}_0=0,
\end{equation*}
$$
<p>&nbsp;<br>
or consider the system
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
\end{equation*}
$$
<p>&nbsp;<br>
instead.
</div>
</section>
<section>
<h2 id="___sec44">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
<p>&nbsp;<br>
$$
\begin{equation*}
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
\end{equation*}
$$
<p>&nbsp;<br>
This suggests taking the first basis vector \( \hat{p}_1 \)
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
which equals
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{A}\hat{x}_0-\hat{b},
\end{equation*}
$$
<p>&nbsp;<br>
and
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
The other vectors in the basis will be conjugate to the gradient,
hence the name conjugate gradient method.
</div>
</section>
<section>
<h2 id="___sec45">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
\end{equation*}
$$
<p>&nbsp;<br>
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
\( \hat{x}=\hat{x}_k \),
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
Here, we insist that the directions \( \hat{p}_k \) are conjugate to each other,
so we take the direction closest to the gradient \( \hat{r}_k \)
under the conjugacy constraint.
This gives the following expression
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
\end{equation*}
$$
<p>&nbsp;<br>
</div>
</section>
<section>
<h2 id="___sec46">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
We can also compute the residual iteratively as
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
\end{equation*}
$$
<p>&nbsp;<br>
which equals
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
\end{equation*}
$$
<p>&nbsp;<br>
or
<p>&nbsp;<br>
$$
\begin{equation*}
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
\end{equation*}
$$
<p>&nbsp;<br>
which gives
<p>&nbsp;<br>
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
\end{equation*}
$$
<p>&nbsp;<br>
</div>
</section>
</div> <!-- class="slides" -->
</div> <!-- class="reveal" -->
+161 -373
View File
@@ -85,48 +85,39 @@ div { text-align: justify; text-justify: inter-word; }
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -639,7 +630,7 @@ When we have found the exact solution, \( \hat{r}=0 \).
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec18">Conjugate gradient method </h2>
<h2 id="___sec18">Gradient method </h2>
<p>
The residual is zero when we reach the minimum of the quadratic equation
@@ -657,12 +648,131 @@ given by the second-derivative of the function we want to minimize.
This quantity is always positive definite.
<p>
More details will be added here soon.
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec19">Steepest descent method </h2>
<p>
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
We can assume without loss of generality that
$$
\begin{equation*}
\hat{x}_0=0,
\end{equation*}
$$
or consider the system
$$
\begin{equation*}
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
\end{equation*}
$$
instead.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec19">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
<h2 id="___sec20">Steepest descent method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
$$
\begin{equation*}
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
\end{equation*}
$$
This suggests taking the first basis vector \( \hat{p}_1 \)
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
which equals
$$
\begin{equation*}
\hat{A}\hat{x}_0-\hat{b},
\end{equation*}
$$
and
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec21">Gradient descent method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
$$
\begin{equation*}
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
\end{equation*}
$$
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
\( \hat{x}=\hat{x}_k \),
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
This gives the following expression
$$
\begin{equation*}
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
\end{equation*}
$$
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec22">Final expressions </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
We can also compute the residual iteratively as
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
\end{equation*}
$$
which equals
$$
\begin{equation*}
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
\end{equation*}
$$
or
$$
\begin{equation*}
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
\end{equation*}
$$
which gives
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
\end{equation*}
$$
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec23">The Steepest descent algorithm </h2>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec24">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
@@ -689,11 +799,7 @@ More details will be added here soon.
cout &lt;&lt; <span style="color: #CD5555">&quot;The Matrix A that we are using: &quot;</span> &lt;&lt; endl;
A.Print();
cout &lt;&lt; endl;
x = ConjugateGradient(A,b,x0);
xsd = SteepestDescent(A,b,x0);
cout &lt;&lt; <span style="color: #CD5555">&quot;The approximate solution using Conjugate Gradient is: &quot;</span> &lt;&lt; endl;
x.Print();
cout &lt;&lt; endl;
cout &lt;&lt; <span style="color: #CD5555">&quot;The approximate solution using Steepest Descent is: &quot;</span> &lt;&lt; endl;
xsd.Print();
cout &lt;&lt; endl;
@@ -706,7 +812,7 @@ More details will be added here soon.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec20">The routine for the steepest descent method </h2>
<h2 id="___sec25">The routine for the steepest descent method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
@@ -742,7 +848,7 @@ More details will be added here soon.
<p>
<!-- !split -->
<h2 id="___sec21">Revisiting our first homework </h2>
<h2 id="___sec26">Revisiting our first homework </h2>
<p>
We will use linear regression as a case study for the gradient descent
@@ -775,7 +881,7 @@ $$
<p>
<!-- !split -->
<h2 id="___sec22">Gradient descent example </h2>
<h2 id="___sec27">Gradient descent example </h2>
<p>
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
@@ -800,7 +906,7 @@ and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec23">The derivative of the cost/loss function </h2>
<h2 id="___sec28">The derivative of the cost/loss function </h2>
<p>
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
@@ -815,7 +921,7 @@ where \( X \) is the design matrix defined above.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec24">The Hessian matrix </h2>
<h2 id="___sec29">The Hessian matrix </h2>
The Hessian matrix of \( C(\beta) \) is given by
$$
\hat{H} \equiv \begin{bmatrix}
@@ -829,7 +935,7 @@ This result implies that \( C(\beta) \) is a convex function since the matrix \(
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec25">Simple program </h2>
<h2 id="___sec30">Simple program </h2>
<p>
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
@@ -870,7 +976,7 @@ beta_NE = np.dot(Xt_X_inv,Xt_y)
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec26">Gradient Descent Example </h2>
<h2 id="___sec31">Gradient Descent Example </h2>
<p>
Another simple example is here
@@ -919,7 +1025,7 @@ plt.show()
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec27">And a corresponding example using <b>scikit-learn</b> </h2>
<h2 id="___sec32">And a corresponding example using <b>scikit-learn</b> </h2>
<p>
@@ -943,7 +1049,7 @@ sgdreg.fit(x,y.ravel())
<p>
<!-- !split -->
<h2 id="___sec28">Gradient descent and Ridge </h2>
<h2 id="___sec33">Gradient descent and Ridge </h2>
<p>
We have also discussed Ridge regression where the loss function contains a regularized given by the \( L_2 \) norm of \( \beta \),
@@ -1000,7 +1106,7 @@ beta_ridge = np.dot(Z,np.dot(X.T,y))
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec29">Stochastic Gradient Descent </h2>
<h2 id="___sec34">Stochastic Gradient Descent </h2>
<p>
Stochastic gradient descent (SGD) and variants thereof address some of
@@ -1018,7 +1124,7 @@ $$
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec30">Computation of gradients </h2>
<h2 id="___sec35">Computation of gradients </h2>
<p>
This in turn means that the gradient can be
@@ -1038,7 +1144,7 @@ minibatches. We denote these minibatches by \( B_k \) where
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec31">SGD example </h2>
<h2 id="___sec36">SGD example </h2>
As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \)
and we choose to have \( M=5 \) minibathces,
then each minibatch contains two data points. In particular we have
@@ -1062,7 +1168,7 @@ $$
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec32">The gradient step </h2>
<h2 id="___sec37">The gradient step </h2>
<p>
Thus a gradient descent step now looks like
@@ -1081,7 +1187,7 @@ the number of minibatches, as exemplified in the code below.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec33">Simple example code </h2>
<h2 id="___sec38">Simple example code </h2>
<p>
@@ -1113,7 +1219,7 @@ all \( n \) datapoints.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec34">When do we stop? </h2>
<h2 id="___sec39">When do we stop? </h2>
<p>
A natural question is when do we stop the search for a new minimum?
@@ -1130,7 +1236,7 @@ gave the lowest value.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec35">Slightly different approach </h2>
<h2 id="___sec40">Slightly different approach </h2>
<p>
Another approach is to let the step length \( \gamma_j \) depend on the
@@ -1175,324 +1281,6 @@ j = <span style="color: #B452CD">0</span>
<span style="color: #8B008B; font-weight: bold">print</span>(<span style="color: #CD5555">&quot;gamma_j after %d epochs: %g&quot;</span> % (n_epochs,gamma_j))
</pre></div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec36">Conjugate gradient (CG) method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
The success of the CG method for finding solutions of non-linear problems is based
on the theory of conjugate gradients for linear systems of equations. It belongs
to the class of iterative methods for solving problems from linear algebra of the type
$$
\begin{equation*}
\hat{A}\hat{x} = \hat{b}.
\end{equation*}
$$
In the iterative process we end up with a problem like
$$
\begin{equation*}
\hat{r}= \hat{b}-\hat{A}\hat{x},
\end{equation*}
$$
where \( \hat{r} \) is the so-called residual or error in the iterative process.
<p>
When we have found the exact solution, \( \hat{r}=0 \).
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec37">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
<p>
The residual is zero when we reach the minimum of the quadratic equation
$$
\begin{equation*}
P(\hat{x})=\frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T\hat{b},
\end{equation*}
$$
with the constraint that the matrix \( \hat{A} \) is positive definite and symmetric.
If we search for a minimum of the quantum mechanical variance, then the matrix
\( \hat{A} \), which is called the Hessian, is given by the second-derivative of the function we want to minimize. This quantity is always positive definite. In our case this corresponds normally to the second derivative of the energy.
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec38">Conjugate gradient method, Newton's method first </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
We seek the minimum of the energy or the variance as function of various variational parameters.
In our case we have thus a function \( f \) whose minimum we are seeking.
In Newton's method we set \( \nabla f = 0 \) and we can thus compute the next iteration point
$$
\begin{equation*}
\hat{x}-\hat{x}_i=\hat{A}^{-1}\nabla f(\hat{x}_i).
\end{equation*}
$$
Subtracting this equation from that of \( \hat{x}_{i+1} \) we have
$$
\begin{equation*}
\hat{x}_{i+1}-\hat{x}_i=\hat{A}^{-1}(\nabla f(\hat{x}_{i+1})-\nabla f(\hat{x}_i)).
\end{equation*}
$$
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec39">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
In the CG method we define so-called conjugate directions and two vectors
\( \hat{s} \) and \( \hat{t} \)
are said to be
conjugate if
$$
\begin{equation*}
\hat{s}^T\hat{A}\hat{t}= 0.
\end{equation*}
$$
The philosophy of the CG method is to perform searches in various conjugate directions
of our vectors \( \hat{x}_i \) obeying the above criterion, namely
$$
\begin{equation*}
\hat{x}_i^T\hat{A}\hat{x}_j= 0.
\end{equation*}
$$
Two vectors are conjugate if they are orthogonal with respect to
this inner product. Being conjugate is a symmetric relation: if \( \hat{s} \) is conjugate to \( \hat{t} \), then \( \hat{t} \) is conjugate to \( \hat{s} \).
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec40">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
An example is given by the eigenvectors of the matrix
$$
\begin{equation*}
\hat{v}_i^T\hat{A}\hat{v}_j= \lambda\hat{v}_i^T\hat{v}_j,
\end{equation*}
$$
which is zero unless \( i=j \).
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec41">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
Assume now that we have a symmetric positive-definite matrix \( \hat{A} \) of size
\( n\times n \). At each iteration \( i+1 \) we obtain the conjugate direction of a vector
$$
\begin{equation*}
\hat{x}_{i+1}=\hat{x}_{i}+\alpha_i\hat{p}_{i}.
\end{equation*}
$$
We assume that \( \hat{p}_{i} \) is a sequence of \( n \) mutually conjugate directions.
Then the \( \hat{p}_{i} \) form a basis of \( R^n \) and we can expand the solution
$ \hat{A}\hat{x} = \hat{b}$ in this basis, namely
$$
\begin{equation*}
\hat{x} = \sum^{n}_{i=1} \alpha_i \hat{p}_i.
\end{equation*}
$$
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec42">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
The coefficients are given by
$$
\begin{equation*}
\mathbf{A}\mathbf{x} = \sum^{n}_{i=1} \alpha_i \mathbf{A} \mathbf{p}_i = \mathbf{b}.
\end{equation*}
$$
Multiplying with \( \hat{p}_k^T \) from the left gives
$$
\begin{equation*}
\hat{p}_k^T \hat{A}\hat{x} = \sum^{n}_{i=1} \alpha_i\hat{p}_k^T \hat{A}\hat{p}_i= \hat{p}_k^T \hat{b},
\end{equation*}
$$
and we can define the coefficients \( \alpha_k \) as
$$
\begin{equation*}
\alpha_k = \frac{\hat{p}_k^T \hat{b}}{\hat{p}_k^T \hat{A} \hat{p}_k}
\end{equation*}
$$
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec43">Conjugate gradient method and iterations </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
<p>
If we choose the conjugate vectors \( \hat{p}_k \) carefully,
then we may not need all of them to obtain a good approximation to the solution
\( \hat{x} \).
We want to regard the conjugate gradient method as an iterative method.
This will us to solve systems where \( n \) is so large that the direct
method would take too much time.
<p>
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
We can assume without loss of generality that
$$
\begin{equation*}
\hat{x}_0=0,
\end{equation*}
$$
or consider the system
$$
\begin{equation*}
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
\end{equation*}
$$
instead.
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec44">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
$$
\begin{equation*}
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
\end{equation*}
$$
This suggests taking the first basis vector \( \hat{p}_1 \)
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
which equals
$$
\begin{equation*}
\hat{A}\hat{x}_0-\hat{b},
\end{equation*}
$$
and
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
The other vectors in the basis will be conjugate to the gradient,
hence the name conjugate gradient method.
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec45">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
$$
\begin{equation*}
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
\end{equation*}
$$
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
\( \hat{x}=\hat{x}_k \),
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
Here, we insist that the directions \( \hat{p}_k \) are conjugate to each other,
so we take the direction closest to the gradient \( \hat{r}_k \)
under the conjugacy constraint.
This gives the following expression
$$
\begin{equation*}
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
\end{equation*}
$$
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec46">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
We can also compute the residual iteratively as
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
\end{equation*}
$$
which equals
$$
\begin{equation*}
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
\end{equation*}
$$
or
$$
\begin{equation*}
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
\end{equation*}
$$
which gives
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
\end{equation*}
$$
</div>
<p>
<!-- ------------------- end of main content --------------- -->
+161 -373
View File
@@ -90,48 +90,39 @@ div { text-align: justify; text-justify: inter-word; }
('More on convex functions', 2, None, '___sec15'),
('Some simple problems', 2, None, '___sec16'),
('Standard steepest descent', 2, None, '___sec17'),
('Conjugate gradient method', 2, None, '___sec18'),
('Gradient method', 2, None, '___sec18'),
('Steepest descent method', 2, None, '___sec19'),
('Steepest descent method', 2, None, '___sec20'),
('Gradient descent method', 2, None, '___sec21'),
('Final expressions', 2, None, '___sec22'),
('The Steepest descent algorithm', 2, None, '___sec23'),
('Simple codes for steepest descent and conjugate gradient '
'using a $2\\times 2$ matrix, in c++, Python code to come',
2,
None,
'___sec19'),
'___sec24'),
('The routine for the steepest descent method',
2,
None,
'___sec20'),
('Revisiting our first homework', 2, None, '___sec21'),
('Gradient descent example', 2, None, '___sec22'),
('The derivative of the cost/loss function', 2, None, '___sec23'),
('The Hessian matrix', 2, None, '___sec24'),
('Simple program', 2, None, '___sec25'),
('Gradient Descent Example', 2, None, '___sec26'),
'___sec25'),
('Revisiting our first homework', 2, None, '___sec26'),
('Gradient descent example', 2, None, '___sec27'),
('The derivative of the cost/loss function', 2, None, '___sec28'),
('The Hessian matrix', 2, None, '___sec29'),
('Simple program', 2, None, '___sec30'),
('Gradient Descent Example', 2, None, '___sec31'),
('And a corresponding example using _scikit-learn_',
2,
None,
'___sec27'),
('Gradient descent and Ridge', 2, None, '___sec28'),
('Stochastic Gradient Descent', 2, None, '___sec29'),
('Computation of gradients', 2, None, '___sec30'),
('SGD example', 2, None, '___sec31'),
('The gradient step', 2, None, '___sec32'),
('Simple example code', 2, None, '___sec33'),
('When do we stop?', 2, None, '___sec34'),
('Slightly different approach', 2, None, '___sec35'),
('Conjugate gradient (CG) method', 2, None, '___sec36'),
('Conjugate gradient method', 2, None, '___sec37'),
("Conjugate gradient method, Newton's method first",
2,
None,
'___sec38'),
('Conjugate gradient method', 2, None, '___sec39'),
('Conjugate gradient method', 2, None, '___sec40'),
('Conjugate gradient method', 2, None, '___sec41'),
('Conjugate gradient method', 2, None, '___sec42'),
('Conjugate gradient method and iterations', 2, None, '___sec43'),
('Conjugate gradient method', 2, None, '___sec44'),
('Conjugate gradient method', 2, None, '___sec45'),
('Conjugate gradient method', 2, None, '___sec46')]}
'___sec32'),
('Gradient descent and Ridge', 2, None, '___sec33'),
('Stochastic Gradient Descent', 2, None, '___sec34'),
('Computation of gradients', 2, None, '___sec35'),
('SGD example', 2, None, '___sec36'),
('The gradient step', 2, None, '___sec37'),
('Simple example code', 2, None, '___sec38'),
('When do we stop?', 2, None, '___sec39'),
('Slightly different approach', 2, None, '___sec40')]}
end of tocinfo -->
<body>
@@ -644,7 +635,7 @@ When we have found the exact solution, \( \hat{r}=0 \).
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec18">Conjugate gradient method </h2>
<h2 id="___sec18">Gradient method </h2>
<p>
The residual is zero when we reach the minimum of the quadratic equation
@@ -662,12 +653,131 @@ given by the second-derivative of the function we want to minimize.
This quantity is always positive definite.
<p>
More details will be added here soon.
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec19">Steepest descent method </h2>
<p>
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
We can assume without loss of generality that
$$
\begin{equation*}
\hat{x}_0=0,
\end{equation*}
$$
or consider the system
$$
\begin{equation*}
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
\end{equation*}
$$
instead.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec19">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
<h2 id="___sec20">Steepest descent method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
$$
\begin{equation*}
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
\end{equation*}
$$
This suggests taking the first basis vector \( \hat{p}_1 \)
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
which equals
$$
\begin{equation*}
\hat{A}\hat{x}_0-\hat{b},
\end{equation*}
$$
and
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec21">Gradient descent method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
$$
\begin{equation*}
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
\end{equation*}
$$
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
\( \hat{x}=\hat{x}_k \),
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
This gives the following expression
$$
\begin{equation*}
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
\end{equation*}
$$
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec22">Final expressions </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
We can also compute the residual iteratively as
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
\end{equation*}
$$
which equals
$$
\begin{equation*}
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
\end{equation*}
$$
or
$$
\begin{equation*}
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
\end{equation*}
$$
which gives
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
\end{equation*}
$$
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec23">The Steepest descent algorithm </h2>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec24">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
@@ -694,11 +804,7 @@ More details will be added here soon.
cout <span style="color: #666666">&lt;&lt;</span> <span style="color: #BA2121">&quot;The Matrix A that we are using: &quot;</span> <span style="color: #666666">&lt;&lt;</span> endl;
A.Print();
cout <span style="color: #666666">&lt;&lt;</span> endl;
x <span style="color: #666666">=</span> ConjugateGradient(A,b,x0);
xsd <span style="color: #666666">=</span> SteepestDescent(A,b,x0);
cout <span style="color: #666666">&lt;&lt;</span> <span style="color: #BA2121">&quot;The approximate solution using Conjugate Gradient is: &quot;</span> <span style="color: #666666">&lt;&lt;</span> endl;
x.Print();
cout <span style="color: #666666">&lt;&lt;</span> endl;
cout <span style="color: #666666">&lt;&lt;</span> <span style="color: #BA2121">&quot;The approximate solution using Steepest Descent is: &quot;</span> <span style="color: #666666">&lt;&lt;</span> endl;
xsd.Print();
cout <span style="color: #666666">&lt;&lt;</span> endl;
@@ -711,7 +817,7 @@ More details will be added here soon.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec20">The routine for the steepest descent method </h2>
<h2 id="___sec25">The routine for the steepest descent method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
@@ -747,7 +853,7 @@ More details will be added here soon.
<p>
<!-- !split -->
<h2 id="___sec21">Revisiting our first homework </h2>
<h2 id="___sec26">Revisiting our first homework </h2>
<p>
We will use linear regression as a case study for the gradient descent
@@ -780,7 +886,7 @@ $$
<p>
<!-- !split -->
<h2 id="___sec22">Gradient descent example </h2>
<h2 id="___sec27">Gradient descent example </h2>
<p>
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
@@ -805,7 +911,7 @@ and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec23">The derivative of the cost/loss function </h2>
<h2 id="___sec28">The derivative of the cost/loss function </h2>
<p>
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
@@ -820,7 +926,7 @@ where \( X \) is the design matrix defined above.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec24">The Hessian matrix </h2>
<h2 id="___sec29">The Hessian matrix </h2>
The Hessian matrix of \( C(\beta) \) is given by
$$
\hat{H} \equiv \begin{bmatrix}
@@ -834,7 +940,7 @@ This result implies that \( C(\beta) \) is a convex function since the matrix \(
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec25">Simple program </h2>
<h2 id="___sec30">Simple program </h2>
<p>
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
@@ -875,7 +981,7 @@ beta_NE <span style="color: #666666">=</span> np<span style="color: #666666">.</
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec26">Gradient Descent Example </h2>
<h2 id="___sec31">Gradient Descent Example </h2>
<p>
Another simple example is here
@@ -924,7 +1030,7 @@ plt<span style="color: #666666">.</span>show()
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec27">And a corresponding example using <b>scikit-learn</b> </h2>
<h2 id="___sec32">And a corresponding example using <b>scikit-learn</b> </h2>
<p>
@@ -948,7 +1054,7 @@ sgdreg<span style="color: #666666">.</span>fit(x,y<span style="color: #666666">.
<p>
<!-- !split -->
<h2 id="___sec28">Gradient descent and Ridge </h2>
<h2 id="___sec33">Gradient descent and Ridge </h2>
<p>
We have also discussed Ridge regression where the loss function contains a regularized given by the \( L_2 \) norm of \( \beta \),
@@ -1005,7 +1111,7 @@ beta_ridge <span style="color: #666666">=</span> np<span style="color: #666666">
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec29">Stochastic Gradient Descent </h2>
<h2 id="___sec34">Stochastic Gradient Descent </h2>
<p>
Stochastic gradient descent (SGD) and variants thereof address some of
@@ -1023,7 +1129,7 @@ $$
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec30">Computation of gradients </h2>
<h2 id="___sec35">Computation of gradients </h2>
<p>
This in turn means that the gradient can be
@@ -1043,7 +1149,7 @@ minibatches. We denote these minibatches by \( B_k \) where
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec31">SGD example </h2>
<h2 id="___sec36">SGD example </h2>
As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \)
and we choose to have \( M=5 \) minibathces,
then each minibatch contains two data points. In particular we have
@@ -1067,7 +1173,7 @@ $$
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec32">The gradient step </h2>
<h2 id="___sec37">The gradient step </h2>
<p>
Thus a gradient descent step now looks like
@@ -1086,7 +1192,7 @@ the number of minibatches, as exemplified in the code below.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec33">Simple example code </h2>
<h2 id="___sec38">Simple example code </h2>
<p>
@@ -1118,7 +1224,7 @@ all \( n \) datapoints.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec34">When do we stop? </h2>
<h2 id="___sec39">When do we stop? </h2>
<p>
A natural question is when do we stop the search for a new minimum?
@@ -1135,7 +1241,7 @@ gave the lowest value.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec35">Slightly different approach </h2>
<h2 id="___sec40">Slightly different approach </h2>
<p>
Another approach is to let the step length \( \gamma_j \) depend on the
@@ -1180,324 +1286,6 @@ j <span style="color: #666666">=</span> <span style="color: #666666">0</span>
<span style="color: #008000; font-weight: bold">print</span>(<span style="color: #BA2121">&quot;gamma_j after </span><span style="color: #BB6688; font-weight: bold">%d</span><span style="color: #BA2121"> epochs: </span><span style="color: #BB6688; font-weight: bold">%g</span><span style="color: #BA2121">&quot;</span> <span style="color: #666666">%</span> (n_epochs,gamma_j))
</pre></div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec36">Conjugate gradient (CG) method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
The success of the CG method for finding solutions of non-linear problems is based
on the theory of conjugate gradients for linear systems of equations. It belongs
to the class of iterative methods for solving problems from linear algebra of the type
$$
\begin{equation*}
\hat{A}\hat{x} = \hat{b}.
\end{equation*}
$$
In the iterative process we end up with a problem like
$$
\begin{equation*}
\hat{r}= \hat{b}-\hat{A}\hat{x},
\end{equation*}
$$
where \( \hat{r} \) is the so-called residual or error in the iterative process.
<p>
When we have found the exact solution, \( \hat{r}=0 \).
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec37">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
<p>
The residual is zero when we reach the minimum of the quadratic equation
$$
\begin{equation*}
P(\hat{x})=\frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T\hat{b},
\end{equation*}
$$
with the constraint that the matrix \( \hat{A} \) is positive definite and symmetric.
If we search for a minimum of the quantum mechanical variance, then the matrix
\( \hat{A} \), which is called the Hessian, is given by the second-derivative of the function we want to minimize. This quantity is always positive definite. In our case this corresponds normally to the second derivative of the energy.
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec38">Conjugate gradient method, Newton's method first </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
We seek the minimum of the energy or the variance as function of various variational parameters.
In our case we have thus a function \( f \) whose minimum we are seeking.
In Newton's method we set \( \nabla f = 0 \) and we can thus compute the next iteration point
$$
\begin{equation*}
\hat{x}-\hat{x}_i=\hat{A}^{-1}\nabla f(\hat{x}_i).
\end{equation*}
$$
Subtracting this equation from that of \( \hat{x}_{i+1} \) we have
$$
\begin{equation*}
\hat{x}_{i+1}-\hat{x}_i=\hat{A}^{-1}(\nabla f(\hat{x}_{i+1})-\nabla f(\hat{x}_i)).
\end{equation*}
$$
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec39">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
In the CG method we define so-called conjugate directions and two vectors
\( \hat{s} \) and \( \hat{t} \)
are said to be
conjugate if
$$
\begin{equation*}
\hat{s}^T\hat{A}\hat{t}= 0.
\end{equation*}
$$
The philosophy of the CG method is to perform searches in various conjugate directions
of our vectors \( \hat{x}_i \) obeying the above criterion, namely
$$
\begin{equation*}
\hat{x}_i^T\hat{A}\hat{x}_j= 0.
\end{equation*}
$$
Two vectors are conjugate if they are orthogonal with respect to
this inner product. Being conjugate is a symmetric relation: if \( \hat{s} \) is conjugate to \( \hat{t} \), then \( \hat{t} \) is conjugate to \( \hat{s} \).
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec40">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
An example is given by the eigenvectors of the matrix
$$
\begin{equation*}
\hat{v}_i^T\hat{A}\hat{v}_j= \lambda\hat{v}_i^T\hat{v}_j,
\end{equation*}
$$
which is zero unless \( i=j \).
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec41">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
Assume now that we have a symmetric positive-definite matrix \( \hat{A} \) of size
\( n\times n \). At each iteration \( i+1 \) we obtain the conjugate direction of a vector
$$
\begin{equation*}
\hat{x}_{i+1}=\hat{x}_{i}+\alpha_i\hat{p}_{i}.
\end{equation*}
$$
We assume that \( \hat{p}_{i} \) is a sequence of \( n \) mutually conjugate directions.
Then the \( \hat{p}_{i} \) form a basis of \( R^n \) and we can expand the solution
$ \hat{A}\hat{x} = \hat{b}$ in this basis, namely
$$
\begin{equation*}
\hat{x} = \sum^{n}_{i=1} \alpha_i \hat{p}_i.
\end{equation*}
$$
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec42">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
The coefficients are given by
$$
\begin{equation*}
\mathbf{A}\mathbf{x} = \sum^{n}_{i=1} \alpha_i \mathbf{A} \mathbf{p}_i = \mathbf{b}.
\end{equation*}
$$
Multiplying with \( \hat{p}_k^T \) from the left gives
$$
\begin{equation*}
\hat{p}_k^T \hat{A}\hat{x} = \sum^{n}_{i=1} \alpha_i\hat{p}_k^T \hat{A}\hat{p}_i= \hat{p}_k^T \hat{b},
\end{equation*}
$$
and we can define the coefficients \( \alpha_k \) as
$$
\begin{equation*}
\alpha_k = \frac{\hat{p}_k^T \hat{b}}{\hat{p}_k^T \hat{A} \hat{p}_k}
\end{equation*}
$$
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec43">Conjugate gradient method and iterations </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
<p>
If we choose the conjugate vectors \( \hat{p}_k \) carefully,
then we may not need all of them to obtain a good approximation to the solution
\( \hat{x} \).
We want to regard the conjugate gradient method as an iterative method.
This will us to solve systems where \( n \) is so large that the direct
method would take too much time.
<p>
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
We can assume without loss of generality that
$$
\begin{equation*}
\hat{x}_0=0,
\end{equation*}
$$
or consider the system
$$
\begin{equation*}
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
\end{equation*}
$$
instead.
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec44">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
$$
\begin{equation*}
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
\end{equation*}
$$
This suggests taking the first basis vector \( \hat{p}_1 \)
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
which equals
$$
\begin{equation*}
\hat{A}\hat{x}_0-\hat{b},
\end{equation*}
$$
and
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
The other vectors in the basis will be conjugate to the gradient,
hence the name conjugate gradient method.
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec45">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
$$
\begin{equation*}
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
\end{equation*}
$$
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
\( \hat{x}=\hat{x}_k \),
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
Here, we insist that the directions \( \hat{p}_k \) are conjugate to each other,
so we take the direction closest to the gradient \( \hat{r}_k \)
under the conjugacy constraint.
This gives the following expression
$$
\begin{equation*}
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
\end{equation*}
$$
</div>
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec46">Conjugate gradient method </h2>
<div class="alert alert-block alert-block alert-text-normal">
<b></b>
<p>
We can also compute the residual iteratively as
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
\end{equation*}
$$
which equals
$$
\begin{equation*}
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
\end{equation*}
$$
or
$$
\begin{equation*}
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
\end{equation*}
$$
which gives
$$
\begin{equation*}
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
\end{equation*}
$$
</div>
<p>
<!-- ------------------- end of main content --------------- -->
+184 -454
View File
@@ -585,7 +585,7 @@
"\n",
"When we have found the exact solution, $\\hat{r}=0$.\n",
"\n",
"## Conjugate gradient method\n",
"## Gradient method\n",
"\n",
"The residual is zero when we reach the minimum of the quadratic equation"
]
@@ -609,7 +609,189 @@
"given by the second-derivative of the function we want to minimize.\n",
"This quantity is always positive definite. \n",
"\n",
"More details will be added here soon.\n",
"\n",
"## Steepest descent method\n",
"\n",
"We denote the initial guess for $\\hat{x}$ as $\\hat{x}_0$. \n",
"We can assume without loss of generality that"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{x}_0=0,\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"or consider the system"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{A}\\hat{z} = \\hat{b}-\\hat{A}\\hat{x}_0,\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"instead.\n",
"\n",
"\n",
"## Steepest descent method\n",
"One can show that the solution $\\hat{x}$ is also the unique minimizer of the quadratic form"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"f(\\hat{x}) = \\frac{1}{2}\\hat{x}^T\\hat{A}\\hat{x} - \\hat{x}^T \\hat{x} , \\quad \\hat{x}\\in\\mathbf{R}^n.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"This suggests taking the first basis vector $\\hat{p}_1$ \n",
"to be the gradient of $f$ at $\\hat{x}=\\hat{x}_0$, \n",
"which equals"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{A}\\hat{x}_0-\\hat{b},\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"and \n",
"$\\hat{x}_0=0$ it is equal $-\\hat{b}$.\n",
"\n",
"\n",
"\n",
"\n",
"## Gradient descent method\n",
"Let $\\hat{r}_k$ be the residual at the $k$-th step:"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{r}_k=\\hat{b}-\\hat{A}\\hat{x}_k.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Note that $\\hat{r}_k$ is the negative gradient of $f$ at \n",
"$\\hat{x}=\\hat{x}_k$, \n",
"so the gradient descent method would be to move in the direction $\\hat{r}_k$. \n",
"This gives the following expression"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{p}_{k+1}=\\hat{r}_k-\\frac{\\hat{p}_k^T \\hat{A}\\hat{r}_k}{\\hat{p}_k^T\\hat{A}\\hat{p}_k} \\hat{p}_k.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Final expressions\n",
"We can also compute the residual iteratively as"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{r}_{k+1}=\\hat{b}-\\hat{A}\\hat{x}_{k+1},\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"which equals"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{b}-\\hat{A}(\\hat{x}_k+\\alpha_k\\hat{p}_k),\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"or"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"(\\hat{b}-\\hat{A}\\hat{x}_k)-\\alpha_k\\hat{A}\\hat{p}_k,\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"which gives"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{r}_{k+1}=\\hat{r}_k-\\hat{A}\\hat{p}_{k},\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## The Steepest descent algorithm\n",
"\n",
"\n",
"## Simple codes for steepest descent and conjugate gradient using a $2\\times 2$ matrix, in c++, Python code to come"
]
@@ -638,11 +820,7 @@
" cout << \"The Matrix A that we are using: \" << endl;\n",
" A.Print();\n",
" cout << endl;\n",
" x = ConjugateGradient(A,b,x0);\n",
" xsd = SteepestDescent(A,b,x0);\n",
" cout << \"The approximate solution using Conjugate Gradient is: \" << endl;\n",
" x.Print();\n",
" cout << endl;\n",
" cout << \"The approximate solution using Steepest Descent is: \" << endl;\n",
" xsd.Print();\n",
" cout << endl;\n",
@@ -1289,454 +1467,6 @@
"\n",
"print(\"gamma_j after %d epochs: %g\" % (n_epochs,gamma_j))"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Conjugate gradient (CG) method\n",
"The success of the CG method for finding solutions of non-linear problems is based\n",
"on the theory of conjugate gradients for linear systems of equations. It belongs\n",
"to the class of iterative methods for solving problems from linear algebra of the type"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{A}\\hat{x} = \\hat{b}.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"In the iterative process we end up with a problem like"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{r}= \\hat{b}-\\hat{A}\\hat{x},\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"where $\\hat{r}$ is the so-called residual or error in the iterative process.\n",
"\n",
"When we have found the exact solution, $\\hat{r}=0$.\n",
"\n",
"\n",
"\n",
"\n",
"## Conjugate gradient method\n",
"\n",
"The residual is zero when we reach the minimum of the quadratic equation"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"P(\\hat{x})=\\frac{1}{2}\\hat{x}^T\\hat{A}\\hat{x} - \\hat{x}^T\\hat{b},\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"with the constraint that the matrix $\\hat{A}$ is positive definite and symmetric.\n",
"If we search for a minimum of the quantum mechanical variance, then the matrix \n",
"$\\hat{A}$, which is called the Hessian, is given by the second-derivative of the function we want to minimize. This quantity is always positive definite. In our case this corresponds normally to the second derivative of the energy.\n",
"\n",
"\n",
"\n",
"\n",
"\n",
"\n",
"\n",
"## Conjugate gradient method, Newton's method first\n",
"We seek the minimum of the energy or the variance as function of various variational parameters. \n",
"In our case we have thus a function $f$ whose minimum we are seeking.\n",
"In Newton's method we set $\\nabla f = 0$ and we can thus compute the next iteration point"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{x}-\\hat{x}_i=\\hat{A}^{-1}\\nabla f(\\hat{x}_i).\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Subtracting this equation from that of $\\hat{x}_{i+1}$ we have"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{x}_{i+1}-\\hat{x}_i=\\hat{A}^{-1}(\\nabla f(\\hat{x}_{i+1})-\\nabla f(\\hat{x}_i)).\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Conjugate gradient method\n",
"In the CG method we define so-called conjugate directions and two vectors \n",
"$\\hat{s}$ and $\\hat{t}$\n",
"are said to be\n",
"conjugate if"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{s}^T\\hat{A}\\hat{t}= 0.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"The philosophy of the CG method is to perform searches in various conjugate directions\n",
"of our vectors $\\hat{x}_i$ obeying the above criterion, namely"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{x}_i^T\\hat{A}\\hat{x}_j= 0.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Two vectors are conjugate if they are orthogonal with respect to \n",
"this inner product. Being conjugate is a symmetric relation: if $\\hat{s}$ is conjugate to $\\hat{t}$, then $\\hat{t}$ is conjugate to $\\hat{s}$.\n",
"\n",
"\n",
"\n",
"## Conjugate gradient method\n",
"An example is given by the eigenvectors of the matrix"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{v}_i^T\\hat{A}\\hat{v}_j= \\lambda\\hat{v}_i^T\\hat{v}_j,\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"which is zero unless $i=j$.\n",
"\n",
"\n",
"\n",
"\n",
"## Conjugate gradient method\n",
"Assume now that we have a symmetric positive-definite matrix $\\hat{A}$ of size\n",
"$n\\times n$. At each iteration $i+1$ we obtain the conjugate direction of a vector"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{x}_{i+1}=\\hat{x}_{i}+\\alpha_i\\hat{p}_{i}.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"We assume that $\\hat{p}_{i}$ is a sequence of $n$ mutually conjugate directions. \n",
"Then the $\\hat{p}_{i}$ form a basis of $R^n$ and we can expand the solution \n",
"$ \\hat{A}\\hat{x} = \\hat{b}$ in this basis, namely"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{x} = \\sum^{n}_{i=1} \\alpha_i \\hat{p}_i.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Conjugate gradient method\n",
"The coefficients are given by"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\mathbf{A}\\mathbf{x} = \\sum^{n}_{i=1} \\alpha_i \\mathbf{A} \\mathbf{p}_i = \\mathbf{b}.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Multiplying with $\\hat{p}_k^T$ from the left gives"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{p}_k^T \\hat{A}\\hat{x} = \\sum^{n}_{i=1} \\alpha_i\\hat{p}_k^T \\hat{A}\\hat{p}_i= \\hat{p}_k^T \\hat{b},\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"and we can define the coefficients $\\alpha_k$ as"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\alpha_k = \\frac{\\hat{p}_k^T \\hat{b}}{\\hat{p}_k^T \\hat{A} \\hat{p}_k}\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Conjugate gradient method and iterations\n",
"\n",
"If we choose the conjugate vectors $\\hat{p}_k$ carefully, \n",
"then we may not need all of them to obtain a good approximation to the solution \n",
"$\\hat{x}$. \n",
"We want to regard the conjugate gradient method as an iterative method. \n",
"This will us to solve systems where $n$ is so large that the direct \n",
"method would take too much time.\n",
"\n",
"We denote the initial guess for $\\hat{x}$ as $\\hat{x}_0$. \n",
"We can assume without loss of generality that"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{x}_0=0,\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"or consider the system"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{A}\\hat{z} = \\hat{b}-\\hat{A}\\hat{x}_0,\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"instead.\n",
"\n",
"\n",
"\n",
"\n",
"## Conjugate gradient method\n",
"One can show that the solution $\\hat{x}$ is also the unique minimizer of the quadratic form"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"f(\\hat{x}) = \\frac{1}{2}\\hat{x}^T\\hat{A}\\hat{x} - \\hat{x}^T \\hat{x} , \\quad \\hat{x}\\in\\mathbf{R}^n.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"This suggests taking the first basis vector $\\hat{p}_1$ \n",
"to be the gradient of $f$ at $\\hat{x}=\\hat{x}_0$, \n",
"which equals"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{A}\\hat{x}_0-\\hat{b},\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"and \n",
"$\\hat{x}_0=0$ it is equal $-\\hat{b}$.\n",
"The other vectors in the basis will be conjugate to the gradient, \n",
"hence the name conjugate gradient method.\n",
"\n",
"\n",
"\n",
"\n",
"## Conjugate gradient method\n",
"Let $\\hat{r}_k$ be the residual at the $k$-th step:"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{r}_k=\\hat{b}-\\hat{A}\\hat{x}_k.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Note that $\\hat{r}_k$ is the negative gradient of $f$ at \n",
"$\\hat{x}=\\hat{x}_k$, \n",
"so the gradient descent method would be to move in the direction $\\hat{r}_k$. \n",
"Here, we insist that the directions $\\hat{p}_k$ are conjugate to each other, \n",
"so we take the direction closest to the gradient $\\hat{r}_k$ \n",
"under the conjugacy constraint. \n",
"This gives the following expression"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{p}_{k+1}=\\hat{r}_k-\\frac{\\hat{p}_k^T \\hat{A}\\hat{r}_k}{\\hat{p}_k^T\\hat{A}\\hat{p}_k} \\hat{p}_k.\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Conjugate gradient method\n",
"We can also compute the residual iteratively as"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{r}_{k+1}=\\hat{b}-\\hat{A}\\hat{x}_{k+1},\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"which equals"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{b}-\\hat{A}(\\hat{x}_k+\\alpha_k\\hat{p}_k),\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"or"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"(\\hat{b}-\\hat{A}\\hat{x}_k)-\\alpha_k\\hat{A}\\hat{p}_k,\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"which gives"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"$$\n",
"\\hat{r}_{k+1}=\\hat{r}_k-\\hat{A}\\hat{p}_{k},\n",
"$$"
]
}
],
"metadata": {},
Binary file not shown.
Binary file not shown.
+99 -255
View File
@@ -405,7 +405,7 @@ where $\hat{r}$ is the so-called residual or error in the iterative process.
When we have found the exact solution, $\hat{r}=0$.
!split
===== Conjugate gradient method =====
===== Gradient method =====
The residual is zero when we reach the minimum of the quadratic equation
!bt
@@ -420,7 +420,104 @@ variance, then the matrix $\hat{A}$, which is called the Hessian, is
given by the second-derivative of the function we want to minimize.
This quantity is always positive definite.
More details will be added here soon.
!split
===== Steepest descent method =====
We denote the initial guess for $\hat{x}$ as $\hat{x}_0$.
We can assume without loss of generality that
!bt
\begin{equation*}
\hat{x}_0=0,
\end{equation*}
!et
or consider the system
!bt
\begin{equation*}
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
\end{equation*}
!et
instead.
!split
===== Steepest descent method =====
!bblock
One can show that the solution $\hat{x}$ is also the unique minimizer of the quadratic form
!bt
\begin{equation*}
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
\end{equation*}
!et
This suggests taking the first basis vector $\hat{p}_1$
to be the gradient of $f$ at $\hat{x}=\hat{x}_0$,
which equals
!bt
\begin{equation*}
\hat{A}\hat{x}_0-\hat{b},
\end{equation*}
!et
and
$\hat{x}_0=0$ it is equal $-\hat{b}$.
!eblock
!split
===== Gradient descent method =====
!bblock
Let $\hat{r}_k$ be the residual at the $k$-th step:
!bt
\begin{equation*}
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
\end{equation*}
!et
Note that $\hat{r}_k$ is the negative gradient of $f$ at
$\hat{x}=\hat{x}_k$,
so the gradient descent method would be to move in the direction $\hat{r}_k$.
This gives the following expression
!bt
\begin{equation*}
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
\end{equation*}
!et
!eblock
!split
===== Final expressions =====
!bblock
We can also compute the residual iteratively as
!bt
\begin{equation*}
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
\end{equation*}
!et
which equals
!bt
\begin{equation*}
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
\end{equation*}
!et
or
!bt
\begin{equation*}
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
\end{equation*}
!et
which gives
!bt
\begin{equation*}
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
\end{equation*}
!et
!eblock
!split
===== The Steepest descent algorithm =====
!split
===== Simple codes for steepest descent and conjugate gradient using a $2\times 2$ matrix, in c++, Python code to come =====
@@ -446,11 +543,7 @@ int main(int argc, char * argv[]){
cout << "The Matrix A that we are using: " << endl;
A.Print();
cout << endl;
x = ConjugateGradient(A,b,x0);
xsd = SteepestDescent(A,b,x0);
cout << "The approximate solution using Conjugate Gradient is: " << endl;
x.Print();
cout << endl;
cout << "The approximate solution using Steepest Descent is: " << endl;
xsd.Print();
cout << endl;
@@ -895,255 +988,6 @@ print("gamma_j after %d epochs: %g" % (n_epochs,gamma_j))
!split
===== Conjugate gradient (CG) method =====
!bblock
The success of the CG method for finding solutions of non-linear problems is based
on the theory of conjugate gradients for linear systems of equations. It belongs
to the class of iterative methods for solving problems from linear algebra of the type
!bt
\begin{equation*}
\hat{A}\hat{x} = \hat{b}.
\end{equation*}
!et
In the iterative process we end up with a problem like
!bt
\begin{equation*}
\hat{r}= \hat{b}-\hat{A}\hat{x},
\end{equation*}
!et
where $\hat{r}$ is the so-called residual or error in the iterative process.
When we have found the exact solution, $\hat{r}=0$.
!eblock
!split
===== Conjugate gradient method =====
!bblock
The residual is zero when we reach the minimum of the quadratic equation
!bt
\begin{equation*}
P(\hat{x})=\frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T\hat{b},
\end{equation*}
!et
with the constraint that the matrix $\hat{A}$ is positive definite and symmetric.
If we search for a minimum of the quantum mechanical variance, then the matrix
$\hat{A}$, which is called the Hessian, is given by the second-derivative of the function we want to minimize. This quantity is always positive definite. In our case this corresponds normally to the second derivative of the energy.
!eblock
!split
===== Conjugate gradient method, Newton's method first =====
!bblock
We seek the minimum of the energy or the variance as function of various variational parameters.
In our case we have thus a function $f$ whose minimum we are seeking.
In Newton's method we set $\nabla f = 0$ and we can thus compute the next iteration point
!bt
\begin{equation*}
\hat{x}-\hat{x}_i=\hat{A}^{-1}\nabla f(\hat{x}_i).
\end{equation*}
!et
Subtracting this equation from that of $\hat{x}_{i+1}$ we have
!bt
\begin{equation*}
\hat{x}_{i+1}-\hat{x}_i=\hat{A}^{-1}(\nabla f(\hat{x}_{i+1})-\nabla f(\hat{x}_i)).
\end{equation*}
!et
!eblock
!split
===== Conjugate gradient method =====
!bblock
In the CG method we define so-called conjugate directions and two vectors
$\hat{s}$ and $\hat{t}$
are said to be
conjugate if
!bt
\begin{equation*}
\hat{s}^T\hat{A}\hat{t}= 0.
\end{equation*}
!et
The philosophy of the CG method is to perform searches in various conjugate directions
of our vectors $\hat{x}_i$ obeying the above criterion, namely
!bt
\begin{equation*}
\hat{x}_i^T\hat{A}\hat{x}_j= 0.
\end{equation*}
!et
Two vectors are conjugate if they are orthogonal with respect to
this inner product. Being conjugate is a symmetric relation: if $\hat{s}$ is conjugate to $\hat{t}$, then $\hat{t}$ is conjugate to $\hat{s}$.
!eblock
!split
===== Conjugate gradient method =====
!bblock
An example is given by the eigenvectors of the matrix
!bt
\begin{equation*}
\hat{v}_i^T\hat{A}\hat{v}_j= \lambda\hat{v}_i^T\hat{v}_j,
\end{equation*}
!et
which is zero unless $i=j$.
!eblock
!split
===== Conjugate gradient method =====
!bblock
Assume now that we have a symmetric positive-definite matrix $\hat{A}$ of size
$n\times n$. At each iteration $i+1$ we obtain the conjugate direction of a vector
!bt
\begin{equation*}
\hat{x}_{i+1}=\hat{x}_{i}+\alpha_i\hat{p}_{i}.
\end{equation*}
!et
We assume that $\hat{p}_{i}$ is a sequence of $n$ mutually conjugate directions.
Then the $\hat{p}_{i}$ form a basis of $R^n$ and we can expand the solution
$ \hat{A}\hat{x} = \hat{b}$ in this basis, namely
!bt
\begin{equation*}
\hat{x} = \sum^{n}_{i=1} \alpha_i \hat{p}_i.
\end{equation*}
!et
!eblock
!split
===== Conjugate gradient method =====
!bblock
The coefficients are given by
!bt
\begin{equation*}
\mathbf{A}\mathbf{x} = \sum^{n}_{i=1} \alpha_i \mathbf{A} \mathbf{p}_i = \mathbf{b}.
\end{equation*}
!et
Multiplying with $\hat{p}_k^T$ from the left gives
!bt
\begin{equation*}
\hat{p}_k^T \hat{A}\hat{x} = \sum^{n}_{i=1} \alpha_i\hat{p}_k^T \hat{A}\hat{p}_i= \hat{p}_k^T \hat{b},
\end{equation*}
!et
and we can define the coefficients $\alpha_k$ as
!bt
\begin{equation*}
\alpha_k = \frac{\hat{p}_k^T \hat{b}}{\hat{p}_k^T \hat{A} \hat{p}_k}
\end{equation*}
!et
!eblock
!split
===== Conjugate gradient method and iterations =====
!bblock
If we choose the conjugate vectors $\hat{p}_k$ carefully,
then we may not need all of them to obtain a good approximation to the solution
$\hat{x}$.
We want to regard the conjugate gradient method as an iterative method.
This will us to solve systems where $n$ is so large that the direct
method would take too much time.
We denote the initial guess for $\hat{x}$ as $\hat{x}_0$.
We can assume without loss of generality that
!bt
\begin{equation*}
\hat{x}_0=0,
\end{equation*}
!et
or consider the system
!bt
\begin{equation*}
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
\end{equation*}
!et
instead.
!eblock
!split
===== Conjugate gradient method =====
!bblock
One can show that the solution $\hat{x}$ is also the unique minimizer of the quadratic form
!bt
\begin{equation*}
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
\end{equation*}
!et
This suggests taking the first basis vector $\hat{p}_1$
to be the gradient of $f$ at $\hat{x}=\hat{x}_0$,
which equals
!bt
\begin{equation*}
\hat{A}\hat{x}_0-\hat{b},
\end{equation*}
!et
and
$\hat{x}_0=0$ it is equal $-\hat{b}$.
The other vectors in the basis will be conjugate to the gradient,
hence the name conjugate gradient method.
!eblock
!split
===== Conjugate gradient method =====
!bblock
Let $\hat{r}_k$ be the residual at the $k$-th step:
!bt
\begin{equation*}
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
\end{equation*}
!et
Note that $\hat{r}_k$ is the negative gradient of $f$ at
$\hat{x}=\hat{x}_k$,
so the gradient descent method would be to move in the direction $\hat{r}_k$.
Here, we insist that the directions $\hat{p}_k$ are conjugate to each other,
so we take the direction closest to the gradient $\hat{r}_k$
under the conjugacy constraint.
This gives the following expression
!bt
\begin{equation*}
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
\end{equation*}
!et
!eblock
!split
===== Conjugate gradient method =====
!bblock
We can also compute the residual iteratively as
!bt
\begin{equation*}
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
\end{equation*}
!et
which equals
!bt
\begin{equation*}
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
\end{equation*}
!et
or
!bt
\begin{equation*}
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
\end{equation*}
!et
which gives
!bt
\begin{equation*}
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
\end{equation*}
!et
!eblock