more on steepest descent
This commit is contained in:
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -249,7 +234,7 @@ MathJax.Hub.Config({
|
||||
<li><a href="._Splines-bs008.html">9</a></li>
|
||||
<li><a href="._Splines-bs009.html">10</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs001.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -235,7 +220,7 @@ some approximative/numerical method to compute the minimum.
|
||||
<li><a href="._Splines-bs009.html">10</a></li>
|
||||
<li><a href="._Splines-bs010.html">11</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs002.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -243,7 +228,7 @@ where \( \hat{\beta} \) are the weights we wish to extract from data, in our cas
|
||||
<li><a href="._Splines-bs010.html">11</a></li>
|
||||
<li><a href="._Splines-bs011.html">12</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs003.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -248,7 +233,7 @@ This defines what we call the Hessian.
|
||||
<li><a href="._Splines-bs011.html">12</a></li>
|
||||
<li><a href="._Splines-bs012.html">13</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs004.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -249,7 +234,7 @@ If we can compute these matrices, in particular the Hessian, the above is often
|
||||
<li><a href="._Splines-bs012.html">13</a></li>
|
||||
<li><a href="._Splines-bs013.html">14</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs005.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -243,7 +228,7 @@ discourage the use of this method.
|
||||
<li><a href="._Splines-bs013.html">14</a></li>
|
||||
<li><a href="._Splines-bs014.html">15</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs006.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -261,7 +246,7 @@ $$
|
||||
<li><a href="._Splines-bs014.html">15</a></li>
|
||||
<li><a href="._Splines-bs015.html">16</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs007.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -244,7 +229,7 @@ vanishes, then Newton-Raphson may fail totally
|
||||
<li><a href="._Splines-bs015.html">16</a></li>
|
||||
<li><a href="._Splines-bs016.html">17</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs008.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -282,7 +267,7 @@ more than two non-linear equations. In our case, the Jacobian matrix is given by
|
||||
<li><a href="._Splines-bs016.html">17</a></li>
|
||||
<li><a href="._Splines-bs017.html">18</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs009.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -252,7 +237,7 @@ we are always moving towards smaller function values, i.e a minimum.
|
||||
<li><a href="._Splines-bs017.html">18</a></li>
|
||||
<li><a href="._Splines-bs018.html">19</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs010.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -248,7 +233,7 @@ the learning rate within the context of Machine Learning.
|
||||
<li><a href="._Splines-bs018.html">19</a></li>
|
||||
<li><a href="._Splines-bs019.html">20</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs011.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -255,7 +240,7 @@ Note that the gradient is a function of \( \mathbf{x} =
|
||||
<li><a href="._Splines-bs019.html">20</a></li>
|
||||
<li><a href="._Splines-bs020.html">21</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs012.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -248,7 +233,7 @@ randomness. One such method is that of Stochastic Gradient Descent
|
||||
<li><a href="._Splines-bs020.html">21</a></li>
|
||||
<li><a href="._Splines-bs021.html">22</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs013.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -249,7 +234,7 @@ regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
<li><a href="._Splines-bs021.html">22</a></li>
|
||||
<li><a href="._Splines-bs022.html">23</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs014.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -237,7 +222,7 @@ MathJax.Hub.Config({
|
||||
<li><a href="._Splines-bs022.html">23</a></li>
|
||||
<li><a href="._Splines-bs023.html">24</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs015.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -273,7 +258,7 @@ This condition is particularly useful since it gives us an procedure for determi
|
||||
<li><a href="._Splines-bs023.html">24</a></li>
|
||||
<li><a href="._Splines-bs024.html">25</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs016.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -260,7 +245,7 @@ This result means that if we know that the cost/loss function is convex and we a
|
||||
<li><a href="._Splines-bs024.html">25</a></li>
|
||||
<li><a href="._Splines-bs025.html">26</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs017.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -256,7 +241,7 @@ Using the definition of convexity, try to show that a function satisfying the pr
|
||||
<li><a href="._Splines-bs025.html">26</a></li>
|
||||
<li><a href="._Splines-bs026.html">27</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs018.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -263,7 +248,7 @@ When we have found the exact solution, \( \hat{r}=0 \).
|
||||
<li><a href="._Splines-bs026.html">27</a></li>
|
||||
<li><a href="._Splines-bs027.html">28</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs019.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -206,7 +191,7 @@ MathJax.Hub.Config({
|
||||
<a name="part0019"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec18" class="anchor">Conjugate gradient method </h2>
|
||||
<h2 id="___sec18" class="anchor">Gradient method </h2>
|
||||
|
||||
<p>
|
||||
The residual is zero when we reach the minimum of the quadratic equation
|
||||
@@ -223,9 +208,6 @@ variance, then the matrix \( \hat{A} \), which is called the Hessian, is
|
||||
given by the second-derivative of the function we want to minimize.
|
||||
This quantity is always positive definite.
|
||||
|
||||
<p>
|
||||
More details will be added here soon.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
@@ -252,7 +234,7 @@ More details will be added here soon.
|
||||
<li><a href="._Splines-bs027.html">28</a></li>
|
||||
<li><a href="._Splines-bs028.html">29</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs020.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -206,47 +191,25 @@ MathJax.Hub.Config({
|
||||
<a name="part0020"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec19" class="anchor">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
|
||||
<div class="panel panel-default">
|
||||
<div class="panel-body">
|
||||
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
|
||||
<h2 id="___sec19" class="anchor">Steepest descent method </h2>
|
||||
|
||||
<p>
|
||||
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
|
||||
We can assume without loss of generality that
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_0=0,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
<!-- code=c++ (!bc cppcod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic"><cmath></span><span style="color: #BC7A00"></span>
|
||||
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic"><iostream></span><span style="color: #BC7A00"></span>
|
||||
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic"><fstream></span><span style="color: #BC7A00"></span>
|
||||
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic"><iomanip></span><span style="color: #BC7A00"></span>
|
||||
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic">"vectormatrixclass.h"</span><span style="color: #BC7A00"></span>
|
||||
<span style="color: #008000; font-weight: bold">using</span> <span style="color: #008000; font-weight: bold">namespace</span> std;
|
||||
<span style="color: #408080; font-style: italic">// Main function begins here</span>
|
||||
<span style="color: #B00040">int</span> <span style="color: #0000FF">main</span>(<span style="color: #B00040">int</span> argc, <span style="color: #B00040">char</span> <span style="color: #666666">*</span> argv[]){
|
||||
<span style="color: #B00040">int</span> dim <span style="color: #666666">=</span> <span style="color: #666666">2</span>;
|
||||
Vector x(dim),xsd(dim), b(dim),x0(dim);
|
||||
Matrix A(dim,dim);
|
||||
|
||||
<span style="color: #408080; font-style: italic">// Set our initial guess</span>
|
||||
x0(<span style="color: #666666">0</span>) <span style="color: #666666">=</span> x0(<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">0</span>;
|
||||
<span style="color: #408080; font-style: italic">// Set the matrix</span>
|
||||
A(<span style="color: #666666">0</span>,<span style="color: #666666">0</span>) <span style="color: #666666">=</span> <span style="color: #666666">3</span>; A(<span style="color: #666666">1</span>,<span style="color: #666666">0</span>) <span style="color: #666666">=</span> <span style="color: #666666">2</span>; A(<span style="color: #666666">0</span>,<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">2</span>; A(<span style="color: #666666">1</span>,<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">6</span>;
|
||||
b(<span style="color: #666666">0</span>) <span style="color: #666666">=</span> <span style="color: #666666">2</span>; b(<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">-8</span>;
|
||||
cout <span style="color: #666666"><<</span> <span style="color: #BA2121">"The Matrix A that we are using: "</span> <span style="color: #666666"><<</span> endl;
|
||||
A.Print();
|
||||
cout <span style="color: #666666"><<</span> endl;
|
||||
x <span style="color: #666666">=</span> ConjugateGradient(A,b,x0);
|
||||
xsd <span style="color: #666666">=</span> SteepestDescent(A,b,x0);
|
||||
cout <span style="color: #666666"><<</span> <span style="color: #BA2121">"The approximate solution using Conjugate Gradient is: "</span> <span style="color: #666666"><<</span> endl;
|
||||
x.Print();
|
||||
cout <span style="color: #666666"><<</span> endl;
|
||||
cout <span style="color: #666666"><<</span> <span style="color: #BA2121">"The approximate solution using Steepest Descent is: "</span> <span style="color: #666666"><<</span> endl;
|
||||
xsd.Print();
|
||||
cout <span style="color: #666666"><<</span> endl;
|
||||
}
|
||||
</pre></div>
|
||||
<p>
|
||||
</div>
|
||||
</div>
|
||||
or consider the system
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
instead.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
@@ -274,7 +237,7 @@ MathJax.Hub.Config({
|
||||
<li><a href="._Splines-bs028.html">29</a></li>
|
||||
<li><a href="._Splines-bs029.html">30</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs021.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -206,35 +191,29 @@ MathJax.Hub.Config({
|
||||
<a name="part0021"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec20" class="anchor">The routine for the steepest descent method </h2>
|
||||
<h2 id="___sec20" class="anchor">Steepest descent method </h2>
|
||||
<div class="panel panel-default">
|
||||
<div class="panel-body">
|
||||
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
|
||||
<p>
|
||||
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
|
||||
$$
|
||||
\begin{equation*}
|
||||
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
This suggests taking the first basis vector \( \hat{p}_1 \)
|
||||
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
|
||||
which equals
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x}_0-\hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
and
|
||||
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
|
||||
|
||||
<!-- code=c++ (!bc cppcod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span>Vector <span style="color: #0000FF">SteepestDescent</span>(Matrix A, Vector b, Vector x0){
|
||||
<span style="color: #B00040">int</span> IterMax, i;
|
||||
<span style="color: #B00040">int</span> dim <span style="color: #666666">=</span> x0.Dimension();
|
||||
<span style="color: #008000; font-weight: bold">const</span> <span style="color: #B00040">double</span> tolerance <span style="color: #666666">=</span> <span style="color: #666666">1.0e-14</span>;
|
||||
Vector x(dim),f(dim),z(dim);
|
||||
<span style="color: #B00040">double</span> c,alpha,d;
|
||||
IterMax <span style="color: #666666">=</span> <span style="color: #666666">30</span>;
|
||||
x <span style="color: #666666">=</span> x0;
|
||||
f <span style="color: #666666">=</span> A<span style="color: #666666">*</span>x<span style="color: #666666">-</span>b;
|
||||
i <span style="color: #666666">=</span> <span style="color: #666666">0</span>;
|
||||
<span style="color: #008000; font-weight: bold">while</span> (i <span style="color: #666666"><=</span> IterMax){
|
||||
z <span style="color: #666666">=</span> A<span style="color: #666666">*</span>f;
|
||||
c <span style="color: #666666">=</span> dot(f,f);
|
||||
alpha <span style="color: #666666">=</span> c<span style="color: #666666">/</span>dot(f,z);
|
||||
x <span style="color: #666666">=</span> x <span style="color: #666666">-</span> alpha<span style="color: #666666">*</span>f;
|
||||
f <span style="color: #666666">=</span> A<span style="color: #666666">*</span>x<span style="color: #666666">-</span>b;
|
||||
<span style="color: #008000; font-weight: bold">if</span>(sqrt(dot(f,f)) <span style="color: #666666"><</span> tolerance) <span style="color: #008000; font-weight: bold">break</span>;
|
||||
i<span style="color: #666666">++</span>;
|
||||
}
|
||||
<span style="color: #008000; font-weight: bold">return</span> x;
|
||||
}
|
||||
</pre></div>
|
||||
<p>
|
||||
</div>
|
||||
</div>
|
||||
@@ -266,7 +245,7 @@ MathJax.Hub.Config({
|
||||
<li><a href="._Splines-bs029.html">30</a></li>
|
||||
<li><a href="._Splines-bs030.html">31</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs022.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -204,37 +189,31 @@ MathJax.Hub.Config({
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0022"></a>
|
||||
<!-- !split -->
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec21" class="anchor">Revisiting our first homework </h2>
|
||||
|
||||
<p>
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
<ol>
|
||||
<li> An analytical solution (recall homework set 1).</li>
|
||||
<li> The gradient can be computed analytically.</li>
|
||||
<li> The cost function is convex which guarantees that gradient descent converges for small enough learning rates</li>
|
||||
</ol>
|
||||
|
||||
We revisit the example from homework set 1 where we had
|
||||
<h2 id="___sec21" class="anchor">Gradient descent method </h2>
|
||||
<div class="panel panel-default">
|
||||
<div class="panel-body">
|
||||
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
|
||||
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
|
||||
$$
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
\begin{equation*}
|
||||
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
|
||||
The linear regression model is given by
|
||||
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
|
||||
\( \hat{x}=\hat{x}_k \),
|
||||
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
|
||||
This gives the following expression
|
||||
$$
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
\begin{equation*}
|
||||
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
</div>
|
||||
|
||||
such that
|
||||
$$
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<p>
|
||||
@@ -262,7 +241,7 @@ $$
|
||||
<li><a href="._Splines-bs030.html">31</a></li>
|
||||
<li><a href="._Splines-bs031.html">32</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs023.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -204,29 +189,43 @@ MathJax.Hub.Config({
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0023"></a>
|
||||
<!-- !split -->
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec22" class="anchor">Gradient descent example </h2>
|
||||
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
|
||||
<p>
|
||||
It is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by
|
||||
<h2 id="___sec22" class="anchor">Final expressions </h2>
|
||||
<div class="panel panel-default">
|
||||
<div class="panel-body">
|
||||
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
|
||||
We can also compute the residual iteratively as
|
||||
$$
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
The loss function is given by
|
||||
which equals
|
||||
$$
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
\begin{equation*}
|
||||
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
or
|
||||
$$
|
||||
\begin{equation*}
|
||||
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
which gives
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<p>
|
||||
@@ -254,7 +253,7 @@ and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
<li><a href="._Splines-bs031.html">32</a></li>
|
||||
<li><a href="._Splines-bs032.html">33</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs024.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -206,17 +191,7 @@ MathJax.Hub.Config({
|
||||
<a name="part0024"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec23" class="anchor">The derivative of the cost/loss function </h2>
|
||||
|
||||
<p>
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
$$
|
||||
\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
$$
|
||||
|
||||
where \( X \) is the design matrix defined above.
|
||||
<h2 id="___sec23" class="anchor">The Steepest descent algorithm </h2>
|
||||
|
||||
<p>
|
||||
<p>
|
||||
@@ -244,7 +219,7 @@ where \( X \) is the design matrix defined above.
|
||||
<li><a href="._Splines-bs032.html">33</a></li>
|
||||
<li><a href="._Splines-bs033.html">34</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs025.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -206,16 +191,43 @@ MathJax.Hub.Config({
|
||||
<a name="part0025"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec24" class="anchor">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
$$
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
$$
|
||||
<h2 id="___sec24" class="anchor">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
|
||||
<div class="panel panel-default">
|
||||
<div class="panel-body">
|
||||
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
|
||||
<p>
|
||||
|
||||
<!-- code=c++ (!bc cppcod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic"><cmath></span><span style="color: #BC7A00"></span>
|
||||
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic"><iostream></span><span style="color: #BC7A00"></span>
|
||||
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic"><fstream></span><span style="color: #BC7A00"></span>
|
||||
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic"><iomanip></span><span style="color: #BC7A00"></span>
|
||||
<span style="color: #BC7A00">#include</span> <span style="color: #408080; font-style: italic">"vectormatrixclass.h"</span><span style="color: #BC7A00"></span>
|
||||
<span style="color: #008000; font-weight: bold">using</span> <span style="color: #008000; font-weight: bold">namespace</span> std;
|
||||
<span style="color: #408080; font-style: italic">// Main function begins here</span>
|
||||
<span style="color: #B00040">int</span> <span style="color: #0000FF">main</span>(<span style="color: #B00040">int</span> argc, <span style="color: #B00040">char</span> <span style="color: #666666">*</span> argv[]){
|
||||
<span style="color: #B00040">int</span> dim <span style="color: #666666">=</span> <span style="color: #666666">2</span>;
|
||||
Vector x(dim),xsd(dim), b(dim),x0(dim);
|
||||
Matrix A(dim,dim);
|
||||
|
||||
<span style="color: #408080; font-style: italic">// Set our initial guess</span>
|
||||
x0(<span style="color: #666666">0</span>) <span style="color: #666666">=</span> x0(<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">0</span>;
|
||||
<span style="color: #408080; font-style: italic">// Set the matrix</span>
|
||||
A(<span style="color: #666666">0</span>,<span style="color: #666666">0</span>) <span style="color: #666666">=</span> <span style="color: #666666">3</span>; A(<span style="color: #666666">1</span>,<span style="color: #666666">0</span>) <span style="color: #666666">=</span> <span style="color: #666666">2</span>; A(<span style="color: #666666">0</span>,<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">2</span>; A(<span style="color: #666666">1</span>,<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">6</span>;
|
||||
b(<span style="color: #666666">0</span>) <span style="color: #666666">=</span> <span style="color: #666666">2</span>; b(<span style="color: #666666">1</span>) <span style="color: #666666">=</span> <span style="color: #666666">-8</span>;
|
||||
cout <span style="color: #666666"><<</span> <span style="color: #BA2121">"The Matrix A that we are using: "</span> <span style="color: #666666"><<</span> endl;
|
||||
A.Print();
|
||||
cout <span style="color: #666666"><<</span> endl;
|
||||
xsd <span style="color: #666666">=</span> SteepestDescent(A,b,x0);
|
||||
cout <span style="color: #666666"><<</span> <span style="color: #BA2121">"The approximate solution using Steepest Descent is: "</span> <span style="color: #666666"><<</span> endl;
|
||||
xsd.Print();
|
||||
cout <span style="color: #666666"><<</span> endl;
|
||||
}
|
||||
</pre></div>
|
||||
<p>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
@@ -243,7 +255,7 @@ This result implies that \( C(\beta) \) is a convex function since the matrix \(
|
||||
<li><a href="._Splines-bs033.html">34</a></li>
|
||||
<li><a href="._Splines-bs034.html">35</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs026.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -206,44 +191,40 @@ MathJax.Hub.Config({
|
||||
<a name="part0026"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec25" class="anchor">Simple program </h2>
|
||||
|
||||
<p>
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
$$
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
$$
|
||||
|
||||
<p>
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
|
||||
when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \).
|
||||
|
||||
<p>
|
||||
And finally we can compare our solution for \( \beta \) with the analytic result given by
|
||||
\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
|
||||
<h2 id="___sec25" class="anchor">The routine for the steepest descent method </h2>
|
||||
<div class="panel panel-default">
|
||||
<div class="panel-body">
|
||||
<p> <!-- subsequent paragraphs come in larger fonts, so start with a paragraph -->
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
<span style="color: #BA2121; font-style: italic">The following setup is just a suggestion, feel free to write it the way you like.</span>
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Setup problem described in the exercise</span>
|
||||
N <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#Nr of datapoints</span>
|
||||
M <span style="color: #666666">=</span> <span style="color: #666666">2</span> <span style="color: #408080; font-style: italic">#Nr of features</span>
|
||||
x <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(N) <span style="color: #408080; font-style: italic">#Uniformly generated x-values in [0,1]</span>
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">5*</span>x<span style="color: #666666">**2</span> <span style="color: #666666">+</span> <span style="color: #666666">0.1*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(N)
|
||||
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones(N),x] <span style="color: #408080; font-style: italic">#Construct design matrix</span>
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Compute beta according to normal equations to compare with GD solution</span>
|
||||
Xt_X_inv <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,X))
|
||||
Xt_y <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>transpose(),y)
|
||||
beta_NE <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(Xt_X_inv,Xt_y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_NE)
|
||||
<!-- code=c++ (!bc cppcod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span>Vector <span style="color: #0000FF">SteepestDescent</span>(Matrix A, Vector b, Vector x0){
|
||||
<span style="color: #B00040">int</span> IterMax, i;
|
||||
<span style="color: #B00040">int</span> dim <span style="color: #666666">=</span> x0.Dimension();
|
||||
<span style="color: #008000; font-weight: bold">const</span> <span style="color: #B00040">double</span> tolerance <span style="color: #666666">=</span> <span style="color: #666666">1.0e-14</span>;
|
||||
Vector x(dim),f(dim),z(dim);
|
||||
<span style="color: #B00040">double</span> c,alpha,d;
|
||||
IterMax <span style="color: #666666">=</span> <span style="color: #666666">30</span>;
|
||||
x <span style="color: #666666">=</span> x0;
|
||||
f <span style="color: #666666">=</span> A<span style="color: #666666">*</span>x<span style="color: #666666">-</span>b;
|
||||
i <span style="color: #666666">=</span> <span style="color: #666666">0</span>;
|
||||
<span style="color: #008000; font-weight: bold">while</span> (i <span style="color: #666666"><=</span> IterMax){
|
||||
z <span style="color: #666666">=</span> A<span style="color: #666666">*</span>f;
|
||||
c <span style="color: #666666">=</span> dot(f,f);
|
||||
alpha <span style="color: #666666">=</span> c<span style="color: #666666">/</span>dot(f,z);
|
||||
x <span style="color: #666666">=</span> x <span style="color: #666666">-</span> alpha<span style="color: #666666">*</span>f;
|
||||
f <span style="color: #666666">=</span> A<span style="color: #666666">*</span>x<span style="color: #666666">-</span>b;
|
||||
<span style="color: #008000; font-weight: bold">if</span>(sqrt(dot(f,f)) <span style="color: #666666"><</span> tolerance) <span style="color: #008000; font-weight: bold">break</span>;
|
||||
i<span style="color: #666666">++</span>;
|
||||
}
|
||||
<span style="color: #008000; font-weight: bold">return</span> x;
|
||||
}
|
||||
</pre></div>
|
||||
<p>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
@@ -270,7 +251,7 @@ beta_NE <span style="color: #666666">=</span> np<span style="color: #666666">.</
|
||||
<li><a href="._Splines-bs034.html">35</a></li>
|
||||
<li><a href="._Splines-bs035.html">36</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs027.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -204,54 +189,38 @@ MathJax.Hub.Config({
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0027"></a>
|
||||
<!-- !split -->
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec26" class="anchor">Gradient Descent Example </h2>
|
||||
<h2 id="___sec26" class="anchor">Revisiting our first homework </h2>
|
||||
|
||||
<p>
|
||||
Another simple example is here
|
||||
<p>
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Importing various packages</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">random</span> <span style="color: #008000; font-weight: bold">import</span> random, seed
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">mpl_toolkits.mplot3d</span> <span style="color: #008000; font-weight: bold">import</span> Axes3D
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">matplotlib</span> <span style="color: #008000; font-weight: bold">import</span> cm
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">matplotlib.ticker</span> <span style="color: #008000; font-weight: bold">import</span> LinearLocator, FormatStrFormatter
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">sys</span>
|
||||
<ol>
|
||||
<li> An analytical solution (recall homework set 1).</li>
|
||||
<li> The gradient can be computed analytically.</li>
|
||||
<li> The cost function is convex which guarantees that gradient descent converges for small enough learning rates</li>
|
||||
</ol>
|
||||
|
||||
x <span style="color: #666666">=</span> <span style="color: #666666">2*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">4+3*</span>x<span style="color: #666666">+</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
We revisit the example from homework set 1 where we had
|
||||
$$
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
$$
|
||||
|
||||
xb <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)), x]
|
||||
beta_linreg <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb))<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>T)<span style="color: #666666">.</span>dot(y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_linreg)
|
||||
beta <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">2</span>,<span style="color: #666666">1</span>)
|
||||
with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
|
||||
The linear regression model is given by
|
||||
$$
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
$$
|
||||
|
||||
eta <span style="color: #666666">=</span> <span style="color: #666666">0.1</span>
|
||||
Niterations <span style="color: #666666">=</span> <span style="color: #666666">1000</span>
|
||||
m <span style="color: #666666">=</span> <span style="color: #666666">100</span>
|
||||
such that
|
||||
$$
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
$$
|
||||
|
||||
<span style="color: #008000; font-weight: bold">for</span> <span style="color: #008000">iter</span> <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(Niterations):
|
||||
gradients <span style="color: #666666">=</span> <span style="color: #666666">2.0/</span>m<span style="color: #666666">*</span>xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>dot(beta)<span style="color: #666666">-</span>y)
|
||||
beta <span style="color: #666666">-=</span> eta<span style="color: #666666">*</span>gradients
|
||||
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta)
|
||||
xnew <span style="color: #666666">=</span> np<span style="color: #666666">.</span>array([[<span style="color: #666666">0</span>],[<span style="color: #666666">2</span>]])
|
||||
xbnew <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">2</span>,<span style="color: #666666">1</span>)), xnew]
|
||||
ypredict <span style="color: #666666">=</span> xbnew<span style="color: #666666">.</span>dot(beta)
|
||||
ypredict2 <span style="color: #666666">=</span> xbnew<span style="color: #666666">.</span>dot(beta_linreg)
|
||||
plt<span style="color: #666666">.</span>plot(xnew, ypredict, <span style="color: #BA2121">"r-"</span>)
|
||||
plt<span style="color: #666666">.</span>plot(xnew, ypredict2, <span style="color: #BA2121">"b-"</span>)
|
||||
plt<span style="color: #666666">.</span>plot(x, y ,<span style="color: #BA2121">'ro'</span>)
|
||||
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">0</span>,<span style="color: #666666">2.0</span>,<span style="color: #666666">0</span>, <span style="color: #666666">15.0</span>])
|
||||
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r'$x$'</span>)
|
||||
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r'$y$'</span>)
|
||||
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">r'Gradient descent example'</span>)
|
||||
plt<span style="color: #666666">.</span>show()
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
@@ -278,7 +247,7 @@ plt<span style="color: #666666">.</span>show()
|
||||
<li><a href="._Splines-bs035.html">36</a></li>
|
||||
<li><a href="._Splines-bs036.html">37</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs028.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -204,29 +189,30 @@ MathJax.Hub.Config({
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0028"></a>
|
||||
<!-- !split -->
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec27" class="anchor">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
<h2 id="___sec27" class="anchor">Gradient descent example </h2>
|
||||
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Importing various packages</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">random</span> <span style="color: #008000; font-weight: bold">import</span> random, seed
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.linear_model</span> <span style="color: #008000; font-weight: bold">import</span> SGDRegressor
|
||||
<p>
|
||||
It is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by
|
||||
$$
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
$$
|
||||
|
||||
x <span style="color: #666666">=</span> <span style="color: #666666">2*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">4+3*</span>x<span style="color: #666666">+</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
The loss function is given by
|
||||
$$
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
$$
|
||||
|
||||
and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
|
||||
xb <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)), x]
|
||||
beta_linreg <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb))<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>T)<span style="color: #666666">.</span>dot(y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_linreg)
|
||||
sgdreg <span style="color: #666666">=</span> SGDRegressor(n_iter <span style="color: #666666">=</span> <span style="color: #666666">50</span>, penalty<span style="color: #666666">=</span><span style="color: #008000">None</span>, eta0<span style="color: #666666">=0.1</span>)
|
||||
sgdreg<span style="color: #666666">.</span>fit(x,y<span style="color: #666666">.</span>ravel())
|
||||
<span style="color: #008000; font-weight: bold">print</span>(sgdreg<span style="color: #666666">.</span>intercept_, sgdreg<span style="color: #666666">.</span>coef_)
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
@@ -253,7 +239,7 @@ sgdreg<span style="color: #666666">.</span>fit(x,y<span style="color: #666666">.
|
||||
<li><a href="._Splines-bs036.html">37</a></li>
|
||||
<li><a href="._Splines-bs037.html">38</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs029.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -204,62 +189,20 @@ MathJax.Hub.Config({
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0029"></a>
|
||||
<!-- !split -->
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec28" class="anchor">Gradient descent and Ridge </h2>
|
||||
<h2 id="___sec28" class="anchor">The derivative of the cost/loss function </h2>
|
||||
|
||||
<p>
|
||||
We have also discussed Ridge regression where the loss function contains a regularized given by the \( L_2 \) norm of \( \beta \),
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
$$
|
||||
C_{\text{ridge}}(\beta) = ||X\beta -\mathbf{y}||^2 + \lambda ||\beta||^2, \ \lambda \geq 0.
|
||||
$$
|
||||
|
||||
<p>
|
||||
In order to minimize \( C_{\text{ridge}}(\beta) \) using GD we only have adjust the gradient as follows
|
||||
$$
|
||||
\nabla_\beta C_{\text{ridge}}(\beta) = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} + 2\lambda\begin{bmatrix} \beta_0 \\ \beta_1\end{bmatrix} = 2 (X^T(X\beta - \mathbf{y})+\lambda \beta).
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
$$
|
||||
|
||||
<p>
|
||||
We can now extend our program to minimize \( C_{\text{ridge}}(\beta) \) using gradient descent and compare with the analytical solution given by
|
||||
$$
|
||||
\beta_{\text{ridge}} = \left(X^T X + \lambda I_{2 \times 2} \right)^{-1} X^T \mathbf{y},
|
||||
$$
|
||||
where \( X \) is the design matrix defined above.
|
||||
|
||||
for \( \lambda = {0,1,10,50,100} \) (\( \lambda = 0 \) corresponds to ordinary least squares).
|
||||
We can then compute \( ||\beta_{\text{ridge}}|| \) for each \( \lambda \).
|
||||
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
<span style="color: #BA2121; font-style: italic">The following setup is just a suggestion, feel free to write it the way you like.</span>
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Setup problem described in the exercise</span>
|
||||
N <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#Nr of datapoints</span>
|
||||
M <span style="color: #666666">=</span> <span style="color: #666666">2</span> <span style="color: #408080; font-style: italic">#Nr of features</span>
|
||||
x <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(N)
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">5*</span>x<span style="color: #666666">**2</span> <span style="color: #666666">+</span> <span style="color: #666666">0.1*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(N)
|
||||
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Compute analytic beta for Ridge regression </span>
|
||||
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones(N),x]
|
||||
XT_X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,X)
|
||||
|
||||
l <span style="color: #666666">=</span> <span style="color: #666666">0.1</span> <span style="color: #408080; font-style: italic">#Ridge parameter lambda</span>
|
||||
Id <span style="color: #666666">=</span> np<span style="color: #666666">.</span>eye(XT_X<span style="color: #666666">.</span>shape[<span style="color: #666666">0</span>])
|
||||
|
||||
Z <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(XT_X<span style="color: #666666">+</span>l<span style="color: #666666">*</span>Id)
|
||||
beta_ridge <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(Z,np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,y))
|
||||
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_ridge)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>norm(beta_ridge)) <span style="color: #408080; font-style: italic">#||beta||</span>
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
@@ -286,7 +229,7 @@ beta_ridge <span style="color: #666666">=</span> np<span style="color: #666666">
|
||||
<li><a href="._Splines-bs037.html">38</a></li>
|
||||
<li><a href="._Splines-bs038.html">39</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs030.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -206,21 +191,17 @@ MathJax.Hub.Config({
|
||||
<a name="part0030"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec29" class="anchor">Stochastic Gradient Descent </h2>
|
||||
|
||||
<p>
|
||||
Stochastic gradient descent (SGD) and variants thereof address some of
|
||||
the shortcomings of the Gradient descent method discussed above.
|
||||
|
||||
<p>
|
||||
The underlying idea of SGD comes from the observation that the cost
|
||||
function, which we want to minimize, can almost always be written as a
|
||||
sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \),
|
||||
<h2 id="___sec29" class="anchor">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
$$
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}).
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
$$
|
||||
|
||||
This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
@@ -247,7 +228,7 @@ $$
|
||||
<li><a href="._Splines-bs038.html">39</a></li>
|
||||
<li><a href="._Splines-bs039.html">40</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs031.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -206,23 +191,44 @@ MathJax.Hub.Config({
|
||||
<a name="part0031"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec30" class="anchor">Computation of gradients </h2>
|
||||
<h2 id="___sec30" class="anchor">Simple program </h2>
|
||||
|
||||
<p>
|
||||
This in turn means that the gradient can be
|
||||
computed as a sum over \( i \)-gradients
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
$$
|
||||
\nabla_\beta C(\mathbf{\beta}) = \sum_i^n \nabla_\beta c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}).
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
$$
|
||||
|
||||
<p>
|
||||
Stochasticity/randomness is introduced by only taking the
|
||||
gradient on a subset of the data called minibatches. If there are \( n \)
|
||||
data points and the size of each minibatch is \( M \), there will be \( n/M \)
|
||||
minibatches. We denote these minibatches by \( B_k \) where
|
||||
\( k=1,\cdots,n/M \).
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
|
||||
when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \).
|
||||
|
||||
<p>
|
||||
And finally we can compare our solution for \( \beta \) with the analytic result given by
|
||||
\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
<span style="color: #BA2121; font-style: italic">The following setup is just a suggestion, feel free to write it the way you like.</span>
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Setup problem described in the exercise</span>
|
||||
N <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#Nr of datapoints</span>
|
||||
M <span style="color: #666666">=</span> <span style="color: #666666">2</span> <span style="color: #408080; font-style: italic">#Nr of features</span>
|
||||
x <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(N) <span style="color: #408080; font-style: italic">#Uniformly generated x-values in [0,1]</span>
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">5*</span>x<span style="color: #666666">**2</span> <span style="color: #666666">+</span> <span style="color: #666666">0.1*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(N)
|
||||
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones(N),x] <span style="color: #408080; font-style: italic">#Construct design matrix</span>
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Compute beta according to normal equations to compare with GD solution</span>
|
||||
Xt_X_inv <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,X))
|
||||
Xt_y <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>transpose(),y)
|
||||
beta_NE <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(Xt_X_inv,Xt_y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_NE)
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
@@ -249,7 +255,7 @@ minibatches. We denote these minibatches by \( B_k \) where
|
||||
<li><a href="._Splines-bs039.html">40</a></li>
|
||||
<li><a href="._Splines-bs040.html">41</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs032.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -206,27 +191,52 @@ MathJax.Hub.Config({
|
||||
<a name="part0032"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec31" class="anchor">SGD example </h2>
|
||||
As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \)
|
||||
and we choose to have \( M=5 \) minibathces,
|
||||
then each minibatch contains two data points. In particular we have
|
||||
\( B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 =
|
||||
(\mathbf{x}_9,\mathbf{x}_{10}) \). Note that if you choose \( M=1 \) you
|
||||
have only a single batch with all data points and on the other extreme,
|
||||
you may choose \( M=n \) resulting in a minibatch for each datapoint, i.e
|
||||
\( B_k = \mathbf{x}_k \).
|
||||
<h2 id="___sec31" class="anchor">Gradient Descent Example </h2>
|
||||
|
||||
<p>
|
||||
The idea is now to approximate the gradient by replacing the sum over
|
||||
all data points with a sum over the data points in one the minibatches
|
||||
picked at random in each gradient descent step
|
||||
$$
|
||||
\nabla_{\beta}
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta
|
||||
c_i(\mathbf{x}_i, \mathbf{\beta}).
|
||||
$$
|
||||
Another simple example is here
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Importing various packages</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">random</span> <span style="color: #008000; font-weight: bold">import</span> random, seed
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">mpl_toolkits.mplot3d</span> <span style="color: #008000; font-weight: bold">import</span> Axes3D
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">matplotlib</span> <span style="color: #008000; font-weight: bold">import</span> cm
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">matplotlib.ticker</span> <span style="color: #008000; font-weight: bold">import</span> LinearLocator, FormatStrFormatter
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">sys</span>
|
||||
|
||||
x <span style="color: #666666">=</span> <span style="color: #666666">2*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">4+3*</span>x<span style="color: #666666">+</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
|
||||
xb <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)), x]
|
||||
beta_linreg <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb))<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>T)<span style="color: #666666">.</span>dot(y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_linreg)
|
||||
beta <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">2</span>,<span style="color: #666666">1</span>)
|
||||
|
||||
eta <span style="color: #666666">=</span> <span style="color: #666666">0.1</span>
|
||||
Niterations <span style="color: #666666">=</span> <span style="color: #666666">1000</span>
|
||||
m <span style="color: #666666">=</span> <span style="color: #666666">100</span>
|
||||
|
||||
<span style="color: #008000; font-weight: bold">for</span> <span style="color: #008000">iter</span> <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(Niterations):
|
||||
gradients <span style="color: #666666">=</span> <span style="color: #666666">2.0/</span>m<span style="color: #666666">*</span>xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>dot(beta)<span style="color: #666666">-</span>y)
|
||||
beta <span style="color: #666666">-=</span> eta<span style="color: #666666">*</span>gradients
|
||||
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta)
|
||||
xnew <span style="color: #666666">=</span> np<span style="color: #666666">.</span>array([[<span style="color: #666666">0</span>],[<span style="color: #666666">2</span>]])
|
||||
xbnew <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">2</span>,<span style="color: #666666">1</span>)), xnew]
|
||||
ypredict <span style="color: #666666">=</span> xbnew<span style="color: #666666">.</span>dot(beta)
|
||||
ypredict2 <span style="color: #666666">=</span> xbnew<span style="color: #666666">.</span>dot(beta_linreg)
|
||||
plt<span style="color: #666666">.</span>plot(xnew, ypredict, <span style="color: #BA2121">"r-"</span>)
|
||||
plt<span style="color: #666666">.</span>plot(xnew, ypredict2, <span style="color: #BA2121">"b-"</span>)
|
||||
plt<span style="color: #666666">.</span>plot(x, y ,<span style="color: #BA2121">'ro'</span>)
|
||||
plt<span style="color: #666666">.</span>axis([<span style="color: #666666">0</span>,<span style="color: #666666">2.0</span>,<span style="color: #666666">0</span>, <span style="color: #666666">15.0</span>])
|
||||
plt<span style="color: #666666">.</span>xlabel(<span style="color: #BA2121">r'$x$'</span>)
|
||||
plt<span style="color: #666666">.</span>ylabel(<span style="color: #BA2121">r'$y$'</span>)
|
||||
plt<span style="color: #666666">.</span>title(<span style="color: #BA2121">r'Gradient descent example'</span>)
|
||||
plt<span style="color: #666666">.</span>show()
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
@@ -252,8 +262,6 @@ $$
|
||||
<li><a href="._Splines-bs039.html">40</a></li>
|
||||
<li><a href="._Splines-bs040.html">41</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs033.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -206,22 +191,27 @@ MathJax.Hub.Config({
|
||||
<a name="part0033"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec32" class="anchor">The gradient step </h2>
|
||||
<h2 id="___sec32" class="anchor">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
|
||||
<p>
|
||||
Thus a gradient descent step now looks like
|
||||
$$
|
||||
\beta_{j+1} = \beta_j - \gamma_j \sum_{i \in B_k}^n \nabla_\beta c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta})
|
||||
$$
|
||||
|
||||
<p>
|
||||
where \( k \) is picked at random with equal
|
||||
probability from \( [1,n/M] \). An iteration over the number of
|
||||
minibathces (n/M) is commonly referred to as an epoch. Thus it is
|
||||
typical to choose a number of epochs and for each epoch iterate over
|
||||
the number of minibatches, as exemplified in the code below.
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #408080; font-style: italic"># Importing various packages</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">random</span> <span style="color: #008000; font-weight: bold">import</span> random, seed
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
<span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">matplotlib.pyplot</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">plt</span>
|
||||
<span style="color: #008000; font-weight: bold">from</span> <span style="color: #0000FF; font-weight: bold">sklearn.linear_model</span> <span style="color: #008000; font-weight: bold">import</span> SGDRegressor
|
||||
|
||||
x <span style="color: #666666">=</span> <span style="color: #666666">2*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">4+3*</span>x<span style="color: #666666">+</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)
|
||||
|
||||
xb <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones((<span style="color: #666666">100</span>,<span style="color: #666666">1</span>)), x]
|
||||
beta_linreg <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(xb<span style="color: #666666">.</span>T<span style="color: #666666">.</span>dot(xb))<span style="color: #666666">.</span>dot(xb<span style="color: #666666">.</span>T)<span style="color: #666666">.</span>dot(y)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_linreg)
|
||||
sgdreg <span style="color: #666666">=</span> SGDRegressor(n_iter <span style="color: #666666">=</span> <span style="color: #666666">50</span>, penalty<span style="color: #666666">=</span><span style="color: #008000">None</span>, eta0<span style="color: #666666">=0.1</span>)
|
||||
sgdreg<span style="color: #666666">.</span>fit(x,y<span style="color: #666666">.</span>ravel())
|
||||
<span style="color: #008000; font-weight: bold">print</span>(sgdreg<span style="color: #666666">.</span>intercept_, sgdreg<span style="color: #666666">.</span>coef_)
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
@@ -246,9 +236,6 @@ the number of minibatches, as exemplified in the code below.
|
||||
<li><a href="._Splines-bs039.html">40</a></li>
|
||||
<li><a href="._Splines-bs040.html">41</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs042.html">43</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs034.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -204,37 +189,62 @@ MathJax.Hub.Config({
|
||||
<p> </p><p> </p><p> </p> <!-- add vertical space -->
|
||||
|
||||
<a name="part0034"></a>
|
||||
<!-- !split -->
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec33" class="anchor">Simple example code </h2>
|
||||
<h2 id="___sec33" class="anchor">Gradient descent and Ridge </h2>
|
||||
|
||||
<p>
|
||||
We have also discussed Ridge regression where the loss function contains a regularized given by the \( L_2 \) norm of \( \beta \),
|
||||
$$
|
||||
C_{\text{ridge}}(\beta) = ||X\beta -\mathbf{y}||^2 + \lambda ||\beta||^2, \ \lambda \geq 0.
|
||||
$$
|
||||
|
||||
<p>
|
||||
In order to minimize \( C_{\text{ridge}}(\beta) \) using GD we only have adjust the gradient as follows
|
||||
$$
|
||||
\nabla_\beta C_{\text{ridge}}(\beta) = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} + 2\lambda\begin{bmatrix} \beta_0 \\ \beta_1\end{bmatrix} = 2 (X^T(X\beta - \mathbf{y})+\lambda \beta).
|
||||
$$
|
||||
|
||||
<p>
|
||||
We can now extend our program to minimize \( C_{\text{ridge}}(\beta) \) using gradient descent and compare with the analytical solution given by
|
||||
$$
|
||||
\beta_{\text{ridge}} = \left(X^T X + \lambda I_{2 \times 2} \right)^{-1} X^T \mathbf{y},
|
||||
$$
|
||||
|
||||
for \( \lambda = {0,1,10,50,100} \) (\( \lambda = 0 \) corresponds to ordinary least squares).
|
||||
We can then compute \( ||\beta_{\text{ridge}}|| \) for each \( \lambda \).
|
||||
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
|
||||
n <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#100 datapoints </span>
|
||||
M <span style="color: #666666">=</span> <span style="color: #666666">5</span> <span style="color: #408080; font-style: italic">#size of each minibatch</span>
|
||||
m <span style="color: #666666">=</span> <span style="color: #008000">int</span>(n<span style="color: #666666">/</span>M) <span style="color: #408080; font-style: italic">#number of minibatches</span>
|
||||
n_epochs <span style="color: #666666">=</span> <span style="color: #666666">10</span> <span style="color: #408080; font-style: italic">#number of epochs</span>
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
<span style="color: #BA2121; font-style: italic">The following setup is just a suggestion, feel free to write it the way you like.</span>
|
||||
<span style="color: #BA2121; font-style: italic">"""</span>
|
||||
|
||||
j <span style="color: #666666">=</span> <span style="color: #666666">0</span>
|
||||
<span style="color: #008000; font-weight: bold">for</span> epoch <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(<span style="color: #666666">1</span>,n_epochs<span style="color: #666666">+1</span>):
|
||||
<span style="color: #008000; font-weight: bold">for</span> i <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(m):
|
||||
k <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randint(m) <span style="color: #408080; font-style: italic">#Pick the k-th minibatch at random</span>
|
||||
<span style="color: #408080; font-style: italic">#Compute the gradient using the data in minibatch Bk</span>
|
||||
<span style="color: #408080; font-style: italic">#Compute new suggestion for </span>
|
||||
j <span style="color: #666666">+=</span> <span style="color: #666666">1</span>
|
||||
<span style="color: #408080; font-style: italic">#Setup problem described in the exercise</span>
|
||||
N <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#Nr of datapoints</span>
|
||||
M <span style="color: #666666">=</span> <span style="color: #666666">2</span> <span style="color: #408080; font-style: italic">#Nr of features</span>
|
||||
x <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>rand(N)
|
||||
y <span style="color: #666666">=</span> <span style="color: #666666">5*</span>x<span style="color: #666666">**2</span> <span style="color: #666666">+</span> <span style="color: #666666">0.1*</span>np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randn(N)
|
||||
|
||||
|
||||
<span style="color: #408080; font-style: italic">#Compute analytic beta for Ridge regression </span>
|
||||
X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>c_[np<span style="color: #666666">.</span>ones(N),x]
|
||||
XT_X <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,X)
|
||||
|
||||
l <span style="color: #666666">=</span> <span style="color: #666666">0.1</span> <span style="color: #408080; font-style: italic">#Ridge parameter lambda</span>
|
||||
Id <span style="color: #666666">=</span> np<span style="color: #666666">.</span>eye(XT_X<span style="color: #666666">.</span>shape[<span style="color: #666666">0</span>])
|
||||
|
||||
Z <span style="color: #666666">=</span> np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>inv(XT_X<span style="color: #666666">+</span>l<span style="color: #666666">*</span>Id)
|
||||
beta_ridge <span style="color: #666666">=</span> np<span style="color: #666666">.</span>dot(Z,np<span style="color: #666666">.</span>dot(X<span style="color: #666666">.</span>T,y))
|
||||
|
||||
<span style="color: #008000; font-weight: bold">print</span>(beta_ridge)
|
||||
<span style="color: #008000; font-weight: bold">print</span>(np<span style="color: #666666">.</span>linalg<span style="color: #666666">.</span>norm(beta_ridge)) <span style="color: #408080; font-style: italic">#||beta||</span>
|
||||
</pre></div>
|
||||
<p>
|
||||
Taking the gradient only on a subset of the data has two important
|
||||
benefits. First, it introduces randomness which decreases the chance
|
||||
that our opmization scheme gets stuck in a local minima. Second, if
|
||||
the size of the minibatches are small relative to the number of
|
||||
datapoints (\( M < n \)), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the \( k-th \) minibatch and not
|
||||
all \( n \) datapoints.
|
||||
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
@@ -258,10 +268,6 @@ all \( n \) datapoints.
|
||||
<li><a href="._Splines-bs039.html">40</a></li>
|
||||
<li><a href="._Splines-bs040.html">41</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs042.html">43</a></li>
|
||||
<li><a href="._Splines-bs043.html">44</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs035.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -206,19 +191,20 @@ MathJax.Hub.Config({
|
||||
<a name="part0035"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec34" class="anchor">When do we stop? </h2>
|
||||
<h2 id="___sec34" class="anchor">Stochastic Gradient Descent </h2>
|
||||
|
||||
<p>
|
||||
A natural question is when do we stop the search for a new minimum?
|
||||
One possibility is to compute the full gradient after a given number
|
||||
of epochs and check if the norm of the gradient is smaller than some
|
||||
threshold and stop if true. However, the condition that the gradient
|
||||
is zero is valid also for local minima, so this would only tell us
|
||||
that we are close to a local/global minimum. However, we could also
|
||||
evaluate the cost function at this point, store the result and
|
||||
continue the search. If the test kicks in at a later stage we can
|
||||
compare the values of the cost function and keep the \( \beta \) that
|
||||
gave the lowest value.
|
||||
Stochastic gradient descent (SGD) and variants thereof address some of
|
||||
the shortcomings of the Gradient descent method discussed above.
|
||||
|
||||
<p>
|
||||
The underlying idea of SGD comes from the observation that the cost
|
||||
function, which we want to minimize, can almost always be written as a
|
||||
sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \),
|
||||
$$
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}).
|
||||
$$
|
||||
|
||||
<p>
|
||||
<p>
|
||||
@@ -242,11 +228,6 @@ gave the lowest value.
|
||||
<li><a href="._Splines-bs039.html">40</a></li>
|
||||
<li><a href="._Splines-bs040.html">41</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs042.html">43</a></li>
|
||||
<li><a href="._Splines-bs043.html">44</a></li>
|
||||
<li><a href="._Splines-bs044.html">45</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs036.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -206,51 +191,23 @@ MathJax.Hub.Config({
|
||||
<a name="part0036"></a>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec35" class="anchor">Slightly different approach </h2>
|
||||
<h2 id="___sec35" class="anchor">Computation of gradients </h2>
|
||||
|
||||
<p>
|
||||
Another approach is to let the step length \( \gamma_j \) depend on the
|
||||
number of epochs in such a way that it becomes very small after a
|
||||
reasonable time such that we do not move at all.
|
||||
This in turn means that the gradient can be
|
||||
computed as a sum over \( i \)-gradients
|
||||
$$
|
||||
\nabla_\beta C(\mathbf{\beta}) = \sum_i^n \nabla_\beta c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}).
|
||||
$$
|
||||
|
||||
<p>
|
||||
As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in <em>time</em> \( t \).
|
||||
Stochasticity/randomness is introduced by only taking the
|
||||
gradient on a subset of the data called minibatches. If there are \( n \)
|
||||
data points and the size of each minibatch is \( M \), there will be \( n/M \)
|
||||
minibatches. We denote these minibatches by \( B_k \) where
|
||||
\( k=1,\cdots,n/M \).
|
||||
|
||||
<p>
|
||||
In this way we can fix the number of epochs, compute \( \beta \) and
|
||||
evaluate the cost function at the end. Repeating the computation will
|
||||
give a different result since the scheme is random by design. Then we
|
||||
pick the final \( \beta \) that gives the lowest value of the cost
|
||||
function.
|
||||
|
||||
<p>
|
||||
|
||||
<!-- code=python (!bc pycod) typeset with pygments style "default" -->
|
||||
<div class="highlight" style="background: #f8f8f8"><pre style="line-height: 125%"><span></span><span style="color: #008000; font-weight: bold">import</span> <span style="color: #0000FF; font-weight: bold">numpy</span> <span style="color: #008000; font-weight: bold">as</span> <span style="color: #0000FF; font-weight: bold">np</span>
|
||||
|
||||
<span style="color: #008000; font-weight: bold">def</span> <span style="color: #0000FF">step_length</span>(t,t0,t1):
|
||||
<span style="color: #008000; font-weight: bold">return</span> t0<span style="color: #666666">/</span>(t<span style="color: #666666">+</span>t1)
|
||||
|
||||
n <span style="color: #666666">=</span> <span style="color: #666666">100</span> <span style="color: #408080; font-style: italic">#100 datapoints </span>
|
||||
M <span style="color: #666666">=</span> <span style="color: #666666">5</span> <span style="color: #408080; font-style: italic">#size of each minibatch</span>
|
||||
m <span style="color: #666666">=</span> <span style="color: #008000">int</span>(n<span style="color: #666666">/</span>M) <span style="color: #408080; font-style: italic">#number of minibatches</span>
|
||||
n_epochs <span style="color: #666666">=</span> <span style="color: #666666">500</span> <span style="color: #408080; font-style: italic">#number of epochs</span>
|
||||
t0 <span style="color: #666666">=</span> <span style="color: #666666">1.0</span>
|
||||
t1 <span style="color: #666666">=</span> <span style="color: #666666">10</span>
|
||||
|
||||
gamma_j <span style="color: #666666">=</span> t0<span style="color: #666666">/</span>t1
|
||||
j <span style="color: #666666">=</span> <span style="color: #666666">0</span>
|
||||
<span style="color: #008000; font-weight: bold">for</span> epoch <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(<span style="color: #666666">1</span>,n_epochs<span style="color: #666666">+1</span>):
|
||||
<span style="color: #008000; font-weight: bold">for</span> i <span style="color: #AA22FF; font-weight: bold">in</span> <span style="color: #008000">range</span>(m):
|
||||
k <span style="color: #666666">=</span> np<span style="color: #666666">.</span>random<span style="color: #666666">.</span>randint(m) <span style="color: #408080; font-style: italic">#Pick the k-th minibatch at random</span>
|
||||
<span style="color: #408080; font-style: italic">#Compute the gradient using the data in minibatch Bk</span>
|
||||
<span style="color: #408080; font-style: italic">#Compute new suggestion for beta</span>
|
||||
t <span style="color: #666666">=</span> epoch<span style="color: #666666">*</span>m<span style="color: #666666">+</span>i
|
||||
gamma_j <span style="color: #666666">=</span> step_length(t,t0,t1)
|
||||
j <span style="color: #666666">+=</span> <span style="color: #666666">1</span>
|
||||
|
||||
<span style="color: #008000; font-weight: bold">print</span>(<span style="color: #BA2121">"gamma_j after </span><span style="color: #BB6688; font-weight: bold">%d</span><span style="color: #BA2121"> epochs: </span><span style="color: #BB6688; font-weight: bold">%g</span><span style="color: #BA2121">"</span> <span style="color: #666666">%</span> (n_epochs,gamma_j))
|
||||
</pre></div>
|
||||
<p>
|
||||
<p>
|
||||
<!-- navigation buttons at the bottom of the page -->
|
||||
@@ -272,12 +229,6 @@ j <span style="color: #666666">=</span> <span style="color: #666666">0</span>
|
||||
<li><a href="._Splines-bs039.html">40</a></li>
|
||||
<li><a href="._Splines-bs040.html">41</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs042.html">43</a></li>
|
||||
<li><a href="._Splines-bs043.html">44</a></li>
|
||||
<li><a href="._Splines-bs044.html">45</a></li>
|
||||
<li><a href="._Splines-bs045.html">46</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs037.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -65,48 +65,39 @@ Automatically generated HTML file from DocOnce source
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -162,35 +153,29 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs016.html#___sec15" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs017.html#___sec16" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs018.html#___sec17" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">Conjugate gradient (CG) method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Conjugate gradient method, Newton's method first</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs042.html#___sec41" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs043.html#___sec42" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs044.html#___sec43" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs045.html#___sec44" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs046.html#___sec45" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs047.html#___sec46" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs019.html#___sec18" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs020.html#___sec19" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs021.html#___sec20" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs022.html#___sec21" style="font-size: 80%;">Gradient descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs023.html#___sec22" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs024.html#___sec23" style="font-size: 80%;">The Steepest descent algorithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs025.html#___sec24" style="font-size: 80%;">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs026.html#___sec25" style="font-size: 80%;">The routine for the steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs027.html#___sec26" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs028.html#___sec27" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs029.html#___sec28" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs030.html#___sec29" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs031.html#___sec30" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs032.html#___sec31" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs033.html#___sec32" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs034.html#___sec33" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs035.html#___sec34" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs036.html#___sec35" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs037.html#___sec36" style="font-size: 80%;">SGD example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs038.html#___sec37" style="font-size: 80%;">The gradient step</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs039.html#___sec38" style="font-size: 80%;">Simple example code</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs040.html#___sec39" style="font-size: 80%;">When do we stop?</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._Splines-bs041.html#___sec40" style="font-size: 80%;">Slightly different approach</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -249,7 +234,7 @@ MathJax.Hub.Config({
|
||||
<li><a href="._Splines-bs008.html">9</a></li>
|
||||
<li><a href="._Splines-bs009.html">10</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._Splines-bs047.html">48</a></li>
|
||||
<li><a href="._Splines-bs041.html">42</a></li>
|
||||
<li><a href="._Splines-bs001.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -662,7 +662,7 @@ When we have found the exact solution, \( \hat{r}=0 \).
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec18">Conjugate gradient method </h2>
|
||||
<h2 id="___sec18">Gradient method </h2>
|
||||
|
||||
<p>
|
||||
The residual is zero when we reach the minimum of the quadratic equation
|
||||
@@ -680,14 +680,150 @@ symmetric. If we search for a minimum of the quantum mechanical
|
||||
variance, then the matrix \( \hat{A} \), which is called the Hessian, is
|
||||
given by the second-derivative of the function we want to minimize.
|
||||
This quantity is always positive definite.
|
||||
|
||||
<p>
|
||||
More details will be added here soon.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec19">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
|
||||
<h2 id="___sec19">Steepest descent method </h2>
|
||||
|
||||
<p>
|
||||
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
|
||||
We can assume without loss of generality that
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_0=0,
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
or consider the system
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
instead.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec20">Steepest descent method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
This suggests taking the first basis vector \( \hat{p}_1 \)
|
||||
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
|
||||
which equals
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x}_0-\hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
and
|
||||
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
|
||||
|
||||
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec21">Gradient descent method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
|
||||
\( \hat{x}=\hat{x}_k \),
|
||||
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
|
||||
This gives the following expression
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec22">Final expressions </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
We can also compute the residual iteratively as
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
which equals
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
or
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
which gives
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec23">The Steepest descent algorithm </h2>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec24">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
@@ -713,11 +849,7 @@ More details will be added here soon.
|
||||
cout << <span style="color: #CD5555">"The Matrix A that we are using: "</span> << endl;
|
||||
A.Print();
|
||||
cout << endl;
|
||||
x = ConjugateGradient(A,b,x0);
|
||||
xsd = SteepestDescent(A,b,x0);
|
||||
cout << <span style="color: #CD5555">"The approximate solution using Conjugate Gradient is: "</span> << endl;
|
||||
x.Print();
|
||||
cout << endl;
|
||||
cout << <span style="color: #CD5555">"The approximate solution using Steepest Descent is: "</span> << endl;
|
||||
xsd.Print();
|
||||
cout << endl;
|
||||
@@ -729,7 +861,7 @@ More details will be added here soon.
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec20">The routine for the steepest descent method </h2>
|
||||
<h2 id="___sec25">The routine for the steepest descent method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
@@ -763,7 +895,7 @@ More details will be added here soon.
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec21">Revisiting our first homework </h2>
|
||||
<h2 id="___sec26">Revisiting our first homework </h2>
|
||||
|
||||
<p>
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
@@ -803,7 +935,7 @@ $$
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec22">Gradient descent example </h2>
|
||||
<h2 id="___sec27">Gradient descent example </h2>
|
||||
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
@@ -832,7 +964,7 @@ and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec23">The derivative of the cost/loss function </h2>
|
||||
<h2 id="___sec28">The derivative of the cost/loss function </h2>
|
||||
|
||||
<p>
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
@@ -849,7 +981,7 @@ where \( X \) is the design matrix defined above.
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec24">The Hessian matrix </h2>
|
||||
<h2 id="___sec29">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
<p> <br>
|
||||
$$
|
||||
@@ -865,7 +997,7 @@ This result implies that \( C(\beta) \) is a convex function since the matrix \(
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec25">Simple program </h2>
|
||||
<h2 id="___sec30">Simple program </h2>
|
||||
|
||||
<p>
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
@@ -909,7 +1041,7 @@ beta_NE = np.dot(Xt_X_inv,Xt_y)
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec26">Gradient Descent Example </h2>
|
||||
<h2 id="___sec31">Gradient Descent Example </h2>
|
||||
|
||||
<p>
|
||||
Another simple example is here
|
||||
@@ -959,7 +1091,7 @@ plt.show()
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec27">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
<h2 id="___sec32">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
|
||||
<p>
|
||||
|
||||
@@ -984,7 +1116,7 @@ sgdreg.fit(x,y.ravel())
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec28">Gradient descent and Ridge </h2>
|
||||
<h2 id="___sec33">Gradient descent and Ridge </h2>
|
||||
|
||||
<p>
|
||||
We have also discussed Ridge regression where the loss function contains a regularized given by the \( L_2 \) norm of \( \beta \),
|
||||
@@ -1048,7 +1180,7 @@ beta_ridge = np.dot(Z,np.dot(X.T,y))
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec29">Stochastic Gradient Descent </h2>
|
||||
<h2 id="___sec34">Stochastic Gradient Descent </h2>
|
||||
|
||||
<p>
|
||||
Stochastic gradient descent (SGD) and variants thereof address some of
|
||||
@@ -1068,7 +1200,7 @@ $$
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec30">Computation of gradients </h2>
|
||||
<h2 id="___sec35">Computation of gradients </h2>
|
||||
|
||||
<p>
|
||||
This in turn means that the gradient can be
|
||||
@@ -1090,7 +1222,7 @@ minibatches. We denote these minibatches by \( B_k \) where
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec31">SGD example </h2>
|
||||
<h2 id="___sec36">SGD example </h2>
|
||||
As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \)
|
||||
and we choose to have \( M=5 \) minibathces,
|
||||
then each minibatch contains two data points. In particular we have
|
||||
@@ -1116,7 +1248,7 @@ $$
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec32">The gradient step </h2>
|
||||
<h2 id="___sec37">The gradient step </h2>
|
||||
|
||||
<p>
|
||||
Thus a gradient descent step now looks like
|
||||
@@ -1137,7 +1269,7 @@ the number of minibatches, as exemplified in the code below.
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec33">Simple example code </h2>
|
||||
<h2 id="___sec38">Simple example code </h2>
|
||||
|
||||
<p>
|
||||
|
||||
@@ -1169,7 +1301,7 @@ all \( n \) datapoints.
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec34">When do we stop? </h2>
|
||||
<h2 id="___sec39">When do we stop? </h2>
|
||||
|
||||
<p>
|
||||
A natural question is when do we stop the search for a new minimum?
|
||||
@@ -1186,7 +1318,7 @@ gave the lowest value.
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec35">Slightly different approach </h2>
|
||||
<h2 id="___sec40">Slightly different approach </h2>
|
||||
|
||||
<p>
|
||||
Another approach is to let the step length \( \gamma_j \) depend on the
|
||||
@@ -1236,355 +1368,6 @@ j = <span style="color: #B452CD">0</span>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec36">Conjugate gradient (CG) method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
The success of the CG method for finding solutions of non-linear problems is based
|
||||
on the theory of conjugate gradients for linear systems of equations. It belongs
|
||||
to the class of iterative methods for solving problems from linear algebra of the type
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x} = \hat{b}.
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
In the iterative process we end up with a problem like
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}= \hat{b}-\hat{A}\hat{x},
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
where \( \hat{r} \) is the so-called residual or error in the iterative process.
|
||||
|
||||
<p>
|
||||
When we have found the exact solution, \( \hat{r}=0 \).
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec37">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
The residual is zero when we reach the minimum of the quadratic equation
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
P(\hat{x})=\frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T\hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
with the constraint that the matrix \( \hat{A} \) is positive definite and symmetric.
|
||||
If we search for a minimum of the quantum mechanical variance, then the matrix
|
||||
\( \hat{A} \), which is called the Hessian, is given by the second-derivative of the function we want to minimize. This quantity is always positive definite. In our case this corresponds normally to the second derivative of the energy.
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec38">Conjugate gradient method, Newton's method first </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
We seek the minimum of the energy or the variance as function of various variational parameters.
|
||||
In our case we have thus a function \( f \) whose minimum we are seeking.
|
||||
In Newton's method we set \( \nabla f = 0 \) and we can thus compute the next iteration point
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}-\hat{x}_i=\hat{A}^{-1}\nabla f(\hat{x}_i).
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
Subtracting this equation from that of \( \hat{x}_{i+1} \) we have
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_{i+1}-\hat{x}_i=\hat{A}^{-1}(\nabla f(\hat{x}_{i+1})-\nabla f(\hat{x}_i)).
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec39">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
In the CG method we define so-called conjugate directions and two vectors
|
||||
\( \hat{s} \) and \( \hat{t} \)
|
||||
are said to be
|
||||
conjugate if
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{s}^T\hat{A}\hat{t}= 0.
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
The philosophy of the CG method is to perform searches in various conjugate directions
|
||||
of our vectors \( \hat{x}_i \) obeying the above criterion, namely
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_i^T\hat{A}\hat{x}_j= 0.
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
Two vectors are conjugate if they are orthogonal with respect to
|
||||
this inner product. Being conjugate is a symmetric relation: if \( \hat{s} \) is conjugate to \( \hat{t} \), then \( \hat{t} \) is conjugate to \( \hat{s} \).
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec40">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
An example is given by the eigenvectors of the matrix
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{v}_i^T\hat{A}\hat{v}_j= \lambda\hat{v}_i^T\hat{v}_j,
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
which is zero unless \( i=j \).
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec41">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
Assume now that we have a symmetric positive-definite matrix \( \hat{A} \) of size
|
||||
\( n\times n \). At each iteration \( i+1 \) we obtain the conjugate direction of a vector
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_{i+1}=\hat{x}_{i}+\alpha_i\hat{p}_{i}.
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
We assume that \( \hat{p}_{i} \) is a sequence of \( n \) mutually conjugate directions.
|
||||
Then the \( \hat{p}_{i} \) form a basis of \( R^n \) and we can expand the solution
|
||||
$ \hat{A}\hat{x} = \hat{b}$ in this basis, namely
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x} = \sum^{n}_{i=1} \alpha_i \hat{p}_i.
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec42">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
The coefficients are given by
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\mathbf{A}\mathbf{x} = \sum^{n}_{i=1} \alpha_i \mathbf{A} \mathbf{p}_i = \mathbf{b}.
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
Multiplying with \( \hat{p}_k^T \) from the left gives
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{p}_k^T \hat{A}\hat{x} = \sum^{n}_{i=1} \alpha_i\hat{p}_k^T \hat{A}\hat{p}_i= \hat{p}_k^T \hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
and we can define the coefficients \( \alpha_k \) as
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\alpha_k = \frac{\hat{p}_k^T \hat{b}}{\hat{p}_k^T \hat{A} \hat{p}_k}
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec43">Conjugate gradient method and iterations </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
If we choose the conjugate vectors \( \hat{p}_k \) carefully,
|
||||
then we may not need all of them to obtain a good approximation to the solution
|
||||
\( \hat{x} \).
|
||||
We want to regard the conjugate gradient method as an iterative method.
|
||||
This will us to solve systems where \( n \) is so large that the direct
|
||||
method would take too much time.
|
||||
|
||||
<p>
|
||||
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
|
||||
We can assume without loss of generality that
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_0=0,
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
or consider the system
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
instead.
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec44">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
This suggests taking the first basis vector \( \hat{p}_1 \)
|
||||
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
|
||||
which equals
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x}_0-\hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
and
|
||||
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
|
||||
The other vectors in the basis will be conjugate to the gradient,
|
||||
hence the name conjugate gradient method.
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec45">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
|
||||
\( \hat{x}=\hat{x}_k \),
|
||||
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
|
||||
Here, we insist that the directions \( \hat{p}_k \) are conjugate to each other,
|
||||
so we take the direction closest to the gradient \( \hat{r}_k \)
|
||||
under the conjugacy constraint.
|
||||
This gives the following expression
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="___sec46">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
We can also compute the residual iteratively as
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
which equals
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
or
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
which gives
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
|
||||
\end{equation*}
|
||||
$$
|
||||
<p> <br>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
|
||||
|
||||
</div> <!-- class="slides" -->
|
||||
</div> <!-- class="reveal" -->
|
||||
|
||||
@@ -85,48 +85,39 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -639,7 +630,7 @@ When we have found the exact solution, \( \hat{r}=0 \).
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec18">Conjugate gradient method </h2>
|
||||
<h2 id="___sec18">Gradient method </h2>
|
||||
|
||||
<p>
|
||||
The residual is zero when we reach the minimum of the quadratic equation
|
||||
@@ -657,12 +648,131 @@ given by the second-derivative of the function we want to minimize.
|
||||
This quantity is always positive definite.
|
||||
|
||||
<p>
|
||||
More details will be added here soon.
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec19">Steepest descent method </h2>
|
||||
|
||||
<p>
|
||||
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
|
||||
We can assume without loss of generality that
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_0=0,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
or consider the system
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
instead.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec19">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
|
||||
<h2 id="___sec20">Steepest descent method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
|
||||
$$
|
||||
\begin{equation*}
|
||||
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
This suggests taking the first basis vector \( \hat{p}_1 \)
|
||||
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
|
||||
which equals
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x}_0-\hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
and
|
||||
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
|
||||
|
||||
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec21">Gradient descent method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
|
||||
\( \hat{x}=\hat{x}_k \),
|
||||
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
|
||||
This gives the following expression
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec22">Final expressions </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
We can also compute the residual iteratively as
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
which equals
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
or
|
||||
$$
|
||||
\begin{equation*}
|
||||
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
which gives
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec23">The Steepest descent algorithm </h2>
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec24">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
@@ -689,11 +799,7 @@ More details will be added here soon.
|
||||
cout << <span style="color: #CD5555">"The Matrix A that we are using: "</span> << endl;
|
||||
A.Print();
|
||||
cout << endl;
|
||||
x = ConjugateGradient(A,b,x0);
|
||||
xsd = SteepestDescent(A,b,x0);
|
||||
cout << <span style="color: #CD5555">"The approximate solution using Conjugate Gradient is: "</span> << endl;
|
||||
x.Print();
|
||||
cout << endl;
|
||||
cout << <span style="color: #CD5555">"The approximate solution using Steepest Descent is: "</span> << endl;
|
||||
xsd.Print();
|
||||
cout << endl;
|
||||
@@ -706,7 +812,7 @@ More details will be added here soon.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec20">The routine for the steepest descent method </h2>
|
||||
<h2 id="___sec25">The routine for the steepest descent method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
@@ -742,7 +848,7 @@ More details will be added here soon.
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec21">Revisiting our first homework </h2>
|
||||
<h2 id="___sec26">Revisiting our first homework </h2>
|
||||
|
||||
<p>
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
@@ -775,7 +881,7 @@ $$
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec22">Gradient descent example </h2>
|
||||
<h2 id="___sec27">Gradient descent example </h2>
|
||||
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
@@ -800,7 +906,7 @@ and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec23">The derivative of the cost/loss function </h2>
|
||||
<h2 id="___sec28">The derivative of the cost/loss function </h2>
|
||||
|
||||
<p>
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
@@ -815,7 +921,7 @@ where \( X \) is the design matrix defined above.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec24">The Hessian matrix </h2>
|
||||
<h2 id="___sec29">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
$$
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
@@ -829,7 +935,7 @@ This result implies that \( C(\beta) \) is a convex function since the matrix \(
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec25">Simple program </h2>
|
||||
<h2 id="___sec30">Simple program </h2>
|
||||
|
||||
<p>
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
@@ -870,7 +976,7 @@ beta_NE = np.dot(Xt_X_inv,Xt_y)
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec26">Gradient Descent Example </h2>
|
||||
<h2 id="___sec31">Gradient Descent Example </h2>
|
||||
|
||||
<p>
|
||||
Another simple example is here
|
||||
@@ -919,7 +1025,7 @@ plt.show()
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec27">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
<h2 id="___sec32">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
|
||||
<p>
|
||||
|
||||
@@ -943,7 +1049,7 @@ sgdreg.fit(x,y.ravel())
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec28">Gradient descent and Ridge </h2>
|
||||
<h2 id="___sec33">Gradient descent and Ridge </h2>
|
||||
|
||||
<p>
|
||||
We have also discussed Ridge regression where the loss function contains a regularized given by the \( L_2 \) norm of \( \beta \),
|
||||
@@ -1000,7 +1106,7 @@ beta_ridge = np.dot(Z,np.dot(X.T,y))
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec29">Stochastic Gradient Descent </h2>
|
||||
<h2 id="___sec34">Stochastic Gradient Descent </h2>
|
||||
|
||||
<p>
|
||||
Stochastic gradient descent (SGD) and variants thereof address some of
|
||||
@@ -1018,7 +1124,7 @@ $$
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec30">Computation of gradients </h2>
|
||||
<h2 id="___sec35">Computation of gradients </h2>
|
||||
|
||||
<p>
|
||||
This in turn means that the gradient can be
|
||||
@@ -1038,7 +1144,7 @@ minibatches. We denote these minibatches by \( B_k \) where
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec31">SGD example </h2>
|
||||
<h2 id="___sec36">SGD example </h2>
|
||||
As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \)
|
||||
and we choose to have \( M=5 \) minibathces,
|
||||
then each minibatch contains two data points. In particular we have
|
||||
@@ -1062,7 +1168,7 @@ $$
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec32">The gradient step </h2>
|
||||
<h2 id="___sec37">The gradient step </h2>
|
||||
|
||||
<p>
|
||||
Thus a gradient descent step now looks like
|
||||
@@ -1081,7 +1187,7 @@ the number of minibatches, as exemplified in the code below.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec33">Simple example code </h2>
|
||||
<h2 id="___sec38">Simple example code </h2>
|
||||
|
||||
<p>
|
||||
|
||||
@@ -1113,7 +1219,7 @@ all \( n \) datapoints.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec34">When do we stop? </h2>
|
||||
<h2 id="___sec39">When do we stop? </h2>
|
||||
|
||||
<p>
|
||||
A natural question is when do we stop the search for a new minimum?
|
||||
@@ -1130,7 +1236,7 @@ gave the lowest value.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec35">Slightly different approach </h2>
|
||||
<h2 id="___sec40">Slightly different approach </h2>
|
||||
|
||||
<p>
|
||||
Another approach is to let the step length \( \gamma_j \) depend on the
|
||||
@@ -1175,324 +1281,6 @@ j = <span style="color: #B452CD">0</span>
|
||||
|
||||
<span style="color: #8B008B; font-weight: bold">print</span>(<span style="color: #CD5555">"gamma_j after %d epochs: %g"</span> % (n_epochs,gamma_j))
|
||||
</pre></div>
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec36">Conjugate gradient (CG) method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
The success of the CG method for finding solutions of non-linear problems is based
|
||||
on the theory of conjugate gradients for linear systems of equations. It belongs
|
||||
to the class of iterative methods for solving problems from linear algebra of the type
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x} = \hat{b}.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
In the iterative process we end up with a problem like
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}= \hat{b}-\hat{A}\hat{x},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
where \( \hat{r} \) is the so-called residual or error in the iterative process.
|
||||
|
||||
<p>
|
||||
When we have found the exact solution, \( \hat{r}=0 \).
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec37">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
|
||||
<p>
|
||||
The residual is zero when we reach the minimum of the quadratic equation
|
||||
$$
|
||||
\begin{equation*}
|
||||
P(\hat{x})=\frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T\hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
with the constraint that the matrix \( \hat{A} \) is positive definite and symmetric.
|
||||
If we search for a minimum of the quantum mechanical variance, then the matrix
|
||||
\( \hat{A} \), which is called the Hessian, is given by the second-derivative of the function we want to minimize. This quantity is always positive definite. In our case this corresponds normally to the second derivative of the energy.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec38">Conjugate gradient method, Newton's method first </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
We seek the minimum of the energy or the variance as function of various variational parameters.
|
||||
In our case we have thus a function \( f \) whose minimum we are seeking.
|
||||
In Newton's method we set \( \nabla f = 0 \) and we can thus compute the next iteration point
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}-\hat{x}_i=\hat{A}^{-1}\nabla f(\hat{x}_i).
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
Subtracting this equation from that of \( \hat{x}_{i+1} \) we have
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_{i+1}-\hat{x}_i=\hat{A}^{-1}(\nabla f(\hat{x}_{i+1})-\nabla f(\hat{x}_i)).
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec39">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
In the CG method we define so-called conjugate directions and two vectors
|
||||
\( \hat{s} \) and \( \hat{t} \)
|
||||
are said to be
|
||||
conjugate if
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{s}^T\hat{A}\hat{t}= 0.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
The philosophy of the CG method is to perform searches in various conjugate directions
|
||||
of our vectors \( \hat{x}_i \) obeying the above criterion, namely
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_i^T\hat{A}\hat{x}_j= 0.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
Two vectors are conjugate if they are orthogonal with respect to
|
||||
this inner product. Being conjugate is a symmetric relation: if \( \hat{s} \) is conjugate to \( \hat{t} \), then \( \hat{t} \) is conjugate to \( \hat{s} \).
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec40">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
An example is given by the eigenvectors of the matrix
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{v}_i^T\hat{A}\hat{v}_j= \lambda\hat{v}_i^T\hat{v}_j,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
which is zero unless \( i=j \).
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec41">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
Assume now that we have a symmetric positive-definite matrix \( \hat{A} \) of size
|
||||
\( n\times n \). At each iteration \( i+1 \) we obtain the conjugate direction of a vector
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_{i+1}=\hat{x}_{i}+\alpha_i\hat{p}_{i}.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
We assume that \( \hat{p}_{i} \) is a sequence of \( n \) mutually conjugate directions.
|
||||
Then the \( \hat{p}_{i} \) form a basis of \( R^n \) and we can expand the solution
|
||||
$ \hat{A}\hat{x} = \hat{b}$ in this basis, namely
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x} = \sum^{n}_{i=1} \alpha_i \hat{p}_i.
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec42">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
The coefficients are given by
|
||||
$$
|
||||
\begin{equation*}
|
||||
\mathbf{A}\mathbf{x} = \sum^{n}_{i=1} \alpha_i \mathbf{A} \mathbf{p}_i = \mathbf{b}.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
Multiplying with \( \hat{p}_k^T \) from the left gives
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{p}_k^T \hat{A}\hat{x} = \sum^{n}_{i=1} \alpha_i\hat{p}_k^T \hat{A}\hat{p}_i= \hat{p}_k^T \hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
and we can define the coefficients \( \alpha_k \) as
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\alpha_k = \frac{\hat{p}_k^T \hat{b}}{\hat{p}_k^T \hat{A} \hat{p}_k}
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec43">Conjugate gradient method and iterations </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
|
||||
<p>
|
||||
If we choose the conjugate vectors \( \hat{p}_k \) carefully,
|
||||
then we may not need all of them to obtain a good approximation to the solution
|
||||
\( \hat{x} \).
|
||||
We want to regard the conjugate gradient method as an iterative method.
|
||||
This will us to solve systems where \( n \) is so large that the direct
|
||||
method would take too much time.
|
||||
|
||||
<p>
|
||||
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
|
||||
We can assume without loss of generality that
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_0=0,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
or consider the system
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
instead.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec44">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
|
||||
$$
|
||||
\begin{equation*}
|
||||
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
This suggests taking the first basis vector \( \hat{p}_1 \)
|
||||
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
|
||||
which equals
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x}_0-\hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
and
|
||||
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
|
||||
The other vectors in the basis will be conjugate to the gradient,
|
||||
hence the name conjugate gradient method.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec45">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
|
||||
\( \hat{x}=\hat{x}_k \),
|
||||
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
|
||||
Here, we insist that the directions \( \hat{p}_k \) are conjugate to each other,
|
||||
so we take the direction closest to the gradient \( \hat{r}_k \)
|
||||
under the conjugacy constraint.
|
||||
This gives the following expression
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec46">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
We can also compute the residual iteratively as
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
which equals
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
or
|
||||
$$
|
||||
\begin{equation*}
|
||||
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
which gives
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
+161
-373
@@ -90,48 +90,39 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
('More on convex functions', 2, None, '___sec15'),
|
||||
('Some simple problems', 2, None, '___sec16'),
|
||||
('Standard steepest descent', 2, None, '___sec17'),
|
||||
('Conjugate gradient method', 2, None, '___sec18'),
|
||||
('Gradient method', 2, None, '___sec18'),
|
||||
('Steepest descent method', 2, None, '___sec19'),
|
||||
('Steepest descent method', 2, None, '___sec20'),
|
||||
('Gradient descent method', 2, None, '___sec21'),
|
||||
('Final expressions', 2, None, '___sec22'),
|
||||
('The Steepest descent algorithm', 2, None, '___sec23'),
|
||||
('Simple codes for steepest descent and conjugate gradient '
|
||||
'using a $2\\times 2$ matrix, in c++, Python code to come',
|
||||
2,
|
||||
None,
|
||||
'___sec19'),
|
||||
'___sec24'),
|
||||
('The routine for the steepest descent method',
|
||||
2,
|
||||
None,
|
||||
'___sec20'),
|
||||
('Revisiting our first homework', 2, None, '___sec21'),
|
||||
('Gradient descent example', 2, None, '___sec22'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec23'),
|
||||
('The Hessian matrix', 2, None, '___sec24'),
|
||||
('Simple program', 2, None, '___sec25'),
|
||||
('Gradient Descent Example', 2, None, '___sec26'),
|
||||
'___sec25'),
|
||||
('Revisiting our first homework', 2, None, '___sec26'),
|
||||
('Gradient descent example', 2, None, '___sec27'),
|
||||
('The derivative of the cost/loss function', 2, None, '___sec28'),
|
||||
('The Hessian matrix', 2, None, '___sec29'),
|
||||
('Simple program', 2, None, '___sec30'),
|
||||
('Gradient Descent Example', 2, None, '___sec31'),
|
||||
('And a corresponding example using _scikit-learn_',
|
||||
2,
|
||||
None,
|
||||
'___sec27'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec28'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec29'),
|
||||
('Computation of gradients', 2, None, '___sec30'),
|
||||
('SGD example', 2, None, '___sec31'),
|
||||
('The gradient step', 2, None, '___sec32'),
|
||||
('Simple example code', 2, None, '___sec33'),
|
||||
('When do we stop?', 2, None, '___sec34'),
|
||||
('Slightly different approach', 2, None, '___sec35'),
|
||||
('Conjugate gradient (CG) method', 2, None, '___sec36'),
|
||||
('Conjugate gradient method', 2, None, '___sec37'),
|
||||
("Conjugate gradient method, Newton's method first",
|
||||
2,
|
||||
None,
|
||||
'___sec38'),
|
||||
('Conjugate gradient method', 2, None, '___sec39'),
|
||||
('Conjugate gradient method', 2, None, '___sec40'),
|
||||
('Conjugate gradient method', 2, None, '___sec41'),
|
||||
('Conjugate gradient method', 2, None, '___sec42'),
|
||||
('Conjugate gradient method and iterations', 2, None, '___sec43'),
|
||||
('Conjugate gradient method', 2, None, '___sec44'),
|
||||
('Conjugate gradient method', 2, None, '___sec45'),
|
||||
('Conjugate gradient method', 2, None, '___sec46')]}
|
||||
'___sec32'),
|
||||
('Gradient descent and Ridge', 2, None, '___sec33'),
|
||||
('Stochastic Gradient Descent', 2, None, '___sec34'),
|
||||
('Computation of gradients', 2, None, '___sec35'),
|
||||
('SGD example', 2, None, '___sec36'),
|
||||
('The gradient step', 2, None, '___sec37'),
|
||||
('Simple example code', 2, None, '___sec38'),
|
||||
('When do we stop?', 2, None, '___sec39'),
|
||||
('Slightly different approach', 2, None, '___sec40')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -644,7 +635,7 @@ When we have found the exact solution, \( \hat{r}=0 \).
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec18">Conjugate gradient method </h2>
|
||||
<h2 id="___sec18">Gradient method </h2>
|
||||
|
||||
<p>
|
||||
The residual is zero when we reach the minimum of the quadratic equation
|
||||
@@ -662,12 +653,131 @@ given by the second-derivative of the function we want to minimize.
|
||||
This quantity is always positive definite.
|
||||
|
||||
<p>
|
||||
More details will be added here soon.
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec19">Steepest descent method </h2>
|
||||
|
||||
<p>
|
||||
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
|
||||
We can assume without loss of generality that
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_0=0,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
or consider the system
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
instead.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec19">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
|
||||
<h2 id="___sec20">Steepest descent method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
|
||||
$$
|
||||
\begin{equation*}
|
||||
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
This suggests taking the first basis vector \( \hat{p}_1 \)
|
||||
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
|
||||
which equals
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x}_0-\hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
and
|
||||
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
|
||||
|
||||
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec21">Gradient descent method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
|
||||
\( \hat{x}=\hat{x}_k \),
|
||||
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
|
||||
This gives the following expression
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec22">Final expressions </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
We can also compute the residual iteratively as
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
which equals
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
or
|
||||
$$
|
||||
\begin{equation*}
|
||||
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
which gives
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec23">The Steepest descent algorithm </h2>
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec24">Simple codes for steepest descent and conjugate gradient using a \( 2\times 2 \) matrix, in c++, Python code to come </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
@@ -694,11 +804,7 @@ More details will be added here soon.
|
||||
cout <span style="color: #666666"><<</span> <span style="color: #BA2121">"The Matrix A that we are using: "</span> <span style="color: #666666"><<</span> endl;
|
||||
A.Print();
|
||||
cout <span style="color: #666666"><<</span> endl;
|
||||
x <span style="color: #666666">=</span> ConjugateGradient(A,b,x0);
|
||||
xsd <span style="color: #666666">=</span> SteepestDescent(A,b,x0);
|
||||
cout <span style="color: #666666"><<</span> <span style="color: #BA2121">"The approximate solution using Conjugate Gradient is: "</span> <span style="color: #666666"><<</span> endl;
|
||||
x.Print();
|
||||
cout <span style="color: #666666"><<</span> endl;
|
||||
cout <span style="color: #666666"><<</span> <span style="color: #BA2121">"The approximate solution using Steepest Descent is: "</span> <span style="color: #666666"><<</span> endl;
|
||||
xsd.Print();
|
||||
cout <span style="color: #666666"><<</span> endl;
|
||||
@@ -711,7 +817,7 @@ More details will be added here soon.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec20">The routine for the steepest descent method </h2>
|
||||
<h2 id="___sec25">The routine for the steepest descent method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
@@ -747,7 +853,7 @@ More details will be added here soon.
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec21">Revisiting our first homework </h2>
|
||||
<h2 id="___sec26">Revisiting our first homework </h2>
|
||||
|
||||
<p>
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
@@ -780,7 +886,7 @@ $$
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec22">Gradient descent example </h2>
|
||||
<h2 id="___sec27">Gradient descent example </h2>
|
||||
|
||||
<p>
|
||||
Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
|
||||
@@ -805,7 +911,7 @@ and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec23">The derivative of the cost/loss function </h2>
|
||||
<h2 id="___sec28">The derivative of the cost/loss function </h2>
|
||||
|
||||
<p>
|
||||
Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
|
||||
@@ -820,7 +926,7 @@ where \( X \) is the design matrix defined above.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec24">The Hessian matrix </h2>
|
||||
<h2 id="___sec29">The Hessian matrix </h2>
|
||||
The Hessian matrix of \( C(\beta) \) is given by
|
||||
$$
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
@@ -834,7 +940,7 @@ This result implies that \( C(\beta) \) is a convex function since the matrix \(
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec25">Simple program </h2>
|
||||
<h2 id="___sec30">Simple program </h2>
|
||||
|
||||
<p>
|
||||
We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
|
||||
@@ -875,7 +981,7 @@ beta_NE <span style="color: #666666">=</span> np<span style="color: #666666">.</
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec26">Gradient Descent Example </h2>
|
||||
<h2 id="___sec31">Gradient Descent Example </h2>
|
||||
|
||||
<p>
|
||||
Another simple example is here
|
||||
@@ -924,7 +1030,7 @@ plt<span style="color: #666666">.</span>show()
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec27">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
<h2 id="___sec32">And a corresponding example using <b>scikit-learn</b> </h2>
|
||||
|
||||
<p>
|
||||
|
||||
@@ -948,7 +1054,7 @@ sgdreg<span style="color: #666666">.</span>fit(x,y<span style="color: #666666">.
|
||||
<p>
|
||||
<!-- !split -->
|
||||
|
||||
<h2 id="___sec28">Gradient descent and Ridge </h2>
|
||||
<h2 id="___sec33">Gradient descent and Ridge </h2>
|
||||
|
||||
<p>
|
||||
We have also discussed Ridge regression where the loss function contains a regularized given by the \( L_2 \) norm of \( \beta \),
|
||||
@@ -1005,7 +1111,7 @@ beta_ridge <span style="color: #666666">=</span> np<span style="color: #666666">
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec29">Stochastic Gradient Descent </h2>
|
||||
<h2 id="___sec34">Stochastic Gradient Descent </h2>
|
||||
|
||||
<p>
|
||||
Stochastic gradient descent (SGD) and variants thereof address some of
|
||||
@@ -1023,7 +1129,7 @@ $$
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec30">Computation of gradients </h2>
|
||||
<h2 id="___sec35">Computation of gradients </h2>
|
||||
|
||||
<p>
|
||||
This in turn means that the gradient can be
|
||||
@@ -1043,7 +1149,7 @@ minibatches. We denote these minibatches by \( B_k \) where
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec31">SGD example </h2>
|
||||
<h2 id="___sec36">SGD example </h2>
|
||||
As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \)
|
||||
and we choose to have \( M=5 \) minibathces,
|
||||
then each minibatch contains two data points. In particular we have
|
||||
@@ -1067,7 +1173,7 @@ $$
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec32">The gradient step </h2>
|
||||
<h2 id="___sec37">The gradient step </h2>
|
||||
|
||||
<p>
|
||||
Thus a gradient descent step now looks like
|
||||
@@ -1086,7 +1192,7 @@ the number of minibatches, as exemplified in the code below.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec33">Simple example code </h2>
|
||||
<h2 id="___sec38">Simple example code </h2>
|
||||
|
||||
<p>
|
||||
|
||||
@@ -1118,7 +1224,7 @@ all \( n \) datapoints.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec34">When do we stop? </h2>
|
||||
<h2 id="___sec39">When do we stop? </h2>
|
||||
|
||||
<p>
|
||||
A natural question is when do we stop the search for a new minimum?
|
||||
@@ -1135,7 +1241,7 @@ gave the lowest value.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec35">Slightly different approach </h2>
|
||||
<h2 id="___sec40">Slightly different approach </h2>
|
||||
|
||||
<p>
|
||||
Another approach is to let the step length \( \gamma_j \) depend on the
|
||||
@@ -1180,324 +1286,6 @@ j <span style="color: #666666">=</span> <span style="color: #666666">0</span>
|
||||
|
||||
<span style="color: #008000; font-weight: bold">print</span>(<span style="color: #BA2121">"gamma_j after </span><span style="color: #BB6688; font-weight: bold">%d</span><span style="color: #BA2121"> epochs: </span><span style="color: #BB6688; font-weight: bold">%g</span><span style="color: #BA2121">"</span> <span style="color: #666666">%</span> (n_epochs,gamma_j))
|
||||
</pre></div>
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec36">Conjugate gradient (CG) method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
The success of the CG method for finding solutions of non-linear problems is based
|
||||
on the theory of conjugate gradients for linear systems of equations. It belongs
|
||||
to the class of iterative methods for solving problems from linear algebra of the type
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x} = \hat{b}.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
In the iterative process we end up with a problem like
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}= \hat{b}-\hat{A}\hat{x},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
where \( \hat{r} \) is the so-called residual or error in the iterative process.
|
||||
|
||||
<p>
|
||||
When we have found the exact solution, \( \hat{r}=0 \).
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec37">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
|
||||
<p>
|
||||
The residual is zero when we reach the minimum of the quadratic equation
|
||||
$$
|
||||
\begin{equation*}
|
||||
P(\hat{x})=\frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T\hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
with the constraint that the matrix \( \hat{A} \) is positive definite and symmetric.
|
||||
If we search for a minimum of the quantum mechanical variance, then the matrix
|
||||
\( \hat{A} \), which is called the Hessian, is given by the second-derivative of the function we want to minimize. This quantity is always positive definite. In our case this corresponds normally to the second derivative of the energy.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec38">Conjugate gradient method, Newton's method first </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
We seek the minimum of the energy or the variance as function of various variational parameters.
|
||||
In our case we have thus a function \( f \) whose minimum we are seeking.
|
||||
In Newton's method we set \( \nabla f = 0 \) and we can thus compute the next iteration point
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}-\hat{x}_i=\hat{A}^{-1}\nabla f(\hat{x}_i).
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
Subtracting this equation from that of \( \hat{x}_{i+1} \) we have
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_{i+1}-\hat{x}_i=\hat{A}^{-1}(\nabla f(\hat{x}_{i+1})-\nabla f(\hat{x}_i)).
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec39">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
In the CG method we define so-called conjugate directions and two vectors
|
||||
\( \hat{s} \) and \( \hat{t} \)
|
||||
are said to be
|
||||
conjugate if
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{s}^T\hat{A}\hat{t}= 0.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
The philosophy of the CG method is to perform searches in various conjugate directions
|
||||
of our vectors \( \hat{x}_i \) obeying the above criterion, namely
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_i^T\hat{A}\hat{x}_j= 0.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
Two vectors are conjugate if they are orthogonal with respect to
|
||||
this inner product. Being conjugate is a symmetric relation: if \( \hat{s} \) is conjugate to \( \hat{t} \), then \( \hat{t} \) is conjugate to \( \hat{s} \).
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec40">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
An example is given by the eigenvectors of the matrix
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{v}_i^T\hat{A}\hat{v}_j= \lambda\hat{v}_i^T\hat{v}_j,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
which is zero unless \( i=j \).
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec41">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
Assume now that we have a symmetric positive-definite matrix \( \hat{A} \) of size
|
||||
\( n\times n \). At each iteration \( i+1 \) we obtain the conjugate direction of a vector
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_{i+1}=\hat{x}_{i}+\alpha_i\hat{p}_{i}.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
We assume that \( \hat{p}_{i} \) is a sequence of \( n \) mutually conjugate directions.
|
||||
Then the \( \hat{p}_{i} \) form a basis of \( R^n \) and we can expand the solution
|
||||
$ \hat{A}\hat{x} = \hat{b}$ in this basis, namely
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x} = \sum^{n}_{i=1} \alpha_i \hat{p}_i.
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec42">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
The coefficients are given by
|
||||
$$
|
||||
\begin{equation*}
|
||||
\mathbf{A}\mathbf{x} = \sum^{n}_{i=1} \alpha_i \mathbf{A} \mathbf{p}_i = \mathbf{b}.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
Multiplying with \( \hat{p}_k^T \) from the left gives
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{p}_k^T \hat{A}\hat{x} = \sum^{n}_{i=1} \alpha_i\hat{p}_k^T \hat{A}\hat{p}_i= \hat{p}_k^T \hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
and we can define the coefficients \( \alpha_k \) as
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\alpha_k = \frac{\hat{p}_k^T \hat{b}}{\hat{p}_k^T \hat{A} \hat{p}_k}
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec43">Conjugate gradient method and iterations </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
|
||||
<p>
|
||||
If we choose the conjugate vectors \( \hat{p}_k \) carefully,
|
||||
then we may not need all of them to obtain a good approximation to the solution
|
||||
\( \hat{x} \).
|
||||
We want to regard the conjugate gradient method as an iterative method.
|
||||
This will us to solve systems where \( n \) is so large that the direct
|
||||
method would take too much time.
|
||||
|
||||
<p>
|
||||
We denote the initial guess for \( \hat{x} \) as \( \hat{x}_0 \).
|
||||
We can assume without loss of generality that
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{x}_0=0,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
or consider the system
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
instead.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec44">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
One can show that the solution \( \hat{x} \) is also the unique minimizer of the quadratic form
|
||||
$$
|
||||
\begin{equation*}
|
||||
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
This suggests taking the first basis vector \( \hat{p}_1 \)
|
||||
to be the gradient of \( f \) at \( \hat{x}=\hat{x}_0 \),
|
||||
which equals
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x}_0-\hat{b},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
and
|
||||
\( \hat{x}_0=0 \) it is equal \( -\hat{b} \).
|
||||
The other vectors in the basis will be conjugate to the gradient,
|
||||
hence the name conjugate gradient method.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec45">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
Let \( \hat{r}_k \) be the residual at the \( k \)-th step:
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
Note that \( \hat{r}_k \) is the negative gradient of \( f \) at
|
||||
\( \hat{x}=\hat{x}_k \),
|
||||
so the gradient descent method would be to move in the direction \( \hat{r}_k \).
|
||||
Here, we insist that the directions \( \hat{p}_k \) are conjugate to each other,
|
||||
so we take the direction closest to the gradient \( \hat{r}_k \)
|
||||
under the conjugacy constraint.
|
||||
This gives the following expression
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec46">Conjugate gradient method </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
We can also compute the residual iteratively as
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
which equals
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
or
|
||||
$$
|
||||
\begin{equation*}
|
||||
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
|
||||
\end{equation*}
|
||||
$$
|
||||
|
||||
which gives
|
||||
|
||||
$$
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
|
||||
\end{equation*}
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
+184
-454
@@ -585,7 +585,7 @@
|
||||
"\n",
|
||||
"When we have found the exact solution, $\\hat{r}=0$.\n",
|
||||
"\n",
|
||||
"## Conjugate gradient method\n",
|
||||
"## Gradient method\n",
|
||||
"\n",
|
||||
"The residual is zero when we reach the minimum of the quadratic equation"
|
||||
]
|
||||
@@ -609,7 +609,189 @@
|
||||
"given by the second-derivative of the function we want to minimize.\n",
|
||||
"This quantity is always positive definite. \n",
|
||||
"\n",
|
||||
"More details will be added here soon.\n",
|
||||
"\n",
|
||||
"## Steepest descent method\n",
|
||||
"\n",
|
||||
"We denote the initial guess for $\\hat{x}$ as $\\hat{x}_0$. \n",
|
||||
"We can assume without loss of generality that"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{x}_0=0,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"or consider the system"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{A}\\hat{z} = \\hat{b}-\\hat{A}\\hat{x}_0,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"instead.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Steepest descent method\n",
|
||||
"One can show that the solution $\\hat{x}$ is also the unique minimizer of the quadratic form"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"f(\\hat{x}) = \\frac{1}{2}\\hat{x}^T\\hat{A}\\hat{x} - \\hat{x}^T \\hat{x} , \\quad \\hat{x}\\in\\mathbf{R}^n.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"This suggests taking the first basis vector $\\hat{p}_1$ \n",
|
||||
"to be the gradient of $f$ at $\\hat{x}=\\hat{x}_0$, \n",
|
||||
"which equals"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{A}\\hat{x}_0-\\hat{b},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and \n",
|
||||
"$\\hat{x}_0=0$ it is equal $-\\hat{b}$.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Gradient descent method\n",
|
||||
"Let $\\hat{r}_k$ be the residual at the $k$-th step:"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{r}_k=\\hat{b}-\\hat{A}\\hat{x}_k.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Note that $\\hat{r}_k$ is the negative gradient of $f$ at \n",
|
||||
"$\\hat{x}=\\hat{x}_k$, \n",
|
||||
"so the gradient descent method would be to move in the direction $\\hat{r}_k$. \n",
|
||||
"This gives the following expression"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{p}_{k+1}=\\hat{r}_k-\\frac{\\hat{p}_k^T \\hat{A}\\hat{r}_k}{\\hat{p}_k^T\\hat{A}\\hat{p}_k} \\hat{p}_k.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Final expressions\n",
|
||||
"We can also compute the residual iteratively as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{r}_{k+1}=\\hat{b}-\\hat{A}\\hat{x}_{k+1},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"which equals"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{b}-\\hat{A}(\\hat{x}_k+\\alpha_k\\hat{p}_k),\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"or"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"(\\hat{b}-\\hat{A}\\hat{x}_k)-\\alpha_k\\hat{A}\\hat{p}_k,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"which gives"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{r}_{k+1}=\\hat{r}_k-\\hat{A}\\hat{p}_{k},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## The Steepest descent algorithm\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Simple codes for steepest descent and conjugate gradient using a $2\\times 2$ matrix, in c++, Python code to come"
|
||||
]
|
||||
@@ -638,11 +820,7 @@
|
||||
" cout << \"The Matrix A that we are using: \" << endl;\n",
|
||||
" A.Print();\n",
|
||||
" cout << endl;\n",
|
||||
" x = ConjugateGradient(A,b,x0);\n",
|
||||
" xsd = SteepestDescent(A,b,x0);\n",
|
||||
" cout << \"The approximate solution using Conjugate Gradient is: \" << endl;\n",
|
||||
" x.Print();\n",
|
||||
" cout << endl;\n",
|
||||
" cout << \"The approximate solution using Steepest Descent is: \" << endl;\n",
|
||||
" xsd.Print();\n",
|
||||
" cout << endl;\n",
|
||||
@@ -1289,454 +1467,6 @@
|
||||
"\n",
|
||||
"print(\"gamma_j after %d epochs: %g\" % (n_epochs,gamma_j))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Conjugate gradient (CG) method\n",
|
||||
"The success of the CG method for finding solutions of non-linear problems is based\n",
|
||||
"on the theory of conjugate gradients for linear systems of equations. It belongs\n",
|
||||
"to the class of iterative methods for solving problems from linear algebra of the type"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{A}\\hat{x} = \\hat{b}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"In the iterative process we end up with a problem like"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{r}= \\hat{b}-\\hat{A}\\hat{x},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"where $\\hat{r}$ is the so-called residual or error in the iterative process.\n",
|
||||
"\n",
|
||||
"When we have found the exact solution, $\\hat{r}=0$.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Conjugate gradient method\n",
|
||||
"\n",
|
||||
"The residual is zero when we reach the minimum of the quadratic equation"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"P(\\hat{x})=\\frac{1}{2}\\hat{x}^T\\hat{A}\\hat{x} - \\hat{x}^T\\hat{b},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"with the constraint that the matrix $\\hat{A}$ is positive definite and symmetric.\n",
|
||||
"If we search for a minimum of the quantum mechanical variance, then the matrix \n",
|
||||
"$\\hat{A}$, which is called the Hessian, is given by the second-derivative of the function we want to minimize. This quantity is always positive definite. In our case this corresponds normally to the second derivative of the energy.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Conjugate gradient method, Newton's method first\n",
|
||||
"We seek the minimum of the energy or the variance as function of various variational parameters. \n",
|
||||
"In our case we have thus a function $f$ whose minimum we are seeking.\n",
|
||||
"In Newton's method we set $\\nabla f = 0$ and we can thus compute the next iteration point"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{x}-\\hat{x}_i=\\hat{A}^{-1}\\nabla f(\\hat{x}_i).\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Subtracting this equation from that of $\\hat{x}_{i+1}$ we have"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{x}_{i+1}-\\hat{x}_i=\\hat{A}^{-1}(\\nabla f(\\hat{x}_{i+1})-\\nabla f(\\hat{x}_i)).\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Conjugate gradient method\n",
|
||||
"In the CG method we define so-called conjugate directions and two vectors \n",
|
||||
"$\\hat{s}$ and $\\hat{t}$\n",
|
||||
"are said to be\n",
|
||||
"conjugate if"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{s}^T\\hat{A}\\hat{t}= 0.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"The philosophy of the CG method is to perform searches in various conjugate directions\n",
|
||||
"of our vectors $\\hat{x}_i$ obeying the above criterion, namely"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{x}_i^T\\hat{A}\\hat{x}_j= 0.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Two vectors are conjugate if they are orthogonal with respect to \n",
|
||||
"this inner product. Being conjugate is a symmetric relation: if $\\hat{s}$ is conjugate to $\\hat{t}$, then $\\hat{t}$ is conjugate to $\\hat{s}$.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Conjugate gradient method\n",
|
||||
"An example is given by the eigenvectors of the matrix"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{v}_i^T\\hat{A}\\hat{v}_j= \\lambda\\hat{v}_i^T\\hat{v}_j,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"which is zero unless $i=j$.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Conjugate gradient method\n",
|
||||
"Assume now that we have a symmetric positive-definite matrix $\\hat{A}$ of size\n",
|
||||
"$n\\times n$. At each iteration $i+1$ we obtain the conjugate direction of a vector"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{x}_{i+1}=\\hat{x}_{i}+\\alpha_i\\hat{p}_{i}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"We assume that $\\hat{p}_{i}$ is a sequence of $n$ mutually conjugate directions. \n",
|
||||
"Then the $\\hat{p}_{i}$ form a basis of $R^n$ and we can expand the solution \n",
|
||||
"$ \\hat{A}\\hat{x} = \\hat{b}$ in this basis, namely"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{x} = \\sum^{n}_{i=1} \\alpha_i \\hat{p}_i.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Conjugate gradient method\n",
|
||||
"The coefficients are given by"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\mathbf{A}\\mathbf{x} = \\sum^{n}_{i=1} \\alpha_i \\mathbf{A} \\mathbf{p}_i = \\mathbf{b}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Multiplying with $\\hat{p}_k^T$ from the left gives"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{p}_k^T \\hat{A}\\hat{x} = \\sum^{n}_{i=1} \\alpha_i\\hat{p}_k^T \\hat{A}\\hat{p}_i= \\hat{p}_k^T \\hat{b},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and we can define the coefficients $\\alpha_k$ as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\alpha_k = \\frac{\\hat{p}_k^T \\hat{b}}{\\hat{p}_k^T \\hat{A} \\hat{p}_k}\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Conjugate gradient method and iterations\n",
|
||||
"\n",
|
||||
"If we choose the conjugate vectors $\\hat{p}_k$ carefully, \n",
|
||||
"then we may not need all of them to obtain a good approximation to the solution \n",
|
||||
"$\\hat{x}$. \n",
|
||||
"We want to regard the conjugate gradient method as an iterative method. \n",
|
||||
"This will us to solve systems where $n$ is so large that the direct \n",
|
||||
"method would take too much time.\n",
|
||||
"\n",
|
||||
"We denote the initial guess for $\\hat{x}$ as $\\hat{x}_0$. \n",
|
||||
"We can assume without loss of generality that"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{x}_0=0,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"or consider the system"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{A}\\hat{z} = \\hat{b}-\\hat{A}\\hat{x}_0,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"instead.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Conjugate gradient method\n",
|
||||
"One can show that the solution $\\hat{x}$ is also the unique minimizer of the quadratic form"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"f(\\hat{x}) = \\frac{1}{2}\\hat{x}^T\\hat{A}\\hat{x} - \\hat{x}^T \\hat{x} , \\quad \\hat{x}\\in\\mathbf{R}^n.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"This suggests taking the first basis vector $\\hat{p}_1$ \n",
|
||||
"to be the gradient of $f$ at $\\hat{x}=\\hat{x}_0$, \n",
|
||||
"which equals"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{A}\\hat{x}_0-\\hat{b},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and \n",
|
||||
"$\\hat{x}_0=0$ it is equal $-\\hat{b}$.\n",
|
||||
"The other vectors in the basis will be conjugate to the gradient, \n",
|
||||
"hence the name conjugate gradient method.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Conjugate gradient method\n",
|
||||
"Let $\\hat{r}_k$ be the residual at the $k$-th step:"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{r}_k=\\hat{b}-\\hat{A}\\hat{x}_k.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Note that $\\hat{r}_k$ is the negative gradient of $f$ at \n",
|
||||
"$\\hat{x}=\\hat{x}_k$, \n",
|
||||
"so the gradient descent method would be to move in the direction $\\hat{r}_k$. \n",
|
||||
"Here, we insist that the directions $\\hat{p}_k$ are conjugate to each other, \n",
|
||||
"so we take the direction closest to the gradient $\\hat{r}_k$ \n",
|
||||
"under the conjugacy constraint. \n",
|
||||
"This gives the following expression"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{p}_{k+1}=\\hat{r}_k-\\frac{\\hat{p}_k^T \\hat{A}\\hat{r}_k}{\\hat{p}_k^T\\hat{A}\\hat{p}_k} \\hat{p}_k.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Conjugate gradient method\n",
|
||||
"We can also compute the residual iteratively as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{r}_{k+1}=\\hat{b}-\\hat{A}\\hat{x}_{k+1},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"which equals"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{b}-\\hat{A}(\\hat{x}_k+\\alpha_k\\hat{p}_k),\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"or"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"(\\hat{b}-\\hat{A}\\hat{x}_k)-\\alpha_k\\hat{A}\\hat{p}_k,\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"which gives"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{r}_{k+1}=\\hat{r}_k-\\hat{A}\\hat{p}_{k},\n",
|
||||
"$$"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {},
|
||||
|
||||
Binary file not shown.
Binary file not shown.
+99
-255
@@ -405,7 +405,7 @@ where $\hat{r}$ is the so-called residual or error in the iterative process.
|
||||
When we have found the exact solution, $\hat{r}=0$.
|
||||
|
||||
!split
|
||||
===== Conjugate gradient method =====
|
||||
===== Gradient method =====
|
||||
|
||||
The residual is zero when we reach the minimum of the quadratic equation
|
||||
!bt
|
||||
@@ -420,7 +420,104 @@ variance, then the matrix $\hat{A}$, which is called the Hessian, is
|
||||
given by the second-derivative of the function we want to minimize.
|
||||
This quantity is always positive definite.
|
||||
|
||||
More details will be added here soon.
|
||||
|
||||
!split
|
||||
===== Steepest descent method =====
|
||||
|
||||
We denote the initial guess for $\hat{x}$ as $\hat{x}_0$.
|
||||
We can assume without loss of generality that
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{x}_0=0,
|
||||
\end{equation*}
|
||||
!et
|
||||
or consider the system
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
|
||||
\end{equation*}
|
||||
!et
|
||||
instead.
|
||||
|
||||
|
||||
!split
|
||||
===== Steepest descent method =====
|
||||
!bblock
|
||||
One can show that the solution $\hat{x}$ is also the unique minimizer of the quadratic form
|
||||
!bt
|
||||
\begin{equation*}
|
||||
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
|
||||
\end{equation*}
|
||||
!et
|
||||
This suggests taking the first basis vector $\hat{p}_1$
|
||||
to be the gradient of $f$ at $\hat{x}=\hat{x}_0$,
|
||||
which equals
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x}_0-\hat{b},
|
||||
\end{equation*}
|
||||
!et
|
||||
and
|
||||
$\hat{x}_0=0$ it is equal $-\hat{b}$.
|
||||
|
||||
!eblock
|
||||
|
||||
|
||||
!split
|
||||
===== Gradient descent method =====
|
||||
!bblock
|
||||
Let $\hat{r}_k$ be the residual at the $k$-th step:
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
|
||||
\end{equation*}
|
||||
!et
|
||||
Note that $\hat{r}_k$ is the negative gradient of $f$ at
|
||||
$\hat{x}=\hat{x}_k$,
|
||||
so the gradient descent method would be to move in the direction $\hat{r}_k$.
|
||||
This gives the following expression
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
|
||||
\end{equation*}
|
||||
!et
|
||||
!eblock
|
||||
|
||||
!split
|
||||
===== Final expressions =====
|
||||
!bblock
|
||||
We can also compute the residual iteratively as
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
|
||||
\end{equation*}
|
||||
!et
|
||||
which equals
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
|
||||
\end{equation*}
|
||||
!et
|
||||
or
|
||||
!bt
|
||||
\begin{equation*}
|
||||
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
|
||||
\end{equation*}
|
||||
!et
|
||||
which gives
|
||||
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
|
||||
\end{equation*}
|
||||
!et
|
||||
!eblock
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== The Steepest descent algorithm =====
|
||||
|
||||
|
||||
!split
|
||||
===== Simple codes for steepest descent and conjugate gradient using a $2\times 2$ matrix, in c++, Python code to come =====
|
||||
@@ -446,11 +543,7 @@ int main(int argc, char * argv[]){
|
||||
cout << "The Matrix A that we are using: " << endl;
|
||||
A.Print();
|
||||
cout << endl;
|
||||
x = ConjugateGradient(A,b,x0);
|
||||
xsd = SteepestDescent(A,b,x0);
|
||||
cout << "The approximate solution using Conjugate Gradient is: " << endl;
|
||||
x.Print();
|
||||
cout << endl;
|
||||
cout << "The approximate solution using Steepest Descent is: " << endl;
|
||||
xsd.Print();
|
||||
cout << endl;
|
||||
@@ -895,255 +988,6 @@ print("gamma_j after %d epochs: %g" % (n_epochs,gamma_j))
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== Conjugate gradient (CG) method =====
|
||||
!bblock
|
||||
The success of the CG method for finding solutions of non-linear problems is based
|
||||
on the theory of conjugate gradients for linear systems of equations. It belongs
|
||||
to the class of iterative methods for solving problems from linear algebra of the type
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x} = \hat{b}.
|
||||
\end{equation*}
|
||||
!et
|
||||
In the iterative process we end up with a problem like
|
||||
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{r}= \hat{b}-\hat{A}\hat{x},
|
||||
\end{equation*}
|
||||
!et
|
||||
where $\hat{r}$ is the so-called residual or error in the iterative process.
|
||||
|
||||
When we have found the exact solution, $\hat{r}=0$.
|
||||
!eblock
|
||||
|
||||
|
||||
!split
|
||||
===== Conjugate gradient method =====
|
||||
!bblock
|
||||
|
||||
The residual is zero when we reach the minimum of the quadratic equation
|
||||
!bt
|
||||
\begin{equation*}
|
||||
P(\hat{x})=\frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T\hat{b},
|
||||
\end{equation*}
|
||||
!et
|
||||
with the constraint that the matrix $\hat{A}$ is positive definite and symmetric.
|
||||
If we search for a minimum of the quantum mechanical variance, then the matrix
|
||||
$\hat{A}$, which is called the Hessian, is given by the second-derivative of the function we want to minimize. This quantity is always positive definite. In our case this corresponds normally to the second derivative of the energy.
|
||||
!eblock
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== Conjugate gradient method, Newton's method first =====
|
||||
!bblock
|
||||
We seek the minimum of the energy or the variance as function of various variational parameters.
|
||||
In our case we have thus a function $f$ whose minimum we are seeking.
|
||||
In Newton's method we set $\nabla f = 0$ and we can thus compute the next iteration point
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{x}-\hat{x}_i=\hat{A}^{-1}\nabla f(\hat{x}_i).
|
||||
\end{equation*}
|
||||
!et
|
||||
Subtracting this equation from that of $\hat{x}_{i+1}$ we have
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{x}_{i+1}-\hat{x}_i=\hat{A}^{-1}(\nabla f(\hat{x}_{i+1})-\nabla f(\hat{x}_i)).
|
||||
\end{equation*}
|
||||
!et
|
||||
!eblock
|
||||
|
||||
|
||||
!split
|
||||
===== Conjugate gradient method =====
|
||||
!bblock
|
||||
In the CG method we define so-called conjugate directions and two vectors
|
||||
$\hat{s}$ and $\hat{t}$
|
||||
are said to be
|
||||
conjugate if
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{s}^T\hat{A}\hat{t}= 0.
|
||||
\end{equation*}
|
||||
!et
|
||||
The philosophy of the CG method is to perform searches in various conjugate directions
|
||||
of our vectors $\hat{x}_i$ obeying the above criterion, namely
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{x}_i^T\hat{A}\hat{x}_j= 0.
|
||||
\end{equation*}
|
||||
!et
|
||||
Two vectors are conjugate if they are orthogonal with respect to
|
||||
this inner product. Being conjugate is a symmetric relation: if $\hat{s}$ is conjugate to $\hat{t}$, then $\hat{t}$ is conjugate to $\hat{s}$.
|
||||
!eblock
|
||||
|
||||
!split
|
||||
===== Conjugate gradient method =====
|
||||
!bblock
|
||||
An example is given by the eigenvectors of the matrix
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{v}_i^T\hat{A}\hat{v}_j= \lambda\hat{v}_i^T\hat{v}_j,
|
||||
\end{equation*}
|
||||
!et
|
||||
which is zero unless $i=j$.
|
||||
!eblock
|
||||
|
||||
|
||||
!split
|
||||
===== Conjugate gradient method =====
|
||||
!bblock
|
||||
Assume now that we have a symmetric positive-definite matrix $\hat{A}$ of size
|
||||
$n\times n$. At each iteration $i+1$ we obtain the conjugate direction of a vector
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{x}_{i+1}=\hat{x}_{i}+\alpha_i\hat{p}_{i}.
|
||||
\end{equation*}
|
||||
!et
|
||||
We assume that $\hat{p}_{i}$ is a sequence of $n$ mutually conjugate directions.
|
||||
Then the $\hat{p}_{i}$ form a basis of $R^n$ and we can expand the solution
|
||||
$ \hat{A}\hat{x} = \hat{b}$ in this basis, namely
|
||||
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{x} = \sum^{n}_{i=1} \alpha_i \hat{p}_i.
|
||||
\end{equation*}
|
||||
!et
|
||||
!eblock
|
||||
|
||||
!split
|
||||
===== Conjugate gradient method =====
|
||||
!bblock
|
||||
The coefficients are given by
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\mathbf{A}\mathbf{x} = \sum^{n}_{i=1} \alpha_i \mathbf{A} \mathbf{p}_i = \mathbf{b}.
|
||||
\end{equation*}
|
||||
!et
|
||||
Multiplying with $\hat{p}_k^T$ from the left gives
|
||||
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{p}_k^T \hat{A}\hat{x} = \sum^{n}_{i=1} \alpha_i\hat{p}_k^T \hat{A}\hat{p}_i= \hat{p}_k^T \hat{b},
|
||||
\end{equation*}
|
||||
!et
|
||||
and we can define the coefficients $\alpha_k$ as
|
||||
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\alpha_k = \frac{\hat{p}_k^T \hat{b}}{\hat{p}_k^T \hat{A} \hat{p}_k}
|
||||
\end{equation*}
|
||||
!et
|
||||
!eblock
|
||||
|
||||
!split
|
||||
===== Conjugate gradient method and iterations =====
|
||||
!bblock
|
||||
|
||||
If we choose the conjugate vectors $\hat{p}_k$ carefully,
|
||||
then we may not need all of them to obtain a good approximation to the solution
|
||||
$\hat{x}$.
|
||||
We want to regard the conjugate gradient method as an iterative method.
|
||||
This will us to solve systems where $n$ is so large that the direct
|
||||
method would take too much time.
|
||||
|
||||
We denote the initial guess for $\hat{x}$ as $\hat{x}_0$.
|
||||
We can assume without loss of generality that
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{x}_0=0,
|
||||
\end{equation*}
|
||||
!et
|
||||
or consider the system
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{z} = \hat{b}-\hat{A}\hat{x}_0,
|
||||
\end{equation*}
|
||||
!et
|
||||
instead.
|
||||
!eblock
|
||||
|
||||
|
||||
!split
|
||||
===== Conjugate gradient method =====
|
||||
!bblock
|
||||
One can show that the solution $\hat{x}$ is also the unique minimizer of the quadratic form
|
||||
!bt
|
||||
\begin{equation*}
|
||||
f(\hat{x}) = \frac{1}{2}\hat{x}^T\hat{A}\hat{x} - \hat{x}^T \hat{x} , \quad \hat{x}\in\mathbf{R}^n.
|
||||
\end{equation*}
|
||||
!et
|
||||
This suggests taking the first basis vector $\hat{p}_1$
|
||||
to be the gradient of $f$ at $\hat{x}=\hat{x}_0$,
|
||||
which equals
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{A}\hat{x}_0-\hat{b},
|
||||
\end{equation*}
|
||||
!et
|
||||
and
|
||||
$\hat{x}_0=0$ it is equal $-\hat{b}$.
|
||||
The other vectors in the basis will be conjugate to the gradient,
|
||||
hence the name conjugate gradient method.
|
||||
!eblock
|
||||
|
||||
|
||||
!split
|
||||
===== Conjugate gradient method =====
|
||||
!bblock
|
||||
Let $\hat{r}_k$ be the residual at the $k$-th step:
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{r}_k=\hat{b}-\hat{A}\hat{x}_k.
|
||||
\end{equation*}
|
||||
!et
|
||||
Note that $\hat{r}_k$ is the negative gradient of $f$ at
|
||||
$\hat{x}=\hat{x}_k$,
|
||||
so the gradient descent method would be to move in the direction $\hat{r}_k$.
|
||||
Here, we insist that the directions $\hat{p}_k$ are conjugate to each other,
|
||||
so we take the direction closest to the gradient $\hat{r}_k$
|
||||
under the conjugacy constraint.
|
||||
This gives the following expression
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{p}_{k+1}=\hat{r}_k-\frac{\hat{p}_k^T \hat{A}\hat{r}_k}{\hat{p}_k^T\hat{A}\hat{p}_k} \hat{p}_k.
|
||||
\end{equation*}
|
||||
!et
|
||||
!eblock
|
||||
|
||||
!split
|
||||
===== Conjugate gradient method =====
|
||||
!bblock
|
||||
We can also compute the residual iteratively as
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{b}-\hat{A}\hat{x}_{k+1},
|
||||
\end{equation*}
|
||||
!et
|
||||
which equals
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{b}-\hat{A}(\hat{x}_k+\alpha_k\hat{p}_k),
|
||||
\end{equation*}
|
||||
!et
|
||||
or
|
||||
!bt
|
||||
\begin{equation*}
|
||||
(\hat{b}-\hat{A}\hat{x}_k)-\alpha_k\hat{A}\hat{p}_k,
|
||||
\end{equation*}
|
||||
!et
|
||||
which gives
|
||||
|
||||
!bt
|
||||
\begin{equation*}
|
||||
\hat{r}_{k+1}=\hat{r}_k-\hat{A}\hat{p}_{k},
|
||||
\end{equation*}
|
||||
!et
|
||||
!eblock
|
||||
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user