update
This commit is contained in:
@@ -59,7 +59,6 @@ Automatically generated HTML file from DocOnce source
|
||||
None,
|
||||
'revisiting-our-logistic-regression-case'),
|
||||
('The equations to solve', 2, None, 'the-equations-to-solve'),
|
||||
('To be added', 2, None, 'to-be-added'),
|
||||
("Solving using Newton-Raphson's method",
|
||||
2,
|
||||
None,
|
||||
@@ -161,6 +160,7 @@ Automatically generated HTML file from DocOnce source
|
||||
2,
|
||||
None,
|
||||
'using-gradient-descent-methods-limitations'),
|
||||
('To be added', 2, None, 'to-be-added'),
|
||||
('Friday October 1', 2, None, 'friday-october-1'),
|
||||
('Stochastic Gradient Descent',
|
||||
2,
|
||||
@@ -224,45 +224,45 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs006.html#optimization-the-central-part-of-any-machine-learning-algortithm" style="font-size: 80%;">Optimization, the central part of any Machine Learning algortithm</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs007.html#revisiting-our-logistic-regression-case" style="font-size: 80%;">Revisiting our Logistic Regression case</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs008.html#the-equations-to-solve" style="font-size: 80%;">The equations to solve</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs009.html#to-be-added" style="font-size: 80%;">To be added</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs010.html#solving-using-newton-raphson-s-method" style="font-size: 80%;">Solving using Newton-Raphson's method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs011.html#brief-reminder-on-newton-raphson-s-method" style="font-size: 80%;">Brief reminder on Newton-Raphson's method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs012.html#the-equations" style="font-size: 80%;">The equations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs013.html#simple-geometric-interpretation" style="font-size: 80%;">Simple geometric interpretation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs014.html#extending-to-more-than-one-variable" style="font-size: 80%;">Extending to more than one variable</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs015.html#steepest-descent" style="font-size: 80%;">Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs016.html#more-on-steepest-descent" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs017.html#the-ideal" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs018.html#the-sensitiveness-of-the-gradient-descent" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs019.html#convex-functions" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs020.html#convex-function" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs021.html#conditions-on-convex-functions" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs022.html#more-on-convex-functions" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs023.html#some-simple-problems" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs024.html#standard-steepest-descent" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs025.html#gradient-method" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs027.html#steepest-descent-method" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs027.html#steepest-descent-method" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs028.html#final-expressions" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs029.html#steepest-descent-example" style="font-size: 80%;">Steepest descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs037.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs037.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs037.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs037.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs034.html#conjugate-gradient-method-and-iterations" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs037.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs037.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs037.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs038.html#revisiting-our-first-homework" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs043.html#gradient-descent-example" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs040.html#the-derivative-of-the-cost-loss-function" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs041.html#the-hessian-matrix" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs042.html#simple-program" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs043.html#gradient-descent-example" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs044.html#and-a-corresponding-example-using-_scikit-learn_" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs045.html#gradient-descent-and-ridge" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs046.html#program-example-for-gradient-descent-with-ridge-regression" style="font-size: 80%;">Program example for gradient descent with Ridge Regression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs047.html#using-gradient-descent-methods-limitations" style="font-size: 80%;">Using gradient descent methods, limitations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs009.html#solving-using-newton-raphson-s-method" style="font-size: 80%;">Solving using Newton-Raphson's method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs010.html#brief-reminder-on-newton-raphson-s-method" style="font-size: 80%;">Brief reminder on Newton-Raphson's method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs011.html#the-equations" style="font-size: 80%;">The equations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs012.html#simple-geometric-interpretation" style="font-size: 80%;">Simple geometric interpretation</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs013.html#extending-to-more-than-one-variable" style="font-size: 80%;">Extending to more than one variable</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs014.html#steepest-descent" style="font-size: 80%;">Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs015.html#more-on-steepest-descent" style="font-size: 80%;">More on Steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs016.html#the-ideal" style="font-size: 80%;">The ideal</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs017.html#the-sensitiveness-of-the-gradient-descent" style="font-size: 80%;">The sensitiveness of the gradient descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs018.html#convex-functions" style="font-size: 80%;">Convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs019.html#convex-function" style="font-size: 80%;">Convex function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs020.html#conditions-on-convex-functions" style="font-size: 80%;">Conditions on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs021.html#more-on-convex-functions" style="font-size: 80%;">More on convex functions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs022.html#some-simple-problems" style="font-size: 80%;">Some simple problems</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs023.html#standard-steepest-descent" style="font-size: 80%;">Standard steepest descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs024.html#gradient-method" style="font-size: 80%;">Gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs026.html#steepest-descent-method" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs026.html#steepest-descent-method" style="font-size: 80%;">Steepest descent method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs027.html#final-expressions" style="font-size: 80%;">Final expressions</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs028.html#steepest-descent-example" style="font-size: 80%;">Steepest descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs036.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs036.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs036.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs036.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs033.html#conjugate-gradient-method-and-iterations" style="font-size: 80%;">Conjugate gradient method and iterations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs036.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs036.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs036.html#conjugate-gradient-method" style="font-size: 80%;">Conjugate gradient method</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs037.html#revisiting-our-first-homework" style="font-size: 80%;">Revisiting our first homework</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs042.html#gradient-descent-example" style="font-size: 80%;">Gradient descent example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs039.html#the-derivative-of-the-cost-loss-function" style="font-size: 80%;">The derivative of the cost/loss function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs040.html#the-hessian-matrix" style="font-size: 80%;">The Hessian matrix</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs041.html#simple-program" style="font-size: 80%;">Simple program</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs042.html#gradient-descent-example" style="font-size: 80%;">Gradient Descent Example</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs043.html#and-a-corresponding-example-using-_scikit-learn_" style="font-size: 80%;">And a corresponding example using <b>scikit-learn</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs044.html#gradient-descent-and-ridge" style="font-size: 80%;">Gradient descent and Ridge</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs045.html#program-example-for-gradient-descent-with-ridge-regression" style="font-size: 80%;">Program example for gradient descent with Ridge Regression</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs046.html#using-gradient-descent-methods-limitations" style="font-size: 80%;">Using gradient descent methods, limitations</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs047.html#to-be-added" style="font-size: 80%;">To be added</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs048.html#friday-october-1" style="font-size: 80%;">Friday October 1</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs049.html#stochastic-gradient-descent" style="font-size: 80%;">Stochastic Gradient Descent</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week39-bs050.html#computation-of-gradients" style="font-size: 80%;">Computation of gradients</a></li>
|
||||
@@ -306,7 +306,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Sep 29, 2021</h4></center> <!-- date -->
|
||||
<center><h4>Sep 30, 2021</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
|
||||
|
||||
@@ -148,7 +148,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p> <br>
|
||||
<center><h4>Sep 29, 2021</h4></center> <!-- date -->
|
||||
<center><h4>Sep 30, 2021</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
|
||||
@@ -170,6 +170,10 @@ MathJax.Hub.Config({
|
||||
|
||||
See <a href="https://compphysics.github.io/MachineLearning/doc/web/course.html" target="_blank">lecture notes for week 39</a>.
|
||||
For a good discussion on gradient methods, see Goodfellow et al section 4.3-4.5 and chapter 8. We will come back to the latter chapter in our discussion of Neural networks as well.
|
||||
|
||||
<p>
|
||||
<b>For project 1, chapter 5 of Goodfellow et al is a good read, in particular sections 5.1-5.55 and 5.7-5.11</b>.
|
||||
These sections summarize neatly what we have done till now and point to what is coming with respect to deep learning.
|
||||
</section>
|
||||
|
||||
|
||||
@@ -444,17 +448,6 @@ This defines what is called the Hessian matrix.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="to-be-added">To be added </h2>
|
||||
|
||||
<p>
|
||||
We will add here an example which computes the likelihood \( p_i \), sets up the gradient and the Hessian matrix.
|
||||
|
||||
<p>
|
||||
Make link with linear regression and the Hessian matrix from linear regression.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="solving-using-newton-raphson-s-method">Solving using Newton-Raphson's method </h2>
|
||||
|
||||
@@ -1662,6 +1655,14 @@ plt.show()
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="to-be-added">To be added </h2>
|
||||
|
||||
<p>
|
||||
We will add here an example which computes the likelihood \( p_i \), sets up the gradient and the Hessian matrix.
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="friday-october-1">Friday October 1 </h2>
|
||||
</section>
|
||||
|
||||
@@ -79,7 +79,6 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
None,
|
||||
'revisiting-our-logistic-regression-case'),
|
||||
('The equations to solve', 2, None, 'the-equations-to-solve'),
|
||||
('To be added', 2, None, 'to-be-added'),
|
||||
("Solving using Newton-Raphson's method",
|
||||
2,
|
||||
None,
|
||||
@@ -181,6 +180,7 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
2,
|
||||
None,
|
||||
'using-gradient-descent-methods-limitations'),
|
||||
('To be added', 2, None, 'to-be-added'),
|
||||
('Friday October 1', 2, None, 'friday-october-1'),
|
||||
('Stochastic Gradient Descent',
|
||||
2,
|
||||
@@ -240,7 +240,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Sep 29, 2021</h4></center> <!-- date -->
|
||||
<center><h4>Sep 30, 2021</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -256,6 +256,10 @@ MathJax.Hub.Config({
|
||||
See <a href="https://compphysics.github.io/MachineLearning/doc/web/course.html" target="_blank">lecture notes for week 39</a>.
|
||||
For a good discussion on gradient methods, see Goodfellow et al section 4.3-4.5 and chapter 8. We will come back to the latter chapter in our discussion of Neural networks as well.
|
||||
|
||||
<p>
|
||||
<b>For project 1, chapter 5 of Goodfellow et al is a good read, in particular sections 5.1-5.55 and 5.7-5.11</b>.
|
||||
These sections summarize neatly what we have done till now and point to what is coming with respect to deep learning.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
@@ -523,17 +527,6 @@ This defines what is called the Hessian matrix.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="to-be-added">To be added </h2>
|
||||
|
||||
<p>
|
||||
We will add here an example which computes the likelihood \( p_i \), sets up the gradient and the Hessian matrix.
|
||||
|
||||
<p>
|
||||
Make link with linear regression and the Hessian matrix from linear regression.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="solving-using-newton-raphson-s-method">Solving using Newton-Raphson's method </h2>
|
||||
|
||||
<p>
|
||||
@@ -1640,6 +1633,14 @@ plt.show()
|
||||
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="to-be-added">To be added </h2>
|
||||
|
||||
<p>
|
||||
We will add here an example which computes the likelihood \( p_i \), sets up the gradient and the Hessian matrix.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="friday-october-1">Friday October 1 </h2>
|
||||
|
||||
<p>
|
||||
|
||||
@@ -84,7 +84,6 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
None,
|
||||
'revisiting-our-logistic-regression-case'),
|
||||
('The equations to solve', 2, None, 'the-equations-to-solve'),
|
||||
('To be added', 2, None, 'to-be-added'),
|
||||
("Solving using Newton-Raphson's method",
|
||||
2,
|
||||
None,
|
||||
@@ -186,6 +185,7 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
2,
|
||||
None,
|
||||
'using-gradient-descent-methods-limitations'),
|
||||
('To be added', 2, None, 'to-be-added'),
|
||||
('Friday October 1', 2, None, 'friday-october-1'),
|
||||
('Stochastic Gradient Descent',
|
||||
2,
|
||||
@@ -245,7 +245,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Sep 29, 2021</h4></center> <!-- date -->
|
||||
<center><h4>Sep 30, 2021</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -261,6 +261,10 @@ MathJax.Hub.Config({
|
||||
See <a href="https://compphysics.github.io/MachineLearning/doc/web/course.html" target="_blank">lecture notes for week 39</a>.
|
||||
For a good discussion on gradient methods, see Goodfellow et al section 4.3-4.5 and chapter 8. We will come back to the latter chapter in our discussion of Neural networks as well.
|
||||
|
||||
<p>
|
||||
<b>For project 1, chapter 5 of Goodfellow et al is a good read, in particular sections 5.1-5.55 and 5.7-5.11</b>.
|
||||
These sections summarize neatly what we have done till now and point to what is coming with respect to deep learning.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
@@ -528,17 +532,6 @@ This defines what is called the Hessian matrix.
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="to-be-added">To be added </h2>
|
||||
|
||||
<p>
|
||||
We will add here an example which computes the likelihood \( p_i \), sets up the gradient and the Hessian matrix.
|
||||
|
||||
<p>
|
||||
Make link with linear regression and the Hessian matrix from linear regression.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="solving-using-newton-raphson-s-method">Solving using Newton-Raphson's method </h2>
|
||||
|
||||
<p>
|
||||
@@ -1645,6 +1638,14 @@ plt<span style="color: #666666">.</span>show()
|
||||
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="to-be-added">To be added </h2>
|
||||
|
||||
<p>
|
||||
We will add here an example which computes the likelihood \( p_i \), sets up the gradient and the Hessian matrix.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="friday-october-1">Friday October 1 </h2>
|
||||
|
||||
<p>
|
||||
|
||||
Binary file not shown.
@@ -10,7 +10,7 @@
|
||||
"<!-- Author: --> \n",
|
||||
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
|
||||
"\n",
|
||||
"Date: **Sep 29, 2021**\n",
|
||||
"Date: **Sep 30, 2021**\n",
|
||||
"\n",
|
||||
"Copyright 1999-2021, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
|
||||
"\n",
|
||||
@@ -27,6 +27,9 @@
|
||||
"See [lecture notes for week 39](https://compphysics.github.io/MachineLearning/doc/web/course.html).\n",
|
||||
"For a good discussion on gradient methods, see Goodfellow et al section 4.3-4.5 and chapter 8. We will come back to the latter chapter in our discussion of Neural networks as well.\n",
|
||||
"\n",
|
||||
"**For project 1, chapter 5 of Goodfellow et al is a good read, in particular sections 5.1-5.55 and 5.7-5.11**.\n",
|
||||
"These sections summarize neatly what we have done till now and point to what is coming with respect to deep learning. \n",
|
||||
"\n",
|
||||
"## Thursday September 30\n",
|
||||
"\n",
|
||||
"[Overview Video, why do we care about gradient methods?](https://www.uio.no/studier/emner/matnat/fys/FYS-STK3155/h20/forelesningsvideoer/OverarchingAimsWeek39.mp4?vrtx=view-as-webpage)\n",
|
||||
@@ -334,11 +337,6 @@
|
||||
"This defines what is called the Hessian matrix.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## To be added\n",
|
||||
"\n",
|
||||
"We will add here an example which computes the likelihood $p_i$, sets up the gradient and the Hessian matrix.\n",
|
||||
"\n",
|
||||
"Make link with linear regression and the Hessian matrix from linear regression.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
@@ -1891,6 +1889,13 @@
|
||||
"\n",
|
||||
"* GD can take exponential time to escape saddle points, even with random initialization. As we mentioned, GD is extremely sensitive to initial condition since it determines the particular local minimum GD would eventually reach. However, even with a good initialization scheme, through the introduction of randomness, GD can still take exponential time to escape saddle points.\n",
|
||||
"\n",
|
||||
"## To be added\n",
|
||||
"\n",
|
||||
"We will add here an example which computes the likelihood $p_i$, sets up the gradient and the Hessian matrix.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"## Friday October 1\n",
|
||||
"\n",
|
||||
"\n",
|
||||
|
||||
@@ -14,6 +14,9 @@ DATE: today
|
||||
See "lecture notes for week 39":"https://compphysics.github.io/MachineLearning/doc/web/course.html".
|
||||
For a good discussion on gradient methods, see Goodfellow et al section 4.3-4.5 and chapter 8. We will come back to the latter chapter in our discussion of Neural networks as well.
|
||||
|
||||
_For project 1, chapter 5 of Goodfellow et al is a good read, in particular sections 5.1-5.55 and 5.7-5.11_.
|
||||
These sections summarize neatly what we have done till now and point to what is coming with respect to deep learning.
|
||||
|
||||
!split
|
||||
===== Thursday September 30 =====
|
||||
|
||||
@@ -264,12 +267,6 @@ $p(y_i\vert x_i,\bm{\beta})(1-p(y_i\vert x_i,\bm{\beta})$, we can obtain a compa
|
||||
This defines what is called the Hessian matrix.
|
||||
|
||||
|
||||
!split
|
||||
===== To be added =====
|
||||
|
||||
We will add here an example which computes the likelihood $p_i$, sets up the gradient and the Hessian matrix.
|
||||
|
||||
Make link with linear regression and the Hessian matrix from linear regression.
|
||||
|
||||
|
||||
|
||||
@@ -1231,6 +1228,14 @@ plt.show()
|
||||
* GD can take exponential time to escape saddle points, even with random initialization. As we mentioned, GD is extremely sensitive to initial condition since it determines the particular local minimum GD would eventually reach. However, even with a good initialization scheme, through the introduction of randomness, GD can still take exponential time to escape saddle points.
|
||||
|
||||
|
||||
!split
|
||||
===== To be added =====
|
||||
|
||||
We will add here an example which computes the likelihood $p_i$, sets up the gradient and the Hessian matrix.
|
||||
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== Friday October 1 =====
|
||||
|
||||
|
||||
Reference in New Issue
Block a user