diff --git a/doc/pub/week38/html/._week38-bs000.html b/doc/pub/week38/html/._week38-bs000.html
index 5321c2f13..152fc4214 100644
--- a/doc/pub/week38/html/._week38-bs000.html
+++ b/doc/pub/week38/html/._week38-bs000.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
19
20
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs011.html b/doc/pub/week38/html/._week38-bs011.html
index a2a36a0d1..e65be3a79 100644
--- a/doc/pub/week38/html/._week38-bs011.html
+++ b/doc/pub/week38/html/._week38-bs011.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -480,7 +374,7 @@ $$
20
21
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs012.html b/doc/pub/week38/html/._week38-bs012.html
index fd50ab60c..3e88a9b5a 100644
--- a/doc/pub/week38/html/._week38-bs012.html
+++ b/doc/pub/week38/html/._week38-bs012.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -506,7 +400,7 @@ $$
21
22
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs013.html b/doc/pub/week38/html/._week38-bs013.html
index 981ac928b..913629684 100644
--- a/doc/pub/week38/html/._week38-bs013.html
+++ b/doc/pub/week38/html/._week38-bs013.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -466,7 +360,7 @@ What does this mean? And why do we insist on all this? Let us look at some examp
22
23
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs014.html b/doc/pub/week38/html/._week38-bs014.html
index fa21e3ef4..f1c81bc3f 100644
--- a/doc/pub/week38/html/._week38-bs014.html
+++ b/doc/pub/week38/html/._week38-bs014.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -573,7 +467,7 @@ It means that, when scaling the design matrix and the outputs/targets, by subtra
23
24
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs015.html b/doc/pub/week38/html/._week38-bs015.html
index d40055597..7e44f680a 100644
--- a/doc/pub/week38/html/._week38-bs015.html
+++ b/doc/pub/week38/html/._week38-bs015.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -524,7 +418,7 @@ Let us see how we can change this code by zero centering (thanks to Stian Bilek
24
25
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs016.html b/doc/pub/week38/html/._week38-bs016.html
index 7f0b084f4..4590f9fc8 100644
--- a/doc/pub/week38/html/._week38-bs016.html
+++ b/doc/pub/week38/html/._week38-bs016.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -542,7 +436,7 @@ The next example is indeed an example where all these discussions about the role
25
26
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs017.html b/doc/pub/week38/html/._week38-bs017.html
index 81818dadc..3a912b331 100644
--- a/doc/pub/week38/html/._week38-bs017.html
+++ b/doc/pub/week38/html/._week38-bs017.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -498,7 +392,7 @@ the coupling constant to achieve this.
26
27
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs018.html b/doc/pub/week38/html/._week38-bs018.html
index f0f4a3def..73608dbd4 100644
--- a/doc/pub/week38/html/._week38-bs018.html
+++ b/doc/pub/week38/html/._week38-bs018.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -491,7 +385,7 @@ X_train, X_test, y_train, y_test = train_tes
27
28
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs019.html b/doc/pub/week38/html/._week38-bs019.html
index 29fa795a5..9d93c8742 100644
--- a/doc/pub/week38/html/._week38-bs019.html
+++ b/doc/pub/week38/html/._week38-bs019.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -489,7 +383,7 @@ beta = ols_inv(X_train_own, y_train)
28
29
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs020.html b/doc/pub/week38/html/._week38-bs020.html
index 8c4fc9923..61025097f 100644
--- a/doc/pub/week38/html/._week38-bs020.html
+++ b/doc/pub/week38/html/._week38-bs020.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -528,7 +422,7 @@ In this case our matrix inversion was actually possible. The obvious question no
29
30
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs021.html b/doc/pub/week38/html/._week38-bs021.html
index 10aaddbc0..77888069f 100644
--- a/doc/pub/week38/html/._week38-bs021.html
+++ b/doc/pub/week38/html/._week38-bs021.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -575,7 +469,7 @@ The results perfectly with our previous discussion where we used our own code.
30
31
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs022.html b/doc/pub/week38/html/._week38-bs022.html
index d1b4e1c9a..130cd0508 100644
--- a/doc/pub/week38/html/._week38-bs022.html
+++ b/doc/pub/week38/html/._week38-bs022.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -476,7 +370,7 @@ plt.show()
31
32
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs023.html b/doc/pub/week38/html/._week38-bs023.html
index c7b61759f..6a639c67b 100644
--- a/doc/pub/week38/html/._week38-bs023.html
+++ b/doc/pub/week38/html/._week38-bs023.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -479,7 +373,7 @@ constant as opposed to ridge and OLS. We get a sparse solution with
32
33
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs024.html b/doc/pub/week38/html/._week38-bs024.html
index 7ac38c92b..a3285fb77 100644
--- a/doc/pub/week38/html/._week38-bs024.html
+++ b/doc/pub/week38/html/._week38-bs024.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -495,7 +389,7 @@ much. Ridge is more stable over a larger range of values for
33
34
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs025.html b/doc/pub/week38/html/._week38-bs025.html
index 3b5c8f105..05d315d1d 100644
--- a/doc/pub/week38/html/._week38-bs025.html
+++ b/doc/pub/week38/html/._week38-bs025.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -493,7 +387,7 @@ other models for all values of \( \lambda \).
34
35
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs026.html b/doc/pub/week38/html/._week38-bs026.html
index 8019c1d72..e7482877a 100644
--- a/doc/pub/week38/html/._week38-bs026.html
+++ b/doc/pub/week38/html/._week38-bs026.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -459,7 +353,7 @@ simple recipe for fitting our data.
35
36
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs027.html b/doc/pub/week38/html/._week38-bs027.html
index ec191791b..92916b3b6 100644
--- a/doc/pub/week38/html/._week38-bs027.html
+++ b/doc/pub/week38/html/._week38-bs027.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -464,7 +358,7 @@ failure etc.
36
37
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs028.html b/doc/pub/week38/html/._week38-bs028.html
index c17ab46b3..e4758edaa 100644
--- a/doc/pub/week38/html/._week38-bs028.html
+++ b/doc/pub/week38/html/._week38-bs028.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -462,7 +356,7 @@ models, as we will see later.
37
38
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs029.html b/doc/pub/week38/html/._week38-bs029.html
index 1633e4aa3..d5bd4ebbf 100644
--- a/doc/pub/week38/html/._week38-bs029.html
+++ b/doc/pub/week38/html/._week38-bs029.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -468,7 +362,7 @@ $$
38
39
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs030.html b/doc/pub/week38/html/._week38-bs030.html
index ecae8af76..c39c8ffc9 100644
--- a/doc/pub/week38/html/._week38-bs030.html
+++ b/doc/pub/week38/html/._week38-bs030.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -465,7 +359,7 @@ where \( \boldsymbol{y} \) is a vector representing the possible outcomes, \( \b
39
40
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs031.html b/doc/pub/week38/html/._week38-bs031.html
index 806addd4a..17392aed4 100644
--- a/doc/pub/week38/html/._week38-bs031.html
+++ b/doc/pub/week38/html/._week38-bs031.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -463,7 +357,7 @@ the probability of a given category. This leads us to the logistic function.
40
41
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs032.html b/doc/pub/week38/html/._week38-bs032.html
index d4bbd8380..657669d6b 100644
--- a/doc/pub/week38/html/._week38-bs032.html
+++ b/doc/pub/week38/html/._week38-bs032.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -507,7 +401,7 @@ plt.show()
41
42
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs033.html b/doc/pub/week38/html/._week38-bs033.html
index aae61091f..64e693c7e 100644
--- a/doc/pub/week38/html/._week38-bs033.html
+++ b/doc/pub/week38/html/._week38-bs033.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -480,7 +374,7 @@ representing the probability for finding a value of \( y_i \) with a given
42
43
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs034.html b/doc/pub/week38/html/._week38-bs034.html
index 850014b15..75e652a46 100644
--- a/doc/pub/week38/html/._week38-bs034.html
+++ b/doc/pub/week38/html/._week38-bs034.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -464,7 +358,7 @@ Note that \( 1-p(t)= p(-t) \).
43
44
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs035.html b/doc/pub/week38/html/._week38-bs035.html
index 4acbc04d7..90a8f9605 100644
--- a/doc/pub/week38/html/._week38-bs035.html
+++ b/doc/pub/week38/html/._week38-bs035.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -507,7 +401,7 @@ plt.show()
44
45
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs036.html b/doc/pub/week38/html/._week38-bs036.html
index e6c59bd28..f799b6e68 100644
--- a/doc/pub/week38/html/._week38-bs036.html
+++ b/doc/pub/week38/html/._week38-bs036.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -463,7 +357,7 @@ $$
45
46
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs037.html b/doc/pub/week38/html/._week38-bs037.html
index 63291841f..8246b7af5 100644
--- a/doc/pub/week38/html/._week38-bs037.html
+++ b/doc/pub/week38/html/._week38-bs037.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -464,7 +358,7 @@ $$
46
47
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs038.html b/doc/pub/week38/html/._week38-bs038.html
index 58c4b249a..c0d924b40 100644
--- a/doc/pub/week38/html/._week38-bs038.html
+++ b/doc/pub/week38/html/._week38-bs038.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -462,7 +356,7 @@ in practice we often supplement the cross-entropy with additional regularization
47
48
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs039.html b/doc/pub/week38/html/._week38-bs039.html
index e177ad1ab..7667eb9f8 100644
--- a/doc/pub/week38/html/._week38-bs039.html
+++ b/doc/pub/week38/html/._week38-bs039.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -463,7 +357,7 @@ $$
48
49
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs040.html b/doc/pub/week38/html/._week38-bs040.html
index 6010937b7..5bcc161ff 100644
--- a/doc/pub/week38/html/._week38-bs040.html
+++ b/doc/pub/week38/html/._week38-bs040.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -464,7 +358,7 @@ $$
49
50
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs041.html b/doc/pub/week38/html/._week38-bs041.html
index 1024a0957..6e1fe75fc 100644
--- a/doc/pub/week38/html/._week38-bs041.html
+++ b/doc/pub/week38/html/._week38-bs041.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -457,7 +351,7 @@ $$
50
51
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/._week38-bs042.html b/doc/pub/week38/html/._week38-bs042.html
index 6e8f03803..9b2d90522 100644
--- a/doc/pub/week38/html/._week38-bs042.html
+++ b/doc/pub/week38/html/._week38-bs042.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -469,7 +363,7 @@ and the model is specified in term of \( K-1 \) so-called log-odds or
51
52
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/week38-bs.html b/doc/pub/week38/html/week38-bs.html
index 5321c2f13..152fc4214 100644
--- a/doc/pub/week38/html/week38-bs.html
+++ b/doc/pub/week38/html/week38-bs.html
@@ -164,6 +164,7 @@ Automatically generated HTML file from DocOnce source
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -197,85 +198,7 @@ Automatically generated HTML file from DocOnce source
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -361,48 +284,19 @@ MathJax.Hub.Config({
Using the correlation matrix
Discussing the correlation data
Other measures in classification studies: Cancer Data again
- Optimization, the central part of any Machine Learning algortithm
- Revisiting our Logistic Regression case
- The equations to solve
- Solving using Newton-Raphson's method
- Brief reminder on Newton-Raphson's method
- The equations
- Simple geometric interpretation
- Extending to more than one variable
- Steepest descent
- More on Steepest descent
- The ideal
- The sensitiveness of the gradient descent
- Convex functions
- Convex function
- Conditions on convex functions
- More on convex functions
- Some simple problems
- Friday September 25
- Standard steepest descent
- Gradient method
- Steepest descent method
- Steepest descent method
- Final expressions
- Steepest descent example
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method and iterations
- Conjugate gradient method
- Conjugate gradient method
- Conjugate gradient method
- Revisiting some of our first Linear Regression Encounters
- Gradient descent example
- The derivative of the cost/loss function
- The Hessian matrix
- Simple program
- Gradient Descent Example
- And a corresponding example using scikit-learn
- Gradient descent and Ridge
- Program example for gradient descent with Ridge Regression
- Using gradient descent methods, limitations
+ Friday September 25
+ Optimization, the central part of any Machine Learning algortithm
+ Revisiting our Logistic Regression case
+ The equations to solve
+ Solving using Newton-Raphson's method
+ Brief reminder on Newton-Raphson's method
+ The equations
+ Simple geometric interpretation
+ Extending to more than one variable
+ Steepest descent
+ More on Steepest descent
+ The ideal
+ The sensitiveness of the gradient descent
@@ -437,7 +331,7 @@ MathJax.Hub.Config({
[2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University
-
Sep 24, 2021
+Sep 28, 2021
@@ -461,7 +355,7 @@ MathJax.Hub.Config({
9
10
...
- 91
+ 62
»
diff --git a/doc/pub/week38/html/week38-reveal.html b/doc/pub/week38/html/week38-reveal.html
index f451e3a23..f469ac57f 100644
--- a/doc/pub/week38/html/week38-reveal.html
+++ b/doc/pub/week38/html/week38-reveal.html
@@ -148,7 +148,7 @@ MathJax.Hub.Config({
[2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University
-
Sep 24, 2021
+Sep 28, 2021
@@ -2327,6 +2327,11 @@ plt.show()
+
+
+
Optimization, the central part of any Machine Learning algortithm
@@ -2665,955 +2670,6 @@ randomness. One such method is that of Stochastic Gradient Descent
-
-Convex functions
-
-
-Ideally we want our cost/loss function to be convex(concave).
-
-
-First we give the definition of a convex set: A set \( C \) in
-\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and
-all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to
-C. Geometrically this means that every point on the line segment
-connecting \( x \) and \( y \) is in \( C \) as discussed below.
-
-
-The convex subsets of \( \mathbb{R} \) are the intervals of
-\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the
-regular polygons (triangles, rectangles, pentagons, etc...).
-
-
-
-
-Convex function
-
-
-Convex function: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if
-$$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$
-
for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below.
-
-
-
-
-Conditions on convex functions
-
-
-In the following we state first and second-order conditions which
-ensures convexity of a function \( f \). We write \( D_f \) to denote the
-domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more
-details and proofs we refer to: S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press.
-
-
-
-
First order condition
-
-Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for
-all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \)
-is a convex set and
-$$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$
-
holds
-for all \( x,y \in D_f \). This condition means that for a convex function
-the first order Taylor expansion (right hand side above) at any point
-a global under estimator of the function. To convince yourself you can
-make a drawing of \( f(x) = x^2+1 \) and draw the tangent line to \( f(x) \) and
-note that it is always below the graph.
-
-
-
-
-
Second order condition
-
-Assume that \( f \) is twice
-differentiable, i.e the Hessian matrix exists at each point in
-\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its
-Hessian is positive semi-definite for all \( x\in D_f \). For a
-single-variable function this reduces to \( f''(x) \geq 0 \). Geometrically this means that \( f \) has nonnegative curvature
-everywhere.
-
-
-
-This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
-
-
-
-
-More on convex functions
-
-
-The next result is of great importance to us and the reason why we are
-going on about convex functions. In machine learning we frequently
-have to minimize a loss/cost function in order to find the best
-parameters for the model we are considering.
-
-
-Ideally we want the
-global minimum (for high-dimensional models it is hard to know
-if we have local or global minimum). However, if the cost/loss function
-is convex the following result provides invaluable information:
-
-
-
-
Any minimum is global for convex functions
-
-Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \)
-is minimal, where \( f \) is convex and differentiable. Then, any point
-\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum.
-
-
-
-This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
-
-
-
-
-Some simple problems
-
-
-- Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.
-- Using the second order condition show that the following functions are convex on the specified domain.
-
-
- - \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).
- - \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).
-
-- Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.
-- A norm is any function that satisfy the following properties
-
-
- - \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).
- - \( f(x+y) \leq f(x) + f(y) \)
- - \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)
-
-
-
-
-
-Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
-
-
-
-
-
-
-
-Standard steepest descent
-
-
-Before we proceed, we would like to discuss the approach called the
-standard Steepest descent (different from the above steepest descent discussion), which again leads to us having to be able
-to compute a matrix. It belongs to the class of Conjugate Gradient methods (CG).
-
-
-The success of the CG method
-for finding solutions of non-linear problems is based on the theory
-of conjugate gradients for linear systems of equations. It belongs to
-the class of iterative methods for solving problems from linear
-algebra of the type
-
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{x} = \boldsymbol{b}.
-\end{equation*}
-$$
-
-
-
-In the iterative process we end up with a problem like
-
-
-$$
-\begin{equation*}
- \boldsymbol{r}= \boldsymbol{b}-\boldsymbol{A}\boldsymbol{x},
-\end{equation*}
-$$
-
-
-where \( \boldsymbol{r} \) is the so-called residual or error in the iterative process.
-
-
-When we have found the exact solution, \( \boldsymbol{r}=0 \).
-
-
-
-
-Gradient method
-
-
-The residual is zero when we reach the minimum of the quadratic equation
-
-$$
-\begin{equation*}
- P(\boldsymbol{x})=\frac{1}{2}\boldsymbol{x}^T\boldsymbol{A}\boldsymbol{x} - \boldsymbol{x}^T\boldsymbol{b},
-\end{equation*}
-$$
-
-
-
-with the constraint that the matrix \( \boldsymbol{A} \) is positive definite and
-symmetric. This defines also the Hessian and we want it to be positive definite.
-
-
-
-
-Steepest descent method
-
-
-We denote the initial guess for \( \boldsymbol{x} \) as \( \boldsymbol{x}_0 \).
-We can assume without loss of generality that
-
-$$
-\begin{equation*}
-\boldsymbol{x}_0=0,
-\end{equation*}
-$$
-
-
-or consider the system
-
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{z} = \boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_0,
-\end{equation*}
-$$
-
-
-instead.
-
-
-
-
-Steepest descent method
-
-
-
-One can show that the solution \( \boldsymbol{x} \) is also the unique minimizer of the quadratic form
-
-$$
-\begin{equation*}
- f(\boldsymbol{x}) = \frac{1}{2}\boldsymbol{x}^T\boldsymbol{A}\boldsymbol{x} - \boldsymbol{x}^T \boldsymbol{x} , \quad \boldsymbol{x}\in\mathbf{R}^n.
-\end{equation*}
-$$
-
-
-This suggests taking the first basis vector \( \boldsymbol{r}_1 \) (see below for definition)
-to be the gradient of \( f \) at \( \boldsymbol{x}=\boldsymbol{x}_0 \),
-which equals
-
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{x}_0-\boldsymbol{b},
-\end{equation*}
-$$
-
-
-and
-\( \boldsymbol{x}_0=0 \) it is equal \( -\boldsymbol{b} \).
-
-
-
-
-
-
-
-Final expressions
-
-
-
-We can compute the residual iteratively as
-
-$$
-\begin{equation*}
-\boldsymbol{r}_{k+1}=\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_{k+1},
- \end{equation*}
-$$
-
-
-which equals
-
-$$
-\begin{equation*}
-\boldsymbol{b}-\boldsymbol{A}(\boldsymbol{x}_k+\alpha_k\boldsymbol{r}_k),
- \end{equation*}
-$$
-
-
-or
-
-$$
-\begin{equation*}
-(\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_k)-\alpha_k\boldsymbol{A}\boldsymbol{r}_k,
- \end{equation*}
-$$
-
-
-which gives
-
-
-$$
-\alpha_k = \frac{\boldsymbol{r}_k^T\boldsymbol{r}_k}{\boldsymbol{r}_k^T\boldsymbol{A}\boldsymbol{r}_k}
-$$
-
-
-leading to the iterative scheme
-
-$$
-\begin{equation*}
-\boldsymbol{x}_{k+1}=\boldsymbol{x}_k-\alpha_k\boldsymbol{r}_{k},
- \end{equation*}
-$$
-
-
-
-
-
-
-Steepest descent example
-
-
-
-
-
import numpy as np
-import numpy.linalg as la
-
-import scipy.optimize as sopt
-
-import matplotlib.pyplot as pt
-from mpl_toolkits.mplot3d import axes3d
-
-def f(x):
- return 0.5*x[0]**2 + 2.5*x[1]**2
-
-def df(x):
- return np.array([x[0], 5*x[1]])
-
-fig = pt.figure()
-ax = fig.gca(projection="3d")
-
-xmesh, ymesh = np.mgrid[-2:2:50j,-2:2:50j]
-fmesh = f(np.array([xmesh, ymesh]))
-ax.plot_surface(xmesh, ymesh, fmesh)
-
-
-And then as countor plot
-
-
-
-
pt.axis("equal")
-pt.contour(xmesh, ymesh, fmesh)
-guesses = [np.array([2, 2./5])]
-
-
-Find guesses
-
-
-
-
x = guesses[-1]
-s = -df(x)
-
-
-Run it!
-
-
-
-
def f1d(alpha):
- return f(x + alpha*s)
-
-alpha_opt = sopt.golden(f1d)
-next_guess = x + alpha_opt * s
-guesses.append(next_guess)
-print(next_guess)
-
-
-What happened?
-
-
-
-
pt.axis("equal")
-pt.contour(xmesh, ymesh, fmesh, 50)
-it_array = np.array(guesses)
-pt.plot(it_array.T[0], it_array.T[1], "x-")
-
-
-
-
-
-Conjugate gradient method
-
-
-
-In the CG method we define so-called conjugate directions and two vectors
-\( \boldsymbol{s} \) and \( \boldsymbol{t} \)
-are said to be
-conjugate if
-
-$$
-\begin{equation*}
-\boldsymbol{s}^T\boldsymbol{A}\boldsymbol{t}= 0.
-\end{equation*}
-$$
-
-
-The philosophy of the CG method is to perform searches in various conjugate directions
-of our vectors \( \boldsymbol{x}_i \) obeying the above criterion, namely
-
-$$
-\begin{equation*}
-\boldsymbol{x}_i^T\boldsymbol{A}\boldsymbol{x}_j= 0.
-\end{equation*}
-$$
-
-
-Two vectors are conjugate if they are orthogonal with respect to
-this inner product. Being conjugate is a symmetric relation: if \( \boldsymbol{s} \) is conjugate to \( \boldsymbol{t} \), then \( \boldsymbol{t} \) is conjugate to \( \boldsymbol{s} \).
-
-
-
-
-
-Conjugate gradient method
-
-
-
-An example is given by the eigenvectors of the matrix
-
-$$
-\begin{equation*}
-\boldsymbol{v}_i^T\boldsymbol{A}\boldsymbol{v}_j= \lambda\boldsymbol{v}_i^T\boldsymbol{v}_j,
-\end{equation*}
-$$
-
-
-which is zero unless \( i=j \).
-
-
-
-
-
-Conjugate gradient method
-
-
-
-Assume now that we have a symmetric positive-definite matrix \( \boldsymbol{A} \) of size
-\( n\times n \). At each iteration \( i+1 \) we obtain the conjugate direction of a vector
-
-$$
-\begin{equation*}
-\boldsymbol{x}_{i+1}=\boldsymbol{x}_{i}+\alpha_i\boldsymbol{p}_{i}.
-\end{equation*}
-$$
-
-
-We assume that \( \boldsymbol{p}_{i} \) is a sequence of \( n \) mutually conjugate directions.
-Then the \( \boldsymbol{p}_{i} \) form a basis of \( R^n \) and we can expand the solution
-$ \boldsymbol{A}\boldsymbol{x} = \boldsymbol{b}$ in this basis, namely
-
-
-$$
-\begin{equation*}
- \boldsymbol{x} = \sum^{n}_{i=1} \alpha_i \boldsymbol{p}_i.
-\end{equation*}
-$$
-
-
-
-
-
-
-Conjugate gradient method
-
-
-
-The coefficients are given by
-
-$$
-\begin{equation*}
- \mathbf{A}\mathbf{x} = \sum^{n}_{i=1} \alpha_i \mathbf{A} \mathbf{p}_i = \mathbf{b}.
-\end{equation*}
-$$
-
-
-Multiplying with \( \boldsymbol{p}_k^T \) from the left gives
-
-
-$$
-\begin{equation*}
- \boldsymbol{p}_k^T \boldsymbol{A}\boldsymbol{x} = \sum^{n}_{i=1} \alpha_i\boldsymbol{p}_k^T \boldsymbol{A}\boldsymbol{p}_i= \boldsymbol{p}_k^T \boldsymbol{b},
-\end{equation*}
-$$
-
-
-and we can define the coefficients \( \alpha_k \) as
-
-
-$$
-\begin{equation*}
- \alpha_k = \frac{\boldsymbol{p}_k^T \boldsymbol{b}}{\boldsymbol{p}_k^T \boldsymbol{A} \boldsymbol{p}_k}
-\end{equation*}
-$$
-
-
-
-
-
-
-Conjugate gradient method and iterations
-
-
-
-If we choose the conjugate vectors \( \boldsymbol{p}_k \) carefully,
-then we may not need all of them to obtain a good approximation to the solution
-\( \boldsymbol{x} \).
-We want to regard the conjugate gradient method as an iterative method.
-This will us to solve systems where \( n \) is so large that the direct
-method would take too much time.
-
-
-We denote the initial guess for \( \boldsymbol{x} \) as \( \boldsymbol{x}_0 \).
-We can assume without loss of generality that
-
-$$
-\begin{equation*}
-\boldsymbol{x}_0=0,
-\end{equation*}
-$$
-
-
-or consider the system
-
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{z} = \boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_0,
-\end{equation*}
-$$
-
-
-instead.
-
-
-
-
-
-Conjugate gradient method
-
-
-
-One can show that the solution \( \boldsymbol{x} \) is also the unique minimizer of the quadratic form
-
-$$
-\begin{equation*}
- f(\boldsymbol{x}) = \frac{1}{2}\boldsymbol{x}^T\boldsymbol{A}\boldsymbol{x} - \boldsymbol{x}^T \boldsymbol{x} , \quad \boldsymbol{x}\in\mathbf{R}^n.
-\end{equation*}
-$$
-
-
-This suggests taking the first basis vector \( \boldsymbol{p}_1 \)
-to be the gradient of \( f \) at \( \boldsymbol{x}=\boldsymbol{x}_0 \),
-which equals
-
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{x}_0-\boldsymbol{b},
-\end{equation*}
-$$
-
-
-and
-\( \boldsymbol{x}_0=0 \) it is equal \( -\boldsymbol{b} \).
-The other vectors in the basis will be conjugate to the gradient,
-hence the name conjugate gradient method.
-
-
-
-
-
-Conjugate gradient method
-
-
-
-Let \( \boldsymbol{r}_k \) be the residual at the \( k \)-th step:
-
-$$
-\begin{equation*}
-\boldsymbol{r}_k=\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_k.
-\end{equation*}
-$$
-
-
-Note that \( \boldsymbol{r}_k \) is the negative gradient of \( f \) at
-\( \boldsymbol{x}=\boldsymbol{x}_k \),
-so the gradient descent method would be to move in the direction \( \boldsymbol{r}_k \).
-Here, we insist that the directions \( \boldsymbol{p}_k \) are conjugate to each other,
-so we take the direction closest to the gradient \( \boldsymbol{r}_k \)
-under the conjugacy constraint.
-This gives the following expression
-
-$$
-\begin{equation*}
-\boldsymbol{p}_{k+1}=\boldsymbol{r}_k-\frac{\boldsymbol{p}_k^T \boldsymbol{A}\boldsymbol{r}_k}{\boldsymbol{p}_k^T\boldsymbol{A}\boldsymbol{p}_k} \boldsymbol{p}_k.
-\end{equation*}
-$$
-
-
-
-
-
-
-Conjugate gradient method
-
-
-
-We can also compute the residual iteratively as
-
-$$
-\begin{equation*}
-\boldsymbol{r}_{k+1}=\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_{k+1},
- \end{equation*}
-$$
-
-
-which equals
-
-$$
-\begin{equation*}
-\boldsymbol{b}-\boldsymbol{A}(\boldsymbol{x}_k+\alpha_k\boldsymbol{p}_k),
- \end{equation*}
-$$
-
-
-or
-
-$$
-\begin{equation*}
-(\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_k)-\alpha_k\boldsymbol{A}\boldsymbol{p}_k,
- \end{equation*}
-$$
-
-
-which gives
-
-
-$$
-\begin{equation*}
-\boldsymbol{r}_{k+1}=\boldsymbol{r}_k-\boldsymbol{A}\boldsymbol{p}_{k},
- \end{equation*}
-$$
-
-
-
-
-
-
-Revisiting some of our first Linear Regression Encounters
-
-
-We will use linear regression as a case study for the gradient descent
-methods. Linear regression is a great test case for the gradient
-descent methods discussed in the lectures since it has several
-desirable properties such as:
-
-
-- An analytical solution (recall homework set 1).
-- The gradient can be computed analytically.
-- The cost function is convex which guarantees that gradient descent converges for small enough learning rates
-
-
-
-We revisit an example similar to what we had in the first homework set. We had a function of the type
-
-
-
-
-
x = 2*np.random.rand(m,1)
-y = 4+3*x+np.random.randn(m,1)
-
-
-with \( x_i \in [0,1] \) is chosen randomly using a uniform distribution. Additionally we have a stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
-The linear regression model is given by
-
-$$
-h_\beta(x) = \boldsymbol{y} = \beta_0 + \beta_1 x,
-$$
-
-
-such that
-
-$$
-\boldsymbol{y}_i = \beta_0 + \beta_1 x_i.
-$$
-
-
-
-
-
-Gradient descent example
-
-
-Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\boldsymbol{y}} = (\boldsymbol{y}_1,\cdots,\boldsymbol{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
-
-
-It is convenient to write \( \mathbf{\boldsymbol{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by (we keep the intercept here)
-
-$$
-X \equiv \begin{bmatrix}
-1 & x_1 \\
-\vdots & \vdots \\
-1 & x_{100} & \\
-\end{bmatrix}.
-$$
-
-
-The cost/loss/risk function is given by (
-
-$$
-C(\beta) = \frac{1}{n}||X\beta-\mathbf{y}||_{2}^{2} = \frac{1}{n}\sum_{i=1}^{100}\left[ (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2\right]
-$$
-
-
-and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
-
-
-
-
-The derivative of the cost/loss function
-
-
-Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
-
-$$
-\nabla_{\beta} C(\beta) = \frac{2}{n}\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
-\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
-\end{bmatrix} = \frac{2}{n}X^T(X\beta - \mathbf{y}),
-$$
-
-
-where \( X \) is the design matrix defined above.
-
-
-
-
-The Hessian matrix
-The Hessian matrix of \( C(\beta) \) is given by
-
-$$
-\boldsymbol{H} \equiv \begin{bmatrix}
-\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
-\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
-\end{bmatrix} = \frac{2}{n}X^T X.
-$$
-
-
-This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
-
-
-
-
-Simple program
-
-
-We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
-
-$$
-\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
-$$
-
-
-
-We can use the expression we computed for the gradient and let use a
-\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
-when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \). Note that the code below does not include the latter stop criterion.
-
-
-And finally we can compare our solution for \( \beta \) with the analytic result given by
-\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
-
-
-
-
-Gradient Descent Example
-
-
-Here our simple example
-
-
-
-
# Importing various packages
-from random import random, seed
-import numpy as np
-import matplotlib.pyplot as plt
-from mpl_toolkits.mplot3d import Axes3D
-from matplotlib import cm
-from matplotlib.ticker import LinearLocator, FormatStrFormatter
-import sys
-
-# the number of datapoints
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-# Hessian matrix
-H = (2.0/n)* X.T @ X
-# Get the eigenvalues
-EigValues, EigVectors = np.linalg.eig(H)
-print(EigValues)
-
-beta_linreg = np.linalg.inv(X.T @ X) @ X.T @ y
-print(beta_linreg)
-beta = np.random.randn(2,1)
-
-eta = 1.0/np.max(EigValues)
-Niterations = 1000
-
-for iter in range(Niterations):
- gradient = (2.0/n)*X.T @ (X @ beta-y)
- beta -= eta*gradient
-
-print(beta)
-xnew = np.array([[0],[2]])
-xbnew = np.c_[np.ones((2,1)), xnew]
-ypredict = xbnew.dot(beta)
-ypredict2 = xbnew.dot(beta_linreg)
-plt.plot(xnew, ypredict, "r-")
-plt.plot(xnew, ypredict2, "b-")
-plt.plot(x, y ,'ro')
-plt.axis([0,2.0,0, 15.0])
-plt.xlabel(r'$x$')
-plt.ylabel(r'$y$')
-plt.title(r'Gradient descent example')
-plt.show()
-
-
-
-
-
-And a corresponding example using scikit-learn
-
-
-
-
-
# Importing various packages
-from random import random, seed
-import numpy as np
-import matplotlib.pyplot as plt
-from sklearn.linear_model import SGDRegressor
-
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-beta_linreg = np.linalg.inv(X.T @ X) @ (X.T @ y)
-print(beta_linreg)
-sgdreg = SGDRegressor(max_iter = 50, penalty=None, eta0=0.1)
-sgdreg.fit(x,y.ravel())
-print(sgdreg.intercept_, sgdreg.coef_)
-
-
-
-
-
-Gradient descent and Ridge
-
-
-We have also discussed Ridge regression where the loss function contains a regularized term given by the \( L_2 \) norm of \( \beta \),
-
-$$
-C_{\text{ridge}}(\beta) = \frac{1}{n}||X\beta -\mathbf{y}||^2 + \lambda ||\beta||^2, \ \lambda \geq 0.
-$$
-
-
-
-In order to minimize \( C_{\text{ridge}}(\beta) \) using GD we only have adjust the gradient as follows
-
-$$
-\nabla_\beta C_{\text{ridge}}(\beta) = \frac{2}{n}\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
-\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
-\end{bmatrix} + 2\lambda\begin{bmatrix} \beta_0 \\ \beta_1\end{bmatrix} = 2 (X^T(X\beta - \mathbf{y})+\lambda \beta).
-$$
-
-
-
-We can easily extend our program to minimize \( C_{\text{ridge}}(\beta) \) using gradient descent and compare with the analytical solution given by
-
-$$
-\beta_{\text{ridge}} = \left(X^T X + \lambda I_{2 \times 2} \right)^{-1} X^T \mathbf{y}.
-$$
-
-
-
-
-
-Program example for gradient descent with Ridge Regression
-
-
-
-
from random import random, seed
-import numpy as np
-import matplotlib.pyplot as plt
-from mpl_toolkits.mplot3d import Axes3D
-from matplotlib import cm
-from matplotlib.ticker import LinearLocator, FormatStrFormatter
-import sys
-
-# the number of datapoints
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-XT_X = X.T @ X
-
-#Ridge parameter lambda
-lmbda = 0.001
-Id = lmbda* np.eye(XT_X.shape[0])
-
-beta_linreg = np.linalg.inv(XT_X+Id) @ X.T @ y
-print(beta_linreg)
-# Start plain gradient descent
-beta = np.random.randn(2,1)
-
-eta = 0.1
-Niterations = 100
-
-for iter in range(Niterations):
- gradients = 2.0/n*X.T @ (X @ (beta)-y)+2*lmbda*beta
- beta -= eta*gradients
-
-print(beta)
-ypredict = X @ beta
-ypredict2 = X @ beta_linreg
-plt.plot(x, ypredict, "r-")
-plt.plot(x, ypredict2, "b-")
-plt.plot(x, y ,'ro')
-plt.axis([0,2.0,0, 15.0])
-plt.xlabel(r'$x$')
-plt.ylabel(r'$y$')
-plt.title(r'Gradient descent example for Ridge')
-plt.show()
-
-
-
-
-
-Using gradient descent methods, limitations
-
-
-- Gradient descent (GD) finds local minima of our function. Since the GD algorithm is deterministic, if it converges, it will converge to a local minimum of our cost/loss/risk function. Because in ML we are often dealing with extremely rugged landscapes with many local minima, this can lead to poor performance.
-- GD is sensitive to initial conditions. One consequence of the local nature of GD is that initial conditions matter. Depending on where one starts, one will end up at a different local minima. Therefore, it is very important to think about how one initializes the training process. This is true for GD as well as more complicated variants of GD.
-- Gradients are computationally expensive to calculate for large datasets. In many cases in statistics and ML, the cost/loss/risk function is a sum of terms, with one term for each data point. For example, in linear regression, \( E \propto \sum_{i=1}^n (y_i - \mathbf{w}^T\cdot\mathbf{x}_i)^2 \); for logistic regression, the square error is replaced by the cross entropy. To calculate the gradient we have to sum over all \( n \) data points. Doing this at every GD step becomes extremely computationally expensive. An ingenious solution to this, is to calculate the gradients using small subsets of the data called "mini batches". This has the added benefit of introducing stochasticity into our algorithm.
-- GD is very sensitive to choices of learning rates. GD is extremely sensitive to the choice of learning rates. If the learning rate is very small, the training process take an extremely long time. For larger learning rates, GD can diverge and give poor results. Furthermore, depending on what the local landscape looks like, we have to modify the learning rates to ensure convergence. Ideally, we would adaptively choose the learning rates to match the landscape.
-- GD treats all directions in parameter space uniformly. Another major drawback of GD is that unlike Newton's method, the learning rate for GD is the same in all directions in parameter space. For this reason, the maximum learning rate is set by the behavior of the steepest direction and this can significantly slow down training. Ideally, we would like to take large steps in flat directions and small steps in steep directions. Since we are exploring rugged landscapes where curvatures change, this requires us to keep track of not only the gradient but second derivatives. The ideal scenario would be to calculate the Hessian but this proves to be too computationally expensive.
-- GD can take exponential time to escape saddle points, even with random initialization. As we mentioned, GD is extremely sensitive to initial condition since it determines the particular local minimum GD would eventually reach. However, even with a good initialization scheme, through the introduction of randomness, GD can still take exponential time to escape saddle points. This leads us to our next topic, Stochastic Gradient Methods.
-
-
-
-
diff --git a/doc/pub/week38/html/week38-solarized.html b/doc/pub/week38/html/week38-solarized.html
index 25e6892c5..8cfbf13d9 100644
--- a/doc/pub/week38/html/week38-solarized.html
+++ b/doc/pub/week38/html/week38-solarized.html
@@ -26,32 +26,6 @@ pre {
border: 0pt solid #93a1a1;
box-shadow: none;
}
-.alert-text-small { font-size: 80%; }
-.alert-text-large { font-size: 130%; }
-.alert-text-normal { font-size: 90%; }
-.alert {
- padding:8px 35px 8px 14px; margin-bottom:18px;
- text-shadow:0 1px 0 rgba(255,255,255,0.5);
- border:1px solid #93a1a1;
- border-radius: 4px;
- -webkit-border-radius: 4px;
- -moz-border-radius: 4px;
- color: #555;
- background-color: #eee8d5;
- background-position: 10px 5px;
- background-repeat: no-repeat;
- background-size: 38px;
- padding-left: 55px;
- width: 75%;
- }
-.alert-block {padding-top:14px; padding-bottom:14px}
-.alert-block > p, .alert-block > ul {margin-bottom:1em}
-.alert li {margin-top: 1em}
-.alert-block p+p {margin-top:5px}
-.alert-notice { background-image: url(https://cdn.rawgit.com/doconce/doconce/master/bundled/html_images/small_yellow_notice.png); }
-.alert-summary { background-image:url(https://cdn.rawgit.com/doconce/doconce/master/bundled/html_images/small_yellow_summary.png); }
-.alert-warning { background-image: url(https://cdn.rawgit.com/doconce/doconce/master/bundled/html_images/small_yellow_warning.png); }
-.alert-question {background-image:url(https://cdn.rawgit.com/doconce/doconce/master/bundled/html_images/small_yellow_question.png); }
div { text-align: justify; text-justify: inter-word; }
@@ -184,6 +158,7 @@ div { text-align: justify; text-justify: inter-word; }
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -217,85 +192,7 @@ div { text-align: justify; text-justify: inter-word; }
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -337,7 +234,7 @@ MathJax.Hub.Config({
[2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University
-
Sep 24, 2021
+Sep 28, 2021
@@ -2376,6 +2273,11 @@ plt.show()
+
Friday September 25
+
+
+
+
Optimization, the central part of any Machine Learning algortithm
@@ -2680,880 +2582,6 @@ randomness. One such method is that of Stochastic Gradient Descent
(SGD), see below.
-
-
-
Convex functions
-
-
-Ideally we want our cost/loss function to be convex(concave).
-
-
-First we give the definition of a convex set: A set \( C \) in
-\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and
-all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to
-C. Geometrically this means that every point on the line segment
-connecting \( x \) and \( y \) is in \( C \) as discussed below.
-
-
-The convex subsets of \( \mathbb{R} \) are the intervals of
-\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the
-regular polygons (triangles, rectangles, pentagons, etc...).
-
-
-
-
-
Convex function
-
-
-Convex function: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below.
-
-
-
-
-
Conditions on convex functions
-
-
-In the following we state first and second-order conditions which
-ensures convexity of a function \( f \). We write \( D_f \) to denote the
-domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more
-details and proofs we refer to: S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press.
-
-
-
-
First order condition
-
-Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for
-all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \)
-is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
-for all \( x,y \in D_f \). This condition means that for a convex function
-the first order Taylor expansion (right hand side above) at any point
-a global under estimator of the function. To convince yourself you can
-make a drawing of \( f(x) = x^2+1 \) and draw the tangent line to \( f(x) \) and
-note that it is always below the graph.
-
-
-
-
-
-
Second order condition
-
-Assume that \( f \) is twice
-differentiable, i.e the Hessian matrix exists at each point in
-\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its
-Hessian is positive semi-definite for all \( x\in D_f \). For a
-single-variable function this reduces to \( f''(x) \geq 0 \). Geometrically this means that \( f \) has nonnegative curvature
-everywhere.
-
-
-
-
-This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
-
-
-
-
-
More on convex functions
-
-
-The next result is of great importance to us and the reason why we are
-going on about convex functions. In machine learning we frequently
-have to minimize a loss/cost function in order to find the best
-parameters for the model we are considering.
-
-
-Ideally we want the
-global minimum (for high-dimensional models it is hard to know
-if we have local or global minimum). However, if the cost/loss function
-is convex the following result provides invaluable information:
-
-
-
-
Any minimum is global for convex functions
-
-Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \)
-is minimal, where \( f \) is convex and differentiable. Then, any point
-\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum.
-
-
-
-
-This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
-
-
-
-
-
Some simple problems
-
-
-- Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.
-- Using the second order condition show that the following functions are convex on the specified domain.
-
-
- - \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).
- - \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).
-
-
-- Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.
-- A norm is any function that satisfy the following properties
-
-
- - \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).
- - \( f(x+y) \leq f(x) + f(y) \)
- - \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)
-
-
-
-
-Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
-
-
-
-
-
Friday September 25
-
-
-Video of Lecture and link to handwritten notes.
-
-
-
-
-
Standard steepest descent
-
-
-Before we proceed, we would like to discuss the approach called the
-standard Steepest descent (different from the above steepest descent discussion), which again leads to us having to be able
-to compute a matrix. It belongs to the class of Conjugate Gradient methods (CG).
-
-
-The success of the CG method
-for finding solutions of non-linear problems is based on the theory
-of conjugate gradients for linear systems of equations. It belongs to
-the class of iterative methods for solving problems from linear
-algebra of the type
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{x} = \boldsymbol{b}.
-\end{equation*}
-$$
-
-
-In the iterative process we end up with a problem like
-
-$$
-\begin{equation*}
- \boldsymbol{r}= \boldsymbol{b}-\boldsymbol{A}\boldsymbol{x},
-\end{equation*}
-$$
-
-where \( \boldsymbol{r} \) is the so-called residual or error in the iterative process.
-
-
-When we have found the exact solution, \( \boldsymbol{r}=0 \).
-
-
-
-
-
Gradient method
-
-
-The residual is zero when we reach the minimum of the quadratic equation
-$$
-\begin{equation*}
- P(\boldsymbol{x})=\frac{1}{2}\boldsymbol{x}^T\boldsymbol{A}\boldsymbol{x} - \boldsymbol{x}^T\boldsymbol{b},
-\end{equation*}
-$$
-
-
-with the constraint that the matrix \( \boldsymbol{A} \) is positive definite and
-symmetric. This defines also the Hessian and we want it to be positive definite.
-
-
-
-
-
Steepest descent method
-
-
-We denote the initial guess for \( \boldsymbol{x} \) as \( \boldsymbol{x}_0 \).
-We can assume without loss of generality that
-$$
-\begin{equation*}
-\boldsymbol{x}_0=0,
-\end{equation*}
-$$
-
-or consider the system
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{z} = \boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_0,
-\end{equation*}
-$$
-
-instead.
-
-
-
-
-
Steepest descent method
-
-
-
-One can show that the solution \( \boldsymbol{x} \) is also the unique minimizer of the quadratic form
-$$
-\begin{equation*}
- f(\boldsymbol{x}) = \frac{1}{2}\boldsymbol{x}^T\boldsymbol{A}\boldsymbol{x} - \boldsymbol{x}^T \boldsymbol{x} , \quad \boldsymbol{x}\in\mathbf{R}^n.
-\end{equation*}
-$$
-
-This suggests taking the first basis vector \( \boldsymbol{r}_1 \) (see below for definition)
-to be the gradient of \( f \) at \( \boldsymbol{x}=\boldsymbol{x}_0 \),
-which equals
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{x}_0-\boldsymbol{b},
-\end{equation*}
-$$
-
-and
-\( \boldsymbol{x}_0=0 \) it is equal \( -\boldsymbol{b} \).
-
-
-
-
-
-
-
-
-
Final expressions
-
-
-
-We can compute the residual iteratively as
-$$
-\begin{equation*}
-\boldsymbol{r}_{k+1}=\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_{k+1},
- \end{equation*}
-$$
-
-which equals
-$$
-\begin{equation*}
-\boldsymbol{b}-\boldsymbol{A}(\boldsymbol{x}_k+\alpha_k\boldsymbol{r}_k),
- \end{equation*}
-$$
-
-or
-$$
-\begin{equation*}
-(\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_k)-\alpha_k\boldsymbol{A}\boldsymbol{r}_k,
- \end{equation*}
-$$
-
-which gives
-
-$$
-\alpha_k = \frac{\boldsymbol{r}_k^T\boldsymbol{r}_k}{\boldsymbol{r}_k^T\boldsymbol{A}\boldsymbol{r}_k}
-$$
-
-leading to the iterative scheme
-$$
-\begin{equation*}
-\boldsymbol{x}_{k+1}=\boldsymbol{x}_k-\alpha_k\boldsymbol{r}_{k},
- \end{equation*}
-$$
-
-
-
-
-
-
-
Steepest descent example
-
-
-
-
-
import numpy as np
-import numpy.linalg as la
-
-import scipy.optimize as sopt
-
-import matplotlib.pyplot as pt
-from mpl_toolkits.mplot3d import axes3d
-
-def f(x):
- return 0.5*x[0]**2 + 2.5*x[1]**2
-
-def df(x):
- return np.array([x[0], 5*x[1]])
-
-fig = pt.figure()
-ax = fig.gca(projection="3d")
-
-xmesh, ymesh = np.mgrid[-2:2:50j,-2:2:50j]
-fmesh = f(np.array([xmesh, ymesh]))
-ax.plot_surface(xmesh, ymesh, fmesh)
-
-
-And then as countor plot
-
-
-
-
pt.axis("equal")
-pt.contour(xmesh, ymesh, fmesh)
-guesses = [np.array([2, 2./5])]
-
-
-Find guesses
-
-
-
-
x = guesses[-1]
-s = -df(x)
-
-
-Run it!
-
-
-
-
def f1d(alpha):
- return f(x + alpha*s)
-
-alpha_opt = sopt.golden(f1d)
-next_guess = x + alpha_opt * s
-guesses.append(next_guess)
-print(next_guess)
-
-
-What happened?
-
-
-
-
pt.axis("equal")
-pt.contour(xmesh, ymesh, fmesh, 50)
-it_array = np.array(guesses)
-pt.plot(it_array.T[0], it_array.T[1], "x-")
-
-
-
-
-
Conjugate gradient method
-
-
-
-In the CG method we define so-called conjugate directions and two vectors
-\( \boldsymbol{s} \) and \( \boldsymbol{t} \)
-are said to be
-conjugate if
-$$
-\begin{equation*}
-\boldsymbol{s}^T\boldsymbol{A}\boldsymbol{t}= 0.
-\end{equation*}
-$$
-
-The philosophy of the CG method is to perform searches in various conjugate directions
-of our vectors \( \boldsymbol{x}_i \) obeying the above criterion, namely
-$$
-\begin{equation*}
-\boldsymbol{x}_i^T\boldsymbol{A}\boldsymbol{x}_j= 0.
-\end{equation*}
-$$
-
-Two vectors are conjugate if they are orthogonal with respect to
-this inner product. Being conjugate is a symmetric relation: if \( \boldsymbol{s} \) is conjugate to \( \boldsymbol{t} \), then \( \boldsymbol{t} \) is conjugate to \( \boldsymbol{s} \).
-
-
-
-
-
-
-
Conjugate gradient method
-
-
-
-An example is given by the eigenvectors of the matrix
-$$
-\begin{equation*}
-\boldsymbol{v}_i^T\boldsymbol{A}\boldsymbol{v}_j= \lambda\boldsymbol{v}_i^T\boldsymbol{v}_j,
-\end{equation*}
-$$
-
-which is zero unless \( i=j \).
-
-
-
-
-
-
-
Conjugate gradient method
-
-
-
-Assume now that we have a symmetric positive-definite matrix \( \boldsymbol{A} \) of size
-\( n\times n \). At each iteration \( i+1 \) we obtain the conjugate direction of a vector
-$$
-\begin{equation*}
-\boldsymbol{x}_{i+1}=\boldsymbol{x}_{i}+\alpha_i\boldsymbol{p}_{i}.
-\end{equation*}
-$$
-
-We assume that \( \boldsymbol{p}_{i} \) is a sequence of \( n \) mutually conjugate directions.
-Then the \( \boldsymbol{p}_{i} \) form a basis of \( R^n \) and we can expand the solution
-$ \boldsymbol{A}\boldsymbol{x} = \boldsymbol{b}$ in this basis, namely
-
-$$
-\begin{equation*}
- \boldsymbol{x} = \sum^{n}_{i=1} \alpha_i \boldsymbol{p}_i.
-\end{equation*}
-$$
-
-
-
-
-
-
-
Conjugate gradient method
-
-
-
-The coefficients are given by
-$$
-\begin{equation*}
- \mathbf{A}\mathbf{x} = \sum^{n}_{i=1} \alpha_i \mathbf{A} \mathbf{p}_i = \mathbf{b}.
-\end{equation*}
-$$
-
-Multiplying with \( \boldsymbol{p}_k^T \) from the left gives
-
-$$
-\begin{equation*}
- \boldsymbol{p}_k^T \boldsymbol{A}\boldsymbol{x} = \sum^{n}_{i=1} \alpha_i\boldsymbol{p}_k^T \boldsymbol{A}\boldsymbol{p}_i= \boldsymbol{p}_k^T \boldsymbol{b},
-\end{equation*}
-$$
-
-and we can define the coefficients \( \alpha_k \) as
-
-$$
-\begin{equation*}
- \alpha_k = \frac{\boldsymbol{p}_k^T \boldsymbol{b}}{\boldsymbol{p}_k^T \boldsymbol{A} \boldsymbol{p}_k}
-\end{equation*}
-$$
-
-
-
-
-
-
-
Conjugate gradient method and iterations
-
-
-
-
-
-If we choose the conjugate vectors \( \boldsymbol{p}_k \) carefully,
-then we may not need all of them to obtain a good approximation to the solution
-\( \boldsymbol{x} \).
-We want to regard the conjugate gradient method as an iterative method.
-This will us to solve systems where \( n \) is so large that the direct
-method would take too much time.
-
-
-We denote the initial guess for \( \boldsymbol{x} \) as \( \boldsymbol{x}_0 \).
-We can assume without loss of generality that
-$$
-\begin{equation*}
-\boldsymbol{x}_0=0,
-\end{equation*}
-$$
-
-or consider the system
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{z} = \boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_0,
-\end{equation*}
-$$
-
-instead.
-
-
-
-
-
-
-
Conjugate gradient method
-
-
-
-One can show that the solution \( \boldsymbol{x} \) is also the unique minimizer of the quadratic form
-$$
-\begin{equation*}
- f(\boldsymbol{x}) = \frac{1}{2}\boldsymbol{x}^T\boldsymbol{A}\boldsymbol{x} - \boldsymbol{x}^T \boldsymbol{x} , \quad \boldsymbol{x}\in\mathbf{R}^n.
-\end{equation*}
-$$
-
-This suggests taking the first basis vector \( \boldsymbol{p}_1 \)
-to be the gradient of \( f \) at \( \boldsymbol{x}=\boldsymbol{x}_0 \),
-which equals
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{x}_0-\boldsymbol{b},
-\end{equation*}
-$$
-
-and
-\( \boldsymbol{x}_0=0 \) it is equal \( -\boldsymbol{b} \).
-The other vectors in the basis will be conjugate to the gradient,
-hence the name conjugate gradient method.
-
-
-
-
-
-
-
Conjugate gradient method
-
-
-
-Let \( \boldsymbol{r}_k \) be the residual at the \( k \)-th step:
-$$
-\begin{equation*}
-\boldsymbol{r}_k=\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_k.
-\end{equation*}
-$$
-
-Note that \( \boldsymbol{r}_k \) is the negative gradient of \( f \) at
-\( \boldsymbol{x}=\boldsymbol{x}_k \),
-so the gradient descent method would be to move in the direction \( \boldsymbol{r}_k \).
-Here, we insist that the directions \( \boldsymbol{p}_k \) are conjugate to each other,
-so we take the direction closest to the gradient \( \boldsymbol{r}_k \)
-under the conjugacy constraint.
-This gives the following expression
-$$
-\begin{equation*}
-\boldsymbol{p}_{k+1}=\boldsymbol{r}_k-\frac{\boldsymbol{p}_k^T \boldsymbol{A}\boldsymbol{r}_k}{\boldsymbol{p}_k^T\boldsymbol{A}\boldsymbol{p}_k} \boldsymbol{p}_k.
-\end{equation*}
-$$
-
-
-
-
-
-
-
Conjugate gradient method
-
-
-
-We can also compute the residual iteratively as
-$$
-\begin{equation*}
-\boldsymbol{r}_{k+1}=\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_{k+1},
- \end{equation*}
-$$
-
-which equals
-$$
-\begin{equation*}
-\boldsymbol{b}-\boldsymbol{A}(\boldsymbol{x}_k+\alpha_k\boldsymbol{p}_k),
- \end{equation*}
-$$
-
-or
-$$
-\begin{equation*}
-(\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_k)-\alpha_k\boldsymbol{A}\boldsymbol{p}_k,
- \end{equation*}
-$$
-
-which gives
-
-$$
-\begin{equation*}
-\boldsymbol{r}_{k+1}=\boldsymbol{r}_k-\boldsymbol{A}\boldsymbol{p}_{k},
- \end{equation*}
-$$
-
-
-
-
-
-
-
Revisiting some of our first Linear Regression Encounters
-
-
-We will use linear regression as a case study for the gradient descent
-methods. Linear regression is a great test case for the gradient
-descent methods discussed in the lectures since it has several
-desirable properties such as:
-
-
-- An analytical solution (recall homework set 1).
-- The gradient can be computed analytically.
-- The cost function is convex which guarantees that gradient descent converges for small enough learning rates
-
-
-We revisit an example similar to what we had in the first homework set. We had a function of the type
-
-
-
-
-
x = 2*np.random.rand(m,1)
-y = 4+3*x+np.random.randn(m,1)
-
-
-with \( x_i \in [0,1] \) is chosen randomly using a uniform distribution. Additionally we have a stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
-The linear regression model is given by
-$$
-h_\beta(x) = \boldsymbol{y} = \beta_0 + \beta_1 x,
-$$
-
-such that
-$$
-\boldsymbol{y}_i = \beta_0 + \beta_1 x_i.
-$$
-
-
-
-
-
Gradient descent example
-
-
-Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\boldsymbol{y}} = (\boldsymbol{y}_1,\cdots,\boldsymbol{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
-
-
-It is convenient to write \( \mathbf{\boldsymbol{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by (we keep the intercept here)
-$$
-X \equiv \begin{bmatrix}
-1 & x_1 \\
-\vdots & \vdots \\
-1 & x_{100} & \\
-\end{bmatrix}.
-$$
-
-The cost/loss/risk function is given by (
-$$
-C(\beta) = \frac{1}{n}||X\beta-\mathbf{y}||_{2}^{2} = \frac{1}{n}\sum_{i=1}^{100}\left[ (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2\right]
-$$
-
-and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
-
-
-
-
-
The derivative of the cost/loss function
-
-
-Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
-$$
-\nabla_{\beta} C(\beta) = \frac{2}{n}\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
-\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
-\end{bmatrix} = \frac{2}{n}X^T(X\beta - \mathbf{y}),
-$$
-
-where \( X \) is the design matrix defined above.
-
-
-
-
-
The Hessian matrix
-The Hessian matrix of \( C(\beta) \) is given by
-$$
-\boldsymbol{H} \equiv \begin{bmatrix}
-\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
-\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
-\end{bmatrix} = \frac{2}{n}X^T X.
-$$
-
-This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
-
-
-
-
-
Simple program
-
-
-We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
-$$
-\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
-$$
-
-
-We can use the expression we computed for the gradient and let use a
-\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
-when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \). Note that the code below does not include the latter stop criterion.
-
-
-And finally we can compare our solution for \( \beta \) with the analytic result given by
-\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
-
-
-
-
-
Gradient Descent Example
-
-
-Here our simple example
-
-
-
-
# Importing various packages
-from random import random, seed
-import numpy as np
-import matplotlib.pyplot as plt
-from mpl_toolkits.mplot3d import Axes3D
-from matplotlib import cm
-from matplotlib.ticker import LinearLocator, FormatStrFormatter
-import sys
-
-# the number of datapoints
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-# Hessian matrix
-H = (2.0/n)* X.T @ X
-# Get the eigenvalues
-EigValues, EigVectors = np.linalg.eig(H)
-print(EigValues)
-
-beta_linreg = np.linalg.inv(X.T @ X) @ X.T @ y
-print(beta_linreg)
-beta = np.random.randn(2,1)
-
-eta = 1.0/np.max(EigValues)
-Niterations = 1000
-
-for iter in range(Niterations):
- gradient = (2.0/n)*X.T @ (X @ beta-y)
- beta -= eta*gradient
-
-print(beta)
-xnew = np.array([[0],[2]])
-xbnew = np.c_[np.ones((2,1)), xnew]
-ypredict = xbnew.dot(beta)
-ypredict2 = xbnew.dot(beta_linreg)
-plt.plot(xnew, ypredict, "r-")
-plt.plot(xnew, ypredict2, "b-")
-plt.plot(x, y ,'ro')
-plt.axis([0,2.0,0, 15.0])
-plt.xlabel(r'$x$')
-plt.ylabel(r'$y$')
-plt.title(r'Gradient descent example')
-plt.show()
-
-
-
-
-
And a corresponding example using scikit-learn
-
-
-
-
-
# Importing various packages
-from random import random, seed
-import numpy as np
-import matplotlib.pyplot as plt
-from sklearn.linear_model import SGDRegressor
-
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-beta_linreg = np.linalg.inv(X.T @ X) @ (X.T @ y)
-print(beta_linreg)
-sgdreg = SGDRegressor(max_iter = 50, penalty=None, eta0=0.1)
-sgdreg.fit(x,y.ravel())
-print(sgdreg.intercept_, sgdreg.coef_)
-
-
-
-
-
Gradient descent and Ridge
-
-
-We have also discussed Ridge regression where the loss function contains a regularized term given by the \( L_2 \) norm of \( \beta \),
-$$
-C_{\text{ridge}}(\beta) = \frac{1}{n}||X\beta -\mathbf{y}||^2 + \lambda ||\beta||^2, \ \lambda \geq 0.
-$$
-
-
-In order to minimize \( C_{\text{ridge}}(\beta) \) using GD we only have adjust the gradient as follows
-$$
-\nabla_\beta C_{\text{ridge}}(\beta) = \frac{2}{n}\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
-\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
-\end{bmatrix} + 2\lambda\begin{bmatrix} \beta_0 \\ \beta_1\end{bmatrix} = 2 (X^T(X\beta - \mathbf{y})+\lambda \beta).
-$$
-
-
-We can easily extend our program to minimize \( C_{\text{ridge}}(\beta) \) using gradient descent and compare with the analytical solution given by
-$$
-\beta_{\text{ridge}} = \left(X^T X + \lambda I_{2 \times 2} \right)^{-1} X^T \mathbf{y}.
-$$
-
-
-
-
-
Program example for gradient descent with Ridge Regression
-
-
-
-
from random import random, seed
-import numpy as np
-import matplotlib.pyplot as plt
-from mpl_toolkits.mplot3d import Axes3D
-from matplotlib import cm
-from matplotlib.ticker import LinearLocator, FormatStrFormatter
-import sys
-
-# the number of datapoints
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-XT_X = X.T @ X
-
-#Ridge parameter lambda
-lmbda = 0.001
-Id = lmbda* np.eye(XT_X.shape[0])
-
-beta_linreg = np.linalg.inv(XT_X+Id) @ X.T @ y
-print(beta_linreg)
-# Start plain gradient descent
-beta = np.random.randn(2,1)
-
-eta = 0.1
-Niterations = 100
-
-for iter in range(Niterations):
- gradients = 2.0/n*X.T @ (X @ (beta)-y)+2*lmbda*beta
- beta -= eta*gradients
-
-print(beta)
-ypredict = X @ beta
-ypredict2 = X @ beta_linreg
-plt.plot(x, ypredict, "r-")
-plt.plot(x, ypredict2, "b-")
-plt.plot(x, y ,'ro')
-plt.axis([0,2.0,0, 15.0])
-plt.xlabel(r'$x$')
-plt.ylabel(r'$y$')
-plt.title(r'Gradient descent example for Ridge')
-plt.show()
-
-
-
-
-
Using gradient descent methods, limitations
-
-
-- Gradient descent (GD) finds local minima of our function. Since the GD algorithm is deterministic, if it converges, it will converge to a local minimum of our cost/loss/risk function. Because in ML we are often dealing with extremely rugged landscapes with many local minima, this can lead to poor performance.
-- GD is sensitive to initial conditions. One consequence of the local nature of GD is that initial conditions matter. Depending on where one starts, one will end up at a different local minima. Therefore, it is very important to think about how one initializes the training process. This is true for GD as well as more complicated variants of GD.
-- Gradients are computationally expensive to calculate for large datasets. In many cases in statistics and ML, the cost/loss/risk function is a sum of terms, with one term for each data point. For example, in linear regression, \( E \propto \sum_{i=1}^n (y_i - \mathbf{w}^T\cdot\mathbf{x}_i)^2 \); for logistic regression, the square error is replaced by the cross entropy. To calculate the gradient we have to sum over all \( n \) data points. Doing this at every GD step becomes extremely computationally expensive. An ingenious solution to this, is to calculate the gradients using small subsets of the data called "mini batches". This has the added benefit of introducing stochasticity into our algorithm.
-- GD is very sensitive to choices of learning rates. GD is extremely sensitive to the choice of learning rates. If the learning rate is very small, the training process take an extremely long time. For larger learning rates, GD can diverge and give poor results. Furthermore, depending on what the local landscape looks like, we have to modify the learning rates to ensure convergence. Ideally, we would adaptively choose the learning rates to match the landscape.
-- GD treats all directions in parameter space uniformly. Another major drawback of GD is that unlike Newton's method, the learning rate for GD is the same in all directions in parameter space. For this reason, the maximum learning rate is set by the behavior of the steepest direction and this can significantly slow down training. Ideally, we would like to take large steps in flat directions and small steps in steep directions. Since we are exploring rugged landscapes where curvatures change, this requires us to keep track of not only the gradient but second derivatives. The ideal scenario would be to calculate the Hessian but this proves to be too computationally expensive.
-- GD can take exponential time to escape saddle points, even with random initialization. As we mentioned, GD is extremely sensitive to initial condition since it determines the particular local minimum GD would eventually reach. However, even with a good initialization scheme, through the introduction of randomness, GD can still take exponential time to escape saddle points. This leads us to our next topic, Stochastic Gradient Methods.
-
-
diff --git a/doc/pub/week38/html/week38.html b/doc/pub/week38/html/week38.html
index 35131b829..f6d73981c 100644
--- a/doc/pub/week38/html/week38.html
+++ b/doc/pub/week38/html/week38.html
@@ -31,32 +31,6 @@ p { text-indent: 0px; }
hr { border: 0; width: 80%; border-bottom: 1px solid #aaa}
p.caption { width: 80%; font-style: normal; text-align: left; }
hr.figure { border: 0; width: 80%; border-bottom: 1px solid #aaa}
-.alert-text-small { font-size: 80%; }
-.alert-text-large { font-size: 130%; }
-.alert-text-normal { font-size: 90%; }
-.alert {
- padding:8px 35px 8px 14px; margin-bottom:18px;
- text-shadow:0 1px 0 rgba(255,255,255,0.5);
- border:1px solid #bababa;
- border-radius: 4px;
- -webkit-border-radius: 4px;
- -moz-border-radius: 4px;
- color: #555;
- background-color: #f8f8f8;
- background-position: 10px 5px;
- background-repeat: no-repeat;
- background-size: 38px;
- padding-left: 55px;
- width: 75%;
- }
-.alert-block {padding-top:14px; padding-bottom:14px}
-.alert-block > p, .alert-block > ul {margin-bottom:1em}
-.alert li {margin-top: 1em}
-.alert-block p+p {margin-top:5px}
-.alert-notice { background-image: url(https://cdn.rawgit.com/doconce/doconce/master/bundled/html_images/small_gray_notice.png); }
-.alert-summary { background-image:url(https://cdn.rawgit.com/doconce/doconce/master/bundled/html_images/small_gray_summary.png); }
-.alert-warning { background-image: url(https://cdn.rawgit.com/doconce/doconce/master/bundled/html_images/small_gray_warning.png); }
-.alert-question {background-image:url(https://cdn.rawgit.com/doconce/doconce/master/bundled/html_images/small_gray_question.png); }
div { text-align: justify; text-justify: inter-word; }
@@ -189,6 +163,7 @@ div { text-align: justify; text-justify: inter-word; }
2,
None,
'other-measures-in-classification-studies-cancer-data-again'),
+ ('Friday September 25', 2, None, 'friday-september-25'),
('Optimization, the central part of any Machine Learning '
'algortithm',
2,
@@ -222,85 +197,7 @@ div { text-align: justify; text-justify: inter-word; }
('The sensitiveness of the gradient descent',
2,
None,
- 'the-sensitiveness-of-the-gradient-descent'),
- ('Convex functions', 2, None, 'convex-functions'),
- ('Convex function', 2, None, 'convex-function'),
- ('Conditions on convex functions',
- 2,
- None,
- 'conditions-on-convex-functions'),
- ('More on convex functions', 2, None, 'more-on-convex-functions'),
- ('Some simple problems', 2, None, 'some-simple-problems'),
- ('Friday September 25', 2, None, 'friday-september-25'),
- ('Standard steepest descent',
- 2,
- None,
- 'standard-steepest-descent'),
- ('Gradient method', 2, None, 'gradient-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Steepest descent method', 2, None, 'steepest-descent-method'),
- ('Final expressions', 2, None, 'final-expressions'),
- ('Steepest descent example', 2, None, 'steepest-descent-example'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method and iterations',
- 2,
- None,
- 'conjugate-gradient-method-and-iterations'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Conjugate gradient method',
- 2,
- None,
- 'conjugate-gradient-method'),
- ('Revisiting some of our first Linear Regression Encounters',
- 2,
- None,
- 'revisiting-some-of-our-first-linear-regression-encounters'),
- ('Gradient descent example', 2, None, 'gradient-descent-example'),
- ('The derivative of the cost/loss function',
- 2,
- None,
- 'the-derivative-of-the-cost-loss-function'),
- ('The Hessian matrix', 2, None, 'the-hessian-matrix'),
- ('Simple program', 2, None, 'simple-program'),
- ('Gradient Descent Example', 2, None, 'gradient-descent-example'),
- ('And a corresponding example using _scikit-learn_',
- 2,
- None,
- 'and-a-corresponding-example-using-_scikit-learn_'),
- ('Gradient descent and Ridge',
- 2,
- None,
- 'gradient-descent-and-ridge'),
- ('Program example for gradient descent with Ridge Regression',
- 2,
- None,
- 'program-example-for-gradient-descent-with-ridge-regression'),
- ('Using gradient descent methods, limitations',
- 2,
- None,
- 'using-gradient-descent-methods-limitations')]}
+ 'the-sensitiveness-of-the-gradient-descent')]}
end of tocinfo -->
@@ -342,7 +239,7 @@ MathJax.Hub.Config({
[2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University
-
Sep 24, 2021
+Sep 28, 2021
@@ -2381,6 +2278,11 @@ plt.show()
+
Friday September 25
+
+
+
+
Optimization, the central part of any Machine Learning algortithm
@@ -2685,880 +2587,6 @@ randomness. One such method is that of Stochastic Gradient Descent
(SGD), see below.
-
-
-
Convex functions
-
-
-Ideally we want our cost/loss function to be convex(concave).
-
-
-First we give the definition of a convex set: A set \( C \) in
-\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and
-all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to
-C. Geometrically this means that every point on the line segment
-connecting \( x \) and \( y \) is in \( C \) as discussed below.
-
-
-The convex subsets of \( \mathbb{R} \) are the intervals of
-\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the
-regular polygons (triangles, rectangles, pentagons, etc...).
-
-
-
-
-
Convex function
-
-
-Convex function: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below.
-
-
-
-
-
Conditions on convex functions
-
-
-In the following we state first and second-order conditions which
-ensures convexity of a function \( f \). We write \( D_f \) to denote the
-domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more
-details and proofs we refer to: S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press.
-
-
-
-
First order condition
-
-Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for
-all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \)
-is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
-for all \( x,y \in D_f \). This condition means that for a convex function
-the first order Taylor expansion (right hand side above) at any point
-a global under estimator of the function. To convince yourself you can
-make a drawing of \( f(x) = x^2+1 \) and draw the tangent line to \( f(x) \) and
-note that it is always below the graph.
-
-
-
-
-
-
Second order condition
-
-Assume that \( f \) is twice
-differentiable, i.e the Hessian matrix exists at each point in
-\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its
-Hessian is positive semi-definite for all \( x\in D_f \). For a
-single-variable function this reduces to \( f''(x) \geq 0 \). Geometrically this means that \( f \) has nonnegative curvature
-everywhere.
-
-
-
-
-This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
-
-
-
-
-
More on convex functions
-
-
-The next result is of great importance to us and the reason why we are
-going on about convex functions. In machine learning we frequently
-have to minimize a loss/cost function in order to find the best
-parameters for the model we are considering.
-
-
-Ideally we want the
-global minimum (for high-dimensional models it is hard to know
-if we have local or global minimum). However, if the cost/loss function
-is convex the following result provides invaluable information:
-
-
-
-
Any minimum is global for convex functions
-
-Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \)
-is minimal, where \( f \) is convex and differentiable. Then, any point
-\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum.
-
-
-
-
-This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
-
-
-
-
-
Some simple problems
-
-
-- Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.
-- Using the second order condition show that the following functions are convex on the specified domain.
-
-
- - \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).
- - \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).
-
-
-- Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.
-- A norm is any function that satisfy the following properties
-
-
- - \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).
- - \( f(x+y) \leq f(x) + f(y) \)
- - \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)
-
-
-
-
-Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
-
-
-
-
-
Friday September 25
-
-
-Video of Lecture and link to handwritten notes.
-
-
-
-
-
Standard steepest descent
-
-
-Before we proceed, we would like to discuss the approach called the
-standard Steepest descent (different from the above steepest descent discussion), which again leads to us having to be able
-to compute a matrix. It belongs to the class of Conjugate Gradient methods (CG).
-
-
-The success of the CG method
-for finding solutions of non-linear problems is based on the theory
-of conjugate gradients for linear systems of equations. It belongs to
-the class of iterative methods for solving problems from linear
-algebra of the type
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{x} = \boldsymbol{b}.
-\end{equation*}
-$$
-
-
-In the iterative process we end up with a problem like
-
-$$
-\begin{equation*}
- \boldsymbol{r}= \boldsymbol{b}-\boldsymbol{A}\boldsymbol{x},
-\end{equation*}
-$$
-
-where \( \boldsymbol{r} \) is the so-called residual or error in the iterative process.
-
-
-When we have found the exact solution, \( \boldsymbol{r}=0 \).
-
-
-
-
-
Gradient method
-
-
-The residual is zero when we reach the minimum of the quadratic equation
-$$
-\begin{equation*}
- P(\boldsymbol{x})=\frac{1}{2}\boldsymbol{x}^T\boldsymbol{A}\boldsymbol{x} - \boldsymbol{x}^T\boldsymbol{b},
-\end{equation*}
-$$
-
-
-with the constraint that the matrix \( \boldsymbol{A} \) is positive definite and
-symmetric. This defines also the Hessian and we want it to be positive definite.
-
-
-
-
-
Steepest descent method
-
-
-We denote the initial guess for \( \boldsymbol{x} \) as \( \boldsymbol{x}_0 \).
-We can assume without loss of generality that
-$$
-\begin{equation*}
-\boldsymbol{x}_0=0,
-\end{equation*}
-$$
-
-or consider the system
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{z} = \boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_0,
-\end{equation*}
-$$
-
-instead.
-
-
-
-
-
Steepest descent method
-
-
-
-One can show that the solution \( \boldsymbol{x} \) is also the unique minimizer of the quadratic form
-$$
-\begin{equation*}
- f(\boldsymbol{x}) = \frac{1}{2}\boldsymbol{x}^T\boldsymbol{A}\boldsymbol{x} - \boldsymbol{x}^T \boldsymbol{x} , \quad \boldsymbol{x}\in\mathbf{R}^n.
-\end{equation*}
-$$
-
-This suggests taking the first basis vector \( \boldsymbol{r}_1 \) (see below for definition)
-to be the gradient of \( f \) at \( \boldsymbol{x}=\boldsymbol{x}_0 \),
-which equals
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{x}_0-\boldsymbol{b},
-\end{equation*}
-$$
-
-and
-\( \boldsymbol{x}_0=0 \) it is equal \( -\boldsymbol{b} \).
-
-
-
-
-
-
-
-
-
Final expressions
-
-
-
-We can compute the residual iteratively as
-$$
-\begin{equation*}
-\boldsymbol{r}_{k+1}=\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_{k+1},
- \end{equation*}
-$$
-
-which equals
-$$
-\begin{equation*}
-\boldsymbol{b}-\boldsymbol{A}(\boldsymbol{x}_k+\alpha_k\boldsymbol{r}_k),
- \end{equation*}
-$$
-
-or
-$$
-\begin{equation*}
-(\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_k)-\alpha_k\boldsymbol{A}\boldsymbol{r}_k,
- \end{equation*}
-$$
-
-which gives
-
-$$
-\alpha_k = \frac{\boldsymbol{r}_k^T\boldsymbol{r}_k}{\boldsymbol{r}_k^T\boldsymbol{A}\boldsymbol{r}_k}
-$$
-
-leading to the iterative scheme
-$$
-\begin{equation*}
-\boldsymbol{x}_{k+1}=\boldsymbol{x}_k-\alpha_k\boldsymbol{r}_{k},
- \end{equation*}
-$$
-
-
-
-
-
-
-
Steepest descent example
-
-
-
-
-
import numpy as np
-import numpy.linalg as la
-
-import scipy.optimize as sopt
-
-import matplotlib.pyplot as pt
-from mpl_toolkits.mplot3d import axes3d
-
-def f(x):
- return 0.5*x[0]**2 + 2.5*x[1]**2
-
-def df(x):
- return np.array([x[0], 5*x[1]])
-
-fig = pt.figure()
-ax = fig.gca(projection="3d")
-
-xmesh, ymesh = np.mgrid[-2:2:50j,-2:2:50j]
-fmesh = f(np.array([xmesh, ymesh]))
-ax.plot_surface(xmesh, ymesh, fmesh)
-
-
-And then as countor plot
-
-
-
-
pt.axis("equal")
-pt.contour(xmesh, ymesh, fmesh)
-guesses = [np.array([2, 2./5])]
-
-
-Find guesses
-
-
-
-
x = guesses[-1]
-s = -df(x)
-
-
-Run it!
-
-
-
-
def f1d(alpha):
- return f(x + alpha*s)
-
-alpha_opt = sopt.golden(f1d)
-next_guess = x + alpha_opt * s
-guesses.append(next_guess)
-print(next_guess)
-
-
-What happened?
-
-
-
-
pt.axis("equal")
-pt.contour(xmesh, ymesh, fmesh, 50)
-it_array = np.array(guesses)
-pt.plot(it_array.T[0], it_array.T[1], "x-")
-
-
-
-
-
Conjugate gradient method
-
-
-
-In the CG method we define so-called conjugate directions and two vectors
-\( \boldsymbol{s} \) and \( \boldsymbol{t} \)
-are said to be
-conjugate if
-$$
-\begin{equation*}
-\boldsymbol{s}^T\boldsymbol{A}\boldsymbol{t}= 0.
-\end{equation*}
-$$
-
-The philosophy of the CG method is to perform searches in various conjugate directions
-of our vectors \( \boldsymbol{x}_i \) obeying the above criterion, namely
-$$
-\begin{equation*}
-\boldsymbol{x}_i^T\boldsymbol{A}\boldsymbol{x}_j= 0.
-\end{equation*}
-$$
-
-Two vectors are conjugate if they are orthogonal with respect to
-this inner product. Being conjugate is a symmetric relation: if \( \boldsymbol{s} \) is conjugate to \( \boldsymbol{t} \), then \( \boldsymbol{t} \) is conjugate to \( \boldsymbol{s} \).
-
-
-
-
-
-
-
Conjugate gradient method
-
-
-
-An example is given by the eigenvectors of the matrix
-$$
-\begin{equation*}
-\boldsymbol{v}_i^T\boldsymbol{A}\boldsymbol{v}_j= \lambda\boldsymbol{v}_i^T\boldsymbol{v}_j,
-\end{equation*}
-$$
-
-which is zero unless \( i=j \).
-
-
-
-
-
-
-
Conjugate gradient method
-
-
-
-Assume now that we have a symmetric positive-definite matrix \( \boldsymbol{A} \) of size
-\( n\times n \). At each iteration \( i+1 \) we obtain the conjugate direction of a vector
-$$
-\begin{equation*}
-\boldsymbol{x}_{i+1}=\boldsymbol{x}_{i}+\alpha_i\boldsymbol{p}_{i}.
-\end{equation*}
-$$
-
-We assume that \( \boldsymbol{p}_{i} \) is a sequence of \( n \) mutually conjugate directions.
-Then the \( \boldsymbol{p}_{i} \) form a basis of \( R^n \) and we can expand the solution
-$ \boldsymbol{A}\boldsymbol{x} = \boldsymbol{b}$ in this basis, namely
-
-$$
-\begin{equation*}
- \boldsymbol{x} = \sum^{n}_{i=1} \alpha_i \boldsymbol{p}_i.
-\end{equation*}
-$$
-
-
-
-
-
-
-
Conjugate gradient method
-
-
-
-The coefficients are given by
-$$
-\begin{equation*}
- \mathbf{A}\mathbf{x} = \sum^{n}_{i=1} \alpha_i \mathbf{A} \mathbf{p}_i = \mathbf{b}.
-\end{equation*}
-$$
-
-Multiplying with \( \boldsymbol{p}_k^T \) from the left gives
-
-$$
-\begin{equation*}
- \boldsymbol{p}_k^T \boldsymbol{A}\boldsymbol{x} = \sum^{n}_{i=1} \alpha_i\boldsymbol{p}_k^T \boldsymbol{A}\boldsymbol{p}_i= \boldsymbol{p}_k^T \boldsymbol{b},
-\end{equation*}
-$$
-
-and we can define the coefficients \( \alpha_k \) as
-
-$$
-\begin{equation*}
- \alpha_k = \frac{\boldsymbol{p}_k^T \boldsymbol{b}}{\boldsymbol{p}_k^T \boldsymbol{A} \boldsymbol{p}_k}
-\end{equation*}
-$$
-
-
-
-
-
-
-
Conjugate gradient method and iterations
-
-
-
-
-
-If we choose the conjugate vectors \( \boldsymbol{p}_k \) carefully,
-then we may not need all of them to obtain a good approximation to the solution
-\( \boldsymbol{x} \).
-We want to regard the conjugate gradient method as an iterative method.
-This will us to solve systems where \( n \) is so large that the direct
-method would take too much time.
-
-
-We denote the initial guess for \( \boldsymbol{x} \) as \( \boldsymbol{x}_0 \).
-We can assume without loss of generality that
-$$
-\begin{equation*}
-\boldsymbol{x}_0=0,
-\end{equation*}
-$$
-
-or consider the system
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{z} = \boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_0,
-\end{equation*}
-$$
-
-instead.
-
-
-
-
-
-
-
Conjugate gradient method
-
-
-
-One can show that the solution \( \boldsymbol{x} \) is also the unique minimizer of the quadratic form
-$$
-\begin{equation*}
- f(\boldsymbol{x}) = \frac{1}{2}\boldsymbol{x}^T\boldsymbol{A}\boldsymbol{x} - \boldsymbol{x}^T \boldsymbol{x} , \quad \boldsymbol{x}\in\mathbf{R}^n.
-\end{equation*}
-$$
-
-This suggests taking the first basis vector \( \boldsymbol{p}_1 \)
-to be the gradient of \( f \) at \( \boldsymbol{x}=\boldsymbol{x}_0 \),
-which equals
-$$
-\begin{equation*}
-\boldsymbol{A}\boldsymbol{x}_0-\boldsymbol{b},
-\end{equation*}
-$$
-
-and
-\( \boldsymbol{x}_0=0 \) it is equal \( -\boldsymbol{b} \).
-The other vectors in the basis will be conjugate to the gradient,
-hence the name conjugate gradient method.
-
-
-
-
-
-
-
Conjugate gradient method
-
-
-
-Let \( \boldsymbol{r}_k \) be the residual at the \( k \)-th step:
-$$
-\begin{equation*}
-\boldsymbol{r}_k=\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_k.
-\end{equation*}
-$$
-
-Note that \( \boldsymbol{r}_k \) is the negative gradient of \( f \) at
-\( \boldsymbol{x}=\boldsymbol{x}_k \),
-so the gradient descent method would be to move in the direction \( \boldsymbol{r}_k \).
-Here, we insist that the directions \( \boldsymbol{p}_k \) are conjugate to each other,
-so we take the direction closest to the gradient \( \boldsymbol{r}_k \)
-under the conjugacy constraint.
-This gives the following expression
-$$
-\begin{equation*}
-\boldsymbol{p}_{k+1}=\boldsymbol{r}_k-\frac{\boldsymbol{p}_k^T \boldsymbol{A}\boldsymbol{r}_k}{\boldsymbol{p}_k^T\boldsymbol{A}\boldsymbol{p}_k} \boldsymbol{p}_k.
-\end{equation*}
-$$
-
-
-
-
-
-
-
Conjugate gradient method
-
-
-
-We can also compute the residual iteratively as
-$$
-\begin{equation*}
-\boldsymbol{r}_{k+1}=\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_{k+1},
- \end{equation*}
-$$
-
-which equals
-$$
-\begin{equation*}
-\boldsymbol{b}-\boldsymbol{A}(\boldsymbol{x}_k+\alpha_k\boldsymbol{p}_k),
- \end{equation*}
-$$
-
-or
-$$
-\begin{equation*}
-(\boldsymbol{b}-\boldsymbol{A}\boldsymbol{x}_k)-\alpha_k\boldsymbol{A}\boldsymbol{p}_k,
- \end{equation*}
-$$
-
-which gives
-
-$$
-\begin{equation*}
-\boldsymbol{r}_{k+1}=\boldsymbol{r}_k-\boldsymbol{A}\boldsymbol{p}_{k},
- \end{equation*}
-$$
-
-
-
-
-
-
-
Revisiting some of our first Linear Regression Encounters
-
-
-We will use linear regression as a case study for the gradient descent
-methods. Linear regression is a great test case for the gradient
-descent methods discussed in the lectures since it has several
-desirable properties such as:
-
-
-- An analytical solution (recall homework set 1).
-- The gradient can be computed analytically.
-- The cost function is convex which guarantees that gradient descent converges for small enough learning rates
-
-
-We revisit an example similar to what we had in the first homework set. We had a function of the type
-
-
-
-
-
x = 2*np.random.rand(m,1)
-y = 4+3*x+np.random.randn(m,1)
-
-
-with \( x_i \in [0,1] \) is chosen randomly using a uniform distribution. Additionally we have a stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \).
-The linear regression model is given by
-$$
-h_\beta(x) = \boldsymbol{y} = \beta_0 + \beta_1 x,
-$$
-
-such that
-$$
-\boldsymbol{y}_i = \beta_0 + \beta_1 x_i.
-$$
-
-
-
-
-
Gradient descent example
-
-
-Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\boldsymbol{y}} = (\boldsymbol{y}_1,\cdots,\boldsymbol{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \)
-
-
-It is convenient to write \( \mathbf{\boldsymbol{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by (we keep the intercept here)
-$$
-X \equiv \begin{bmatrix}
-1 & x_1 \\
-\vdots & \vdots \\
-1 & x_{100} & \\
-\end{bmatrix}.
-$$
-
-The cost/loss/risk function is given by (
-$$
-C(\beta) = \frac{1}{n}||X\beta-\mathbf{y}||_{2}^{2} = \frac{1}{n}\sum_{i=1}^{100}\left[ (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2\right]
-$$
-
-and we want to find \( \beta \) such that \( C(\beta) \) is minimized.
-
-
-
-
-
The derivative of the cost/loss function
-
-
-Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as
-$$
-\nabla_{\beta} C(\beta) = \frac{2}{n}\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
-\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
-\end{bmatrix} = \frac{2}{n}X^T(X\beta - \mathbf{y}),
-$$
-
-where \( X \) is the design matrix defined above.
-
-
-
-
-
The Hessian matrix
-The Hessian matrix of \( C(\beta) \) is given by
-$$
-\boldsymbol{H} \equiv \begin{bmatrix}
-\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
-\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
-\end{bmatrix} = \frac{2}{n}X^T X.
-$$
-
-This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.
-
-
-
-
-
Simple program
-
-
-We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to
-$$
-\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
-$$
-
-
-We can use the expression we computed for the gradient and let use a
-\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating
-when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \). Note that the code below does not include the latter stop criterion.
-
-
-And finally we can compare our solution for \( \beta \) with the analytic result given by
-\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).
-
-
-
-
-
Gradient Descent Example
-
-
-Here our simple example
-
-
-
-
# Importing various packages
-from random import random, seed
-import numpy as np
-import matplotlib.pyplot as plt
-from mpl_toolkits.mplot3d import Axes3D
-from matplotlib import cm
-from matplotlib.ticker import LinearLocator, FormatStrFormatter
-import sys
-
-# the number of datapoints
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-# Hessian matrix
-H = (2.0/n)* X.T @ X
-# Get the eigenvalues
-EigValues, EigVectors = np.linalg.eig(H)
-print(EigValues)
-
-beta_linreg = np.linalg.inv(X.T @ X) @ X.T @ y
-print(beta_linreg)
-beta = np.random.randn(2,1)
-
-eta = 1.0/np.max(EigValues)
-Niterations = 1000
-
-for iter in range(Niterations):
- gradient = (2.0/n)*X.T @ (X @ beta-y)
- beta -= eta*gradient
-
-print(beta)
-xnew = np.array([[0],[2]])
-xbnew = np.c_[np.ones((2,1)), xnew]
-ypredict = xbnew.dot(beta)
-ypredict2 = xbnew.dot(beta_linreg)
-plt.plot(xnew, ypredict, "r-")
-plt.plot(xnew, ypredict2, "b-")
-plt.plot(x, y ,'ro')
-plt.axis([0,2.0,0, 15.0])
-plt.xlabel(r'$x$')
-plt.ylabel(r'$y$')
-plt.title(r'Gradient descent example')
-plt.show()
-
-
-
-
-
And a corresponding example using scikit-learn
-
-
-
-
-
# Importing various packages
-from random import random, seed
-import numpy as np
-import matplotlib.pyplot as plt
-from sklearn.linear_model import SGDRegressor
-
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-beta_linreg = np.linalg.inv(X.T @ X) @ (X.T @ y)
-print(beta_linreg)
-sgdreg = SGDRegressor(max_iter = 50, penalty=None, eta0=0.1)
-sgdreg.fit(x,y.ravel())
-print(sgdreg.intercept_, sgdreg.coef_)
-
-
-
-
-
Gradient descent and Ridge
-
-
-We have also discussed Ridge regression where the loss function contains a regularized term given by the \( L_2 \) norm of \( \beta \),
-$$
-C_{\text{ridge}}(\beta) = \frac{1}{n}||X\beta -\mathbf{y}||^2 + \lambda ||\beta||^2, \ \lambda \geq 0.
-$$
-
-
-In order to minimize \( C_{\text{ridge}}(\beta) \) using GD we only have adjust the gradient as follows
-$$
-\nabla_\beta C_{\text{ridge}}(\beta) = \frac{2}{n}\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
-\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
-\end{bmatrix} + 2\lambda\begin{bmatrix} \beta_0 \\ \beta_1\end{bmatrix} = 2 (X^T(X\beta - \mathbf{y})+\lambda \beta).
-$$
-
-
-We can easily extend our program to minimize \( C_{\text{ridge}}(\beta) \) using gradient descent and compare with the analytical solution given by
-$$
-\beta_{\text{ridge}} = \left(X^T X + \lambda I_{2 \times 2} \right)^{-1} X^T \mathbf{y}.
-$$
-
-
-
-
-
Program example for gradient descent with Ridge Regression
-
-
-
-
from random import random, seed
-import numpy as np
-import matplotlib.pyplot as plt
-from mpl_toolkits.mplot3d import Axes3D
-from matplotlib import cm
-from matplotlib.ticker import LinearLocator, FormatStrFormatter
-import sys
-
-# the number of datapoints
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-XT_X = X.T @ X
-
-#Ridge parameter lambda
-lmbda = 0.001
-Id = lmbda* np.eye(XT_X.shape[0])
-
-beta_linreg = np.linalg.inv(XT_X+Id) @ X.T @ y
-print(beta_linreg)
-# Start plain gradient descent
-beta = np.random.randn(2,1)
-
-eta = 0.1
-Niterations = 100
-
-for iter in range(Niterations):
- gradients = 2.0/n*X.T @ (X @ (beta)-y)+2*lmbda*beta
- beta -= eta*gradients
-
-print(beta)
-ypredict = X @ beta
-ypredict2 = X @ beta_linreg
-plt.plot(x, ypredict, "r-")
-plt.plot(x, ypredict2, "b-")
-plt.plot(x, y ,'ro')
-plt.axis([0,2.0,0, 15.0])
-plt.xlabel(r'$x$')
-plt.ylabel(r'$y$')
-plt.title(r'Gradient descent example for Ridge')
-plt.show()
-
-
-
-
-
Using gradient descent methods, limitations
-
-
-- Gradient descent (GD) finds local minima of our function. Since the GD algorithm is deterministic, if it converges, it will converge to a local minimum of our cost/loss/risk function. Because in ML we are often dealing with extremely rugged landscapes with many local minima, this can lead to poor performance.
-- GD is sensitive to initial conditions. One consequence of the local nature of GD is that initial conditions matter. Depending on where one starts, one will end up at a different local minima. Therefore, it is very important to think about how one initializes the training process. This is true for GD as well as more complicated variants of GD.
-- Gradients are computationally expensive to calculate for large datasets. In many cases in statistics and ML, the cost/loss/risk function is a sum of terms, with one term for each data point. For example, in linear regression, \( E \propto \sum_{i=1}^n (y_i - \mathbf{w}^T\cdot\mathbf{x}_i)^2 \); for logistic regression, the square error is replaced by the cross entropy. To calculate the gradient we have to sum over all \( n \) data points. Doing this at every GD step becomes extremely computationally expensive. An ingenious solution to this, is to calculate the gradients using small subsets of the data called "mini batches". This has the added benefit of introducing stochasticity into our algorithm.
-- GD is very sensitive to choices of learning rates. GD is extremely sensitive to the choice of learning rates. If the learning rate is very small, the training process take an extremely long time. For larger learning rates, GD can diverge and give poor results. Furthermore, depending on what the local landscape looks like, we have to modify the learning rates to ensure convergence. Ideally, we would adaptively choose the learning rates to match the landscape.
-- GD treats all directions in parameter space uniformly. Another major drawback of GD is that unlike Newton's method, the learning rate for GD is the same in all directions in parameter space. For this reason, the maximum learning rate is set by the behavior of the steepest direction and this can significantly slow down training. Ideally, we would like to take large steps in flat directions and small steps in steep directions. Since we are exploring rugged landscapes where curvatures change, this requires us to keep track of not only the gradient but second derivatives. The ideal scenario would be to calculate the Hessian but this proves to be too computationally expensive.
-- GD can take exponential time to escape saddle points, even with random initialization. As we mentioned, GD is extremely sensitive to initial condition since it determines the particular local minimum GD would eventually reach. However, even with a good initialization scheme, through the introduction of randomness, GD can still take exponential time to escape saddle points. This leads us to our next topic, Stochastic Gradient Methods.
-
-
diff --git a/doc/pub/week38/ipynb/ipynb-week38-src.tar.gz b/doc/pub/week38/ipynb/ipynb-week38-src.tar.gz
index 743909c47..d20d5f0c3 100644
Binary files a/doc/pub/week38/ipynb/ipynb-week38-src.tar.gz and b/doc/pub/week38/ipynb/ipynb-week38-src.tar.gz differ
diff --git a/doc/pub/week38/ipynb/week38.ipynb b/doc/pub/week38/ipynb/week38.ipynb
index 9af201d31..919581e58 100644
--- a/doc/pub/week38/ipynb/week38.ipynb
+++ b/doc/pub/week38/ipynb/week38.ipynb
@@ -10,7 +10,7 @@
" \n",
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
"\n",
- "Date: **Sep 24, 2021**\n",
+ "Date: **Sep 28, 2021**\n",
"\n",
"Copyright 1999-2021, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
"\n",
@@ -2916,6 +2916,11 @@
"cell_type": "markdown",
"metadata": {},
"source": [
+ "## Friday September 25\n",
+ "\n",
+ "\n",
+ "\n",
+ "\n",
"## Optimization, the central part of any Machine Learning algortithm\n",
"\n",
"[Overview Video, why do we care about gradient methods?](https://www.uio.no/studier/emner/matnat/fys/FYS-STK3155/h20/forelesningsvideoer/OverarchingAimsWeek39.mp4?vrtx=view-as-webpage)\n",
@@ -3339,1219 +3344,7 @@
"\n",
"Many of these shortcomings can be alleviated by introducing\n",
"randomness. One such method is that of Stochastic Gradient Descent\n",
- "(SGD), see below.\n",
- "\n",
- "\n",
- "\n",
- "## Convex functions\n",
- "\n",
- "Ideally we want our cost/loss function to be convex(concave).\n",
- "\n",
- "First we give the definition of a convex set: A set $C$ in\n",
- "$\\mathbb{R}^n$ is said to be convex if, for all $x$ and $y$ in $C$ and\n",
- "all $t \\in (0,1)$ , the point $(1 − t)x + ty$ also belongs to\n",
- "C. Geometrically this means that every point on the line segment\n",
- "connecting $x$ and $y$ is in $C$ as discussed below.\n",
- "\n",
- "The convex subsets of $\\mathbb{R}$ are the intervals of\n",
- "$\\mathbb{R}$. Examples of convex sets of $\\mathbb{R}^2$ are the\n",
- "regular polygons (triangles, rectangles, pentagons, etc...).\n",
- "\n",
- "## Convex function\n",
- "\n",
- "**Convex function**: Let $X \\subset \\mathbb{R}^n$ be a convex set. Assume that the function $f: X \\rightarrow \\mathbb{R}$ is continuous, then $f$ is said to be convex if $$f(tx_1 + (1-t)x_2) \\leq tf(x_1) + (1-t)f(x_2) $$ for all $x_1, x_2 \\in X$ and for all $t \\in [0,1]$. If $\\leq$ is replaced with a strict inequaltiy in the definition, we demand $x_1 \\neq x_2$ and $t\\in(0,1)$ then $f$ is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting $f(x_1)$ and $f(x_2)$, the value of the function on the interval $[x_1,x_2]$ is always below the line as illustrated below.\n",
- "\n",
- "## Conditions on convex functions\n",
- "\n",
- "In the following we state first and second-order conditions which\n",
- "ensures convexity of a function $f$. We write $D_f$ to denote the\n",
- "domain of $f$, i.e the subset of $R^n$ where $f$ is defined. For more\n",
- "details and proofs we refer to: [S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press](http://stanford.edu/boyd/cvxbook/, 2004).\n",
- "\n",
- "**First order condition.**\n",
- "\n",
- "Suppose $f$ is differentiable (i.e $\\nabla f(x)$ is well defined for\n",
- "all $x$ in the domain of $f$). Then $f$ is convex if and only if $D_f$\n",
- "is a convex set and $$f(y) \\geq f(x) + \\nabla f(x)^T (y-x) $$ holds\n",
- "for all $x,y \\in D_f$. This condition means that for a convex function\n",
- "the first order Taylor expansion (right hand side above) at any point\n",
- "a global under estimator of the function. To convince yourself you can\n",
- "make a drawing of $f(x) = x^2+1$ and draw the tangent line to $f(x)$ and\n",
- "note that it is always below the graph.\n",
- "\n",
- "\n",
- "\n",
- "**Second order condition.**\n",
- "\n",
- "Assume that $f$ is twice\n",
- "differentiable, i.e the Hessian matrix exists at each point in\n",
- "$D_f$. Then $f$ is convex if and only if $D_f$ is a convex set and its\n",
- "Hessian is positive semi-definite for all $x\\in D_f$. For a\n",
- "single-variable function this reduces to $f''(x) \\geq 0$. Geometrically this means that $f$ has nonnegative curvature\n",
- "everywhere.\n",
- "\n",
- "\n",
- "\n",
- "This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.\n",
- "\n",
- "## More on convex functions\n",
- "\n",
- "The next result is of great importance to us and the reason why we are\n",
- "going on about convex functions. In machine learning we frequently\n",
- "have to minimize a loss/cost function in order to find the best\n",
- "parameters for the model we are considering. \n",
- "\n",
- "Ideally we want the\n",
- "global minimum (for high-dimensional models it is hard to know\n",
- "if we have local or global minimum). However, if the cost/loss function\n",
- "is convex the following result provides invaluable information:\n",
- "\n",
- "**Any minimum is global for convex functions.**\n",
- "\n",
- "Consider the problem of finding $x \\in \\mathbb{R}^n$ such that $f(x)$\n",
- "is minimal, where $f$ is convex and differentiable. Then, any point\n",
- "$x^*$ that satisfies $\\nabla f(x^*) = 0$ is a global minimum.\n",
- "\n",
- "\n",
- "\n",
- "This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.\n",
- "\n",
- "## Some simple problems\n",
- "\n",
- "1. Show that $f(x)=x^2$ is convex for $x \\in \\mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \\in D_f$ and any $\\lambda \\in [0,1]$ $\\lambda f(x)+(1-\\lambda)f(y)-f(\\lambda x + (1-\\lambda) y ) \\geq 0$.\n",
- "\n",
- "2. Using the second order condition show that the following functions are convex on the specified domain.\n",
- "\n",
- " * $f(x) = e^x$ is convex for $x \\in \\mathbb{R}$.\n",
- "\n",
- " * $g(x) = -\\ln(x)$ is convex for $x \\in (0,\\infty)$.\n",
- "\n",
- "\n",
- "3. Let $f(x) = x^2$ and $g(x) = e^x$. Show that $f(g(x))$ and $g(f(x))$ is convex for $x \\in \\mathbb{R}$. Also show that if $f(x)$ is any convex function than $h(x) = e^{f(x)}$ is convex.\n",
- "\n",
- "4. A norm is any function that satisfy the following properties\n",
- "\n",
- " * $f(\\alpha x) = |\\alpha| f(x)$ for all $\\alpha \\in \\mathbb{R}$.\n",
- "\n",
- " * $f(x+y) \\leq f(x) + f(y)$\n",
- "\n",
- " * $f(x) \\leq 0$ for all $x \\in \\mathbb{R}^n$ with equality if and only if $x = 0$\n",
- "\n",
- "\n",
- "Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).\n",
- "\n",
- "\n",
- "## Friday September 25\n",
- "\n",
- "[Video of Lecture](https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureSeptember25.mp4?vrtx=view-as-webpage) and [link to handwritten notes](https://github.com/CompPhysics/MachineLearning/blob/master/doc/HandWrittenNotes/NotesSeptember25.pdf).\n",
- "\n",
- "\n",
- "## Standard steepest descent\n",
- "\n",
- "\n",
- "Before we proceed, we would like to discuss the approach called the\n",
- "**standard Steepest descent** (different from the above steepest descent discussion), which again leads to us having to be able\n",
- "to compute a matrix. It belongs to the class of Conjugate Gradient methods (CG).\n",
- "\n",
- "[The success of the CG method](https://www.cs.cmu.edu/~quake-papers/painless-conjugate-gradient.pdf)\n",
- "for finding solutions of non-linear problems is based on the theory\n",
- "of conjugate gradients for linear systems of equations. It belongs to\n",
- "the class of iterative methods for solving problems from linear\n",
- "algebra of the type"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{A}\\boldsymbol{x} = \\boldsymbol{b}.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "In the iterative process we end up with a problem like"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{r}= \\boldsymbol{b}-\\boldsymbol{A}\\boldsymbol{x},\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "where $\\boldsymbol{r}$ is the so-called residual or error in the iterative process.\n",
- "\n",
- "When we have found the exact solution, $\\boldsymbol{r}=0$.\n",
- "\n",
- "## Gradient method\n",
- "\n",
- "The residual is zero when we reach the minimum of the quadratic equation"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "P(\\boldsymbol{x})=\\frac{1}{2}\\boldsymbol{x}^T\\boldsymbol{A}\\boldsymbol{x} - \\boldsymbol{x}^T\\boldsymbol{b},\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "with the constraint that the matrix $\\boldsymbol{A}$ is positive definite and\n",
- "symmetric. This defines also the Hessian and we want it to be positive definite. \n",
- "\n",
- "\n",
- "## Steepest descent method\n",
- "\n",
- "We denote the initial guess for $\\boldsymbol{x}$ as $\\boldsymbol{x}_0$. \n",
- "We can assume without loss of generality that"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{x}_0=0,\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "or consider the system"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{A}\\boldsymbol{z} = \\boldsymbol{b}-\\boldsymbol{A}\\boldsymbol{x}_0,\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "instead.\n",
- "\n",
- "\n",
- "## Steepest descent method\n",
- "One can show that the solution $\\boldsymbol{x}$ is also the unique minimizer of the quadratic form"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "f(\\boldsymbol{x}) = \\frac{1}{2}\\boldsymbol{x}^T\\boldsymbol{A}\\boldsymbol{x} - \\boldsymbol{x}^T \\boldsymbol{x} , \\quad \\boldsymbol{x}\\in\\mathbf{R}^n.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "This suggests taking the first basis vector $\\boldsymbol{r}_1$ (see below for definition) \n",
- "to be the gradient of $f$ at $\\boldsymbol{x}=\\boldsymbol{x}_0$, \n",
- "which equals"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{A}\\boldsymbol{x}_0-\\boldsymbol{b},\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "and \n",
- "$\\boldsymbol{x}_0=0$ it is equal $-\\boldsymbol{b}$.\n",
- "\n",
- "\n",
- "\n",
- "## Final expressions\n",
- "We can compute the residual iteratively as"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{r}_{k+1}=\\boldsymbol{b}-\\boldsymbol{A}\\boldsymbol{x}_{k+1},\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "which equals"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{b}-\\boldsymbol{A}(\\boldsymbol{x}_k+\\alpha_k\\boldsymbol{r}_k),\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "or"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "(\\boldsymbol{b}-\\boldsymbol{A}\\boldsymbol{x}_k)-\\alpha_k\\boldsymbol{A}\\boldsymbol{r}_k,\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "which gives"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\alpha_k = \\frac{\\boldsymbol{r}_k^T\\boldsymbol{r}_k}{\\boldsymbol{r}_k^T\\boldsymbol{A}\\boldsymbol{r}_k}\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "leading to the iterative scheme"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{x}_{k+1}=\\boldsymbol{x}_k-\\alpha_k\\boldsymbol{r}_{k},\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "## Steepest descent example"
- ]
- },
- {
- "cell_type": "code",
- "execution_count": null,
- "metadata": {
- "collapsed": false,
- "editable": true
- },
- "outputs": [],
- "source": [
- "import numpy as np\n",
- "import numpy.linalg as la\n",
- "\n",
- "import scipy.optimize as sopt\n",
- "\n",
- "import matplotlib.pyplot as pt\n",
- "from mpl_toolkits.mplot3d import axes3d\n",
- "\n",
- "def f(x):\n",
- " return 0.5*x[0]**2 + 2.5*x[1]**2\n",
- "\n",
- "def df(x):\n",
- " return np.array([x[0], 5*x[1]])\n",
- "\n",
- "fig = pt.figure()\n",
- "ax = fig.gca(projection=\"3d\")\n",
- "\n",
- "xmesh, ymesh = np.mgrid[-2:2:50j,-2:2:50j]\n",
- "fmesh = f(np.array([xmesh, ymesh]))\n",
- "ax.plot_surface(xmesh, ymesh, fmesh)"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "And then as countor plot"
- ]
- },
- {
- "cell_type": "code",
- "execution_count": null,
- "metadata": {
- "collapsed": false,
- "editable": true
- },
- "outputs": [],
- "source": [
- "pt.axis(\"equal\")\n",
- "pt.contour(xmesh, ymesh, fmesh)\n",
- "guesses = [np.array([2, 2./5])]"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "Find guesses"
- ]
- },
- {
- "cell_type": "code",
- "execution_count": null,
- "metadata": {
- "collapsed": false,
- "editable": true
- },
- "outputs": [],
- "source": [
- "x = guesses[-1]\n",
- "s = -df(x)"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "Run it!"
- ]
- },
- {
- "cell_type": "code",
- "execution_count": null,
- "metadata": {
- "collapsed": false,
- "editable": true
- },
- "outputs": [],
- "source": [
- "def f1d(alpha):\n",
- " return f(x + alpha*s)\n",
- "\n",
- "alpha_opt = sopt.golden(f1d)\n",
- "next_guess = x + alpha_opt * s\n",
- "guesses.append(next_guess)\n",
- "print(next_guess)"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "What happened?"
- ]
- },
- {
- "cell_type": "code",
- "execution_count": null,
- "metadata": {
- "collapsed": false,
- "editable": true
- },
- "outputs": [],
- "source": [
- "pt.axis(\"equal\")\n",
- "pt.contour(xmesh, ymesh, fmesh, 50)\n",
- "it_array = np.array(guesses)\n",
- "pt.plot(it_array.T[0], it_array.T[1], \"x-\")"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "## Conjugate gradient method\n",
- "In the CG method we define so-called conjugate directions and two vectors \n",
- "$\\boldsymbol{s}$ and $\\boldsymbol{t}$\n",
- "are said to be\n",
- "conjugate if"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{s}^T\\boldsymbol{A}\\boldsymbol{t}= 0.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "The philosophy of the CG method is to perform searches in various conjugate directions\n",
- "of our vectors $\\boldsymbol{x}_i$ obeying the above criterion, namely"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{x}_i^T\\boldsymbol{A}\\boldsymbol{x}_j= 0.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "Two vectors are conjugate if they are orthogonal with respect to \n",
- "this inner product. Being conjugate is a symmetric relation: if $\\boldsymbol{s}$ is conjugate to $\\boldsymbol{t}$, then $\\boldsymbol{t}$ is conjugate to $\\boldsymbol{s}$.\n",
- "\n",
- "\n",
- "\n",
- "## Conjugate gradient method\n",
- "An example is given by the eigenvectors of the matrix"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{v}_i^T\\boldsymbol{A}\\boldsymbol{v}_j= \\lambda\\boldsymbol{v}_i^T\\boldsymbol{v}_j,\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "which is zero unless $i=j$.\n",
- "\n",
- "\n",
- "\n",
- "\n",
- "## Conjugate gradient method\n",
- "Assume now that we have a symmetric positive-definite matrix $\\boldsymbol{A}$ of size\n",
- "$n\\times n$. At each iteration $i+1$ we obtain the conjugate direction of a vector"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{x}_{i+1}=\\boldsymbol{x}_{i}+\\alpha_i\\boldsymbol{p}_{i}.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "We assume that $\\boldsymbol{p}_{i}$ is a sequence of $n$ mutually conjugate directions. \n",
- "Then the $\\boldsymbol{p}_{i}$ form a basis of $R^n$ and we can expand the solution \n",
- "$ \\boldsymbol{A}\\boldsymbol{x} = \\boldsymbol{b}$ in this basis, namely"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{x} = \\sum^{n}_{i=1} \\alpha_i \\boldsymbol{p}_i.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "## Conjugate gradient method\n",
- "The coefficients are given by"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\mathbf{A}\\mathbf{x} = \\sum^{n}_{i=1} \\alpha_i \\mathbf{A} \\mathbf{p}_i = \\mathbf{b}.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "Multiplying with $\\boldsymbol{p}_k^T$ from the left gives"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{p}_k^T \\boldsymbol{A}\\boldsymbol{x} = \\sum^{n}_{i=1} \\alpha_i\\boldsymbol{p}_k^T \\boldsymbol{A}\\boldsymbol{p}_i= \\boldsymbol{p}_k^T \\boldsymbol{b},\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "and we can define the coefficients $\\alpha_k$ as"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\alpha_k = \\frac{\\boldsymbol{p}_k^T \\boldsymbol{b}}{\\boldsymbol{p}_k^T \\boldsymbol{A} \\boldsymbol{p}_k}\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "## Conjugate gradient method and iterations\n",
- "\n",
- "If we choose the conjugate vectors $\\boldsymbol{p}_k$ carefully, \n",
- "then we may not need all of them to obtain a good approximation to the solution \n",
- "$\\boldsymbol{x}$. \n",
- "We want to regard the conjugate gradient method as an iterative method. \n",
- "This will us to solve systems where $n$ is so large that the direct \n",
- "method would take too much time.\n",
- "\n",
- "We denote the initial guess for $\\boldsymbol{x}$ as $\\boldsymbol{x}_0$. \n",
- "We can assume without loss of generality that"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{x}_0=0,\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "or consider the system"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{A}\\boldsymbol{z} = \\boldsymbol{b}-\\boldsymbol{A}\\boldsymbol{x}_0,\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "instead.\n",
- "\n",
- "\n",
- "\n",
- "\n",
- "## Conjugate gradient method\n",
- "One can show that the solution $\\boldsymbol{x}$ is also the unique minimizer of the quadratic form"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "f(\\boldsymbol{x}) = \\frac{1}{2}\\boldsymbol{x}^T\\boldsymbol{A}\\boldsymbol{x} - \\boldsymbol{x}^T \\boldsymbol{x} , \\quad \\boldsymbol{x}\\in\\mathbf{R}^n.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "This suggests taking the first basis vector $\\boldsymbol{p}_1$ \n",
- "to be the gradient of $f$ at $\\boldsymbol{x}=\\boldsymbol{x}_0$, \n",
- "which equals"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{A}\\boldsymbol{x}_0-\\boldsymbol{b},\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "and \n",
- "$\\boldsymbol{x}_0=0$ it is equal $-\\boldsymbol{b}$.\n",
- "The other vectors in the basis will be conjugate to the gradient, \n",
- "hence the name conjugate gradient method.\n",
- "\n",
- "\n",
- "\n",
- "\n",
- "## Conjugate gradient method\n",
- "Let $\\boldsymbol{r}_k$ be the residual at the $k$-th step:"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{r}_k=\\boldsymbol{b}-\\boldsymbol{A}\\boldsymbol{x}_k.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "Note that $\\boldsymbol{r}_k$ is the negative gradient of $f$ at \n",
- "$\\boldsymbol{x}=\\boldsymbol{x}_k$, \n",
- "so the gradient descent method would be to move in the direction $\\boldsymbol{r}_k$. \n",
- "Here, we insist that the directions $\\boldsymbol{p}_k$ are conjugate to each other, \n",
- "so we take the direction closest to the gradient $\\boldsymbol{r}_k$ \n",
- "under the conjugacy constraint. \n",
- "This gives the following expression"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{p}_{k+1}=\\boldsymbol{r}_k-\\frac{\\boldsymbol{p}_k^T \\boldsymbol{A}\\boldsymbol{r}_k}{\\boldsymbol{p}_k^T\\boldsymbol{A}\\boldsymbol{p}_k} \\boldsymbol{p}_k.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "## Conjugate gradient method\n",
- "We can also compute the residual iteratively as"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{r}_{k+1}=\\boldsymbol{b}-\\boldsymbol{A}\\boldsymbol{x}_{k+1},\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "which equals"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{b}-\\boldsymbol{A}(\\boldsymbol{x}_k+\\alpha_k\\boldsymbol{p}_k),\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "or"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "(\\boldsymbol{b}-\\boldsymbol{A}\\boldsymbol{x}_k)-\\alpha_k\\boldsymbol{A}\\boldsymbol{p}_k,\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "which gives"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{r}_{k+1}=\\boldsymbol{r}_k-\\boldsymbol{A}\\boldsymbol{p}_{k},\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "\n",
- "## Revisiting some of our first Linear Regression Encounters\n",
- "\n",
- "We will use linear regression as a case study for the gradient descent\n",
- "methods. Linear regression is a great test case for the gradient\n",
- "descent methods discussed in the lectures since it has several\n",
- "desirable properties such as:\n",
- "\n",
- "1. An analytical solution (recall homework set 1).\n",
- "\n",
- "2. The gradient can be computed analytically.\n",
- "\n",
- "3. The cost function is convex which guarantees that gradient descent converges for small enough learning rates\n",
- "\n",
- "We revisit an example similar to what we had in the first homework set. We had a function of the type"
- ]
- },
- {
- "cell_type": "code",
- "execution_count": null,
- "metadata": {
- "collapsed": false,
- "editable": true
- },
- "outputs": [],
- "source": [
- "x = 2*np.random.rand(m,1)\n",
- "y = 4+3*x+np.random.randn(m,1)"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "with $x_i \\in [0,1] $ is chosen randomly using a uniform distribution. Additionally we have a stochastic noise chosen according to a normal distribution $\\cal {N}(0,1)$. \n",
- "The linear regression model is given by"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "h_\\beta(x) = \\boldsymbol{y} = \\beta_0 + \\beta_1 x,\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "such that"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{y}_i = \\beta_0 + \\beta_1 x_i.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "\n",
- "## Gradient descent example\n",
- "\n",
- "Let $\\mathbf{y} = (y_1,\\cdots,y_n)^T$, $\\mathbf{\\boldsymbol{y}} = (\\boldsymbol{y}_1,\\cdots,\\boldsymbol{y}_n)^T$ and $\\beta = (\\beta_0, \\beta_1)^T$\n",
- "\n",
- "It is convenient to write $\\mathbf{\\boldsymbol{y}} = X\\beta$ where $X \\in \\mathbb{R}^{100 \\times 2} $ is the design matrix given by (we keep the intercept here)"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "X \\equiv \\begin{bmatrix}\n",
- "1 & x_1 \\\\\n",
- "\\vdots & \\vdots \\\\\n",
- "1 & x_{100} & \\\\\n",
- "\\end{bmatrix}.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "The cost/loss/risk function is given by ("
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "C(\\beta) = \\frac{1}{n}||X\\beta-\\mathbf{y}||_{2}^{2} = \\frac{1}{n}\\sum_{i=1}^{100}\\left[ (\\beta_0 + \\beta_1 x_i)^2 - 2 y_i (\\beta_0 + \\beta_1 x_i) + y_i^2\\right]\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "and we want to find $\\beta$ such that $C(\\beta)$ is minimized.\n",
- "\n",
- "## The derivative of the cost/loss function\n",
- "\n",
- "Computing $\\partial C(\\beta) / \\partial \\beta_0$ and $\\partial C(\\beta) / \\partial \\beta_1$ we can show that the gradient can be written as"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\nabla_{\\beta} C(\\beta) = \\frac{2}{n}\\begin{bmatrix} \\sum_{i=1}^{100} \\left(\\beta_0+\\beta_1x_i-y_i\\right) \\\\\n",
- "\\sum_{i=1}^{100}\\left( x_i (\\beta_0+\\beta_1x_i)-y_ix_i\\right) \\\\\n",
- "\\end{bmatrix} = \\frac{2}{n}X^T(X\\beta - \\mathbf{y}),\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "where $X$ is the design matrix defined above.\n",
- "\n",
- "## The Hessian matrix\n",
- "The Hessian matrix of $C(\\beta)$ is given by"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\boldsymbol{H} \\equiv \\begin{bmatrix}\n",
- "\\frac{\\partial^2 C(\\beta)}{\\partial \\beta_0^2} & \\frac{\\partial^2 C(\\beta)}{\\partial \\beta_0 \\partial \\beta_1} \\\\\n",
- "\\frac{\\partial^2 C(\\beta)}{\\partial \\beta_0 \\partial \\beta_1} & \\frac{\\partial^2 C(\\beta)}{\\partial \\beta_1^2} & \\\\\n",
- "\\end{bmatrix} = \\frac{2}{n}X^T X.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "This result implies that $C(\\beta)$ is a convex function since the matrix $X^T X$ always is positive semi-definite.\n",
- "\n",
- "\n",
- "\n",
- "\n",
- "## Simple program\n",
- "\n",
- "We can now write a program that minimizes $C(\\beta)$ using the gradient descent method with a constant learning rate $\\gamma$ according to"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\beta_{k+1} = \\beta_k - \\gamma \\nabla_\\beta C(\\beta_k), \\ k=0,1,\\cdots\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "We can use the expression we computed for the gradient and let use a\n",
- "$\\beta_0$ be chosen randomly and let $\\gamma = 0.001$. Stop iterating\n",
- "when $||\\nabla_\\beta C(\\beta_k) || \\leq \\epsilon = 10^{-8}$. **Note that the code below does not include the latter stop criterion**.\n",
- "\n",
- "And finally we can compare our solution for $\\beta$ with the analytic result given by \n",
- "$\\beta= (X^TX)^{-1} X^T \\mathbf{y}$.\n",
- "\n",
- "## Gradient Descent Example\n",
- "\n",
- "Here our simple example"
- ]
- },
- {
- "cell_type": "code",
- "execution_count": null,
- "metadata": {
- "collapsed": false,
- "editable": true
- },
- "outputs": [],
- "source": [
- "\n",
- "# Importing various packages\n",
- "from random import random, seed\n",
- "import numpy as np\n",
- "import matplotlib.pyplot as plt\n",
- "from mpl_toolkits.mplot3d import Axes3D\n",
- "from matplotlib import cm\n",
- "from matplotlib.ticker import LinearLocator, FormatStrFormatter\n",
- "import sys\n",
- "\n",
- "# the number of datapoints\n",
- "n = 100\n",
- "x = 2*np.random.rand(n,1)\n",
- "y = 4+3*x+np.random.randn(n,1)\n",
- "\n",
- "X = np.c_[np.ones((n,1)), x]\n",
- "# Hessian matrix\n",
- "H = (2.0/n)* X.T @ X\n",
- "# Get the eigenvalues\n",
- "EigValues, EigVectors = np.linalg.eig(H)\n",
- "print(EigValues)\n",
- "\n",
- "beta_linreg = np.linalg.inv(X.T @ X) @ X.T @ y\n",
- "print(beta_linreg)\n",
- "beta = np.random.randn(2,1)\n",
- "\n",
- "eta = 1.0/np.max(EigValues)\n",
- "Niterations = 1000\n",
- "\n",
- "for iter in range(Niterations):\n",
- " gradient = (2.0/n)*X.T @ (X @ beta-y)\n",
- " beta -= eta*gradient\n",
- "\n",
- "print(beta)\n",
- "xnew = np.array([[0],[2]])\n",
- "xbnew = np.c_[np.ones((2,1)), xnew]\n",
- "ypredict = xbnew.dot(beta)\n",
- "ypredict2 = xbnew.dot(beta_linreg)\n",
- "plt.plot(xnew, ypredict, \"r-\")\n",
- "plt.plot(xnew, ypredict2, \"b-\")\n",
- "plt.plot(x, y ,'ro')\n",
- "plt.axis([0,2.0,0, 15.0])\n",
- "plt.xlabel(r'$x$')\n",
- "plt.ylabel(r'$y$')\n",
- "plt.title(r'Gradient descent example')\n",
- "plt.show()"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "## And a corresponding example using **scikit-learn**"
- ]
- },
- {
- "cell_type": "code",
- "execution_count": null,
- "metadata": {
- "collapsed": false,
- "editable": true
- },
- "outputs": [],
- "source": [
- "# Importing various packages\n",
- "from random import random, seed\n",
- "import numpy as np\n",
- "import matplotlib.pyplot as plt\n",
- "from sklearn.linear_model import SGDRegressor\n",
- "\n",
- "n = 100\n",
- "x = 2*np.random.rand(n,1)\n",
- "y = 4+3*x+np.random.randn(n,1)\n",
- "\n",
- "X = np.c_[np.ones((n,1)), x]\n",
- "beta_linreg = np.linalg.inv(X.T @ X) @ (X.T @ y)\n",
- "print(beta_linreg)\n",
- "sgdreg = SGDRegressor(max_iter = 50, penalty=None, eta0=0.1)\n",
- "sgdreg.fit(x,y.ravel())\n",
- "print(sgdreg.intercept_, sgdreg.coef_)"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "\n",
- "## Gradient descent and Ridge\n",
- "\n",
- "We have also discussed Ridge regression where the loss function contains a regularized term given by the $L_2$ norm of $\\beta$,"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "C_{\\text{ridge}}(\\beta) = \\frac{1}{n}||X\\beta -\\mathbf{y}||^2 + \\lambda ||\\beta||^2, \\ \\lambda \\geq 0.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "In order to minimize $C_{\\text{ridge}}(\\beta)$ using GD we only have adjust the gradient as follows"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\nabla_\\beta C_{\\text{ridge}}(\\beta) = \\frac{2}{n}\\begin{bmatrix} \\sum_{i=1}^{100} \\left(\\beta_0+\\beta_1x_i-y_i\\right) \\\\\n",
- "\\sum_{i=1}^{100}\\left( x_i (\\beta_0+\\beta_1x_i)-y_ix_i\\right) \\\\\n",
- "\\end{bmatrix} + 2\\lambda\\begin{bmatrix} \\beta_0 \\\\ \\beta_1\\end{bmatrix} = 2 (X^T(X\\beta - \\mathbf{y})+\\lambda \\beta).\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "We can easily extend our program to minimize $C_{\\text{ridge}}(\\beta)$ using gradient descent and compare with the analytical solution given by"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "$$\n",
- "\\beta_{\\text{ridge}} = \\left(X^T X + \\lambda I_{2 \\times 2} \\right)^{-1} X^T \\mathbf{y}.\n",
- "$$"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "## Program example for gradient descent with Ridge Regression"
- ]
- },
- {
- "cell_type": "code",
- "execution_count": null,
- "metadata": {
- "collapsed": false,
- "editable": true
- },
- "outputs": [],
- "source": [
- "from random import random, seed\n",
- "import numpy as np\n",
- "import matplotlib.pyplot as plt\n",
- "from mpl_toolkits.mplot3d import Axes3D\n",
- "from matplotlib import cm\n",
- "from matplotlib.ticker import LinearLocator, FormatStrFormatter\n",
- "import sys\n",
- "\n",
- "# the number of datapoints\n",
- "n = 100\n",
- "x = 2*np.random.rand(n,1)\n",
- "y = 4+3*x+np.random.randn(n,1)\n",
- "\n",
- "X = np.c_[np.ones((n,1)), x]\n",
- "XT_X = X.T @ X\n",
- "\n",
- "#Ridge parameter lambda\n",
- "lmbda = 0.001\n",
- "Id = lmbda* np.eye(XT_X.shape[0])\n",
- "\n",
- "beta_linreg = np.linalg.inv(XT_X+Id) @ X.T @ y\n",
- "print(beta_linreg)\n",
- "# Start plain gradient descent\n",
- "beta = np.random.randn(2,1)\n",
- "\n",
- "eta = 0.1\n",
- "Niterations = 100\n",
- "\n",
- "for iter in range(Niterations):\n",
- " gradients = 2.0/n*X.T @ (X @ (beta)-y)+2*lmbda*beta\n",
- " beta -= eta*gradients\n",
- "\n",
- "print(beta)\n",
- "ypredict = X @ beta\n",
- "ypredict2 = X @ beta_linreg\n",
- "plt.plot(x, ypredict, \"r-\")\n",
- "plt.plot(x, ypredict2, \"b-\")\n",
- "plt.plot(x, y ,'ro')\n",
- "plt.axis([0,2.0,0, 15.0])\n",
- "plt.xlabel(r'$x$')\n",
- "plt.ylabel(r'$y$')\n",
- "plt.title(r'Gradient descent example for Ridge')\n",
- "plt.show()"
- ]
- },
- {
- "cell_type": "markdown",
- "metadata": {},
- "source": [
- "## Using gradient descent methods, limitations\n",
- "\n",
- "* **Gradient descent (GD) finds local minima of our function**. Since the GD algorithm is deterministic, if it converges, it will converge to a local minimum of our cost/loss/risk function. Because in ML we are often dealing with extremely rugged landscapes with many local minima, this can lead to poor performance.\n",
- "\n",
- "* **GD is sensitive to initial conditions**. One consequence of the local nature of GD is that initial conditions matter. Depending on where one starts, one will end up at a different local minima. Therefore, it is very important to think about how one initializes the training process. This is true for GD as well as more complicated variants of GD.\n",
- "\n",
- "* **Gradients are computationally expensive to calculate for large datasets**. In many cases in statistics and ML, the cost/loss/risk function is a sum of terms, with one term for each data point. For example, in linear regression, $E \\propto \\sum_{i=1}^n (y_i - \\mathbf{w}^T\\cdot\\mathbf{x}_i)^2$; for logistic regression, the square error is replaced by the cross entropy. To calculate the gradient we have to sum over *all* $n$ data points. Doing this at every GD step becomes extremely computationally expensive. An ingenious solution to this, is to calculate the gradients using small subsets of the data called \"mini batches\". This has the added benefit of introducing stochasticity into our algorithm.\n",
- "\n",
- "* **GD is very sensitive to choices of learning rates**. GD is extremely sensitive to the choice of learning rates. If the learning rate is very small, the training process take an extremely long time. For larger learning rates, GD can diverge and give poor results. Furthermore, depending on what the local landscape looks like, we have to modify the learning rates to ensure convergence. Ideally, we would *adaptively* choose the learning rates to match the landscape.\n",
- "\n",
- "* **GD treats all directions in parameter space uniformly.** Another major drawback of GD is that unlike Newton's method, the learning rate for GD is the same in all directions in parameter space. For this reason, the maximum learning rate is set by the behavior of the steepest direction and this can significantly slow down training. Ideally, we would like to take large steps in flat directions and small steps in steep directions. Since we are exploring rugged landscapes where curvatures change, this requires us to keep track of not only the gradient but second derivatives. The ideal scenario would be to calculate the Hessian but this proves to be too computationally expensive. \n",
- "\n",
- "* GD can take exponential time to escape saddle points, even with random initialization. As we mentioned, GD is extremely sensitive to initial condition since it determines the particular local minimum GD would eventually reach. However, even with a good initialization scheme, through the introduction of randomness, GD can still take exponential time to escape saddle points. This leads us to our next topic, Stochastic Gradient Methods."
+ "(SGD), see below."
]
}
],
diff --git a/doc/src/week38/week38.do.txt b/doc/src/week38/week38.do.txt
index 7b129742c..4d07e43f6 100644
--- a/doc/src/week38/week38.do.txt
+++ b/doc/src/week38/week38.do.txt
@@ -1905,6 +1905,8 @@ plt.show()
!ec
+!split
+===== Friday September 25 =====
@@ -2195,738 +2197,3 @@ randomness. One such method is that of Stochastic Gradient Descent
(SGD), see below.
-!split
-===== Convex functions =====
-
-Ideally we want our cost/loss function to be convex(concave).
-
-First we give the definition of a convex set: A set $C$ in
-$\mathbb{R}^n$ is said to be convex if, for all $x$ and $y$ in $C$ and
-all $t \in (0,1)$ , the point $(1 − t)x + ty$ also belongs to
-C. Geometrically this means that every point on the line segment
-connecting $x$ and $y$ is in $C$ as discussed below.
-
-The convex subsets of $\mathbb{R}$ are the intervals of
-$\mathbb{R}$. Examples of convex sets of $\mathbb{R}^2$ are the
-regular polygons (triangles, rectangles, pentagons, etc...).
-
-!split
-===== Convex function =====
-
-_Convex function_: Let $X \subset \mathbb{R}^n$ be a convex set. Assume that the function $f: X \rightarrow \mathbb{R}$ is continuous, then $f$ is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all $x_1, x_2 \in X$ and for all $t \in [0,1]$. If $\leq$ is replaced with a strict inequaltiy in the definition, we demand $x_1 \neq x_2$ and $t\in(0,1)$ then $f$ is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting $f(x_1)$ and $f(x_2)$, the value of the function on the interval $[x_1,x_2]$ is always below the line as illustrated below.
-
-!split
-===== Conditions on convex functions =====
-
-In the following we state first and second-order conditions which
-ensures convexity of a function $f$. We write $D_f$ to denote the
-domain of $f$, i.e the subset of $R^n$ where $f$ is defined. For more
-details and proofs we refer to: "S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press":"http://stanford.edu/boyd/cvxbook/, 2004".
-
-!bblock First order condition
-Suppose $f$ is differentiable (i.e $\nabla f(x)$ is well defined for
-all $x$ in the domain of $f$). Then $f$ is convex if and only if $D_f$
-is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
-for all $x,y \in D_f$. This condition means that for a convex function
-the first order Taylor expansion (right hand side above) at any point
-a global under estimator of the function. To convince yourself you can
-make a drawing of $f(x) = x^2+1$ and draw the tangent line to $f(x)$ and
-note that it is always below the graph.
-!eblock
-
-!bblock Second order condition
-Assume that $f$ is twice
-differentiable, i.e the Hessian matrix exists at each point in
-$D_f$. Then $f$ is convex if and only if $D_f$ is a convex set and its
-Hessian is positive semi-definite for all $x\in D_f$. For a
-single-variable function this reduces to $f''(x) \geq 0$. Geometrically this means that $f$ has nonnegative curvature
-everywhere.
-!eblock
-
-This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
-
-!split
-===== More on convex functions =====
-
-The next result is of great importance to us and the reason why we are
-going on about convex functions. In machine learning we frequently
-have to minimize a loss/cost function in order to find the best
-parameters for the model we are considering.
-
-Ideally we want the
-global minimum (for high-dimensional models it is hard to know
-if we have local or global minimum). However, if the cost/loss function
-is convex the following result provides invaluable information:
-
-!bblock Any minimum is global for convex functions
-Consider the problem of finding $x \in \mathbb{R}^n$ such that $f(x)$
-is minimal, where $f$ is convex and differentiable. Then, any point
-$x^*$ that satisfies $\nabla f(x^*) = 0$ is a global minimum.
-!eblock
-
-This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
-
-!split
-===== Some simple problems =====
-
-o Show that $f(x)=x^2$ is convex for $x \in \mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \in D_f$ and any $\lambda \in [0,1]$ $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.
-
-o Using the second order condition show that the following functions are convex on the specified domain.
- * $f(x) = e^x$ is convex for $x \in \mathbb{R}$.
- * $g(x) = -\ln(x)$ is convex for $x \in (0,\infty)$.
-o Let $f(x) = x^2$ and $g(x) = e^x$. Show that $f(g(x))$ and $g(f(x))$ is convex for $x \in \mathbb{R}$. Also show that if $f(x)$ is any convex function than $h(x) = e^{f(x)}$ is convex.
-
-o A norm is any function that satisfy the following properties
- * $f(\alpha x) = |\alpha| f(x)$ for all $\alpha \in \mathbb{R}$.
- * $f(x+y) \leq f(x) + f(y)$
- * $f(x) \leq 0$ for all $x \in \mathbb{R}^n$ with equality if and only if $x = 0$
-
-Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
-
-
-!split
-===== Friday September 25 =====
-
-"Video of Lecture":"https://www.uio.no/studier/emner/matnat/fys/FYS-STK4155/h20/forelesningsvideoer/LectureSeptember25.mp4?vrtx=view-as-webpage" and "link to handwritten notes":"https://github.com/CompPhysics/MachineLearning/blob/master/doc/HandWrittenNotes/NotesSeptember25.pdf".
-
-
-!split
-===== Standard steepest descent =====
-
-
-Before we proceed, we would like to discuss the approach called the
-_standard Steepest descent_ (different from the above steepest descent discussion), which again leads to us having to be able
-to compute a matrix. It belongs to the class of Conjugate Gradient methods (CG).
-
-"The success of the CG method":"https://www.cs.cmu.edu/~quake-papers/painless-conjugate-gradient.pdf"
-for finding solutions of non-linear problems is based on the theory
-of conjugate gradients for linear systems of equations. It belongs to
-the class of iterative methods for solving problems from linear
-algebra of the type
-!bt
-\begin{equation*}
-\bm{A}\bm{x} = \bm{b}.
-\end{equation*}
-!et
-
-In the iterative process we end up with a problem like
-
-!bt
-\begin{equation*}
- \bm{r}= \bm{b}-\bm{A}\bm{x},
-\end{equation*}
-!et
-where $\bm{r}$ is the so-called residual or error in the iterative process.
-
-When we have found the exact solution, $\bm{r}=0$.
-
-!split
-===== Gradient method =====
-
-The residual is zero when we reach the minimum of the quadratic equation
-!bt
-\begin{equation*}
- P(\bm{x})=\frac{1}{2}\bm{x}^T\bm{A}\bm{x} - \bm{x}^T\bm{b},
-\end{equation*}
-!et
-
-with the constraint that the matrix $\bm{A}$ is positive definite and
-symmetric. This defines also the Hessian and we want it to be positive definite.
-
-
-!split
-===== Steepest descent method =====
-
-We denote the initial guess for $\bm{x}$ as $\bm{x}_0$.
-We can assume without loss of generality that
-!bt
-\begin{equation*}
-\bm{x}_0=0,
-\end{equation*}
-!et
-or consider the system
-!bt
-\begin{equation*}
-\bm{A}\bm{z} = \bm{b}-\bm{A}\bm{x}_0,
-\end{equation*}
-!et
-instead.
-
-
-!split
-===== Steepest descent method =====
-!bblock
-One can show that the solution $\bm{x}$ is also the unique minimizer of the quadratic form
-!bt
-\begin{equation*}
- f(\bm{x}) = \frac{1}{2}\bm{x}^T\bm{A}\bm{x} - \bm{x}^T \bm{x} , \quad \bm{x}\in\mathbf{R}^n.
-\end{equation*}
-!et
-This suggests taking the first basis vector $\bm{r}_1$ (see below for definition)
-to be the gradient of $f$ at $\bm{x}=\bm{x}_0$,
-which equals
-!bt
-\begin{equation*}
-\bm{A}\bm{x}_0-\bm{b},
-\end{equation*}
-!et
-and
-$\bm{x}_0=0$ it is equal $-\bm{b}$.
-
-!eblock
-
-!split
-===== Final expressions =====
-!bblock
-We can compute the residual iteratively as
-!bt
-\begin{equation*}
-\bm{r}_{k+1}=\bm{b}-\bm{A}\bm{x}_{k+1},
- \end{equation*}
-!et
-which equals
-!bt
-\begin{equation*}
-\bm{b}-\bm{A}(\bm{x}_k+\alpha_k\bm{r}_k),
- \end{equation*}
-!et
-or
-!bt
-\begin{equation*}
-(\bm{b}-\bm{A}\bm{x}_k)-\alpha_k\bm{A}\bm{r}_k,
- \end{equation*}
-!et
-which gives
-
-!bt
-\[
-\alpha_k = \frac{\bm{r}_k^T\bm{r}_k}{\bm{r}_k^T\bm{A}\bm{r}_k}
-\]
-!et
-leading to the iterative scheme
-!bt
-\begin{equation*}
-\bm{x}_{k+1}=\bm{x}_k-\alpha_k\bm{r}_{k},
- \end{equation*}
-!et
-!eblock
-
-
-
-!split
-===== Steepest descent example =====
-
-!bc pycod
-import numpy as np
-import numpy.linalg as la
-
-import scipy.optimize as sopt
-
-import matplotlib.pyplot as pt
-from mpl_toolkits.mplot3d import axes3d
-
-def f(x):
- return 0.5*x[0]**2 + 2.5*x[1]**2
-
-def df(x):
- return np.array([x[0], 5*x[1]])
-
-fig = pt.figure()
-ax = fig.gca(projection="3d")
-
-xmesh, ymesh = np.mgrid[-2:2:50j,-2:2:50j]
-fmesh = f(np.array([xmesh, ymesh]))
-ax.plot_surface(xmesh, ymesh, fmesh)
-!ec
-And then as countor plot
-!bc pycod
-pt.axis("equal")
-pt.contour(xmesh, ymesh, fmesh)
-guesses = [np.array([2, 2./5])]
-!ec
-Find guesses
-!bc pycod
-x = guesses[-1]
-s = -df(x)
-!ec
-Run it!
-!bc pycod
-def f1d(alpha):
- return f(x + alpha*s)
-
-alpha_opt = sopt.golden(f1d)
-next_guess = x + alpha_opt * s
-guesses.append(next_guess)
-print(next_guess)
-!ec
-What happened?
-!bc pycod
-pt.axis("equal")
-pt.contour(xmesh, ymesh, fmesh, 50)
-it_array = np.array(guesses)
-pt.plot(it_array.T[0], it_array.T[1], "x-")
-!ec
-
-!split
-===== Conjugate gradient method =====
-!bblock
-In the CG method we define so-called conjugate directions and two vectors
-$\bm{s}$ and $\bm{t}$
-are said to be
-conjugate if
-!bt
-\begin{equation*}
-\bm{s}^T\bm{A}\bm{t}= 0.
-\end{equation*}
-!et
-The philosophy of the CG method is to perform searches in various conjugate directions
-of our vectors $\bm{x}_i$ obeying the above criterion, namely
-!bt
-\begin{equation*}
-\bm{x}_i^T\bm{A}\bm{x}_j= 0.
-\end{equation*}
-!et
-Two vectors are conjugate if they are orthogonal with respect to
-this inner product. Being conjugate is a symmetric relation: if $\bm{s}$ is conjugate to $\bm{t}$, then $\bm{t}$ is conjugate to $\bm{s}$.
-!eblock
-
-!split
-===== Conjugate gradient method =====
-!bblock
-An example is given by the eigenvectors of the matrix
-!bt
-\begin{equation*}
-\bm{v}_i^T\bm{A}\bm{v}_j= \lambda\bm{v}_i^T\bm{v}_j,
-\end{equation*}
-!et
-which is zero unless $i=j$.
-!eblock
-
-
-!split
-===== Conjugate gradient method =====
-!bblock
-Assume now that we have a symmetric positive-definite matrix $\bm{A}$ of size
-$n\times n$. At each iteration $i+1$ we obtain the conjugate direction of a vector
-!bt
-\begin{equation*}
-\bm{x}_{i+1}=\bm{x}_{i}+\alpha_i\bm{p}_{i}.
-\end{equation*}
-!et
-We assume that $\bm{p}_{i}$ is a sequence of $n$ mutually conjugate directions.
-Then the $\bm{p}_{i}$ form a basis of $R^n$ and we can expand the solution
-$ \bm{A}\bm{x} = \bm{b}$ in this basis, namely
-
-!bt
-\begin{equation*}
- \bm{x} = \sum^{n}_{i=1} \alpha_i \bm{p}_i.
-\end{equation*}
-!et
-!eblock
-
-!split
-===== Conjugate gradient method =====
-!bblock
-The coefficients are given by
-!bt
-\begin{equation*}
- \mathbf{A}\mathbf{x} = \sum^{n}_{i=1} \alpha_i \mathbf{A} \mathbf{p}_i = \mathbf{b}.
-\end{equation*}
-!et
-Multiplying with $\bm{p}_k^T$ from the left gives
-
-!bt
-\begin{equation*}
- \bm{p}_k^T \bm{A}\bm{x} = \sum^{n}_{i=1} \alpha_i\bm{p}_k^T \bm{A}\bm{p}_i= \bm{p}_k^T \bm{b},
-\end{equation*}
-!et
-and we can define the coefficients $\alpha_k$ as
-
-!bt
-\begin{equation*}
- \alpha_k = \frac{\bm{p}_k^T \bm{b}}{\bm{p}_k^T \bm{A} \bm{p}_k}
-\end{equation*}
-!et
-!eblock
-
-!split
-===== Conjugate gradient method and iterations =====
-!bblock
-
-If we choose the conjugate vectors $\bm{p}_k$ carefully,
-then we may not need all of them to obtain a good approximation to the solution
-$\bm{x}$.
-We want to regard the conjugate gradient method as an iterative method.
-This will us to solve systems where $n$ is so large that the direct
-method would take too much time.
-
-We denote the initial guess for $\bm{x}$ as $\bm{x}_0$.
-We can assume without loss of generality that
-!bt
-\begin{equation*}
-\bm{x}_0=0,
-\end{equation*}
-!et
-or consider the system
-!bt
-\begin{equation*}
-\bm{A}\bm{z} = \bm{b}-\bm{A}\bm{x}_0,
-\end{equation*}
-!et
-instead.
-!eblock
-
-
-!split
-===== Conjugate gradient method =====
-!bblock
-One can show that the solution $\bm{x}$ is also the unique minimizer of the quadratic form
-!bt
-\begin{equation*}
- f(\bm{x}) = \frac{1}{2}\bm{x}^T\bm{A}\bm{x} - \bm{x}^T \bm{x} , \quad \bm{x}\in\mathbf{R}^n.
-\end{equation*}
-!et
-This suggests taking the first basis vector $\bm{p}_1$
-to be the gradient of $f$ at $\bm{x}=\bm{x}_0$,
-which equals
-!bt
-\begin{equation*}
-\bm{A}\bm{x}_0-\bm{b},
-\end{equation*}
-!et
-and
-$\bm{x}_0=0$ it is equal $-\bm{b}$.
-The other vectors in the basis will be conjugate to the gradient,
-hence the name conjugate gradient method.
-!eblock
-
-
-!split
-===== Conjugate gradient method =====
-!bblock
-Let $\bm{r}_k$ be the residual at the $k$-th step:
-!bt
-\begin{equation*}
-\bm{r}_k=\bm{b}-\bm{A}\bm{x}_k.
-\end{equation*}
-!et
-Note that $\bm{r}_k$ is the negative gradient of $f$ at
-$\bm{x}=\bm{x}_k$,
-so the gradient descent method would be to move in the direction $\bm{r}_k$.
-Here, we insist that the directions $\bm{p}_k$ are conjugate to each other,
-so we take the direction closest to the gradient $\bm{r}_k$
-under the conjugacy constraint.
-This gives the following expression
-!bt
-\begin{equation*}
-\bm{p}_{k+1}=\bm{r}_k-\frac{\bm{p}_k^T \bm{A}\bm{r}_k}{\bm{p}_k^T\bm{A}\bm{p}_k} \bm{p}_k.
-\end{equation*}
-!et
-!eblock
-
-!split
-===== Conjugate gradient method =====
-!bblock
-We can also compute the residual iteratively as
-!bt
-\begin{equation*}
-\bm{r}_{k+1}=\bm{b}-\bm{A}\bm{x}_{k+1},
- \end{equation*}
-!et
-which equals
-!bt
-\begin{equation*}
-\bm{b}-\bm{A}(\bm{x}_k+\alpha_k\bm{p}_k),
- \end{equation*}
-!et
-or
-!bt
-\begin{equation*}
-(\bm{b}-\bm{A}\bm{x}_k)-\alpha_k\bm{A}\bm{p}_k,
- \end{equation*}
-!et
-which gives
-
-!bt
-\begin{equation*}
-\bm{r}_{k+1}=\bm{r}_k-\bm{A}\bm{p}_{k},
- \end{equation*}
-!et
-!eblock
-
-
-
-
-
-!split
-===== Revisiting some of our first Linear Regression Encounters =====
-
-We will use linear regression as a case study for the gradient descent
-methods. Linear regression is a great test case for the gradient
-descent methods discussed in the lectures since it has several
-desirable properties such as:
-
-o An analytical solution (recall homework set 1).
-o The gradient can be computed analytically.
-o The cost function is convex which guarantees that gradient descent converges for small enough learning rates
-
-We revisit an example similar to what we had in the first homework set. We had a function of the type
-
-!bc pycod
-x = 2*np.random.rand(m,1)
-y = 4+3*x+np.random.randn(m,1)
-!ec
-with $x_i \in [0,1] $ is chosen randomly using a uniform distribution. Additionally we have a stochastic noise chosen according to a normal distribution $\cal {N}(0,1)$.
-The linear regression model is given by
-!bt
-\[
-h_\beta(x) = \bm{y} = \beta_0 + \beta_1 x,
-\]
-!et
-such that
-!bt
-\[
-\bm{y}_i = \beta_0 + \beta_1 x_i.
-\]
-!et
-
-!split
-===== Gradient descent example =====
-
-Let $\mathbf{y} = (y_1,\cdots,y_n)^T$, $\mathbf{\bm{y}} = (\bm{y}_1,\cdots,\bm{y}_n)^T$ and $\beta = (\beta_0, \beta_1)^T$
-
-It is convenient to write $\mathbf{\bm{y}} = X\beta$ where $X \in \mathbb{R}^{100 \times 2} $ is the design matrix given by (we keep the intercept here)
-!bt
-\[
-X \equiv \begin{bmatrix}
-1 & x_1 \\
-\vdots & \vdots \\
-1 & x_{100} & \\
-\end{bmatrix}.
-\]
-!et
-The cost/loss/risk function is given by (
-!bt
-\[
-C(\beta) = \frac{1}{n}||X\beta-\mathbf{y}||_{2}^{2} = \frac{1}{n}\sum_{i=1}^{100}\left[ (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2\right]
-\]
-!et
-and we want to find $\beta$ such that $C(\beta)$ is minimized.
-
-!split
-===== The derivative of the cost/loss function =====
-
-Computing $\partial C(\beta) / \partial \beta_0$ and $\partial C(\beta) / \partial \beta_1$ we can show that the gradient can be written as
-!bt
-\[
-\nabla_{\beta} C(\beta) = \frac{2}{n}\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
-\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
-\end{bmatrix} = \frac{2}{n}X^T(X\beta - \mathbf{y}),
-\]
-!et
-where $X$ is the design matrix defined above.
-
-!split
-===== The Hessian matrix =====
-The Hessian matrix of $C(\beta)$ is given by
-!bt
-\[
-\bm{H} \equiv \begin{bmatrix}
-\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
-\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
-\end{bmatrix} = \frac{2}{n}X^T X.
-\]
-!et
-This result implies that $C(\beta)$ is a convex function since the matrix $X^T X$ always is positive semi-definite.
-
-
-
-
-!split
-===== Simple program =====
-
-We can now write a program that minimizes $C(\beta)$ using the gradient descent method with a constant learning rate $\gamma$ according to
-!bt
-\[
-\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
-\]
-!et
-
-We can use the expression we computed for the gradient and let use a
-$\beta_0$ be chosen randomly and let $\gamma = 0.001$. Stop iterating
-when $||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8}$. _Note that the code below does not include the latter stop criterion_.
-
-And finally we can compare our solution for $\beta$ with the analytic result given by
-$\beta= (X^TX)^{-1} X^T \mathbf{y}$.
-
-!split
-===== Gradient Descent Example =====
-
-Here our simple example
-!bc pycod
-
-# Importing various packages
-from random import random, seed
-import numpy as np
-import matplotlib.pyplot as plt
-from mpl_toolkits.mplot3d import Axes3D
-from matplotlib import cm
-from matplotlib.ticker import LinearLocator, FormatStrFormatter
-import sys
-
-# the number of datapoints
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-# Hessian matrix
-H = (2.0/n)* X.T @ X
-# Get the eigenvalues
-EigValues, EigVectors = np.linalg.eig(H)
-print(EigValues)
-
-beta_linreg = np.linalg.inv(X.T @ X) @ X.T @ y
-print(beta_linreg)
-beta = np.random.randn(2,1)
-
-eta = 1.0/np.max(EigValues)
-Niterations = 1000
-
-for iter in range(Niterations):
- gradient = (2.0/n)*X.T @ (X @ beta-y)
- beta -= eta*gradient
-
-print(beta)
-xnew = np.array([[0],[2]])
-xbnew = np.c_[np.ones((2,1)), xnew]
-ypredict = xbnew.dot(beta)
-ypredict2 = xbnew.dot(beta_linreg)
-plt.plot(xnew, ypredict, "r-")
-plt.plot(xnew, ypredict2, "b-")
-plt.plot(x, y ,'ro')
-plt.axis([0,2.0,0, 15.0])
-plt.xlabel(r'$x$')
-plt.ylabel(r'$y$')
-plt.title(r'Gradient descent example')
-plt.show()
-
-!ec
-
-!split
-===== And a corresponding example using _scikit-learn_ =====
-
-!bc pycod
-# Importing various packages
-from random import random, seed
-import numpy as np
-import matplotlib.pyplot as plt
-from sklearn.linear_model import SGDRegressor
-
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-beta_linreg = np.linalg.inv(X.T @ X) @ (X.T @ y)
-print(beta_linreg)
-sgdreg = SGDRegressor(max_iter = 50, penalty=None, eta0=0.1)
-sgdreg.fit(x,y.ravel())
-print(sgdreg.intercept_, sgdreg.coef_)
-
-!ec
-
-
-
-!split
-===== Gradient descent and Ridge =====
-
-We have also discussed Ridge regression where the loss function contains a regularized term given by the $L_2$ norm of $\beta$,
-!bt
-\[
-C_{\text{ridge}}(\beta) = \frac{1}{n}||X\beta -\mathbf{y}||^2 + \lambda ||\beta||^2, \ \lambda \geq 0.
-\]
-!et
-
-In order to minimize $C_{\text{ridge}}(\beta)$ using GD we only have adjust the gradient as follows
-!bt
-\[
-\nabla_\beta C_{\text{ridge}}(\beta) = \frac{2}{n}\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
-\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
-\end{bmatrix} + 2\lambda\begin{bmatrix} \beta_0 \\ \beta_1\end{bmatrix} = 2 (X^T(X\beta - \mathbf{y})+\lambda \beta).
-\]
-!et
-
-We can easily extend our program to minimize $C_{\text{ridge}}(\beta)$ using gradient descent and compare with the analytical solution given by
-!bt
-\[
-\beta_{\text{ridge}} = \left(X^T X + \lambda I_{2 \times 2} \right)^{-1} X^T \mathbf{y}.
-\]
-!et
-
-
-!split
-===== Program example for gradient descent with Ridge Regression =====
-!bc pycod
-from random import random, seed
-import numpy as np
-import matplotlib.pyplot as plt
-from mpl_toolkits.mplot3d import Axes3D
-from matplotlib import cm
-from matplotlib.ticker import LinearLocator, FormatStrFormatter
-import sys
-
-# the number of datapoints
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-XT_X = X.T @ X
-
-#Ridge parameter lambda
-lmbda = 0.001
-Id = lmbda* np.eye(XT_X.shape[0])
-
-beta_linreg = np.linalg.inv(XT_X+Id) @ X.T @ y
-print(beta_linreg)
-# Start plain gradient descent
-beta = np.random.randn(2,1)
-
-eta = 0.1
-Niterations = 100
-
-for iter in range(Niterations):
- gradients = 2.0/n*X.T @ (X @ (beta)-y)+2*lmbda*beta
- beta -= eta*gradients
-
-print(beta)
-ypredict = X @ beta
-ypredict2 = X @ beta_linreg
-plt.plot(x, ypredict, "r-")
-plt.plot(x, ypredict2, "b-")
-plt.plot(x, y ,'ro')
-plt.axis([0,2.0,0, 15.0])
-plt.xlabel(r'$x$')
-plt.ylabel(r'$y$')
-plt.title(r'Gradient descent example for Ridge')
-plt.show()
-
-
-!ec
-
-!split
-===== Using gradient descent methods, limitations =====
-
-* _Gradient descent (GD) finds local minima of our function_. Since the GD algorithm is deterministic, if it converges, it will converge to a local minimum of our cost/loss/risk function. Because in ML we are often dealing with extremely rugged landscapes with many local minima, this can lead to poor performance.
-
-* _GD is sensitive to initial conditions_. One consequence of the local nature of GD is that initial conditions matter. Depending on where one starts, one will end up at a different local minima. Therefore, it is very important to think about how one initializes the training process. This is true for GD as well as more complicated variants of GD.
-
-* _Gradients are computationally expensive to calculate for large datasets_. In many cases in statistics and ML, the cost/loss/risk function is a sum of terms, with one term for each data point. For example, in linear regression, $E \propto \sum_{i=1}^n (y_i - \mathbf{w}^T\cdot\mathbf{x}_i)^2$; for logistic regression, the square error is replaced by the cross entropy. To calculate the gradient we have to sum over *all* $n$ data points. Doing this at every GD step becomes extremely computationally expensive. An ingenious solution to this, is to calculate the gradients using small subsets of the data called ``mini batches''. This has the added benefit of introducing stochasticity into our algorithm.
-
-* _GD is very sensitive to choices of learning rates_. GD is extremely sensitive to the choice of learning rates. If the learning rate is very small, the training process take an extremely long time. For larger learning rates, GD can diverge and give poor results. Furthermore, depending on what the local landscape looks like, we have to modify the learning rates to ensure convergence. Ideally, we would *adaptively* choose the learning rates to match the landscape.
-
-* _GD treats all directions in parameter space uniformly._ Another major drawback of GD is that unlike Newton's method, the learning rate for GD is the same in all directions in parameter space. For this reason, the maximum learning rate is set by the behavior of the steepest direction and this can significantly slow down training. Ideally, we would like to take large steps in flat directions and small steps in steep directions. Since we are exploring rugged landscapes where curvatures change, this requires us to keep track of not only the gradient but second derivatives. The ideal scenario would be to calculate the Hessian but this proves to be too computationally expensive.
-
-* GD can take exponential time to escape saddle points, even with random initialization. As we mentioned, GD is extremely sensitive to initial condition since it determines the particular local minimum GD would eventually reach. However, even with a good initialization scheme, through the introduction of randomness, GD can still take exponential time to escape saddle points. This leads us to our next topic, Stochastic Gradient Methods.
-