Week 39: Optimization and Gradient Methods
Contents
Plan for week 39, September 23-27, 2024
Lecture Monday September 23
Lab sessions week 39
Lecture Monday September 23, Optimization, the central part of any Machine Learning algortithm
Revisiting our Logistic Regression case
The equations to solve
Solving using Newton-Raphson's method
Brief reminder on Newton-Raphson's method
The equations
Simple geometric interpretation
Extending to more than one variable
Steepest descent
More on Steepest descent
The ideal
The sensitiveness of the gradient descent
Convex functions
Convex function
Conditions on convex functions
More on convex functions
Some simple problems
Standard steepest descent
Gradient method
Steepest descent method
Steepest descent method
Final expressions
Steepest descent example
Conjugate gradient method
Conjugate gradient method
Conjugate gradient method
Conjugate gradient method
Conjugate gradient method and iterations
Conjugate gradient method
Conjugate gradient method
Conjugate gradient method
Revisiting our first homework
Gradient descent example
The derivative of the cost/loss function
The Hessian matrix
Simple program
Gradient Descent Example
And a corresponding example using
scikit-learn
Gradient descent and Ridge
The Hessian matrix for Ridge Regression
Program example for gradient descent with Ridge Regression
Using gradient descent methods, limitations
Improving gradient descent with momentum
Same code but now with momentum gradient descent
Overview video on Stochastic Gradient Descent
Batches and mini-batches
Stochastic Gradient Descent (SGD)
Stochastic Gradient Descent
Computation of gradients
SGD example
The gradient step
Simple example code
When do we stop?
Slightly different approach
Time decay rate
Code with a Number of Minibatches which varies
Replace or not
Momentum based GD
More on momentum based approaches
Momentum parameter
Second moment of the gradient
RMS prop
"ADAM optimizer":"https://arxiv.org/abs/1412.6980"
Algorithms and codes for Adagrad, RMSprop and Adam
Practical tips
Automatic differentiation
Using autograd
Autograd with more complicated functions
More complicated functions using the elements of their arguments directly
Functions using mathematical functions from Numpy
More autograd
And with loops
Using recursion
Unsupported functions
The syntax a.dot(b) when finding the dot product
Using Autograd with OLS
Same code but now with momentum gradient descent
But none of these can compete with Newton's method
Including Stochastic Gradient Descent with Autograd
Same code but now with momentum gradient descent
Similar (second order function now) problem but now with AdaGrad
RMSprop for adaptive learning rate with Stochastic Gradient Descent
And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
And Logistic Regression
Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
Overview video on Stochastic Gradient Descent
What is Stochastic Gradient Descent
«
1
...
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
...
89
»